# JustOneCacophony — JOC [中文版](README.md) > JustOneCacophony is an experimental/test implementation of E-AC-3 JOC for studying JOC parsing, reconstruction, rendering, and the associated mathematics. The project can extract and parse EMDF, ID14 JOC parameters, and ID11 OAMD metadata from common E-AC-3 JOC streams. It combines those data with the core 5.1 PCM decoded by FFmpeg, reconstructs LFE plus 15 object channels, and writes ADM BWF, a WAV file for a selected speaker layout, or direct DLL-free Rosella binaural stereo. This is research code, not a complete, standards-compliant, or production-grade JOC decoder. It covers only the stream forms currently implemented. Unknown variants fail explicitly—because when the math goes wrong, all that may remain is the cacophony. ## Current features - Scan common contiguous EMDF containers in E-AC-3 sync frames. - Parse ID14 dense JOC parameters, Huffman data, differential matrices, and `joc_clipgain`. - Parse ID11 OAMD position updates and build object trajectories. - Reconstruct LFE plus 15 object channels through analysis QMF, parameter interpolation, the object matrix, and inverse QMF. - Write a 25-channel ADM BWF: a 10-channel 7.1.2 bed (silent except for LFE) plus 15 objects. - Render directly to `2.0`, `3.1`, `5.1`, `7.1`, `5.1.2`, `5.1.4`, `7.1.2`, `7.1.4`, `9.1.4`, or `9.1.6`. - Run DLL-free Rosella binaural rendering directly from `pcm16 + ID11/OAMD`, without a temporary ADM BWF. - Keep the binaural DSP in float64/complex128, including 961-sample latency compensation, cross-frame state, and the room tail. - Use a shared float32/PCM24 WAV writer and explicit PCM24 clipping policy for direct outputs. - Use the NumPy backend or an optional C++20 core through `ctypes`; `auto` falls back to Python when the native library is unavailable. - Read or write metadata sidecars and produce metadata, timing, and output reports. ## Processing flow ```text M4A / E-AC-3 ├─ FFmpeg extracts E-AC-3 and decodes the core 5.1 PCM ├─ EMDF → ID14 JOC parameters → object matrix ├─ core PCM → analysis QMF → parameter interpolation → inverse QMF ├─ ID11 OAMD → object positions and timing └─ LFE + 15 objects ├─ 25ch ADM BWF ├─ speaker WAV for the selected layout └─ direct ID11 timeline → Rosella → binaural WAV ``` The Python and C++ backends follow the same mathematics. The native core handles object reconstruction, speaker rendering, and binaural QMF/hybrid/room/synthesis; bitstream parsing, the OAMD timeline, model parsing, and CLI behavior remain in Python. ## Requirements - Python 3.10+ - NumPy 1.24+ - A standalone FFmpeg executable; `ffmpeg-python` is not required. FFmpeg is discovered through `PATH` by default or selected with `--ffmpeg` - Optional: CMake and a C++20 toolchain to build the native core Install the Python dependency in a project-specific environment: ```powershell python -m pip install -r requirements.txt ``` If FFmpeg is not on `PATH`: ```powershell python main.py input.m4a --ffmpeg C:\path\to\ffmpeg.exe ``` ## Usage Write a 25-channel ADM BWF by default: ```powershell python main.py input.m4a ``` Select a backend or output path: ```powershell python main.py input.eac3 -o output.adm.wav --backend python python main.py input.m4a --backend native --native-threads 2 python main.py input.m4a --native-library lib/eac3joc_core.dll ``` Write a speaker-layout WAV directly: ```powershell python main.py input.m4a --speaker-layout 2.0 --speaker-format float32 python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 python main.py input.m4a --speaker-layout 7.1.2 --speaker-output output.7.1.2.wav ``` Write Rosella binaural stereo directly (ordinary objects are Near/Mid/Far only; Mid is the default): ```powershell python main.py input.m4a --binaural python main.py input.m4a --binaural --binaural-mode near python main.py input.m4a --binaural --binaural-mode far --binaural-format int24 python main.py input.m4a --binaural --binaural-output output.binaural.wav ` --personalized-headphone C:\HRTF\my.personalized_headphone ``` The default model path is `HRTF/binaural.personalized_headphone`. An example HRTF file is available from: https://professionalsupport.dolby.com/s/question/0D54u0000AAT85HCQT/the-state-of-personalized-binaural-rendering?language=en_US See [Binaural Rendering Mathematics](docs/binaural.en.md) for the formulas, state, and timing model. Speaker and binaural output share peak analysis, the WAV writer, and clipping policy. When PCM24 may clip in a non-interactive environment, select a policy explicitly: ```powershell python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action abort python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action float32 python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action continue python main.py input.m4a --binaural --binaural-format int24 --clip-action abort ``` Metadata and diagnostics: ```powershell python main.py input.m4a --print-metadata summary python main.py input.m4a --metadata-only --print-metadata frames python main.py input.m4a --metadata-cache metadata_cache python main.py input.m4a --metadata-dir metadata_cache ``` ### Experimental binaural mode settings for JOC objects The binaural mode written here is a user-selected, experimental rendering hint for downstream ADM renderers. It is **not original binaural metadata extracted or recovered from the input E-AC-3 JOC bitstream**, nor does it represent the original mix's per-object binaural settings. The selected mode is applied uniformly to all 15 JOC objects; the default `unspecified` is this tool's default, not a mode detected in the source file. Use `--joc-binaural-mode off|near|far|mid|unspecified` to select a mode, encoded as `0|1|2|3|4` respectively. The default is `unspecified`: ```powershell python main.py input.m4a --joc-binaural-mode mid ``` This option only sets the low 3 binaural-render-mode bits of the last 15 JOC object entries in ADM BWF DBMD segment 10, leaving the first 10 bed entries unchanged. It does not change PCM, object trajectories, direct speaker rendering, or direct Rosella binaural rendering. The adjacent `.report.json` records the mode name and value in `joc_binaural_mode` and `joc_binaural_mode_value`; both are `null` for direct speaker or direct binaural output, where the option does not apply. ### Binaural calculation See [Binaural Rendering Mathematics](docs/binaural.en.md) for QMF, hybrid processing, direction fields, distance, ITD, room processing, 512-sample parameter updates, and 961-sample latency compensation. For all options: ```powershell python main.py --help ``` Without `-o`, output still goes to `output/` at the repository root. The directory move intentionally preserves this behavior. ## Native core The repository does not include native binaries by default. Download a prebuilt runtime for the current platform from a project Release, or build one locally, then place the runtime library under `lib/` at the repository root; create the directory if it is absent. To build it yourself, run CMake from the repository root: ```powershell cmake -S native -B build/cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX="$PWD/lib" cmake --build build/cmake --config Release cmake --install build/cmake --config Release ``` The runtime lookup order is: 1. `--native-library`; 2. `EAC3JOC_NATIVE_LIBRARY`; 3. the standard platform library name under `lib/`. See the [native-core notes](docs/native.en.md) for ABI, state, and precision details. ## Repository layout ```text JustOneCacophony/ ├─ main.py command-line entry point ├─ src/ Python implementation modules ├─ native/ C/C++ acceleration core, C ABI, and required table data ├─ data/ Python runtime table data ├─ lib/ native runtime drop-in directory (create as needed) ├─ docs/ math and native-core notes in both languages ├─ requirements.txt Python dependency ├─ README.md Chinese documentation └─ README.en.md English documentation ``` ## Mathematical implementation The main documented stages are: - dense JOC differential reconstruction and dequantization; - parameter-band mapping to 64 QMF subbands; - cross-frame parameter interpolation; - analysis/inverse QMF, surround delay, and FIR state; - the 1217-sample LFE delay; - OAMD Q15 coordinate conversion; - equal-power panning over target-layout regions; - layout-dependent position compensation and sample-wise gain ramps; - float32 and PCM24 output quantization; - Rosella 64-band QMF, 77-band hybrid processing, `77×36` direction fields, distance/ITD, room FIR, and special LFE. See the [mathematical notes](docs/math.en.md) for the equations used by the decoding and rendering process. ## Known limitations - Only the common contiguous EMDF transport is covered. Fragmented transport across multiple audio-block skip fields is not covered. - Dense JOC is the main path. The Sparse JOC branch should not be treated as supported. - The speaker and Rosella binaural paths currently cover ordinary point objects; extent, spread, diffuse, divergence, channel lock, and similar controls are outside the supported scope. - OAMD trim elements are boundary-checked and skipped; warp, balance, and trim parameters are not applied to raw object trajectories or speaker rendering. - Multi-data-point streams, uncommon band configurations, and unusual OAMD scheduling have less coverage than common 12-band, single-data-point material. - A speaker limiter is outside the current primary formula. - Rosella requires a user-supplied compatible `.personalized_headphone`; arbitrary SOFA data cannot become a valid Rosella rp through JSON rearrangement alone. - ADM output, native binaries, speaker layouts, and binaural models still need broader interoperability checks across platforms, players, and real material. ## Documentation - [Mathematical notes](docs/math.en.md) · [中文](docs/math.md) - [Native-core notes](docs/native.en.md) · [中文](docs/native.md) - [DLL-free Rosella binaural](docs/binaural.en.md) · [中文](docs/binaural.md)