10 KiB
JustOneCacophony — JOC
JustOneCacophony is an experimental/test implementation of E-AC-3 JOC for studying JOC parsing, reconstruction, rendering, and the associated mathematics.
The project can extract and parse EMDF, ID14 JOC parameters, and ID11 OAMD metadata from common E-AC-3 JOC streams. It combines those data with the core 5.1 PCM decoded by FFmpeg, reconstructs LFE plus 15 object channels, and writes ADM BWF, a WAV file for a selected speaker layout, or direct DLL-free Rosella binaural stereo.
This is research code, not a complete, standards-compliant, or production-grade JOC decoder. It covers only the stream forms currently implemented. Unknown variants fail explicitly—because when the math goes wrong, all that may remain is the cacophony.
Current features
- Scan common contiguous EMDF containers in E-AC-3 sync frames.
- Parse ID14 dense JOC parameters, Huffman data, differential matrices, and
joc_clipgain. - Parse ID11 OAMD position updates and build object trajectories.
- Reconstruct LFE plus 15 object channels through analysis QMF, parameter interpolation, the object matrix, and inverse QMF.
- Write a 25-channel ADM BWF: a 10-channel 7.1.2 bed (silent except for LFE) plus 15 objects.
- Render directly to
2.0,3.1,5.1,7.1,5.1.2,5.1.4,7.1.2,7.1.4,9.1.4, or9.1.6. - Run DLL-free Rosella binaural rendering directly from
pcm16 + ID11/OAMD, without a temporary ADM BWF. - Keep the binaural DSP in float64/complex128, including 961-sample latency compensation, cross-frame state, and the room tail.
- Use a shared float32/PCM24 WAV writer and explicit PCM24 clipping policy for direct outputs.
- Use the NumPy backend or an optional C++20 core through
ctypes;autofalls back to Python when the native library is unavailable. - Read or write metadata sidecars and produce metadata, timing, and output reports.
Processing flow
M4A / E-AC-3
├─ FFmpeg extracts E-AC-3 and decodes the core 5.1 PCM
├─ EMDF → ID14 JOC parameters → object matrix
├─ core PCM → analysis QMF → parameter interpolation → inverse QMF
├─ ID11 OAMD → object positions and timing
└─ LFE + 15 objects
├─ 25ch ADM BWF
├─ speaker WAV for the selected layout
└─ direct ID11 timeline → Rosella → binaural WAV
The Python and C++ backends follow the same mathematics. The native core handles object reconstruction, speaker rendering, and binaural QMF/hybrid/room/synthesis; bitstream parsing, the OAMD timeline, model parsing, and CLI behavior remain in Python.
Requirements
- Python 3.10+
- NumPy 1.24+
- A standalone FFmpeg executable;
ffmpeg-pythonis not required. FFmpeg is discovered throughPATHby default or selected with--ffmpeg - Optional: CMake and a C++20 toolchain to build the native core
Install the Python dependency in a project-specific environment:
python -m pip install -r requirements.txt
If FFmpeg is not on PATH:
python main.py input.m4a --ffmpeg C:\path\to\ffmpeg.exe
Usage
Write a 25-channel ADM BWF by default:
python main.py input.m4a
Select a backend or output path:
python main.py input.eac3 -o output.adm.wav --backend python
python main.py input.m4a --backend native --native-threads 2
python main.py input.m4a --native-library lib/eac3joc_core.dll
Write a speaker-layout WAV directly:
python main.py input.m4a --speaker-layout 2.0 --speaker-format float32
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24
python main.py input.m4a --speaker-layout 7.1.2 --speaker-output output.7.1.2.wav
Write Rosella binaural stereo directly (ordinary objects are Near/Mid/Far only; Mid is the default):
python main.py input.m4a --binaural
python main.py input.m4a --binaural --binaural-mode near
python main.py input.m4a --binaural --binaural-mode far --binaural-format int24
python main.py input.m4a --binaural --binaural-output output.binaural.wav `
--personalized-headphone C:\HRTF\my.personalized_headphone
The default model path is HRTF/binaural.personalized_headphone. An example HRTF file is available from:
See Binaural Rendering Mathematics for the formulas, state, and timing model.
Speaker and binaural output share peak analysis, the WAV writer, and clipping policy. When PCM24 may clip in a non-interactive environment, select a policy explicitly:
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action abort
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action float32
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action continue
python main.py input.m4a --binaural --binaural-format int24 --clip-action abort
Metadata and diagnostics:
python main.py input.m4a --print-metadata summary
python main.py input.m4a --metadata-only --print-metadata frames
python main.py input.m4a --metadata-cache metadata_cache
python main.py input.m4a --metadata-dir metadata_cache
Experimental binaural mode settings for JOC objects
The binaural mode written here is a user-selected, experimental rendering hint for downstream ADM renderers. It is not original binaural metadata extracted or recovered from the input E-AC-3 JOC bitstream, nor does it represent the original mix's per-object binaural settings. The selected mode is applied uniformly to all 15 JOC objects; the default unspecified is this tool's default, not a mode detected in the source file.
Use --joc-binaural-mode off|near|far|mid|unspecified to select a mode, encoded as 0|1|2|3|4 respectively. The default is unspecified:
python main.py input.m4a --joc-binaural-mode mid
This option only sets the low 3 binaural-render-mode bits of the last 15 JOC object entries in ADM BWF DBMD segment 10, leaving the first 10 bed entries unchanged. It does not change PCM, object trajectories, direct speaker rendering, or direct Rosella binaural rendering. The adjacent .report.json records the mode name and value in joc_binaural_mode and joc_binaural_mode_value; both are null for direct speaker or direct binaural output, where the option does not apply.
Binaural calculation
See Binaural Rendering Mathematics for QMF, hybrid processing, direction fields, distance, ITD, room processing, 512-sample parameter updates, and 961-sample latency compensation.
For all options:
python main.py --help
Without -o, output still goes to output/ at the repository root. The directory move intentionally preserves this behavior.
Native core
The repository does not include native binaries by default. Download a prebuilt runtime for the current platform from a project Release, or build one locally, then place the runtime library under lib/ at the repository root; create the directory if it is absent. To build it yourself, run CMake from the repository root:
cmake -S native -B build/cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX="$PWD/lib"
cmake --build build/cmake --config Release
cmake --install build/cmake --config Release
The runtime lookup order is:
--native-library;EAC3JOC_NATIVE_LIBRARY;- the standard platform library name under
lib/.
See the native-core notes for ABI, state, and precision details.
Repository layout
JustOneCacophony/
├─ main.py command-line entry point
├─ src/ Python implementation modules
├─ native/ C/C++ acceleration core, C ABI, and required table data
├─ data/ Python runtime table data
├─ lib/ native runtime drop-in directory (create as needed)
├─ docs/ math and native-core notes in both languages
├─ requirements.txt Python dependency
├─ README.md Chinese documentation
└─ README.en.md English documentation
Mathematical implementation
The main documented stages are:
- dense JOC differential reconstruction and dequantization;
- parameter-band mapping to 64 QMF subbands;
- cross-frame parameter interpolation;
- analysis/inverse QMF, surround delay, and FIR state;
- the 1217-sample LFE delay;
- OAMD Q15 coordinate conversion;
- equal-power panning over target-layout regions;
- layout-dependent position compensation and sample-wise gain ramps;
- float32 and PCM24 output quantization;
- Rosella 64-band QMF, 77-band hybrid processing,
77×36direction fields, distance/ITD, room FIR, and special LFE.
See the mathematical notes for the equations used by the decoding and rendering process.
Known limitations
- Only the common contiguous EMDF transport is covered. Fragmented transport across multiple audio-block skip fields is not covered.
- Dense JOC is the main path. The Sparse JOC branch should not be treated as supported.
- The speaker and Rosella binaural paths currently cover ordinary point objects; extent, spread, diffuse, divergence, channel lock, and similar controls are outside the supported scope.
- OAMD trim elements are boundary-checked and skipped; warp, balance, and trim parameters are not applied to raw object trajectories or speaker rendering.
- Multi-data-point streams, uncommon band configurations, and unusual OAMD scheduling have less coverage than common 12-band, single-data-point material.
- A speaker limiter is outside the current primary formula.
- Rosella requires a user-supplied compatible
.personalized_headphone; arbitrary SOFA data cannot become a valid Rosella rp through JSON rearrangement alone. - ADM output, native binaries, speaker layouts, and binaural models still need broader interoperability checks across platforms, players, and real material.