Archive DLL-free personalized HRTF binaural renderer

This commit is contained in:
2026-09-05 02:43:54 +08:00
parent 329445ed25
commit 22a16ab60d
25 changed files with 4285 additions and 109 deletions
+36 -19
View File
@@ -4,9 +4,9 @@
> JustOneCacophony is an experimental/test implementation of E-AC-3 JOC for studying JOC parsing, reconstruction, rendering, and the associated mathematics.
The project can extract and parse EMDF, ID14 JOC parameters, and ID11 OAMD metadata from common E-AC-3 JOC streams. It combines those data with the core 5.1 PCM decoded by FFmpeg, reconstructs LFE plus 15 object channels, and writes either ADM BWF or a WAV file for a selected speaker layout.
The project can extract and parse EMDF, ID14 JOC parameters, and ID11 OAMD metadata from common E-AC-3 JOC streams. It combines those data with the core 5.1 PCM decoded by FFmpeg, reconstructs LFE plus 15 object channels, and writes ADM BWF, a WAV file for a selected speaker layout, or direct DLL-free Rosella binaural stereo.
This is research code, not a complete, standards-compliant, or production-grade Dolby JOC decoder. It covers only the stream forms currently implemented. Unknown variants fail explicitly—because when the math goes wrong, all that may remain is the cacophony.
This is research code, not a complete, standards-compliant, or production-grade JOC decoder. It covers only the stream forms currently implemented. Unknown variants fail explicitly—because when the math goes wrong, all that may remain is the cacophony.
## Current features
@@ -16,7 +16,9 @@ This is research code, not a complete, standards-compliant, or production-grade
- Reconstruct LFE plus 15 object channels through analysis QMF, parameter interpolation, the object matrix, and inverse QMF.
- Write a 25-channel ADM BWF: a 10-channel 7.1.2 bed (silent except for LFE) plus 15 objects.
- Render directly to `2.0`, `3.1`, `5.1`, `7.1`, `5.1.2`, `5.1.4`, `7.1.2`, `7.1.4`, `9.1.4`, or `9.1.6`.
- Write float32 or PCM24 WAV and require an explicit policy when PCM24 would clip.
- Run DLL-free Rosella binaural rendering directly from `pcm16 + ID11/OAMD`, without a temporary ADM BWF.
- Keep the binaural DSP in float64/complex128, including 961-sample latency compensation, cross-frame state, and the room tail.
- Use a shared float32/PCM24 WAV writer and explicit PCM24 clipping policy for direct outputs.
- Use the NumPy backend or an optional C++20 core through `ctypes`; `auto` falls back to Python when the native library is unavailable.
- Read or write metadata sidecars and produce metadata, timing, and output reports.
@@ -30,10 +32,11 @@ M4A / E-AC-3
├─ ID11 OAMD → object positions and timing
└─ LFE + 15 objects
├─ 25ch ADM BWF
└─ speaker WAV for the selected layout
├─ speaker WAV for the selected layout
└─ direct ID11 timeline → Rosella → binaural WAV
```
The Python and C++ backends follow the same documented mathematics. The native core handles the state-heavy DSP and speaker rendering; high-level bitstream parsing, ADM assembly, and CLI behavior remain in Python.
The Python and C++ backends follow the same mathematics. The native core handles object reconstruction, speaker rendering, and binaural QMF/hybrid/room/synthesis; bitstream parsing, the OAMD timeline, model parsing, and CLI behavior remain in Python.
## Requirements
@@ -78,12 +81,29 @@ python main.py input.m4a --speaker-layout 5.1 --speaker-format int24
python main.py input.m4a --speaker-layout 7.1.2 --speaker-output output.7.1.2.wav
```
When PCM24 may clip in a non-interactive environment, select a policy explicitly:
Write Rosella binaural stereo directly (ordinary objects are Near/Mid/Far only; Mid is the default):
```powershell
python main.py input.m4a --binaural
python main.py input.m4a --binaural --binaural-mode near
python main.py input.m4a --binaural --binaural-mode far --binaural-format int24
python main.py input.m4a --binaural --binaural-output output.binaural.wav `
--personalized-headphone C:\HRTF\my.personalized_headphone
```
The default model path is `HRTF/binaural.personalized_headphone`. An example HRTF file is available from:
https://professionalsupport.dolby.com/s/question/0D54u0000AAT85HCQT/the-state-of-personalized-binaural-rendering?language=en_US
See [Binaural Rendering Mathematics](docs/binaural.en.md) for the formulas, state, and timing model.
Speaker and binaural output share peak analysis, the WAV writer, and clipping policy. When PCM24 may clip in a non-interactive environment, select a policy explicitly:
```powershell
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action abort
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action float32
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action continue
python main.py input.m4a --binaural --binaural-format int24 --clip-action abort
```
Metadata and diagnostics:
@@ -105,17 +125,11 @@ Use `--joc-binaural-mode off|near|far|mid|unspecified` to select a mode, encoded
python main.py input.m4a --joc-binaural-mode mid
```
This option only sets the low 3 binaural-render-mode bits of the last 15 JOC object entries in ADM BWF DBMD segment 10, leaving the first 10 bed entries unchanged. It does not change PCM, object trajectories, or direct speaker rendering, and does not itself produce binaural stereo audio. The adjacent `.report.json` records the mode name and value in `joc_binaural_mode` and `joc_binaural_mode_value`; both are `null` for direct speaker output, where the option does not apply.
This option only sets the low 3 binaural-render-mode bits of the last 15 JOC object entries in ADM BWF DBMD segment 10, leaving the first 10 bed entries unchanged. It does not change PCM, object trajectories, direct speaker rendering, or direct Rosella binaural rendering. The adjacent `.report.json` records the mode name and value in `joc_binaural_mode` and `joc_binaural_mode_value`; both are `null` for direct speaker or direct binaural output, where the option does not apply.
### OAMD time alignment
### Binaural calculation
Object trajectories and direct speaker rendering both default to a metadata delay of `1473 samples`. This value describes the theoretical mapping between decoder-output PCM and OAMD updates. The speaker renderer retains its existing 32-sample control block, so the default update lands on effective block boundary `1472`:
```text
align32(1473) = 1472
```
Override the two paths with `--object-delay-samples` and `--speaker-metadata-offset`, respectively. The 1473-sample timing offset is distinct from the 640-value inverse-QMF filter/window state; 640 is a QMF state length, not a metadata delay.
See [Binaural Rendering Mathematics](docs/binaural.en.md) for QMF, hybrid processing, direction fields, distance, ITD, room processing, 512-sample parameter updates, and 961-sample latency compensation.
For all options:
@@ -150,7 +164,7 @@ JustOneCacophony/
├─ main.py command-line entry point
├─ src/ Python implementation modules
├─ native/ C/C++ acceleration core, C ABI, and required table data
├─ data/ runtime table data for Python
├─ data/ Python runtime table data
├─ lib/ native runtime drop-in directory (create as needed)
├─ docs/ math and native-core notes in both languages
├─ requirements.txt Python dependency
@@ -170,7 +184,8 @@ The main documented stages are:
- OAMD Q15 coordinate conversion;
- equal-power panning over target-layout regions;
- layout-dependent position compensation and sample-wise gain ramps;
- float32 and PCM24 output quantization.
- float32 and PCM24 output quantization;
- Rosella 64-band QMF, 77-band hybrid processing, `77×36` direction fields, distance/ITD, room FIR, and special LFE.
See the [mathematical notes](docs/math.en.md) for the equations used by the decoding and rendering process.
@@ -178,13 +193,15 @@ See the [mathematical notes](docs/math.en.md) for the equations used by the deco
- Only the common contiguous EMDF transport is covered. Fragmented transport across multiple audio-block skip fields is not covered.
- Dense JOC is the main path. The Sparse JOC branch should not be treated as supported.
- The speaker path currently covers ordinary point objects; extent, spread, divergence, and similar modes are outside the supported scope.
- The speaker and Rosella binaural paths currently cover ordinary point objects; extent, spread, diffuse, divergence, channel lock, and similar controls are outside the supported scope.
- OAMD trim elements are boundary-checked and skipped; warp, balance, and trim parameters are not applied to raw object trajectories or speaker rendering.
- Multi-data-point streams, uncommon band configurations, and unusual OAMD scheduling have less coverage than common 12-band, single-data-point material.
- A speaker limiter is outside the current primary formula.
- ADM output, native binaries, and speaker layouts still need broader interoperability checks across platforms, players, and real material.
- Rosella requires a user-supplied compatible `.personalized_headphone`; arbitrary SOFA data cannot become a valid Rosella rp through JSON rearrangement alone.
- ADM output, native binaries, speaker layouts, and binaural models still need broader interoperability checks across platforms, players, and real material.
## Documentation
- [Mathematical notes](docs/math.en.md) · [中文](docs/math.md)
- [Native-core notes](docs/native.en.md) · [中文](docs/native.md)
- [DLL-free Rosella binaural](docs/binaural.en.md) · [中文](docs/binaural.md)