Compare commits
1 Commits
main
..
22a16ab60d
| Author | SHA1 | Date | |
|---|---|---|---|
| 22a16ab60d |
+4
-4
@@ -1,6 +1,5 @@
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
.pytest_cache/
|
||||
|
||||
.venv/
|
||||
venv/
|
||||
@@ -18,6 +17,7 @@ metadata_cache/
|
||||
|
||||
HRTF/
|
||||
|
||||
*.sofa
|
||||
*.personalized_headphone
|
||||
*.jochrtf
|
||||
# Keep the production binaural regression test while local research fixtures stay ignored.
|
||||
!tests/
|
||||
tests/*
|
||||
!tests/test_binaural_production.py
|
||||
|
||||
+27
-123
@@ -4,19 +4,19 @@
|
||||
|
||||
> JustOneCacophony is an experimental/test implementation of E-AC-3 JOC for studying JOC parsing, reconstruction, rendering, and the associated mathematics.
|
||||
|
||||
The project can extract and parse EMDF, ID14 JOC parameters, and ID11 OAMD metadata from common E-AC-3 JOC streams. It combines those data with the core 5.1 PCM decoded by FFmpeg, reconstructs LFE plus 15 object channels, and writes ADM BWF, a WAV file for a selected speaker layout, or direct binaural stereo using a standard SOFA HRTF.
|
||||
The project can extract and parse EMDF, ID14 JOC parameters, and ID11 OAMD metadata from common E-AC-3 JOC streams. It combines those data with the core 5.1 PCM decoded by FFmpeg, reconstructs LFE plus 15 object channels, and writes ADM BWF, a WAV file for a selected speaker layout, or direct DLL-free Rosella binaural stereo.
|
||||
|
||||
This is research code, not a complete, standards-compliant, or production-grade JOC decoder. It covers only the stream forms currently implemented. Unknown variants fail explicitly—because when the math goes wrong, all that may remain is the cacophony.
|
||||
|
||||
## Current features
|
||||
|
||||
- Scan common contiguous EMDF containers in E-AC-3 sync frames.
|
||||
- Parse ID14 dense / sparse JOC parameters, Huffman data, differential matrices, and `joc_clipgain`.
|
||||
- Parse ID14 dense JOC parameters, Huffman data, differential matrices, and `joc_clipgain`.
|
||||
- Parse ID11 OAMD position updates and build object trajectories.
|
||||
- Reconstruct LFE plus 15 object channels through analysis QMF, parameter interpolation, the object matrix, and inverse QMF.
|
||||
- Write a 25-channel ADM BWF: a 10-channel 7.1.2 bed (silent except for LFE) plus 15 objects.
|
||||
- Render directly to `2.0`, `3.1`, `5.1`, `7.1`, `5.1.2`, `5.1.4`, `7.1.2`, `7.1.4`, `9.1.4`, or `9.1.6`.
|
||||
- Run public SOFA binaural rendering directly from `pcm16 + ID11/OAMD`, without a temporary ADM BWF.
|
||||
- Run DLL-free Rosella binaural rendering directly from `pcm16 + ID11/OAMD`, without a temporary ADM BWF.
|
||||
- Keep the binaural DSP in float64/complex128, including 961-sample latency compensation, cross-frame state, and the room tail.
|
||||
- Use a shared float32/PCM24 WAV writer and explicit PCM24 clipping policy for direct outputs.
|
||||
- Use the NumPy backend or an optional C++20 core through `ctypes`; `auto` falls back to Python when the native library is unavailable.
|
||||
@@ -33,18 +33,16 @@ M4A / E-AC-3
|
||||
└─ LFE + 15 objects
|
||||
├─ 25ch ADM BWF
|
||||
├─ speaker WAV for the selected layout
|
||||
└─ direct ID11 timeline + SOFA HRTF → binaural WAV
|
||||
└─ direct ID11 timeline → Rosella → binaural WAV
|
||||
```
|
||||
|
||||
The Python and C++ backends follow the same mathematics for JOC object reconstruction and speaker rendering. The public SOFA binaural backend currently runs in Python; bitstream parsing, the OAMD timeline, and CLI behavior also remain in Python.
|
||||
The Python and C++ backends follow the same mathematics. The native core handles object reconstruction, speaker rendering, and binaural QMF/hybrid/room/synthesis; bitstream parsing, the OAMD timeline, model parsing, and CLI behavior remain in Python.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Python 3.10+
|
||||
- NumPy 1.24+
|
||||
- h5py 3.8+
|
||||
- SciPy 1.10+
|
||||
- A standalone FFmpeg executable; `ffmpeg-python` is not required. FFmpeg is discovered through `PATH` by default or selected with `--ffmpeg`. On startup the decoder options are probed with `ffmpeg -h decoder=eac3`: a missing E-AC-3 decoder or `-drc_scale` is a hard error, while a missing `-target_level` only fails when `--eac3-target-level` is used
|
||||
- A standalone FFmpeg executable; `ffmpeg-python` is not required. FFmpeg is discovered through `PATH` by default or selected with `--ffmpeg`
|
||||
- Optional: CMake and a C++20 toolchain to build the native core
|
||||
|
||||
Install the Python dependency in a project-specific environment:
|
||||
@@ -83,49 +81,21 @@ python main.py input.m4a --speaker-layout 5.1 --speaker-format int24
|
||||
python main.py input.m4a --speaker-layout 7.1.2 --speaker-output output.7.1.2.wav
|
||||
```
|
||||
|
||||
Write binaural stereo directly (ordinary objects are Near/Mid/Far only; Mid is
|
||||
the default). The HRTF input accepts three sources:
|
||||
Write Rosella binaural stereo directly (ordinary objects are Near/Mid/Far only; Mid is the default):
|
||||
|
||||
```powershell
|
||||
# 1) SOFA (defaults to HRTF/binaural.sofa, or an explicit path)
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone (defaults to HRTF/binaural.personalized_headphone)
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf compiled cache
|
||||
python main.py input.m4a --binaural --compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# Common options
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--binaural-mode near
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
python main.py input.m4a --binaural --binaural-output output.binaural.wav
|
||||
python main.py input.m4a --binaural --binaural-mode near
|
||||
python main.py input.m4a --binaural --binaural-mode far --binaural-format int24
|
||||
python main.py input.m4a --binaural --binaural-output output.binaural.wav `
|
||||
--personalized-headphone C:\HRTF\my.personalized_headphone
|
||||
```
|
||||
|
||||
With none of the three specified, resolution tries, in order:
|
||||
`HRTF/binaural.sofa`, the unique `.jochrtf` under `output/hrtf-cache`, then
|
||||
`HRTF/binaural.personalized_headphone`; if none exist, an error asks for an
|
||||
explicit path.
|
||||
The default model path is `HRTF/binaural.personalized_headphone`. An example HRTF file is available from:
|
||||
|
||||
- `.sofa` is the portable source of truth; it can hold self-scanned or any
|
||||
generic HRTF data.
|
||||
- `.personalized_headphone` is a model produced by Dolby's official
|
||||
personalization scan; its JSON parsing is implemented by this project
|
||||
(`src/rosella_model.py`) and does not invoke any Dolby software.
|
||||
- `.jochrtf` is a project-internal cache compiled from SOFA; it is disposable,
|
||||
rebuildable, and written to `output/hrtf-cache` by default.
|
||||
https://professionalsupport.dolby.com/s/question/0D54u0000AAT85HCQT/the-state-of-personalized-binaural-rendering?language=en_US
|
||||
|
||||
HRTF data lives under `HRTF/` (git-ignored): the default SOFA
|
||||
`HRTF/binaural.sofa` and the default model
|
||||
`HRTF/binaural.personalized_headphone`. Because the cache contains transformed
|
||||
HRTF data, its use and redistribution remain subject to the source dataset's
|
||||
terms. See [Binaural Rendering](docs/binaural.en.md) and
|
||||
[Third-party notices](THIRD_PARTY_NOTICES.md) for format boundaries, formulas,
|
||||
state, timing, and distribution considerations.
|
||||
See [Binaural Rendering Mathematics](docs/binaural.en.md) for the formulas, state, and timing model.
|
||||
|
||||
Speaker and binaural output share peak analysis, the WAV writer, and clipping policy. When PCM24 may clip in a non-interactive environment, select a policy explicitly:
|
||||
|
||||
@@ -133,8 +103,7 @@ Speaker and binaural output share peak analysis, the WAV writer, and clipping po
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action abort
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action float32
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action continue
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--binaural-format int24 --clip-action abort
|
||||
python main.py input.m4a --binaural --binaural-format int24 --clip-action abort
|
||||
```
|
||||
|
||||
Metadata and diagnostics:
|
||||
@@ -146,84 +115,21 @@ python main.py input.m4a --metadata-cache metadata_cache
|
||||
python main.py input.m4a --metadata-dir metadata_cache
|
||||
```
|
||||
|
||||
### E-AC-3 decode-side dynamic range and level
|
||||
### Experimental binaural mode settings for JOC objects
|
||||
|
||||
By default FFmpeg applies the stream `dynrng` dynamic range compression when
|
||||
decoding E-AC-3 (`-drc_scale 1`). The core 5.1 PCM is the input of JOC object
|
||||
reconstruction, and `dynrng` is playback-time gain, so it is inherited linearly
|
||||
by every object and every output (ADM, speaker, binaural). This tool therefore
|
||||
decodes at **full dynamic range** by default:
|
||||
The binaural mode written here is a user-selected, experimental rendering hint for downstream ADM renderers. It is **not original binaural metadata extracted or recovered from the input E-AC-3 JOC bitstream**, nor does it represent the original mix's per-object binaural settings. The selected mode is applied uniformly to all 15 JOC objects; the default `unspecified` is this tool's default, not a mode detected in the source file.
|
||||
|
||||
Use `--joc-binaural-mode off|near|far|mid|unspecified` to select a mode, encoded as `0|1|2|3|4` respectively. The default is `unspecified`:
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a # default: -drc_scale 0, full range
|
||||
python main.py input.m4a --eac3-drc-scale 1 # reproduce consumer playback
|
||||
python main.py input.m4a --eac3-drc-scale 0.5 # apply half of it
|
||||
python main.py input.m4a --eac3-target-level -27 # dialnorm-referenced level
|
||||
python main.py input.m4a --joc-binaural-mode mid
|
||||
```
|
||||
|
||||
- `--eac3-drc-scale` (`0`–`6`, default `0`) maps to FFmpeg `-drc_scale`: the gain
|
||||
of each E-AC-3 block is `dynrng factor ^ value`. `0` disables DRC, `1` is the
|
||||
author's intent, and `>1` is asymmetric (loud parts fully compressed, quiet
|
||||
parts enhanced).
|
||||
- `--eac3-target-level` (`-31`–`0`, default `0` = off) maps to FFmpeg
|
||||
`-target_level`: a static per-frame gain of about `target_level - dialnorm` dB,
|
||||
independent of and stackable with `--eac3-drc-scale`. dialnorm is a per-stream
|
||||
property (measured Apple Music Atmos streams are about `-18` to `-19` dB, so
|
||||
`-27` is roughly `8`–`9` dB of attenuation).
|
||||
- The level change is expected: compared with the FFmpeg default, measured
|
||||
tracks move by `0` to `-2.15` dB peak and `0` to `-1.69` dB RMS (direction
|
||||
depends on the stream `dynrng`), so `output_clip.peak` and the PCM24 clipping
|
||||
decision in `.report.json` change accordingly.
|
||||
- `--gain-db` is a static gain applied **after** reconstruction (float64 on the
|
||||
binaural path) and is not the same thing as decode-side DRC, which is
|
||||
block-varying; do not use `--gain-db` to cancel it.
|
||||
- The `ffmpeg` field of `.report.json` records the FFmpeg version and the decode
|
||||
options that were actually passed (`version`, `eac3_decode_options`).
|
||||
|
||||
### Binaural render mode
|
||||
|
||||
`--binaural-mode off|near|mid|far` selects the binaural render mode; the default
|
||||
is `mid`, and both outputs share this single option:
|
||||
|
||||
- **Direct binaural rendering** (`--binaural`): `off` is rejected (error);
|
||||
near/mid/far apply, defaulting to `mid`;
|
||||
- **ADM BWF**: the low 3 binaural-render-mode bits of the last 15 JOC object
|
||||
entries in DBMD segment 10 carry `off=0/near=1/far=2/mid=3`, leaving the first
|
||||
10 bed entries unchanged; the default is `mid`, and `off` explicitly disables
|
||||
the binaural metadata hint.
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a --binaural-mode mid
|
||||
python main.py input.m4a --binaural-mode off # ADM BWF only: disable the DBMD hint
|
||||
```
|
||||
|
||||
**The default `mid` is a human-specified rendering hint**; it is not original
|
||||
binaural metadata extracted or recovered from the input E-AC-3 JOC bitstream,
|
||||
nor does it represent the original mix's per-object binaural settings. The hint
|
||||
does not change PCM, object trajectories, or direct speaker rendering. The
|
||||
adjacent `.report.json` records `binaural_mode` (the mode name) and
|
||||
`binaural_mode_value` (the ADM code; `null` for direct binaural output).
|
||||
|
||||
### OAMD time alignment
|
||||
|
||||
Object trajectories and direct speaker rendering both default to a metadata delay of `1473 samples`. This value describes the theoretical mapping between decoder-output PCM and OAMD updates. The speaker renderer retains its existing 32-sample control block, so the default update lands on effective block boundary `1472`:
|
||||
|
||||
```text
|
||||
align32(1473) = 1472
|
||||
```
|
||||
|
||||
Override the two paths with `--object-delay-samples` and `--speaker-metadata-offset`, respectively. The 1473-sample timing offset is distinct from the 640-value inverse-QMF filter/window state; 640 is a QMF state length, not a metadata delay.
|
||||
|
||||
The direct binaural path uses `--object-delay-samples`. Each ID11/OAMD event is
|
||||
placed on an absolute sample timeline from its frame start, outer-subpayload
|
||||
offset, and block offset, then shifted by that delay. Each 1536-sample input
|
||||
frame is processed as three consecutive 512-sample blocks; the interpolated
|
||||
position, direction, and profile are updated at each block's absolute starting
|
||||
sample.
|
||||
This option only sets the low 3 binaural-render-mode bits of the last 15 JOC object entries in ADM BWF DBMD segment 10, leaving the first 10 bed entries unchanged. It does not change PCM, object trajectories, direct speaker rendering, or direct Rosella binaural rendering. The adjacent `.report.json` records the mode name and value in `joc_binaural_mode` and `joc_binaural_mode_value`; both are `null` for direct speaker or direct binaural output, where the option does not apply.
|
||||
|
||||
### Binaural calculation
|
||||
|
||||
See [Binaural Rendering Mathematics](docs/binaural.en.md) for QMF, hybrid processing, direction fields, distance, ITD, room processing, the 512-sample parameter updates above, and 961-sample latency compensation.
|
||||
See [Binaural Rendering Mathematics](docs/binaural.en.md) for QMF, hybrid processing, direction fields, distance, ITD, room processing, 512-sample parameter updates, and 961-sample latency compensation.
|
||||
|
||||
For all options:
|
||||
|
||||
@@ -260,8 +166,6 @@ JustOneCacophony/
|
||||
├─ native/ C/C++ acceleration core, C ABI, and required table data
|
||||
├─ data/ Python runtime table data
|
||||
├─ lib/ native runtime drop-in directory (create as needed)
|
||||
├─ HRTF/ user HRTF data directory (create as needed, git-ignored)
|
||||
├─ output/ output directory (create as needed; the .jochrtf cache defaults to its hrtf-cache subdirectory)
|
||||
├─ docs/ math and native-core notes in both languages
|
||||
├─ requirements.txt Python dependency
|
||||
├─ README.md Chinese documentation
|
||||
@@ -281,23 +185,23 @@ The main documented stages are:
|
||||
- equal-power panning over target-layout regions;
|
||||
- layout-dependent position compensation and sample-wise gain ramps;
|
||||
- float32 and PCM24 output quantization;
|
||||
- SOFA canonical import, 64-QMF/77-hybrid projection, `36×2×77` fifth-order fields, exactly-once delay/phase, project early/late room behavior, and special LFE.
|
||||
- Rosella 64-band QMF, 77-band hybrid processing, `77×36` direction fields, distance/ITD, room FIR, and special LFE.
|
||||
|
||||
See the [mathematical notes](docs/math.en.md) for the equations used by the decoding and rendering process.
|
||||
|
||||
## Known limitations
|
||||
|
||||
- Only the common contiguous EMDF transport is covered. Fragmented transport across multiple audio-block skip fields is not covered.
|
||||
- The speaker and SOFA binaural paths currently cover ordinary point objects; extent, spread, diffuse, divergence, channel lock, and similar controls are outside the supported scope.
|
||||
- Dense JOC is the main path. The Sparse JOC branch should not be treated as supported.
|
||||
- The speaker and Rosella binaural paths currently cover ordinary point objects; extent, spread, diffuse, divergence, channel lock, and similar controls are outside the supported scope.
|
||||
- OAMD trim elements are boundary-checked and skipped; warp, balance, and trim parameters are not applied to raw object trajectories or speaker rendering.
|
||||
- Multi-data-point streams, uncommon band configurations, and unusual OAMD scheduling have less coverage than common 12-band, single-data-point material.
|
||||
- A speaker limiter is outside the current primary formula.
|
||||
- The SOFA importer currently supports the strict `SimpleFreeFieldHRIR` FIR subset; other SOFA conventions require explicit adapters.
|
||||
- The binaural runtime is fixed at 48 kHz, fifth order, and one measurement-radius shell at a time; the public binaural backend defaults to the native accelerator and falls back to Python when the native library is unavailable.
|
||||
- Rosella requires a user-supplied compatible `.personalized_headphone`; arbitrary SOFA data cannot become a valid Rosella rp through JSON rearrangement alone.
|
||||
- ADM output, native binaries, speaker layouts, and binaural models still need broader interoperability checks across platforms, players, and real material.
|
||||
|
||||
## Documentation
|
||||
|
||||
- [Mathematical notes](docs/math.en.md) · [中文](docs/math.md)
|
||||
- [Native-core notes](docs/native.en.md) · [中文](docs/native.md)
|
||||
- [Binaural rendering](docs/binaural.en.md) · [中文](docs/binaural.md)
|
||||
- [DLL-free Rosella binaural](docs/binaural.en.md) · [中文](docs/binaural.md)
|
||||
|
||||
@@ -4,19 +4,19 @@
|
||||
|
||||
> JustOneCacophony 是一个 E-AC-3 JOC 的实验性 / 测试实现,用于研究 JOC 的解析、重建、渲染以及相关数学过程。
|
||||
|
||||
项目可以从常见 E-AC-3 JOC 码流中提取并解析 EMDF、ID14 JOC 参数和 ID11 OAMD 元数据,结合 FFmpeg 解码出的核心 5.1 PCM 重建 LFE 与 15 路对象 PCM,并输出 ADM BWF、指定扬声器布局的 WAV,或使用标准 SOFA HRTF 直接输出双耳 WAV。
|
||||
项目可以从常见 E-AC-3 JOC 码流中提取并解析 EMDF、ID14 JOC 参数和 ID11 OAMD 元数据,结合 FFmpeg 解码出的核心 5.1 PCM 重建 LFE 与 15 路对象 PCM,并输出 ADM BWF、指定扬声器布局的 WAV,或直接输出 DLL-free Rosella 双耳 WAV。
|
||||
|
||||
这是研究代码,不是完整、标准兼容或生产级的 JOC 解码器。它只覆盖当前已实现的码流形态;遇到未知变体时会明确报错,而不是假装一切都很和谐——如果哪里算错了,它可能就真的只剩 cacophony 了。
|
||||
|
||||
## 当前功能
|
||||
|
||||
- 扫描 E-AC-3 同步帧中的常见连续 EMDF 容器;
|
||||
- 解析 ID14 dense / sparse JOC 参数、Huffman 数据、差分矩阵与 `joc_clipgain`;
|
||||
- 解析 ID14 dense JOC 参数、Huffman 数据、差分矩阵与 `joc_clipgain`;
|
||||
- 解析 ID11 OAMD 位置更新并生成对象轨迹;
|
||||
- 通过 analysis QMF、参数插值、对象矩阵和 inverse QMF 重建 LFE + 15 路对象 PCM;
|
||||
- 输出 25 声道 ADM BWF:10 声道 7.1.2 bed(除 LFE 外静音)+ 15 个对象;
|
||||
- 直接渲染 `2.0`、`3.1`、`5.1`、`7.1`、`5.1.2`、`5.1.4`、`7.1.2`、`7.1.4`、`9.1.4`、`9.1.6`;
|
||||
- 从 `pcm16 + ID11/OAMD` 直接运行公开 SOFA 双耳渲染,不生成临时 ADM BWF;
|
||||
- 从 `pcm16 + ID11/OAMD` 直接运行 DLL-free Rosella 双耳渲染,不生成临时 ADM BWF;
|
||||
- 双耳 DSP 全程使用 float64/complex128,并保留 961-sample latency compensation、跨帧状态和 room 尾声;
|
||||
- 直接输出统一支持 float32 或 PCM24 WAV,并在 PCM24 削波前提供明确处理策略;
|
||||
- 使用 NumPy 后端,或通过 `ctypes` 调用可选的 C++20 原生核;`auto` 模式在原生库不可用时回退到 Python;
|
||||
@@ -33,20 +33,16 @@ M4A / E-AC-3
|
||||
└─ LFE + 15 objects
|
||||
├─ 25ch ADM BWF
|
||||
├─ 指定布局的扬声器 WAV
|
||||
└─ ID11 直接时间轴 + SOFA HRTF → 双耳 WAV
|
||||
└─ ID11 直接时间轴 → Rosella → 双耳 WAV
|
||||
```
|
||||
|
||||
Python 与 C++ 后端在 JOC 对象重建、扬声器渲染和公开 SOFA 双耳渲染中使用同一组
|
||||
数学过程;native 双耳后端与 Python 参考实现逐值一致(差异 < 1e-9)。位流解析、
|
||||
OAMD 时间轴和命令行逻辑在 Python 中。
|
||||
Python 与 C++ 后端使用同一组数学过程。原生核处理对象重建、扬声器渲染以及双耳 QMF/hybrid/room/synthesis;位流解析、OAMD 时间轴、模型解析和命令行逻辑仍在 Python 中。
|
||||
|
||||
## 环境
|
||||
|
||||
- Python 3.10+
|
||||
- NumPy 1.24+
|
||||
- h5py 3.8+
|
||||
- SciPy 1.10+
|
||||
- 独立的 FFmpeg 可执行程序;不需要 `ffmpeg-python`。默认从 `PATH` 查找,也可通过 `--ffmpeg` 指定可执行文件路径。启动时会探测 `ffmpeg -h decoder=eac3`:缺 E-AC-3 解码器或 `-drc_scale` 直接报错,缺 `-target_level` 只在使用 `--eac3-target-level` 时报错
|
||||
- 独立的 FFmpeg 可执行程序;不需要 `ffmpeg-python`。默认从 `PATH` 查找,也可通过 `--ffmpeg` 指定可执行文件路径
|
||||
- 可选:支持 C++20 的 CMake 工具链,用于自行构建原生核
|
||||
|
||||
建议在项目专用虚拟环境中安装依赖:
|
||||
@@ -85,41 +81,21 @@ python main.py input.m4a --speaker-layout 5.1 --speaker-format int24
|
||||
python main.py input.m4a --speaker-layout 7.1.2 --speaker-output output.7.1.2.wav
|
||||
```
|
||||
|
||||
直接输出双耳渲染 WAV(普通对象仅 Near/Mid/Far,默认 Mid)。HRTF 输入支持三种来源:
|
||||
直接输出 Rosella 双耳 WAV(普通对象仅 Near/Mid/Far,默认 Mid):
|
||||
|
||||
```powershell
|
||||
# 1) SOFA(缺省取 HRTF/binaural.sofa,也可显式指定)
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone(缺省取 HRTF/binaural.personalized_headphone)
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf 编译缓存
|
||||
python main.py input.m4a --binaural --compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# 常用选项
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--binaural-mode near
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
python main.py input.m4a --binaural --binaural-output output.binaural.wav
|
||||
python main.py input.m4a --binaural --binaural-mode near
|
||||
python main.py input.m4a --binaural --binaural-mode far --binaural-format int24
|
||||
python main.py input.m4a --binaural --binaural-output output.binaural.wav `
|
||||
--personalized-headphone C:\HRTF\my.personalized_headphone
|
||||
```
|
||||
|
||||
三者都不指定时的自动选择顺序:`HRTF/binaural.sofa` → `output/hrtf-cache` 下唯一的
|
||||
`.jochrtf` → `HRTF/binaural.personalized_headphone`;都没有则报错并提示显式指定。
|
||||
默认模型路径为 `HRTF/binaural.personalized_headphone`。示例 HRTF 文件见:
|
||||
|
||||
- `.sofa` 是可移植的 source of truth;可以是自行扫描或任何来源的通用 HRTF 数据。
|
||||
- `.personalized_headphone` 是杜比官方软件个性化扫描得到的模型,其 JSON 解析由
|
||||
本项目自行实现(`src/rosella_model.py`),不调用杜比软件。
|
||||
- `.jochrtf` 是从 SOFA 编译出的项目内部 cache,可删除、可从 SOFA 重建,默认写在
|
||||
`output/hrtf-cache`。
|
||||
https://professionalsupport.dolby.com/s/question/0D54u0000AAT85HCQT/the-state-of-personalized-binaural-rendering?language=en_US
|
||||
|
||||
HRTF 数据统一放在 `HRTF/`(git 忽略):默认 SOFA `HRTF/binaural.sofa`、默认模型
|
||||
`HRTF/binaural.personalized_headphone`。cache 含有源 HRTF 的变换数据,使用与再分发
|
||||
仍受源数据许可约束;格式边界、计算公式、状态、时间轴及发布注意事项见
|
||||
[双耳渲染](docs/binaural.md) 和 [第三方通知](THIRD_PARTY_NOTICES.md)。
|
||||
计算公式、状态和时间轴见[双耳渲染数学](docs/binaural.md)。
|
||||
|
||||
扬声器和双耳输出共享峰值检查、writer 与削波策略。在非交互环境请求 PCM24 且可能削波时,需要显式选择处理方式:
|
||||
|
||||
@@ -127,8 +103,7 @@ HRTF 数据统一放在 `HRTF/`(git 忽略):默认 SOFA `HRTF/binaural.sof
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action abort
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action float32
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action continue
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--binaural-format int24 --clip-action abort
|
||||
python main.py input.m4a --binaural --binaural-format int24 --clip-action abort
|
||||
```
|
||||
|
||||
元数据与诊断:
|
||||
@@ -140,60 +115,21 @@ python main.py input.m4a --metadata-cache metadata_cache
|
||||
python main.py input.m4a --metadata-dir metadata_cache
|
||||
```
|
||||
|
||||
### E-AC-3 解码级动态范围与电平
|
||||
### 实验性 JOC 对象双耳模式设置
|
||||
|
||||
FFmpeg 解码 E-AC-3 时默认施加码流 `dynrng` 动态范围压缩(`-drc_scale 1`)。核心 5.1 PCM 是 JOC 对象重建的输入,而 `dynrng` 属于回放期增益,会被线性继承到全部对象与成品(ADM/扬声器/双耳),因此本工具默认按**全动态范围**解码:
|
||||
这里写入的双耳模式是用户手动指定、供下游 ADM 渲染器使用的实验性渲染提示,**不是从输入 E-AC-3 JOC 码流中提取或还原的原始双耳元数据**,也不代表原始混音中各对象的双耳设置。所选模式会统一应用到 15 个 JOC 对象;默认 `unspecified` 只是本工具的默认值,并非从源文件检测到的模式。
|
||||
|
||||
使用 `--joc-binaural-mode off|near|far|mid|unspecified` 选择模式,编码分别为 `0|1|2|3|4`,默认 `unspecified`:
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a # 默认:-drc_scale 0,全动态范围
|
||||
python main.py input.m4a --eac3-drc-scale 1 # 复现消费者回放(码流作者意图)
|
||||
python main.py input.m4a --eac3-drc-scale 0.5 # 施加一半
|
||||
python main.py input.m4a --eac3-target-level -27 # 按码流 dialnorm 归一化电平
|
||||
python main.py input.m4a --joc-binaural-mode mid
|
||||
```
|
||||
|
||||
- `--eac3-drc-scale`(`0`~`6`,默认 `0`)对应 FFmpeg 的 `-drc_scale`:每个 E-AC-3 block 的增益为 `dynrng 因子 ^ 该值`。`0` 关闭 DRC;`1` 为码流作者意图;`>1` 非对称(响处全压、轻处增强)。
|
||||
- `--eac3-target-level`(`-31`~`0`,默认 `0` 不施加)对应 FFmpeg 的 `-target_level`:按每帧 dialnorm 施加静态增益,约 `target_level - dialnorm` dB,与 `--eac3-drc-scale` 相互独立、可叠加。dialnorm 是逐码流属性(实测 Apple Music Atmos 流约 `-18`~`-19` dB,故 `-27` 约等于衰减 `8`~`9` dB)。
|
||||
- 电平变化是预期的:与 FFmpeg 默认值相比,实测曲目峰值变化 `0`~`-2.15` dB、RMS `0`~`-1.69` dB(方向取决于码流 `dynrng`),`.report.json` 的 `output_clip.peak` 与 int24 削波判定会随之变化。
|
||||
- `--gain-db` 是**重建之后**的静态增益(双耳路径 float64),与解码级 DRC 不是一回事;解码级 DRC 是按 block 时变的,不要用 `--gain-db` 去抵消它。
|
||||
- `.report.json` 的 `ffmpeg` 字段记录 FFmpeg 版本与实际下发的解码选项(`version`、`eac3_decode_options`)。
|
||||
|
||||
### 双耳渲染模式
|
||||
|
||||
`--binaural-mode off|near|mid|far` 选择双耳渲染模式,默认 `mid`,两种输出共用这一个选项:
|
||||
|
||||
- **直接双耳渲染**(`--binaural`):`off` 不可用(报错),near/mid/far 生效,默认 `mid`;
|
||||
- **ADM BWF**:DBMD segment 10 中后 15 个 JOC 对象的 binaural render mode 写
|
||||
`off=0/near=1/far=2/mid=3`,前 10 个 bed 保持不变,默认 `mid`;`off` 用于显式
|
||||
关闭双耳元数据提示。
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a --binaural-mode mid
|
||||
python main.py input.m4a --binaural-mode off # 仅 ADM BWF:关闭 DBMD 双耳提示
|
||||
```
|
||||
|
||||
**默认 `mid` 是本工具人为指定的渲染提示**,不是从输入 E-AC-3 JOC 码流中提取或
|
||||
还原的原始双耳元数据,也不代表原始混音中各对象的双耳设置。该提示不改变 PCM、
|
||||
对象轨迹或直接扬声器渲染。输出旁的 `.report.json` 用 `binaural_mode`(模式名)
|
||||
和 `binaural_mode_value`(ADM 编码值,直接双耳输出时为 `null`)记录。
|
||||
|
||||
### OAMD 时间对齐
|
||||
|
||||
对象轨迹和直接扬声器渲染的 metadata delay 默认均为 `1473 samples`。该值描述 decoder 输出 PCM 与 OAMD 更新之间的理论时间映射;扬声器 renderer 仍使用现有的 32-sample control block,因此默认更新的实际 block boundary 为 `1472`:
|
||||
|
||||
```text
|
||||
align32(1473) = 1472
|
||||
```
|
||||
|
||||
可分别用 `--object-delay-samples` 和 `--speaker-metadata-offset` 覆盖默认值。这里的 1473 不应与 inverse-QMF 的 640 项 filter/window state 混淆;后者是 QMF 状态长度,不是 metadata delay。
|
||||
|
||||
直接双耳路径使用 `--object-delay-samples`。每个 ID11/OAMD event 先按 frame start、
|
||||
outer subpayload offset 与 block offset 落到绝对 sample timeline,再加该 delay;每个
|
||||
1536-sample 输入帧按三个连续 512-sample block 处理,并在每块的绝对起始 sample
|
||||
查询插值后的位置、更新方向和 profile。
|
||||
此选项仅设置 ADM BWF 的 DBMD segment 10 中后 15 个 JOC 对象的 binaural render mode 低 3 bit;前 10 个 bed 保持不变。它不改变 PCM、对象轨迹、直接扬声器渲染或直接 Rosella 双耳渲染。输出旁的 `.report.json` 用 `joc_binaural_mode` 和 `joc_binaural_mode_value` 记录模式名称与数值;直接扬声器或直接双耳输出时两者为 `null`,表示不适用。
|
||||
|
||||
### 双耳计算
|
||||
|
||||
双耳路径的 QMF、hybrid、方向场、距离、ITD、room、上述 512-sample 参数更新和 961-sample 延迟补偿见[双耳渲染数学](docs/binaural.md)。
|
||||
双耳路径的 QMF、hybrid、方向场、距离、ITD、room、512-sample 参数更新和 961-sample 延迟补偿见[双耳渲染数学](docs/binaural.md)。
|
||||
|
||||
更多参数可查看:
|
||||
|
||||
@@ -230,8 +166,6 @@ JustOneCacophony/
|
||||
├─ native/ C/C++ 加速核、C ABI 与必要表数据
|
||||
├─ data/ Python 运行时表数据
|
||||
├─ lib/ 原生运行库投放目录(按需创建)
|
||||
├─ HRTF/ 用户 HRTF 数据目录(按需创建,git 忽略)
|
||||
├─ output/ 输出目录(按需创建;.jochrtf 缓存默认在其 hrtf-cache 子目录)
|
||||
├─ docs/ 数学与原生核文档(中英文)
|
||||
├─ requirements.txt Python 依赖
|
||||
├─ README.md 中文说明
|
||||
@@ -251,23 +185,23 @@ JustOneCacophony/
|
||||
- 基于目标布局 region 的等功率声像;
|
||||
- 布局位置补偿与逐样本增益斜坡;
|
||||
- float32 与 PCM24 输出量化;
|
||||
- SOFA canonical importer、64-QMF/77-hybrid 投影、`36×2×77` 五阶方向 field、exactly-once delay/phase、项目 early/late room 与 special LFE。
|
||||
- Rosella 64-band QMF、77-band hybrid、`77×36` 方向 field、距离/ITD、room FIR 与 special LFE。
|
||||
|
||||
解码与渲染过程使用的公式见[数学说明](docs/math.md)。
|
||||
|
||||
## 已知限制
|
||||
|
||||
- 当前只覆盖常见 continuous EMDF transport;跨多个 audio-block skip field 的碎片化 transport 尚未覆盖。
|
||||
- 扬声器与 SOFA 双耳路径当前只覆盖普通点对象;extent、spread、diffuse、divergence、channel lock 等对象控制不在支持范围内。
|
||||
- Dense JOC 是当前主要路径;Sparse JOC 分支不应视为受支持能力。
|
||||
- 扬声器与 Rosella 双耳路径当前只覆盖普通点对象;extent、spread、diffuse、divergence、channel lock 等对象控制不在支持范围内。
|
||||
- OAMD trim element 会按声明边界校验并跳过;warp、balance 和 trim 参数不应用于当前原始对象轨迹或扬声器渲染。
|
||||
- 多数据点、少见参数带配置和特殊 OAMD 调度的覆盖度低于常见 12-band、单数据点素材。
|
||||
- 扬声器 limiter 不属于当前实现的主公式。
|
||||
- SOFA importer 当前严格支持 `SimpleFreeFieldHRIR` FIR;其它 SOFA convention 需要显式 adapter。
|
||||
- 双耳 runtime 固定 48 kHz、五阶和一次选择一个 measurement-radius shell;公开双耳默认走 native 加速,原生库不可用时自动回退 Python。
|
||||
- Rosella 路径需要用户提供兼容的 `.personalized_headphone`;任意 SOFA 不能仅靠 JSON 重排成为有效 Rosella rp。
|
||||
- ADM 输出、原生库、扬声器布局和双耳模型仍需在更多平台、播放器与真实素材上确认互操作性。
|
||||
|
||||
## 文档
|
||||
|
||||
- [数学说明](docs/math.md) · [English](docs/math.en.md)
|
||||
- [原生核说明](docs/native.md) · [English](docs/native.en.md)
|
||||
- [双耳渲染](docs/binaural.md) · [English](docs/binaural.en.md)
|
||||
- [DLL-free Rosella 双耳](docs/binaural.md) · [English](docs/binaural.en.md)
|
||||
|
||||
@@ -1,41 +0,0 @@
|
||||
# Third-party notices / 第三方通知
|
||||
|
||||
本文件记录 `data/rosella_kernels.npz`(`src/public_filterbank.py` 使用的滤波器组表)
|
||||
的公开标准来源,以及 HRTF 数据与专利的边界说明。
|
||||
|
||||
## 公开标准来源
|
||||
|
||||
64-QMF → 77-hybrid 结构与 13-tap 低带 prototype 定义于
|
||||
[3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
第 5.2.2 节(Table 1 的 $Q=8$/
|
||||
$Q=4$ 系数,delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr)$$
|
||||
|
||||
64-band QMF analysis 即 ISO/IEC 14496-3/AMD1:2003 第 4.B.18.2 节的 MPEG-4
|
||||
AAC/SBR 64 complex QMF bank;打包的 $64\times10$ 表是公开 640-tap prototype 的
|
||||
多相重排:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t}$$
|
||||
|
||||
QMF synthesis 表为 analysis 多相矩阵 $\mathbf{A}$ 的因果左逆
|
||||
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$(
|
||||
$\mathbf{P}$ 为 577-sample 延迟置换;
|
||||
全链 $961 = 577 + 6\times64$),rank-4 分解存储:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
hybrid synthesis 表为 77→64 重组:高频带恒等 $Y_{3+b}=X_{16+b}$,低频带:
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\mathrm{Re}X_q + j\,s_q\,\mathrm{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
相同数值可在 FFmpeg(`aacps_tablegen.h`、`aacsbrdata.h`)等公开实现中查到。
|
||||
|
||||
## HRTF 数据与 `.jochrtf`
|
||||
|
||||
`.jochrtf` 含有特定源 SOFA/HRTF 数据集的变换系数与 delay;其使用、复制与再分发
|
||||
仍受源数据集许可约束,权限不明确时应作为私有 cache 保存。
|
||||
|
||||
## 专利说明
|
||||
|
||||
标准可公开获取不等于获准实施相关专利。
|
||||
+2
-48
@@ -2,7 +2,7 @@
|
||||
|
||||
[中文](README.md)
|
||||
|
||||
This directory contains static production tables. It does not contain user HRTFs.
|
||||
This directory contains static production tables and the user-model directory.
|
||||
|
||||
`tables.npz` contains the JOC core decoding tables:
|
||||
|
||||
@@ -23,11 +23,9 @@ The corresponding native data are stored in `native/src/qmf_tables.h` and `nativ
|
||||
|
||||
## Binaural rendering tables
|
||||
|
||||
`rosella_kernels.npz` contains the fixed 64-QMF/77-hybrid tables used by the
|
||||
public SOFA binaural path:
|
||||
`rosella_kernels.npz` contains the fixed QMF/hybrid tables:
|
||||
|
||||
```text
|
||||
format_version little-endian int32[1]
|
||||
qmf_analysis_coefficients float32[64,10]
|
||||
hybrid_analysis_low_kernel float32[3,2,13,16,2]
|
||||
hybrid_synthesis_indices int16[154,4]
|
||||
@@ -37,47 +35,3 @@ qmf_synthesis_taps float64[64,10,4]
|
||||
```
|
||||
|
||||
The float32 table values are promoted to float64 when loaded.
|
||||
`src/public_filterbank.py` verifies the archive and every array by SHA-256.
|
||||
Those hashes, the table version, and the 77 reference band-center values all
|
||||
participate in the `.jochrtf` cache key. The full analysis/synthesis latency is
|
||||
961 samples.
|
||||
|
||||
The packaged tables implement publicly standardized filter banks, computable
|
||||
from the following formulas.
|
||||
|
||||
The 64-QMF → 77-hybrid structure, the 13-tap low-band prototypes, and their
|
||||
half-bin complex modulation are defined in
|
||||
[3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf),
|
||||
Section 5.2.2 (Table 1 $Q=8$/$Q=4$ coefficients, delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\!\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
The 64-band QMF analysis is the MPEG-4 AAC/SBR 64 complex QMF analysis bank of
|
||||
ISO/IEC 14496-3/AMD1:2003, subclause 4.B.18.2; the packaged $64\times10$ table
|
||||
is the polyphase reordering of the public 640-tap prototype $c_0,\dots,c_{639}$:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
The QMF synthesis table is the causal left inverse of the analysis polyphase
|
||||
matrix $\mathbf{A}$, i.e. the solution of $\mathbf{A}\,\mathbf{W}=\mathbf{P}$
|
||||
($\mathbf{P}$ is the 577-sample delay permutation; total latency
|
||||
$961 = 577 + 6\times64$), stored as a rank-4 factorization:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
The hybrid synthesis table is the 77→64 recombination: identity for the high
|
||||
bands, $Y_{3+b}=X_{16+b}$, and for the low bands ($C_p$ is the $8+4+4$ child
|
||||
partition):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\operatorname{Re}X_q + j\,s_q\,\operatorname{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
The same values also appear in other public implementations of these standards
|
||||
(for example FFmpeg's `aacps_tablegen.h` and `aacsbrdata.h`).
|
||||
|
||||
Public availability of a standard does not by itself grant permission to
|
||||
practice related patent claims.
|
||||
|
||||
SOFA is the user-visible source of truth. A `.jochrtf` file is a disposable JOC
|
||||
compiled HRTF cache that can be rebuilt from SOFA. The cache contains
|
||||
transformed source-HRTF data and remains subject to the source SOFA/HRTF
|
||||
dataset's licence and redistribution restrictions.
|
||||
|
||||
+3
-39
@@ -2,7 +2,7 @@
|
||||
|
||||
[English](README.en.md)
|
||||
|
||||
本目录保存 Python 生产路径使用的静态表数据,不保存用户 HRTF。
|
||||
本目录保存 Python 生产路径使用的静态表数据与用户模型目录。
|
||||
|
||||
`tables.npz` 保存 JOC 核心解码表:
|
||||
|
||||
@@ -23,10 +23,9 @@ joc_huff_code_7ch_pos_index_sparse int64[6,2]
|
||||
|
||||
## 双耳渲染表
|
||||
|
||||
`rosella_kernels.npz` 保存公开 SOFA 双耳路径使用的 64-QMF/77-hybrid 固定表:
|
||||
`rosella_kernels.npz` 保存双耳 QMF/hybrid 固定表:
|
||||
|
||||
```text
|
||||
format_version little-endian int32[1]
|
||||
qmf_analysis_coefficients float32[64,10]
|
||||
hybrid_analysis_low_kernel float32[3,2,13,16,2]
|
||||
hybrid_synthesis_indices int16[154,4]
|
||||
@@ -35,39 +34,4 @@ qmf_synthesis_basis float64[64,4,128]
|
||||
qmf_synthesis_taps float64[64,10,4]
|
||||
```
|
||||
|
||||
float32 表值载入后提升为 float64。`src/public_filterbank.py` 在读取时校验 archive
|
||||
及每个数组的 SHA-256;这些 hash、table version 和 77 个 band-center 参考值共同进入
|
||||
`.jochrtf` cache key。analysis/synthesis 全链 latency 为 961 samples。
|
||||
|
||||
打包表实现的是公开标准化的滤波器组,各表可由如下公式计算。
|
||||
|
||||
64-QMF → 77-hybrid 结构、13-tap 低带 prototype 与半 bin 复调制定义于
|
||||
[3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
第 5.2.2 节(Table 1 的 $Q=8$/$Q=4$ 系数,delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\!\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
64-band QMF analysis 即 ISO/IEC 14496-3/AMD1:2003 第 4.B.18.2 节的 MPEG-4
|
||||
AAC/SBR 64 complex QMF bank;打包的 $64\times10$ 表是公开 640-tap prototype
|
||||
$c_0,\dots,c_{639}$ 的多相重排:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
QMF synthesis 表为上述 analysis 多相矩阵 $\mathbf{A}$ 的因果左逆,即求解
|
||||
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$($\mathbf{P}$ 为 577-sample 延迟置换;
|
||||
全链 $961 = 577 + 6\times64$),以 rank-4 分解形式存储:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
hybrid synthesis 表为 77→64 重组:高频带恒等 $Y_{3+b}=X_{16+b}$;低频带
|
||||
($C_p$ 为 $8+4+4$ 子带划分):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\operatorname{Re}X_q + j\,s_q\,\operatorname{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
相同数值可在 FFmpeg(`aacps_tablegen.h`、`aacsbrdata.h`)等公开实现中查到。
|
||||
|
||||
标准可公开获取不等于获准实施相关专利。
|
||||
|
||||
`.sofa` 是用户可见的 source of truth;`.jochrtf` 是可删除、可从 SOFA 重建的
|
||||
JOC compiled HRTF cache。cache 含有源 HRTF 的变换数据,仍受源 SOFA/HRTF
|
||||
数据集的许可与再分发限制约束。
|
||||
float32 表值载入后提升为 float64。
|
||||
|
||||
+328
-264
@@ -1,284 +1,348 @@
|
||||
# Binaural rendering
|
||||
# JustOneCacophony — Binaural Rendering Mathematics
|
||||
|
||||
[中文](binaural.md) · [Back to README](../README.en.md)
|
||||
|
||||
JustOneCacophony's binaural backend supports three HRTF sources:
|
||||
`SimpleFreeFieldHRIR` SOFA, the Rosella `.personalized_headphone` model exported
|
||||
by Dolby's official personalization scan (its JSON parsing is implemented by
|
||||
this project and invokes no Dolby software), and the `.jochrtf` cache compiled
|
||||
from SOFA. SOFA is compiled into an in-memory directional field when the model
|
||||
is loaded. A `.jochrtf` file is only a disposable, reproducible JOC compiled
|
||||
HRTF cache; it is neither an interchange format nor a prerequisite for using
|
||||
SOFA.
|
||||
This document defines the `pcm16 + ID11/OAMD → stereo` calculation. The path begins after object reconstruction and does not pass through ADM BWF or AXML.
|
||||
|
||||
## 1. Signal path and notation
|
||||
|
||||
```text
|
||||
SOFA FIR
|
||||
-> CanonicalHrtf
|
||||
-> 48 kHz / one radius shell / delay-phase policy
|
||||
-> 64-QMF / 77-hybrid projection
|
||||
-> fifth-order ACN/N3D real-SH field
|
||||
-> per-object direct + early reflections
|
||||
-> shared unitary-FDN late room
|
||||
-> float64 stereo
|
||||
LFE + 15 object PCM channels
|
||||
→ 64-band QMF analysis
|
||||
→ 77-band hybrid analysis
|
||||
→ per-object geometry, transfer functions, and room send
|
||||
→ direct accumulation + room network
|
||||
→ hybrid synthesis
|
||||
→ QMF synthesis
|
||||
→ 961-sample latency compensation
|
||||
→ stereo WAV
|
||||
```
|
||||
|
||||
## Inputs
|
||||
|
||||
The CLI has three mutually exclusive HRTF input sources; with none given, a
|
||||
default rule resolves the input:
|
||||
|
||||
```powershell
|
||||
# 1) SOFA: defaults to HRTF/binaural.sofa, or an explicit path
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone: defaults to HRTF/binaural.personalized_headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf: explicitly load a compiled cache
|
||||
python main.py input.m4a --binaural `
|
||||
--compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# Optional: create/reuse a transparent disk cache for SOFA
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
```
|
||||
|
||||
The default order is `HRTF/binaural.sofa`, then the unique `.jochrtf` under
|
||||
`output/hrtf-cache`, then `HRTF/binaural.personalized_headphone`; if none of
|
||||
the three exist, an error asks for an explicit path. Multiple `.jochrtf` files
|
||||
under `output/hrtf-cache` are also an error requiring an explicit choice.
|
||||
|
||||
The `.personalized_headphone` JSON parsing is implemented by this project
|
||||
(`src/rosella_model.py`) and does not invoke any Dolby software.
|
||||
|
||||
`--hrtf-cache-policy` accepts `none`, `memory`, or `disk`. The default is
|
||||
`memory`; neither `none` nor `memory` creates a file. `disk` writes to
|
||||
`output/hrtf-cache` by default, or to `--hrtf-cache-dir`. `--hrtf-radius-m`
|
||||
selects the nearest measurement-radius shell.
|
||||
|
||||
The Python API also uses explicit factories:
|
||||
|
||||
```python
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
|
||||
renderer = SofaBinauralBackend.from_sofa(
|
||||
"subject.sofa",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
cache_policy="memory",
|
||||
)
|
||||
|
||||
cached = SofaBinauralBackend.from_compiled_cache(
|
||||
"subject.jochrtf",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
)
|
||||
```
|
||||
|
||||
The factories never guess a format from an unknown suffix: SOFA and `.jochrtf`
|
||||
always use distinct loaders.
|
||||
|
||||
## Binaural render mode
|
||||
|
||||
`--binaural-mode off|near|mid|far` (default `mid`) is a **human-specified
|
||||
rendering hint**, not original binaural metadata extracted or recovered from the
|
||||
input E-AC-3 JOC bitstream:
|
||||
|
||||
- Direct binaural rendering (`--binaural`): near/mid/far apply, default `mid`;
|
||||
`off` is an error;
|
||||
- ADM BWF: the low 3 binaural-render-mode bits of the last 15 JOC object entries
|
||||
in DBMD segment 10 carry `off=0/near=1/far=2/mid=3`, leaving the first 10 bed
|
||||
entries unchanged; the default is `mid`, and `off` explicitly disables the
|
||||
binaural metadata hint.
|
||||
|
||||
## Canonical SOFA contract
|
||||
|
||||
The strict importer currently accepts:
|
||||
|
||||
- `Conventions=SOFA`;
|
||||
- `SOFAConventions=SimpleFreeFieldHRIR`, version `0.4`, `1.0`, or `1.1`;
|
||||
- `DataType=FIR` and `Data.IR[M,2,N]`;
|
||||
- one positive finite `Data.SamplingRate` in hertz/Hz;
|
||||
- spherical or Cartesian `SourcePosition`;
|
||||
- singleton or per-measurement `ListenerPosition/View/Up`;
|
||||
- two receivers whose listener-local lateral geometry uniquely identifies L/R;
|
||||
- one zero-offset emitter;
|
||||
- causal `Data.Delay[I,2]` or `[M,2]`;
|
||||
- an explicitly free-field/anechoic `RoomType`.
|
||||
|
||||
Receiver order comes from geometry, never from the receiver array index. SOFA
|
||||
listener coordinates are $+X$ front,
|
||||
$+Y$ left,
|
||||
$+Z$ up; ADM coordinates are
|
||||
$+X$ right,
|
||||
$+Y$ front,
|
||||
$+Z$ up:
|
||||
|
||||
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
|
||||
|
||||
`CanonicalHrtf` keeps `Data.IR` and `Data.Delay` separate. Only a time-domain
|
||||
baseline calls `materialized_measurement()` to apply delay once; the runtime SH
|
||||
path never materializes and then restores the delay. Non-48-kHz HRIRs are
|
||||
normalized with float64 `scipy.signal.resample_poly`, and delay samples scale by
|
||||
the same ratio.
|
||||
|
||||
GeneralFIR, BRIR, TF, multiple emitters, ambiguous receivers, and non-free-field
|
||||
data require convention-specific adapters. They cannot enter the core importer
|
||||
through a reshape.
|
||||
|
||||
## Exactly-once delay and phase
|
||||
|
||||
The compiler recognizes three mutually exclusive representations:
|
||||
|
||||
1. Nonzero `Data.Delay` is external to `Data.IR`; the FIR is not de-rotated and
|
||||
runtime applies the delay once.
|
||||
2. With `Data.Delay=0` and an ordinary positive-onset HRIR, each ear's main peak
|
||||
supplies arrival time. Compilation separates it and runtime restores it once.
|
||||
The current threshold is a peak index greater than two samples.
|
||||
3. With `Data.Delay=0` and both FIRs at a shared sample-zero origin, no external
|
||||
delay is invented. The authored complex phase stays in the fifth-order field.
|
||||
|
||||
No path may add a second ear delay or phase-group delay.
|
||||
|
||||
## Public filterbank and directional field
|
||||
|
||||
The runtime is fixed at:
|
||||
|
||||
- 48 kHz;
|
||||
- a 64-sample QMF hop;
|
||||
- 64-QMF / 77 hybrid bands;
|
||||
- 961 samples of analysis/synthesis latency;
|
||||
- fifth order, 36 terms, ACN/N3D real spherical harmonics;
|
||||
- float64 PCM, delay, SH, and room state; complex128 band transfers and spectra.
|
||||
|
||||
Real and imaginary unit gains for every hybrid band pass through the same
|
||||
analysis/synthesis chain to form a 154-real-parameter impulse dictionary. The
|
||||
compiler does not sample 77 FFT bins. Defaults are `1e-3` projection ridge and
|
||||
`1e-5` SH ridge. Coincident directions are merged before a spherical-Voronoi
|
||||
weighted ridge fit.
|
||||
|
||||
The fixed resource is `data/rosella_kernels.npz`, which implements publicly
|
||||
standardized filter banks, computable from the following formulas.
|
||||
|
||||
The hybrid analysis kernels are defined in [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf),
|
||||
Section 5.2.2 (Table 1 $Q=8$/
|
||||
$Q=4$ coefficients, delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
The QMF analysis table is the MPEG-4 AAC/SBR 64 complex QMF bank of
|
||||
ISO/IEC 14496-3/AMD1:2003, subclause 4.B.18.2, stored as the polyphase
|
||||
reordering of the public 640-tap prototype $c_0,\dots,c_{639}$:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
The QMF synthesis table is the causal left inverse of the analysis polyphase
|
||||
matrix $\mathbf{A}$, i.e. the solution of
|
||||
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$
|
||||
($\mathbf{P}$ is the 577-sample delay permutation; total latency
|
||||
$961 = 577 + 6\times64$), stored as a rank-4 factorization:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
The hybrid synthesis table is the 77→64 recombination: identity for the high
|
||||
bands, $Y_{3+b}=X_{16+b}$, and for the low bands(
|
||||
$C_p$ is the
|
||||
$8+4+4$ child partition):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\mathrm{Re}X_q + j\,s_q\,\mathrm{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
The loader verifies the archive and every array by SHA-256; the table version,
|
||||
all array hashes, and the 77 reference band-center values are part of the cache
|
||||
key. Public availability of a standard does not by itself grant permission to
|
||||
practice related patent claims. See
|
||||
[`data/README.en.md`](../data/README.en.md) and
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the sources and the
|
||||
rights boundary.
|
||||
|
||||
## `.jochrtf`
|
||||
|
||||
A `.jochrtf` file is a pickle-free compressed NumPy archive with an exact member set:
|
||||
|
||||
| key | dtype / shape |
|
||||
| Symbol | Meaning |
|
||||
|---|---|
|
||||
| `metadata_json` | NumPy Unicode scalar containing JSON text (`dtype.kind == "U"`) |
|
||||
| `band_center_frequencies_hz` | little-endian `float64[77]` |
|
||||
| `coefficients` | little-endian `complex128[36,2,77]` |
|
||||
| `delay_coefficients` | little-endian `float64[36,2]` |
|
||||
| `delay_bounds` | little-endian `float64[2,2]` |
|
||||
| $s=0\ldots15$ | input source; source 0 is LFE |
|
||||
| $e\in\{L,R\}$ | output ear |
|
||||
| $k=0\ldots63$ | QMF band |
|
||||
| $h=0\ldots76$ | hybrid band |
|
||||
| $j=0\ldots35$ | direction-basis term |
|
||||
| $m$ | 64-sample QMF slot |
|
||||
|
||||
Metadata uses the `JOC-HRTF-CACHE` magic and records the schema, compiler and
|
||||
phase-policy versions, ACN/N3D convention, filterbank hashes, SOFA content
|
||||
SHA-256, sample rate, radius, order, both ridge values, payload hash, and fit
|
||||
report. Every setting that changes compilation participates in the cache key.
|
||||
Metadata never persists an absolute local `source_path`; it may keep a display
|
||||
name only.
|
||||
A control block is
|
||||
|
||||
Before constructing a field, the loader uses `allow_pickle=False` and validates
|
||||
ZIP members and expanded sizes, shapes, dtypes, byte order, contiguous layout,
|
||||
finite values, delay bounds, band centers, payload hash, and cache key. The
|
||||
writer uses a same-directory temporary file, `fsync`, a process-held OS file
|
||||
lock, and atomic `os.replace`. Its hidden `.lock` sidecar may remain and does not
|
||||
mean that a writer still owns the lock. Outdated, damaged, or mismatched
|
||||
caches cannot hit. SOFA input rebuilds an invalid cache; an explicitly selected
|
||||
cache reports the error.
|
||||
$$N_b=512=8\times64,$$
|
||||
|
||||
Deleting a disk cache must not change the field or render produced from the same
|
||||
SOFA and compiler configuration.
|
||||
and an input frame is
|
||||
|
||||
A `.jochrtf` file contains directional-field coefficients and delay data
|
||||
transformed from the source HRIRs. Its reproducibility therefore does not make
|
||||
it licence-free. Creating a cache does not enlarge the rights granted by the
|
||||
source SOFA/HRTF dataset: use, copying, and redistribution remain subject to
|
||||
that dataset's terms. If those terms are unclear, keep `.jochrtf` as a private
|
||||
local cache and do not ship it with the program or another build artifact.
|
||||
`source_sha256` is only a content-integrity identifier, not proof of provenance
|
||||
or permission.
|
||||
$$N_f=1536=3N_b.$$
|
||||
|
||||
## JOC objects and room behavior
|
||||
All filter and room state continues across frame boundaries.
|
||||
|
||||
The production adapter retains the existing JOC schedule:
|
||||
## 2. QMF analysis
|
||||
|
||||
- `[1536,16]` input per frame;
|
||||
- channel 0 is special LFE and channels 1..15 are JOC objects;
|
||||
- ID11/OAMD positions use a sample-timed timeline;
|
||||
- source parameters update every 512 samples;
|
||||
- every object owns independent direct/early history while one late FDN is shared;
|
||||
- `finish()` drains early/late tails; output gain is explicit, with no implicit
|
||||
limiter or programme loudness normalization.
|
||||
Let $a_{p,\ell}$ be the fixed 64×10 polyphase coefficients and $r_{s,\ell,p}[m]$ the current and previous nine phase vectors:
|
||||
|
||||
Near/Mid/Far, equal-power direct level, six first-order shoebox image sources,
|
||||
late sends, the unitary FDN, the 120–180 Hz cosine-squared LFE low-pass, and room
|
||||
calibration are JOC project-defined behavior, not constants published by SOFA or
|
||||
Dolby.
|
||||
$$
|
||||
E_{s,p}[m]=\sum_{\ell\text{ even}}a_{p,\ell}r_{s,\ell,p}[m],
|
||||
$$
|
||||
|
||||
The public SOFA binaural renderer defaults to the C++20 native core under
|
||||
`--backend auto/native` (`ejoc_sofa_binaural_*` in `lib/eac3joc_core.dll`): the
|
||||
filterbank, the SH direction-field evaluation, the per-object direct/early
|
||||
histories and the shared FDN all run natively, while Python only compiles the
|
||||
SOFA source and issues the per-512-sample metadata updates. When the native
|
||||
library is unavailable the renderer falls back to the Python/NumPy reference
|
||||
implementation; the two agree to better than 1e-9. `--backend python` forces
|
||||
the Python backend.
|
||||
`--backend` still selects native/Python JOC reconstruction and speaker rendering;
|
||||
native acceleration for the public binaural DSP is outside the current API.
|
||||
$$
|
||||
O_{s,p}[m]=\sum_{\ell\text{ odd}}a_{p,\ell}r_{s,\ell,p}[m].
|
||||
$$
|
||||
|
||||
## Technical references and rights boundary
|
||||
Define
|
||||
|
||||
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
|
||||
- [3GPP TS 26.405 / ETSI TS 126 405 (64-QMF/77-hybrid definition)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
- [Dolby binaural render-mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
|
||||
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en), used only as
|
||||
architectural background for direct/early/late, subbands, and FDNs; it does
|
||||
not establish that any product uses a particular embodiment.
|
||||
$$
|
||||
\mathcal Q(v)_k=
|
||||
\operatorname{FFT}_{128}
|
||||
\left([v[p]e^{-j\pi p/128}]_{p=0}^{63},0_{64}\right)_k
|
||||
e^{-j3\pi(k+1/2)/128}.
|
||||
$$
|
||||
|
||||
Public availability of a specification, source file, or patent document does
|
||||
not by itself authorize copying its contents, redistribution of derivatives,
|
||||
or practice of patent claims. These technical references grant no patent
|
||||
licence and make no non-infringement representation. Anyone preparing a release
|
||||
or product integration must assess the applicable data and software licences,
|
||||
patent permissions, and freedom to operate. See
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the public-standard
|
||||
provenance and rights boundary.
|
||||
The complex QMF output is
|
||||
|
||||
$$
|
||||
X_{s,k}[m]=\mathcal Q(O_s)_k+j(-1)^k\mathcal Q(E_s)_k.
|
||||
$$
|
||||
|
||||
## 3. Hybrid analysis
|
||||
|
||||
The lowest three QMF bands are split into sixteen hybrid bands by a 13-slot FIR:
|
||||
|
||||
$$
|
||||
H_{s,h,o}[m]
|
||||
=
|
||||
\sum_{p=0}^{2}\sum_{i=0}^{1}\sum_{\ell=0}^{12}
|
||||
X_{s,p,i}[m-\ell]K_{p,i,\ell,h,o},
|
||||
\qquad h=0\ldots15.
|
||||
$$
|
||||
|
||||
The remaining bands are delayed QMF bands 3..63:
|
||||
|
||||
$$
|
||||
H_{s,16+q}[m]=X_{s,3+q}[m-6],
|
||||
\qquad q=0\ldots60.
|
||||
$$
|
||||
|
||||
## 4. OAMD coordinates and time
|
||||
|
||||
The Q15 object fields are restored to their discrete grids:
|
||||
|
||||
$$
|
||||
u_1=\min\left(1,\frac{\operatorname{round}(62q_1/32767)}{62}\right),$$
|
||||
|
||||
$$
|
||||
u_2=\min\left(1,\frac{\operatorname{round}(62q_2/32767)}{62}\right),$$
|
||||
|
||||
$$
|
||||
u_3=\operatorname{clip}\left(
|
||||
\frac{\operatorname{round}(15q_3/32767)}{15},-1,1\right),$$
|
||||
|
||||
$$
|
||||
(X,Y,Z)=(2u_1-1,\ 1-2u_2,\ u_3).
|
||||
$$
|
||||
|
||||
An update is coded at
|
||||
|
||||
$$
|
||||
n_{\mathrm{coded}}
|
||||
=n_{\mathrm{frame}}+n_{\mathrm{outer}}+n_{\mathrm{block}}.
|
||||
$$
|
||||
|
||||
The first valid state is the position at sample 0. Later updates add the object delay $D_o=1473$. For $R>64$:
|
||||
|
||||
$$
|
||||
n_{\mathrm{start}}=n_{\mathrm{coded}}+D_o+64,$$
|
||||
|
||||
$$R_{\mathrm{eff}}=R-64,$$
|
||||
|
||||
$$
|
||||
\mathbf p[n]=(1-\alpha)\mathbf p_0+\alpha\mathbf p_1,
|
||||
\qquad
|
||||
\alpha=\frac{n-n_{\mathrm{start}}}{R_{\mathrm{eff}}}.
|
||||
$$
|
||||
|
||||
The position is evaluated at each 512-sample block boundary.
|
||||
|
||||
## 5. Distance profile and direction
|
||||
|
||||
Each Near, Mid, or Far profile contains six bounds, distance scale $D$, inverse scale $D^{-1}$, three axis scales, and minimum radius $\rho_{\min}$.
|
||||
|
||||
After axis conversion and scale:
|
||||
|
||||
$$
|
||||
\mathbf s=(a_zq_f,a_xq_l,a_yq_v).
|
||||
$$
|
||||
|
||||
A single ray factor $\lambda\le1$ keeps the point inside the profile bounds:
|
||||
|
||||
$$
|
||||
\mathbf s'=\lambda\mathbf s.
|
||||
$$
|
||||
|
||||
Then
|
||||
|
||||
$$
|
||||
\rho=\|\mathbf s'\|_2,
|
||||
\quad
|
||||
\rho_c=\max(\rho,\rho_{\min}),
|
||||
\quad
|
||||
\alpha=\rho/\rho_c,
|
||||
$$
|
||||
|
||||
$$
|
||||
\mathbf d=\mathbf s'/\rho,
|
||||
\qquad
|
||||
R=D\rho.
|
||||
$$
|
||||
|
||||
## 6. Direction basis and ear paths
|
||||
|
||||
The direction is expanded into a fixed 36-term polynomial basis:
|
||||
|
||||
$$
|
||||
\mathbf b(\mathbf d)=
|
||||
[1,x,y,z,x^2-\tfrac13,xy,xz,y^2-\tfrac13,yz,\ldots]^T.
|
||||
$$
|
||||
|
||||
For ear offset $e$:
|
||||
|
||||
$$
|
||||
\epsilon=\frac{eD^{-1}}{\rho_c},
|
||||
$$
|
||||
|
||||
$$
|
||||
\mathbf d_{\mp}=
|
||||
\frac{(x,y\mp\epsilon,z)}{\|(x,y\mp\epsilon,z)\|_2}.
|
||||
$$
|
||||
|
||||
The normalized paths are
|
||||
|
||||
$$
|
||||
\ell_{\mp}=\rho_c\sqrt{x^2+(y\mp\epsilon)^2+z^2}.
|
||||
$$
|
||||
|
||||
A model direction vector may add a non-negative path correction:
|
||||
|
||||
$$
|
||||
\ell'_e=\ell_e+
|
||||
\max(\mathbf v_e^T\mathbf b_e,0)\,2cD^{-1}.
|
||||
$$
|
||||
|
||||
The interaural delay is
|
||||
|
||||
$$
|
||||
\tau=|\ell'_+-\ell'_-|D\frac{48000}{343.3}\alpha.
|
||||
$$
|
||||
|
||||
The longer path receives the hybrid phase
|
||||
|
||||
$$P_h=e^{j\omega_h\tau}.$$
|
||||
|
||||
## 7. Direction fields and direct gains
|
||||
|
||||
Each ear has a 77×36 complex field:
|
||||
|
||||
$$
|
||||
C_{e,h}(\mathbf d_e)=
|
||||
\sum_{j=0}^{35}F_{e,h,j}b_j(\mathbf d_e).
|
||||
$$
|
||||
|
||||
Path weights are
|
||||
|
||||
$$
|
||||
w_L=\frac{\ell_+}{\sqrt{\ell_-^2+\ell_+^2}},
|
||||
\qquad
|
||||
w_R=\frac{\ell_-}{\sqrt{\ell_-^2+\ell_+^2}}.
|
||||
$$
|
||||
|
||||
For effective distance $R_e=\rho s_dD$, Mid and Far use
|
||||
|
||||
$$
|
||||
g_c=\frac{1}{\sqrt{1+s_rR_e^2}},
|
||||
\qquad
|
||||
g_{\mathrm{room}}=R_eg_c.
|
||||
$$
|
||||
|
||||
Near uses $g_c=1$ and $g_{\mathrm{room}}=0$. With field term zero denoted by $C^{(0)}$:
|
||||
|
||||
$$
|
||||
G_{L,h}=g_c[C_{L,h}w_L\alpha+C_{L,h}^{(0)}c_L(1-\alpha)],
|
||||
$$
|
||||
|
||||
$$
|
||||
G_{R,h}=g_c[C_{R,h}w_R\alpha+C_{R,h}^{(0)}c_R(1-\alpha)].
|
||||
$$
|
||||
|
||||
## 8. LFE
|
||||
|
||||
LFE bypasses ordinary-object geometry:
|
||||
|
||||
$$
|
||||
G_{L,h}^{\mathrm{LFE}}=G_{R,h}^{\mathrm{LFE}}=
|
||||
\begin{cases}
|
||||
g_h,&0\le h<16,\\0,&16\le h<77.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
```text
|
||||
2.60290003, 1.80741799, 0.659342408, -0.0275855921,
|
||||
-0.105803289, -0.0699509233, 0.0749056414, -0.00919809937,
|
||||
0.00349014648,-0.0158600751,-0.000723021978,0.00188189559,
|
||||
-0.000421735429,0.0000329252762,0.0000317397971,0.000000580376991
|
||||
```
|
||||
|
||||
Its room send is zero.
|
||||
|
||||
## 9. Source accumulation and room network
|
||||
|
||||
Direct output and room input are
|
||||
|
||||
$$
|
||||
Y^{\mathrm{direct}}_{e,h}=
|
||||
\sum_{s=0}^{15}H_{s,h}G_{s,e,h},
|
||||
$$
|
||||
|
||||
$$
|
||||
U_h=\sum_{s=1}^{15}H_{s,h}g_{\mathrm{room},s}.
|
||||
$$
|
||||
|
||||
The room input is scaled by $0.70710677$. Each all-pass stage uses
|
||||
|
||||
$$r[n]=x[n]-ad[n],$$
|
||||
|
||||
$$y[n]=ar[n]+d[n].$$
|
||||
|
||||
For the four-branch delay network:
|
||||
|
||||
$$
|
||||
\mathbf b_h[m]=U_h[m]\mathbf1+M\mathbf d_h[m],
|
||||
$$
|
||||
|
||||
$$m_{h,i}[m]=f_{h,i}b_{h,i}[m].$$
|
||||
|
||||
The main tap, optional extra taps, and ear output matrices produce
|
||||
|
||||
$$
|
||||
Y^{\mathrm{room}}_{e,h}[m]=
|
||||
\sum_{i=0}^{3}O_{e,h,i}z_{h,i}[m].
|
||||
$$
|
||||
|
||||
The final hybrid signal is
|
||||
|
||||
$$Y_{e,h}=Y^{\mathrm{direct}}_{e,h}+Y^{\mathrm{room}}_{e,h}.$$
|
||||
|
||||
The Python backend uses a finite complex FIR/overlap-add realization. The C++ backend keeps the recursive room state directly.
|
||||
|
||||
## 10. Hybrid and QMF synthesis
|
||||
|
||||
Hybrid synthesis is a 154-entry sparse map. For an entry $(h,i,k,o,w)$:
|
||||
|
||||
$$Q_{e,k,o}[m]\mathrel{+}=Y_{e,h,i}[m]w.$$
|
||||
|
||||
The complex QMF vector is flattened to
|
||||
|
||||
$$
|
||||
\mathbf q_e=[\Re Q_{e,0},\Im Q_{e,0},\ldots,\Re Q_{e,63},\Im Q_{e,63}]^T.
|
||||
$$
|
||||
|
||||
Rank-four features and ten-slot synthesis are
|
||||
|
||||
$$f_{e,p,r}[m]=\mathbf b_{p,r}^T\mathbf q_e[m],$$
|
||||
|
||||
$$
|
||||
y_e[64m+p]=
|
||||
\sum_{\ell=0}^{9}\sum_{r=0}^{3}
|
||||
t_{p,\ell,r}f_{e,p,r}[m-\ell].
|
||||
$$
|
||||
|
||||
## 11. Latency, tail, and precision
|
||||
|
||||
The filterbank latency is 961 samples and is removed once at the beginning of the continuous stream. Zero input is then processed to release filterbank and room state. Tail trimming keeps the final sample satisfying
|
||||
|
||||
$$
|
||||
\max(|y_L[n]|,|y_R[n]|)>10^{-8},
|
||||
$$
|
||||
|
||||
while never shortening the output below the source PCM length.
|
||||
|
||||
All internal state, geometry, field products, room processing, source accumulation, and tail processing use `float64/complex128`. Conversion to float32 or PCM24 occurs only in the final writer.
|
||||
|
||||
## 12. Backends and model path
|
||||
|
||||
Python and C++ use the same fixed tables, parsed model parameters, 512-sample control timeline, direct gains, room sends, latency compensation, and tail policy.
|
||||
|
||||
The C++ backend owns QMF, hybrid, recursive room, and synthesis state. Python supplies parsed parameters and per-block gains.
|
||||
|
||||
The default model path is
|
||||
|
||||
```text
|
||||
HRTF/binaural.personalized_headphone
|
||||
```
|
||||
|
||||
Override it with `--personalized-headphone PATH`.
|
||||
|
||||
A SOFA FIR cannot be converted into this parameter model by array rearrangement alone. A conversion requires fitting the direction fields, ITD, distance profiles, ear geometry, and room parameters.
|
||||
|
||||
## 13. Scope
|
||||
|
||||
The current path covers fifteen point objects and one special LFE source. Extent, spread, diffuse, divergence, channel lock, and unsupported OAMD element variants are outside this model.
|
||||
|
||||
+515
-224
@@ -1,245 +1,536 @@
|
||||
# 双耳渲染
|
||||
# JustOneCacophony — 双耳渲染数学
|
||||
|
||||
[English](binaural.en.md) · [返回 README](../README.md)
|
||||
|
||||
JustOneCacophony 的双耳后端支持三种 HRTF 来源:`SimpleFreeFieldHRIR` SOFA、
|
||||
杜比官方软件个性化扫描导出的 Rosella `.personalized_headphone`(JSON 解析由本项目
|
||||
自行实现,不调用杜比软件),以及从 SOFA 编译出的 `.jochrtf` 缓存。SOFA 在模型加载
|
||||
时编译成内存方向场;`.jochrtf` 只是可删除、可重建的 JOC compiled HRTF cache,
|
||||
不是交换格式,也不是使用 SOFA 的前置步骤。
|
||||
本文说明 `pcm16 + ID11/OAMD → stereo` 路径中的计算、状态和时间对齐。双耳渲染直接接在对象重建之后,不经过 ADM BWF 或 AXML。
|
||||
|
||||
## 1. 总体路径与记号
|
||||
|
||||
```text
|
||||
SOFA FIR
|
||||
-> CanonicalHrtf
|
||||
-> 48 kHz / 单 radius shell / delay-phase policy
|
||||
-> 64-QMF / 77-hybrid projection
|
||||
-> 五阶 ACN/N3D 实球谐场
|
||||
-> 逐对象 direct + early reflections
|
||||
-> shared unitary-FDN late room
|
||||
-> stereo float64
|
||||
pcm16:LFE + 15 路对象 PCM
|
||||
→ 64-band QMF analysis
|
||||
→ 77-band hybrid analysis
|
||||
→ 逐对象方向、距离、双耳传递函数和 room send
|
||||
→ 对象累加 + room network
|
||||
→ hybrid synthesis
|
||||
→ QMF synthesis
|
||||
→ 961-sample 延迟补偿
|
||||
→ stereo WAV
|
||||
```
|
||||
|
||||
## 输入接口
|
||||
主要记号:
|
||||
|
||||
CLI 有三个互斥的 HRTF 输入来源;都不指定时按默认规则自动选择:
|
||||
|
||||
```powershell
|
||||
# 1) SOFA:缺省取 HRTF/binaural.sofa,也可显式指定
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone:缺省取 HRTF/binaural.personalized_headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf:显式读取预编译 cache
|
||||
python main.py input.m4a --binaural `
|
||||
--compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# 可选:SOFA 透明生成/复用磁盘 cache
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
```
|
||||
|
||||
默认选择顺序:`HRTF/binaural.sofa` → `output/hrtf-cache` 下唯一的 `.jochrtf` →
|
||||
`HRTF/binaural.personalized_headphone`;三者都没有时报错并提示显式指定。
|
||||
`output/hrtf-cache` 下有多个 `.jochrtf` 时同样报错,要求显式选择。
|
||||
|
||||
`.personalized_headphone` 的 JSON 解析由本项目自行实现(`src/rosella_model.py`),
|
||||
不调用任何杜比软件。
|
||||
|
||||
`--hrtf-cache-policy` 可取 `none`、`memory`、`disk`。默认是 `memory`;`none` 和
|
||||
`memory` 都不会创建磁盘文件。`disk` 默认写入 `output/hrtf-cache`,也可用
|
||||
`--hrtf-cache-dir` 指定。`--hrtf-radius-m` 选择距离目标最近的 measurement shell。
|
||||
|
||||
Python API 使用显式 factory:
|
||||
|
||||
```python
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
|
||||
renderer = SofaBinauralBackend.from_sofa(
|
||||
"subject.sofa",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
cache_policy="memory",
|
||||
)
|
||||
|
||||
cached = SofaBinauralBackend.from_compiled_cache(
|
||||
"subject.jochrtf",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
)
|
||||
```
|
||||
|
||||
文件工厂不会按“未知后缀”猜格式:SOFA 和 `.jochrtf` 始终走不同 loader。
|
||||
|
||||
## 双耳渲染模式
|
||||
|
||||
`--binaural-mode off|near|mid|far`(默认 `mid`)是**人为指定的渲染提示**,不是
|
||||
从输入 E-AC-3 JOC 码流提取或还原的原始双耳元数据:
|
||||
|
||||
- 直接双耳渲染(`--binaural`):near/mid/far 生效,默认 `mid`;`off` 报错;
|
||||
- ADM BWF:DBMD segment 10 中后 15 个 JOC 对象的 binaural render mode 写
|
||||
`off=0/near=1/far=2/mid=3`,前 10 个 bed 保持不变,默认 `mid`;`off` 用于显式
|
||||
关闭双耳元数据提示。
|
||||
|
||||
## Canonical SOFA 契约
|
||||
|
||||
当前 strict importer 接受:
|
||||
|
||||
- `Conventions=SOFA`;
|
||||
- `SOFAConventions=SimpleFreeFieldHRIR`,version `0.4`、`1.0` 或 `1.1`;
|
||||
- `DataType=FIR`,`Data.IR[M,2,N]`;
|
||||
- 单一正有限 `Data.SamplingRate`,单位为 hertz/Hz;
|
||||
- spherical 或 Cartesian `SourcePosition`;
|
||||
- 单值或 per-measurement 的 `ListenerPosition/View/Up`;
|
||||
- 两个能由 listener-local lateral 坐标唯一识别左右的 receiver;
|
||||
- 单一且零偏移的 emitter;
|
||||
- causal `Data.Delay[I,2]` 或 `[M,2]`;
|
||||
- 明确的 free-field/anechoic `RoomType`。
|
||||
|
||||
receiver 左右顺序由几何决定,不能假定 `Data.IR` 的 receiver index。SOFA listener
|
||||
坐标为 $+X$ front、
|
||||
$+Y$ left、
|
||||
$+Z$ up;ADM 坐标为
|
||||
$+X$ right、
|
||||
$+Y$ front、
|
||||
$+Z$ up,转换为:
|
||||
|
||||
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
|
||||
|
||||
`CanonicalHrtf` 将 `Data.IR` 与 `Data.Delay` 分开保存。只有时域 baseline 才调用
|
||||
`materialized_measurement()` 将 delay 应用一次;运行时 SH 路径不先 materialize。
|
||||
非 48 kHz HRIR 使用 float64 `scipy.signal.resample_poly` 规范化,delay samples 按
|
||||
相同比例缩放。
|
||||
|
||||
GeneralFIR、BRIR、TF、多 emitter、多义 receiver 或非 free-field 数据需要单独的
|
||||
convention adapter,不能只通过 reshape 进入核心 importer。
|
||||
|
||||
## Delay/phase:exactly once
|
||||
|
||||
编译器只允许三种互斥语义:
|
||||
|
||||
1. 非零 `Data.Delay` 是 `Data.IR` 外部 delay;FIR 不去旋,运行时应用一次。
|
||||
2. `Data.Delay=0` 且 HRIR 有普通正 onset:以每耳 main peak 分离 arrival,拟合后
|
||||
在运行时恢复一次;当前阈值为 peak index 大于 2 samples。
|
||||
3. `Data.Delay=0` 且双耳 FIR 共享 sample-0 起点:不发明外部 delay,原 complex
|
||||
phase 直接进入五阶场。
|
||||
|
||||
任何路径都不能再叠加第二套 ear delay 或 phase-group delay。
|
||||
|
||||
## 公开 filterbank 与方向场
|
||||
|
||||
运行时固定为:
|
||||
|
||||
- 48 kHz;
|
||||
- 64-sample QMF hop;
|
||||
- 64-QMF / 77-hybrid;
|
||||
- analysis/synthesis latency 961 samples;
|
||||
- 五阶、36 项、ACN/N3D real spherical harmonics;
|
||||
- PCM、delay、SH、room state 为 float64;频带传递和频域状态为 complex128。
|
||||
|
||||
每个 hybrid band 的 real/imaginary 单位增益都通过同一套 analysis/synthesis 链生成
|
||||
脉冲字典,共 154 个实参数;编译不是直接读取 77 个 FFT bin。默认 projection
|
||||
ridge 为 `1e-3`,SH ridge 为 `1e-5`。同方向 measurement 先合并,再用球面 Voronoi
|
||||
面积权重做 ridge fit。
|
||||
|
||||
固定表位于 `data/rosella_kernels.npz`,实现公开标准化的滤波器组,各表可由如下
|
||||
公式计算。
|
||||
|
||||
hybrid 分析核定义于 [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
第 5.2.2 节(Table 1 的 $Q=8$/
|
||||
$Q=4$ 系数,delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
QMF analysis 表即 MPEG-4 AAC/SBR(ISO/IEC 14496-3/AMD1:2003 第 4.B.18.2 节)
|
||||
的 64 complex QMF bank;打包的 $64\times10$ 表是公开 640-tap prototype
|
||||
$c_0,\dots,c_{639}$ 的多相重排:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
QMF synthesis 表为上述 analysis 多相矩阵 $\mathbf{A}$ 的因果左逆,即求解
|
||||
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$(
|
||||
$\mathbf{P}$ 为 577-sample 延迟置换;
|
||||
全链 $961 = 577 + 6\times64$),以 rank-4 分解形式存储:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
hybrid synthesis 表为 77→64 重组:高频带恒等 $Y_{3+b}=X_{16+b}$;低频带(
|
||||
$C_p$ 为
|
||||
$8+4+4$ 子带划分):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\mathrm{Re}X_q + j\,s_q\,\mathrm{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
loader 校验 archive 和每个数组的 SHA-256;table version、所有数组 hash 与
|
||||
77 个 band-center 参考值都属于 cache key。标准可公开获取不等于获准实施相关
|
||||
专利;更多来源信息见 [`data/README.md`](../data/README.md) 与
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
|
||||
|
||||
## `.jochrtf`
|
||||
|
||||
`.jochrtf` 是无 pickle 的压缩 NumPy archive,固定包含:
|
||||
|
||||
| key | dtype / shape |
|
||||
| 符号 | 含义 |
|
||||
|---|---|
|
||||
| `metadata_json` | 含 JSON 文本的 NumPy Unicode scalar(`dtype.kind == "U"`) |
|
||||
| `band_center_frequencies_hz` | little-endian `float64[77]` |
|
||||
| `coefficients` | little-endian `complex128[36,2,77]` |
|
||||
| `delay_coefficients` | little-endian `float64[36,2]` |
|
||||
| `delay_bounds` | little-endian `float64[2,2]` |
|
||||
| $s=0\ldots15$ | 输入源;$s=0$ 为 LFE,$s=1\ldots15$ 为对象 |
|
||||
| $e\in\{L,R\}$ | 左右输出耳 |
|
||||
| $p=0\ldots63$ | QMF phase / 时域 hop 内采样 |
|
||||
| $k=0\ldots63$ | QMF 子带 |
|
||||
| $h=0\ldots76$ | hybrid 子带 |
|
||||
| $j=0\ldots35$ | 方向 basis 项 |
|
||||
| $m$ | 64-sample QMF 时槽 |
|
||||
| $n$ | 时域采样位置 |
|
||||
|
||||
metadata magic 固定为 `JOC-HRTF-CACHE`,并记录 schema/compiler/phase-policy、
|
||||
ACN/N3D、filterbank table hashes、SOFA content SHA-256、采样率、radius、order、
|
||||
两个 ridge、payload hash 和 fit report。cache key 覆盖所有会改变编译结果的字段。
|
||||
metadata 不保存本机绝对 `source_path`,仅可保存 source display name。
|
||||
每个 QMF hop 为 64 samples,每个双耳控制块为
|
||||
|
||||
loader 使用 `allow_pickle=False`,并在构造对象前检查 ZIP 成员集、解压大小、shape、
|
||||
dtype、端序、连续布局、有限值、delay bounds、band centers、payload hash 和 cache
|
||||
key。writer 使用同目录临时文件、`fsync`、进程持有的 OS 文件锁和原子
|
||||
`os.replace`;对应的隐藏 `.lock` sidecar 可保留,但不代表仍有 writer 持锁。
|
||||
旧版本、损坏或配置不匹配的 cache 不能命中;从 SOFA 启动时会重建,显式 cache
|
||||
入口则直接报错。
|
||||
$$
|
||||
N_b=512=8\times64,
|
||||
$$
|
||||
|
||||
删除磁盘 cache 后,从同一 SOFA 和同一编译配置得到的场与渲染结果不得改变。
|
||||
每个 E-AC-3/JOC 音频帧为
|
||||
|
||||
`.jochrtf` 包含由源 HRIR 变换得到的方向场系数与 delay 数据,因此“可以重建”不表示
|
||||
它不受数据许可约束。生成 cache 不会扩大源 SOFA/HRTF 数据集授予的权利;cache 的
|
||||
使用、复制和再分发仍须遵守源数据集条款。不能确认条款时,应把 `.jochrtf` 作为本地
|
||||
私有 cache,不随程序或构建产物发布。`source_sha256` 只用于内容一致性校验,不是许可
|
||||
或来源证明。
|
||||
$$
|
||||
N_f=1536=3N_b.
|
||||
$$
|
||||
|
||||
## JOC 对象与房间
|
||||
## 2. 输入与控制块
|
||||
|
||||
生产适配器继续使用现有 JOC 调度:
|
||||
输入矩阵为
|
||||
|
||||
- 每帧输入 `[1536,16]`;
|
||||
- channel 0 是 special LFE,channel 1..15 是 JOC objects;
|
||||
- ID11/OAMD position 使用 sample-timed timeline;
|
||||
- 每 512 samples 更新方向/profile;
|
||||
- 每个对象拥有独立 direct/early history,late FDN 全局共享;
|
||||
- `finish()` 排空 early/late tail;输出增益显式应用,不隐含 limiter 或节目响度归一化。
|
||||
$$
|
||||
x_s[n],\qquad s=0\ldots15.
|
||||
$$
|
||||
|
||||
Near/Mid/Far、equal-power direct level、六面 shoebox 一阶 image source、late send、
|
||||
unitary FDN、LFE 120–180 Hz cosine-squared 低通及 room calibration 都是 JOC
|
||||
项目定义行为,不是 SOFA 或 Dolby 公布常数。
|
||||
`pcm16` 的通道约定为:
|
||||
|
||||
公开 SOFA 双耳渲染在 `--backend auto/native` 下默认走 C++20 原生核
|
||||
(`lib/eac3joc_core.dll` 的 `ejoc_sofa_binaural_*` 接口:filterbank、SH 方向场求值、
|
||||
逐对象 early/direct 历史与共享 FDN 全部在原生侧执行,Python 只做 SOFA 编译与每
|
||||
512-sample 的元数据更新);原生库不可用时自动回退 Python/NumPy 参考实现,两者
|
||||
逐值一致(差异 < 1e-9)。`--backend python` 强制使用 Python 后端。
|
||||
```text
|
||||
ch0 special LFE
|
||||
ch1..15 JOC 对象 1..15
|
||||
```
|
||||
|
||||
## 技术引用与权利边界
|
||||
渲染器按连续采样流推进。QMF、hybrid、room 和 synthesis 状态不会在 1536-sample 帧边界清零。
|
||||
|
||||
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
|
||||
- [3GPP TS 26.405 / ETSI TS 126 405(64-QMF/77-hybrid 定义)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
- [Dolby binaural render mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
|
||||
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en),仅作 direct/early/late、
|
||||
subband 与 FDN 架构背景,不证明某个产品使用特定实施例。
|
||||
## 3. 64-band QMF analysis
|
||||
|
||||
规范、源码或专利文献可公开获取,不等于获准复制其内容、再分发派生产物或实施其中的
|
||||
专利权利要求。本项目的技术引用本身不授予专利许可,也不作不侵权保证;准备发布或集成
|
||||
到产品的一方应自行审查适用的数据许可、软件许可、专利许可及 freedom-to-operate。
|
||||
公开标准来源与权利边界见
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
|
||||
令 $a_{p,\ell}$ 为固定的 64×10 polyphase 系数,$r_{s,\ell,p}[m]$ 为当前和前 9 个 hop 的 phase 历史。奇偶 lag 分别累加:
|
||||
|
||||
$$
|
||||
E_{s,p}[m]
|
||||
=\sum_{\substack{\ell=0\\\ell\text{ even}}}^{9}
|
||||
a_{p,\ell}r_{s,\ell,p}[m],
|
||||
$$
|
||||
|
||||
$$
|
||||
O_{s,p}[m]
|
||||
=\sum_{\substack{\ell=0\\\ell\text{ odd}}}^{9}
|
||||
a_{p,\ell}r_{s,\ell,p}[m].
|
||||
$$
|
||||
|
||||
对任一 64-vector $v[p]$,定义调制变换
|
||||
|
||||
$$
|
||||
\mathcal Q(v)_k
|
||||
=
|
||||
\operatorname{FFT}_{128}
|
||||
\left(
|
||||
\left[v[p]e^{-j\pi p/128}\right]_{p=0}^{63},
|
||||
0_{64}
|
||||
\right)_k
|
||||
|
||||
e^{-j3\pi(k+1/2)/128}.
|
||||
$$
|
||||
|
||||
analysis 输出为
|
||||
|
||||
$$
|
||||
X_{s,k}[m]
|
||||
=
|
||||
\mathcal Q(O_s)_k
|
||||
+j(-1)^k\mathcal Q(E_s)_k.
|
||||
$$
|
||||
|
||||
所有历史、乘加和 FFT 结果使用 `float64/complex128`。
|
||||
|
||||
## 4. 77-band hybrid analysis
|
||||
|
||||
低 3 个 QMF 子带使用 13-slot FIR 拆分为 16 个 hybrid bands。把复数的实部和虚部分量记为 $i,o\in\{0,1\}$,固定核为 $K_{p,i,\ell,h,o}$:
|
||||
|
||||
$$
|
||||
H_{s,h,o}[m]
|
||||
=
|
||||
\sum_{p=0}^{2}
|
||||
\sum_{i=0}^{1}
|
||||
\sum_{\ell=0}^{12}
|
||||
X_{s,p,i}[m-\ell]K_{p,i,\ell,h,o},
|
||||
\qquad h=0\ldots15.
|
||||
$$
|
||||
|
||||
其余 61 个 hybrid bands 是 QMF 3..63 的 6-slot 延迟:
|
||||
|
||||
$$
|
||||
H_{s,16+q}[m]=X_{s,3+q}[m-6],
|
||||
\qquad q=0\ldots60.
|
||||
$$
|
||||
|
||||
因此 hybrid vector 的顺序为:
|
||||
|
||||
```text
|
||||
0..15 低 3 个 QMF bands 的细分
|
||||
16..76 延迟后的 QMF bands 3..63
|
||||
```
|
||||
|
||||
## 5. OAMD 坐标与时间轴
|
||||
|
||||
### 5.1 Q15 坐标到 Cartesian
|
||||
|
||||
对象状态中的 $q_1,q_2,q_3$ 先恢复到离散位置网格:
|
||||
|
||||
$$
|
||||
u_1=\min\left(1,\frac{\operatorname{round}(62q_1/32767)}{62}\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
u_2=\min\left(1,\frac{\operatorname{round}(62q_2/32767)}{62}\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
u_3=\operatorname{clip}\left(
|
||||
\frac{\operatorname{round}(15q_3/32767)}{15},-1,1\right).
|
||||
$$
|
||||
|
||||
ADM Cartesian 坐标为
|
||||
|
||||
$$
|
||||
(X,Y,Z)=(2u_1-1,\ 1-2u_2,\ u_3).
|
||||
$$
|
||||
|
||||
### 5.2 更新时间
|
||||
|
||||
一条位置更新的编码时刻为
|
||||
|
||||
$$
|
||||
n_{\mathrm{coded}}
|
||||
=n_{\mathrm{frame}}
|
||||
+n_{\mathrm{outer}}
|
||||
+n_{\mathrm{block}}.
|
||||
$$
|
||||
|
||||
首个有效状态作为 sample 0 的初始位置。后续更新加入对象 PCM 延迟 $D_o$,默认
|
||||
|
||||
$$
|
||||
D_o=1473.
|
||||
$$
|
||||
|
||||
若 ramp duration 为 $R>64$,连续运动为
|
||||
|
||||
$$
|
||||
n_{\mathrm{start}}=n_{\mathrm{coded}}+D_o+64,
|
||||
$$
|
||||
|
||||
$$
|
||||
R_{\mathrm{eff}}=R-64,
|
||||
$$
|
||||
|
||||
$$
|
||||
\mathbf p[n]
|
||||
=(1-\alpha)\mathbf p_0+\alpha\mathbf p_1,
|
||||
\qquad
|
||||
\alpha=\frac{n-n_{\mathrm{start}}}{R_{\mathrm{eff}}}.
|
||||
$$
|
||||
|
||||
当 $R\le64$ 时,目标位置在 $n_{\mathrm{coded}}+D_o$ 直接生效。
|
||||
|
||||
Rosella 参数在每个 512-sample block 起点求值,并用于该块的 8 个 hybrid slots。
|
||||
|
||||
## 6. 距离 profile 与方向
|
||||
|
||||
普通对象只使用 Near、Mid、Far 三个 profile。每个 profile 包含:
|
||||
|
||||
- 三轴负/正边界 $b_{x-},b_{x+},b_{y-},b_{y+},b_{z-},b_{z+}$;
|
||||
- 距离尺度 $D$ 和倒数尺度 $D^{-1}$;
|
||||
- 三轴内部尺度 $a_x,a_y,a_z$;
|
||||
- 最小归一化半径 $\rho_{\min}$。
|
||||
|
||||
Cartesian 坐标经过 Q15 metadata grid 后换成内部前、侧、上轴,乘以 profile 尺度:
|
||||
|
||||
$$
|
||||
\mathbf s=(a_zq_f,\ a_xq_l,\ a_yq_v).
|
||||
$$
|
||||
|
||||
若射线超出 profile 边界,则用单一比例 $\lambda\le1$ 缩放:
|
||||
|
||||
$$
|
||||
\mathbf s' = \lambda\mathbf s.
|
||||
$$
|
||||
|
||||
随后
|
||||
|
||||
$$
|
||||
\rho=\|\mathbf s'\|_2,
|
||||
\qquad
|
||||
\rho_c=\max(\rho,\rho_{\min}),
|
||||
\qquad
|
||||
\alpha=\frac{\rho}{\rho_c},
|
||||
$$
|
||||
|
||||
$$
|
||||
\mathbf d=
|
||||
\begin{cases}
|
||||
\mathbf s'/\rho,&\rho>0,\\
|
||||
(1,0,0),&\rho=0,
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
物理半径为
|
||||
|
||||
$$
|
||||
R=D\rho.
|
||||
$$
|
||||
|
||||
## 7. 36 项方向 basis
|
||||
|
||||
方向 $\mathbf d=(x,y,z)$ 被展开为 36 项实值多项式:
|
||||
|
||||
$$
|
||||
\mathbf b(\mathbf d)=
|
||||
[1,x,y,z,x^2-\tfrac13,xy,xz,y^2-\tfrac13,yz,\ldots]^T.
|
||||
$$
|
||||
|
||||
完整顺序由 `rosella_model.direction_basis()` 固定。最高次数为 5;所有 field 系数必须按该顺序点积,不能交换 basis 项。
|
||||
|
||||
逐耳 basis 会根据耳偏移重新归一化。令耳偏移标量为 $e$:
|
||||
|
||||
$$
|
||||
\epsilon=\frac{eD^{-1}}{\rho_c},
|
||||
$$
|
||||
|
||||
$$
|
||||
\mathbf d_-=
|
||||
\frac{(x,y-\epsilon,z)}{\|(x,y-\epsilon,z)\|_2},
|
||||
\qquad
|
||||
\mathbf d_+=
|
||||
\frac{(x,y+\epsilon,z)}{\|(x,y+\epsilon,z)\|_2}.
|
||||
$$
|
||||
|
||||
## 8. 逐耳路径与 ITD
|
||||
|
||||
归一化路径长度为
|
||||
|
||||
$$
|
||||
\ell_-=\rho_c\sqrt{x^2+(y-\epsilon)^2+z^2},
|
||||
$$
|
||||
|
||||
$$
|
||||
\ell_+=\rho_c\sqrt{x^2+(y+\epsilon)^2+z^2}.
|
||||
$$
|
||||
|
||||
模型允许通过 36-vector 对路径加入非负方向修正:
|
||||
|
||||
$$
|
||||
\ell'_e
|
||||
=
|
||||
\ell_e
|
||||
+
|
||||
\max(\mathbf v_e^T\mathbf b_e,0)\,2cD^{-1}.
|
||||
$$
|
||||
|
||||
耳间延迟为
|
||||
|
||||
$$
|
||||
\tau
|
||||
=|\ell'_+-\ell'_-|\,D\frac{f_s}{343.3}\alpha,
|
||||
\qquad f_s=48000.
|
||||
$$
|
||||
|
||||
路径较长的一耳应用 hybrid-band 相位:
|
||||
|
||||
$$
|
||||
P_h=e^{j\omega_h\tau},
|
||||
$$
|
||||
|
||||
其中 $\omega_h$ 由模型的 20 个 hybrid group 参数递推到 77 个 bands。
|
||||
|
||||
## 9. 方向 field 与直达增益
|
||||
|
||||
左右耳各有一个 77×36 complex field:
|
||||
|
||||
$$
|
||||
C_{e,h}(\mathbf d_e)
|
||||
=
|
||||
\sum_{j=0}^{35}F_{e,h,j}b_j(\mathbf d_e).
|
||||
$$
|
||||
|
||||
另一次耳路径计算给出左右权重:
|
||||
|
||||
$$
|
||||
w_L=\frac{\ell_+}{\sqrt{\ell_-^2+\ell_+^2}},
|
||||
\qquad
|
||||
w_R=\frac{\ell_-}{\sqrt{\ell_-^2+\ell_+^2}}.
|
||||
$$
|
||||
|
||||
有效距离为
|
||||
|
||||
$$
|
||||
R_e=\rho\,s_dD,
|
||||
$$
|
||||
|
||||
其中 $s_d$ 为模型距离标量。Mid/Far 的公共衰减和 room send 为
|
||||
|
||||
$$
|
||||
g_c=\frac{1}{\sqrt{1+s_rR_e^2}},
|
||||
$$
|
||||
|
||||
$$
|
||||
g_{\mathrm{room}}=R_eg_c.
|
||||
$$
|
||||
|
||||
Near 使用
|
||||
|
||||
$$
|
||||
g_c=1,
|
||||
\qquad
|
||||
g_{\mathrm{room}}=0.
|
||||
$$
|
||||
|
||||
令 $C_{e,h}^{(0)}$ 为 field 的第 0 个 basis 系数,中心保护项为 $1-\alpha$。普通直达传递函数可写成
|
||||
|
||||
$$
|
||||
G_{L,h}
|
||||
=g_c\left[C_{L,h}w_L\alpha+C_{L,h}^{(0)}c_L(1-\alpha)\right],
|
||||
$$
|
||||
|
||||
$$
|
||||
G_{R,h}
|
||||
=g_c\left[C_{R,h}w_R\alpha+C_{R,h}^{(0)}c_R(1-\alpha)\right].
|
||||
$$
|
||||
|
||||
$c_L,c_R$ 由模型的耳权重配置选择;路径较长的一耳再乘 $P_h$。
|
||||
|
||||
## 10. LFE 传递函数
|
||||
|
||||
LFE 不进入普通对象方向计算。其传递函数为
|
||||
|
||||
$$
|
||||
G_{L,h}^{\mathrm{LFE}}=G_{R,h}^{\mathrm{LFE}}=
|
||||
\begin{cases}
|
||||
g_h,&0\le h<16,\\
|
||||
0,&16\le h<77.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
前 16 个固定系数为
|
||||
|
||||
```text
|
||||
2.60290003, 1.80741799, 0.659342408, -0.0275855921,
|
||||
-0.105803289, -0.0699509233, 0.0749056414, -0.00919809937,
|
||||
0.00349014648,-0.0158600751,-0.000723021978,0.00188189559,
|
||||
-0.000421735429,0.0000329252762,0.0000317397971,0.000000580376991
|
||||
```
|
||||
|
||||
LFE 的 room send 恒为 0。
|
||||
|
||||
## 11. 对象累加与 room input
|
||||
|
||||
每个 hybrid slot 的直达输出为
|
||||
|
||||
$$
|
||||
Y^{\mathrm{direct}}_{e,h}
|
||||
=
|
||||
\sum_{s=0}^{15}H_{s,h}G_{s,e,h}.
|
||||
$$
|
||||
|
||||
room 输入为
|
||||
|
||||
$$
|
||||
U_h
|
||||
=
|
||||
\sum_{s=1}^{15}H_{s,h}g_{\mathrm{room},s}.
|
||||
$$
|
||||
|
||||
LFE 不进入该和式。
|
||||
|
||||
## 12. Room network
|
||||
|
||||
room 只处理前 64 个 hybrid bands。输入先乘
|
||||
|
||||
$$
|
||||
g_0=0.70710677.
|
||||
$$
|
||||
|
||||
对每级 all-pass,设延迟样本为 $d[n]$、系数为 $a$:
|
||||
|
||||
$$
|
||||
r[n]=x[n]-ad[n],
|
||||
$$
|
||||
|
||||
$$
|
||||
y[n]=ar[n]+d[n].
|
||||
$$
|
||||
|
||||
all-pass 输出复制到 4 个 FDN branches。设延迟输出为 $\mathbf d_h[m]$、4×4 混合矩阵为 $M$:
|
||||
|
||||
$$
|
||||
\mathbf b_h[m]
|
||||
=U_h[m]\mathbf 1+M\mathbf d_h[m].
|
||||
$$
|
||||
|
||||
每个 branch 使用复反馈系数 $f_{h,i}$:
|
||||
|
||||
$$
|
||||
m_{h,i}[m]=f_{h,i}b_{h,i}[m].
|
||||
$$
|
||||
|
||||
主 tap、可选额外 tap 和左右输出矩阵合成为
|
||||
|
||||
$$
|
||||
Y^{\mathrm{room}}_{e,h}[m]
|
||||
=
|
||||
\sum_{i=0}^{3}O_{e,h,i}z_{h,i}[m].
|
||||
$$
|
||||
|
||||
最终 hybrid 输出为
|
||||
|
||||
$$
|
||||
Y_{e,h}=Y^{\mathrm{direct}}_{e,h}+Y^{\mathrm{room}}_{e,h}.
|
||||
$$
|
||||
|
||||
Python 后端把该递归网络展开为有限 complex FIR 并使用 overlap-add;C++ 后端直接保持递归状态。两者均跨帧连续。
|
||||
|
||||
## 13. Hybrid synthesis
|
||||
|
||||
hybrid synthesis 是 154 项稀疏即时映射。令映射项为 $(h,i,k,o,w)$,其中 $i,o$ 表示实部或虚部,则
|
||||
|
||||
$$
|
||||
Q_{e,k,o}[m]
|
||||
\mathrel{+}=
|
||||
Y_{e,h,i}[m]w.
|
||||
$$
|
||||
|
||||
输出是每耳 64 个 complex QMF bands。
|
||||
|
||||
## 14. QMF synthesis
|
||||
|
||||
每耳 QMF vector 先展开为 128 项实向量
|
||||
|
||||
$$
|
||||
\mathbf q_e=[\Re Q_{e,0},\Im Q_{e,0},\ldots,\Re Q_{e,63},\Im Q_{e,63}]^T.
|
||||
$$
|
||||
|
||||
对每个 phase $p$ 和 rank $r=0\ldots3$:
|
||||
|
||||
$$
|
||||
f_{e,p,r}[m]
|
||||
=\mathbf b_{p,r}^T\mathbf q_e[m].
|
||||
$$
|
||||
|
||||
使用 10-slot taps 合成时域样本:
|
||||
|
||||
$$
|
||||
y_e[64m+p]
|
||||
=
|
||||
\sum_{\ell=0}^{9}
|
||||
\sum_{r=0}^{3}
|
||||
t_{p,\ell,r}f_{e,p,r}[m-\ell].
|
||||
$$
|
||||
|
||||
## 15. 延迟、尾声和输出
|
||||
|
||||
完整 filterbank 的固定延迟为
|
||||
|
||||
$$
|
||||
L=961\ \text{samples}.
|
||||
$$
|
||||
|
||||
只在连续流起点丢弃一次前 $L$ 个输出 samples。输入结束后继续送零,以释放 QMF、hybrid 和 room 状态。尾声裁切只作用于文件末端:
|
||||
|
||||
$$
|
||||
\max(|y_L[n]|,|y_R[n]|) > 10^{-8}
|
||||
$$
|
||||
|
||||
的最后一个 sample 被保留,同时输出长度不得短于源 PCM 长度。
|
||||
|
||||
所有内部状态、参数计算和对象累加使用 `float64/complex128`。最终 writer 才转换为 float32 或 PCM24。
|
||||
|
||||
## 16. Python 与 C++ 后端
|
||||
|
||||
两套后端共享:
|
||||
|
||||
- 同一份 QMF/hybrid 固定表;
|
||||
- 同一份模型解析结果;
|
||||
- 同一套 512-sample 参数更新时间轴;
|
||||
- 同一组逐对象 complex gains 和 room sends;
|
||||
- 同一 961-sample 延迟补偿与尾声策略。
|
||||
|
||||
C++ 后端以 512-sample block 为处理单位,内部持有 QMF、hybrid、room 和 synthesis 状态。Python 只负责模型解析、OAMD 时间轴和每块参数更新。
|
||||
|
||||
## 17. 模型文件
|
||||
|
||||
默认路径为
|
||||
|
||||
```text
|
||||
HRTF/binaural.personalized_headphone
|
||||
```
|
||||
|
||||
也可通过
|
||||
|
||||
```text
|
||||
--personalized-headphone PATH
|
||||
```
|
||||
|
||||
指定其它文件。
|
||||
|
||||
`.personalized_headphone` 中的 int32/Q15 参数在解析后提升为 float64。任意 SOFA FIR 不能只通过数组重排变成该参数模型;若要转换,需要拟合方向 fields、ITD、距离 profile、耳几何和 room 参数。
|
||||
|
||||
## 18. 适用范围
|
||||
|
||||
当前路径处理 15 个普通点对象和 1 路 special LFE。对象 extent、spread、diffuse、divergence、channel lock,以及未实现的 OAMD element 变体不在本公式范围内。
|
||||
|
||||
+78
-97
@@ -4,7 +4,7 @@
|
||||
|
||||
This document covers only the signal model and formulas used in the JustOneCacophony research path: how JOC parameters combine with core PCM to reconstruct object signals, and how OAMD coordinates become speaker gains.
|
||||
|
||||
The formulas describe the JOC matrix parameters (both the dense and the sparse differential syntax) and the ordinary point-object paths studied by the project. They are not a complete definition of every E-AC-3 JOC variant.
|
||||
The formulas describe the dense-JOC and ordinary point-object paths studied by the project. They are not a complete definition of every E-AC-3 JOC variant.
|
||||
|
||||
## 1. Overall path and notation
|
||||
|
||||
@@ -52,11 +52,9 @@ $$
|
||||
N_f=1536=24\times64.
|
||||
$$
|
||||
|
||||
## 2. JOC matrix parameters
|
||||
## 2. Dense-JOC matrix parameters
|
||||
|
||||
For every object and data point, the quantized matrix `joc_mix_mtx_q` is defined on $N_q$ quantization levels. The `b_joc_sparse` flag selects one of two differential syntaxes: dense sends one MTX difference per core channel, while sparse sends one active channel plus one coefficient difference per parameter band.
|
||||
|
||||
### 2.1 Dense differential reconstruction
|
||||
### 2.1 Differential reconstruction
|
||||
|
||||
Let `quant_idx` be $q_i\in\{0,1\}$. The number of quantization levels is
|
||||
|
||||
@@ -77,76 +75,38 @@ $$
|
||||
For object $o$, data point $d$, core channel $c$, and parameter band $p$, the coded difference $\Delta_{o,d,c,p}$ reconstructs to
|
||||
|
||||
$$
|
||||
Q_{o,d,c,0}=
|
||||
Q_{o,d,c,0}
|
||||
=
|
||||
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}=
|
||||
Q_{o,d,c,p}
|
||||
=
|
||||
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
|
||||
\qquad p>0.
|
||||
$$
|
||||
|
||||
### 2.2 Sparse differential reconstruction
|
||||
|
||||
Let $I_{o,d,p}$ be the `joc_channel_idx` symbol (IDX), $V_{o,d,p}$ the `joc_vec` symbol (VEC), and $N_c\in\lbrace5,7\rbrace$ the number of core channels. Each parameter band has exactly one active channel:
|
||||
|
||||
$$
|
||||
A_{o,d,p}=
|
||||
\begin{cases}
|
||||
I_{o,d,0}, & p=0,\\
|
||||
\left(A_{o,d,p-1}+I_{o,d,p}\right)\bmod N_c, & p>0,
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
where $I_{o,d,0}$ is a 3-bit absolute channel index and every later IDX symbol is an increment relative to the previous **active channel**. The coefficient is a single accumulator running across parameter bands:
|
||||
|
||||
$$
|
||||
\kappa_{o,d,-1}=O^{(s)}_q,\qquad
|
||||
\kappa_{o,d,p}=
|
||||
\left(\kappa_{o,d,p-1}+V_{o,d,p}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
with a sparse starting point two quantization levels above the dense center offset:
|
||||
|
||||
$$
|
||||
O^{(s)}_q=
|
||||
\begin{cases}
|
||||
50, & q_i=0,\\
|
||||
100, & q_i=1.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
The accumulator is **not** reset when the active channel changes. The complete matrix is
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}=
|
||||
\begin{cases}
|
||||
\kappa_{o,d,p}, & c=A_{o,d,p},\\
|
||||
\dfrac{N_q}{2}, & c\neq A_{o,d,p}.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
Non-active entries take $N_q/2$, which dequantizes to exactly 0.
|
||||
|
||||
### 2.3 Dequantization
|
||||
### 2.2 Dequantization
|
||||
|
||||
The dequantized matrix coefficient is
|
||||
|
||||
$$
|
||||
D_{o,d,c,p}=
|
||||
D_{o,d,c,p}
|
||||
=
|
||||
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
|
||||
\frac{820}{4096(1+q_i)}.
|
||||
$$
|
||||
|
||||
The effective denominator is therefore 4096 in coarse mode and 8192 in fine mode.
|
||||
|
||||
### 2.4 JOC clipgain
|
||||
### 2.3 JOC clipgain
|
||||
|
||||
If the clipgain field consists of integer $x$ and mantissa $y$, then
|
||||
|
||||
$$
|
||||
G_{\mathrm{clip}}=
|
||||
G_{\mathrm{clip}}
|
||||
=
|
||||
1+\frac{y}{32}2^{x-4}.
|
||||
$$
|
||||
|
||||
@@ -192,7 +152,8 @@ $$
|
||||
$$
|
||||
|
||||
$$
|
||||
M_{o,c,b,t}=
|
||||
M_{o,c,b,t}
|
||||
=
|
||||
(1-\alpha_t)P_{o,c,b}
|
||||
+\alpha_tD_{o,c,p(b)}.
|
||||
$$
|
||||
@@ -220,8 +181,9 @@ $$
|
||||
Let $\mathcal A_b$ denote the 64-band analysis-QMF operator with polyphase history state. Then
|
||||
|
||||
$$
|
||||
X_{c,b,t}=
|
||||
\mathcal A_b\left(
|
||||
X_{c,b,t}
|
||||
=
|
||||
\mathcal A_b\!\left(
|
||||
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
|
||||
\mathbf s^{\mathrm A}_{c,t}
|
||||
\right).
|
||||
@@ -248,7 +210,8 @@ $$
|
||||
Band 0 of each surround channel additionally passes through a 21-tap complex FIR:
|
||||
|
||||
$$
|
||||
\widehat X_{c,0,t}=
|
||||
\widehat X_{c,0,t}
|
||||
=
|
||||
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
|
||||
$$
|
||||
|
||||
@@ -259,7 +222,8 @@ These delays and filter histories are decoder state and cannot be reset independ
|
||||
For each object $o$, subband $b$, and slot $t$, the object's frequency-domain value is a linear combination of the five core channels:
|
||||
|
||||
$$
|
||||
Z_{o,b,t}=
|
||||
Z_{o,b,t}
|
||||
=
|
||||
\sum_{c=0}^{4}
|
||||
M_{o,c,b,t}\widehat X_{c,b,t}.
|
||||
$$
|
||||
@@ -274,20 +238,21 @@ Write the 64 complex subbands as 128 interleaved real values in `src`. For $k=0\
|
||||
|
||||
$$
|
||||
\begin{aligned}
|
||||
\mathrm{zone}[2k] &= \mathrm{src}[4k],\\
|
||||
\mathrm{zone}[2k+1] &= -\mathrm{src}[4k+1],\\
|
||||
\mathrm{zone}[126-2k] &= \mathrm{src}[4k+2],\\
|
||||
\mathrm{zone}[127-2k] &= \mathrm{src}[4k+3].
|
||||
\operatorname{zone}[2k] &= \operatorname{src}[4k],\\
|
||||
\operatorname{zone}[2k+1] &= -\operatorname{src}[4k+1],\\
|
||||
\operatorname{zone}[126-2k] &= \operatorname{src}[4k+2],\\
|
||||
\operatorname{zone}[127-2k] &= \operatorname{src}[4k+3].
|
||||
\end{aligned}
|
||||
$$
|
||||
|
||||
Treat `zone` as 64 complex values and apply an unnormalized 64-point FFT:
|
||||
|
||||
$$
|
||||
F_k=
|
||||
F_k
|
||||
=
|
||||
\sum_{n=0}^{63}
|
||||
\mathrm{zone}_n
|
||||
\exp\left(-j\frac{2\pi kn}{64}\right).
|
||||
\operatorname{zone}_n
|
||||
\exp\!\left(-j\frac{2\pi kn}{64}\right).
|
||||
$$
|
||||
|
||||
### 7.2 Modulation and synthesis
|
||||
@@ -295,7 +260,8 @@ $$
|
||||
Define the rotation coefficient
|
||||
|
||||
$$
|
||||
r_k=
|
||||
r_k
|
||||
=
|
||||
\frac12\left(
|
||||
\sin\frac{\pi k}{128}
|
||||
+j\cos\frac{\pi k}{128}
|
||||
@@ -311,8 +277,9 @@ $$
|
||||
Let $\mathcal S$ denote polyphase synthesis with a 640-value synthesis window and cross-slot state:
|
||||
|
||||
$$
|
||||
\mathbf y_{o,t}=
|
||||
\mathcal S\left(
|
||||
\mathbf y_{o,t}
|
||||
=
|
||||
\mathcal S\!\left(
|
||||
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
|
||||
\right).
|
||||
$$
|
||||
@@ -320,8 +287,9 @@ $$
|
||||
Object output is
|
||||
|
||||
$$
|
||||
y_o[64t+r]=
|
||||
\mathrm{clip}\left(
|
||||
y_o[64t+r]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
16\,\mathbf y_{o,t}[r],-1,1
|
||||
\right)G_{\mathrm{clip}},
|
||||
$$
|
||||
@@ -333,8 +301,9 @@ where $r=0\ldots63$. Synthesis state must advance continuously by slot.
|
||||
LFE bypasses the object matrix and inverse QMF and uses a 1217-sample delay. After the input and output scale factors cancel:
|
||||
|
||||
$$
|
||||
y_{\mathrm{LFE}}[n]=
|
||||
\mathrm{clip}\left(
|
||||
y_{\mathrm{LFE}}[n]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
x_{\mathrm{LFE,core}}[n-1217],-1,1
|
||||
\right).
|
||||
$$
|
||||
@@ -344,8 +313,9 @@ $$
|
||||
The lateral and longitudinal grids use $N=62$; the height grid uses $N=15$. The quantizer is
|
||||
|
||||
$$
|
||||
q_N(k)=
|
||||
\min\left(
|
||||
q_N(k)
|
||||
=
|
||||
\min\!\left(
|
||||
32767,
|
||||
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
|
||||
\right).
|
||||
@@ -366,11 +336,11 @@ Their maximum runtime value is $32767/32768$, not exactly 1.
|
||||
For conversion to the ADM grid:
|
||||
|
||||
$$
|
||||
k_1=\mathrm{round}\left(\frac{62q_1}{32767}\right),
|
||||
k_1=\operatorname{round}\!\left(\frac{62q_1}{32767}\right),
|
||||
\quad
|
||||
k_2=\mathrm{round}\left(\frac{62q_2}{32767}\right),
|
||||
k_2=\operatorname{round}\!\left(\frac{62q_2}{32767}\right),
|
||||
\quad
|
||||
k_3=\mathrm{round}\left(\frac{15q_3}{32767}\right),
|
||||
k_3=\operatorname{round}\!\left(\frac{15q_3}{32767}\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
@@ -434,15 +404,17 @@ $$
|
||||
The two-dimensional point gain is
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{2D}}(u,v)=
|
||||
\mathbf G_{\mathrm{2D}}(u,v)
|
||||
=
|
||||
\mathbf h(u)\odot\mathbf v(v).
|
||||
$$
|
||||
|
||||
For 5.1-family layouts with one horizontal surround pair rather than separate side and rear pairs, the longitudinal coordinate is
|
||||
|
||||
$$
|
||||
v_{\mathrm{floor}}=
|
||||
\mathrm{clamp}(2v,0,1).
|
||||
v_{\mathrm{floor}}
|
||||
=
|
||||
\operatorname{clamp}(2v,0,1).
|
||||
$$
|
||||
|
||||
Other layouts use $v_{\mathrm{floor}}=v$.
|
||||
@@ -452,7 +424,8 @@ Other layouts use $v_{\mathrm{floor}}=v$.
|
||||
Three-dimensional layouts compute floor gain $\mathbf G_f$ and height gain $\mathbf G_h$ separately:
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)=
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)
|
||||
=
|
||||
\cos\left(\frac\pi2w\right)\mathbf G_f
|
||||
+
|
||||
\sin\left(\frac\pi2w\right)\mathbf G_h.
|
||||
@@ -477,7 +450,8 @@ $$
|
||||
Maximum position compensation is
|
||||
|
||||
$$
|
||||
A_{\max}=
|
||||
A_{\max}
|
||||
=
|
||||
-\max\left(4.5-1.5H-3F,0\right)
|
||||
\quad\text{dB}.
|
||||
$$
|
||||
@@ -485,15 +459,15 @@ $$
|
||||
Longitudinal and height weights are
|
||||
|
||||
$$
|
||||
p_v=\mathrm{clamp}\left(\frac v{0.6},0,1\right),
|
||||
p_v=\operatorname{clamp}\left(\frac v{0.6},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p_w=\mathrm{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
p_w=\operatorname{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p=\mathrm{clamp}(p_v+p_w,0,1).
|
||||
p=\operatorname{clamp}(p_v+p_w,0,1).
|
||||
$$
|
||||
|
||||
The linear compensation gain is
|
||||
@@ -505,7 +479,8 @@ $$
|
||||
The object's target-gain vector is
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{target}}=
|
||||
\mathbf G_{\mathrm{target}}
|
||||
=
|
||||
G_{\mathrm{object}}
|
||||
G_{\mathrm{pos}}
|
||||
\mathbf G_{\mathrm{point}}.
|
||||
@@ -516,7 +491,8 @@ $$
|
||||
The coded position of an OAMD update is
|
||||
|
||||
$$
|
||||
s_{\mathrm{coded}}=
|
||||
s_{\mathrm{coded}}
|
||||
=
|
||||
s_{\mathrm{frame}}
|
||||
+s_{\mathrm{outer}}
|
||||
+s_{\mathrm{OAMD}}
|
||||
@@ -526,15 +502,16 @@ $$
|
||||
The theoretical update position on the decoder-output PCM timeline is
|
||||
|
||||
$$
|
||||
s_{\mathrm{theoretical}}=
|
||||
s_{\mathrm{coded}}+d_{\mathrm{decoder}},
|
||||
s_{\mathrm{theoretical}}
|
||||
=s_{\mathrm{coded}}+d_{\mathrm{decoder}},
|
||||
\qquad d_{\mathrm{decoder}}=1473.
|
||||
$$
|
||||
|
||||
The speaker renderer retains the existing processing-block length $B=32$, so the aligned update point is
|
||||
|
||||
$$
|
||||
\widehat s=
|
||||
\widehat s
|
||||
=
|
||||
B\left\lfloor
|
||||
\frac{s_{\mathrm{theoretical}}+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
@@ -545,7 +522,8 @@ Thus, for frame-aligned updates, `align32(1473)=1472`. The 1473 value is the the
|
||||
For ramp duration $D$, the number of blocks is
|
||||
|
||||
$$
|
||||
K=
|
||||
K
|
||||
=
|
||||
\left\lfloor
|
||||
\frac{D+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
@@ -572,7 +550,8 @@ If no new metadata update intervenes, this is equivalent to a sample-wise linear
|
||||
For target output channel $c$:
|
||||
|
||||
$$
|
||||
y_c[n]=
|
||||
y_c[n]
|
||||
=
|
||||
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
|
||||
+
|
||||
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
|
||||
@@ -581,7 +560,8 @@ $$
|
||||
Here
|
||||
|
||||
$$
|
||||
\delta_{c,\mathrm{LFE}}=
|
||||
\delta_{c,\mathrm{LFE}}
|
||||
=
|
||||
\begin{cases}
|
||||
1, & c\text{ is the target layout's LFE channel},\\
|
||||
0, & \text{otherwise}.
|
||||
@@ -593,15 +573,16 @@ A layout without LFE output does not mix input LFE into other channels. After ob
|
||||
For PCM24 output, quantization is
|
||||
|
||||
$$
|
||||
y_{24}[n]=
|
||||
\mathrm{trunc}\left(
|
||||
8388607\,\mathrm{clip}(y[n],-1,1)
|
||||
y_{24}[n]
|
||||
=
|
||||
\operatorname{trunc}\left(
|
||||
8388607\,\operatorname{clip}(y[n],-1,1)
|
||||
\right).
|
||||
$$
|
||||
|
||||
## 14. Scope of the formulas
|
||||
|
||||
- The JOC matrix section covers both the dense MTX and the sparse IDX/VEC differential syntax.
|
||||
- The JOC matrix section describes dense JOC; Sparse JOC uses a different sparse coefficient/index path.
|
||||
- The speaker-panning section describes ordinary point objects; extent, spread, divergence, and similar modes require additional models.
|
||||
- Multiple OAMD position blocks must be scheduled in time order.
|
||||
- A limiter is separate post-processing and is not included in the mixing equations above.
|
||||
|
||||
+78
-97
@@ -4,7 +4,7 @@
|
||||
|
||||
本文只说明 JustOneCacophony 研究路径中使用的信号模型和公式:JOC 参数如何与核心 PCM 结合并重建对象信号,以及 OAMD 坐标如何转换为扬声器增益。
|
||||
|
||||
这些公式描述项目当前研究的 JOC 矩阵参数(dense 与 sparse 两条差分语法)与普通点对象路径,不代表对所有 E-AC-3 JOC 变体的完整定义。
|
||||
这些公式描述项目当前研究的 dense JOC 与普通点对象路径,不代表对所有 E-AC-3 JOC 变体的完整定义。
|
||||
|
||||
## 1. 总体路径与记号
|
||||
|
||||
@@ -52,11 +52,9 @@ $$
|
||||
N_f=1536=24\times64.
|
||||
$$
|
||||
|
||||
## 2. JOC 矩阵参数
|
||||
## 2. Dense JOC 矩阵参数
|
||||
|
||||
每个对象、每个数据点的量化矩阵 `joc_mix_mtx_q` 都定义在 $N_q$ 个量化级上。标志位 `b_joc_sparse` 选择两条差分语法之一:dense 为每个核心声道各送一路 MTX 差分,sparse 每参数带只送一个 active 声道与一路系数差分。
|
||||
|
||||
### 2.1 Dense 差分还原
|
||||
### 2.1 差分还原
|
||||
|
||||
令 `quant_idx` 为 $q_i\in\{0,1\}$,量化级数为
|
||||
|
||||
@@ -77,76 +75,38 @@ $$
|
||||
对对象 $o$、数据点 $d$、核心声道 $c$ 和参数带 $p$,编码差分 $\Delta_{o,d,c,p}$ 还原为
|
||||
|
||||
$$
|
||||
Q_{o,d,c,0}=
|
||||
Q_{o,d,c,0}
|
||||
=
|
||||
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}=
|
||||
Q_{o,d,c,p}
|
||||
=
|
||||
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
|
||||
\qquad p>0.
|
||||
$$
|
||||
|
||||
### 2.2 Sparse 差分还原
|
||||
|
||||
令 $I_{o,d,p}$ 为 `joc_channel_idx` 符号(IDX), $V_{o,d,p}$ 为 `joc_vec` 符号(VEC), $N_c\in\lbrace5,7\rbrace$ 为核心声道数。每参数带只有一个 active 声道
|
||||
|
||||
$$
|
||||
A_{o,d,p}=
|
||||
\begin{cases}
|
||||
I_{o,d,0}, & p=0,\\
|
||||
\left(A_{o,d,p-1}+I_{o,d,p}\right)\bmod N_c, & p>0,
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
其中 $I_{o,d,0}$ 是 3 bit 绝对声道号,其余 IDX 符号是相对上一个 **active 声道**的增量。系数是一个跨参数带连续的单累加器
|
||||
|
||||
$$
|
||||
\kappa_{o,d,-1}=O^{(s)}_q,\qquad
|
||||
\kappa_{o,d,p}=
|
||||
\left(\kappa_{o,d,p-1}+V_{o,d,p}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
sparse 起点比 dense 的中心偏移高两个量化级:
|
||||
|
||||
$$
|
||||
O^{(s)}_q=
|
||||
\begin{cases}
|
||||
50, & q_i=0,\\
|
||||
100, & q_i=1.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
active 声道切换时累加器**不**重置。完整矩阵为
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}=
|
||||
\begin{cases}
|
||||
\kappa_{o,d,p}, & c=A_{o,d,p},\\
|
||||
\dfrac{N_q}{2}, & c\neq A_{o,d,p}.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
非 active 项取 $N_q/2$,即去量化后恰为 0。
|
||||
|
||||
### 2.3 去量化
|
||||
### 2.2 去量化
|
||||
|
||||
矩阵系数的去量化值为
|
||||
|
||||
$$
|
||||
D_{o,d,c,p}=
|
||||
D_{o,d,c,p}
|
||||
=
|
||||
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
|
||||
\frac{820}{4096(1+q_i)}.
|
||||
$$
|
||||
|
||||
因此 coarse 模式的有效分母为 4096,fine 模式为 8192。
|
||||
|
||||
### 2.4 JOC clipgain
|
||||
### 2.3 JOC clipgain
|
||||
|
||||
若 clipgain 字段由整数 $x$ 和尾数 $y$ 组成,则
|
||||
|
||||
$$
|
||||
G_{\mathrm{clip}}=
|
||||
G_{\mathrm{clip}}
|
||||
=
|
||||
1+\frac{y}{32}2^{x-4}.
|
||||
$$
|
||||
|
||||
@@ -192,7 +152,8 @@ $$
|
||||
$$
|
||||
|
||||
$$
|
||||
M_{o,c,b,t}=
|
||||
M_{o,c,b,t}
|
||||
=
|
||||
(1-\alpha_t)P_{o,c,b}
|
||||
+\alpha_tD_{o,c,p(b)}.
|
||||
$$
|
||||
@@ -220,8 +181,9 @@ $$
|
||||
令 $\mathcal A_b$ 表示带 polyphase 历史状态的 64-band analysis-QMF 算子,则
|
||||
|
||||
$$
|
||||
X_{c,b,t}=
|
||||
\mathcal A_b\left(
|
||||
X_{c,b,t}
|
||||
=
|
||||
\mathcal A_b\!\left(
|
||||
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
|
||||
\mathbf s^{\mathrm A}_{c,t}
|
||||
\right).
|
||||
@@ -248,7 +210,8 @@ $$
|
||||
环绕声道的 band 0 还经过 21-tap 复 FIR:
|
||||
|
||||
$$
|
||||
\widehat X_{c,0,t}=
|
||||
\widehat X_{c,0,t}
|
||||
=
|
||||
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
|
||||
$$
|
||||
|
||||
@@ -259,7 +222,8 @@ $$
|
||||
对每个对象 $o$、子带 $b$ 和时槽 $t$,对象频域值为五个核心声道的线性组合:
|
||||
|
||||
$$
|
||||
Z_{o,b,t}=
|
||||
Z_{o,b,t}
|
||||
=
|
||||
\sum_{c=0}^{4}
|
||||
M_{o,c,b,t}\widehat X_{c,b,t}.
|
||||
$$
|
||||
@@ -274,20 +238,21 @@ analysis 输入的 $1/16$ 缩放会在 inverse QMF 输出端由 $\times16$ 抵
|
||||
|
||||
$$
|
||||
\begin{aligned}
|
||||
\mathrm{zone}[2k] &= \mathrm{src}[4k],\\
|
||||
\mathrm{zone}[2k+1] &= -\mathrm{src}[4k+1],\\
|
||||
\mathrm{zone}[126-2k] &= \mathrm{src}[4k+2],\\
|
||||
\mathrm{zone}[127-2k] &= \mathrm{src}[4k+3].
|
||||
\operatorname{zone}[2k] &= \operatorname{src}[4k],\\
|
||||
\operatorname{zone}[2k+1] &= -\operatorname{src}[4k+1],\\
|
||||
\operatorname{zone}[126-2k] &= \operatorname{src}[4k+2],\\
|
||||
\operatorname{zone}[127-2k] &= \operatorname{src}[4k+3].
|
||||
\end{aligned}
|
||||
$$
|
||||
|
||||
把 `zone` 重新视为 64 个复数后执行未归一化 64 点 FFT:
|
||||
|
||||
$$
|
||||
F_k=
|
||||
F_k
|
||||
=
|
||||
\sum_{n=0}^{63}
|
||||
\mathrm{zone}_n
|
||||
\exp\left(-j\frac{2\pi kn}{64}\right).
|
||||
\operatorname{zone}_n
|
||||
\exp\!\left(-j\frac{2\pi kn}{64}\right).
|
||||
$$
|
||||
|
||||
### 7.2 调制与合成
|
||||
@@ -295,7 +260,8 @@ $$
|
||||
定义旋转系数
|
||||
|
||||
$$
|
||||
r_k=
|
||||
r_k
|
||||
=
|
||||
\frac12\left(
|
||||
\sin\frac{\pi k}{128}
|
||||
+j\cos\frac{\pi k}{128}
|
||||
@@ -311,8 +277,9 @@ $$
|
||||
令 $\mathcal S$ 表示带 640 项 synthesis window 和跨时槽状态的 polyphase 合成算子:
|
||||
|
||||
$$
|
||||
\mathbf y_{o,t}=
|
||||
\mathcal S\left(
|
||||
\mathbf y_{o,t}
|
||||
=
|
||||
\mathcal S\!\left(
|
||||
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
|
||||
\right).
|
||||
$$
|
||||
@@ -320,8 +287,9 @@ $$
|
||||
对象输出为
|
||||
|
||||
$$
|
||||
y_o[64t+r]=
|
||||
\mathrm{clip}\left(
|
||||
y_o[64t+r]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
16\,\mathbf y_{o,t}[r],-1,1
|
||||
\right)G_{\mathrm{clip}},
|
||||
$$
|
||||
@@ -333,8 +301,9 @@ $$
|
||||
LFE 不经过对象矩阵或 inverse QMF,而是使用 1217-sample 延迟。输入与输出端的比例因子抵消后:
|
||||
|
||||
$$
|
||||
y_{\mathrm{LFE}}[n]=
|
||||
\mathrm{clip}\left(
|
||||
y_{\mathrm{LFE}}[n]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
x_{\mathrm{LFE,core}}[n-1217],-1,1
|
||||
\right).
|
||||
$$
|
||||
@@ -344,8 +313,9 @@ $$
|
||||
横向和纵向网格使用 $N=62$,高度网格使用 $N=15$。量化函数为
|
||||
|
||||
$$
|
||||
q_N(k)=
|
||||
\min\left(
|
||||
q_N(k)
|
||||
=
|
||||
\min\!\left(
|
||||
32767,
|
||||
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
|
||||
\right).
|
||||
@@ -366,11 +336,11 @@ $$
|
||||
转换为 ADM 网格时:
|
||||
|
||||
$$
|
||||
k_1=\mathrm{round}\left(\frac{62q_1}{32767}\right),
|
||||
k_1=\operatorname{round}\!\left(\frac{62q_1}{32767}\right),
|
||||
\quad
|
||||
k_2=\mathrm{round}\left(\frac{62q_2}{32767}\right),
|
||||
k_2=\operatorname{round}\!\left(\frac{62q_2}{32767}\right),
|
||||
\quad
|
||||
k_3=\mathrm{round}\left(\frac{15q_3}{32767}\right),
|
||||
k_3=\operatorname{round}\!\left(\frac{15q_3}{32767}\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
@@ -434,15 +404,17 @@ $$
|
||||
二维点增益为
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{2D}}(u,v)=
|
||||
\mathbf G_{\mathrm{2D}}(u,v)
|
||||
=
|
||||
\mathbf h(u)\odot\mathbf v(v).
|
||||
$$
|
||||
|
||||
对于只有一对水平环绕、没有独立 side/rear 两对的 5.1 系列布局,纵向坐标使用
|
||||
|
||||
$$
|
||||
v_{\mathrm{floor}}=
|
||||
\mathrm{clamp}(2v,0,1).
|
||||
v_{\mathrm{floor}}
|
||||
=
|
||||
\operatorname{clamp}(2v,0,1).
|
||||
$$
|
||||
|
||||
其他布局使用 $v_{\mathrm{floor}}=v$。
|
||||
@@ -452,7 +424,8 @@ $$
|
||||
三维布局分别计算地面层增益 $\mathbf G_f$ 和高度层增益 $\mathbf G_h$:
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)=
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)
|
||||
=
|
||||
\cos\left(\frac\pi2w\right)\mathbf G_f
|
||||
+
|
||||
\sin\left(\frac\pi2w\right)\mathbf G_h.
|
||||
@@ -477,7 +450,8 @@ $$
|
||||
最大位置补偿为
|
||||
|
||||
$$
|
||||
A_{\max}=
|
||||
A_{\max}
|
||||
=
|
||||
-\max\left(4.5-1.5H-3F,0\right)
|
||||
\quad\text{dB}.
|
||||
$$
|
||||
@@ -485,15 +459,15 @@ $$
|
||||
前后与高度位置权重为
|
||||
|
||||
$$
|
||||
p_v=\mathrm{clamp}\left(\frac v{0.6},0,1\right),
|
||||
p_v=\operatorname{clamp}\left(\frac v{0.6},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p_w=\mathrm{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
p_w=\operatorname{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p=\mathrm{clamp}(p_v+p_w,0,1).
|
||||
p=\operatorname{clamp}(p_v+p_w,0,1).
|
||||
$$
|
||||
|
||||
线性补偿增益为
|
||||
@@ -505,7 +479,8 @@ $$
|
||||
对象的目标增益向量为
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{target}}=
|
||||
\mathbf G_{\mathrm{target}}
|
||||
=
|
||||
G_{\mathrm{object}}
|
||||
G_{\mathrm{pos}}
|
||||
\mathbf G_{\mathrm{point}}.
|
||||
@@ -516,7 +491,8 @@ $$
|
||||
OAMD 更新的编码位置为
|
||||
|
||||
$$
|
||||
s_{\mathrm{coded}}=
|
||||
s_{\mathrm{coded}}
|
||||
=
|
||||
s_{\mathrm{frame}}
|
||||
+s_{\mathrm{outer}}
|
||||
+s_{\mathrm{OAMD}}
|
||||
@@ -526,15 +502,16 @@ $$
|
||||
decoder 输出 PCM timeline 上的理论更新位置为
|
||||
|
||||
$$
|
||||
s_{\mathrm{theoretical}}=
|
||||
s_{\mathrm{coded}}+d_{\mathrm{decoder}},
|
||||
s_{\mathrm{theoretical}}
|
||||
=s_{\mathrm{coded}}+d_{\mathrm{decoder}},
|
||||
\qquad d_{\mathrm{decoder}}=1473.
|
||||
$$
|
||||
|
||||
扬声器 renderer 保留现有的处理块长度 $B=32$,更新点对齐为
|
||||
|
||||
$$
|
||||
\widehat s=
|
||||
\widehat s
|
||||
=
|
||||
B\left\lfloor
|
||||
\frac{s_{\mathrm{theoretical}}+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
@@ -545,7 +522,8 @@ $$
|
||||
给定 ramp duration $D$,block 数为
|
||||
|
||||
$$
|
||||
K=
|
||||
K
|
||||
=
|
||||
\left\lfloor
|
||||
\frac{D+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
@@ -572,7 +550,8 @@ $$
|
||||
对目标输出声道 $c$:
|
||||
|
||||
$$
|
||||
y_c[n]=
|
||||
y_c[n]
|
||||
=
|
||||
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
|
||||
+
|
||||
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
|
||||
@@ -581,7 +560,8 @@ $$
|
||||
其中
|
||||
|
||||
$$
|
||||
\delta_{c,\mathrm{LFE}}=
|
||||
\delta_{c,\mathrm{LFE}}
|
||||
=
|
||||
\begin{cases}
|
||||
1, & c\text{ 为目标布局的 LFE},\\
|
||||
0, & \text{其他声道}.
|
||||
@@ -593,15 +573,16 @@ $$
|
||||
若输出 PCM24,量化关系为
|
||||
|
||||
$$
|
||||
y_{24}[n]=
|
||||
\mathrm{trunc}\left(
|
||||
8388607\,\mathrm{clip}(y[n],-1,1)
|
||||
y_{24}[n]
|
||||
=
|
||||
\operatorname{trunc}\left(
|
||||
8388607\,\operatorname{clip}(y[n],-1,1)
|
||||
\right).
|
||||
$$
|
||||
|
||||
## 14. 公式适用范围
|
||||
|
||||
- JOC 矩阵部分同时描述 dense MTX 与 sparse IDX/VEC 两条差分语法。
|
||||
- JOC 矩阵部分描述 dense JOC;Sparse JOC 使用不同的稀疏系数/索引路径。
|
||||
- 扬声器声像部分描述普通点对象;extent、spread、divergence 等模式需要额外模型。
|
||||
- 多个 OAMD position block 必须按其时间顺序调度。
|
||||
- limiter 属于独立后处理,不包含在上述混音公式中。
|
||||
|
||||
+1
-1
@@ -42,7 +42,7 @@ int ejoc_renderer_process(
|
||||
float* output16_planar); /* [16][1536] */
|
||||
```
|
||||
|
||||
Python performs JOC Huffman decoding, differential reconstruction, and dequantization before the call, with the dense and sparse syntaxes sharing one entry point. The native core consumes the already dequantized `dq` in double precision, and both syntaxes have the same layout at that ABI.
|
||||
Python performs dense-JOC Huffman decoding, differential reconstruction, and dequantization before the call. Sparse JOC is not silently passed to the dense native path.
|
||||
|
||||
Thread control is exposed as:
|
||||
|
||||
|
||||
+1
-14
@@ -42,7 +42,7 @@ int ejoc_renderer_process(
|
||||
float* output16_planar); /* [16][1536] */
|
||||
```
|
||||
|
||||
JOC 的 Huffman 解码、差分还原和去量化先在 Python 中完成,dense 与 sparse 两条语法共用同一条入口。原生核心消费已去量化的 `dq`(double),两条语法在该 ABI 上布局一致。
|
||||
Dense JOC 的 Huffman 解码、差分还原和去量化先在 Python 中完成。Sparse JOC 不会被静默送入 dense 原生路径。
|
||||
|
||||
线程接口为:
|
||||
|
||||
@@ -135,19 +135,6 @@ int ejoc_binaural_renderer_process(
|
||||
|
||||
Python 负责模型解析、OAMD 时间轴和每 512 samples 的 complex gains/room sends。C++ handle 保存 QMF、hybrid、递归 room 和 QMF synthesis 状态。全部输入、状态、乘加和输出均为 double/complex double。
|
||||
|
||||
## 6.1 公开 SOFA 双耳渲染 ABI
|
||||
|
||||
共享库同时提供完整的原生 SOFA 双耳渲染器(`ejoc_sofa_binaural_*`),它镜像
|
||||
Python `SofaBinauralBackend` 的全部数学:64-QMF/77-hybrid analysis/synthesis、
|
||||
五阶 ACN/N3D 实球谐方向场求值、whole-QMF-slot 逐对象 delay 历史、六面一阶
|
||||
image-source early reflections、共享 unitary FDN late room、LFE 120–180 Hz
|
||||
低通与 961-sample latency 语义。kernel 表、编译好的 HRTF 场与房间常数通过
|
||||
`configure_kernels/configure_field/configure_room` 一次上传;每 512-sample
|
||||
block 先 `set_source` 更新 16 个 source,再 `process` 输入 PCM;`process` 返回
|
||||
裁剪后的 stereo 样本数(首个 961 samples 被丢弃)。`finish` 以 64-sample 对齐的
|
||||
块排空尾音。Python 桥位于 `src/sofa_native_backend.py`,与 Python 参考实现逐值
|
||||
一致(差异 < 1e-9);原生库缺失时 `main.py` 自动回退 Python。
|
||||
|
||||
## 7. 构建
|
||||
|
||||
CMake 定义位于 `native/CMakeLists.txt`。从仓库根目录运行:
|
||||
|
||||
@@ -6,7 +6,6 @@ import math
|
||||
import os
|
||||
from pathlib import Path
|
||||
import platform
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
@@ -28,19 +27,11 @@ import oamd_tracks
|
||||
from renderer import JocRenderer
|
||||
from native_renderer import NativeBackendUnavailable, NativeJocRenderer
|
||||
from binaural_renderer import (
|
||||
DEFAULT_SOFA_HRTF,
|
||||
SofaBinauralRenderer,
|
||||
resolve_compiled_hrtf_cache,
|
||||
resolve_sofa_hrtf,
|
||||
)
|
||||
from rosella_binaural_renderer import (
|
||||
DEFAULT_PERSONALIZED_HEADPHONE,
|
||||
ROSSELLA_BLOCK_SAMPLES,
|
||||
ROSSELLA_LATENCY_SAMPLES,
|
||||
RosellaBinauralRenderer,
|
||||
resolve_personalized_headphone,
|
||||
)
|
||||
from sofa_hrtf_field import DEFAULT_HRTF_CACHE_DIR
|
||||
from speaker_backend import create_speaker_renderer
|
||||
from speaker_layouts import (SPEAKER_LAYOUT_CHOICES, get_speaker_layout,
|
||||
speaker_layout_display_name)
|
||||
@@ -51,9 +42,6 @@ from variant_error import UnsupportedVariantError, write_variant_report
|
||||
RATE = 48000
|
||||
FRAME_SAMPLES = 1536
|
||||
DEFAULT_OUTPUT_DIR = PROJECT_DIR / "output"
|
||||
EAC3_DRC_SCALE_MAX = 6.0
|
||||
EAC3_TARGET_LEVEL_RANGE = (-31, 0)
|
||||
EAC3_DECODER_OPTION_RE = re.compile(r"(?m)^\s*-([A-Za-z0-9_]+)\s+<")
|
||||
|
||||
|
||||
def resolve_output(source, requested=None, speaker_layout=None, *, binaural=False):
|
||||
@@ -70,97 +58,6 @@ def resolve_output(source, requested=None, speaker_layout=None, *, binaural=Fals
|
||||
return target.expanduser().resolve()
|
||||
|
||||
|
||||
def _find_default_compiled_hrtf_cache():
|
||||
"""在默认 cache 目录寻找唯一的 .jochrtf;无文件返回 None,多个则报错。"""
|
||||
directory = DEFAULT_HRTF_CACHE_DIR
|
||||
if not directory.is_dir():
|
||||
return None
|
||||
candidates = sorted(directory.glob("*.jochrtf"))
|
||||
if not candidates:
|
||||
return None
|
||||
if len(candidates) > 1:
|
||||
listing = ", ".join(path.name for path in candidates[:8])
|
||||
raise ValueError(
|
||||
f"{directory} 下有多个 .jochrtf 缓存({listing}…),无法自动选择;"
|
||||
"请用 --compiled-hrtf-cache PATH 或 --sofa-hrtf PATH 显式指定")
|
||||
return candidates[0]
|
||||
|
||||
|
||||
def resolve_binaural_hrtf_input(args, *, required):
|
||||
"""解析 binaural 的 HRTF 输入。
|
||||
|
||||
无显式输入时按顺序回退:默认 HRTF/binaural.sofa → 默认 cache 目录下唯一的
|
||||
.jochrtf → 默认 HRTF/binaural.personalized_headphone → 报错。
|
||||
只校验路径,不做编译。
|
||||
"""
|
||||
sofa = args.sofa_hrtf
|
||||
compiled = args.compiled_hrtf_cache
|
||||
private = args.personalized_headphone
|
||||
cache_policy = args.hrtf_cache_policy
|
||||
cache_dir = args.hrtf_cache_dir
|
||||
radius = args.hrtf_radius_m
|
||||
|
||||
if compiled is not None and cache_policy is not None:
|
||||
raise ValueError("显式 .jochrtf 输入不能再指定 --hrtf-cache-policy")
|
||||
if compiled is not None and radius != 1.0:
|
||||
raise ValueError("显式 .jochrtf 输入不能再选择 SOFA radius shell")
|
||||
if private is not None and (cache_policy is not None or cache_dir is not None
|
||||
or radius != 1.0):
|
||||
raise ValueError(
|
||||
"Rosella 模型输入不能使用 "
|
||||
"--hrtf-cache-policy/--hrtf-cache-dir/--hrtf-radius-m")
|
||||
|
||||
if required and sofa is None and compiled is None and private is None:
|
||||
if DEFAULT_SOFA_HRTF.is_file():
|
||||
sofa = DEFAULT_SOFA_HRTF
|
||||
else:
|
||||
compiled = _find_default_compiled_hrtf_cache()
|
||||
if compiled is None and DEFAULT_PERSONALIZED_HEADPHONE.is_file():
|
||||
private = DEFAULT_PERSONALIZED_HEADPHONE
|
||||
|
||||
if sofa is None and compiled is None and private is None:
|
||||
if cache_policy is not None or cache_dir is not None or radius != 1.0:
|
||||
raise ValueError("HRTF cache/radius 选项需要 --sofa-hrtf")
|
||||
if required:
|
||||
raise ValueError(
|
||||
"--binaural 未找到 HRTF 输入:默认 "
|
||||
f"{DEFAULT_SOFA_HRTF}、{DEFAULT_PERSONALIZED_HEADPHONE} 与 "
|
||||
f"{DEFAULT_HRTF_CACHE_DIR} 下的 .jochrtf 缓存都不存在;请用 "
|
||||
"--sofa-hrtf PATH、--compiled-hrtf-cache PATH 或 "
|
||||
"--personalized-headphone PATH 指定")
|
||||
return None
|
||||
|
||||
if sofa is None and (cache_policy is not None or cache_dir is not None
|
||||
or radius != 1.0):
|
||||
raise ValueError("HRTF cache/radius 选项需要 --sofa-hrtf")
|
||||
effective_policy = "memory" if cache_policy is None else cache_policy
|
||||
if cache_dir is not None and (sofa is None or effective_policy != "disk"):
|
||||
raise ValueError("--hrtf-cache-dir 仅与 SOFA 的 disk cache policy 一起使用")
|
||||
if sofa is not None:
|
||||
return {
|
||||
"kind": "sofa",
|
||||
"path": resolve_sofa_hrtf(sofa),
|
||||
"cache_policy": effective_policy,
|
||||
"cache_dir": (DEFAULT_HRTF_CACHE_DIR if cache_dir is None else
|
||||
cache_dir.expanduser().resolve()),
|
||||
}
|
||||
if compiled is not None:
|
||||
return {
|
||||
"kind": "compiled_cache",
|
||||
"path": resolve_compiled_hrtf_cache(compiled),
|
||||
"cache_policy": None,
|
||||
"cache_dir": None,
|
||||
}
|
||||
if private is not None:
|
||||
return {
|
||||
"kind": "rosella",
|
||||
"path": resolve_personalized_headphone(private),
|
||||
"cache_policy": None,
|
||||
"cache_dir": None,
|
||||
}
|
||||
return None
|
||||
|
||||
|
||||
def executable(value, name):
|
||||
path = shutil.which(value) if value else None
|
||||
if path is None and value and Path(value).is_file():
|
||||
@@ -187,53 +84,6 @@ def timed_call(timings, name, function, *args, **kwargs):
|
||||
timings[name] = time.perf_counter() - started
|
||||
|
||||
|
||||
def probe_eac3_decoder_options(ffmpeg):
|
||||
"""读取 ``ffmpeg -h decoder=eac3`` 暴露的 AVOption 名。"""
|
||||
result = subprocess.run(
|
||||
[ffmpeg, "-hide_banner", "-h", "decoder=eac3"],
|
||||
stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
|
||||
text=True, encoding="utf-8", errors="replace")
|
||||
options = frozenset(EAC3_DECODER_OPTION_RE.findall(result.stdout or ""))
|
||||
# decoder 名不存在时 ffmpeg 依然返回 0,因此以“解析不到任何选项”为失败。
|
||||
if not options:
|
||||
raise RuntimeError(
|
||||
"无法读取 FFmpeg 的 eac3 解码器选项(ffmpeg -h decoder=eac3);"
|
||||
"需要带 E-AC-3 解码器的构建")
|
||||
return options
|
||||
|
||||
|
||||
def ffmpeg_version(ffmpeg):
|
||||
"""FFmpeg 版本字符串;探测失败返回空串,不影响渲染。"""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
[ffmpeg, "-hide_banner", "-version"],
|
||||
stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
|
||||
text=True, encoding="utf-8", errors="replace")
|
||||
except OSError:
|
||||
return ""
|
||||
lines = (result.stdout or "").splitlines()
|
||||
line = lines[0].strip() if lines else ""
|
||||
prefix = "ffmpeg version "
|
||||
return line[len(prefix):].strip() if line.startswith(prefix) else line
|
||||
|
||||
|
||||
def eac3_decode_options(drc_scale, target_level, available):
|
||||
"""构造 ``-i`` 之前的 E-AC-3 解码选项,返回 ``(argv, report 片段)``。"""
|
||||
if "drc_scale" not in available:
|
||||
raise RuntimeError(
|
||||
"FFmpeg 的 eac3 解码器缺少 -drc_scale,无法关闭码流 DRC")
|
||||
# -drc_scale 始终显式下发:0(全动态范围)不是 ffmpeg 的默认值。
|
||||
argv = ["-drc_scale", format(float(drc_scale), ".10g")]
|
||||
if target_level:
|
||||
if "target_level" not in available:
|
||||
raise RuntimeError(
|
||||
"FFmpeg 的 eac3 解码器不支持 -target_level;请升级 FFmpeg "
|
||||
"或去掉 --eac3-target-level")
|
||||
argv += ["-target_level", str(int(target_level))]
|
||||
applied = {"drc_scale": float(drc_scale), "target_level": int(target_level)}
|
||||
return argv, applied
|
||||
|
||||
|
||||
def extract_eac3(ffmpeg, source, target):
|
||||
if source.suffix.lower() in (".eac3", ".ec3"):
|
||||
return source
|
||||
@@ -243,10 +93,10 @@ def extract_eac3(ffmpeg, source, target):
|
||||
return target
|
||||
|
||||
|
||||
def decode_core(ffmpeg, eac3, target, duration_sec=None, *, options=()):
|
||||
def decode_core(ffmpeg, eac3, target, duration_sec=None):
|
||||
# 5.1(side) 的 f32le 顺序为 FL FR FC LFE SL SR;JOC 使用其中 0,1,2,4,5。
|
||||
command = [ffmpeg, "-hide_banner", "-loglevel", "error", "-y", *options,
|
||||
"-i", str(eac3), "-map", "0:a:0", "-vn"]
|
||||
command = [ffmpeg, "-hide_banner", "-loglevel", "error", "-y", "-i", str(eac3),
|
||||
"-map", "0:a:0", "-vn"]
|
||||
if duration_sec is not None:
|
||||
command.extend(["-t", f"{duration_sec:.9f}"])
|
||||
command.extend(["-ac", "6", "-ar", str(RATE),
|
||||
@@ -468,7 +318,7 @@ def render(index, bed_path, frame_count, raw_path, gain, progress_every,
|
||||
def build_parser():
|
||||
parser = argparse.ArgumentParser(
|
||||
description=("JustOneCacophony (JOC):E-AC-3 JOC → 25ch ADM BWF、"
|
||||
"扬声器 WAV 或公开 SOFA 双耳 WAV"))
|
||||
"扬声器 WAV 或 DLL-free Rosella 双耳 WAV"))
|
||||
parser.add_argument("input", type=Path, help="输入 .m4a/.eac3/.ec3")
|
||||
parser.add_argument("-o", "--output", type=Path, help="输出文件;默认按模式和布局命名")
|
||||
parser.add_argument("--speaker-output", type=Path,
|
||||
@@ -479,7 +329,7 @@ def build_parser():
|
||||
direct_mode.add_argument("--speaker-layout", choices=SPEAKER_LAYOUT_CHOICES,
|
||||
help="直接扬声器渲染布局,例如 2.0、5.1、7.1.2")
|
||||
direct_mode.add_argument("--binaural", action="store_true",
|
||||
help="直接 SOFA 双耳渲染;不生成临时 ADM BWF")
|
||||
help="直接 DLL-free Rosella 双耳渲染;不生成临时 ADM BWF")
|
||||
parser.add_argument("--speaker-format", choices=("float32", "int24"), default="float32",
|
||||
help="扬声器 WAV 格式,默认 float32")
|
||||
parser.add_argument("--binaural-format", choices=("float32", "int24"), default="float32",
|
||||
@@ -489,34 +339,10 @@ def build_parser():
|
||||
help="int24 削波处理:交互询问、继续截断、改 float32 或中止")
|
||||
parser.add_argument("--speaker-metadata-offset", type=int, default=1473,
|
||||
help="扬声器渲染 metadata 相对帧偏移,默认 1473 samples")
|
||||
parser.add_argument("--binaural-mode", choices=("off", "near", "mid", "far"),
|
||||
default="mid",
|
||||
help="双耳渲染模式,默认 mid(人为指定的渲染提示,非码流 "
|
||||
"原始元数据);直接双耳渲染与 ADM BWF 的 DBMD 提示共用。"
|
||||
"off 仅用于 ADM BWF:关闭 DBMD 双耳提示(编码 0)")
|
||||
hrtf_input = parser.add_mutually_exclusive_group()
|
||||
hrtf_input.add_argument(
|
||||
"--sofa-hrtf", type=Path,
|
||||
help="SimpleFreeFieldHRIR SOFA;缺省时依次尝试 HRTF/binaural.sofa、"
|
||||
"output/hrtf-cache 下唯一的 .jochrtf、"
|
||||
"HRTF/binaural.personalized_headphone,均无则报错")
|
||||
hrtf_input.add_argument(
|
||||
"--compiled-hrtf-cache", type=Path,
|
||||
help="高级入口:显式读取 JOC .jochrtf compiled cache")
|
||||
hrtf_input.add_argument(
|
||||
"--personalized-headphone", type=Path, nargs="?",
|
||||
const=DEFAULT_PERSONALIZED_HEADPHONE,
|
||||
help="Rosella .personalized_headphone 模型;不带路径时默认 "
|
||||
"HRTF/binaural.personalized_headphone")
|
||||
parser.add_argument(
|
||||
"--hrtf-cache-policy", choices=("none", "memory", "disk"), default=None,
|
||||
help="SOFA 编译缓存;默认 memory,disk 写入可删除的 .jochrtf")
|
||||
parser.add_argument(
|
||||
"--hrtf-cache-dir", type=Path,
|
||||
help="disk cache 目录;默认 output/hrtf-cache")
|
||||
parser.add_argument(
|
||||
"--hrtf-radius-m", type=float, default=1.0,
|
||||
help="选择最近的 SOFA measurement-radius shell,默认 1.0 m")
|
||||
parser.add_argument("--binaural-mode", choices=("near", "mid", "far"), default="mid",
|
||||
help="普通对象 Rosella 距离模式,默认 mid;LFE 始终走 special 低通")
|
||||
parser.add_argument("--personalized-headphone", "--binaural-hrtf", dest="personalized_headphone",
|
||||
type=Path, help="覆盖 HRTF/binaural.personalized_headphone")
|
||||
parser.add_argument("--binaural-tail-seconds", type=float, default=5.0,
|
||||
help="双耳 room/filterbank flush 上限,默认 5 秒")
|
||||
parser.add_argument("--binaural-tail-threshold", type=float, default=1.0e-8,
|
||||
@@ -528,17 +354,15 @@ def build_parser():
|
||||
parser.add_argument("--duration", type=float, help="只处理开头指定秒数")
|
||||
parser.add_argument("--object-delay-samples", type=int, default=1473,
|
||||
help="对象 PCM/OAMD 时间补偿;ADM 与双耳默认 1473 samples")
|
||||
parser.add_argument(
|
||||
"--joc-binaural-mode", choices=tuple(adm_atmos.JOC_BINAURAL_MODES),
|
||||
default=adm_atmos.JOC_BINAURAL_MODE_DEFAULT,
|
||||
help="ADM DBMD JOC 模式:off/near/far/mid/unspecified;仅影响 ADM BWF")
|
||||
parser.add_argument("--trajectory-mode", choices=("compact", "dense64"), default="compact",
|
||||
help="ADM 对象轨迹表示;直接双耳路径不序列化 AXML")
|
||||
parser.add_argument("--ffmpeg", default=os.environ.get("FFMPEG", "ffmpeg"))
|
||||
parser.add_argument("--eac3-drc-scale", type=float, default=0.0,
|
||||
help="E-AC-3 解码器 -drc_scale:0=关闭码流 dynrng(全动态范围),"
|
||||
"1=码流作者意图,>1 非对称;默认 0")
|
||||
parser.add_argument("--eac3-target-level", type=int, default=0,
|
||||
help="E-AC-3 解码器 -target_level:按码流 dialnorm 归一化电平,"
|
||||
"增益约 target_level - dialnorm dB;0=不施加,默认 0")
|
||||
parser.add_argument("--backend", choices=("auto", "native", "python"), default="auto",
|
||||
help="JOC/扬声器 DSP 后端;SOFA 双耳 DSP 当前使用 Python")
|
||||
help="DSP 后端;auto 优先 C++,不可用时回退 Python")
|
||||
parser.add_argument("--native-library", type=Path,
|
||||
help="显式指定原生库;默认从单层 lib 目录选择当前平台文件")
|
||||
parser.add_argument("--native-threads", type=int,
|
||||
@@ -573,11 +397,6 @@ def main(argv=None):
|
||||
raise FileNotFoundError(source)
|
||||
speaker_mode = args.speaker_layout is not None
|
||||
binaural_mode = bool(args.binaural)
|
||||
binaural_render_mode = args.binaural_mode
|
||||
if binaural_mode and binaural_render_mode == "off":
|
||||
raise ValueError(
|
||||
"--binaural-mode off 仅用于 ADM BWF 输出(关闭 DBMD 双耳提示);"
|
||||
"直接双耳渲染请使用 near/mid/far")
|
||||
if args.speaker_output is not None and not speaker_mode:
|
||||
raise ValueError("--speaker-output 必须与 --speaker-layout 一起使用")
|
||||
if args.binaural_output is not None and not binaural_mode:
|
||||
@@ -590,26 +409,14 @@ def main(argv=None):
|
||||
raise ValueError("--speaker-output 与 --binaural-output 不能同时使用")
|
||||
if args.speaker_metadata_offset < 0:
|
||||
raise ValueError("speaker-metadata-offset 不能为负数")
|
||||
hrtf_options_used = any((
|
||||
args.sofa_hrtf is not None,
|
||||
args.compiled_hrtf_cache is not None,
|
||||
args.personalized_headphone is not None,
|
||||
args.hrtf_cache_policy is not None,
|
||||
args.hrtf_cache_dir is not None,
|
||||
args.hrtf_radius_m != 1.0,
|
||||
))
|
||||
if hrtf_options_used and not binaural_mode:
|
||||
raise ValueError("SOFA/HRTF 选项仅与 --binaural 一起使用")
|
||||
if (not math.isfinite(args.binaural_tail_seconds)
|
||||
or args.binaural_tail_seconds < 0):
|
||||
raise ValueError("binaural-tail-seconds 必须是非负有限值")
|
||||
if (not math.isfinite(args.binaural_tail_threshold)
|
||||
or args.binaural_tail_threshold < 0):
|
||||
raise ValueError("binaural-tail-threshold 必须是非负有限值")
|
||||
if args.personalized_headphone is not None and not binaural_mode:
|
||||
raise ValueError("--personalized-headphone 仅与 --binaural 一起使用")
|
||||
if args.binaural_tail_seconds < 0:
|
||||
raise ValueError("binaural-tail-seconds 不能为负数")
|
||||
if args.binaural_tail_threshold < 0:
|
||||
raise ValueError("binaural-tail-threshold 不能为负数")
|
||||
if args.binaural_chunk_frames <= 0:
|
||||
raise ValueError("binaural-chunk-frames 必须大于 0")
|
||||
if not math.isfinite(args.hrtf_radius_m) or args.hrtf_radius_m <= 0.0:
|
||||
raise ValueError("hrtf-radius-m 必须是正有限值")
|
||||
requested_output = (args.speaker_output if args.speaker_output is not None
|
||||
else args.binaural_output if args.binaural_output is not None
|
||||
else args.output)
|
||||
@@ -625,28 +432,10 @@ def main(argv=None):
|
||||
gain = np.float32(gain_float64)
|
||||
if not math.isfinite(gain_float64) or not np.isfinite(gain):
|
||||
raise ValueError("gain-db 超出支持范围")
|
||||
if (not math.isfinite(args.eac3_drc_scale)
|
||||
or not 0.0 <= args.eac3_drc_scale <= EAC3_DRC_SCALE_MAX):
|
||||
raise ValueError(f"eac3-drc-scale 必须在 0..{EAC3_DRC_SCALE_MAX:g} 之间")
|
||||
if not (EAC3_TARGET_LEVEL_RANGE[0] <= args.eac3_target_level
|
||||
<= EAC3_TARGET_LEVEL_RANGE[1]):
|
||||
raise ValueError("eac3-target-level 必须在 -31..0 之间")
|
||||
binaural_hrtf_input = resolve_binaural_hrtf_input(
|
||||
args, required=binaural_mode and not args.metadata_only)
|
||||
binaural_model_path = (
|
||||
resolve_personalized_headphone(args.personalized_headphone)
|
||||
if binaural_mode and not args.metadata_only else None)
|
||||
ffmpeg = executable(args.ffmpeg, "FFmpeg")
|
||||
decode_options = ()
|
||||
decode_option_info = None
|
||||
if not args.metadata_only:
|
||||
available = probe_eac3_decoder_options(ffmpeg)
|
||||
decode_options, applied = eac3_decode_options(
|
||||
args.eac3_drc_scale, args.eac3_target_level, available)
|
||||
decode_option_info = {
|
||||
"version": ffmpeg_version(ffmpeg),
|
||||
"eac3_decode_options": applied,
|
||||
}
|
||||
print(f"[decode] ffmpeg {decode_option_info['version']} "
|
||||
f"drc_scale={args.eac3_drc_scale:g} "
|
||||
f"target_level={args.eac3_target_level}", flush=True)
|
||||
|
||||
total_started = time.perf_counter()
|
||||
timings = {}
|
||||
@@ -682,8 +471,7 @@ def main(argv=None):
|
||||
|
||||
bed_path = timed_call(
|
||||
timings, "decode_core", decode_core,
|
||||
ffmpeg, eac3, temp_dir / "core51_f32le.raw", duration_sec,
|
||||
options=decode_options)
|
||||
ffmpeg, eac3, temp_dir / "core51_f32le.raw", duration_sec)
|
||||
raw_path = (output.with_name(output.name + ".objects16.f32le")
|
||||
if args.keep_raw else None)
|
||||
master = None
|
||||
@@ -692,7 +480,6 @@ def main(argv=None):
|
||||
speaker_clip_info = None
|
||||
speaker_actual_format = None
|
||||
binaural_backend_info = None
|
||||
binaural_hrtf_report = None
|
||||
binaural_wav_info = None
|
||||
binaural_clip_info = None
|
||||
binaural_actual_format = None
|
||||
@@ -740,95 +527,26 @@ def main(argv=None):
|
||||
info = (f"speaker layout={speaker_name}, format={speaker_actual_format}, "
|
||||
f"peak={speaker_clip_info['peak']:.9g}")
|
||||
elif binaural_mode:
|
||||
hrtf_source = binaural_hrtf_input
|
||||
common_options = {
|
||||
"mode": binaural_render_mode,
|
||||
"object_delay_samples": args.object_delay_samples,
|
||||
"tail_seconds": args.binaural_tail_seconds,
|
||||
"output_gain": gain_float64,
|
||||
"chunk_frames": args.binaural_chunk_frames,
|
||||
}
|
||||
if hrtf_source["kind"] == "sofa":
|
||||
binaural_decoder = None
|
||||
if args.backend in ("auto", "native"):
|
||||
try:
|
||||
from sofa_native_backend import create_native_sofa_renderer
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
create_native_sofa_renderer,
|
||||
hrtf_source["path"],
|
||||
cache_policy=hrtf_source["cache_policy"],
|
||||
cache_dir=hrtf_source["cache_dir"],
|
||||
shell_radius_m=args.hrtf_radius_m,
|
||||
**common_options)
|
||||
except (ImportError, OSError, RuntimeError, ValueError) as exc:
|
||||
print(
|
||||
f"[binaural] native SOFA backend unavailable "
|
||||
f"({exc.__class__.__name__}: {exc}); "
|
||||
f"falling back to Python", flush=True)
|
||||
binaural_decoder = None
|
||||
if binaural_decoder is None:
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
SofaBinauralRenderer.from_sofa,
|
||||
hrtf_source["path"],
|
||||
cache_policy=hrtf_source["cache_policy"],
|
||||
cache_dir=hrtf_source["cache_dir"],
|
||||
shell_radius_m=args.hrtf_radius_m,
|
||||
**common_options)
|
||||
elif hrtf_source["kind"] == "rosella":
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
RosellaBinauralRenderer,
|
||||
hrtf_source["path"],
|
||||
mode=binaural_render_mode,
|
||||
object_delay_samples=args.object_delay_samples,
|
||||
tail_seconds=args.binaural_tail_seconds,
|
||||
output_gain=gain_float64,
|
||||
chunk_frames=args.binaural_chunk_frames,
|
||||
backend=args.backend,
|
||||
native_library=args.native_library)
|
||||
else:
|
||||
binaural_decoder = None
|
||||
if args.backend in ("auto", "native"):
|
||||
try:
|
||||
from sofa_native_backend import (
|
||||
create_native_compiled_cache_renderer)
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
create_native_compiled_cache_renderer,
|
||||
hrtf_source["path"],
|
||||
**common_options)
|
||||
except (ImportError, OSError, RuntimeError, ValueError) as exc:
|
||||
print(
|
||||
f"[binaural] native SOFA backend unavailable "
|
||||
f"({exc.__class__.__name__}: {exc}); "
|
||||
f"falling back to Python", flush=True)
|
||||
binaural_decoder = None
|
||||
if binaural_decoder is None:
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
SofaBinauralRenderer.from_compiled_cache,
|
||||
hrtf_source["path"],
|
||||
**common_options)
|
||||
model_path = binaural_model_path
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer", RosellaBinauralRenderer,
|
||||
model_path, mode=args.binaural_mode,
|
||||
object_delay_samples=args.object_delay_samples,
|
||||
tail_seconds=args.binaural_tail_seconds,
|
||||
output_gain=gain_float64,
|
||||
chunk_frames=args.binaural_chunk_frames,
|
||||
backend=args.backend, native_library=args.native_library)
|
||||
print(
|
||||
f"[binaural] mode={binaural_render_mode} "
|
||||
f"[binaural] mode={args.binaural_mode} "
|
||||
f"backend={binaural_decoder.dsp_backend} "
|
||||
f"precision=float64/complex128 "
|
||||
f"hrtf={hrtf_source['kind']}:{hrtf_source['path']}", flush=True)
|
||||
if hrtf_source["kind"] == "rosella":
|
||||
flush_samples = math.ceil(
|
||||
(args.binaural_tail_seconds * RATE
|
||||
+ ROSSELLA_LATENCY_SAMPLES + ROSSELLA_BLOCK_SAMPLES)
|
||||
/ ROSSELLA_BLOCK_SAMPLES) * ROSSELLA_BLOCK_SAMPLES
|
||||
spool_capacity = frame_count * FRAME_SAMPLES + flush_samples
|
||||
else:
|
||||
spool_capacity = (
|
||||
frame_count * FRAME_SAMPLES
|
||||
+ binaural_decoder.finish_capacity_samples)
|
||||
f"precision=float64/complex128 model={model_path}", flush=True)
|
||||
flush_samples = math.ceil(
|
||||
(args.binaural_tail_seconds * RATE
|
||||
+ ROSSELLA_LATENCY_SAMPLES + ROSSELLA_BLOCK_SAMPLES)
|
||||
/ ROSSELLA_BLOCK_SAMPLES) * ROSSELLA_BLOCK_SAMPLES
|
||||
spool = BinauralPcmSpool(
|
||||
temp_dir / "binaural_interleaved_f64.raw",
|
||||
spool_capacity,
|
||||
frame_count * FRAME_SAMPLES + flush_samples,
|
||||
tail_threshold=args.binaural_tail_threshold)
|
||||
try:
|
||||
render_seconds, renderer_backend, render_breakdown = timed_call(
|
||||
@@ -857,32 +575,12 @@ def main(argv=None):
|
||||
"kept_samples": spool.sample_count,
|
||||
}
|
||||
binaural_backend_info = binaural_decoder.backend_info
|
||||
if hrtf_source["kind"] == "rosella":
|
||||
binaural_hrtf_report = {
|
||||
"input_kind": "rosella",
|
||||
"input_path": str(binaural_decoder.model_path.resolve()),
|
||||
"model_coefficient_sha256": (
|
||||
binaural_decoder.model.coefficient_sha256),
|
||||
"cache_policy": None,
|
||||
}
|
||||
else:
|
||||
binaural_hrtf_report = {
|
||||
"input_kind": binaural_backend_info["hrtf_input_kind"],
|
||||
"input_path": binaural_backend_info["hrtf_input_path"],
|
||||
"source_sha256": (
|
||||
binaural_backend_info["field"]["source_sha256"]),
|
||||
"cache_policy": binaural_backend_info["cache_policy"],
|
||||
"cache_key": (
|
||||
binaural_backend_info["field"]["cache_key"]),
|
||||
"format_version": (
|
||||
binaural_backend_info["field"]["format_version"]),
|
||||
}
|
||||
finally:
|
||||
spool.close()
|
||||
timings["build_adm_tracks"] = 0.0
|
||||
timings["finalize_adm"] = 0.0
|
||||
timings["validate_adm"] = 0.0
|
||||
info = (f"binaural mode={binaural_render_mode}, "
|
||||
info = (f"binaural mode={args.binaural_mode}, "
|
||||
f"format={binaural_actual_format}, "
|
||||
f"peak={binaural_clip_info['peak']:.9g}, "
|
||||
f"samples={binaural_clip_info['kept_samples']}")
|
||||
@@ -890,7 +588,7 @@ def main(argv=None):
|
||||
timings["create_binaural_renderer"] = 0.0
|
||||
master = adm_assemble.StreamingMaster(
|
||||
output, duration_sec, rate=RATE,
|
||||
joc_binaural_mode=adm_atmos.JOC_BINAURAL_MODES[args.binaural_mode])
|
||||
joc_binaural_mode=adm_atmos.JOC_BINAURAL_MODES[args.joc_binaural_mode])
|
||||
try:
|
||||
render_seconds, renderer_backend, render_breakdown = timed_call(
|
||||
timings, "render_and_stream", variant_call,
|
||||
@@ -935,8 +633,9 @@ def main(argv=None):
|
||||
"gain_float64": float(gain_float64),
|
||||
"object_delay_samples": (None if speaker_mode else args.object_delay_samples),
|
||||
"trajectory_mode": args.trajectory_mode if mode_name == "adm" else None,
|
||||
"binaural_mode_value": (
|
||||
adm_atmos.JOC_BINAURAL_MODES[args.binaural_mode]
|
||||
"joc_binaural_mode": args.joc_binaural_mode if mode_name == "adm" else None,
|
||||
"joc_binaural_mode_value": (
|
||||
adm_atmos.JOC_BINAURAL_MODES[args.joc_binaural_mode]
|
||||
if mode_name == "adm" else None),
|
||||
"render_seconds": render_seconds,
|
||||
"render_breakdown": render_breakdown,
|
||||
@@ -947,10 +646,9 @@ def main(argv=None):
|
||||
"speaker_clip": speaker_clip_info,
|
||||
"speaker_wav": speaker_wav_info,
|
||||
"binaural_renderer_backend": binaural_backend_info,
|
||||
"binaural_mode": (
|
||||
args.binaural_mode if (binaural_mode or mode_name == "adm") else None),
|
||||
"binaural_hrtf": (
|
||||
binaural_hrtf_report if binaural_backend_info else None),
|
||||
"binaural_mode": args.binaural_mode if binaural_mode else None,
|
||||
"personalized_headphone": (
|
||||
binaural_backend_info["model"] if binaural_backend_info else None),
|
||||
"binaural_clip": binaural_clip_info,
|
||||
"binaural_wav": binaural_wav_info,
|
||||
"output_clip": speaker_clip_info if speaker_mode else binaural_clip_info,
|
||||
@@ -964,7 +662,6 @@ def main(argv=None):
|
||||
"sha256": output_sha,
|
||||
"python": platform.python_version(),
|
||||
"numpy": np.__version__,
|
||||
"ffmpeg": decode_option_info,
|
||||
}
|
||||
report_path = Path(str(output) + ".report.json")
|
||||
report_path.write_text(json.dumps(report, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
|
||||
@@ -9,7 +9,6 @@ add_library(eac3joc_core SHARED
|
||||
src/eac3joc_core.cpp
|
||||
src/speaker_renderer.cpp
|
||||
src/binaural_renderer.cpp
|
||||
src/sofa_binaural_renderer.cpp
|
||||
src/joc_huffman_tables.h
|
||||
src/qmf_tables.h
|
||||
src/speaker_layouts.h
|
||||
|
||||
@@ -53,9 +53,8 @@ Fixed array layouts used by ejoc_renderer_process():
|
||||
output16 [16][1536]
|
||||
|
||||
Only objects selected by object_mask are read from the descriptor arrays.
|
||||
dq carries already dequantized matrix coefficients in double precision; the
|
||||
caller performs the JOC bitstream differential decoding for both dense and
|
||||
sparse objects, so this ABI is identical for both syntaxes.
|
||||
Sparse JOC must be rejected by the caller; this ABI accepts already dequantized
|
||||
dense matrix coefficients.
|
||||
*/
|
||||
|
||||
EJOC_API uint32_t EJOC_CALL ejoc_abi_version(void);
|
||||
@@ -64,15 +63,6 @@ EJOC_API ejoc_renderer_handle EJOC_CALL ejoc_renderer_create(void);
|
||||
EJOC_API void EJOC_CALL ejoc_renderer_destroy(ejoc_renderer_handle handle);
|
||||
EJOC_API int EJOC_CALL ejoc_renderer_reset(ejoc_renderer_handle handle);
|
||||
EJOC_API int EJOC_CALL ejoc_renderer_set_threads(ejoc_renderer_handle handle, uint32_t total_threads);
|
||||
/*
|
||||
Enables or disables the Ls/Rs band-0 21-tap DC compensation. The caller derives
|
||||
it from the JOC downmix configuration: only configurations 3 and 4 enable the
|
||||
filter. When disabled, band 0 keeps the common per-band processing (surround
|
||||
delay plus -j rotation) instead of being overwritten by the FIR. The delay line
|
||||
and DC history advance either way, so the flag may change between frames.
|
||||
Defaults to enabled when never called.
|
||||
*/
|
||||
EJOC_API int EJOC_CALL ejoc_renderer_set_dc_filter(ejoc_renderer_handle handle, uint32_t enabled);
|
||||
EJOC_API uint32_t EJOC_CALL ejoc_renderer_thread_count(ejoc_renderer_handle handle);
|
||||
EJOC_API const char* EJOC_CALL ejoc_renderer_last_error(ejoc_renderer_handle handle);
|
||||
|
||||
@@ -168,77 +158,6 @@ EJOC_API int EJOC_CALL ejoc_binaural_renderer_process(
|
||||
double output_gain,
|
||||
double* output_stereo_interleaved);
|
||||
|
||||
/*
|
||||
Native SOFA binaural renderer.
|
||||
|
||||
The handle owns the complete runtime: 64-QMF/77-hybrid analysis and synthesis,
|
||||
fifth-order ACN/N3D real spherical-harmonic direction-field evaluation,
|
||||
per-object whole-QMF-slot delay histories, six first-order image-source early
|
||||
reflections, the shared unitary-FDN late room, the 120-180 Hz LFE low-pass and
|
||||
the 961-sample latency compensation. The caller configures the filterbank
|
||||
tables, the compiled HRTF field and the room constants once, then per 512-sample
|
||||
block updates every source with ejoc_sofa_binaural_set_source() and calls
|
||||
ejoc_sofa_binaural_process(). Process returns the number of trimmed stereo
|
||||
samples written; the first 961 processed samples across calls are discarded.
|
||||
*/
|
||||
typedef void* ejoc_sofa_binaural_handle;
|
||||
|
||||
EJOC_API ejoc_sofa_binaural_handle EJOC_CALL ejoc_sofa_binaural_create(void);
|
||||
EJOC_API void EJOC_CALL ejoc_sofa_binaural_destroy(ejoc_sofa_binaural_handle handle);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_reset(ejoc_sofa_binaural_handle handle);
|
||||
EJOC_API const char* EJOC_CALL ejoc_sofa_binaural_last_error(
|
||||
ejoc_sofa_binaural_handle handle);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_configure_kernels(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
const double* qmf_analysis,
|
||||
const double* hybrid_low,
|
||||
const int16_t* hybrid_indices,
|
||||
const double* hybrid_values,
|
||||
uint32_t hybrid_count,
|
||||
const double* qmf_basis,
|
||||
const double* qmf_taps);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_configure_field(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
const double* coefficients,
|
||||
const double* delay_coefficients,
|
||||
const double* delay_bounds,
|
||||
const double* band_centers,
|
||||
double measurement_radius_m);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_configure_room(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
const double* room_dims,
|
||||
const double* listener_pos,
|
||||
const double* wall_gains,
|
||||
double speed_of_sound,
|
||||
const uint32_t* fdn_delays,
|
||||
const double* fdn_feedback,
|
||||
double damping,
|
||||
double fdn_output_gain,
|
||||
const uint32_t* allpass_delays,
|
||||
const double* allpass_gains,
|
||||
uint32_t enable_early_reflections,
|
||||
uint32_t enable_late_room);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_set_source(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
uint32_t source,
|
||||
const double* position_adm,
|
||||
uint32_t profile,
|
||||
double gain,
|
||||
uint32_t enabled,
|
||||
uint32_t special_lfe,
|
||||
uint32_t fade);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_process(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
const double* input16_interleaved,
|
||||
uint32_t sample_count,
|
||||
double output_gain,
|
||||
double* output_stereo_interleaved);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_finish(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
uint32_t flush_samples,
|
||||
double* output_stereo_interleaved,
|
||||
uint32_t capacity);
|
||||
|
||||
#ifdef __cplusplus
|
||||
}
|
||||
#endif
|
||||
|
||||
@@ -125,10 +125,6 @@ public:
|
||||
return error_[0] ? error_ : "";
|
||||
}
|
||||
|
||||
void set_dc_filter(const bool enabled) noexcept {
|
||||
dc_filter_enabled_ = enabled;
|
||||
}
|
||||
|
||||
int process(
|
||||
const float* bed5,
|
||||
const float* lfe,
|
||||
@@ -309,18 +305,16 @@ private:
|
||||
for (int i = 0; i < 4; ++i) {
|
||||
dc_buffer[20 + i] = current[i][0];
|
||||
}
|
||||
if (dc_filter_enabled_) {
|
||||
for (int slot = 0; slot < 4; ++slot) {
|
||||
Complex sum{0.0, 0.0};
|
||||
for (int tap = 0; tap < 21; ++tap) {
|
||||
const Complex sample = dc_buffer[slot + tap];
|
||||
const double cr = kDcB[tap];
|
||||
const double ci = kDcA[tap];
|
||||
sum.re += sample.re * cr - sample.im * ci;
|
||||
sum.im += sample.re * ci + sample.im * cr;
|
||||
}
|
||||
x_[channel][0][group + slot] = {2.0 * sum.re, 2.0 * sum.im};
|
||||
for (int slot = 0; slot < 4; ++slot) {
|
||||
Complex sum{0.0, 0.0};
|
||||
for (int tap = 0; tap < 21; ++tap) {
|
||||
const Complex sample = dc_buffer[slot + tap];
|
||||
const double cr = kDcB[tap];
|
||||
const double ci = kDcA[tap];
|
||||
sum.re += sample.re * cr - sample.im * ci;
|
||||
sum.im += sample.re * ci + sample.im * cr;
|
||||
}
|
||||
x_[channel][0][group + slot] = {2.0 * sum.re, 2.0 * sum.im};
|
||||
}
|
||||
for (int i = 0; i < 20; ++i) {
|
||||
surround_history_[surround][i] = dc_buffer[i + 4];
|
||||
@@ -647,8 +641,6 @@ private:
|
||||
float analysis_phase_;
|
||||
alignas(64) Complex surround_delay_[2][10][64];
|
||||
alignas(64) Complex surround_history_[2][20];
|
||||
// band-0 的 21-tap DC 补偿开关;仅 downmix 配置 3/4 由调用方置位。
|
||||
bool dc_filter_enabled_ = true;
|
||||
alignas(64) double lfe_delay_[kLfeDelay];
|
||||
alignas(64) double matrix_previous_[15][5][64];
|
||||
alignas(64) double synthesis_state_[15][640];
|
||||
@@ -704,14 +696,6 @@ int EJOC_CALL ejoc_renderer_set_threads(ejoc_renderer_handle handle, uint32_t to
|
||||
return static_cast<ejoc::Renderer*>(handle)->set_threads(total_threads);
|
||||
}
|
||||
|
||||
int EJOC_CALL ejoc_renderer_set_dc_filter(ejoc_renderer_handle handle, uint32_t enabled) {
|
||||
if (!handle) {
|
||||
return -1;
|
||||
}
|
||||
static_cast<ejoc::Renderer*>(handle)->set_dc_filter(enabled != 0);
|
||||
return 0;
|
||||
}
|
||||
|
||||
uint32_t EJOC_CALL ejoc_renderer_thread_count(ejoc_renderer_handle handle) {
|
||||
if (!handle) {
|
||||
return 0;
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,3 +1 @@
|
||||
numpy>=1.24
|
||||
scipy>=1.10
|
||||
h5py>=3.8
|
||||
|
||||
+9
-17
@@ -234,21 +234,11 @@ def build_dbmd(object_count=25, joc_binaural_mode=4):
|
||||
return bytes(out)
|
||||
|
||||
class Sink25:
|
||||
"""RF64 ADM BWF writer.
|
||||
|
||||
Header layout is fixed so that sizes can be patched without rereading the
|
||||
file: RF64+size+WAVE (12) + ds64 chunk (8+28) + fmt chunk (8+16) + data
|
||||
chunk header (8). Sizes beyond 32 bits follow the RF64 convention: the
|
||||
chunk size field holds 0xFFFFFFFF and the true value lives in ds64.
|
||||
"""
|
||||
_DS64_BODY_OFFSET = 20
|
||||
_DATA_SIZE_OFFSET = 76
|
||||
|
||||
def __init__(self, path, channels, rate):
|
||||
self.ch = channels; self.rate = rate; self.frames = 0
|
||||
self.fp = open(path, "wb+")
|
||||
self.fp.write(b"RF64" + struct.pack("<I", 0xFFFFFFFF) + b"WAVE")
|
||||
self._chunk(b"ds64", b"\x00" * 28)
|
||||
self._chunk(b"ds64", b"\x00" * 64)
|
||||
self._chunk(b"fmt ", self._fmt())
|
||||
self._chunk(b"data", b"")
|
||||
def _chunk(self, cid, body):
|
||||
@@ -269,12 +259,14 @@ class Sink25:
|
||||
self._chunk(b"chna", chna_bytes)
|
||||
self._chunk(b"dbmd", dbmd_bytes)
|
||||
self.fp.seek(0, 2); total = self.fp.tell()
|
||||
# RF64: 超过 32-bit 的 chunk size 字段写 0xFFFFFFFF,真实大小回填 ds64。
|
||||
self.fp.seek(self._DATA_SIZE_OFFSET)
|
||||
self.fp.write(struct.pack(
|
||||
"<I", data_len if data_len <= 0xFFFFFFFF else 0xFFFFFFFF))
|
||||
self.fp.seek(self._DS64_BODY_OFFSET)
|
||||
self.fp.write(struct.pack("<QQQI", total - 8, data_len, self.frames, 0))
|
||||
self.fp.seek(0); head = self.fp.read()
|
||||
m = head.find(b"data")
|
||||
if m >= 0:
|
||||
self.fp.seek(m + 4); self.fp.write(struct.pack("<I", data_len))
|
||||
m = head.find(b"ds64")
|
||||
if m >= 0:
|
||||
self.fp.seek(m + 8)
|
||||
self.fp.write(struct.pack("<QQQI", total - 8, data_len, self.frames, 0))
|
||||
self.fp.flush()
|
||||
self.fp.close()
|
||||
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
"""Direct ID11/OAMD position scheduling for the binaural render path."""
|
||||
"""Direct ID11/OAMD position scheduling for the Rosella binaural path."""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
@@ -7,9 +7,10 @@ import numpy as np
|
||||
|
||||
from adm_atmos import q_to_adm_xyz
|
||||
from oamd_bits import JocFieldState, frame_update
|
||||
from oamd_tracks import align_metadata_sample
|
||||
from variant_error import UnsupportedVariantError
|
||||
|
||||
OAMD_UPDATE_QUANTUM_SAMPLES = 64
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PositionTransition:
|
||||
@@ -153,8 +154,11 @@ class OamdPositionTimeline:
|
||||
self.payload_count += 1
|
||||
return
|
||||
|
||||
effective_ramp = max(0, int(update["ramp_duration_samples"]))
|
||||
transition_start = align_metadata_sample(coded_event + object_delay)
|
||||
ramp_duration = int(update["ramp_duration_samples"])
|
||||
effective_ramp = max(0, ramp_duration - OAMD_UPDATE_QUANTUM_SAMPLES)
|
||||
transition_start = coded_event + object_delay
|
||||
if effective_ramp:
|
||||
transition_start += OAMD_UPDATE_QUANTUM_SAMPLES
|
||||
for index, target in enumerate(targets):
|
||||
if self.previous_targets[index] == target:
|
||||
continue
|
||||
|
||||
+173
-144
@@ -1,92 +1,134 @@
|
||||
"""JOC frame adapter for the public SOFA binaural backend."""
|
||||
"""Stateful binaural renderer for reconstructed JOC objects."""
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import math
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
from binaural_metadata import OamdPositionTimeline
|
||||
from public_filterbank import ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
from sofa_hrtf_field import (
|
||||
DEFAULT_HRTF_CACHE_DIR,
|
||||
DEFAULT_PROJECTION_RIDGE,
|
||||
DEFAULT_SH_RIDGE,
|
||||
from binaural_native_renderer import NativeBinauralDsp
|
||||
from rosella_core import RosellaRenderer
|
||||
from rosella_direct import BINAURAL_PROFILE_NAMES
|
||||
from rosella_filterbank import (
|
||||
DEFAULT_KERNEL_DATA,
|
||||
HybridAnalysis,
|
||||
HybridSynthesis,
|
||||
QmfAnalysis,
|
||||
QmfSynthesis,
|
||||
)
|
||||
|
||||
from rosella_model import RosellaModel, load_personalized_headphone
|
||||
|
||||
SAMPLE_RATE = 48000
|
||||
FRAME_SAMPLES = 1536
|
||||
BINAURAL_BLOCK_SAMPLES = 512
|
||||
ROSSELLA_BLOCK_SAMPLES = 512
|
||||
QMF_HOP_SAMPLES = 64
|
||||
BINAURAL_LATENCY_SAMPLES = ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
ROSSELLA_LATENCY_SAMPLES = 961
|
||||
SOURCE_CHANNELS = 16
|
||||
OUTPUT_CHANNELS = 2
|
||||
PROJECT_DIR = Path(__file__).resolve().parent.parent
|
||||
DEFAULT_HRTF_DIR = PROJECT_DIR / "HRTF"
|
||||
DEFAULT_SOFA_HRTF = DEFAULT_HRTF_DIR / "binaural.sofa"
|
||||
DEFAULT_PERSONALIZED_HEADPHONE = (
|
||||
PROJECT_DIR / "HRTF" / "binaural.personalized_headphone")
|
||||
|
||||
|
||||
def _resolve_hrtf_file(path: str | Path, suffix: str, label: str) -> Path:
|
||||
target = Path(path).expanduser().resolve()
|
||||
if target.suffix.lower() != suffix:
|
||||
raise ValueError(f"{label} must use the {suffix} extension: {target}")
|
||||
def _sha256_file(path: Path) -> str:
|
||||
digest = hashlib.sha256()
|
||||
with path.open("rb") as stream:
|
||||
for block in iter(lambda: stream.read(1 << 20), b""):
|
||||
digest.update(block)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def resolve_personalized_headphone(path: str | Path | None = None) -> Path:
|
||||
target = (DEFAULT_PERSONALIZED_HEADPHONE if path is None
|
||||
else Path(path).expanduser().resolve())
|
||||
if not target.is_file():
|
||||
raise FileNotFoundError(f"{label} not found: {target}")
|
||||
raise FileNotFoundError(
|
||||
f"未找到双耳模型:{target}\n"
|
||||
"请将兼容模型保存为 HRTF/binaural.personalized_headphone,"
|
||||
"或使用 --personalized-headphone PATH 指定文件。"
|
||||
)
|
||||
return target
|
||||
|
||||
|
||||
def resolve_sofa_hrtf(path: str | Path) -> Path:
|
||||
"""Resolve an explicitly selected public SOFA source."""
|
||||
return _resolve_hrtf_file(path, ".sofa", "SOFA HRTF")
|
||||
|
||||
|
||||
def resolve_compiled_hrtf_cache(path: str | Path) -> Path:
|
||||
"""Resolve an explicitly selected JOC compiled HRTF cache."""
|
||||
return _resolve_hrtf_file(path, ".jochrtf", "compiled HRTF cache")
|
||||
|
||||
|
||||
class SofaBinauralRenderer:
|
||||
"""Render interleaved LFE plus fifteen JOC objects to stereo.
|
||||
|
||||
The adapter owns frame buffering and sample-timed OAMD updates. The
|
||||
backend owns the 64-QMF/77-hybrid state, the 961-sample latency policy,
|
||||
per-object direct/early state, and the shared late room.
|
||||
"""
|
||||
class RosellaBinauralRenderer:
|
||||
"""Render interleaved LFE plus fifteen objects to stereo."""
|
||||
|
||||
def __init__(
|
||||
self, backend, *,
|
||||
self,
|
||||
personalized_headphone: str | Path | RosellaModel,
|
||||
*,
|
||||
mode: str = "mid",
|
||||
kernel_data: str | Path = DEFAULT_KERNEL_DATA,
|
||||
object_delay_samples: int = 1473,
|
||||
tail_seconds: float = 5.0,
|
||||
chunk_frames: int = 64):
|
||||
required_interface = (
|
||||
"source_count", "default_profile", "set_source", "process",
|
||||
"finish", "finish_output_capacity", "info")
|
||||
missing = [name for name in required_interface if not hasattr(backend, name)]
|
||||
if missing:
|
||||
raise TypeError(
|
||||
f"backend must implement the binaural backend interface; "
|
||||
f"missing: {', '.join(missing)}")
|
||||
if backend.source_count != SOURCE_CHANNELS:
|
||||
raise ValueError(f"JOC binaural backend must have {SOURCE_CHANNELS} sources")
|
||||
if backend.default_profile != str(mode).lower():
|
||||
raise ValueError("backend default profile does not match renderer mode")
|
||||
output_gain: float = 1.0,
|
||||
chunk_frames: int = 64,
|
||||
room_impulse_slots: int = 4096,
|
||||
backend: str = "python",
|
||||
native_library=None):
|
||||
if mode not in BINAURAL_PROFILE_NAMES:
|
||||
raise ValueError("binaural mode must be near, mid, or far")
|
||||
if int(object_delay_samples) < 0:
|
||||
raise ValueError("object_delay_samples must be non-negative")
|
||||
if not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0:
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
if float(tail_seconds) < 0.0:
|
||||
raise ValueError("tail_seconds must be non-negative")
|
||||
if int(chunk_frames) <= 0:
|
||||
raise ValueError("chunk_frames must be positive")
|
||||
if not math.isfinite(float(output_gain)):
|
||||
raise ValueError("output_gain must be finite")
|
||||
if backend not in ("auto", "native", "python"):
|
||||
raise ValueError("backend must be auto, native, or python")
|
||||
|
||||
self.backend = backend
|
||||
self.mode = str(mode).lower()
|
||||
if isinstance(personalized_headphone, RosellaModel):
|
||||
self.model = personalized_headphone
|
||||
self.model_path = Path(self.model.source_path)
|
||||
else:
|
||||
self.model_path = resolve_personalized_headphone(personalized_headphone)
|
||||
self.model = load_personalized_headphone(self.model_path)
|
||||
if self.model.sample_rate != SAMPLE_RATE:
|
||||
raise ValueError(
|
||||
f"Rosella model sample rate must be {SAMPLE_RATE}, got {self.model.sample_rate}")
|
||||
|
||||
self.mode = mode
|
||||
self.profile_index = BINAURAL_PROFILE_NAMES[mode]
|
||||
self.kernel_data = Path(kernel_data).expanduser().resolve()
|
||||
self.kernel_data_sha256 = _sha256_file(self.kernel_data)
|
||||
self.object_delay_samples = int(object_delay_samples)
|
||||
self.tail_seconds = float(tail_seconds)
|
||||
self.output_gain = np.float64(output_gain)
|
||||
self.chunk_frames = int(chunk_frames)
|
||||
self.chunk_samples = self.chunk_frames * FRAME_SAMPLES
|
||||
self.dsp_backend = getattr(backend, "dsp_backend", "python-sofa")
|
||||
|
||||
self.native_dsp = None
|
||||
self.backend_fallback = None
|
||||
if backend in ("auto", "native"):
|
||||
try:
|
||||
self.native_dsp = NativeBinauralDsp(
|
||||
self.model, library_path=native_library,
|
||||
kernel_data=self.kernel_data)
|
||||
except (AttributeError, OSError, RuntimeError) as exc:
|
||||
if backend == "native":
|
||||
raise RuntimeError(f"native binaural backend unavailable: {exc}") from exc
|
||||
self.backend_fallback = str(exc)
|
||||
if self.native_dsp is not None:
|
||||
self.dsp_backend = "native"
|
||||
self.qmf_analysis = None
|
||||
self.hybrid_analysis = None
|
||||
self.hybrid_synthesis = None
|
||||
self.qmf_synthesis = None
|
||||
self.core = RosellaRenderer(
|
||||
self.model, SOURCE_CHANNELS, create_room=False)
|
||||
else:
|
||||
self.dsp_backend = "python"
|
||||
self.qmf_analysis = QmfAnalysis(SOURCE_CHANNELS, self.kernel_data)
|
||||
self.hybrid_analysis = HybridAnalysis(SOURCE_CHANNELS, self.kernel_data)
|
||||
self.core = RosellaRenderer(
|
||||
self.model, SOURCE_CHANNELS,
|
||||
room_impulse_slots=room_impulse_slots)
|
||||
self.hybrid_synthesis = HybridSynthesis(OUTPUT_CHANNELS, self.kernel_data)
|
||||
self.qmf_synthesis = QmfSynthesis(OUTPUT_CHANNELS, self.kernel_data)
|
||||
self.timeline = OamdPositionTimeline(15)
|
||||
|
||||
self._input_buffer = np.empty(
|
||||
@@ -94,72 +136,18 @@ class SofaBinauralRenderer:
|
||||
self._buffer_used = 0
|
||||
self.input_samples = 0
|
||||
self.processed_input_samples = 0
|
||||
self.raw_output_samples = 0
|
||||
self.output_samples = 0
|
||||
self.finished = False
|
||||
self.metadata_block_updates = 0
|
||||
|
||||
@classmethod
|
||||
def from_sofa(
|
||||
cls, sofa: str | Path, *,
|
||||
mode: str = "mid",
|
||||
cache_policy: str = "memory",
|
||||
cache_dir: str | Path | None = DEFAULT_HRTF_CACHE_DIR,
|
||||
shell_radius_m: float = 1.0,
|
||||
projection_ridge: float = DEFAULT_PROJECTION_RIDGE,
|
||||
sh_ridge: float = DEFAULT_SH_RIDGE,
|
||||
object_delay_samples: int = 1473,
|
||||
tail_seconds: float = 5.0,
|
||||
output_gain: float = 1.0,
|
||||
chunk_frames: int = 64) -> "SofaBinauralRenderer":
|
||||
source = resolve_sofa_hrtf(sofa)
|
||||
backend = SofaBinauralBackend.from_sofa(
|
||||
source,
|
||||
source_count=SOURCE_CHANNELS,
|
||||
default_profile=mode,
|
||||
output_gain=output_gain,
|
||||
cache_policy=cache_policy,
|
||||
cache_dir=cache_dir,
|
||||
shell_radius_m=shell_radius_m,
|
||||
projection_ridge=projection_ridge,
|
||||
sh_ridge=sh_ridge)
|
||||
return cls(
|
||||
backend,
|
||||
mode=mode,
|
||||
object_delay_samples=object_delay_samples,
|
||||
tail_seconds=tail_seconds,
|
||||
chunk_frames=chunk_frames)
|
||||
|
||||
@classmethod
|
||||
def from_compiled_cache(
|
||||
cls, cache: str | Path, *,
|
||||
mode: str = "mid",
|
||||
object_delay_samples: int = 1473,
|
||||
tail_seconds: float = 5.0,
|
||||
output_gain: float = 1.0,
|
||||
chunk_frames: int = 64) -> "SofaBinauralRenderer":
|
||||
source = resolve_compiled_hrtf_cache(cache)
|
||||
backend = SofaBinauralBackend.from_compiled_cache(
|
||||
source,
|
||||
source_count=SOURCE_CHANNELS,
|
||||
default_profile=mode,
|
||||
output_gain=output_gain)
|
||||
return cls(
|
||||
backend,
|
||||
mode=mode,
|
||||
object_delay_samples=object_delay_samples,
|
||||
tail_seconds=tail_seconds,
|
||||
chunk_frames=chunk_frames)
|
||||
|
||||
@property
|
||||
def finish_capacity_samples(self) -> int:
|
||||
return self.backend.finish_output_capacity(self.tail_seconds)
|
||||
|
||||
def _append_input(self, samples: np.ndarray) -> list[np.ndarray]:
|
||||
outputs = []
|
||||
source = np.asarray(samples, dtype=np.float64)
|
||||
position = 0
|
||||
while position < len(source):
|
||||
count = min(self.chunk_samples - self._buffer_used, len(source) - position)
|
||||
count = min(self.chunk_samples - self._buffer_used,
|
||||
len(source) - position)
|
||||
self._input_buffer[self._buffer_used:self._buffer_used + count] = (
|
||||
source[position:position + count])
|
||||
self._buffer_used += count
|
||||
@@ -199,40 +187,67 @@ class SofaBinauralRenderer:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
return np.concatenate(chunks, axis=0) if len(chunks) > 1 else chunks[0]
|
||||
|
||||
def _set_block_parameters(self, sample: int) -> None:
|
||||
def _set_block_parameters(self, sample: int):
|
||||
positions = self.timeline.positions_at(sample)
|
||||
self.backend.set_source(
|
||||
0, (0.0, 1.0, 0.0), profile=self.mode,
|
||||
special_lfe=True)
|
||||
self.core.set_source(0, (0.0, 1.0, 0.0), special_lfe=True)
|
||||
for object_index in range(15):
|
||||
self.backend.set_source(
|
||||
object_index + 1,
|
||||
positions[object_index],
|
||||
profile=self.mode)
|
||||
self.core.set_source(
|
||||
object_index + 1, positions[object_index], self.profile_index)
|
||||
|
||||
def _process_samples(self, source: np.ndarray) -> np.ndarray:
|
||||
values = np.asarray(source, dtype=np.float64)
|
||||
if values.ndim != 2 or values.shape[1] != SOURCE_CHANNELS:
|
||||
raise ValueError(f"expected [samples,{SOURCE_CHANNELS}], got {values.shape}")
|
||||
if len(values) % BINAURAL_BLOCK_SAMPLES:
|
||||
if len(values) % ROSSELLA_BLOCK_SAMPLES:
|
||||
raise ValueError("binaural input must be divisible by 512 samples")
|
||||
outputs = []
|
||||
blocks = len(values) // ROSSELLA_BLOCK_SAMPLES
|
||||
block_base = self.processed_input_samples
|
||||
for start in range(0, len(values), BINAURAL_BLOCK_SAMPLES):
|
||||
sample = block_base + start
|
||||
self._set_block_parameters(sample)
|
||||
outputs.append(self.backend.process(
|
||||
values[start:start + BINAURAL_BLOCK_SAMPLES]))
|
||||
|
||||
if self.native_dsp is not None:
|
||||
stereo = np.empty((len(values), OUTPUT_CHANNELS), dtype=np.float64)
|
||||
for block in range(blocks):
|
||||
sample = block_base + block * ROSSELLA_BLOCK_SAMPLES
|
||||
self._set_block_parameters(sample)
|
||||
start = block * ROSSELLA_BLOCK_SAMPLES
|
||||
stop = start + ROSSELLA_BLOCK_SAMPLES
|
||||
stereo[start:stop] = self.native_dsp.process_block(
|
||||
values[start:stop], self.core.gains, self.core.room_sends,
|
||||
self.output_gain)
|
||||
else:
|
||||
hops = values.reshape(
|
||||
blocks, ROSSELLA_BLOCK_SAMPLES // QMF_HOP_SAMPLES,
|
||||
QMF_HOP_SAMPLES, SOURCE_CHANNELS,
|
||||
).transpose(0, 1, 3, 2).reshape(
|
||||
blocks * (ROSSELLA_BLOCK_SAMPLES // QMF_HOP_SAMPLES),
|
||||
SOURCE_CHANNELS, QMF_HOP_SAMPLES)
|
||||
hybrid = self.hybrid_analysis.process_chunk(
|
||||
self.qmf_analysis.process_chunk(hops))
|
||||
direct = np.empty((blocks * 8, OUTPUT_CHANNELS, 77), dtype=np.complex128)
|
||||
room_send = np.empty((blocks * 8, 77), dtype=np.complex128)
|
||||
for block in range(blocks):
|
||||
sample = block_base + block * ROSSELLA_BLOCK_SAMPLES
|
||||
self._set_block_parameters(sample)
|
||||
start = block * 8
|
||||
stop = start + 8
|
||||
direct[start:stop], room_send[start:stop] = (
|
||||
self.core.direct_and_send_static(hybrid[start:stop]))
|
||||
rendered = direct + self.core.room.process_chunk(room_send)
|
||||
time_bands = self.qmf_synthesis.process_chunk(
|
||||
self.hybrid_synthesis.process_chunk(rendered))
|
||||
stereo = time_bands.transpose(0, 2, 1).reshape(
|
||||
blocks * ROSSELLA_BLOCK_SAMPLES, OUTPUT_CHANNELS)
|
||||
stereo *= self.output_gain
|
||||
|
||||
skip = max(0, min(
|
||||
len(stereo), ROSSELLA_LATENCY_SAMPLES - self.raw_output_samples))
|
||||
self.raw_output_samples += len(stereo)
|
||||
self.processed_input_samples += len(values)
|
||||
nonempty = [value for value in outputs if len(value)]
|
||||
if not nonempty:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
output = np.concatenate(nonempty, axis=0)
|
||||
output = stereo[skip:]
|
||||
self.output_samples += len(output)
|
||||
return output
|
||||
|
||||
def finish(self) -> np.ndarray:
|
||||
"""Process pending source samples and drain early/late room state once."""
|
||||
"""Process pending source samples and preserve the configured room tail."""
|
||||
if self.finished:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
outputs: list[np.ndarray] = []
|
||||
@@ -240,33 +255,47 @@ class SofaBinauralRenderer:
|
||||
outputs.append(self._process_samples(
|
||||
self._input_buffer[:self._buffer_used]))
|
||||
self._buffer_used = 0
|
||||
outputs.append(self.backend.finish(tail_seconds=self.tail_seconds))
|
||||
flush_samples = math.ceil(
|
||||
(self.tail_seconds * SAMPLE_RATE
|
||||
+ ROSSELLA_LATENCY_SAMPLES + ROSSELLA_BLOCK_SAMPLES)
|
||||
/ ROSSELLA_BLOCK_SAMPLES) * ROSSELLA_BLOCK_SAMPLES
|
||||
while flush_samples:
|
||||
count = min(flush_samples, self.chunk_samples)
|
||||
zero = np.zeros((count, SOURCE_CHANNELS), dtype=np.float64)
|
||||
outputs.append(self._process_samples(zero))
|
||||
flush_samples -= count
|
||||
self.finished = True
|
||||
nonempty = [value for value in outputs if len(value)]
|
||||
if not nonempty:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
output = np.concatenate(nonempty, axis=0)
|
||||
self.output_samples += len(outputs[-1])
|
||||
return output
|
||||
return np.concatenate(nonempty, axis=0)
|
||||
|
||||
def close(self) -> None:
|
||||
def close(self):
|
||||
if self.native_dsp is not None:
|
||||
self.native_dsp.close()
|
||||
self.finished = True
|
||||
|
||||
@property
|
||||
def backend_info(self) -> dict:
|
||||
info = self.backend.info()
|
||||
info.update({
|
||||
"adapter": "JOC 1536-frame / 512-sample metadata",
|
||||
"dsp_backend": self.dsp_backend,
|
||||
return {
|
||||
"name": self.dsp_backend,
|
||||
"precision": "float64/complex128",
|
||||
"fallback_reason": self.backend_fallback,
|
||||
"library": (str(self.native_dsp.library_path)
|
||||
if self.native_dsp is not None else None),
|
||||
"model": str(self.model_path.resolve()),
|
||||
"model_coefficients": int(len(self.model.coefficients)),
|
||||
"model_coefficient_sha256": self.model.coefficient_sha256,
|
||||
"model_version": self.model.coefficient_version,
|
||||
"kernel_data": str(self.kernel_data),
|
||||
"kernel_data_sha256": self.kernel_data_sha256,
|
||||
"mode": self.mode,
|
||||
"latency_compensated_samples": BINAURAL_LATENCY_SAMPLES,
|
||||
"latency_compensated_samples": ROSSELLA_LATENCY_SAMPLES,
|
||||
"object_delay_samples": self.object_delay_samples,
|
||||
"tail_seconds": self.tail_seconds,
|
||||
"metadata_payloads": self.timeline.payload_count,
|
||||
"metadata_position_transitions": self.timeline.transition_count,
|
||||
"input_samples": self.input_samples,
|
||||
"source_samples_processed": self.processed_input_samples,
|
||||
"processed_samples_including_flush": self.processed_input_samples,
|
||||
"output_samples_before_tail_trim": self.output_samples,
|
||||
"thread_safe": False,
|
||||
})
|
||||
return info
|
||||
}
|
||||
|
||||
+44
-98
@@ -2,11 +2,11 @@
|
||||
|
||||
范围:
|
||||
- EMDF ID14 的 joc_header、joc_info 和 Huffman joc_data;
|
||||
- 差分还原得到 joc_mix_mtx_q(dense MTX 与 sparse IDX/VEC 两条语法);
|
||||
- 差分解码得到 joc_mix_mtx_q;
|
||||
- 去量化得到 joc_mix_mtx_dq;
|
||||
- 位流自洽验证(joc_data 后剩余 = padding_bits 0..7 + 可能 joc_ext_data)
|
||||
后续的时间插值、QMF/时域重建和 ``joc_clipgain`` 位于 ``renderer.py``。
|
||||
公式与符号定义见 ``docs/math.md`` 第 2 节。
|
||||
Sparse 分支仍缺少实际样本验证。
|
||||
"""
|
||||
from pathlib import Path
|
||||
|
||||
@@ -23,7 +23,6 @@ _HUFF_NAMES = (
|
||||
"joc_huff_code_7ch_pos_index_sparse",
|
||||
)
|
||||
|
||||
|
||||
def _load_huff_tables():
|
||||
with np.load(_TABLES_PATH) as tables:
|
||||
return {
|
||||
@@ -31,26 +30,18 @@ def _load_huff_tables():
|
||||
for name in _HUFF_NAMES
|
||||
}
|
||||
|
||||
|
||||
H = _load_huff_tables()
|
||||
|
||||
JOC_NUM_CHANNELS = {0: 5, 1: 7, 2: 7, 3: 5, 4: 7} # Table 33
|
||||
# band-0 的 21-tap DC 补偿只在 downmix 配置 3/4 下启用。
|
||||
DC_FILTER_DMX_CONFIGS = (3, 4)
|
||||
JOC_NUM_BANDS = {0: 1, 1: 3, 2: 5, 3: 7, 4: 9, 5: 12, 6: 15, 7: 23} # Table 35
|
||||
JOC_NUM_QUANT = {0: 96, 1: 192} # Table 51
|
||||
# dense 的量化零点就是 nquant/2;sparse 的递推起点比它高 2 个量化步。
|
||||
JOC_DENSE_OFFSET = {0: 48, 1: 96}
|
||||
JOC_SPARSE_OFFSET = {0: 50, 1: 100}
|
||||
|
||||
_PUBLIC_TABLE39_EXCERPT_UNUSED = { # 仅保留作表格差异说明;渲染映射在 joc_qmf.py。
|
||||
23: [0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22],
|
||||
}
|
||||
|
||||
class BR:
|
||||
"""MSB-first 位读取器;位置以载荷内的 bit offset 表示。"""
|
||||
|
||||
def __init__(self, data, pos=0):
|
||||
self.d = data
|
||||
self.p = pos
|
||||
|
||||
def bits(self, n):
|
||||
v = 0
|
||||
for _ in range(n):
|
||||
@@ -62,19 +53,18 @@ class BR:
|
||||
def huff_decode(tree, br):
|
||||
node = 0
|
||||
while node >= 0:
|
||||
node = tree[node][br.bits(1)]
|
||||
b = br.bits(1)
|
||||
node = tree[node][b]
|
||||
return -node - 1
|
||||
|
||||
|
||||
def get_huff_code(mode, typ, nch):
|
||||
if typ == "IDX":
|
||||
return H["joc_huff_code_5ch_pos_index_sparse" if nch == 5
|
||||
else "joc_huff_code_7ch_pos_index_sparse"]
|
||||
return H["joc_huff_code_5ch_pos_index_sparse" if nch == 5 else "joc_huff_code_7ch_pos_index_sparse"]
|
||||
if typ == "VEC":
|
||||
return H["joc_huff_code_coarse_coeff_sparse" if mode == 0
|
||||
else "joc_huff_code_fine_coeff_sparse"]
|
||||
return H["joc_huff_code_coarse_generic" if mode == 0
|
||||
else "joc_huff_code_fine_generic"] # MTX
|
||||
return H["joc_huff_code_coarse_coeff_sparse" if mode == 0 else "joc_huff_code_fine_coeff_sparse"]
|
||||
# MTX
|
||||
return H["joc_huff_code_coarse_generic" if mode == 0 else "joc_huff_code_fine_generic"]
|
||||
|
||||
|
||||
def parse_joc(payload):
|
||||
@@ -86,8 +76,6 @@ def parse_joc(payload):
|
||||
out["ext_config_idx"] = br.bits(3)
|
||||
n_objects = out["num_objects_bits"] + 1
|
||||
n_channels = JOC_NUM_CHANNELS.get(out["dmx_config_idx"])
|
||||
if n_channels is None:
|
||||
raise ValueError(f"未知的 JOC downmix 配置 {out['dmx_config_idx']}")
|
||||
out["n_objects"], out["n_channels"] = n_objects, n_channels
|
||||
out["clipgain_x_bits"] = br.bits(3)
|
||||
out["clipgain_y_bits"] = br.bits(5)
|
||||
@@ -110,120 +98,78 @@ def parse_joc(payload):
|
||||
o["offset_ts"] = [br.bits(5) + 1 for _ in range(o["n_dpoints"])]
|
||||
objs.append(o)
|
||||
out["objs"] = objs
|
||||
# joc_data(Huffman):dense 逐声道逐带读 MTX;sparse 读 IDX 后读 VEC。
|
||||
# joc_data(Huffman)
|
||||
for obj, o in enumerate(objs):
|
||||
if not o["present"]:
|
||||
continue
|
||||
nquant = 96 if o["quant_idx"] == 0 else 192
|
||||
o["channel_idx"] = []
|
||||
o["vec"] = []
|
||||
o["mtx"] = []
|
||||
for dp in range(o["n_dpoints"]):
|
||||
if o["sparse"] == 1:
|
||||
tree = get_huff_code(o["quant_idx"], "IDX", n_channels)
|
||||
idx = [br.bits(3)]
|
||||
idx += [huff_decode(tree, br) for _ in range(o["n_bands"] - 1)]
|
||||
# Sparse JOC 使用 VEC/IDX Huffman 树;此分支尚无真实码流验证。
|
||||
ci0 = br.bits(3)
|
||||
tree = get_huff_code(n_channels, "IDX", n_channels)
|
||||
ci = [ci0] + [huff_decode(tree, br) for _ in range(o["n_bands"] - 1)]
|
||||
o["channel_idx"].append(ci)
|
||||
tree = get_huff_code(o["quant_idx"], "VEC", n_channels)
|
||||
o["channel_idx"].append(idx)
|
||||
o["vec"].append([huff_decode(tree, br) for _ in range(o["n_bands"])])
|
||||
o["mtx"].append(None)
|
||||
vec = [huff_decode(tree, br) for _ in range(o["n_bands"])]
|
||||
o["vec"].append(vec)
|
||||
else:
|
||||
tree = get_huff_code(o["quant_idx"], "MTX", n_channels)
|
||||
mtx = [[huff_decode(tree, br) for _ in range(o["n_bands"])]
|
||||
for _ in range(n_channels)]
|
||||
o["mtx"].append(mtx)
|
||||
o["channel_idx"].append(None)
|
||||
o["vec"].append(None)
|
||||
out["data_end_bits"] = br.p
|
||||
out["remaining_bits"] = len(payload) * 8 - br.p
|
||||
out["tail_bytes"] = payload[br.p // 8:]
|
||||
return out
|
||||
|
||||
|
||||
def reconstruct_dense(o, dp, n_ch):
|
||||
"""Dense 差分还原 → 量化矩阵 ``[ch][pb]``。
|
||||
|
||||
每个核心声道各自从 ``nquant/2``(去量化 0)出发,沿参数带累加 MTX 符号。
|
||||
"""
|
||||
nquant = JOC_NUM_QUANT[o["quant_idx"]]
|
||||
offset = JOC_DENSE_OFFSET[o["quant_idx"]]
|
||||
mtx = o["mtx"][dp]
|
||||
if mtx is None:
|
||||
raise ValueError("Dense JOC 对象缺少 MTX 符号")
|
||||
q = np.zeros((n_ch, o["n_bands"]), dtype=np.int64)
|
||||
for ch in range(n_ch):
|
||||
q[ch][0] = (offset + mtx[ch][0]) % nquant
|
||||
for pb in range(1, o["n_bands"]):
|
||||
q[ch][pb] = (q[ch][pb - 1] + mtx[ch][pb]) % nquant
|
||||
return q
|
||||
|
||||
|
||||
def reconstruct_sparse(o, dp, n_ch):
|
||||
"""Sparse 差分还原 → 量化矩阵 ``[ch][pb]``。
|
||||
|
||||
每个参数带只有一个 active 声道:
|
||||
- ``active[0]`` 是 3 bit 绝对声道号,其后由 IDX 符号累加得到,
|
||||
因此递推锚点是**已重建**的 active 声道,而不是编码器送的符号本身;
|
||||
- 系数是一个跨参数带连续的单累加器(active 声道切换时**不**重置),
|
||||
起点为 sparse offset,增量为 VEC 符号;
|
||||
- 非 active 项取 ``nquant/2``,即去量化后的 0。
|
||||
|
||||
IDX 符号取值恒为 ``0..n_ch-1``(Huffman 叶数即声道数),
|
||||
故 ``(active + idx) % n_ch`` 与单次条件减等价。
|
||||
"""
|
||||
if n_ch not in (5, 7):
|
||||
raise ValueError(f"Sparse JOC 需要 5 或 7 个核心声道,实际 {n_ch}")
|
||||
nquant = JOC_NUM_QUANT[o["quant_idx"]]
|
||||
offset = JOC_SPARSE_OFFSET[o["quant_idx"]]
|
||||
n_bands = o["n_bands"]
|
||||
idx = o["channel_idx"][dp]
|
||||
vec = o["vec"][dp]
|
||||
if idx is None or vec is None:
|
||||
raise ValueError("Sparse JOC 对象缺少 channel_idx/vec 符号")
|
||||
if len(idx) != n_bands or len(vec) != n_bands:
|
||||
raise ValueError(
|
||||
f"Sparse JOC 维度不符: bands={n_bands}, idx={len(idx)}, vec={len(vec)}")
|
||||
if not 0 <= idx[0] < n_ch:
|
||||
raise ValueError(f"Sparse JOC 初始声道 {idx[0]} 超出 {n_ch} 个声道")
|
||||
|
||||
q = np.full((n_ch, n_bands), nquant // 2, dtype=np.int64)
|
||||
active = idx[0]
|
||||
coefficient = offset
|
||||
for pb in range(n_bands):
|
||||
if pb:
|
||||
active = (active + idx[pb]) % n_ch
|
||||
coefficient = (coefficient + vec[pb]) % nquant
|
||||
q[active][pb] = coefficient
|
||||
return q
|
||||
|
||||
|
||||
def diff_decode(out):
|
||||
"""差分还原 → joc_mix_mtx_q[obj][dp][ch][pb]。"""
|
||||
"""6.6.2:差分解码 → joc_mix_mtx_q[obj][dp][ch][pb]。"""
|
||||
mix_q = {}
|
||||
n_ch = out["n_channels"]
|
||||
for obj, o in enumerate(out["objs"]):
|
||||
if not o["present"]:
|
||||
continue
|
||||
nquant = 96 if o["quant_idx"] == 0 else 192
|
||||
q = np.zeros((o["n_dpoints"], n_ch, o["n_bands"]), dtype=np.int64)
|
||||
for dp in range(o["n_dpoints"]):
|
||||
if o["sparse"] == 1:
|
||||
q[dp] = reconstruct_sparse(o, dp, n_ch)
|
||||
# Sparse 差分路径尚无真实码流验证。
|
||||
offset = 50 if o["quant_idx"] == 0 else 100
|
||||
ci = o["channel_idx"][dp]
|
||||
vec = o["vec"][dp]
|
||||
for pb in range(o["n_bands"]):
|
||||
ci_mod = ci[0] if pb == 0 else (ci[pb - 1] + ci[pb]) % n_ch
|
||||
for ch in range(n_ch):
|
||||
if ch == ci_mod:
|
||||
if pb == 0:
|
||||
q[dp][ch][pb] = (offset + vec[pb]) % nquant
|
||||
else:
|
||||
q[dp][ch][pb] = (q[dp][ch][pb - 1] + vec[pb]) % nquant
|
||||
else:
|
||||
q[dp][ch][pb] = offset
|
||||
else:
|
||||
q[dp] = reconstruct_dense(o, dp, n_ch)
|
||||
offset = 48 if o["quant_idx"] == 0 else 96
|
||||
mtx = o["mtx"][dp]
|
||||
for ch in range(n_ch):
|
||||
q[dp][ch][0] = (offset + mtx[ch][0]) % nquant
|
||||
for pb in range(1, o["n_bands"]):
|
||||
q[dp][ch][pb] = (q[dp][ch][pb - 1] + mtx[ch][pb]) % nquant
|
||||
mix_q[obj] = q
|
||||
return mix_q
|
||||
|
||||
|
||||
def dequantize(out, mix_q):
|
||||
"""去量化 → joc_mix_mtx_dq。
|
||||
|
||||
Sparse 的非 active 项在 ``joc_mix_mtx_q`` 中取 ``nquant/2``,
|
||||
因此与 dense 共用同一条去量化公式即得到 0。
|
||||
"""
|
||||
"""6.6.4:去量化 → joc_mix_mtx_dq。"""
|
||||
mix_dq = {}
|
||||
for obj, o in enumerate(out["objs"]):
|
||||
if not o["present"]:
|
||||
continue
|
||||
nquant = JOC_NUM_QUANT[o["quant_idx"]]
|
||||
nquant = 96 if o["quant_idx"] == 0 else 192
|
||||
q = mix_q[obj]
|
||||
dq = (q.astype(np.float64) - nquant / 2) * 820 / (4096 * (1 + o["quant_idx"]))
|
||||
mix_dq[obj] = dq
|
||||
|
||||
+4
-9
@@ -130,12 +130,8 @@ _SURROUND_DC_B = np.array([
|
||||
_SURROUND_DC_C = _SURROUND_DC_B + 1j * _SURROUND_DC_A
|
||||
|
||||
|
||||
def surround_post_frame(x, delay, dc_hist, apply_dc_filter=True):
|
||||
"""处理 Ls/Rs 的 10 槽延迟、-j 旋转和 band-0 FIR。
|
||||
|
||||
``apply_dc_filter=False`` 时跳过 band-0 的 21-tap DC 补偿,只做延迟与
|
||||
-j 旋转;延迟线与 DC 历史仍照常推进,便于逐帧切换。
|
||||
"""
|
||||
def surround_post_frame(x, delay, dc_hist):
|
||||
"""处理 Ls/Rs 的 10 槽延迟、-j 旋转和 band-0 FIR。"""
|
||||
src = np.asarray(x, dtype=np.complex128)
|
||||
qdelay = np.asarray(delay, dtype=np.complex128).copy()
|
||||
hist = np.asarray(dc_hist, dtype=np.complex128).copy()
|
||||
@@ -148,9 +144,8 @@ def surround_post_frame(x, delay, dc_hist, apply_dc_filter=True):
|
||||
block = -1j * queued[:, :4, :]
|
||||
qdelay = queued[:, 4:, :]
|
||||
dc_buf = np.concatenate((hist, current[:, :, 0]), axis=1)
|
||||
if apply_dc_filter:
|
||||
windows = np.lib.stride_tricks.sliding_window_view(dc_buf, 21, axis=1)
|
||||
block[:, :, 0] = 2.0 * np.sum(windows * _SURROUND_DC_C[None, None, :], axis=2)
|
||||
windows = np.lib.stride_tricks.sliding_window_view(dc_buf, 21, axis=1)
|
||||
block[:, :, 0] = 2.0 * np.sum(windows * _SURROUND_DC_C[None, None, :], axis=2)
|
||||
hist = dc_buf[:, 4:]
|
||||
out[:, :, group:group + 4] = block.transpose(0, 2, 1)
|
||||
return out, qdelay, hist
|
||||
|
||||
@@ -246,6 +246,20 @@ def inspect(index, limit=None, print_frames=False):
|
||||
"parser_error": str(exc),
|
||||
"repair_hint": "检查 JOC header、对象数、参数带、Huffman 或扩展字段",
|
||||
}) from exc
|
||||
sparse = [i for i, obj in enumerate(parsed["objs"]) if obj["present"] and obj["sparse"]]
|
||||
if sparse:
|
||||
raise UnsupportedVariantError(
|
||||
"joc", "sparse_joc",
|
||||
"发现尚未验证的 Sparse JOC 帧",
|
||||
frame=frame_number,
|
||||
details={
|
||||
"sparse_objects": sparse,
|
||||
"downmix_config": parsed["dmx_config_idx"],
|
||||
"extension_config": parsed["ext_config_idx"],
|
||||
"objects": parsed["n_objects"],
|
||||
"payload": bytes_descriptor(subs[14]),
|
||||
"repair_hint": "需要 Sparse JOC 实际样本及对应输出建立回归后再启用",
|
||||
})
|
||||
if parsed["n_channels"] != 5 or parsed["n_objects"] > 15:
|
||||
raise UnsupportedVariantError(
|
||||
"joc", "unsupported_configuration",
|
||||
|
||||
+3
-20
@@ -14,7 +14,7 @@ import sys
|
||||
import numpy as np
|
||||
|
||||
from evo_unpack import unpack_evolution
|
||||
from joc_decode import DC_FILTER_DMX_CONFIGS, dequantize, diff_decode, parse_joc
|
||||
from joc_decode import dequantize, diff_decode, parse_joc
|
||||
|
||||
FRAME_SAMPLES = 1536
|
||||
MAX_OBJECTS = 15
|
||||
@@ -88,9 +88,6 @@ def _load_library(path):
|
||||
lib.ejoc_renderer_reset.restype = ctypes.c_int
|
||||
lib.ejoc_renderer_set_threads.argtypes = [ctypes.c_void_p, ctypes.c_uint32]
|
||||
lib.ejoc_renderer_set_threads.restype = ctypes.c_int
|
||||
if hasattr(lib, "ejoc_renderer_set_dc_filter"):
|
||||
lib.ejoc_renderer_set_dc_filter.argtypes = [ctypes.c_void_p, ctypes.c_uint32]
|
||||
lib.ejoc_renderer_set_dc_filter.restype = ctypes.c_int
|
||||
lib.ejoc_renderer_thread_count.argtypes = [ctypes.c_void_p]
|
||||
lib.ejoc_renderer_thread_count.restype = ctypes.c_uint32
|
||||
lib.ejoc_renderer_last_error.argtypes = [ctypes.c_void_p]
|
||||
@@ -211,6 +208,8 @@ class NativeJocRenderer:
|
||||
for object_index, info in enumerate(out["objs"]):
|
||||
if not info["present"]:
|
||||
continue
|
||||
if info["sparse"]:
|
||||
raise ValueError("native core does not accept unvalidated Sparse JOC")
|
||||
bands = int(info["n_bands"])
|
||||
points = int(info["n_dpoints"])
|
||||
if bands > MAX_BANDS or points > MAX_DPOINTS:
|
||||
@@ -228,21 +227,6 @@ class NativeJocRenderer:
|
||||
self._dq[object_index, :points, :, :bands] = values
|
||||
return mask
|
||||
|
||||
def _set_dc_filter(self, dmx_config_idx):
|
||||
"""band-0 的 21-tap DC 补偿只在 downmix 配置 3/4 下启用。"""
|
||||
enabled = dmx_config_idx in DC_FILTER_DMX_CONFIGS
|
||||
setter = getattr(self._lib, "ejoc_renderer_set_dc_filter", None)
|
||||
if setter is None:
|
||||
if not enabled:
|
||||
raise RuntimeError(
|
||||
"native library has no ejoc_renderer_set_dc_filter; rebuild "
|
||||
"eac3joc_core to disable the band-0 DC filter for "
|
||||
f"dmx_config_idx={dmx_config_idx}")
|
||||
return
|
||||
result = setter(self._handle, 1 if enabled else 0)
|
||||
if result != 0:
|
||||
self._raise_native("set_dc_filter", result)
|
||||
|
||||
def render_frame(self, payload_bytes, bed5_pcm, lfe_pcm=None):
|
||||
subs, _ = unpack_evolution(payload_bytes, loose=True)
|
||||
return self.render_subpayloads(subs, bed5_pcm, lfe_pcm)
|
||||
@@ -251,7 +235,6 @@ class NativeJocRenderer:
|
||||
self._require_open()
|
||||
out, _, mix_dq = self.decode_subpayloads(subs)
|
||||
object_mask = self._pack_frame(out, mix_dq)
|
||||
self._set_dc_filter(out["dmx_config_idx"])
|
||||
bed5 = np.ascontiguousarray(bed5_pcm, dtype=np.float32)
|
||||
if bed5.shape != (CORE_CHANNELS, FRAME_SAMPLES):
|
||||
raise ValueError(f"core PCM shape must be (5,1536), got {bed5.shape}")
|
||||
|
||||
+232
-419
@@ -1,14 +1,12 @@
|
||||
"""OAMD 位载荷 → 16 个对象槽的 q1/q2/q3 增量状态。
|
||||
|
||||
依据 ETSI TS 103 420 V1.2.1 clause 5。槽 0 为 bed/LFE,槽 1..15 为输出
|
||||
ch1..15;bed/ISF/未激活对象没有位置字段,保持上一帧位置。element 目录按声明长度
|
||||
驱动,未知 element 按边界跳过;个别编码器的 ``oa_element_size`` 比实际内容短时以
|
||||
结构解析为准,差异记入 ``diagnostics``,不算变体错误。
|
||||
槽 0 是 bed/LFE;槽 1..15 对应输出 ch1..15 的对象元数据。
|
||||
"""
|
||||
import numpy as np
|
||||
|
||||
from variant_error import UnsupportedVariantError, bytes_descriptor
|
||||
|
||||
Q1_OFF = 192
|
||||
N_Q12 = 62
|
||||
N_Q3 = 15
|
||||
SAMPLE_OFFSET_INDEX = (8, 16, 18, 24)
|
||||
@@ -18,20 +16,9 @@ RAMP_DURATION_INDEX = (
|
||||
1024, 1600, 1601, 1602, 1920, 2000, 2002, 2048,
|
||||
)
|
||||
OBJECT_ELEMENT_ID = 1
|
||||
# 5.6.0.11 Table 11b:ISF 类型 → 对象数。
|
||||
ISF_OBJECT_COUNTS = (4, 8, 10, 14, 15, 30)
|
||||
# 5.6.1.1.4 Table 12:10 bit 标准 bed 掩码按位对应的声道(LSB = RC_L/RC_R)。
|
||||
BED_CHANNEL_LABELS = (
|
||||
"RC_L/RC_R", "RC_C", "RC_LFE", "RC_LS/RC_RS", "RC_LB/RC_RB",
|
||||
"RC_TFL/RC_TFR", "RC_TSL/RC_TSR", "RC_TBL/RC_TBR", "RC_LW/RC_RW", "RC_LFE2",
|
||||
)
|
||||
# 5.6.1.1.5 Table 13:17 bit 非标准 bed 掩码按位对应的声道标签。
|
||||
NONSTD_BED_LABELS = (
|
||||
"RC_LFE2", "RC_RW", "RC_LW", "RC_TBR", "RC_TBL", "RC_TSR", "RC_TSL",
|
||||
"RC_TFR", "RC_TFL", "RC_RB", "RC_LB", "RC_RS", "RC_LS", "RC_LFE", "RC_C",
|
||||
"RC_R", "RC_L",
|
||||
)
|
||||
MAX_SLOTS = 16
|
||||
TRIM_ELEMENT_ID = 2
|
||||
EXTENDED_OBJECT_ELEMENT_ID = 5
|
||||
POSITION_WINDOW_END_BIT = 112 + 31 * (15 - 3) + 24
|
||||
|
||||
|
||||
def q_of(k, n):
|
||||
@@ -104,23 +91,14 @@ class _BitReader:
|
||||
self.read(count)
|
||||
|
||||
|
||||
def _variable_bits_max(reader, width, max_groups):
|
||||
"""5.5.1 ``variable_bits_max(n, max_num_groups)``。"""
|
||||
value = reader.read(width)
|
||||
more = reader.read(1)
|
||||
num_group = 1
|
||||
if max_groups > num_group:
|
||||
if more:
|
||||
value = (value + 1) << width
|
||||
while more:
|
||||
value += reader.read(width)
|
||||
more = reader.read(1)
|
||||
if num_group >= max_groups:
|
||||
break
|
||||
if more:
|
||||
value = (value + 1) << width
|
||||
num_group += 1
|
||||
return value
|
||||
def _variable_bits(reader, width, max_groups=5):
|
||||
value = 0
|
||||
for _ in range(max_groups + 1):
|
||||
value += reader.read(width)
|
||||
if not reader.read(1):
|
||||
return value
|
||||
value = (value + 1) << width
|
||||
raise ValueError(f"OAMD variable_bits({width}) 延伸组过多")
|
||||
|
||||
|
||||
def _element_details(elements):
|
||||
@@ -131,185 +109,166 @@ def _element_details(elements):
|
||||
"header_start_bit": element["header_start_bit"],
|
||||
"body_start_bit": element["body_start_bit"],
|
||||
"body_end_bit": element["body_end_bit"],
|
||||
"parsed_end_bit": element.get("parsed_end_bit"),
|
||||
"discard_unknown": element["discard_unknown"],
|
||||
"alternate_data_id": element["alternate_data_id"],
|
||||
} for element in elements]
|
||||
|
||||
|
||||
def _syntax_error(variant, message, raw_payload, exc=None, **details):
|
||||
payload = {"payload": bytes_descriptor(raw_payload)}
|
||||
if exc is not None:
|
||||
payload["parser_error"] = str(exc)
|
||||
payload.update(details)
|
||||
return UnsupportedVariantError("oamd", variant, message, details=payload)
|
||||
def _parse_elements(bits, header, header_end, raw_payload):
|
||||
"""建立 OAMD element 目录;element size 表示其后 body 的字节数。"""
|
||||
reader = _BitReader(bits, header_end)
|
||||
elements = []
|
||||
for ordinal in range(header["element_count"]):
|
||||
header_start = reader.position
|
||||
try:
|
||||
element_id = reader.read(4)
|
||||
size_bytes = _variable_bits(reader, 4) + 1
|
||||
except ValueError as exc:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "element_header_truncated",
|
||||
"OAMD element header 不完整",
|
||||
details={
|
||||
"element_ordinal": ordinal,
|
||||
"header_start_bit": header_start,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"parser_error": str(exc),
|
||||
}) from exc
|
||||
|
||||
|
||||
def _parse_program_assignment(reader, raw_payload):
|
||||
"""5.5.3 ``program_assignment()``:bed / ISF / dynamic 对象数量。"""
|
||||
program = {
|
||||
"dynamic_object_only": bool(reader.read(1)),
|
||||
"lfe_present": False,
|
||||
"content_description": None,
|
||||
"bed_assignments": [],
|
||||
"num_bed_objects": 0,
|
||||
"isf_idx": None,
|
||||
"num_isf_objects": 0,
|
||||
"num_dynamic_objects": None,
|
||||
}
|
||||
if program["dynamic_object_only"]:
|
||||
program["lfe_present"] = bool(reader.read(1))
|
||||
# 5.6.4.8:对象顺序为 bed → ISF → dynamic,LFE 属于 bed,排在最前。
|
||||
program["num_bed_objects"] = 1 if program["lfe_present"] else 0
|
||||
return program
|
||||
|
||||
mask = reader.read(4)
|
||||
program["content_description"] = {
|
||||
"reserved": bool(mask & 0x8),
|
||||
"dynamic": bool(mask & 0x4),
|
||||
"isf": bool(mask & 0x2),
|
||||
"bed": bool(mask & 0x1),
|
||||
}
|
||||
if mask & 0x1:
|
||||
program["b_bed_chan_distribute"] = bool(reader.read(1))
|
||||
num_instances = (reader.read(3) + 2) if reader.read(1) else 1
|
||||
for instance in range(num_instances):
|
||||
if reader.read(1): # b_lfe_only
|
||||
program["bed_assignments"].append({
|
||||
"instance": instance, "lfe_only": True, "channels": ["RC_LFE"],
|
||||
body_start = reader.position
|
||||
body_end = body_start + size_bytes * 8
|
||||
if body_end > len(bits):
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "element_bounds",
|
||||
"OAMD element 声明长度超过 payload 边界",
|
||||
details={
|
||||
"element_ordinal": ordinal,
|
||||
"element_id": element_id,
|
||||
"size_bytes": size_bytes,
|
||||
"body_start_bit": body_start,
|
||||
"body_end_bit": body_end,
|
||||
"payload_bits": len(bits),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
continue
|
||||
if reader.read(1): # b_standard_chan_assign
|
||||
bed_mask = reader.read(10)
|
||||
channels = []
|
||||
for bit in range(10):
|
||||
if (bed_mask >> bit) & 1:
|
||||
channels.extend(BED_CHANNEL_LABELS[bit].split("/"))
|
||||
program["bed_assignments"].append({
|
||||
"instance": instance, "lfe_only": False, "standard": True,
|
||||
"mask": bed_mask, "channels": channels,
|
||||
|
||||
control_bits = 5 if header["alternate_object_present"] else 1
|
||||
if body_start + control_bits > body_end:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "element_control_bounds",
|
||||
"OAMD element 太短,无法容纳控制字段",
|
||||
details={
|
||||
"element_ordinal": ordinal,
|
||||
"element_id": element_id,
|
||||
"size_bytes": size_bytes,
|
||||
"control_bits": control_bits,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
else:
|
||||
bed_mask = reader.read(17)
|
||||
channels = [NONSTD_BED_LABELS[bit]
|
||||
for bit in range(17) if (bed_mask >> bit) & 1]
|
||||
program["bed_assignments"].append({
|
||||
"instance": instance, "lfe_only": False, "standard": False,
|
||||
"mask": bed_mask, "channels": channels,
|
||||
control = _BitReader(bits, body_start, body_end)
|
||||
alternate_data_id = (control.read(4)
|
||||
if header["alternate_object_present"] else None)
|
||||
discard_unknown = bool(control.read(1))
|
||||
elements.append({
|
||||
"ordinal": ordinal,
|
||||
"element_id": element_id,
|
||||
"size_bytes": size_bytes,
|
||||
"header_start_bit": header_start,
|
||||
"body_start_bit": body_start,
|
||||
"data_start_bit": control.position,
|
||||
"body_end_bit": body_end,
|
||||
"alternate_data_id": alternate_data_id,
|
||||
"discard_unknown": discard_unknown,
|
||||
})
|
||||
reader.position = body_end
|
||||
|
||||
padding = bits[reader.position:]
|
||||
if len(padding) > 7:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "trailing_payload_data",
|
||||
"OAMD element 结束后仍有超过一个字节的未声明数据",
|
||||
details={
|
||||
"elements": _element_details(elements),
|
||||
"trailing_bits": len(padding),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
if np.any(padding):
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "nonzero_padding",
|
||||
"OAMD payload 尾部 padding 含非零位",
|
||||
details={
|
||||
"elements": _element_details(elements),
|
||||
"padding_start_bit": reader.position,
|
||||
"padding_bits": "".join(str(int(bit)) for bit in padding),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
|
||||
object_elements = [element for element in elements
|
||||
if element["element_id"] == OBJECT_ELEMENT_ID]
|
||||
if len(object_elements) != 1:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "object_element_count",
|
||||
"OAMD 必须包含且只能包含一个 object element",
|
||||
details={
|
||||
"object_element_count": len(object_elements),
|
||||
"elements": _element_details(elements),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
object_element = object_elements[0]
|
||||
if object_element["ordinal"] != 0 or object_element["header_start_bit"] != header_end:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "object_element_order",
|
||||
"object element 不在当前固定位置窗口支持的首个 element 位置",
|
||||
details={
|
||||
"elements": _element_details(elements),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"repair_hint": "以 object element 的实际位置为基准重新定位对象窗口",
|
||||
})
|
||||
|
||||
for element in elements:
|
||||
element_id = element["element_id"]
|
||||
if element_id in (OBJECT_ELEMENT_ID, TRIM_ELEMENT_ID):
|
||||
continue
|
||||
if element_id == EXTENDED_OBJECT_ELEMENT_ID:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "extended_object_element",
|
||||
"OAMD 含可能改变坐标语义的 extended object element",
|
||||
details={
|
||||
"element": _element_details([element])[0],
|
||||
"elements": _element_details(elements),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"repair_hint": "解析 divergence/extended-precision position 后再应用轨迹",
|
||||
})
|
||||
program["num_bed_objects"] = sum(
|
||||
len(assignment["channels"]) for assignment in program["bed_assignments"])
|
||||
if mask & 0x2:
|
||||
isf_idx = reader.read(3)
|
||||
program["isf_idx"] = isf_idx
|
||||
program["num_isf_objects"] = (ISF_OBJECT_COUNTS[isf_idx]
|
||||
if isf_idx < len(ISF_OBJECT_COUNTS) else 0)
|
||||
if mask & 0x4:
|
||||
count = reader.read(5)
|
||||
if count == 0x1F:
|
||||
count += reader.read(7)
|
||||
program["num_dynamic_objects"] = count + 1
|
||||
if mask & 0x8:
|
||||
# 5.6.4.1:reserved_data_size = reserved_data_size_bits + 1(字节)。
|
||||
reader.skip((reader.read(4) + 1) * 8)
|
||||
return program
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", f"unsupported_element_{element_id}",
|
||||
f"OAMD 含当前未覆盖的 element id {element_id}",
|
||||
details={
|
||||
"element": _element_details([element])[0],
|
||||
"elements": _element_details(elements),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
return object_element, elements
|
||||
|
||||
|
||||
def _parse_object_info_block(reader, object_index, in_bed_or_isf, raw_payload):
|
||||
"""5.5.9 ``object_info_block(0)``:一个对象的属性更新。"""
|
||||
info = {
|
||||
"object": object_index,
|
||||
"in_bed_or_isf": bool(in_bed_or_isf),
|
||||
"not_active": bool(reader.read(1)),
|
||||
}
|
||||
# blk == 0 且对象激活时 basic info 恒为 full update(status 0b01)。
|
||||
basic_status = 0 if info["not_active"] else 1
|
||||
info["basic_status"] = basic_status
|
||||
if basic_status in (1, 3):
|
||||
if basic_status == 1:
|
||||
# 5.5.10:object_basic_info[] = {true, true};
|
||||
# 5.6.4.12:array[1] = object_gain_idx,array[0] = b_default_object_priority。
|
||||
gain_present = priority_present = True
|
||||
else:
|
||||
flags = reader.read(2)
|
||||
gain_present = bool(flags & 1)
|
||||
priority_present = bool(flags & 2)
|
||||
if gain_present:
|
||||
gain_idx = reader.read(2)
|
||||
info["gain_idx"] = gain_idx
|
||||
if gain_idx == 2:
|
||||
info["gain_bits"] = reader.read(6)
|
||||
if priority_present:
|
||||
default_priority = bool(reader.read(1))
|
||||
info["default_priority"] = default_priority
|
||||
if not default_priority:
|
||||
info["priority_bits"] = reader.read(5)
|
||||
render_status = 0 if (info["not_active"] or info["in_bed_or_isf"]) else 1
|
||||
info["render_status"] = render_status
|
||||
if render_status in (1, 3):
|
||||
if render_status == 1:
|
||||
# 5.5.11:obj_render_info[] = {true, true, true, true}。
|
||||
position_present = zone_present = size_present = screen_present = True
|
||||
else:
|
||||
flags = reader.read(4)
|
||||
position_present = bool(flags & 0x1)
|
||||
zone_present = bool(flags & 0x2)
|
||||
size_present = bool(flags & 0x4)
|
||||
screen_present = bool(flags & 0x8)
|
||||
if position_present:
|
||||
# blk == 0 时 b_differential_position_specified 恒为 FALSE。
|
||||
x = reader.read(6)
|
||||
y = reader.read(6)
|
||||
z_sign = reader.read(1)
|
||||
z = reader.read(4)
|
||||
info["position"] = (x, y, z if z_sign else -z)
|
||||
if reader.read(1): # b_object_distance_specified
|
||||
if not reader.read(1): # b_object_at_infinity
|
||||
reader.read(4) # distance_factor_idx
|
||||
if zone_present:
|
||||
info["zone_constraints_idx"] = reader.read(3)
|
||||
info["enable_elevation"] = bool(reader.read(1))
|
||||
if size_present:
|
||||
size_idx = reader.read(2)
|
||||
info["object_size_idx"] = size_idx
|
||||
if size_idx == 1:
|
||||
reader.read(5)
|
||||
elif size_idx == 2:
|
||||
reader.read(15)
|
||||
if screen_present:
|
||||
if reader.read(1): # b_object_use_screen_ref
|
||||
reader.read(3) # screen_factor_bits
|
||||
reader.read(2) # depth_factor_idx
|
||||
info["snap"] = bool(reader.read(1))
|
||||
if reader.read(1): # b_additional_table_data_exists
|
||||
info["additional_table_bytes"] = reader.read(4) + 1
|
||||
reader.skip(info["additional_table_bytes"] * 8)
|
||||
return info
|
||||
|
||||
|
||||
def _parse_object_element(reader, element, raw_payload):
|
||||
"""5.5.5 ``object_element()``:时间信息 + 每个对象一条 object_info_block。"""
|
||||
object_count = element["object_count"]
|
||||
def _update_timing(bits, object_element, raw_payload):
|
||||
"""从已定位的 object element 读取位置块开始偏移和 ramp 时长。"""
|
||||
reader = _BitReader(
|
||||
bits, object_element["data_start_bit"], object_element["body_end_bit"])
|
||||
try:
|
||||
sample_offset_code = reader.read(2)
|
||||
if sample_offset_code == 0:
|
||||
offset_code = reader.read(2)
|
||||
if offset_code == 0:
|
||||
sample_offset = 0
|
||||
elif sample_offset_code == 1:
|
||||
elif offset_code == 1:
|
||||
sample_offset = SAMPLE_OFFSET_INDEX[reader.read(2)]
|
||||
elif sample_offset_code == 2:
|
||||
elif offset_code == 2:
|
||||
sample_offset = reader.read(5)
|
||||
else:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "md_sample_offset_mode",
|
||||
"OAMD 使用了当前未覆盖的 MD sample-offset 模式",
|
||||
details={
|
||||
"sample_offset_code": sample_offset_code,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
details={"payload": bytes_descriptor(raw_payload)})
|
||||
|
||||
block_count = reader.read(3) + 1
|
||||
blocks = []
|
||||
for _block in range(block_count):
|
||||
block_offset_factor = reader.read(6)
|
||||
block_offset = sample_offset + block_offset_factor * 32
|
||||
ramp_code = reader.read(2)
|
||||
if ramp_code == 3:
|
||||
if reader.read(1):
|
||||
@@ -318,253 +277,107 @@ def _parse_object_element(reader, element, raw_payload):
|
||||
ramp_duration = reader.read(11)
|
||||
else:
|
||||
ramp_duration = RAMP_DURATIONS[ramp_code]
|
||||
blocks.append({
|
||||
"block_offset_samples": sample_offset + block_offset_factor * 32,
|
||||
"ramp_duration_samples": ramp_duration,
|
||||
})
|
||||
reserved_data_not_present = bool(reader.read(1))
|
||||
if not reserved_data_not_present:
|
||||
reader.read(5)
|
||||
objects = [
|
||||
_parse_object_info_block(
|
||||
reader, index, index < element["bed_isf_objects"], raw_payload)
|
||||
for index in range(object_count)
|
||||
]
|
||||
blocks.append((block_offset, ramp_duration))
|
||||
except UnsupportedVariantError:
|
||||
raise
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"object_element_syntax",
|
||||
"OAMD object element 字段越界或不完整",
|
||||
raw_payload, exc,
|
||||
element=_element_details([element])[0],
|
||||
object_count=object_count,
|
||||
) from exc
|
||||
return {
|
||||
"sample_offset": sample_offset,
|
||||
"blocks": blocks,
|
||||
"reserved_data_not_present": reserved_data_not_present,
|
||||
"objects": objects,
|
||||
}
|
||||
|
||||
|
||||
def _parse_elements(reader, alternate_present, object_count, program, raw_payload):
|
||||
"""5.5.4 ``oa_element_md()`` 目录;对象 element 按结构解析,其余按声明长度跳过。"""
|
||||
try:
|
||||
element_count = reader.read(4)
|
||||
if element_count == 0xF:
|
||||
element_count += reader.read(5)
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"element_count_truncated", "OAMD element 数量字段不完整",
|
||||
raw_payload, exc) from exc
|
||||
if element_count == 0:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "missing_object_element",
|
||||
"OAMD 没有声明任何 element",
|
||||
details={"payload": bytes_descriptor(raw_payload)})
|
||||
|
||||
bed_isf_objects = program["num_bed_objects"] + program["num_isf_objects"]
|
||||
elements = []
|
||||
object_element = None
|
||||
for ordinal in range(element_count):
|
||||
header_start = reader.position
|
||||
try:
|
||||
element_id = reader.read(4)
|
||||
size_bytes = _variable_bits_max(reader, 4, 4) + 1
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"element_header_truncated", "OAMD element header 不完整",
|
||||
raw_payload, exc,
|
||||
element_ordinal=ordinal, header_start_bit=header_start) from exc
|
||||
region_start = reader.position
|
||||
region_end = region_start + size_bytes * 8
|
||||
if region_end > reader.limit:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "element_bounds",
|
||||
"OAMD element 声明长度超过 payload 边界",
|
||||
details={
|
||||
"element_ordinal": ordinal,
|
||||
"element_id": element_id,
|
||||
"size_bytes": size_bytes,
|
||||
"body_start_bit": region_start,
|
||||
"body_end_bit": region_end,
|
||||
"payload_bits": reader.limit,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
try:
|
||||
alternate_data_id = (reader.read(4)
|
||||
if alternate_present else None)
|
||||
discard_unknown = bool(reader.read(1))
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"element_control_bounds", "OAMD element 太短,无法容纳控制字段",
|
||||
raw_payload, exc,
|
||||
element_ordinal=ordinal, element_id=element_id,
|
||||
size_bytes=size_bytes) from exc
|
||||
element = {
|
||||
"ordinal": ordinal,
|
||||
"element_id": element_id,
|
||||
"size_bytes": size_bytes,
|
||||
"header_start_bit": header_start,
|
||||
"body_start_bit": region_start,
|
||||
"body_end_bit": region_end,
|
||||
"data_start_bit": reader.position,
|
||||
"alternate_data_id": alternate_data_id,
|
||||
"discard_unknown": discard_unknown,
|
||||
"object_count": object_count,
|
||||
"bed_isf_objects": bed_isf_objects,
|
||||
}
|
||||
if element_id == OBJECT_ELEMENT_ID:
|
||||
if object_element is not None:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "multiple_object_elements",
|
||||
"OAMD 含多个 object element",
|
||||
details={
|
||||
"elements": _element_details(elements + [element]),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
object_element = _parse_object_element(
|
||||
reader, element, raw_payload)
|
||||
element["parsed_end_bit"] = reader.position
|
||||
element["object_element"] = object_element
|
||||
# 个别编码器声明的 oa_element_size 比实际内容短;以结构解析结果为准。
|
||||
reader.position = max(region_end, reader.position)
|
||||
else:
|
||||
# trim / extended / 未知 element:按声明长度整体跳过(5.5.4)。
|
||||
element["parsed_end_bit"] = region_end
|
||||
reader.position = region_end
|
||||
elements.append(element)
|
||||
|
||||
if object_element is None:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "missing_object_element",
|
||||
"OAMD 缺少 object element",
|
||||
"oamd", "object_element_syntax",
|
||||
"OAMD object element 的 timing 字段越界或不完整",
|
||||
details={
|
||||
"elements": _element_details(elements),
|
||||
"element": _element_details([object_element])[0],
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
padding = reader.bits[reader.position:]
|
||||
if np.any(padding):
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "nonzero_padding",
|
||||
"OAMD payload 尾部 padding 含非零位",
|
||||
details={
|
||||
"elements": _element_details(elements),
|
||||
"padding_start_bit": reader.position,
|
||||
"padding_bits": "".join(str(int(bit)) for bit in padding[:64]),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
return object_element, elements
|
||||
"parser_error": str(exc),
|
||||
}) from exc
|
||||
|
||||
|
||||
def frame_update(bits_one):
|
||||
"""单帧 OAMD → 位置字段增量及其 sample offset/ramp duration。"""
|
||||
bits, raw_payload = _payload_bits(bits_one)
|
||||
try:
|
||||
version = bits[0] << 1 | bits[1]
|
||||
except IndexError:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "header_truncated",
|
||||
"OAMD payload 不足以容纳 header",
|
||||
details={"payload_bits": len(bits),
|
||||
"payload": bytes_descriptor(raw_payload)}) from None
|
||||
reader = _BitReader(bits, 2)
|
||||
try:
|
||||
if version == 3:
|
||||
version += reader.read(3)
|
||||
object_count_bits = reader.read(5)
|
||||
if object_count_bits == 0x1F:
|
||||
object_count_bits += reader.read(7)
|
||||
object_count = object_count_bits + 1
|
||||
program = _parse_program_assignment(reader, raw_payload)
|
||||
alternate_present = bool(reader.read(1))
|
||||
object_element, elements = _parse_elements(
|
||||
reader, alternate_present, object_count, program, raw_payload)
|
||||
except UnsupportedVariantError:
|
||||
raise
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"header_truncated", "OAMD header 字段越界或不完整",
|
||||
raw_payload, exc, payload_bits=len(bits)) from exc
|
||||
|
||||
if version != 0:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "oamd_version",
|
||||
"OAMD 使用了当前未覆盖的 syntax version",
|
||||
details={
|
||||
"version": version,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
if object_count > MAX_SLOTS:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "object_count",
|
||||
"OAMD 对象数超过当前 16 槽对象模型",
|
||||
details={
|
||||
"object_count": object_count,
|
||||
"max_slots": MAX_SLOTS,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
blocks = object_element["blocks"]
|
||||
if len(blocks) != 1:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "multiple_position_blocks",
|
||||
"OAMD 一帧含多个对象位置更新块,单次 frame_update 无法表达",
|
||||
"OAMD 一帧含多个对象位置更新块,固定位置窗口不能安全套用",
|
||||
details={
|
||||
"block_count": len(blocks),
|
||||
"blocks": blocks,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"repair_hint": "按 ObjectInfoBlock 顺序逐块解析坐标,再生成分段 ADM ramp",
|
||||
})
|
||||
return blocks[0]
|
||||
|
||||
# 对外契约:始终给出 16 个槽;本帧没有覆盖到的槽保持上一帧位置。
|
||||
out = {(slot, field): None for slot in range(MAX_SLOTS)
|
||||
for field in ("q1", "q2", "q3")}
|
||||
for info in object_element["objects"]:
|
||||
slot = info["object"]
|
||||
position = info.get("position")
|
||||
if position is None:
|
||||
# bed/ISF 对象与未激活对象没有位置字段:保持上一帧位置。
|
||||
out[(slot, "q1")] = None
|
||||
out[(slot, "q2")] = None
|
||||
out[(slot, "q3")] = None
|
||||
continue
|
||||
x, y, z = position
|
||||
out[(slot, "q1")] = q_of(x, N_Q12) if 0 <= x <= N_Q12 else None
|
||||
out[(slot, "q2")] = q_of(y, N_Q12) if 0 <= y <= N_Q12 else None
|
||||
# 5.6.1.1.10/5.6.1.1.11:pos3D_Z 带符号;当前 16 槽模型只表达非负高度。
|
||||
out[(slot, "q3")] = q_of(z, N_Q3) if 0 <= z <= N_Q3 else (
|
||||
0 if z < 0 else None)
|
||||
|
||||
size_mismatch = [
|
||||
{
|
||||
"element_ordinal": element["ordinal"],
|
||||
"element_id": element["element_id"],
|
||||
"declared_size_bytes": element["size_bytes"],
|
||||
"declared_end_bit": element["body_end_bit"],
|
||||
"parsed_end_bit": element["parsed_end_bit"],
|
||||
}
|
||||
for element in elements
|
||||
if element["parsed_end_bit"] > element["body_end_bit"]
|
||||
]
|
||||
def frame_update(bits_one):
|
||||
"""单帧 OAMD → 位置字段增量及其 sample offset/ramp duration。"""
|
||||
bits, raw_payload = _payload_bits(bits_one)
|
||||
if len(bits) < 14:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "header_truncated",
|
||||
"OAMD payload 不足以容纳受支持的 header",
|
||||
details={"payload_bits": len(bits), "payload": bytes_descriptor(raw_payload)})
|
||||
header = {
|
||||
"version": int((bits[0] << 1) | bits[1]),
|
||||
"objects_minus_one": int(sum(int(bits[2 + i]) << (4 - i) for i in range(5))),
|
||||
"dynamic_object_only": int(bits[7]),
|
||||
"lfe_present": int(bits[8]),
|
||||
"alternate_object_present": int(bits[9]),
|
||||
"element_count": int(sum(int(bits[10 + i]) << (3 - i) for i in range(4))),
|
||||
}
|
||||
if raw_payload[:2] != b"\x1f\x88":
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", f"header_signature_{raw_payload[:2].hex()}",
|
||||
"OAMD header 与当前位置字段布局不一致",
|
||||
details={
|
||||
"supported_header_prefix_hex": "1f88",
|
||||
"header_probe": header,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"repair_hint": "按新 header 的 program assignment 和 element 布局重新定位对象位置字段",
|
||||
})
|
||||
object_element, elements = _parse_elements(bits, header, 14, raw_payload)
|
||||
if object_element["body_end_bit"] < POSITION_WINDOW_END_BIT:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "object_element_too_short",
|
||||
"OAMD object element 无法容纳当前固定位置窗口",
|
||||
details={
|
||||
"element": _element_details([object_element])[0],
|
||||
"required_position_end_bit": POSITION_WINDOW_END_BIT,
|
||||
"elements": _element_details(elements),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
block_offset, ramp_duration = _update_timing(bits, object_element, raw_payload)
|
||||
weights = 1 << np.arange(7, -1, -1)
|
||||
out = {}
|
||||
for obj in range(16):
|
||||
start = 112 + 31 * (obj - 3)
|
||||
wq1 = int((bits[start:start + 8] * weights).sum())
|
||||
wq2 = int((bits[start + 8:start + 16] * weights).sum())
|
||||
wq3 = int((bits[start + 16:start + 24] * weights).sum())
|
||||
if obj and (wq1 >> 6 != 3 or wq2 & 2 != 2 or wq3 & 0x1F != 1):
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "position_layout_signature",
|
||||
"OAMD 对象位置字段标记或位偏移发生变化",
|
||||
details={
|
||||
"object_slot": obj,
|
||||
"position_start_bit": start,
|
||||
"q1_window_hex": f"{wq1:02x}",
|
||||
"q2_window_hex": f"{wq2:02x}",
|
||||
"q3_window_hex": f"{wq3:02x}",
|
||||
"header_probe": header,
|
||||
"elements": _element_details(elements),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"repair_hint": "解析 OAMD element 可选字段并更新每个对象的位置窗口偏移",
|
||||
})
|
||||
k1 = wq1 - Q1_OFF
|
||||
if obj == 0:
|
||||
out[(0, "q1")] = None
|
||||
out[(0, "q2")] = None
|
||||
else:
|
||||
out[(obj, "q1")] = q_of(k1, N_Q12) if 0 <= k1 <= N_Q12 else None
|
||||
k2 = wq2 >> 2
|
||||
if obj != 0:
|
||||
out[(obj, "q2")] = q_of(k2, N_Q12) if 0 <= k2 <= N_Q12 else None
|
||||
k3 = (wq2 & 1) * 8 + (wq3 >> 5)
|
||||
out[(obj, "q3")] = q_of(k3, N_Q3) if 0 <= k3 <= N_Q3 else None
|
||||
return {
|
||||
"values": out,
|
||||
"block_offset_samples": blocks[0]["block_offset_samples"],
|
||||
"ramp_duration_samples": blocks[0]["ramp_duration_samples"],
|
||||
"object_count": object_count,
|
||||
"program": {
|
||||
"dynamic_object_only": program["dynamic_object_only"],
|
||||
"lfe_present": program["lfe_present"],
|
||||
"content_description": program["content_description"],
|
||||
"num_bed_objects": program["num_bed_objects"],
|
||||
"num_isf_objects": program["num_isf_objects"],
|
||||
"num_dynamic_objects": program["num_dynamic_objects"],
|
||||
"bed_channels": [list(assignment["channels"])
|
||||
for assignment in program["bed_assignments"]],
|
||||
},
|
||||
"elements": _element_details(elements),
|
||||
"objects": object_element["objects"],
|
||||
"size_mismatch": size_mismatch,
|
||||
"block_offset_samples": block_offset,
|
||||
"ramp_duration_samples": ramp_duration,
|
||||
}
|
||||
|
||||
|
||||
|
||||
+39
-53
@@ -5,18 +5,6 @@ from adm_atmos import q_to_adm_xyz
|
||||
from oamd_bits import JocFieldState, frame_update
|
||||
from variant_error import UnsupportedVariantError
|
||||
|
||||
# 元数据更新时刻按渲染器处理块对齐,与 Dolby 的 processing_block_size 及
|
||||
# speaker_renderer 的 block_size 落在同一网格。
|
||||
METADATA_BLOCK_SAMPLES = 32
|
||||
|
||||
|
||||
def align_metadata_sample(sample, block_samples=METADATA_BLOCK_SAMPLES):
|
||||
"""把更新时刻量化到处理块边界。"""
|
||||
block = int(block_samples)
|
||||
if block <= 0:
|
||||
raise ValueError("block_samples must be positive")
|
||||
return block * ((int(sample) + block // 2 - 1) // block)
|
||||
|
||||
|
||||
def _lerp_xyz(start, target, amount):
|
||||
return tuple(a + (b - a) * amount for a, b in zip(start, target))
|
||||
@@ -30,19 +18,18 @@ def _append_point(points, sample, xyz, interpolation_samples):
|
||||
points.append(item)
|
||||
|
||||
|
||||
def _expand_events_dense64(events, total_samples, rate, update_block_samples,
|
||||
def _expand_events_dense64(events, total_samples, rate, update_quantum_samples,
|
||||
object_delay_samples, object_index):
|
||||
if not events:
|
||||
return [(0.0, 0.0, 0.0, 0.0, total_samples / float(rate), 0.0)]
|
||||
|
||||
block = int(update_block_samples)
|
||||
# 初始位置从成品 sample 0 起有效;合成延迟只作用于后续位置变化。
|
||||
current = events[0][1]
|
||||
points = []
|
||||
_append_point(points, 0, current, 0)
|
||||
|
||||
for event_index, (coded_start, target, ramp_samples) in enumerate(events[1:], 1):
|
||||
start = align_metadata_sample(coded_start + object_delay_samples, block)
|
||||
start = coded_start + object_delay_samples
|
||||
if start >= total_samples:
|
||||
break
|
||||
if start < points[-1][0]:
|
||||
@@ -56,16 +43,15 @@ def _expand_events_dense64(events, total_samples, rate, update_block_samples,
|
||||
if start > points[-1][0]:
|
||||
_append_point(points, start, current, 0)
|
||||
|
||||
ramp = max(0, int(ramp_samples))
|
||||
if ramp == 0:
|
||||
effective_ramp = max(0, int(ramp_samples) - update_quantum_samples)
|
||||
if effective_ramp == 0:
|
||||
_append_point(points, start, target, 0)
|
||||
current = target
|
||||
continue
|
||||
|
||||
end = start + math.ceil(ramp / block) * block
|
||||
end = start + math.ceil(effective_ramp / update_quantum_samples) * update_quantum_samples
|
||||
if event_index + 1 < len(events):
|
||||
next_start = align_metadata_sample(
|
||||
events[event_index + 1][0] + object_delay_samples, block)
|
||||
next_start = events[event_index + 1][0] + object_delay_samples
|
||||
if next_start < end:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "overlapping_position_ramps",
|
||||
@@ -75,23 +61,23 @@ def _expand_events_dense64(events, total_samples, rate, update_block_samples,
|
||||
"ramp_start_sample": start,
|
||||
"ramp_end_sample": end,
|
||||
"next_update_sample": next_start,
|
||||
"repair_hint": "按 32-sample 状态机截断旧 ramp,再从当前插值位置启动新 ramp",
|
||||
"repair_hint": "按 64-sample 状态机截断旧 ramp,再从当前插值位置启动新 ramp",
|
||||
})
|
||||
|
||||
# 从对齐后的更新点起按处理块推进,整条 ramp 覆盖 ramp_samples 个样本;
|
||||
# 几何式逼近望远镜化简为精确线性,每个中间点都落在真实 ramp 上。
|
||||
future = ramp
|
||||
# 逐位置更新节拍复现状态机。1536-sample ramp 在首次 64-sample
|
||||
# 更新后剩余 1472 samples,因此共有 23 个中间/终点坐标。
|
||||
future = effective_ramp
|
||||
elapsed = 0
|
||||
position = current
|
||||
while future > 0:
|
||||
amount = min(block / float(future), 1.0)
|
||||
amount = min(update_quantum_samples / float(future), 1.0)
|
||||
position = _lerp_xyz(position, target, amount)
|
||||
elapsed += block
|
||||
elapsed += update_quantum_samples
|
||||
sample = start + elapsed
|
||||
if sample >= total_samples:
|
||||
break
|
||||
_append_point(points, sample, position, block)
|
||||
future -= block
|
||||
_append_point(points, sample, position, update_quantum_samples)
|
||||
future -= update_quantum_samples
|
||||
current = target
|
||||
|
||||
blocks = []
|
||||
@@ -108,14 +94,16 @@ def _expand_events_dense64(events, total_samples, rate, update_block_samples,
|
||||
|
||||
|
||||
|
||||
def _compact_events(events, total_samples, rate, update_block_samples,
|
||||
def _compact_events(events, total_samples, rate, update_quantum_samples,
|
||||
object_delay_samples, object_index):
|
||||
"""Represent each linear OAMD ramp with one ADM interpolation block.
|
||||
|
||||
The update instant is quantized to the processing block boundary, the motion
|
||||
starts there immediately, and ``interpolationLength`` spans the full
|
||||
``ramp_duration``. The interpreted motion therefore covers
|
||||
``[align(start), align(start) + ramp_duration]``.
|
||||
The existing dense64 representation keeps the old position at ``start``,
|
||||
writes its first interpolated target at ``start + quantum``, and lets ADM
|
||||
interpolate that block over one quantum. Consequently, the interpreted
|
||||
motion begins at ``start + quantum`` and reaches the final target at
|
||||
``start + ramp_duration``. This compact form preserves that timing with one
|
||||
target block whose interpolationLength is ``ramp_duration - quantum``.
|
||||
"""
|
||||
if not events:
|
||||
return [(0.0, 0.0, 0.0, 0.0, total_samples / float(rate), 0.0)]
|
||||
@@ -124,16 +112,15 @@ def _compact_events(events, total_samples, rate, update_block_samples,
|
||||
current = events[0][1]
|
||||
_append_point(points, 0, current, 0)
|
||||
|
||||
block = int(update_block_samples)
|
||||
for event_index, (coded_start, target, ramp_samples) in enumerate(events[1:], 1):
|
||||
event_start = align_metadata_sample(coded_start + object_delay_samples, block)
|
||||
event_start = coded_start + object_delay_samples
|
||||
if event_start >= total_samples:
|
||||
break
|
||||
ramp = max(0, int(ramp_samples))
|
||||
block_start = event_start
|
||||
effective_ramp = max(0, int(ramp_samples) - update_quantum_samples)
|
||||
block_start = event_start + (update_quantum_samples if effective_ramp else 0)
|
||||
if block_start >= total_samples:
|
||||
break
|
||||
ramp_end = block_start + ramp
|
||||
ramp_end = block_start + effective_ramp
|
||||
|
||||
if block_start < points[-1][0]:
|
||||
raise UnsupportedVariantError(
|
||||
@@ -144,9 +131,9 @@ def _compact_events(events, total_samples, rate, update_block_samples,
|
||||
|
||||
if event_index + 1 < len(events):
|
||||
next_coded_start, _, next_ramp_samples = events[event_index + 1]
|
||||
next_event_start = align_metadata_sample(
|
||||
next_coded_start + object_delay_samples, block)
|
||||
next_block_start = next_event_start
|
||||
next_event_start = next_coded_start + object_delay_samples
|
||||
next_effective = max(0, int(next_ramp_samples) - update_quantum_samples)
|
||||
next_block_start = next_event_start + (update_quantum_samples if next_effective else 0)
|
||||
if next_block_start < ramp_end:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "overlapping_compact_position_ramps",
|
||||
@@ -160,10 +147,10 @@ def _compact_events(events, total_samples, rate, update_block_samples,
|
||||
})
|
||||
|
||||
block_target = target
|
||||
block_interpolation = ramp
|
||||
block_interpolation = effective_ramp
|
||||
available = total_samples - block_start
|
||||
if ramp > available:
|
||||
block_target = _lerp_xyz(current, target, available / float(ramp))
|
||||
if effective_ramp > available:
|
||||
block_target = _lerp_xyz(current, target, available / float(effective_ramp))
|
||||
block_interpolation = available
|
||||
_append_point(points, block_start, block_target, block_interpolation)
|
||||
current = target
|
||||
@@ -181,26 +168,25 @@ def _compact_events(events, total_samples, rate, update_block_samples,
|
||||
return blocks
|
||||
|
||||
|
||||
def _expand_events(events, total_samples, rate, update_block_samples,
|
||||
def _expand_events(events, total_samples, rate, update_quantum_samples,
|
||||
object_delay_samples, object_index, trajectory_mode):
|
||||
if trajectory_mode == "compact":
|
||||
return _compact_events(events, total_samples, rate, update_block_samples,
|
||||
return _compact_events(events, total_samples, rate, update_quantum_samples,
|
||||
object_delay_samples, object_index)
|
||||
if trajectory_mode == "dense64":
|
||||
return _expand_events_dense64(events, total_samples, rate, update_block_samples,
|
||||
return _expand_events_dense64(events, total_samples, rate, update_quantum_samples,
|
||||
object_delay_samples, object_index)
|
||||
raise ValueError(f"未知 trajectory_mode: {trajectory_mode}")
|
||||
|
||||
def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536,
|
||||
update_block_samples=METADATA_BLOCK_SAMPLES,
|
||||
object_delay_samples=1473, trajectory_mode="compact"):
|
||||
update_quantum_samples=64, object_delay_samples=1473,
|
||||
trajectory_mode="compact"):
|
||||
"""从统一 metadata index 构造 15 条 ADM 轨迹。
|
||||
|
||||
返回 ``[(name, [(rtime,x,y,z,duration,interpolation), ...]), ...]``。
|
||||
OAMD 的内外层 sample offset、block offset 和 ramp 均保留;更新时刻量化到
|
||||
``update_block_samples`` 的块边界,ramp 覆盖完整的 ramp_duration。
|
||||
OAMD 的内外层 sample offset、block offset 和 ramp 均保留。
|
||||
``trajectory_mode="compact"`` 用一个长 ADM interpolation block 表示每条
|
||||
线性 ramp;``dense64`` 保留逐块展开作为兼容回退。
|
||||
线性 ramp;``dense64`` 保留逐 64-sample 展开作为兼容回退。
|
||||
``object_delay_samples`` 将位置更新与对象 PCM 的 decoder 输出时刻对齐。
|
||||
slot1..15 与对象 PCM ch1..15 一一对应。
|
||||
"""
|
||||
@@ -235,7 +221,7 @@ def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536,
|
||||
return [
|
||||
(f"JOC_Object_{obj}",
|
||||
_expand_events(events[obj - 1], total_samples, rate,
|
||||
update_block_samples, object_delay_samples, obj,
|
||||
update_quantum_samples, object_delay_samples, obj,
|
||||
trajectory_mode))
|
||||
for obj in range(1, 16)
|
||||
]
|
||||
|
||||
@@ -1,692 +0,0 @@
|
||||
"""Public 64-QMF and 77-band hybrid filterbank for binaural rendering.
|
||||
|
||||
The fixed resource is ``data/rosella_kernels.npz``: the fixed 64-QMF /
|
||||
``3 -> 8+4+4`` 77-hybrid analysis tables and the causal synthesis tables
|
||||
computed from that analysis bank. The filter bank is publicly standardized:
|
||||
the 64-QMF → 77-hybrid structure, the 13-tap low-band prototypes and their
|
||||
half-bin complex modulation follow 3GPP TS 26.405 / ETSI TS 126 405 (Section
|
||||
5.2.2, Table 1, ``Q=8``/``Q=4``); the 64-band QMF analysis is the MPEG-4
|
||||
AAC/SBR 64 complex QMF analysis bank (ISO/IEC 14496-3/AMD1:2003, subclause
|
||||
4.B.18.2), stored here as the polyphase form
|
||||
``A[r,t] = ((-1)**t / 128) * c[63 - r + 64*t]`` of the public 640-tap SBR
|
||||
prototype. The QMF synthesis table is the causal left inverse of that
|
||||
analysis polyphase matrix (``A @ W = P`` with the 577-sample delay
|
||||
permutation; total latency ``961 = 577 + 6*64``), stored as the rank-4
|
||||
factorization ``W[b,l] = sum_r taps[b,l,r] * basis[b,r,:]``; the hybrid
|
||||
synthesis table is the 77->64 recombination (identity for the high bands,
|
||||
signed summation of each 8+4+4 child group for the low bands), stored as a
|
||||
154-entry sparse map. The archive and every array inside it are
|
||||
hash-validated before use, and those hashes participate in every
|
||||
compiled-HRTF cache key. Provenance and rights boundaries are documented in
|
||||
``data/README.md`` and ``THIRD_PARTY_NOTICES.md``.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from functools import lru_cache
|
||||
import hashlib
|
||||
import os
|
||||
from pathlib import Path
|
||||
import zipfile
|
||||
|
||||
import numpy as np
|
||||
|
||||
|
||||
PROJECT_DIR = Path(__file__).resolve().parent.parent
|
||||
DEFAULT_FILTERBANK_DATA = PROJECT_DIR / "data" / "rosella_kernels.npz"
|
||||
FILTERBANK_TABLE_VERSION = "joc-public-64qmf-77hybrid-v1"
|
||||
SAMPLE_RATE = 48000
|
||||
QMF_HOP = 64
|
||||
QMF_BANDS = 64
|
||||
HYBRID_BANDS = 77
|
||||
ANALYSIS_SYNTHESIS_LATENCY_SAMPLES = 961
|
||||
|
||||
_ARCHIVE_SHA256 = "C05BEF4D26E96ECBD4694E2572F05DA400255C777BA5047300B9D3B1F81081CD"
|
||||
_TABLE_SPECS = {
|
||||
"format_version": (np.dtype("<i4"), (1,),
|
||||
"67ABDD721024F0FF4E0B3F4C2FC13BC5BAD42D0B7851D456D88D203D15AAA450",
|
||||
False, False),
|
||||
"qmf_analysis_coefficients": (
|
||||
np.dtype("<f4"), (64, 10),
|
||||
"AEFF6C7117D41664B9C4BF03BBF563F5319EC1B8C551F171ADBB90CF19D9D306",
|
||||
False, False),
|
||||
"hybrid_analysis_low_kernel": (
|
||||
np.dtype("<f4"), (3, 2, 13, 16, 2),
|
||||
"D00D36133B81BA699A7630C4DF1BE203FA1B7E371E595EAAEBBE8957DB322627",
|
||||
False, False),
|
||||
"hybrid_synthesis_indices": (
|
||||
np.dtype("<i2"), (154, 4),
|
||||
"F5BEB3220E4530FCF28E7F4DA7F07E821074265D118C911D61A590E00753A573",
|
||||
True, False),
|
||||
"hybrid_synthesis_values": (
|
||||
np.dtype("<f4"), (154,),
|
||||
"99409FDD9D20D1D7C2BE16BBC1E2159C8042487227C72160850745164C9CEE7F",
|
||||
False, False),
|
||||
"qmf_synthesis_basis": (
|
||||
np.dtype("<f8"), (64, 4, 128),
|
||||
"A0C4A55385F6D6C7C92D7615C83AD5FBDA51046D9EF785CAC0B9AC9A760DC527",
|
||||
False, False),
|
||||
"qmf_synthesis_taps": (
|
||||
np.dtype("<f8"), (64, 10, 4),
|
||||
"CD7756D060D51FBF02F44C1CE53CB6225B221099505C94C3F58D3BEE6F428150",
|
||||
False, False),
|
||||
}
|
||||
|
||||
# Hybrid-band center frequencies measured from the public analysis bank at
|
||||
# 48 kHz (positive-frequency response peaks). They are part of the validated
|
||||
# reference behavior: the runtime uses them only for the fractional-delay band
|
||||
# phase and the project LFE low-pass, never as filterbank coefficients.
|
||||
_BAND_CENTER_FREQUENCIES_HZ = np.asarray([
|
||||
53.19564095937407,
|
||||
26.3876219849709,
|
||||
140.9074183269806,
|
||||
98.55238901464415,
|
||||
234.09258166297573,
|
||||
344.8745321543293,
|
||||
321.80435900819805,
|
||||
401.38762185365727,
|
||||
476.19239745597804,
|
||||
473.552388917001,
|
||||
648.8076026837931,
|
||||
719.8745321201852,
|
||||
780.1254679168173,
|
||||
851.1923973714038,
|
||||
1023.8076024849751,
|
||||
1155.125467695814,
|
||||
1293.0325067063661,
|
||||
1668.032511377253,
|
||||
2043.0325074138086,
|
||||
2456.967490229666,
|
||||
2831.96748495363,
|
||||
3206.967488172826,
|
||||
3581.9675052705525,
|
||||
3918.0325089666067,
|
||||
4331.967500201844,
|
||||
4706.967500350216,
|
||||
5043.032510771545,
|
||||
5456.9674914391635,
|
||||
5831.967490225267,
|
||||
6168.0324885741875,
|
||||
6543.032508972284,
|
||||
6956.96748844665,
|
||||
7293.032503259869,
|
||||
7668.032503551393,
|
||||
8043.032504806491,
|
||||
8418.032498852166,
|
||||
8793.032502508235,
|
||||
9206.967488589786,
|
||||
9543.032511549152,
|
||||
9956.967486913867,
|
||||
10293.032513008677,
|
||||
10668.032507835102,
|
||||
11043.032513641429,
|
||||
11456.967482937946,
|
||||
11793.032513984212,
|
||||
12206.967486015788,
|
||||
12543.032517062376,
|
||||
12956.96748635822,
|
||||
13331.967492165066,
|
||||
13706.967486991198,
|
||||
14043.032513086031,
|
||||
14456.967488451,
|
||||
14793.032511410214,
|
||||
15206.967497492202,
|
||||
15581.967501147887,
|
||||
15956.967495193188,
|
||||
16331.96749644876,
|
||||
16706.967496740173,
|
||||
17043.03251155318,
|
||||
17456.96749102773,
|
||||
17831.967511425748,
|
||||
18168.032509774734,
|
||||
18543.032508560515,
|
||||
18956.967489228293,
|
||||
19293.032499649784,
|
||||
19668.032499798002,
|
||||
20081.967491033392,
|
||||
20418.0324947296,
|
||||
20793.032511827063,
|
||||
21168.032515046092,
|
||||
21543.032509770488,
|
||||
21956.967492586176,
|
||||
22331.96748862257,
|
||||
22706.967493293465,
|
||||
23043.03251375972,
|
||||
23418.03251061199,
|
||||
23831.96749768645,
|
||||
], dtype=np.float64)
|
||||
|
||||
|
||||
def _sha256_bytes(values: bytes) -> str:
|
||||
return hashlib.sha256(values).hexdigest().upper()
|
||||
|
||||
|
||||
def _validate_npy_member_header(
|
||||
archive: zipfile.ZipFile, member: zipfile.ZipInfo,
|
||||
*, name: str, dtype: np.dtype, shape: tuple[int, ...],
|
||||
allow_fortran: bool) -> None:
|
||||
try:
|
||||
with archive.open(member, "r") as payload:
|
||||
version = np.lib.format.read_magic(payload)
|
||||
if version == (1, 0):
|
||||
actual_shape, actual_fortran_order, actual_dtype = (
|
||||
np.lib.format.read_array_header_1_0(
|
||||
payload, max_header_size=4096))
|
||||
elif version == (2, 0):
|
||||
actual_shape, actual_fortran_order, actual_dtype = (
|
||||
np.lib.format.read_array_header_2_0(
|
||||
payload, max_header_size=4096))
|
||||
else:
|
||||
raise ValueError(f"unsupported .npy version {version!r}")
|
||||
header_size = payload.tell()
|
||||
except (EOFError, OSError, ValueError) as exc:
|
||||
raise ValueError(
|
||||
f"invalid public filterbank table .npy header: {member.filename}: "
|
||||
f"{exc}") from exc
|
||||
|
||||
actual_shape = tuple(actual_shape)
|
||||
actual_dtype = np.dtype(actual_dtype)
|
||||
if (actual_shape != shape or actual_dtype != dtype
|
||||
or (bool(actual_fortran_order) and not allow_fortran)):
|
||||
expected_order = "C-order" if not allow_fortran else "C- or Fortran-order"
|
||||
actual_order = "Fortran-order" if actual_fortran_order else "C-order"
|
||||
raise ValueError(
|
||||
f"invalid public filterbank table .npy header for {name}: expected "
|
||||
f"{dtype}{shape} {expected_order}, got "
|
||||
f"{actual_dtype}{actual_shape} {actual_order}")
|
||||
expected_size = header_size + dtype.itemsize * int(np.prod(shape))
|
||||
if member.file_size != expected_size:
|
||||
raise ValueError(
|
||||
f"invalid public filterbank table .npy payload size for {name}: "
|
||||
f"expected {expected_size} bytes including the header, "
|
||||
f"got {member.file_size}")
|
||||
|
||||
|
||||
def _validate_table_members(stream) -> None:
|
||||
expected = {name + ".npy" for name in _TABLE_SPECS}
|
||||
limits = {
|
||||
name + ".npy": dtype.itemsize * int(np.prod(shape)) + 4096
|
||||
for name, (dtype, shape, _, _, _) in _TABLE_SPECS.items()
|
||||
}
|
||||
try:
|
||||
stream.seek(0)
|
||||
with zipfile.ZipFile(stream, "r") as archive:
|
||||
members = archive.infolist()
|
||||
if (len(members) != len(expected)
|
||||
or {member.filename for member in members} != expected):
|
||||
raise ValueError(
|
||||
"public filterbank table archive has an invalid member set")
|
||||
for member in members:
|
||||
if member.flag_bits & 0x1:
|
||||
raise ValueError("encrypted public filterbank tables are unsupported")
|
||||
if member.compress_type not in (
|
||||
zipfile.ZIP_STORED, zipfile.ZIP_DEFLATED):
|
||||
raise ValueError("unsupported public filterbank table compression")
|
||||
if member.file_size > limits[member.filename]:
|
||||
raise ValueError(
|
||||
f"public filterbank table member is unexpectedly large: "
|
||||
f"{member.filename}")
|
||||
if sum(member.file_size for member in members) > sum(limits.values()):
|
||||
raise ValueError("public filterbank tables expand beyond their size limit")
|
||||
for member in members:
|
||||
name = member.filename[:-4]
|
||||
dtype, shape, _, allow_fortran, _ = _TABLE_SPECS[name]
|
||||
_validate_npy_member_header(
|
||||
archive, member, name=name, dtype=dtype, shape=shape,
|
||||
allow_fortran=allow_fortran)
|
||||
except zipfile.BadZipFile as exc:
|
||||
raise ValueError(f"invalid public filterbank table archive: {exc}") from exc
|
||||
|
||||
|
||||
@lru_cache(maxsize=2)
|
||||
def _load_tables(path_string: str) -> dict[str, np.ndarray]:
|
||||
path = Path(path_string)
|
||||
if not path.is_file():
|
||||
raise FileNotFoundError(f"public filterbank table resource not found: {path}")
|
||||
with path.open("rb") as stream:
|
||||
archive_size = os.fstat(stream.fileno()).st_size
|
||||
if archive_size <= 0 or archive_size > 8 << 20:
|
||||
raise ValueError(
|
||||
f"public filterbank table resource is unexpectedly large: {path}")
|
||||
if path.resolve() == DEFAULT_FILTERBANK_DATA.resolve():
|
||||
digest = hashlib.sha256()
|
||||
for block in iter(lambda: stream.read(4 << 20), b""):
|
||||
digest.update(block)
|
||||
actual_archive_hash = digest.hexdigest().upper()
|
||||
if actual_archive_hash != _ARCHIVE_SHA256:
|
||||
raise ValueError(
|
||||
"public filterbank table archive hash mismatch: "
|
||||
f"expected {_ARCHIVE_SHA256}, got {actual_archive_hash}")
|
||||
_validate_table_members(stream)
|
||||
stream.seek(0)
|
||||
with np.load(stream, allow_pickle=False) as archive:
|
||||
if set(archive.files) != set(_TABLE_SPECS):
|
||||
raise ValueError("public filterbank table archive has an invalid key set")
|
||||
result: dict[str, np.ndarray] = {}
|
||||
for name, (dtype, shape, expected_hash, _, _) in _TABLE_SPECS.items():
|
||||
value = np.asarray(archive[name])
|
||||
if value.dtype != dtype or value.shape != shape:
|
||||
raise ValueError(
|
||||
f"invalid public filterbank table {name}: "
|
||||
f"expected {dtype}{shape}, got {value.dtype}{value.shape}")
|
||||
actual_hash = _sha256_bytes(value.tobytes(order="C"))
|
||||
if actual_hash != expected_hash:
|
||||
raise ValueError(f"public filterbank table hash mismatch: {name}")
|
||||
result[name] = np.ascontiguousarray(value)
|
||||
result[name].setflags(write=False)
|
||||
if int(result["format_version"][0]) != 1:
|
||||
raise ValueError("unsupported public filterbank table format version")
|
||||
return result
|
||||
|
||||
|
||||
def load_filterbank_tables(
|
||||
path: str | Path = DEFAULT_FILTERBANK_DATA) -> dict[str, np.ndarray]:
|
||||
"""Load the validated project resource used by the public filterbank."""
|
||||
cached = _load_tables(str(Path(path).expanduser().resolve()))
|
||||
result = {name: value.copy() for name, value in cached.items()}
|
||||
for value in result.values():
|
||||
value.setflags(write=False)
|
||||
return result
|
||||
|
||||
|
||||
def filterbank_fingerprint() -> dict:
|
||||
"""Return stable identifiers used in compiled-HRTF cache keys."""
|
||||
centers = np.ascontiguousarray(_BAND_CENTER_FREQUENCIES_HZ, dtype="<f8")
|
||||
return {
|
||||
"table_version": FILTERBANK_TABLE_VERSION,
|
||||
"archive_sha256": _ARCHIVE_SHA256,
|
||||
"band_centers_sha256": _sha256_bytes(centers.tobytes(order="C")),
|
||||
"array_sha256": {
|
||||
name: spec[2] for name, spec in _TABLE_SPECS.items()
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
class QmfAnalysis:
|
||||
"""Batchable 64-band analysis with float64 state and complex128 FFTs."""
|
||||
|
||||
def __init__(self, channels: int,
|
||||
table_data: str | Path = DEFAULT_FILTERBANK_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_filterbank_tables(table_data)
|
||||
self.coefficients = np.asarray(
|
||||
tables["qmf_analysis_coefficients"], dtype=np.float64)
|
||||
self.channels = int(channels)
|
||||
self.history = np.zeros((9, self.channels, 64), dtype=np.float64)
|
||||
phase = np.arange(64, dtype=np.float64)
|
||||
self.premod = np.exp(-1j * np.pi * phase / 128.0).astype(np.complex128)
|
||||
self.post = np.exp(
|
||||
-1j * 3.0 * (np.arange(64, dtype=np.float64) + 0.5) * np.pi / 128.0
|
||||
).astype(np.complex128)
|
||||
self.even_post = (
|
||||
1j * ((-1.0) ** np.arange(64, dtype=np.float64))
|
||||
).astype(np.complex128)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.history.fill(0.0)
|
||||
|
||||
def process_chunk(self, hops) -> np.ndarray:
|
||||
values = np.asarray(hops, dtype=np.float64)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 64):
|
||||
raise ValueError(f"expected [slots,{self.channels},64], got {values.shape}")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("QMF input contains non-finite values")
|
||||
count = values.shape[0]
|
||||
joined = np.concatenate((self.history, values), axis=0)
|
||||
even = np.zeros_like(values)
|
||||
odd = np.zeros_like(values)
|
||||
for lag in range(10):
|
||||
source = joined[9 - lag:9 - lag + count]
|
||||
target = even if lag % 2 == 0 else odd
|
||||
target += source * self.coefficients[:, lag][None, None, :]
|
||||
self.history[:] = joined[-9:]
|
||||
|
||||
def transform(block):
|
||||
prepared = block.astype(np.complex128, copy=False) * self.premod
|
||||
transformed = np.fft.fft(prepared, n=128, axis=-1)[..., :64]
|
||||
return transformed * self.post
|
||||
|
||||
return np.asarray(transform(odd) + transform(even) * self.even_post,
|
||||
dtype=np.complex128)
|
||||
|
||||
|
||||
class HybridAnalysis:
|
||||
"""Sparse 64-QMF to 77-hybrid analysis in float64/complex128."""
|
||||
|
||||
def __init__(self, channels: int,
|
||||
table_data: str | Path = DEFAULT_FILTERBANK_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_filterbank_tables(table_data)
|
||||
self.low_kernel = np.asarray(
|
||||
tables["hybrid_analysis_low_kernel"], dtype=np.float64)
|
||||
self.channels = int(channels)
|
||||
self.history = np.zeros((12, self.channels, 3, 2), dtype=np.float64)
|
||||
self.high_history = np.zeros(
|
||||
(6, self.channels, 61), dtype=np.complex128)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.history.fill(0.0)
|
||||
self.high_history.fill(0.0)
|
||||
|
||||
def process_chunk(self, qmf) -> np.ndarray:
|
||||
values = np.asarray(qmf, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 64):
|
||||
raise ValueError(f"expected [slots,{self.channels},64], got {values.shape}")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hybrid-analysis input contains non-finite values")
|
||||
count = values.shape[0]
|
||||
low = np.stack((values[:, :, :3].real, values[:, :, :3].imag), axis=-1)
|
||||
joined = np.concatenate((self.history, low), axis=0)
|
||||
output = np.zeros((count, self.channels, 77, 2), dtype=np.float64)
|
||||
for lag in range(13):
|
||||
source = joined[12 - lag:12 - lag + count]
|
||||
output[:, :, :16] += np.einsum(
|
||||
"tcpi,pibo->tcbo", source, self.low_kernel[:, :, lag],
|
||||
dtype=np.float64, optimize=False)
|
||||
self.history[:] = joined[-12:]
|
||||
|
||||
high_joined = np.concatenate((self.high_history, values[:, :, 3:]), axis=0)
|
||||
high = high_joined[:count]
|
||||
output[:, :, 16:, 0] = high.real
|
||||
output[:, :, 16:, 1] = high.imag
|
||||
self.high_history[:] = high_joined[-6:]
|
||||
return np.asarray(output[..., 0] + 1j * output[..., 1], dtype=np.complex128)
|
||||
|
||||
|
||||
class HybridSynthesis:
|
||||
"""Instantaneous sparse 77-hybrid to 64-QMF synthesis map."""
|
||||
|
||||
def __init__(self, channels: int,
|
||||
table_data: str | Path = DEFAULT_FILTERBANK_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_filterbank_tables(table_data)
|
||||
indices = np.asarray(tables["hybrid_synthesis_indices"], dtype=np.int64)
|
||||
values = np.asarray(tables["hybrid_synthesis_values"], dtype=np.float64)
|
||||
if indices.ndim != 2 or indices.shape[1] != 4 or len(indices) != len(values):
|
||||
raise ValueError("invalid hybrid synthesis sparse table")
|
||||
self.mapping = [
|
||||
(int(index[0]), int(index[1]), int(index[2]), int(index[3]), float(value))
|
||||
for index, value in zip(indices, values)
|
||||
]
|
||||
self.channels = int(channels)
|
||||
|
||||
def reset(self) -> None:
|
||||
return None
|
||||
|
||||
def process_chunk(self, hybrid) -> np.ndarray:
|
||||
values = np.asarray(hybrid, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 77):
|
||||
raise ValueError(f"expected [slots,{self.channels},77], got {values.shape}")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hybrid-synthesis input contains non-finite values")
|
||||
source = np.stack((values.real, values.imag), axis=-1)
|
||||
output = np.zeros((values.shape[0], self.channels, 64, 2), dtype=np.float64)
|
||||
for input_band, input_component, output_band, output_component, gain in self.mapping:
|
||||
output[:, :, output_band, output_component] += (
|
||||
source[:, :, input_band, input_component] * gain)
|
||||
return np.asarray(output[..., 0] + 1j * output[..., 1], dtype=np.complex128)
|
||||
|
||||
|
||||
class QmfSynthesis:
|
||||
"""Rank-4 64-band synthesis with float64 state and accumulation."""
|
||||
|
||||
def __init__(self, channels: int,
|
||||
table_data: str | Path = DEFAULT_FILTERBANK_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_filterbank_tables(table_data)
|
||||
self.basis = np.asarray(tables["qmf_synthesis_basis"], dtype=np.float64)
|
||||
self.taps = np.asarray(tables["qmf_synthesis_taps"], dtype=np.float64)
|
||||
if self.basis.shape != (64, 4, 128) or self.taps.shape != (64, 10, 4):
|
||||
raise ValueError("invalid QMF synthesis factorization")
|
||||
self.channels = int(channels)
|
||||
self.rank = 4
|
||||
self.history = np.zeros(
|
||||
(9, self.channels, 64, self.rank), dtype=np.float64)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.history.fill(0.0)
|
||||
|
||||
def process_chunk(self, qmf) -> np.ndarray:
|
||||
values = np.asarray(qmf, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 64):
|
||||
raise ValueError(f"expected [slots,{self.channels},64], got {values.shape}")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("QMF-synthesis input contains non-finite values")
|
||||
count = values.shape[0]
|
||||
flat = np.stack((values.real, values.imag), axis=-1).reshape(
|
||||
count * self.channels, 128)
|
||||
modulation = self.basis.reshape(64 * self.rank, 128)
|
||||
features = (flat @ modulation.T).reshape(
|
||||
count, self.channels, 64, self.rank)
|
||||
joined = np.concatenate((self.history, features), axis=0)
|
||||
output = np.zeros((count, self.channels, 64), dtype=np.float64)
|
||||
for lag in range(10):
|
||||
output += np.sum(
|
||||
joined[9 - lag:9 - lag + count]
|
||||
* self.taps[:, lag, :][None, None, :, :],
|
||||
axis=-1, dtype=np.float64)
|
||||
self.history[:] = joined[-9:]
|
||||
return output
|
||||
|
||||
|
||||
class PublicAnalysis77:
|
||||
"""Full-rate PCM to the public 77-band hybrid representation."""
|
||||
|
||||
def __init__(self, channels: int):
|
||||
self.channels = int(channels)
|
||||
if self.channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
self.qmf = QmfAnalysis(self.channels)
|
||||
self.hybrid = HybridAnalysis(self.channels)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.qmf.reset()
|
||||
self.hybrid.reset()
|
||||
|
||||
def process(self, samples) -> np.ndarray:
|
||||
values = np.asarray(samples, dtype=np.float64)
|
||||
if values.ndim == 1 and self.channels == 1:
|
||||
values = values[:, None]
|
||||
if values.ndim != 2 or values.shape[1] != self.channels:
|
||||
raise ValueError(f"samples must have shape [N,{self.channels}]")
|
||||
if len(values) % QMF_HOP:
|
||||
raise ValueError("sample count must be divisible by the 64-sample QMF hop")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("samples contain non-finite values")
|
||||
hops = values.reshape(-1, QMF_HOP, self.channels).transpose(0, 2, 1)
|
||||
return np.asarray(
|
||||
self.hybrid.process_chunk(self.qmf.process_chunk(hops)),
|
||||
dtype=np.complex128)
|
||||
|
||||
|
||||
class PublicSynthesis77:
|
||||
"""Public 77-band hybrid representation to full-rate PCM."""
|
||||
|
||||
def __init__(self, channels: int):
|
||||
self.channels = int(channels)
|
||||
if self.channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
self.hybrid = HybridSynthesis(self.channels)
|
||||
self.qmf = QmfSynthesis(self.channels)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.hybrid.reset()
|
||||
self.qmf.reset()
|
||||
|
||||
def process(self, hybrid) -> np.ndarray:
|
||||
values = np.asarray(hybrid, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, HYBRID_BANDS):
|
||||
raise ValueError(
|
||||
f"hybrid must have shape [slots,{self.channels},{HYBRID_BANDS}]")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hybrid input contains non-finite values")
|
||||
qmf = self.hybrid.process_chunk(values)
|
||||
time = self.qmf.process_chunk(qmf)
|
||||
return np.asarray(time.transpose(0, 2, 1).reshape(-1, self.channels),
|
||||
dtype=np.float64)
|
||||
|
||||
|
||||
def identity_impulse_response(sample_count: int = 4096) -> np.ndarray:
|
||||
sample_count = int(sample_count)
|
||||
if sample_count <= 0:
|
||||
raise ValueError("sample_count must be positive")
|
||||
total = ((sample_count + QMF_HOP - 1) // QMF_HOP) * QMF_HOP
|
||||
impulse = np.zeros((total, 1), dtype=np.float64)
|
||||
impulse[0, 0] = 1.0
|
||||
analysis = PublicAnalysis77(1)
|
||||
synthesis = PublicSynthesis77(1)
|
||||
return synthesis.process(analysis.process(impulse))[:, 0]
|
||||
|
||||
|
||||
@lru_cache(maxsize=2)
|
||||
def _hybrid_band_center_frequencies_hz_cached(rate: float) -> np.ndarray:
|
||||
centers = np.asarray(
|
||||
_BAND_CENTER_FREQUENCIES_HZ * (rate / SAMPLE_RATE), dtype=np.float64)
|
||||
centers.setflags(write=False)
|
||||
return centers
|
||||
|
||||
|
||||
def hybrid_band_center_frequencies_hz(
|
||||
sample_rate_hz: float = SAMPLE_RATE) -> np.ndarray:
|
||||
"""Return the 77 hybrid-band reference center frequencies."""
|
||||
rate = float(sample_rate_hz)
|
||||
if not np.isfinite(rate) or rate <= 0.0:
|
||||
raise ValueError("sample rate must be positive and finite")
|
||||
centers = _hybrid_band_center_frequencies_hz_cached(rate).copy()
|
||||
centers.setflags(write=False)
|
||||
return centers
|
||||
|
||||
|
||||
def table_info() -> dict:
|
||||
return {
|
||||
"resource": DEFAULT_FILTERBANK_DATA.name,
|
||||
"sample_rate_hz": SAMPLE_RATE,
|
||||
"qmf_bands": QMF_BANDS,
|
||||
"hybrid_bands": HYBRID_BANDS,
|
||||
"hop_samples": QMF_HOP,
|
||||
"analysis_synthesis_latency_samples": ANALYSIS_SYNTHESIS_LATENCY_SAMPLES,
|
||||
"precision": "float64/complex128",
|
||||
"fingerprint": filterbank_fingerprint(),
|
||||
"provenance": {
|
||||
"qmf": (
|
||||
"MPEG-4 AAC/SBR 64 complex QMF analysis (ISO/IEC "
|
||||
"14496-3/AMD1:2003 4.B.18.2), polyphase form of the public "
|
||||
"640-tap SBR prototype"),
|
||||
"hybrid": (
|
||||
"3GPP TS 26.405 / ETSI TS 126 405 5.2.2 Table 1 (Q=8/Q=4) "
|
||||
"with standard half-bin complex modulation"),
|
||||
"synthesis": (
|
||||
"causal left inverse of the public analysis bank "
|
||||
"(A·W = P, 577-sample QMF delay); 77→64 sparse recombination"),
|
||||
"resource": "data/rosella_kernels.npz",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
@lru_cache(maxsize=8)
|
||||
def _hybrid_gain_synthesis_dictionary_cached(count: int) -> np.ndarray:
|
||||
total = int(np.ceil(
|
||||
(ANALYSIS_SYNTHESIS_LATENCY_SAMPLES + count + 512) / QMF_HOP) * QMF_HOP)
|
||||
impulse = np.zeros((total, 1), dtype=np.float64)
|
||||
impulse[0, 0] = 1.0
|
||||
base = PublicAnalysis77(1).process(impulse)[:, 0, :]
|
||||
parameter_count = 2 * HYBRID_BANDS
|
||||
hybrid = np.zeros(
|
||||
(len(base), parameter_count, HYBRID_BANDS), dtype=np.complex128)
|
||||
for band in range(HYBRID_BANDS):
|
||||
hybrid[:, 2 * band, band] = base[:, band]
|
||||
hybrid[:, 2 * band + 1, band] = 1j * base[:, band]
|
||||
rendered = PublicSynthesis77(parameter_count).process(hybrid)
|
||||
start = ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
dictionary = np.asarray(
|
||||
rendered[start:start + count], dtype=np.float64).copy()
|
||||
dictionary.setflags(write=False)
|
||||
return dictionary
|
||||
|
||||
|
||||
def hybrid_gain_synthesis_dictionary(sample_count: int) -> np.ndarray:
|
||||
"""Return the 154-real-parameter analysis/gain/synthesis dictionary.
|
||||
|
||||
Each hybrid band contributes one real-gain and one imaginary-gain column.
|
||||
The common 961-sample filterbank latency is removed from every column.
|
||||
"""
|
||||
count = int(sample_count)
|
||||
if count <= 0:
|
||||
raise ValueError("sample_count must be positive")
|
||||
dictionary = _hybrid_gain_synthesis_dictionary_cached(count).copy()
|
||||
dictionary.setflags(write=False)
|
||||
return dictionary
|
||||
|
||||
|
||||
def project_hrir_to_hybrid_gains(
|
||||
hrir, *, embedded_delay_samples=None,
|
||||
sample_rate_hz: float = SAMPLE_RATE, ridge: float = 1.0e-3,
|
||||
) -> tuple[np.ndarray, dict]:
|
||||
"""Project FIRs and remove only a known embedded arrival delay.
|
||||
|
||||
Non-zero SOFA ``Data.Delay`` is external and must be passed as zero here.
|
||||
A positive onset separated from ``Data.IR`` is de-rotated once, then restored
|
||||
once by the runtime field. A zero-origin FIR keeps its authored complex
|
||||
phase and therefore also passes zero.
|
||||
"""
|
||||
values = np.asarray(hrir, dtype=np.float64)
|
||||
if values.ndim != 3 or values.shape[1] != 2 or values.shape[2] <= 0:
|
||||
raise ValueError("hrir must have shape [M,2,N]")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hrir contains non-finite values")
|
||||
regularization = float(ridge)
|
||||
if not np.isfinite(regularization) or regularization < 0.0:
|
||||
raise ValueError("projection ridge must be finite and non-negative")
|
||||
if embedded_delay_samples is None:
|
||||
delay = np.zeros(values.shape[:2], dtype=np.float64)
|
||||
else:
|
||||
delay = np.asarray(embedded_delay_samples, dtype=np.float64)
|
||||
if delay.shape != values.shape[:2] or not np.isfinite(delay).all():
|
||||
raise ValueError("embedded_delay_samples must have finite shape [M,2]")
|
||||
rate = float(sample_rate_hz)
|
||||
if not np.isfinite(rate) or rate <= 0.0:
|
||||
raise ValueError("sample_rate_hz must be positive and finite")
|
||||
|
||||
dictionary = _hybrid_gain_synthesis_dictionary_cached(values.shape[2])
|
||||
gram = dictionary.T @ dictionary
|
||||
scale = float(np.trace(gram)) / gram.shape[0]
|
||||
system = gram + regularization * scale * np.eye(gram.shape[0], dtype=np.float64)
|
||||
target = values.reshape(-1, values.shape[2]).T
|
||||
parameters = np.linalg.solve(system, dictionary.T @ target).T
|
||||
parts = parameters.reshape(values.shape[0], 2, 2 * HYBRID_BANDS)
|
||||
transfer = np.asarray(parts[..., 0::2] + 1j * parts[..., 1::2],
|
||||
dtype=np.complex128)
|
||||
centers = hybrid_band_center_frequencies_hz(rate)
|
||||
removal_phase = np.exp(
|
||||
2j * np.pi * delay[..., None] * centers[None, None, :] / rate)
|
||||
aligned = np.asarray(transfer * removal_phase, dtype=np.complex128)
|
||||
|
||||
reconstructed = dictionary @ parameters.T
|
||||
error = target - reconstructed
|
||||
reference_energy = np.sum(target * target, axis=0, dtype=np.float64)
|
||||
error_energy = np.sum(error * error, axis=0, dtype=np.float64)
|
||||
snr = 10.0 * np.log10(
|
||||
np.maximum(reference_energy, 1.0e-300)
|
||||
/ np.maximum(error_energy, 1.0e-300))
|
||||
report = {
|
||||
"method": "regularized public analysis/gain/synthesis dictionary",
|
||||
"dictionary_shape": list(dictionary.shape),
|
||||
"real_parameters": 2 * HYBRID_BANDS,
|
||||
"ridge": regularization,
|
||||
"embedded_delay_samples_min": float(np.min(delay)),
|
||||
"embedded_delay_samples_max": float(np.max(delay)),
|
||||
"fir_reconstruction_snr_db_median": float(np.median(snr)),
|
||||
"fir_reconstruction_snr_db_p05": float(np.percentile(snr, 5.0)),
|
||||
"fir_reconstruction_snr_db_min": float(np.min(snr)),
|
||||
"maximum_absolute_hybrid_gain": float(np.max(np.abs(aligned))),
|
||||
"precision": "float64/complex128",
|
||||
}
|
||||
return aligned, report
|
||||
|
||||
|
||||
def project_aligned_hrir_to_hybrid_gains(
|
||||
aligned_hrir, *, ridge: float = 1.0e-3) -> tuple[np.ndarray, dict]:
|
||||
return project_hrir_to_hybrid_gains(aligned_hrir, ridge=ridge)
|
||||
@@ -1,253 +0,0 @@
|
||||
"""Project-owned image-source early reflections and shared unitary FDN."""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
import math
|
||||
import numpy as np
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ShoeboxRoomConfig:
|
||||
dimensions_m: tuple[float, float, float] = (18.0, 18.0, 14.0)
|
||||
listener_position_m: tuple[float, float, float] = (9.0, 9.0, 7.0)
|
||||
wall_reflection_gain: tuple[float, float, float, float, float, float] = (
|
||||
0.62, 0.60, 0.58, 0.61, 0.52, 0.56)
|
||||
speed_of_sound_m_s: float = 343.3
|
||||
|
||||
def validate(self) -> None:
|
||||
dimensions = np.asarray(self.dimensions_m, dtype=np.float64)
|
||||
listener = np.asarray(self.listener_position_m, dtype=np.float64)
|
||||
gains = np.asarray(self.wall_reflection_gain, dtype=np.float64)
|
||||
if (dimensions.shape != (3,) or not np.isfinite(dimensions).all()
|
||||
or np.any(dimensions <= 0.0)):
|
||||
raise ValueError("room dimensions must be three positive finite values")
|
||||
if (listener.shape != (3,) or not np.isfinite(listener).all()
|
||||
or np.any(listener <= 0.0) or np.any(listener >= dimensions)):
|
||||
raise ValueError("listener must be strictly inside the shoebox")
|
||||
if (gains.shape != (6,) or not np.isfinite(gains).all()
|
||||
or np.any(np.abs(gains) >= 1.0)):
|
||||
raise ValueError("six finite wall gains must have magnitude below one")
|
||||
if not math.isfinite(self.speed_of_sound_m_s) or self.speed_of_sound_m_s <= 0.0:
|
||||
raise ValueError("speed of sound must be positive")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class EarlyReflection:
|
||||
wall: str
|
||||
direction_adm: np.ndarray
|
||||
path_distance_m: float
|
||||
extra_delay_samples: float
|
||||
reflection_gain: float
|
||||
|
||||
|
||||
_WALL_NAMES = ("left", "right", "back", "front", "floor", "ceiling")
|
||||
|
||||
|
||||
def first_order_image_sources(direction_adm, source_distance_m: float,
|
||||
sample_rate_hz: float,
|
||||
config: ShoeboxRoomConfig = ShoeboxRoomConfig()
|
||||
) -> tuple[EarlyReflection, ...]:
|
||||
"""Return six first-order image-source paths for one object."""
|
||||
config.validate()
|
||||
direction = np.asarray(direction_adm, dtype=np.float64)
|
||||
if direction.shape != (3,) or not np.isfinite(direction).all():
|
||||
raise ValueError("reflection direction must contain three finite ADM values")
|
||||
norm = float(np.linalg.norm(direction))
|
||||
if norm <= 1.0e-15:
|
||||
direction = np.asarray([0.0, 1.0, 0.0], dtype=np.float64)
|
||||
else:
|
||||
direction = direction / norm
|
||||
distance = float(source_distance_m)
|
||||
rate = float(sample_rate_hz)
|
||||
if not math.isfinite(distance) or distance <= 0.0 or not math.isfinite(rate) or rate <= 0.0:
|
||||
raise ValueError("source distance and sample rate must be positive")
|
||||
dimensions = np.asarray(config.dimensions_m, dtype=np.float64)
|
||||
listener = np.asarray(config.listener_position_m, dtype=np.float64)
|
||||
source = listener + direction * distance
|
||||
if np.any(source <= 0.0) or np.any(source >= dimensions):
|
||||
raise ValueError(
|
||||
"source lies outside the configured public shoebox; enlarge the room")
|
||||
images = []
|
||||
for axis in range(3):
|
||||
low = source.copy()
|
||||
low[axis] = -source[axis]
|
||||
high = source.copy()
|
||||
high[axis] = 2.0 * dimensions[axis] - source[axis]
|
||||
images.extend((low, high))
|
||||
result = []
|
||||
for wall, image, gain in zip(_WALL_NAMES, images, config.wall_reflection_gain):
|
||||
vector = image - listener
|
||||
path_distance = float(np.linalg.norm(vector))
|
||||
path_direction = vector / path_distance
|
||||
extra = max(0.0, (path_distance - distance)
|
||||
* rate / config.speed_of_sound_m_s)
|
||||
air = math.exp(-0.002 * max(path_distance - distance, 0.0))
|
||||
result.append(EarlyReflection(
|
||||
wall=wall,
|
||||
direction_adm=np.asarray(path_direction, dtype=np.float64),
|
||||
path_distance_m=path_distance,
|
||||
extra_delay_samples=extra,
|
||||
reflection_gain=float(gain) * air,
|
||||
))
|
||||
return tuple(result)
|
||||
|
||||
|
||||
def normalized_hadamard4() -> np.ndarray:
|
||||
return 0.5 * np.asarray([
|
||||
[1.0, 1.0, 1.0, 1.0],
|
||||
[1.0, -1.0, 1.0, -1.0],
|
||||
[1.0, 1.0, -1.0, -1.0],
|
||||
[1.0, -1.0, -1.0, 1.0],
|
||||
], dtype=np.float64)
|
||||
|
||||
|
||||
def _is_prime(value: int) -> bool:
|
||||
if value < 2:
|
||||
return False
|
||||
if value % 2 == 0:
|
||||
return value == 2
|
||||
limit = int(math.sqrt(value))
|
||||
return all(value % divisor for divisor in range(3, limit + 1, 2))
|
||||
|
||||
|
||||
def _next_prime(value: int) -> int:
|
||||
candidate = max(2, int(value))
|
||||
while not _is_prime(candidate):
|
||||
candidate += 1
|
||||
return candidate
|
||||
|
||||
|
||||
class SchroederAllpass:
|
||||
def __init__(self, delay_samples: int, gain: float):
|
||||
self.delay_samples = int(delay_samples)
|
||||
self.gain = float(gain)
|
||||
if self.delay_samples <= 0 or not 0.0 <= abs(self.gain) < 1.0:
|
||||
raise ValueError("all-pass delay must be positive and |gain| < 1")
|
||||
self.buffer = np.zeros(self.delay_samples, dtype=np.float64)
|
||||
self.position = 0
|
||||
|
||||
def reset(self) -> None:
|
||||
self.buffer.fill(0.0)
|
||||
self.position = 0
|
||||
|
||||
def process(self, values) -> np.ndarray:
|
||||
source = np.asarray(values, dtype=np.float64)
|
||||
output = np.empty_like(source)
|
||||
for index, value in enumerate(source):
|
||||
delayed = self.buffer[self.position]
|
||||
result = delayed - self.gain * value
|
||||
self.buffer[self.position] = value + self.gain * result
|
||||
self.position = (self.position + 1) % self.delay_samples
|
||||
output[index] = result
|
||||
return output
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class LateFdnConfig:
|
||||
sample_rate_hz: float = 48000.0
|
||||
rt60_seconds: float = 0.85
|
||||
damping: float = 0.32
|
||||
output_gain: float = 0.22
|
||||
delay_seconds: tuple[float, float, float, float] = (
|
||||
0.0297, 0.0371, 0.0411, 0.0437)
|
||||
allpass_seconds: tuple[float, float] = (0.0023, 0.0067)
|
||||
allpass_gain: tuple[float, float] = (0.63, 0.51)
|
||||
|
||||
|
||||
class SharedUnitaryFdn:
|
||||
"""One shared late room driven by the sum of all object room sends."""
|
||||
|
||||
def __init__(self, config: LateFdnConfig = LateFdnConfig()):
|
||||
self.config = config
|
||||
self.sample_rate_hz = float(config.sample_rate_hz)
|
||||
self.rt60_seconds = float(config.rt60_seconds)
|
||||
self.damping = float(config.damping)
|
||||
self.output_gain = float(config.output_gain)
|
||||
if (not math.isfinite(self.sample_rate_hz) or self.sample_rate_hz <= 0.0
|
||||
or not math.isfinite(self.rt60_seconds) or self.rt60_seconds <= 0.0):
|
||||
raise ValueError("FDN sample rate and RT60 must be positive and finite")
|
||||
if (not math.isfinite(self.damping) or not 0.0 <= self.damping < 1.0
|
||||
or not math.isfinite(self.output_gain)):
|
||||
raise ValueError("invalid FDN damping/output gain")
|
||||
delay_seconds = np.asarray(config.delay_seconds, dtype=np.float64)
|
||||
allpass_seconds = np.asarray(config.allpass_seconds, dtype=np.float64)
|
||||
allpass_gain = np.asarray(config.allpass_gain, dtype=np.float64)
|
||||
if (delay_seconds.shape != (4,) or not np.isfinite(delay_seconds).all()
|
||||
or np.any(delay_seconds <= 0.0)):
|
||||
raise ValueError("FDN requires four positive finite delay times")
|
||||
if (allpass_seconds.shape != (2,) or not np.isfinite(allpass_seconds).all()
|
||||
or np.any(allpass_seconds <= 0.0)):
|
||||
raise ValueError("FDN requires two positive finite all-pass delay times")
|
||||
if (allpass_gain.shape != (2,) or not np.isfinite(allpass_gain).all()
|
||||
or np.any(np.abs(allpass_gain) >= 1.0)):
|
||||
raise ValueError("FDN requires two finite all-pass gains with magnitude below one")
|
||||
self.matrix = normalized_hadamard4()
|
||||
self.delays = np.asarray([
|
||||
_next_prime(round(seconds * self.sample_rate_hz))
|
||||
for seconds in delay_seconds
|
||||
], dtype=np.int32)
|
||||
self.feedback_gain = np.power(
|
||||
10.0, -3.0 * self.delays / (self.rt60_seconds * self.sample_rate_hz)
|
||||
).astype(np.float64)
|
||||
self.buffers = [np.zeros(int(delay), dtype=np.float64) for delay in self.delays]
|
||||
self.positions = np.zeros(4, dtype=np.int32)
|
||||
self.damping_state = np.zeros(4, dtype=np.float64)
|
||||
self.input_vector = 0.5 * np.asarray([1.0, -1.0, 1.0, 1.0], dtype=np.float64)
|
||||
self.output_matrix = 0.5 * np.asarray([
|
||||
[1.0, 1.0, -1.0, -1.0],
|
||||
[1.0, -1.0, 1.0, -1.0],
|
||||
], dtype=np.float64)
|
||||
self.diffusers = [
|
||||
SchroederAllpass(
|
||||
_next_prime(round(seconds * self.sample_rate_hz)), gain)
|
||||
for seconds, gain in zip(allpass_seconds, allpass_gain)
|
||||
]
|
||||
|
||||
@property
|
||||
def tail_samples(self) -> int:
|
||||
return int(math.ceil(1.5 * self.rt60_seconds * self.sample_rate_hz))
|
||||
|
||||
def reset(self) -> None:
|
||||
for buffer in self.buffers:
|
||||
buffer.fill(0.0)
|
||||
self.positions.fill(0)
|
||||
self.damping_state.fill(0.0)
|
||||
for diffuser in self.diffusers:
|
||||
diffuser.reset()
|
||||
|
||||
def process(self, mono) -> np.ndarray:
|
||||
values = np.asarray(mono, dtype=np.float64)
|
||||
if values.ndim != 1 or not np.isfinite(values).all():
|
||||
raise ValueError("FDN input must be one finite mono vector")
|
||||
diffused = values
|
||||
for diffuser in self.diffusers:
|
||||
diffused = diffuser.process(diffused)
|
||||
output = np.empty((len(values), 2), dtype=np.float64)
|
||||
for sample, value in enumerate(diffused):
|
||||
delayed = np.asarray([
|
||||
self.buffers[line][int(self.positions[line])]
|
||||
for line in range(4)
|
||||
], dtype=np.float64)
|
||||
self.damping_state = (
|
||||
self.damping * self.damping_state + (1.0 - self.damping) * delayed)
|
||||
output[sample] = self.output_gain * (self.output_matrix @ self.damping_state)
|
||||
feedback = self.matrix @ (self.damping_state * self.feedback_gain)
|
||||
write = self.input_vector * value + feedback
|
||||
for line in range(4):
|
||||
position = int(self.positions[line])
|
||||
self.buffers[line][position] = write[line]
|
||||
self.positions[line] = (position + 1) % int(self.delays[line])
|
||||
return output
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"name": "SharedUnitaryFdn",
|
||||
"sample_rate_hz": self.sample_rate_hz,
|
||||
"rt60_seconds": self.rt60_seconds,
|
||||
"delay_samples": [int(value) for value in self.delays],
|
||||
"feedback_gain": [float(value) for value in self.feedback_gain],
|
||||
"matrix_unitarity_max_error": float(
|
||||
np.max(np.abs(self.matrix.T @ self.matrix - np.eye(4)))),
|
||||
"allpass_delay_samples": [value.delay_samples for value in self.diffusers],
|
||||
"precision": "float64",
|
||||
}
|
||||
@@ -1,115 +0,0 @@
|
||||
"""Project-owned distance policy for the public SOFA renderer."""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
import math
|
||||
import numpy as np
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class DistanceState:
|
||||
profile: str
|
||||
normalized_radius: float
|
||||
reference_distance_m: float
|
||||
physical_distance_m: float
|
||||
direction_adm: np.ndarray
|
||||
|
||||
|
||||
class ReferenceDistanceProfileV1:
|
||||
"""Reference behavior, not a claim about any external public standard."""
|
||||
|
||||
DISTANCE_M = {
|
||||
"near": 1.00000465,
|
||||
"mid": 2.19327927,
|
||||
"far": 6.40177584,
|
||||
}
|
||||
MINIMUM_DISTANCE_M = 0.10
|
||||
# Distance profiles in the reference renderer are presentation presets,
|
||||
# not an instruction to attenuate already-authored programme PCM by 1/r.
|
||||
# Use an energy-normalized dry/room crossfade instead. The coefficient is
|
||||
# an explicit project calibration target.
|
||||
ROOM_ENERGY_COUPLING_PER_M2 = 0.01318359375
|
||||
PUBLIC_ROOM_CALIBRATION_GAIN = 1.4
|
||||
# Public, project-owned room coupling; it is not a SOFA or Dolby constant.
|
||||
LATE_SEND = {
|
||||
"near": 0.06,
|
||||
"mid": 0.16,
|
||||
"far": 0.28,
|
||||
}
|
||||
|
||||
@classmethod
|
||||
def validate_profile(cls, profile: str) -> str:
|
||||
value = str(profile).strip().lower()
|
||||
if value not in cls.DISTANCE_M:
|
||||
raise ValueError("distance profile must be near, mid, or far")
|
||||
return value
|
||||
|
||||
@classmethod
|
||||
def map_adm_position(cls, position, profile: str) -> DistanceState:
|
||||
name = cls.validate_profile(profile)
|
||||
values = np.asarray(position, dtype=np.float64)
|
||||
if values.shape != (3,) or not np.isfinite(values).all():
|
||||
raise ValueError("ADM position must contain three finite Cartesian values")
|
||||
radius = float(np.linalg.norm(values))
|
||||
direction = (values / radius if radius > 1.0e-15
|
||||
else np.asarray([0.0, 1.0, 0.0], dtype=np.float64))
|
||||
reference = float(cls.DISTANCE_M[name])
|
||||
distance = max(float(cls.MINIMUM_DISTANCE_M), radius * reference)
|
||||
return DistanceState(
|
||||
profile=name,
|
||||
normalized_radius=radius,
|
||||
reference_distance_m=reference,
|
||||
physical_distance_m=distance,
|
||||
direction_adm=np.asarray(direction, dtype=np.float64),
|
||||
)
|
||||
|
||||
@staticmethod
|
||||
def inverse_distance_gain(measurement_radius_m: float,
|
||||
path_distance_m: float) -> float:
|
||||
radius = float(measurement_radius_m)
|
||||
distance = float(path_distance_m)
|
||||
if not (math.isfinite(radius) and math.isfinite(distance)):
|
||||
raise ValueError("measurement and path distances must be finite")
|
||||
if radius <= 0.0 or distance <= 0.0:
|
||||
raise ValueError("measurement and path distances must be positive")
|
||||
return radius / distance
|
||||
|
||||
@classmethod
|
||||
def direct_level_gain(cls, state: DistanceState) -> float:
|
||||
"""Programme-normalized direct level for a distance presentation.
|
||||
|
||||
Near is the SOFA reference response. Mid/Far use an equal-power dry
|
||||
coefficient rather than a physical free-field 1/r attenuation. Room
|
||||
distance still changes through image-path lengths and late send.
|
||||
"""
|
||||
if state.profile == "near":
|
||||
return 1.0
|
||||
distance = float(state.physical_distance_m)
|
||||
return 1.0 / math.sqrt(
|
||||
1.0 + cls.ROOM_ENERGY_COUPLING_PER_M2 * distance * distance)
|
||||
|
||||
@classmethod
|
||||
def room_calibration_gain(cls, state: DistanceState) -> float:
|
||||
del state
|
||||
return float(cls.PUBLIC_ROOM_CALIBRATION_GAIN)
|
||||
|
||||
@classmethod
|
||||
def late_send(cls, state: DistanceState) -> float:
|
||||
base = float(cls.LATE_SEND[state.profile])
|
||||
radial = math.sqrt(max(state.normalized_radius, 0.0))
|
||||
return base * min(max(radial, 0.25), 1.5)
|
||||
|
||||
@classmethod
|
||||
def info(cls) -> dict:
|
||||
return {
|
||||
"name": "ReferenceDistanceProfileV1",
|
||||
"reference_distance_m": dict(cls.DISTANCE_M),
|
||||
"minimum_distance_m": cls.MINIMUM_DISTANCE_M,
|
||||
"direct_level_policy": (
|
||||
"Near unity; Mid/Far equal-power dry coefficient, never raw 1/r "
|
||||
"programme attenuation"),
|
||||
"room_energy_coupling_per_m2": cls.ROOM_ENERGY_COUPLING_PER_M2,
|
||||
"public_room_calibration_gain": cls.PUBLIC_ROOM_CALIBRATION_GAIN,
|
||||
"late_send": dict(cls.LATE_SEND),
|
||||
"standard_claim": False,
|
||||
}
|
||||
+4
-8
@@ -13,7 +13,7 @@ output_scale:
|
||||
"""
|
||||
import numpy as np
|
||||
|
||||
from joc_decode import DC_FILTER_DMX_CONFIGS, parse_joc, diff_decode, dequantize
|
||||
from joc_decode import parse_joc, diff_decode, dequantize
|
||||
from joc_qmf import (N, QMF5_WINDOW, qmf_analysis_frame,
|
||||
surround_post_frame, interp_matrix)
|
||||
from evo_unpack import unpack_evolution
|
||||
@@ -60,12 +60,11 @@ class JocRenderer:
|
||||
mix_dq = dequantize(out, mix_q)
|
||||
return out, mix_q, mix_dq
|
||||
|
||||
def qmf_x(self, bed5, phase_new=0.0625, apply_dc_filter=True):
|
||||
def qmf_x(self, bed5, phase_new=0.0625):
|
||||
"""核心 5ch PCM → 对象矩阵使用的复数 QMF ``x``。
|
||||
|
||||
先以 float32 对当前帧应用 phase;phase 变化时仅前 256 个样本从旧值
|
||||
线性过渡。缩放后的 L/R/C 延迟 10 槽,Ls/Rs 不延迟,再进入分析 QMF。
|
||||
``apply_dc_filter`` 控制 Ls/Rs band-0 的 21-tap DC 补偿。
|
||||
"""
|
||||
pcm = np.asarray(bed5, dtype=np.float32)
|
||||
if pcm.shape != (5, 1536):
|
||||
@@ -88,8 +87,7 @@ class JocRenderer:
|
||||
x, self._analysis_fifo = qmf_analysis_frame(self._analysis_fifo, blocks)
|
||||
self._analysis_phase = new_phase
|
||||
x[3:], self._surround_qmf_delay, self._surround_dc_hist = surround_post_frame(
|
||||
x[3:], self._surround_qmf_delay, self._surround_dc_hist,
|
||||
apply_dc_filter=apply_dc_filter)
|
||||
x[3:], self._surround_qmf_delay, self._surround_dc_hist)
|
||||
return x
|
||||
|
||||
def object_z(self, out, mix_dq, x):
|
||||
@@ -246,9 +244,7 @@ class JocRenderer:
|
||||
def render_subpayloads(self, subs, bed5_pcm, lfe_pcm=None):
|
||||
"""以已拆出的 EMDF payload 字典渲染一帧,避免绑定 transport 容器。"""
|
||||
out, mix_q, mix_dq = self.decode_subpayloads(subs)
|
||||
x = self.qmf_x(
|
||||
bed5_pcm,
|
||||
apply_dc_filter=out["dmx_config_idx"] in DC_FILTER_DMX_CONFIGS)
|
||||
x = self.qmf_x(bed5_pcm)
|
||||
self._last_x = x.copy()
|
||||
z_all = self.object_z(out, mix_dq, x)
|
||||
# joc_clipgain 在对象逆 QMF 后应用,并与用户 output_scale 分离。
|
||||
|
||||
@@ -1,308 +0,0 @@
|
||||
"""Rosella .personalized_headphone binaural renderer.
|
||||
|
||||
Rosella JSON 解析由本项目自行实现(src/rosella_model.py),不调用任何 Dolby
|
||||
软件;.personalized_headphone 是用户经官方软件个性化扫描得到的模型文件。
|
||||
该路径与 SOFA 路径各自独立完成 HRTF/room 参数求值,只在最外层的 JOC 调度
|
||||
(1536-sample 帧缓冲、sample-timed OAMD timeline、512-sample 参数更新、输出
|
||||
包装)处汇合。
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import math
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
from binaural_metadata import OamdPositionTimeline
|
||||
from binaural_native_renderer import NativeBinauralDsp
|
||||
from rosella_core import RosellaRenderer
|
||||
from rosella_direct import BINAURAL_PROFILE_NAMES
|
||||
from rosella_filterbank import (
|
||||
DEFAULT_KERNEL_DATA,
|
||||
HybridAnalysis,
|
||||
HybridSynthesis,
|
||||
QmfAnalysis,
|
||||
QmfSynthesis,
|
||||
)
|
||||
from rosella_model import RosellaModel, load_personalized_headphone
|
||||
|
||||
SAMPLE_RATE = 48000
|
||||
FRAME_SAMPLES = 1536
|
||||
ROSSELLA_BLOCK_SAMPLES = 512
|
||||
QMF_HOP_SAMPLES = 64
|
||||
ROSSELLA_LATENCY_SAMPLES = 961
|
||||
SOURCE_CHANNELS = 16
|
||||
OUTPUT_CHANNELS = 2
|
||||
PROJECT_DIR = Path(__file__).resolve().parent.parent
|
||||
DEFAULT_PERSONALIZED_HEADPHONE = (
|
||||
PROJECT_DIR / "HRTF" / "binaural.personalized_headphone")
|
||||
|
||||
|
||||
def _sha256_file(path: Path) -> str:
|
||||
digest = hashlib.sha256()
|
||||
with path.open("rb") as stream:
|
||||
for block in iter(lambda: stream.read(1 << 20), b""):
|
||||
digest.update(block)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def resolve_personalized_headphone(path: str | Path | None = None) -> Path:
|
||||
target = (DEFAULT_PERSONALIZED_HEADPHONE if path is None
|
||||
else Path(path).expanduser().resolve())
|
||||
if not target.is_file():
|
||||
raise FileNotFoundError(
|
||||
f"未找到双耳模型:{target}\n"
|
||||
"请将兼容模型保存为 HRTF/binaural.personalized_headphone,"
|
||||
"或通过参数指定文件。"
|
||||
)
|
||||
return target
|
||||
|
||||
|
||||
class RosellaBinauralRenderer:
|
||||
"""Render interleaved LFE plus fifteen objects to stereo."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
personalized_headphone: str | Path | RosellaModel,
|
||||
*,
|
||||
mode: str = "mid",
|
||||
kernel_data: str | Path = DEFAULT_KERNEL_DATA,
|
||||
object_delay_samples: int = 1473,
|
||||
tail_seconds: float = 5.0,
|
||||
output_gain: float = 1.0,
|
||||
chunk_frames: int = 64,
|
||||
room_impulse_slots: int = 4096,
|
||||
backend: str = "python",
|
||||
native_library=None):
|
||||
if mode not in BINAURAL_PROFILE_NAMES:
|
||||
raise ValueError("binaural mode must be near, mid, or far")
|
||||
if int(object_delay_samples) < 0:
|
||||
raise ValueError("object_delay_samples must be non-negative")
|
||||
if float(tail_seconds) < 0.0:
|
||||
raise ValueError("tail_seconds must be non-negative")
|
||||
if int(chunk_frames) <= 0:
|
||||
raise ValueError("chunk_frames must be positive")
|
||||
if not math.isfinite(float(output_gain)):
|
||||
raise ValueError("output_gain must be finite")
|
||||
if backend not in ("auto", "native", "python"):
|
||||
raise ValueError("backend must be auto, native, or python")
|
||||
|
||||
if isinstance(personalized_headphone, RosellaModel):
|
||||
self.model = personalized_headphone
|
||||
self.model_path = Path(self.model.source_path)
|
||||
else:
|
||||
self.model_path = resolve_personalized_headphone(personalized_headphone)
|
||||
self.model = load_personalized_headphone(self.model_path)
|
||||
if self.model.sample_rate != SAMPLE_RATE:
|
||||
raise ValueError(
|
||||
f"Rosella model sample rate must be {SAMPLE_RATE}, got {self.model.sample_rate}")
|
||||
|
||||
self.mode = mode
|
||||
self.profile_index = BINAURAL_PROFILE_NAMES[mode]
|
||||
self.kernel_data = Path(kernel_data).expanduser().resolve()
|
||||
self.kernel_data_sha256 = _sha256_file(self.kernel_data)
|
||||
self.object_delay_samples = int(object_delay_samples)
|
||||
self.tail_seconds = float(tail_seconds)
|
||||
self.output_gain = np.float64(output_gain)
|
||||
self.chunk_frames = int(chunk_frames)
|
||||
self.chunk_samples = self.chunk_frames * FRAME_SAMPLES
|
||||
|
||||
self.native_dsp = None
|
||||
self.backend_fallback = None
|
||||
if backend in ("auto", "native"):
|
||||
try:
|
||||
self.native_dsp = NativeBinauralDsp(
|
||||
self.model, library_path=native_library,
|
||||
kernel_data=self.kernel_data)
|
||||
except (AttributeError, OSError, RuntimeError) as exc:
|
||||
if backend == "native":
|
||||
raise RuntimeError(f"native binaural backend unavailable: {exc}") from exc
|
||||
self.backend_fallback = str(exc)
|
||||
if self.native_dsp is not None:
|
||||
self.dsp_backend = "native"
|
||||
self.qmf_analysis = None
|
||||
self.hybrid_analysis = None
|
||||
self.hybrid_synthesis = None
|
||||
self.qmf_synthesis = None
|
||||
self.core = RosellaRenderer(
|
||||
self.model, SOURCE_CHANNELS, create_room=False)
|
||||
else:
|
||||
self.dsp_backend = "python"
|
||||
self.qmf_analysis = QmfAnalysis(SOURCE_CHANNELS, self.kernel_data)
|
||||
self.hybrid_analysis = HybridAnalysis(SOURCE_CHANNELS, self.kernel_data)
|
||||
self.core = RosellaRenderer(
|
||||
self.model, SOURCE_CHANNELS,
|
||||
room_impulse_slots=room_impulse_slots)
|
||||
self.hybrid_synthesis = HybridSynthesis(OUTPUT_CHANNELS, self.kernel_data)
|
||||
self.qmf_synthesis = QmfSynthesis(OUTPUT_CHANNELS, self.kernel_data)
|
||||
self.timeline = OamdPositionTimeline(15)
|
||||
|
||||
self._input_buffer = np.empty(
|
||||
(self.chunk_samples, SOURCE_CHANNELS), dtype=np.float64)
|
||||
self._buffer_used = 0
|
||||
self.input_samples = 0
|
||||
self.processed_input_samples = 0
|
||||
self.raw_output_samples = 0
|
||||
self.output_samples = 0
|
||||
self.finished = False
|
||||
self.metadata_block_updates = 0
|
||||
|
||||
def _append_input(self, samples: np.ndarray) -> list[np.ndarray]:
|
||||
outputs = []
|
||||
source = np.asarray(samples, dtype=np.float64)
|
||||
position = 0
|
||||
while position < len(source):
|
||||
count = min(self.chunk_samples - self._buffer_used,
|
||||
len(source) - position)
|
||||
self._input_buffer[self._buffer_used:self._buffer_used + count] = (
|
||||
source[position:position + count])
|
||||
self._buffer_used += count
|
||||
position += count
|
||||
if self._buffer_used == self.chunk_samples:
|
||||
outputs.append(self._process_samples(self._input_buffer))
|
||||
self._buffer_used = 0
|
||||
return outputs
|
||||
|
||||
def render_frame(self, objects16, payload=None, metadata_offset=None,
|
||||
*, outer_sample_offset=0) -> np.ndarray:
|
||||
"""Submit one 1536-sample reconstructed frame and its ID11 payload."""
|
||||
if self.finished:
|
||||
raise RuntimeError("binaural renderer is already finished")
|
||||
source = np.asarray(objects16)
|
||||
if source.shape != (FRAME_SAMPLES, SOURCE_CHANNELS):
|
||||
raise ValueError(
|
||||
f"binaural frame must have shape ({FRAME_SAMPLES},{SOURCE_CHANNELS}), "
|
||||
f"got {source.shape}")
|
||||
frame_start = self.input_samples
|
||||
metadata_delay = (self.object_delay_samples if metadata_offset is None
|
||||
else int(metadata_offset))
|
||||
if metadata_delay < 0:
|
||||
raise ValueError("metadata_offset must be non-negative")
|
||||
if payload is not None:
|
||||
self.timeline.submit_payload(
|
||||
payload,
|
||||
frame_start_sample=frame_start,
|
||||
outer_sample_offset=int(outer_sample_offset),
|
||||
object_delay_samples=metadata_delay,
|
||||
processed_sample=self.processed_input_samples,
|
||||
)
|
||||
self.metadata_block_updates += 1
|
||||
self.input_samples += FRAME_SAMPLES
|
||||
chunks = self._append_input(source)
|
||||
if not chunks:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
return np.concatenate(chunks, axis=0) if len(chunks) > 1 else chunks[0]
|
||||
|
||||
def _set_block_parameters(self, sample: int):
|
||||
positions = self.timeline.positions_at(sample)
|
||||
self.core.set_source(0, (0.0, 1.0, 0.0), special_lfe=True)
|
||||
for object_index in range(15):
|
||||
self.core.set_source(
|
||||
object_index + 1, positions[object_index], self.profile_index)
|
||||
|
||||
def _process_samples(self, source: np.ndarray) -> np.ndarray:
|
||||
values = np.asarray(source, dtype=np.float64)
|
||||
if values.ndim != 2 or values.shape[1] != SOURCE_CHANNELS:
|
||||
raise ValueError(f"expected [samples,{SOURCE_CHANNELS}], got {values.shape}")
|
||||
if len(values) % ROSSELLA_BLOCK_SAMPLES:
|
||||
raise ValueError("binaural input must be divisible by 512 samples")
|
||||
blocks = len(values) // ROSSELLA_BLOCK_SAMPLES
|
||||
block_base = self.processed_input_samples
|
||||
|
||||
if self.native_dsp is not None:
|
||||
stereo = np.empty((len(values), OUTPUT_CHANNELS), dtype=np.float64)
|
||||
for block in range(blocks):
|
||||
sample = block_base + block * ROSSELLA_BLOCK_SAMPLES
|
||||
self._set_block_parameters(sample)
|
||||
start = block * ROSSELLA_BLOCK_SAMPLES
|
||||
stop = start + ROSSELLA_BLOCK_SAMPLES
|
||||
stereo[start:stop] = self.native_dsp.process_block(
|
||||
values[start:stop], self.core.gains, self.core.room_sends,
|
||||
self.output_gain)
|
||||
else:
|
||||
hops = values.reshape(
|
||||
blocks, ROSSELLA_BLOCK_SAMPLES // QMF_HOP_SAMPLES,
|
||||
QMF_HOP_SAMPLES, SOURCE_CHANNELS,
|
||||
).transpose(0, 1, 3, 2).reshape(
|
||||
blocks * (ROSSELLA_BLOCK_SAMPLES // QMF_HOP_SAMPLES),
|
||||
SOURCE_CHANNELS, QMF_HOP_SAMPLES)
|
||||
hybrid = self.hybrid_analysis.process_chunk(
|
||||
self.qmf_analysis.process_chunk(hops))
|
||||
direct = np.empty((blocks * 8, OUTPUT_CHANNELS, 77), dtype=np.complex128)
|
||||
room_send = np.empty((blocks * 8, 77), dtype=np.complex128)
|
||||
for block in range(blocks):
|
||||
sample = block_base + block * ROSSELLA_BLOCK_SAMPLES
|
||||
self._set_block_parameters(sample)
|
||||
start = block * 8
|
||||
stop = start + 8
|
||||
direct[start:stop], room_send[start:stop] = (
|
||||
self.core.direct_and_send_static(hybrid[start:stop]))
|
||||
rendered = direct + self.core.room.process_chunk(room_send)
|
||||
time_bands = self.qmf_synthesis.process_chunk(
|
||||
self.hybrid_synthesis.process_chunk(rendered))
|
||||
stereo = time_bands.transpose(0, 2, 1).reshape(
|
||||
blocks * ROSSELLA_BLOCK_SAMPLES, OUTPUT_CHANNELS)
|
||||
stereo *= self.output_gain
|
||||
|
||||
skip = max(0, min(
|
||||
len(stereo), ROSSELLA_LATENCY_SAMPLES - self.raw_output_samples))
|
||||
self.raw_output_samples += len(stereo)
|
||||
self.processed_input_samples += len(values)
|
||||
output = stereo[skip:]
|
||||
self.output_samples += len(output)
|
||||
return output
|
||||
|
||||
def finish(self) -> np.ndarray:
|
||||
"""Process pending source samples and preserve the configured room tail."""
|
||||
if self.finished:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
outputs: list[np.ndarray] = []
|
||||
if self._buffer_used:
|
||||
outputs.append(self._process_samples(
|
||||
self._input_buffer[:self._buffer_used]))
|
||||
self._buffer_used = 0
|
||||
flush_samples = math.ceil(
|
||||
(self.tail_seconds * SAMPLE_RATE
|
||||
+ ROSSELLA_LATENCY_SAMPLES + ROSSELLA_BLOCK_SAMPLES)
|
||||
/ ROSSELLA_BLOCK_SAMPLES) * ROSSELLA_BLOCK_SAMPLES
|
||||
while flush_samples:
|
||||
count = min(flush_samples, self.chunk_samples)
|
||||
zero = np.zeros((count, SOURCE_CHANNELS), dtype=np.float64)
|
||||
outputs.append(self._process_samples(zero))
|
||||
flush_samples -= count
|
||||
self.finished = True
|
||||
nonempty = [value for value in outputs if len(value)]
|
||||
if not nonempty:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
return np.concatenate(nonempty, axis=0)
|
||||
|
||||
def close(self):
|
||||
if self.native_dsp is not None:
|
||||
self.native_dsp.close()
|
||||
self.finished = True
|
||||
|
||||
@property
|
||||
def backend_info(self) -> dict:
|
||||
return {
|
||||
"name": self.dsp_backend,
|
||||
"precision": "float64/complex128",
|
||||
"fallback_reason": self.backend_fallback,
|
||||
"library": (str(self.native_dsp.library_path)
|
||||
if self.native_dsp is not None else None),
|
||||
"model": str(self.model_path.resolve()),
|
||||
"model_coefficients": int(len(self.model.coefficients)),
|
||||
"model_coefficient_sha256": self.model.coefficient_sha256,
|
||||
"model_version": self.model.coefficient_version,
|
||||
"kernel_data": str(self.kernel_data),
|
||||
"kernel_data_sha256": self.kernel_data_sha256,
|
||||
"mode": self.mode,
|
||||
"latency_compensated_samples": ROSSELLA_LATENCY_SAMPLES,
|
||||
"object_delay_samples": self.object_delay_samples,
|
||||
"tail_seconds": self.tail_seconds,
|
||||
"metadata_payloads": self.timeline.payload_count,
|
||||
"metadata_position_transitions": self.timeline.transition_count,
|
||||
"input_samples": self.input_samples,
|
||||
"processed_samples_including_flush": self.processed_input_samples,
|
||||
"output_samples_before_tail_trim": self.output_samples,
|
||||
}
|
||||
@@ -1,552 +0,0 @@
|
||||
"""Stateful public SOFA binaural renderer.
|
||||
|
||||
The runtime topology mirrors the existing multi-object binaural path:
|
||||
64-QMF -> 77 hybrid -> per-object directional transfer -> stereo synthesis.
|
||||
The HRTF parameter source is a SOFA-derived fifth-order field. Early
|
||||
reflections and the late room use project-owned behavior.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
import math
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
from public_filterbank import (
|
||||
ANALYSIS_SYNTHESIS_LATENCY_SAMPLES,
|
||||
HYBRID_BANDS,
|
||||
QMF_HOP,
|
||||
PublicAnalysis77,
|
||||
PublicSynthesis77,
|
||||
table_info as filterbank_table_info,
|
||||
)
|
||||
from public_room import (
|
||||
LateFdnConfig,
|
||||
SharedUnitaryFdn,
|
||||
ShoeboxRoomConfig,
|
||||
first_order_image_sources,
|
||||
)
|
||||
from reference_distance import DistanceState, ReferenceDistanceProfileV1
|
||||
from sofa_canonical import CanonicalHrtf
|
||||
from sofa_hrtf_field import (
|
||||
DEFAULT_ORDER,
|
||||
DEFAULT_PROJECTION_RIDGE,
|
||||
DEFAULT_SH_RIDGE,
|
||||
SofaHrtfField,
|
||||
compile_sofa_hrtf,
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class HybridPath:
|
||||
label: str
|
||||
delay_slots: np.ndarray # whole-QMF delay per ear, [2]
|
||||
transfer: np.ndarray # [ear,77], includes residual delay and HRTF delay
|
||||
|
||||
def __post_init__(self):
|
||||
slots = np.asarray(self.delay_slots)
|
||||
if slots.shape == ():
|
||||
slots = np.repeat(slots, 2)
|
||||
if slots.shape != (2,) or slots.dtype.kind not in "iu":
|
||||
raise ValueError("hybrid path delay_slots must contain two integers")
|
||||
slots = np.asarray(slots, dtype=np.int64)
|
||||
transfer = np.asarray(self.transfer, dtype=np.complex128)
|
||||
if transfer.shape != (2, HYBRID_BANDS) or not np.isfinite(transfer).all():
|
||||
raise ValueError("hybrid path transfer must have finite shape [2,77]")
|
||||
if np.any(slots < 0):
|
||||
raise ValueError("hybrid path delay_slots must be non-negative")
|
||||
slots.setflags(write=False)
|
||||
transfer.setflags(write=False)
|
||||
object.__setattr__(self, "delay_slots", slots)
|
||||
object.__setattr__(self, "transfer", transfer)
|
||||
|
||||
|
||||
class HybridObjectPathRenderer:
|
||||
"""Per-object hybrid histories for direct and image-source paths."""
|
||||
|
||||
def __init__(self, source_count: int, *, history_slots: int = 256,
|
||||
transition_slots: int = 8):
|
||||
self.source_count = int(source_count)
|
||||
self.history_slots = int(history_slots)
|
||||
self.transition_slots = int(transition_slots)
|
||||
if min(self.source_count, self.history_slots) <= 0 or self.transition_slots < 0:
|
||||
raise ValueError("invalid hybrid path renderer dimensions")
|
||||
self.history = np.zeros(
|
||||
(self.source_count, self.history_slots, HYBRID_BANDS), dtype=np.complex128)
|
||||
self.position = 0
|
||||
self.current: list[tuple[HybridPath, ...]] = [tuple() for _ in range(self.source_count)]
|
||||
self.target: list[tuple[HybridPath, ...] | None] = [None] * self.source_count
|
||||
self.fade_position = np.zeros(self.source_count, dtype=np.int32)
|
||||
self.fade_total = np.zeros(self.source_count, dtype=np.int32)
|
||||
self.processed_slots = 0
|
||||
|
||||
def reset(self) -> None:
|
||||
self.history.fill(0.0)
|
||||
self.position = 0
|
||||
self.target = [None] * self.source_count
|
||||
self.fade_position.fill(0)
|
||||
self.fade_total.fill(0)
|
||||
self.processed_slots = 0
|
||||
|
||||
def set_paths(self, source: int, paths, *, fade_slots: int | None = None) -> None:
|
||||
source = int(source)
|
||||
if not 0 <= source < self.source_count:
|
||||
raise IndexError(source)
|
||||
values = tuple(paths)
|
||||
for path in values:
|
||||
if np.any(path.delay_slots >= self.history_slots):
|
||||
raise ValueError(
|
||||
f"path {path.label!r} needs {path.delay_slots.tolist()} slots, "
|
||||
f"history capacity is {self.history_slots}")
|
||||
fade = self.transition_slots if fade_slots is None else int(fade_slots)
|
||||
if fade < 0:
|
||||
raise ValueError("path fade must be non-negative")
|
||||
if self.target[source] is not None:
|
||||
# Normal 512-sample updates complete an 8-slot transition exactly.
|
||||
# If a caller updates faster, use the previous target as the new
|
||||
# stable side rather than resetting signal history.
|
||||
self.current[source] = self.target[source]
|
||||
self.target[source] = None
|
||||
if not self.processed_slots or fade == 0:
|
||||
self.current[source] = values
|
||||
self.target[source] = None
|
||||
self.fade_position[source] = 0
|
||||
self.fade_total[source] = 0
|
||||
else:
|
||||
self.target[source] = values
|
||||
self.fade_position[source] = 0
|
||||
self.fade_total[source] = fade
|
||||
|
||||
def _render_paths(self, source: int, paths: tuple[HybridPath, ...]) -> np.ndarray:
|
||||
result = np.zeros((2, HYBRID_BANDS), dtype=np.complex128)
|
||||
for path in paths:
|
||||
indices = (self.position - path.delay_slots) % self.history_slots
|
||||
delayed = self.history[source, indices, :]
|
||||
result += delayed * path.transfer
|
||||
return result
|
||||
|
||||
def process(self, hybrid) -> np.ndarray:
|
||||
values = np.asarray(hybrid, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.source_count, HYBRID_BANDS):
|
||||
raise ValueError(
|
||||
f"hybrid input must have shape [slots,{self.source_count},77]")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hybrid input contains non-finite values")
|
||||
output = np.zeros((len(values), 2, HYBRID_BANDS), dtype=np.complex128)
|
||||
for slot in range(len(values)):
|
||||
self.history[:, self.position, :] = values[slot]
|
||||
for source in range(self.source_count - 1, -1, -1):
|
||||
current = self._render_paths(source, self.current[source])
|
||||
target_paths = self.target[source]
|
||||
if target_paths is None:
|
||||
output[slot] += current
|
||||
continue
|
||||
target = self._render_paths(source, target_paths)
|
||||
self.fade_position[source] += 1
|
||||
amount = min(
|
||||
1.0, self.fade_position[source] / float(self.fade_total[source]))
|
||||
output[slot] += current * (1.0 - amount) + target * amount
|
||||
if self.fade_position[source] >= self.fade_total[source]:
|
||||
self.current[source] = target_paths
|
||||
self.target[source] = None
|
||||
self.fade_position[source] = 0
|
||||
self.fade_total[source] = 0
|
||||
self.position = (self.position + 1) % self.history_slots
|
||||
self.processed_slots += 1
|
||||
return output
|
||||
|
||||
|
||||
class _StereoDelay:
|
||||
def __init__(self, delay_samples: int):
|
||||
self.delay_samples = int(delay_samples)
|
||||
if self.delay_samples < 0:
|
||||
raise ValueError("delay must be non-negative")
|
||||
self.state = np.zeros((self.delay_samples, 2), dtype=np.float64)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.state.fill(0.0)
|
||||
|
||||
def process(self, values) -> np.ndarray:
|
||||
source = np.asarray(values, dtype=np.float64)
|
||||
if source.ndim != 2 or source.shape[1] != 2:
|
||||
raise ValueError("stereo delay input must have shape [samples,2]")
|
||||
if self.delay_samples == 0:
|
||||
return source.copy()
|
||||
joined = np.concatenate((self.state, source), axis=0)
|
||||
output = joined[:len(source)].copy()
|
||||
self.state = joined[len(source):len(source) + self.delay_samples].copy()
|
||||
return output
|
||||
|
||||
|
||||
class SofaBinauralBackend:
|
||||
"""SOFA-derived public 77-band/SH renderer with public room processing."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
field: SofaHrtfField,
|
||||
*,
|
||||
source_count: int = 16,
|
||||
sample_rate_hz: float = 48000.0,
|
||||
default_profile: str = "mid",
|
||||
enable_early_reflections: bool = True,
|
||||
enable_late_room: bool = True,
|
||||
room_config: ShoeboxRoomConfig = ShoeboxRoomConfig(),
|
||||
fdn_config: LateFdnConfig | None = None,
|
||||
transition_slots: int = 8,
|
||||
history_slots: int = 256,
|
||||
output_gain: float = 1.0):
|
||||
self.source_count = int(source_count)
|
||||
self.sample_rate_hz = float(sample_rate_hz)
|
||||
self.default_profile = ReferenceDistanceProfileV1.validate_profile(default_profile)
|
||||
self.enable_early_reflections = bool(enable_early_reflections)
|
||||
self.enable_late_room = bool(enable_late_room)
|
||||
self.room_config = room_config
|
||||
self.room_config.validate()
|
||||
self.output_gain = float(output_gain)
|
||||
if (self.source_count <= 0 or not math.isfinite(self.sample_rate_hz)
|
||||
or self.sample_rate_hz <= 0.0):
|
||||
raise ValueError("source_count and sample rate must be positive")
|
||||
if not math.isfinite(self.output_gain):
|
||||
raise ValueError("output gain must be finite")
|
||||
if abs(self.sample_rate_hz - 48000.0) > 1.0e-9:
|
||||
raise ValueError("the public binaural runtime requires 48 kHz")
|
||||
|
||||
if not isinstance(field, SofaHrtfField):
|
||||
raise TypeError(
|
||||
"field must be SofaHrtfField; use from_sofa() or "
|
||||
"from_compiled_cache() for file inputs")
|
||||
self.field = field
|
||||
self.hrtf_input_kind = "field"
|
||||
self.hrtf_input_path: str | None = None
|
||||
self.cache_policy: str | None = None
|
||||
if abs(self.field.sample_rate_hz - self.sample_rate_hz) > 1.0e-9:
|
||||
raise ValueError("HRTF field sample rate does not match the renderer")
|
||||
|
||||
self.early_history_slots = int(history_slots)
|
||||
if self.early_history_slots <= 0:
|
||||
raise ValueError("history_slots must be positive")
|
||||
maximum_hrtf_delay = float(np.max(self.field.delay_bounds[:, 1], initial=0.0))
|
||||
self.maximum_hrtf_delay_samples = maximum_hrtf_delay
|
||||
self.hrtf_history_slots = int(math.ceil(maximum_hrtf_delay / QMF_HOP))
|
||||
self.analysis = PublicAnalysis77(self.source_count)
|
||||
self.paths = HybridObjectPathRenderer(
|
||||
self.source_count,
|
||||
history_slots=self.early_history_slots + self.hrtf_history_slots,
|
||||
transition_slots=transition_slots)
|
||||
self.synthesis = PublicSynthesis77(2)
|
||||
actual_fdn_config = fdn_config or LateFdnConfig(sample_rate_hz=self.sample_rate_hz)
|
||||
if abs(actual_fdn_config.sample_rate_hz - self.sample_rate_hz) > 1.0e-9:
|
||||
raise ValueError("FDN sample rate does not match the renderer")
|
||||
self.fdn = SharedUnitaryFdn(actual_fdn_config)
|
||||
self.late_delay = _StereoDelay(ANALYSIS_SYNTHESIS_LATENCY_SAMPLES)
|
||||
|
||||
self.positions = np.zeros((self.source_count, 3), dtype=np.float64)
|
||||
self.positions[:, 1] = 1.0
|
||||
self.profiles = [self.default_profile] * self.source_count
|
||||
self.user_gain = np.ones(self.source_count, dtype=np.float64)
|
||||
self.special_lfe = np.zeros(self.source_count, dtype=bool)
|
||||
self.distance_state: list[DistanceState | None] = [None] * self.source_count
|
||||
self.late_current = np.zeros(self.source_count, dtype=np.float64)
|
||||
self.late_start = np.zeros(self.source_count, dtype=np.float64)
|
||||
self.late_target = np.zeros(self.source_count, dtype=np.float64)
|
||||
self.late_fade_position = np.zeros(self.source_count, dtype=np.int64)
|
||||
self.late_fade_total = np.zeros(self.source_count, dtype=np.int64)
|
||||
self.maximum_early_delay_samples = 0.0
|
||||
self.latency_to_discard = ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
self.processed_input_samples = 0
|
||||
self.output_samples = 0
|
||||
self.parameter_updates = 0
|
||||
self.finished = False
|
||||
for source in range(self.source_count):
|
||||
self.set_source(
|
||||
source, self.positions[source], profile=self.default_profile,
|
||||
fade=False)
|
||||
|
||||
@classmethod
|
||||
def from_sofa(
|
||||
cls, sofa: str | Path | CanonicalHrtf, *,
|
||||
cache_policy: str = "memory",
|
||||
cache_dir: str | Path | None = None,
|
||||
shell_radius_m: float = 1.0,
|
||||
order: int = DEFAULT_ORDER,
|
||||
projection_ridge: float = DEFAULT_PROJECTION_RIDGE,
|
||||
sh_ridge: float = DEFAULT_SH_RIDGE,
|
||||
**renderer_options) -> "SofaBinauralBackend":
|
||||
"""Compile a SOFA source once and construct the runtime renderer."""
|
||||
field = compile_sofa_hrtf(
|
||||
sofa,
|
||||
target_sample_rate_hz=float(
|
||||
renderer_options.get("sample_rate_hz", 48000.0)),
|
||||
shell_radius_m=shell_radius_m,
|
||||
order=order,
|
||||
projection_ridge=projection_ridge,
|
||||
sh_ridge=sh_ridge,
|
||||
cache_policy=cache_policy,
|
||||
cache_dir=cache_dir)
|
||||
result = cls(field, **renderer_options)
|
||||
result.hrtf_input_kind = "sofa"
|
||||
result.hrtf_input_path = (
|
||||
str(Path(sofa).expanduser().resolve())
|
||||
if not isinstance(sofa, CanonicalHrtf) else sofa.source_path)
|
||||
result.cache_policy = str(cache_policy).lower()
|
||||
return result
|
||||
|
||||
@classmethod
|
||||
def from_compiled_cache(
|
||||
cls, cache: str | Path, **renderer_options
|
||||
) -> "SofaBinauralBackend":
|
||||
"""Load an explicitly selected validated JOC compiled HRTF cache."""
|
||||
path = Path(cache).expanduser().resolve()
|
||||
field = SofaHrtfField.load(path)
|
||||
result = cls(field, **renderer_options)
|
||||
result.hrtf_input_kind = "compiled_cache"
|
||||
result.hrtf_input_path = str(path)
|
||||
result.cache_policy = None
|
||||
return result
|
||||
|
||||
def _make_path(self, label: str, direction_adm, path_distance_m: float,
|
||||
extra_delay_samples: float, amplitude: float) -> HybridPath:
|
||||
del path_distance_m
|
||||
evaluation = self.field.evaluate_adm(direction_adm)
|
||||
extra_delay = float(extra_delay_samples)
|
||||
if not math.isfinite(extra_delay) or extra_delay < 0.0:
|
||||
raise ValueError("path delay must be finite and non-negative")
|
||||
early_delay_slots = int(math.floor(extra_delay / QMF_HOP))
|
||||
if early_delay_slots >= self.early_history_slots:
|
||||
raise ValueError(
|
||||
f"path {label!r} needs {early_delay_slots} early-delay slots, "
|
||||
f"early history capacity is {self.early_history_slots}")
|
||||
total_delay = np.asarray(evaluation.delay_samples, dtype=np.float64) + extra_delay
|
||||
delay_slots = np.floor(total_delay / QMF_HOP).astype(np.int64)
|
||||
residual = total_delay - delay_slots * QMF_HOP
|
||||
propagation_phase = np.exp(
|
||||
-2j * np.pi * self.field.band_center_frequencies_hz[None, :]
|
||||
* residual[:, None]
|
||||
/ self.sample_rate_hz)
|
||||
transfer = np.asarray(
|
||||
evaluation.aligned_gains * propagation_phase * float(amplitude),
|
||||
dtype=np.complex128)
|
||||
return HybridPath(label, delay_slots, transfer)
|
||||
|
||||
def _ordinary_paths(self, state: DistanceState, gain: float) -> tuple[HybridPath, ...]:
|
||||
# Object PCM is programme-normalized. Physical distance controls room
|
||||
# geometry, while the project profile supplies the direct presentation
|
||||
# coefficient instead of applying a second free-field 1/r attenuation.
|
||||
direct_amplitude = gain * ReferenceDistanceProfileV1.direct_level_gain(state)
|
||||
room_gain = ReferenceDistanceProfileV1.room_calibration_gain(state)
|
||||
paths = [self._make_path(
|
||||
"direct", state.direction_adm, state.physical_distance_m, 0.0,
|
||||
direct_amplitude)]
|
||||
if self.enable_early_reflections:
|
||||
reflections = first_order_image_sources(
|
||||
state.direction_adm, state.physical_distance_m,
|
||||
self.sample_rate_hz, self.room_config)
|
||||
for reflection in reflections:
|
||||
amplitude = (
|
||||
gain * room_gain * reflection.reflection_gain
|
||||
* ReferenceDistanceProfileV1.inverse_distance_gain(
|
||||
self.field.measurement_radius_m,
|
||||
reflection.path_distance_m))
|
||||
paths.append(self._make_path(
|
||||
f"early:{reflection.wall}", reflection.direction_adm,
|
||||
reflection.path_distance_m, reflection.extra_delay_samples,
|
||||
amplitude))
|
||||
self.maximum_early_delay_samples = max(
|
||||
self.maximum_early_delay_samples,
|
||||
reflection.extra_delay_samples)
|
||||
return tuple(paths)
|
||||
|
||||
def _lfe_paths(self, gain: float) -> tuple[HybridPath, ...]:
|
||||
frequency = self.field.band_center_frequencies_hz
|
||||
lowpass = np.ones(HYBRID_BANDS, dtype=np.float64)
|
||||
lowpass[frequency >= 180.0] = 0.0
|
||||
transition = (frequency > 120.0) & (frequency < 180.0)
|
||||
amount = (frequency[transition] - 120.0) / 60.0
|
||||
lowpass[transition] = np.cos(0.5 * np.pi * amount) ** 2
|
||||
transfer = np.repeat(
|
||||
(gain * lowpass / math.sqrt(2.0))[None, :], 2, axis=0
|
||||
).astype(np.complex128)
|
||||
return (HybridPath(
|
||||
"public_lfe_lowpass", np.zeros(2, dtype=np.int64), transfer),)
|
||||
|
||||
def _set_late_target(self, source: int, value: float, fade: bool) -> None:
|
||||
value = float(value)
|
||||
fade_samples = (self.paths.transition_slots * QMF_HOP
|
||||
if fade and self.processed_input_samples else 0)
|
||||
if fade_samples == 0:
|
||||
self.late_current[source] = value
|
||||
self.late_start[source] = value
|
||||
self.late_target[source] = value
|
||||
self.late_fade_position[source] = 0
|
||||
self.late_fade_total[source] = 0
|
||||
else:
|
||||
self.late_start[source] = self.late_current[source]
|
||||
self.late_target[source] = value
|
||||
self.late_fade_position[source] = 0
|
||||
self.late_fade_total[source] = fade_samples
|
||||
|
||||
def set_source(self, source: int, position_adm, *, profile: str | None = None,
|
||||
gain: float = 1.0, enabled: bool = True,
|
||||
special_lfe: bool = False, fade: bool = True) -> None:
|
||||
if self.finished:
|
||||
raise RuntimeError("SOFA renderer is finished")
|
||||
source = int(source)
|
||||
if not 0 <= source < self.source_count:
|
||||
raise IndexError(source)
|
||||
gain = float(gain)
|
||||
if not math.isfinite(gain):
|
||||
raise ValueError("source gain must be finite")
|
||||
effective_gain = gain if enabled else 0.0
|
||||
name = self.default_profile if profile is None else profile
|
||||
state = ReferenceDistanceProfileV1.map_adm_position(position_adm, name)
|
||||
path_set = (self._lfe_paths(effective_gain) if special_lfe
|
||||
else self._ordinary_paths(state, effective_gain))
|
||||
self.paths.set_paths(
|
||||
source, path_set,
|
||||
fade_slots=(self.paths.transition_slots if fade else 0))
|
||||
late_send = (0.0 if special_lfe or not self.enable_late_room or not enabled
|
||||
else effective_gain
|
||||
* ReferenceDistanceProfileV1.room_calibration_gain(state)
|
||||
* ReferenceDistanceProfileV1.late_send(state))
|
||||
self._set_late_target(source, late_send, fade)
|
||||
self.positions[source] = np.asarray(position_adm, dtype=np.float64)
|
||||
self.profiles[source] = state.profile
|
||||
self.user_gain[source] = gain
|
||||
self.special_lfe[source] = bool(special_lfe)
|
||||
self.distance_state[source] = state
|
||||
self.parameter_updates += 1
|
||||
|
||||
def _late_send_envelope(self, sample_count: int) -> np.ndarray:
|
||||
envelope = np.empty((sample_count, self.source_count), dtype=np.float64)
|
||||
for source in range(self.source_count):
|
||||
total = int(self.late_fade_total[source])
|
||||
if total == 0:
|
||||
envelope[:, source] = self.late_current[source]
|
||||
continue
|
||||
start_position = int(self.late_fade_position[source])
|
||||
position = start_position + np.arange(1, sample_count + 1)
|
||||
amount = np.clip(position / float(total), 0.0, 1.0)
|
||||
envelope[:, source] = (
|
||||
self.late_start[source] * (1.0 - amount)
|
||||
+ self.late_target[source] * amount)
|
||||
new_position = start_position + sample_count
|
||||
if new_position >= total:
|
||||
self.late_current[source] = self.late_target[source]
|
||||
self.late_start[source] = self.late_target[source]
|
||||
self.late_fade_position[source] = 0
|
||||
self.late_fade_total[source] = 0
|
||||
else:
|
||||
self.late_current[source] = float(envelope[-1, source])
|
||||
self.late_fade_position[source] = new_position
|
||||
return envelope
|
||||
|
||||
def _process(self, sources) -> np.ndarray:
|
||||
values = np.asarray(sources, dtype=np.float64)
|
||||
if values.ndim != 2 or values.shape[1] != self.source_count:
|
||||
raise ValueError(f"sources must have shape [samples,{self.source_count}]")
|
||||
if len(values) % QMF_HOP:
|
||||
raise ValueError("SOFA backend input must be divisible by 64 samples")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("SOFA backend input contains non-finite values")
|
||||
hybrid = self.analysis.process(values)
|
||||
direct_and_early = self.paths.process(hybrid)
|
||||
direct_pcm = self.synthesis.process(direct_and_early)
|
||||
if self.enable_late_room:
|
||||
sends = self._late_send_envelope(len(values))
|
||||
mono = np.sum(values * sends, axis=1, dtype=np.float64)
|
||||
late_pcm = self.late_delay.process(self.fdn.process(mono))
|
||||
else:
|
||||
# Still advance any pending send fade deterministically.
|
||||
self._late_send_envelope(len(values))
|
||||
late_pcm = np.zeros_like(direct_pcm)
|
||||
mixed = np.asarray((direct_pcm + late_pcm) * self.output_gain, dtype=np.float64)
|
||||
skip = min(self.latency_to_discard, len(mixed))
|
||||
self.latency_to_discard -= skip
|
||||
self.processed_input_samples += len(values)
|
||||
output = mixed[skip:]
|
||||
self.output_samples += len(output)
|
||||
return output
|
||||
|
||||
def process(self, sources) -> np.ndarray:
|
||||
if self.finished:
|
||||
raise RuntimeError("SOFA renderer is finished")
|
||||
return self._process(sources)
|
||||
|
||||
def finish(self, *, tail_seconds: float | None = None) -> np.ndarray:
|
||||
if self.finished:
|
||||
return np.zeros((0, 2), dtype=np.float64)
|
||||
if tail_seconds is not None and (
|
||||
not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0):
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
requested = (self.fdn.tail_samples if tail_seconds is None
|
||||
else int(math.ceil(float(tail_seconds) * self.sample_rate_hz)))
|
||||
drain = max(
|
||||
requested if self.enable_late_room else 0,
|
||||
int(math.ceil(
|
||||
self.maximum_hrtf_delay_samples
|
||||
+ self.maximum_early_delay_samples)) + 2048,
|
||||
) + ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
drain = int(math.ceil(drain / QMF_HOP) * QMF_HOP)
|
||||
output = self._process(np.zeros((drain, self.source_count), dtype=np.float64))
|
||||
self.finished = True
|
||||
return output
|
||||
|
||||
def finish_output_capacity(self, tail_seconds: float | None = None) -> int:
|
||||
"""Return a conservative bound for one future :meth:`finish` output."""
|
||||
if tail_seconds is not None and (
|
||||
not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0):
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
requested = (self.fdn.tail_samples if tail_seconds is None
|
||||
else int(math.ceil(float(tail_seconds) * self.sample_rate_hz)))
|
||||
hrtf_bound = self.hrtf_history_slots * QMF_HOP
|
||||
early_bound = hrtf_bound + 2048
|
||||
if self.enable_early_reflections:
|
||||
early_bound += self.early_history_slots * QMF_HOP
|
||||
drain = max(requested if self.enable_late_room else 0, early_bound)
|
||||
drain += ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
return int(math.ceil(drain / QMF_HOP) * QMF_HOP)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.analysis.reset()
|
||||
self.paths.reset()
|
||||
self.synthesis.reset()
|
||||
self.fdn.reset()
|
||||
self.late_delay.reset()
|
||||
self.latency_to_discard = ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
self.processed_input_samples = 0
|
||||
self.output_samples = 0
|
||||
self.finished = False
|
||||
for source in range(self.source_count):
|
||||
self.set_source(
|
||||
source, self.positions[source], profile=self.profiles[source],
|
||||
gain=float(self.user_gain[source]),
|
||||
special_lfe=bool(self.special_lfe[source]), fade=False)
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"name": "SofaBinauralBackend",
|
||||
"source_count": self.source_count,
|
||||
"sample_rate_hz": self.sample_rate_hz,
|
||||
"precision": "float64/complex128",
|
||||
"signal_path": (
|
||||
"public 64-QMF -> public 77-hybrid -> SOFA order-5 real-SH "
|
||||
"direct/early -> public synthesis + shared unitary FDN"),
|
||||
"hrtf_input_kind": self.hrtf_input_kind,
|
||||
"hrtf_input_path": self.hrtf_input_path,
|
||||
"cache_policy": self.cache_policy,
|
||||
"latency_compensated_samples": ANALYSIS_SYNTHESIS_LATENCY_SAMPLES,
|
||||
"enable_early_reflections": self.enable_early_reflections,
|
||||
"enable_late_room": self.enable_late_room,
|
||||
"early_history_slots": self.early_history_slots,
|
||||
"hrtf_history_slots": self.hrtf_history_slots,
|
||||
"maximum_hrtf_delay_samples": self.maximum_hrtf_delay_samples,
|
||||
"maximum_early_delay_samples": self.maximum_early_delay_samples,
|
||||
"parameter_updates": self.parameter_updates,
|
||||
"processed_input_samples_including_flush": self.processed_input_samples,
|
||||
"output_samples_before_trim": self.output_samples,
|
||||
"distance": ReferenceDistanceProfileV1.info(),
|
||||
"filterbank": filterbank_table_info(),
|
||||
"field": self.field.info(),
|
||||
"late_room": self.fdn.info(),
|
||||
}
|
||||
@@ -1,637 +0,0 @@
|
||||
"""Strict SimpleFreeFieldHRIR to canonical HRTF import.
|
||||
|
||||
The canonical representation keeps ``Data.IR`` and ``Data.Delay`` separate.
|
||||
No importer operation silently bakes the SOFA delay into the stored FIRs. A
|
||||
caller must explicitly request :meth:`CanonicalHrtf.materialized_measurement`
|
||||
when a time-domain FIR with ``Data.Delay`` applied exactly once is required.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from contextlib import contextmanager
|
||||
from dataclasses import dataclass, replace
|
||||
from fractions import Fraction
|
||||
import hashlib
|
||||
import math
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import h5py
|
||||
import numpy as np
|
||||
from scipy import signal
|
||||
from scipy.fft import next_fast_len
|
||||
|
||||
|
||||
_SUPPORTED_VERSIONS = {"0.4", "1.0", "1.1"}
|
||||
_FREE_FIELD_ROOM_TYPES = {"free field", "free-field", "anechoic", "hemi-anechoic"}
|
||||
_LENGTH_UNITS = {
|
||||
"m": 1.0,
|
||||
"metre": 1.0,
|
||||
"metres": 1.0,
|
||||
"meter": 1.0,
|
||||
"meters": 1.0,
|
||||
"cm": 1.0e-2,
|
||||
"centimetre": 1.0e-2,
|
||||
"centimetres": 1.0e-2,
|
||||
"centimeter": 1.0e-2,
|
||||
"centimeters": 1.0e-2,
|
||||
"mm": 1.0e-3,
|
||||
"millimetre": 1.0e-3,
|
||||
"millimetres": 1.0e-3,
|
||||
"millimeter": 1.0e-3,
|
||||
"millimeters": 1.0e-3,
|
||||
}
|
||||
_ANGLE_UNITS = {
|
||||
"degree": np.deg2rad,
|
||||
"degrees": np.deg2rad,
|
||||
"radian": lambda value: np.asarray(value, dtype=np.float64),
|
||||
"radians": lambda value: np.asarray(value, dtype=np.float64),
|
||||
}
|
||||
|
||||
|
||||
class SofaImportError(ValueError):
|
||||
"""The file is outside the deliberately narrow public SOFA contract."""
|
||||
|
||||
|
||||
def _text(value: Any) -> str:
|
||||
if isinstance(value, np.ndarray) and value.shape == ():
|
||||
value = value.item()
|
||||
if isinstance(value, (bytes, np.bytes_)):
|
||||
return value.decode("utf-8", "strict")
|
||||
return str(value)
|
||||
|
||||
|
||||
def _sha256_stream(stream) -> str:
|
||||
digest = hashlib.sha256()
|
||||
stream.seek(0)
|
||||
for block in iter(lambda: stream.read(4 << 20), b""):
|
||||
digest.update(block)
|
||||
stream.seek(0)
|
||||
return digest.hexdigest().upper()
|
||||
|
||||
|
||||
@contextmanager
|
||||
def _stable_hdf5_source(path: Path):
|
||||
"""Read arrays and content identity from one stable open-file snapshot."""
|
||||
with path.open("rb") as stream:
|
||||
before = _sha256_stream(stream)
|
||||
with h5py.File(stream, "r") as file:
|
||||
yield file, before
|
||||
after = _sha256_stream(stream)
|
||||
if after != before:
|
||||
raise SofaImportError("SOFA file changed while it was being imported")
|
||||
|
||||
|
||||
def _tokens(units: str) -> list[str]:
|
||||
return [token.strip().lower() for token in units.split(",") if token.strip()]
|
||||
|
||||
|
||||
def _coordinate_attributes(dataset: h5py.Dataset, *, inherit=None) -> tuple[str, str]:
|
||||
source = dataset.attrs
|
||||
if "Type" not in source or "Units" not in source:
|
||||
if inherit is None or "Type" not in inherit.attrs or "Units" not in inherit.attrs:
|
||||
raise SofaImportError(f"{dataset.name} must declare Type and Units")
|
||||
source = inherit.attrs
|
||||
return _text(source["Type"]).strip().lower(), _text(source["Units"]).strip()
|
||||
|
||||
|
||||
def coordinates_to_cartesian_m(values, coordinate_type: str, units: str,
|
||||
*, variable: str) -> np.ndarray:
|
||||
"""Convert a SOFA coordinate array to Cartesian metres without reshaping it."""
|
||||
data = np.asarray(values, dtype=np.float64)
|
||||
if data.shape[-1] != 3 or not np.isfinite(data).all():
|
||||
raise SofaImportError(f"{variable} must contain finite C=3 coordinates")
|
||||
kind = coordinate_type.strip().lower()
|
||||
unit_tokens = _tokens(units)
|
||||
if kind == "cartesian":
|
||||
if len(unit_tokens) == 1:
|
||||
factors = [_LENGTH_UNITS.get(unit_tokens[0])] * 3
|
||||
elif len(unit_tokens) == 3:
|
||||
factors = [_LENGTH_UNITS.get(token) for token in unit_tokens]
|
||||
else:
|
||||
factors = []
|
||||
if len(factors) != 3 or any(value is None for value in factors):
|
||||
raise SofaImportError(f"unsupported Cartesian units for {variable}: {units!r}")
|
||||
return data * np.asarray(factors, dtype=np.float64)
|
||||
if kind != "spherical" or len(unit_tokens) != 3:
|
||||
raise SofaImportError(
|
||||
f"unsupported coordinates for {variable}: Type={coordinate_type!r}, Units={units!r}")
|
||||
if unit_tokens[0] not in _ANGLE_UNITS or unit_tokens[1] not in _ANGLE_UNITS:
|
||||
raise SofaImportError(f"unsupported spherical angle units for {variable}: {units!r}")
|
||||
radius_factor = _LENGTH_UNITS.get(unit_tokens[2])
|
||||
if radius_factor is None:
|
||||
raise SofaImportError(f"unsupported spherical radius unit for {variable}: {units!r}")
|
||||
azimuth = _ANGLE_UNITS[unit_tokens[0]](data[..., 0])
|
||||
elevation = _ANGLE_UNITS[unit_tokens[1]](data[..., 1])
|
||||
radius = data[..., 2] * radius_factor
|
||||
if np.any(radius < 0.0):
|
||||
raise SofaImportError(f"{variable} contains a negative spherical radius")
|
||||
horizontal = np.cos(elevation)
|
||||
return np.stack(
|
||||
(radius * horizontal * np.cos(azimuth),
|
||||
radius * horizontal * np.sin(azimuth),
|
||||
radius * np.sin(elevation)),
|
||||
axis=-1,
|
||||
).astype(np.float64, copy=False)
|
||||
|
||||
|
||||
def _rows(file: h5py.File, name: str, measurements: int, *, inherit=None) -> np.ndarray:
|
||||
if name not in file:
|
||||
raise SofaImportError(f"missing required SOFA variable {name}")
|
||||
dataset = file[name]
|
||||
kind, units = _coordinate_attributes(dataset, inherit=inherit)
|
||||
result = coordinates_to_cartesian_m(dataset[...], kind, units, variable=name)
|
||||
if result.shape not in ((1, 3), (measurements, 3)):
|
||||
raise SofaImportError(
|
||||
f"{name} must have shape [I,C] or [M,C], got {result.shape}")
|
||||
return np.broadcast_to(result, (measurements, 3)).astype(np.float64, copy=True)
|
||||
|
||||
|
||||
def _receiver_rows(file: h5py.File, measurements: int) -> np.ndarray:
|
||||
if "ReceiverPosition" not in file:
|
||||
raise SofaImportError("missing required SOFA variable ReceiverPosition")
|
||||
dataset = file["ReceiverPosition"]
|
||||
raw = np.asarray(dataset[...], dtype=np.float64)
|
||||
if raw.shape not in ((2, 3, 1), (2, 3, measurements)):
|
||||
raise SofaImportError(
|
||||
"ReceiverPosition must have shape [R=2,C=3,I=1 or M]")
|
||||
values = np.moveaxis(raw, 1, -1) # [R,I/M,C]
|
||||
kind, units = _coordinate_attributes(dataset)
|
||||
cartesian = coordinates_to_cartesian_m(
|
||||
values, kind, units, variable="ReceiverPosition")
|
||||
cartesian = np.moveaxis(cartesian, 0, 1) # [I/M,R,C]
|
||||
return np.broadcast_to(cartesian, (measurements, 2, 3)).astype(
|
||||
np.float64, copy=True)
|
||||
|
||||
|
||||
def _emitter_is_origin(file: h5py.File, measurements: int) -> None:
|
||||
if "EmitterPosition" not in file:
|
||||
raise SofaImportError("missing required SOFA variable EmitterPosition")
|
||||
dataset = file["EmitterPosition"]
|
||||
raw = np.asarray(dataset[...], dtype=np.float64)
|
||||
if raw.shape not in ((1, 3, 1), (1, 3, measurements)):
|
||||
raise SofaImportError("SimpleFreeFieldHRIR v1 requires E=1 EmitterPosition[E,C,I/M]")
|
||||
values = np.moveaxis(raw, 1, -1)
|
||||
kind, units = _coordinate_attributes(dataset)
|
||||
cartesian = coordinates_to_cartesian_m(
|
||||
values, kind, units, variable="EmitterPosition")
|
||||
if np.max(np.abs(cartesian), initial=0.0) > 1.0e-9:
|
||||
raise SofaImportError("non-zero EmitterPosition needs a separate source-pose adapter")
|
||||
|
||||
|
||||
def _sampling_rate(file: h5py.File) -> float:
|
||||
if "Data.SamplingRate" not in file:
|
||||
raise SofaImportError("missing Data.SamplingRate")
|
||||
dataset = file["Data.SamplingRate"]
|
||||
values = np.asarray(dataset[...], dtype=np.float64).reshape(-1)
|
||||
if values.size != 1 or not math.isfinite(float(values[0])) or values[0] <= 0.0:
|
||||
raise SofaImportError("Data.SamplingRate must contain one positive finite value")
|
||||
units = _text(dataset.attrs.get("Units", "")).strip().lower()
|
||||
if units not in {"hertz", "hz"}:
|
||||
raise SofaImportError(f"Data.SamplingRate Units must be hertz, got {units!r}")
|
||||
return float(values[0])
|
||||
|
||||
|
||||
def _processing_label(file: h5py.File) -> str:
|
||||
parts = []
|
||||
for key in ("DatabaseName", "Title", "ListenerShortName", "Comment"):
|
||||
value = _text(file.attrs.get(key, "")).strip()
|
||||
if value and value not in parts:
|
||||
parts.append(value)
|
||||
return " | ".join(parts)
|
||||
|
||||
|
||||
def _read_delay(file: h5py.File, measurements: int) -> np.ndarray:
|
||||
if "Data.Delay" not in file:
|
||||
raise SofaImportError("missing Data.Delay")
|
||||
delay = np.asarray(file["Data.Delay"][...], dtype=np.float64)
|
||||
if delay.shape not in ((1, 2), (measurements, 2)) or not np.isfinite(delay).all():
|
||||
raise SofaImportError("Data.Delay must have finite shape [I=1,R=2] or [M,R=2]")
|
||||
delay = np.broadcast_to(delay, (measurements, 2)).astype(np.float64, copy=True)
|
||||
if np.min(delay, initial=0.0) < -1.0e-9:
|
||||
raise SofaImportError("negative Data.Delay is outside the supported causal contract")
|
||||
delay[delay < 0.0] = 0.0
|
||||
return delay
|
||||
|
||||
|
||||
def _readonly(array, dtype) -> np.ndarray:
|
||||
result = np.asarray(array, dtype=dtype)
|
||||
result.setflags(write=False)
|
||||
return result
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class CanonicalHrtf:
|
||||
source_path: str
|
||||
source_sha256: str
|
||||
convention: str
|
||||
convention_version: str
|
||||
sofa_version: str
|
||||
source_sample_rate_hz: float
|
||||
sample_rate_hz: float
|
||||
source_position_cartesian_m: np.ndarray # listener-local [M,3]
|
||||
listener_view: np.ndarray # world, normalized [M,3]
|
||||
listener_up: np.ndarray # world, orthonormal [M,3]
|
||||
receiver_position_cartesian_m: np.ndarray # listener-local, L/R [M,2,3]
|
||||
left_receiver_index: int
|
||||
right_receiver_index: int
|
||||
hrir: np.ndarray # canonical L/R [M,2,N]
|
||||
delay_samples: np.ndarray # canonical L/R [M,2], not applied
|
||||
measurement_radius_m: np.ndarray # [M]
|
||||
processing_label: str
|
||||
resampling_label: str = "none"
|
||||
|
||||
def __post_init__(self):
|
||||
object.__setattr__(self, "source_position_cartesian_m", _readonly(
|
||||
self.source_position_cartesian_m, np.float64))
|
||||
object.__setattr__(self, "listener_view", _readonly(self.listener_view, np.float64))
|
||||
object.__setattr__(self, "listener_up", _readonly(self.listener_up, np.float64))
|
||||
object.__setattr__(self, "receiver_position_cartesian_m", _readonly(
|
||||
self.receiver_position_cartesian_m, np.float64))
|
||||
object.__setattr__(self, "hrir", _readonly(self.hrir, np.float64))
|
||||
object.__setattr__(self, "delay_samples", _readonly(self.delay_samples, np.float64))
|
||||
object.__setattr__(self, "measurement_radius_m", _readonly(
|
||||
self.measurement_radius_m, np.float64))
|
||||
|
||||
@property
|
||||
def measurements(self) -> int:
|
||||
return int(self.hrir.shape[0])
|
||||
|
||||
@property
|
||||
def taps(self) -> int:
|
||||
return int(self.hrir.shape[2])
|
||||
|
||||
@property
|
||||
def unit_directions(self) -> np.ndarray:
|
||||
return self.source_position_cartesian_m / self.measurement_radius_m[:, None]
|
||||
|
||||
@property
|
||||
def shells_m(self) -> np.ndarray:
|
||||
return np.unique(np.round(self.measurement_radius_m, 9))
|
||||
|
||||
def shell_indices(self, radius_m: float) -> np.ndarray:
|
||||
shell = float(self.shells_m[np.argmin(np.abs(self.shells_m - float(radius_m)))])
|
||||
return np.flatnonzero(np.isclose(
|
||||
self.measurement_radius_m, shell, atol=5.0e-7, rtol=0.0))
|
||||
|
||||
def nearest_index(self, direction_sofa, radius_m: float = 1.0) -> tuple[int, float]:
|
||||
direction = np.asarray(direction_sofa, dtype=np.float64)
|
||||
if direction.shape != (3,) or not np.isfinite(direction).all():
|
||||
raise ValueError("direction must contain three finite SOFA Cartesian values")
|
||||
norm = float(np.linalg.norm(direction))
|
||||
if norm <= 1.0e-15:
|
||||
raise ValueError("direction must be non-zero")
|
||||
direction = direction / norm
|
||||
indices = self.shell_indices(radius_m)
|
||||
dots = self.unit_directions[indices] @ direction
|
||||
local = int(np.argmax(dots))
|
||||
error = math.degrees(math.acos(float(np.clip(dots[local], -1.0, 1.0))))
|
||||
return int(indices[local]), float(error)
|
||||
|
||||
def resampled(self, target_sample_rate_hz: float) -> "CanonicalHrtf":
|
||||
target = float(target_sample_rate_hz)
|
||||
if not math.isfinite(target) or target <= 0.0:
|
||||
raise ValueError("target sample rate must be positive and finite")
|
||||
if abs(target - self.sample_rate_hz) <= 1.0e-9:
|
||||
return self
|
||||
ratio = target / self.sample_rate_hz
|
||||
fraction = Fraction(ratio).limit_denominator(100000)
|
||||
if abs(float(fraction) - ratio) > 1.0e-10:
|
||||
raise ValueError("sample-rate ratio cannot be represented safely")
|
||||
converted = signal.resample_poly(
|
||||
np.asarray(self.hrir, dtype=np.float64), fraction.numerator,
|
||||
fraction.denominator, axis=-1, window=("kaiser", 8.6), padtype="constant")
|
||||
converted = np.asarray(converted, dtype=np.float64)
|
||||
return replace(
|
||||
self,
|
||||
sample_rate_hz=target,
|
||||
hrir=converted,
|
||||
delay_samples=np.asarray(self.delay_samples * ratio, dtype=np.float64),
|
||||
resampling_label=(
|
||||
f"scipy.signal.resample_poly {self.sample_rate_hz:g}->{target:g} Hz "
|
||||
f"({fraction.numerator}/{fraction.denominator}, Kaiser beta=8.6)"),
|
||||
)
|
||||
|
||||
def materialized_measurement(self, index: int, *, fractional_half_length: int = 48
|
||||
) -> np.ndarray:
|
||||
"""Return [L/R,taps] with SOFA Data.Delay applied exactly once."""
|
||||
index = int(index)
|
||||
if not 0 <= index < self.measurements:
|
||||
raise IndexError(index)
|
||||
ears = []
|
||||
for ear in range(2):
|
||||
ears.append(apply_fractional_delay(
|
||||
self.hrir[index, ear], float(self.delay_samples[index, ear]),
|
||||
half_length=fractional_half_length))
|
||||
length = max(map(len, ears))
|
||||
result = np.zeros((2, length), dtype=np.float64)
|
||||
for ear, value in enumerate(ears):
|
||||
result[ear, :len(value)] = value
|
||||
return result
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"source_path": self.source_path,
|
||||
"source_sha256": self.source_sha256,
|
||||
"convention": self.convention,
|
||||
"convention_version": self.convention_version,
|
||||
"sofa_version": self.sofa_version,
|
||||
"source_sample_rate_hz": self.source_sample_rate_hz,
|
||||
"sample_rate_hz": self.sample_rate_hz,
|
||||
"measurements": self.measurements,
|
||||
"taps": self.taps,
|
||||
"shells_m": [float(value) for value in self.shells_m],
|
||||
"source_receiver_order": [self.left_receiver_index, self.right_receiver_index],
|
||||
"canonical_ear_order": ["left", "right"],
|
||||
"data_delay_samples_min": float(np.min(self.delay_samples)),
|
||||
"data_delay_samples_max": float(np.max(self.delay_samples)),
|
||||
"data_delay_applied": False,
|
||||
"processing_label": self.processing_label,
|
||||
"resampling": self.resampling_label,
|
||||
"precision": "float64",
|
||||
}
|
||||
|
||||
|
||||
def load_simple_free_field_hrir(path, *, target_sample_rate_hz: float | None = None
|
||||
) -> CanonicalHrtf:
|
||||
"""Strictly import the supported SimpleFreeFieldHRIR subset."""
|
||||
source = Path(path).expanduser().resolve()
|
||||
if not source.is_file():
|
||||
raise FileNotFoundError(source)
|
||||
with _stable_hdf5_source(source) as (file, source_sha256):
|
||||
if _text(file.attrs.get("Conventions", "")) != "SOFA":
|
||||
raise SofaImportError("Conventions must be SOFA")
|
||||
convention = _text(file.attrs.get("SOFAConventions", ""))
|
||||
if convention != "SimpleFreeFieldHRIR":
|
||||
raise SofaImportError(
|
||||
f"unsupported SOFAConventions={convention!r}; convert explicitly first")
|
||||
convention_version = _text(file.attrs.get("SOFAConventionsVersion", ""))
|
||||
if convention_version not in _SUPPORTED_VERSIONS:
|
||||
raise SofaImportError(
|
||||
f"unsupported SimpleFreeFieldHRIR version {convention_version!r}; "
|
||||
f"supported={sorted(_SUPPORTED_VERSIONS)}")
|
||||
if _text(file.attrs.get("DataType", "")) != "FIR":
|
||||
raise SofaImportError("DataType must be FIR")
|
||||
room_type = _text(file.attrs.get("RoomType", "")).strip().lower()
|
||||
if room_type not in _FREE_FIELD_ROOM_TYPES:
|
||||
raise SofaImportError(f"RoomType must explicitly be free-field, got {room_type!r}")
|
||||
if "Data.IR" not in file:
|
||||
raise SofaImportError("missing Data.IR")
|
||||
hrir_source = np.asarray(file["Data.IR"][...], dtype=np.float64)
|
||||
if hrir_source.ndim != 3 or hrir_source.shape[1] != 2 or min(hrir_source.shape) <= 0:
|
||||
raise SofaImportError("Data.IR must have shape [M,R=2,N]")
|
||||
if not np.isfinite(hrir_source).all():
|
||||
raise SofaImportError("Data.IR contains non-finite values")
|
||||
measurements = int(hrir_source.shape[0])
|
||||
sample_rate = _sampling_rate(file)
|
||||
delay_source = _read_delay(file, measurements)
|
||||
_emitter_is_origin(file, measurements)
|
||||
|
||||
listener_position = _rows(file, "ListenerPosition", measurements)
|
||||
listener_view = _rows(file, "ListenerView", measurements)
|
||||
listener_up_raw = _rows(
|
||||
file, "ListenerUp", measurements,
|
||||
inherit=file["ListenerView"] if "ListenerView" in file else None)
|
||||
forward_norm = np.linalg.norm(listener_view, axis=1)
|
||||
if np.any(forward_norm <= 1.0e-12):
|
||||
raise SofaImportError("ListenerView must be non-zero")
|
||||
forward = listener_view / forward_norm[:, None]
|
||||
left = np.cross(listener_up_raw, forward)
|
||||
left_norm = np.linalg.norm(left, axis=1)
|
||||
if np.any(left_norm <= 1.0e-12):
|
||||
raise SofaImportError("ListenerUp must not be parallel to ListenerView")
|
||||
left /= left_norm[:, None]
|
||||
up = np.cross(forward, left)
|
||||
|
||||
if "SourcePosition" not in file:
|
||||
raise SofaImportError("missing required SOFA variable SourcePosition")
|
||||
source_dataset = file["SourcePosition"]
|
||||
source_type, source_units = _coordinate_attributes(source_dataset)
|
||||
source_world = coordinates_to_cartesian_m(
|
||||
source_dataset[...], source_type, source_units, variable="SourcePosition")
|
||||
if source_world.shape != (measurements, 3):
|
||||
raise SofaImportError("SourcePosition must have shape [M,C=3]")
|
||||
relative = source_world - listener_position
|
||||
source_local = np.stack(
|
||||
(np.sum(relative * forward, axis=1),
|
||||
np.sum(relative * left, axis=1),
|
||||
np.sum(relative * up, axis=1)), axis=1)
|
||||
radii = np.linalg.norm(source_local, axis=1)
|
||||
if np.any(radii <= 1.0e-8) or not np.isfinite(radii).all():
|
||||
raise SofaImportError("every source measurement must have a positive radius")
|
||||
|
||||
receiver = _receiver_rows(file, measurements)
|
||||
lateral_difference = receiver[:, 0, 1] - receiver[:, 1, 1]
|
||||
if np.all(lateral_difference > 1.0e-5):
|
||||
left_index, right_index = 0, 1
|
||||
elif np.all(lateral_difference < -1.0e-5):
|
||||
left_index, right_index = 1, 0
|
||||
else:
|
||||
raise SofaImportError(
|
||||
"ReceiverPosition does not identify one consistently-left and one "
|
||||
"consistently-right receiver")
|
||||
ear_order = [left_index, right_index]
|
||||
hrir = hrir_source[:, ear_order, :]
|
||||
delay = delay_source[:, ear_order]
|
||||
receiver = receiver[:, ear_order, :]
|
||||
processing_label = _processing_label(file)
|
||||
sofa_version = _text(file.attrs.get("Version", ""))
|
||||
|
||||
canonical = CanonicalHrtf(
|
||||
source_path=str(source),
|
||||
source_sha256=source_sha256,
|
||||
convention=convention,
|
||||
convention_version=convention_version,
|
||||
sofa_version=sofa_version,
|
||||
source_sample_rate_hz=sample_rate,
|
||||
sample_rate_hz=sample_rate,
|
||||
source_position_cartesian_m=source_local,
|
||||
listener_view=forward,
|
||||
listener_up=up,
|
||||
receiver_position_cartesian_m=receiver,
|
||||
left_receiver_index=left_index,
|
||||
right_receiver_index=right_index,
|
||||
hrir=hrir,
|
||||
delay_samples=delay,
|
||||
measurement_radius_m=radii,
|
||||
processing_label=processing_label,
|
||||
)
|
||||
return (canonical if target_sample_rate_hz is None
|
||||
else canonical.resampled(target_sample_rate_hz))
|
||||
|
||||
|
||||
def apply_fractional_delay(values, delay_samples: float, *, half_length: int = 48
|
||||
) -> np.ndarray:
|
||||
"""Apply one causal non-negative delay to a real FIR using windowed sinc."""
|
||||
source = np.asarray(values, dtype=np.float64)
|
||||
delay = float(delay_samples)
|
||||
if source.ndim != 1 or not np.isfinite(source).all():
|
||||
raise ValueError("fractional delay input must be a finite real vector")
|
||||
if not math.isfinite(delay) or delay < -1.0e-12:
|
||||
raise ValueError("fractional delay must be finite and non-negative")
|
||||
if delay < 1.0e-12:
|
||||
return source.copy()
|
||||
integer = int(math.floor(delay))
|
||||
fraction = delay - integer
|
||||
if fraction < 1.0e-12:
|
||||
return np.pad(source, (integer, 0)).astype(np.float64, copy=False)
|
||||
half = int(half_length)
|
||||
if half < 8:
|
||||
raise ValueError("fractional delay half_length must be at least 8")
|
||||
index = np.arange(-half, half + 1, dtype=np.float64)
|
||||
kernel = np.sinc(index - fraction) * np.kaiser(2 * half + 1, 8.6)
|
||||
kernel /= np.sum(kernel, dtype=np.float64)
|
||||
full = signal.fftconvolve(source, kernel, mode="full")
|
||||
causal = np.asarray(full[half:], dtype=np.float64)
|
||||
return np.pad(causal, (integer, 0)).astype(np.float64, copy=False)
|
||||
|
||||
|
||||
def shift_signal_fft(values, shift_samples: float) -> np.ndarray:
|
||||
"""Band-limited linear shift; positive is delay and negative is advance."""
|
||||
source = np.asarray(values, dtype=np.float64)
|
||||
shift = float(shift_samples)
|
||||
if source.ndim != 1 or not np.isfinite(source).all() or not math.isfinite(shift):
|
||||
raise ValueError("shift input and amount must be finite")
|
||||
if abs(shift) < 1.0e-12:
|
||||
return source.copy()
|
||||
guard = max(128, int(math.ceil(abs(shift))) + 64)
|
||||
needed = len(source) + 2 * guard
|
||||
fft_size = next_fast_len(needed)
|
||||
padded = np.zeros(fft_size, dtype=np.float64)
|
||||
padded[guard:guard + len(source)] = source
|
||||
bins = np.arange(fft_size // 2 + 1, dtype=np.float64)
|
||||
spectrum = np.fft.rfft(padded)
|
||||
spectrum *= np.exp(-2j * np.pi * bins * shift / fft_size)
|
||||
shifted = np.fft.irfft(spectrum, fft_size)
|
||||
return np.asarray(shifted[guard:guard + len(source)], dtype=np.float64)
|
||||
|
||||
|
||||
def estimate_interaural_delay_samples(left, right, sample_rate_hz: float,
|
||||
*, low_hz: float = 200.0,
|
||||
high_hz: float = 1500.0) -> float:
|
||||
"""Estimate L-minus-R delay by low-frequency circular phase coherence.
|
||||
|
||||
A coarse-to-fine delay search avoids the phase-unwrapping branch failures
|
||||
that ordinary straight-line regression can exhibit for strongly filtered
|
||||
Far responses.
|
||||
"""
|
||||
left = np.asarray(left, dtype=np.float64)
|
||||
right = np.asarray(right, dtype=np.float64)
|
||||
if left.shape != right.shape or left.ndim != 1:
|
||||
raise ValueError("ITD inputs must be equal-length vectors")
|
||||
fft_size = next_fast_len(max(4096, 4 * len(left)))
|
||||
left_spectrum = np.fft.rfft(left, fft_size)
|
||||
right_spectrum = np.fft.rfft(right, fft_size)
|
||||
frequency = np.fft.rfftfreq(fft_size, 1.0 / float(sample_rate_hz))
|
||||
selected = (frequency >= low_hz) & (frequency <= high_hz)
|
||||
if np.count_nonzero(selected) < 8:
|
||||
return 0.0
|
||||
cross = left_spectrum[selected] * np.conj(right_spectrum[selected])
|
||||
magnitude = np.abs(cross)
|
||||
maximum = float(np.max(magnitude, initial=0.0))
|
||||
if maximum <= 1.0e-20:
|
||||
return 0.0
|
||||
weighted_unit = cross / np.maximum(magnitude, 1.0e-30)
|
||||
weight = np.sqrt(magnitude / maximum)
|
||||
weighted_unit *= weight
|
||||
omega = 2.0 * np.pi * frequency[selected] / float(sample_rate_hz)
|
||||
limit = 0.0012 * float(sample_rate_hz)
|
||||
|
||||
def best(candidates: np.ndarray) -> float:
|
||||
steering = np.exp(1j * omega[:, None] * candidates[None, :])
|
||||
score = np.abs(weighted_unit @ steering)
|
||||
return float(candidates[int(np.argmax(score))])
|
||||
|
||||
coarse = np.arange(-limit, limit + 0.25, 0.5, dtype=np.float64)
|
||||
estimate = best(coarse)
|
||||
fine = np.arange(estimate - 0.6, estimate + 0.6001, 0.02, dtype=np.float64)
|
||||
return float(np.clip(best(fine), -limit, limit))
|
||||
|
||||
def _subsample_peak(values: np.ndarray) -> float:
|
||||
magnitude = np.abs(np.asarray(values, dtype=np.float64))
|
||||
index = int(np.argmax(magnitude))
|
||||
if index == 0 or index + 1 >= len(magnitude):
|
||||
return float(index)
|
||||
y0, y1, y2 = (float(magnitude[index - 1]), float(magnitude[index]),
|
||||
float(magnitude[index + 1]))
|
||||
denominator = y0 - 2.0 * y1 + y2
|
||||
correction = 0.0 if abs(denominator) < 1.0e-30 else 0.5 * (y0 - y2) / denominator
|
||||
return float(index + np.clip(correction, -0.5, 0.5))
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class TimeAlignedHrtf:
|
||||
canonical: CanonicalHrtf
|
||||
aligned_hrir: np.ndarray
|
||||
runtime_delay_samples: np.ndarray
|
||||
embedded_delay_removed_samples: np.ndarray
|
||||
delay_source: str
|
||||
|
||||
def __post_init__(self):
|
||||
object.__setattr__(self, "aligned_hrir", _readonly(self.aligned_hrir, np.float64))
|
||||
object.__setattr__(self, "runtime_delay_samples", _readonly(
|
||||
self.runtime_delay_samples, np.float64))
|
||||
object.__setattr__(self, "embedded_delay_removed_samples", _readonly(
|
||||
self.embedded_delay_removed_samples, np.float64))
|
||||
|
||||
|
||||
def time_align_hrtf(canonical: CanonicalHrtf) -> TimeAlignedHrtf:
|
||||
"""Separate one delay representation before directional interpolation.
|
||||
|
||||
Trusted non-zero ``Data.Delay`` is external to ``Data.IR`` and is therefore
|
||||
retained without de-rotating the FIR. When ``Data.Delay`` is identically
|
||||
zero, ordinary measured HRIRs with a positive onset use their per-ear main
|
||||
peaks. A zero-origin effective FIR is already expressed at one common
|
||||
time origin; its interaural phase is therefore retained in ``Data.IR``.
|
||||
|
||||
These representations are mutually exclusive. Runtime rendering must
|
||||
restore exactly the delay separated here and must not add any second ear
|
||||
delay or phase-group delay.
|
||||
"""
|
||||
hrir = np.asarray(canonical.hrir, dtype=np.float64)
|
||||
if np.max(np.abs(canonical.delay_samples), initial=0.0) > 1.0e-12:
|
||||
return TimeAlignedHrtf(
|
||||
canonical=canonical,
|
||||
aligned_hrir=hrir.copy(),
|
||||
runtime_delay_samples=np.asarray(canonical.delay_samples, dtype=np.float64),
|
||||
embedded_delay_removed_samples=np.zeros_like(canonical.delay_samples),
|
||||
delay_source="Data.Delay (external; applied once at render time)",
|
||||
)
|
||||
|
||||
measurements = canonical.measurements
|
||||
runtime = np.zeros((measurements, 2), dtype=np.float64)
|
||||
removed = np.zeros_like(runtime)
|
||||
used_peak = 0
|
||||
retained_embedded_phase = 0
|
||||
for measurement in range(measurements):
|
||||
peaks = np.asarray([
|
||||
_subsample_peak(hrir[measurement, 0]),
|
||||
_subsample_peak(hrir[measurement, 1]),
|
||||
], dtype=np.float64)
|
||||
if float(np.max(peaks)) > 2.0:
|
||||
delays = peaks
|
||||
used_peak += 1
|
||||
else:
|
||||
# An effective response can have both ear FIRs beginning at sample
|
||||
# zero while still carrying the correct ITD in complex phase. Do
|
||||
# not invent an external delay which SOFA did not author.
|
||||
delays = np.zeros(2, dtype=np.float64)
|
||||
retained_embedded_phase += 1
|
||||
runtime[measurement] = delays
|
||||
removed[measurement] = delays
|
||||
|
||||
aligned = np.empty_like(hrir)
|
||||
for measurement in range(measurements):
|
||||
for ear in range(2):
|
||||
aligned[measurement, ear] = shift_signal_fft(
|
||||
hrir[measurement, ear], -float(removed[measurement, ear]))
|
||||
source = (
|
||||
f"embedded Data.IR arrival separation: peak={used_peak}, "
|
||||
f"zero-origin embedded phase retained={retained_embedded_phase}; "
|
||||
"positive onset restored once at render time")
|
||||
return TimeAlignedHrtf(
|
||||
canonical=canonical,
|
||||
aligned_hrir=aligned,
|
||||
runtime_delay_samples=runtime,
|
||||
embedded_delay_removed_samples=removed,
|
||||
delay_source=source,
|
||||
)
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,398 +0,0 @@
|
||||
"""Native float64 SOFA binaural DSP (ctypes bridge to eac3joc_core).
|
||||
|
||||
The C++ side mirrors the Python :class:`sofa_binaural_backend.SofaBinauralBackend`
|
||||
mathematics: 64-QMF/77-hybrid analysis and synthesis, fifth-order ACN/N3D real
|
||||
spherical-harmonic field evaluation, whole-QMF-slot per-object delay histories,
|
||||
six image-source early reflections, the shared unitary FDN late room, the LFE
|
||||
low-pass and the 961-sample latency policy. The compiled HRTF field, the
|
||||
filterbank tables and the room constants are uploaded once; per 512-sample
|
||||
block the adapter updates every source and streams PCM through the DLL.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import ctypes
|
||||
import math
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
from native_renderer import ABI_VERSION, find_native_library
|
||||
from public_filterbank import DEFAULT_FILTERBANK_DATA, load_filterbank_tables
|
||||
from public_room import LateFdnConfig, SharedUnitaryFdn, ShoeboxRoomConfig
|
||||
from reference_distance import ReferenceDistanceProfileV1
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
from sofa_hrtf_field import (
|
||||
DEFAULT_HRTF_CACHE_DIR,
|
||||
SofaHrtfField,
|
||||
compile_sofa_hrtf,
|
||||
)
|
||||
|
||||
BLOCK_SAMPLES = 512
|
||||
INPUT_CHANNELS = 16
|
||||
OUTPUT_CHANNELS = 2
|
||||
QMF_HOP = 64
|
||||
LATENCY_SAMPLES = 961
|
||||
|
||||
_PROFILE_INDEX = {"near": 0, "mid": 1, "far": 2}
|
||||
|
||||
|
||||
def _room_numbers(fdn_config: LateFdnConfig) -> dict:
|
||||
"""Derive the FDN delays/feedback with the same arithmetic as the Python room."""
|
||||
fdn = SharedUnitaryFdn(fdn_config)
|
||||
return {
|
||||
"fdn_delays": np.asarray(fdn.delays, dtype=np.uint32),
|
||||
"fdn_feedback": np.asarray(fdn.feedback_gain, dtype=np.float64),
|
||||
"damping": fdn.damping,
|
||||
"output_gain": fdn.output_gain,
|
||||
"allpass_delays": np.asarray(
|
||||
[diffuser.delay_samples for diffuser in fdn.diffusers], dtype=np.uint32),
|
||||
"allpass_gains": np.asarray(fdn_config.allpass_gain, dtype=np.float64),
|
||||
"tail_samples": fdn.tail_samples,
|
||||
}
|
||||
|
||||
|
||||
class NativeSofaBinauralDsp:
|
||||
"""Duck-type compatible with SofaBinauralBackend for the JOC adapter."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
field: SofaHrtfField,
|
||||
*,
|
||||
source_count: int = INPUT_CHANNELS,
|
||||
default_profile: str = "mid",
|
||||
enable_early_reflections: bool = True,
|
||||
enable_late_room: bool = True,
|
||||
room_config: ShoeboxRoomConfig = ShoeboxRoomConfig(),
|
||||
fdn_config: LateFdnConfig | None = None,
|
||||
library_path: str | Path | None = None):
|
||||
if not isinstance(field, SofaHrtfField):
|
||||
raise TypeError("field must be SofaHrtfField")
|
||||
self.source_count = int(source_count)
|
||||
self.default_profile = ReferenceDistanceProfileV1.validate_profile(default_profile)
|
||||
self.enable_early_reflections = bool(enable_early_reflections)
|
||||
self.enable_late_room = bool(enable_late_room)
|
||||
self.field = field
|
||||
if self.source_count != INPUT_CHANNELS:
|
||||
raise ValueError(f"native SOFA backend requires {INPUT_CHANNELS} sources")
|
||||
if abs(self.field.sample_rate_hz - 48000.0) > 1.0e-9:
|
||||
raise ValueError("the native SOFA binaural runtime requires 48 kHz")
|
||||
self.room_config = room_config
|
||||
self.room_config.validate()
|
||||
self.dsp_backend = "native-sofa"
|
||||
self.hrtf_input_kind = "field"
|
||||
self.hrtf_input_path = None
|
||||
self.cache_policy = None
|
||||
|
||||
self.library_path = find_native_library(library_path)
|
||||
self._lib = ctypes.CDLL(str(self.library_path))
|
||||
self._bind()
|
||||
version = int(self._lib.ejoc_abi_version())
|
||||
if version != ABI_VERSION:
|
||||
raise RuntimeError(
|
||||
f"native ABI mismatch: expected {ABI_VERSION}, got {version}")
|
||||
self._handle = self._lib.ejoc_sofa_binaural_create()
|
||||
if not self._handle:
|
||||
raise RuntimeError("native SOFA binaural renderer creation failed")
|
||||
try:
|
||||
self._configure_kernels()
|
||||
self._configure_field()
|
||||
self._configure_room(fdn_config)
|
||||
except Exception:
|
||||
self.close()
|
||||
raise
|
||||
|
||||
self.positions = np.zeros((self.source_count, 3), dtype=np.float64)
|
||||
self.positions[:, 1] = 1.0
|
||||
self.profiles = [self.default_profile] * self.source_count
|
||||
self.user_gain = np.ones(self.source_count, dtype=np.float64)
|
||||
self.special_lfe = np.zeros(self.source_count, dtype=bool)
|
||||
self.parameter_updates = 0
|
||||
self.finished = False
|
||||
for source in range(self.source_count):
|
||||
self.set_source(
|
||||
source, self.positions[source], profile=self.default_profile,
|
||||
fade=False)
|
||||
|
||||
def _bind(self):
|
||||
void_p = ctypes.c_void_p
|
||||
f64_p = ctypes.POINTER(ctypes.c_double)
|
||||
i16_p = ctypes.POINTER(ctypes.c_int16)
|
||||
u32_p = ctypes.POINTER(ctypes.c_uint32)
|
||||
self._lib.ejoc_abi_version.argtypes = []
|
||||
self._lib.ejoc_abi_version.restype = ctypes.c_uint32
|
||||
self._lib.ejoc_sofa_binaural_create.argtypes = []
|
||||
self._lib.ejoc_sofa_binaural_create.restype = void_p
|
||||
self._lib.ejoc_sofa_binaural_destroy.argtypes = [void_p]
|
||||
self._lib.ejoc_sofa_binaural_destroy.restype = None
|
||||
self._lib.ejoc_sofa_binaural_reset.argtypes = [void_p]
|
||||
self._lib.ejoc_sofa_binaural_reset.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_last_error.argtypes = [void_p]
|
||||
self._lib.ejoc_sofa_binaural_last_error.restype = ctypes.c_char_p
|
||||
self._lib.ejoc_sofa_binaural_configure_kernels.argtypes = [
|
||||
void_p, f64_p, f64_p, i16_p, f64_p, ctypes.c_uint32, f64_p, f64_p]
|
||||
self._lib.ejoc_sofa_binaural_configure_kernels.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_configure_field.argtypes = [
|
||||
void_p, f64_p, f64_p, f64_p, f64_p, ctypes.c_double]
|
||||
self._lib.ejoc_sofa_binaural_configure_field.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_configure_room.argtypes = [
|
||||
void_p, f64_p, f64_p, f64_p, ctypes.c_double, u32_p, f64_p,
|
||||
ctypes.c_double, ctypes.c_double, u32_p, f64_p,
|
||||
ctypes.c_uint32, ctypes.c_uint32]
|
||||
self._lib.ejoc_sofa_binaural_configure_room.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_set_source.argtypes = [
|
||||
void_p, ctypes.c_uint32, f64_p, ctypes.c_uint32, ctypes.c_double,
|
||||
ctypes.c_uint32, ctypes.c_uint32, ctypes.c_uint32]
|
||||
self._lib.ejoc_sofa_binaural_set_source.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_process.argtypes = [
|
||||
void_p, f64_p, ctypes.c_uint32, ctypes.c_double, f64_p]
|
||||
self._lib.ejoc_sofa_binaural_process.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_finish.argtypes = [
|
||||
void_p, ctypes.c_uint32, f64_p, ctypes.c_uint32]
|
||||
self._lib.ejoc_sofa_binaural_finish.restype = ctypes.c_int
|
||||
|
||||
def _raise(self, operation, status):
|
||||
message = self._lib.ejoc_sofa_binaural_last_error(self._handle)
|
||||
detail = (message or b"").decode("utf-8", "replace")
|
||||
raise RuntimeError(
|
||||
f"native SOFA binaural renderer {operation} failed ({status}): {detail}")
|
||||
|
||||
@staticmethod
|
||||
def _f64_pointer(values):
|
||||
return values.ctypes.data_as(ctypes.POINTER(ctypes.c_double))
|
||||
|
||||
def _configure_kernels(self):
|
||||
tables = load_filterbank_tables(DEFAULT_FILTERBANK_DATA)
|
||||
qmf_analysis = np.ascontiguousarray(
|
||||
tables["qmf_analysis_coefficients"], dtype=np.float64)
|
||||
hybrid_low = np.ascontiguousarray(
|
||||
tables["hybrid_analysis_low_kernel"], dtype=np.float64)
|
||||
hybrid_indices = np.ascontiguousarray(
|
||||
tables["hybrid_synthesis_indices"], dtype=np.int16)
|
||||
hybrid_values = np.ascontiguousarray(
|
||||
tables["hybrid_synthesis_values"], dtype=np.float64)
|
||||
qmf_basis = np.ascontiguousarray(
|
||||
tables["qmf_synthesis_basis"], dtype=np.float64)
|
||||
qmf_taps = np.ascontiguousarray(
|
||||
tables["qmf_synthesis_taps"], dtype=np.float64)
|
||||
status = self._lib.ejoc_sofa_binaural_configure_kernels(
|
||||
self._handle,
|
||||
self._f64_pointer(qmf_analysis),
|
||||
self._f64_pointer(hybrid_low),
|
||||
hybrid_indices.ctypes.data_as(ctypes.POINTER(ctypes.c_int16)),
|
||||
self._f64_pointer(hybrid_values),
|
||||
len(hybrid_indices),
|
||||
self._f64_pointer(qmf_basis),
|
||||
self._f64_pointer(qmf_taps))
|
||||
if status:
|
||||
self._raise("configure_kernels", status)
|
||||
self._keepalive = (qmf_analysis, hybrid_low, hybrid_indices,
|
||||
hybrid_values, qmf_basis, qmf_taps)
|
||||
|
||||
def _configure_field(self):
|
||||
coefficients = np.ascontiguousarray(
|
||||
self.field.coefficients, dtype=np.complex128).view(np.float64)
|
||||
delay_coefficients = np.ascontiguousarray(
|
||||
self.field.delay_coefficients, dtype=np.float64)
|
||||
delay_bounds = np.ascontiguousarray(
|
||||
self.field.delay_bounds, dtype=np.float64)
|
||||
centers = np.ascontiguousarray(
|
||||
self.field.band_center_frequencies_hz, dtype=np.float64)
|
||||
status = self._lib.ejoc_sofa_binaural_configure_field(
|
||||
self._handle,
|
||||
self._f64_pointer(coefficients),
|
||||
self._f64_pointer(delay_coefficients),
|
||||
self._f64_pointer(delay_bounds),
|
||||
self._f64_pointer(centers),
|
||||
float(self.field.measurement_radius_m))
|
||||
if status:
|
||||
self._raise("configure_field", status)
|
||||
|
||||
def _configure_room(self, fdn_config: LateFdnConfig | None):
|
||||
actual = fdn_config or LateFdnConfig(sample_rate_hz=48000.0)
|
||||
numbers = _room_numbers(actual)
|
||||
dims = np.asarray(self.room_config.dimensions_m, dtype=np.float64)
|
||||
listener = np.asarray(self.room_config.listener_position_m, dtype=np.float64)
|
||||
walls = np.asarray(self.room_config.wall_reflection_gain, dtype=np.float64)
|
||||
status = self._lib.ejoc_sofa_binaural_configure_room(
|
||||
self._handle,
|
||||
self._f64_pointer(dims),
|
||||
self._f64_pointer(listener),
|
||||
self._f64_pointer(walls),
|
||||
float(self.room_config.speed_of_sound_m_s),
|
||||
numbers["fdn_delays"].ctypes.data_as(ctypes.POINTER(ctypes.c_uint32)),
|
||||
self._f64_pointer(numbers["fdn_feedback"]),
|
||||
float(numbers["damping"]),
|
||||
float(numbers["output_gain"]),
|
||||
numbers["allpass_delays"].ctypes.data_as(ctypes.POINTER(ctypes.c_uint32)),
|
||||
self._f64_pointer(numbers["allpass_gains"]),
|
||||
1 if self.enable_early_reflections else 0,
|
||||
1 if self.enable_late_room else 0)
|
||||
if status:
|
||||
self._raise("configure_room", status)
|
||||
self._fdn = SharedUnitaryFdn(actual)
|
||||
self._fdn_tail_samples = numbers["tail_samples"]
|
||||
|
||||
def set_source(self, source: int, position_adm, *, profile: str | None = None,
|
||||
gain: float = 1.0, enabled: bool = True,
|
||||
special_lfe: bool = False, fade: bool = True) -> None:
|
||||
if self.finished:
|
||||
raise RuntimeError("SOFA renderer is finished")
|
||||
source = int(source)
|
||||
if not 0 <= source < self.source_count:
|
||||
raise IndexError(source)
|
||||
name = self.default_profile if profile is None else profile
|
||||
position = np.asarray(position_adm, dtype=np.float64)
|
||||
if position.shape != (3,):
|
||||
raise ValueError("ADM position must contain three Cartesian values")
|
||||
status = self._lib.ejoc_sofa_binaural_set_source(
|
||||
self._handle, source,
|
||||
self._f64_pointer(np.ascontiguousarray(position)),
|
||||
_PROFILE_INDEX[ReferenceDistanceProfileV1.validate_profile(name)],
|
||||
float(gain), 1 if enabled else 0, 1 if special_lfe else 0,
|
||||
1 if fade else 0)
|
||||
if status:
|
||||
self._raise("set_source", status)
|
||||
self.positions[source] = position
|
||||
self.profiles[source] = name
|
||||
self.user_gain[source] = float(gain)
|
||||
self.special_lfe[source] = bool(special_lfe)
|
||||
self.parameter_updates += 1
|
||||
|
||||
def process(self, sources) -> np.ndarray:
|
||||
if self.finished:
|
||||
raise RuntimeError("SOFA renderer is finished")
|
||||
values = np.ascontiguousarray(sources, dtype=np.float64)
|
||||
if values.ndim != 2 or values.shape[1] != self.source_count:
|
||||
raise ValueError(f"sources must have shape [samples,{self.source_count}]")
|
||||
if len(values) % QMF_HOP or len(values) > BLOCK_SAMPLES:
|
||||
raise ValueError("native SOFA backend input must be a 64-aligned block")
|
||||
output = np.empty((len(values), OUTPUT_CHANNELS), dtype=np.float64)
|
||||
count = self._lib.ejoc_sofa_binaural_process(
|
||||
self._handle, self._f64_pointer(values), len(values), 1.0,
|
||||
self._f64_pointer(output))
|
||||
if count < 0:
|
||||
self._raise("process", count)
|
||||
return output[:count]
|
||||
|
||||
def finish(self, *, tail_seconds: float | None = None) -> np.ndarray:
|
||||
if self.finished:
|
||||
return np.zeros((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
if tail_seconds is not None and (
|
||||
not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0):
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
flush = self.finish_output_capacity(tail_seconds)
|
||||
pieces = []
|
||||
remaining = flush
|
||||
while remaining > 0:
|
||||
chunk = min(BLOCK_SAMPLES, remaining)
|
||||
output = np.empty((chunk, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
count = self._lib.ejoc_sofa_binaural_finish(
|
||||
self._handle, chunk, self._f64_pointer(output), chunk)
|
||||
if count < 0:
|
||||
self._raise("finish", count)
|
||||
pieces.append(output[:count])
|
||||
remaining -= chunk
|
||||
self.finished = True
|
||||
nonempty = [piece for piece in pieces if len(piece)]
|
||||
if not nonempty:
|
||||
return np.zeros((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
return np.concatenate(nonempty, axis=0)
|
||||
|
||||
def finish_output_capacity(self, tail_seconds: float | None = None) -> int:
|
||||
if tail_seconds is not None and (
|
||||
not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0):
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
requested = (self._fdn_tail_samples if tail_seconds is None
|
||||
else int(math.ceil(float(tail_seconds) * 48000.0)))
|
||||
maximum_hrtf = float(np.max(self.field.delay_bounds[:, 1], initial=0.0))
|
||||
hrtf_slots = int(math.ceil(maximum_hrtf / QMF_HOP))
|
||||
hrtf_bound = hrtf_slots * QMF_HOP
|
||||
early_bound = hrtf_bound + 2048
|
||||
if self.enable_early_reflections:
|
||||
early_bound += 256 * QMF_HOP
|
||||
drain = max(requested if self.enable_late_room else 0, early_bound)
|
||||
drain += LATENCY_SAMPLES
|
||||
return int(math.ceil(drain / QMF_HOP) * QMF_HOP)
|
||||
|
||||
def reset(self) -> None:
|
||||
if self._lib.ejoc_sofa_binaural_reset(self._handle):
|
||||
self._raise("reset", -1)
|
||||
self.finished = False
|
||||
for source in range(self.source_count):
|
||||
self.set_source(
|
||||
source, self.positions[source], profile=self.profiles[source],
|
||||
gain=float(self.user_gain[source]),
|
||||
special_lfe=bool(self.special_lfe[source]), fade=False)
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"name": "NativeSofaBinauralDsp",
|
||||
"source_count": self.source_count,
|
||||
"sample_rate_hz": 48000.0,
|
||||
"precision": "float64/complex128",
|
||||
"signal_path": (
|
||||
"native 64-QMF -> native 77-hybrid -> SOFA order-5 real-SH "
|
||||
"direct/early -> native synthesis + shared unitary FDN"),
|
||||
"hrtf_input_kind": self.hrtf_input_kind,
|
||||
"hrtf_input_path": self.hrtf_input_path,
|
||||
"cache_policy": self.cache_policy,
|
||||
"latency_compensated_samples": LATENCY_SAMPLES,
|
||||
"enable_early_reflections": self.enable_early_reflections,
|
||||
"enable_late_room": self.enable_late_room,
|
||||
"early_history_slots": 256,
|
||||
"hrtf_history_slots": int(math.ceil(
|
||||
float(np.max(self.field.delay_bounds[:, 1], initial=0.0)) / QMF_HOP)),
|
||||
"maximum_hrtf_delay_samples": float(
|
||||
np.max(self.field.delay_bounds[:, 1], initial=0.0)),
|
||||
"parameter_updates": self.parameter_updates,
|
||||
"distance": ReferenceDistanceProfileV1.info(),
|
||||
"field": self.field.info(),
|
||||
"library_path": str(self.library_path),
|
||||
"native_backend": True,
|
||||
}
|
||||
|
||||
def close(self) -> None:
|
||||
handle = getattr(self, "_handle", None)
|
||||
if handle:
|
||||
self._lib.ejoc_sofa_binaural_destroy(handle)
|
||||
self._handle = None
|
||||
self.finished = True
|
||||
|
||||
|
||||
def create_native_sofa_renderer(
|
||||
sofa, *, mode="mid", cache_policy="memory", cache_dir=None,
|
||||
shell_radius_m=1.0, object_delay_samples=1473, tail_seconds=5.0,
|
||||
output_gain=1.0, chunk_frames=64):
|
||||
"""Compile a SOFA source and build a JOC adapter over the native DSP."""
|
||||
from binaural_renderer import SofaBinauralRenderer, resolve_sofa_hrtf
|
||||
|
||||
source = resolve_sofa_hrtf(sofa)
|
||||
field = compile_sofa_hrtf(
|
||||
source,
|
||||
shell_radius_m=shell_radius_m,
|
||||
cache_policy=cache_policy,
|
||||
cache_dir=cache_dir)
|
||||
backend = NativeSofaBinauralDsp(field, default_profile=mode)
|
||||
backend.hrtf_input_kind = "sofa"
|
||||
backend.hrtf_input_path = str(source)
|
||||
backend.cache_policy = str(cache_policy).lower()
|
||||
return SofaBinauralRenderer(
|
||||
backend, mode=mode, object_delay_samples=object_delay_samples,
|
||||
tail_seconds=tail_seconds, chunk_frames=chunk_frames)
|
||||
|
||||
|
||||
def create_native_compiled_cache_renderer(
|
||||
cache, *, mode="mid", object_delay_samples=1473, tail_seconds=5.0,
|
||||
output_gain=1.0, chunk_frames=64):
|
||||
"""Load a compiled cache and build a JOC adapter over the native DSP."""
|
||||
from binaural_renderer import SofaBinauralRenderer, resolve_compiled_hrtf_cache
|
||||
|
||||
source = resolve_compiled_hrtf_cache(cache)
|
||||
field = SofaHrtfField.load(source)
|
||||
backend = NativeSofaBinauralDsp(field, default_profile=mode)
|
||||
backend.hrtf_input_kind = "compiled_cache"
|
||||
backend.hrtf_input_path = str(source)
|
||||
backend.cache_policy = None
|
||||
return SofaBinauralRenderer(
|
||||
backend, mode=mode, object_delay_samples=object_delay_samples,
|
||||
tail_seconds=tail_seconds, chunk_frames=chunk_frames)
|
||||
+3
-5
@@ -1,7 +1,6 @@
|
||||
"""Shared PCM spool, peak analysis, and WAV writer for direct outputs."""
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
import struct
|
||||
from pathlib import Path
|
||||
|
||||
@@ -40,9 +39,8 @@ class PcmSpool:
|
||||
if self.expected_samples is not None and not (
|
||||
0 <= self.expected_samples <= self.sample_capacity):
|
||||
raise ValueError("expected_samples exceeds sample_capacity")
|
||||
if self.tail_threshold is not None and (
|
||||
not math.isfinite(self.tail_threshold) or self.tail_threshold < 0.0):
|
||||
raise ValueError("tail_threshold must be finite and non-negative")
|
||||
if self.tail_threshold is not None and self.tail_threshold < 0.0:
|
||||
raise ValueError("tail_threshold must be non-negative")
|
||||
self.position = 0
|
||||
self.peak = 0.0
|
||||
self.clipped_values = 0
|
||||
@@ -116,7 +114,7 @@ class SpeakerPcmSpool(PcmSpool):
|
||||
|
||||
|
||||
class BinauralPcmSpool(PcmSpool):
|
||||
"""Float64 variable-tail spool for the binaural renderer."""
|
||||
"""Float64 variable-tail spool for the Rosella binaural renderer."""
|
||||
|
||||
def __init__(self, path, sample_capacity, *, tail_threshold=1.0e-8):
|
||||
super().__init__(
|
||||
|
||||
@@ -1,144 +0,0 @@
|
||||
"""Orthonormal real spherical harmonics in ACN order, through fifth order."""
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
import numpy as np
|
||||
from scipy.spatial import SphericalVoronoi
|
||||
|
||||
|
||||
def _associated_legendre(order: int, degree: int, x: np.ndarray) -> np.ndarray:
|
||||
"""P_degree^order(x), including the Condon-Shortley phase."""
|
||||
m = int(order)
|
||||
l = int(degree)
|
||||
if not 0 <= m <= l:
|
||||
raise ValueError("associated Legendre indices require 0 <= m <= l")
|
||||
x = np.asarray(x, dtype=np.float64)
|
||||
p_mm = np.ones_like(x)
|
||||
if m:
|
||||
double_factorial = 1.0
|
||||
for value in range(1, 2 * m, 2):
|
||||
double_factorial *= value
|
||||
p_mm = ((-1.0) ** m) * double_factorial * np.power(
|
||||
np.maximum(0.0, 1.0 - x * x), 0.5 * m)
|
||||
if l == m:
|
||||
return p_mm
|
||||
p_m1 = x * (2 * m + 1) * p_mm
|
||||
if l == m + 1:
|
||||
return p_m1
|
||||
previous_previous = p_mm
|
||||
previous = p_m1
|
||||
for current_degree in range(m + 2, l + 1):
|
||||
current = (
|
||||
(2 * current_degree - 1) * x * previous
|
||||
- (current_degree + m - 1) * previous_previous
|
||||
) / float(current_degree - m)
|
||||
previous_previous, previous = previous, current
|
||||
return previous
|
||||
|
||||
|
||||
def real_spherical_harmonics(directions, order: int = 5) -> np.ndarray:
|
||||
"""Return [directions,(order+1)^2] ACN/N3D real harmonics.
|
||||
|
||||
Coordinates use SOFA listener axes: +X front, +Y left, +Z up. The basis is
|
||||
orthonormal over the sphere and includes the Condon-Shortley phase.
|
||||
"""
|
||||
maximum_order = int(order)
|
||||
if not 0 <= maximum_order <= 12:
|
||||
raise ValueError("supported spherical-harmonic orders are 0..12")
|
||||
vectors = np.asarray(directions, dtype=np.float64)
|
||||
one = vectors.ndim == 1
|
||||
if one:
|
||||
vectors = vectors[None, :]
|
||||
if vectors.ndim != 2 or vectors.shape[1] != 3 or not np.isfinite(vectors).all():
|
||||
raise ValueError("directions must have finite shape [M,3]")
|
||||
length = np.linalg.norm(vectors, axis=1)
|
||||
if np.any(length <= 1.0e-15):
|
||||
raise ValueError("spherical-harmonic directions must be non-zero")
|
||||
unit = vectors / length[:, None]
|
||||
azimuth = np.arctan2(unit[:, 1], unit[:, 0])
|
||||
cos_colatitude = np.clip(unit[:, 2], -1.0, 1.0)
|
||||
result = np.empty((len(unit), (maximum_order + 1) ** 2), dtype=np.float64)
|
||||
column = 0
|
||||
for degree in range(maximum_order + 1):
|
||||
for m in range(-degree, degree + 1):
|
||||
absolute = abs(m)
|
||||
normalization = math.sqrt(
|
||||
(2 * degree + 1) / (4.0 * math.pi)
|
||||
* math.factorial(degree - absolute)
|
||||
/ math.factorial(degree + absolute))
|
||||
legendre = _associated_legendre(absolute, degree, cos_colatitude)
|
||||
if m < 0:
|
||||
value = math.sqrt(2.0) * normalization * legendre * np.sin(
|
||||
absolute * azimuth)
|
||||
elif m > 0:
|
||||
value = math.sqrt(2.0) * normalization * legendre * np.cos(
|
||||
m * azimuth)
|
||||
else:
|
||||
value = normalization * legendre
|
||||
result[:, column] = value
|
||||
column += 1
|
||||
return result[0] if one else result
|
||||
|
||||
|
||||
def spherical_voronoi_weights(directions) -> np.ndarray:
|
||||
"""Area weights for an irregular full-sphere grid, with uniform fallback."""
|
||||
vectors = np.asarray(directions, dtype=np.float64)
|
||||
if vectors.ndim != 2 or vectors.shape[1] != 3:
|
||||
raise ValueError("directions must have shape [M,3]")
|
||||
unit = vectors / np.linalg.norm(vectors, axis=1)[:, None]
|
||||
if len(unit) < 4:
|
||||
return np.full(len(unit), 1.0 / len(unit), dtype=np.float64)
|
||||
try:
|
||||
voronoi = SphericalVoronoi(unit, radius=1.0, center=np.zeros(3))
|
||||
areas = np.asarray(voronoi.calculate_areas(), dtype=np.float64)
|
||||
if not np.isfinite(areas).all() or np.any(areas <= 0.0):
|
||||
raise ValueError("invalid spherical Voronoi areas")
|
||||
return areas / np.sum(areas, dtype=np.float64)
|
||||
except (ValueError, RuntimeError, np.linalg.LinAlgError):
|
||||
return np.full(len(unit), 1.0 / len(unit), dtype=np.float64)
|
||||
|
||||
|
||||
def fit_real_spherical_harmonics(directions, values, *, order: int = 5,
|
||||
ridge: float = 1.0e-6,
|
||||
weights=None) -> np.ndarray:
|
||||
"""Weighted ridge fit. Output shape is [terms,...value trailing axes]."""
|
||||
basis = real_spherical_harmonics(directions, order=order)
|
||||
target = np.asarray(values)
|
||||
if target.shape[0] != basis.shape[0]:
|
||||
raise ValueError("spherical-harmonic target count does not match directions")
|
||||
if target.dtype.kind == "c":
|
||||
target = np.asarray(target, dtype=np.complex128)
|
||||
solve_dtype = np.complex128
|
||||
else:
|
||||
target = np.asarray(target, dtype=np.float64)
|
||||
solve_dtype = np.float64
|
||||
if weights is None:
|
||||
weight = spherical_voronoi_weights(directions)
|
||||
else:
|
||||
weight = np.asarray(weights, dtype=np.float64)
|
||||
if weight.shape != (len(basis),) or np.any(weight < 0.0) or not np.isfinite(weight).all():
|
||||
raise ValueError("weights must be finite non-negative [M]")
|
||||
total = float(np.sum(weight))
|
||||
if total <= 0.0:
|
||||
raise ValueError("weights must have positive sum")
|
||||
weight = weight / total
|
||||
flat = target.reshape(len(target), -1)
|
||||
weighted_basis = basis * weight[:, None]
|
||||
gram = basis.T @ weighted_basis
|
||||
regularization = float(ridge)
|
||||
if not math.isfinite(regularization) or regularization < 0.0:
|
||||
raise ValueError("ridge must be finite and non-negative")
|
||||
scale = float(np.trace(gram)) / gram.shape[0]
|
||||
system = gram + np.eye(gram.shape[0], dtype=np.float64) * regularization * scale
|
||||
right = basis.T @ (weight[:, None] * flat)
|
||||
coefficients = np.linalg.solve(system.astype(solve_dtype), right.astype(solve_dtype))
|
||||
return coefficients.reshape((basis.shape[1],) + target.shape[1:])
|
||||
|
||||
|
||||
def evaluate_real_spherical_harmonics(coefficients, directions,
|
||||
*, order: int = 5) -> np.ndarray:
|
||||
basis = real_spherical_harmonics(directions, order=order)
|
||||
coeff = np.asarray(coefficients)
|
||||
if coeff.shape[0] != (int(order) + 1) ** 2:
|
||||
raise ValueError("coefficient term count does not match order")
|
||||
return np.tensordot(basis, coeff, axes=([-1], [0]))
|
||||
@@ -0,0 +1,243 @@
|
||||
import importlib.util
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
SRC = ROOT / "src"
|
||||
if str(SRC) not in sys.path:
|
||||
sys.path.insert(0, str(SRC))
|
||||
if str(ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
import main
|
||||
from adm_atmos import q_to_adm_xyz
|
||||
from binaural_metadata import OamdPositionTimeline
|
||||
from binaural_renderer import (
|
||||
DEFAULT_PERSONALIZED_HEADPHONE,
|
||||
RosellaBinauralRenderer,
|
||||
resolve_personalized_headphone,
|
||||
)
|
||||
from oamd_bits import q_of
|
||||
from rosella_direct import BINAURAL_PROFILE_NAMES, direct_and_room_send, special_lfe_direct
|
||||
from rosella_filterbank import HybridAnalysis, HybridSynthesis, QmfAnalysis, QmfSynthesis
|
||||
from rosella_model import load_personalized_headphone
|
||||
from speaker_wav import BinauralPcmSpool, write_pcm_wav
|
||||
|
||||
_MODEL_CANDIDATES = list(
|
||||
(ROOT / "tests" / "binauraltests" / "evidence").glob(
|
||||
"*/test.personalized_headphone"))
|
||||
TEST_MODEL = _MODEL_CANDIDATES[0] if _MODEL_CANDIDATES else Path()
|
||||
|
||||
|
||||
class BinauralProductionTest(unittest.TestCase):
|
||||
def test_cli_exposes_only_near_mid_far_and_defaults_mid_float32(self):
|
||||
parser = main.build_parser()
|
||||
args = parser.parse_args(["input.eac3", "--binaural"])
|
||||
self.assertEqual(args.binaural_mode, "mid")
|
||||
self.assertEqual(args.binaural_format, "float32")
|
||||
action = next(a for a in parser._actions if a.dest == "binaural_mode")
|
||||
self.assertEqual(tuple(action.choices), ("near", "mid", "far"))
|
||||
self.assertNotIn("off", BINAURAL_PROFILE_NAMES)
|
||||
|
||||
def test_default_model_path_and_missing_model_message(self):
|
||||
self.assertEqual(
|
||||
DEFAULT_PERSONALIZED_HEADPHONE,
|
||||
ROOT / "HRTF" / "binaural.personalized_headphone")
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
missing = Path(directory) / "missing.personalized_headphone"
|
||||
with self.assertRaisesRegex(FileNotFoundError, "未找到双耳模型"):
|
||||
resolve_personalized_headphone(missing)
|
||||
|
||||
def test_runtime_filterbanks_are_double_precision(self):
|
||||
qmf = QmfAnalysis(1)
|
||||
hybrid = HybridAnalysis(1)
|
||||
hybrid_synthesis = HybridSynthesis(1)
|
||||
qmf_synthesis = QmfSynthesis(1)
|
||||
self.assertEqual(qmf.coefficients.dtype, np.float64)
|
||||
self.assertEqual(qmf.history.dtype, np.float64)
|
||||
self.assertEqual(hybrid.low_kernel.dtype, np.float64)
|
||||
self.assertEqual(hybrid.high_history.dtype, np.complex128)
|
||||
self.assertEqual(qmf_synthesis.basis.dtype, np.float64)
|
||||
source = np.zeros((1, 1, 64), dtype=np.float64)
|
||||
q = qmf.process_chunk(source)
|
||||
h = hybrid.process_chunk(q)
|
||||
self.assertEqual(q.dtype, np.complex128)
|
||||
self.assertEqual(h.dtype, np.complex128)
|
||||
back = hybrid_synthesis.process_chunk(h)
|
||||
time = qmf_synthesis.process_chunk(back)
|
||||
self.assertEqual(back.dtype, np.complex128)
|
||||
self.assertEqual(time.dtype, np.float64)
|
||||
|
||||
@unittest.skipUnless(
|
||||
(ROOT / "tests" / "binauraltests" / "hqmf.py").is_file(),
|
||||
"research filterbank reference not present")
|
||||
def test_compact_kernel_archive_matches_research_float64_filterbanks(self):
|
||||
reference_root = ROOT / "tests" / "binauraltests"
|
||||
spec = importlib.util.spec_from_file_location(
|
||||
"binaural_reference_hqmf", reference_root / "hqmf.py")
|
||||
reference = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(reference)
|
||||
evidence = reference_root / "evidence"
|
||||
source = np.fromfile(
|
||||
evidence / "hqmf_kernel" / "hqmf_validation_input.f32le",
|
||||
dtype="<f4").astype(np.float64).reshape(-1, 1, 64)
|
||||
production_qmf = QmfAnalysis(1).process_chunk(source)
|
||||
reference_qmf = reference.QmfAnalysisFast(
|
||||
evidence / "hqmf_kernel" / "hqmf_analysis_manifest.json",
|
||||
1, dtype=np.float64).process_chunk(source)
|
||||
np.testing.assert_array_equal(production_qmf, reference_qmf)
|
||||
|
||||
production_hybrid = HybridAnalysis(1).process_chunk(production_qmf)
|
||||
reference_hybrid = reference.HybridAnalysis(
|
||||
evidence / "hybrid_kernel" / "hybrid_analysis_manifest.json",
|
||||
1, dtype=np.float64).process_chunk(reference_qmf)
|
||||
np.testing.assert_array_equal(production_hybrid, reference_hybrid)
|
||||
|
||||
production_qmf_back = HybridSynthesis(1).process_chunk(production_hybrid)
|
||||
reference_qmf_back = reference.HybridSynthesis(
|
||||
evidence / "hybrid_synthesis_kernel" / "hybrid_synthesis_manifest.json",
|
||||
1, dtype=np.float64).process_chunk(reference_hybrid)
|
||||
np.testing.assert_array_equal(production_qmf_back, reference_qmf_back)
|
||||
|
||||
production_time = QmfSynthesis(1).process_chunk(production_qmf_back)
|
||||
reference_time = reference.QmfSynthesisFast(
|
||||
evidence / "qmf_synthesis_kernel" / "qmf_synthesis_manifest.json",
|
||||
1, dtype=np.float64).process_chunk(reference_qmf_back)
|
||||
np.testing.assert_array_equal(production_time, reference_time)
|
||||
|
||||
def test_special_lfe_is_fixed_16_band_complex128_without_room_send(self):
|
||||
result = special_lfe_direct()
|
||||
self.assertEqual(result.gains.dtype, np.complex128)
|
||||
np.testing.assert_array_equal(result.gains[0], result.gains[1])
|
||||
self.assertTrue(np.any(result.gains[:, :16] != 0))
|
||||
np.testing.assert_array_equal(result.gains[:, 16:], 0)
|
||||
self.assertEqual(float(result.room_send), 0.0)
|
||||
|
||||
def test_oamd_timing_retains_outer_block_delay_and_ramp(self):
|
||||
timeline = OamdPositionTimeline()
|
||||
initial = {
|
||||
"values": {
|
||||
(1, "q1"): q_of(0, 62),
|
||||
(1, "q2"): q_of(0, 62),
|
||||
(1, "q3"): q_of(0, 15),
|
||||
},
|
||||
"block_offset_samples": 0,
|
||||
"ramp_duration_samples": 1536,
|
||||
}
|
||||
timeline.submit_update(initial, frame_start_sample=0)
|
||||
old = np.asarray(q_to_adm_xyz(q_of(0, 62), q_of(0, 62), q_of(0, 15)))
|
||||
np.testing.assert_allclose(timeline.positions_at(0)[0], old)
|
||||
|
||||
target_q1 = q_of(62, 62)
|
||||
changed = {
|
||||
"values": {(1, "q1"): target_q1},
|
||||
"block_offset_samples": 32,
|
||||
"ramp_duration_samples": 1536,
|
||||
}
|
||||
timeline.submit_update(
|
||||
changed,
|
||||
frame_start_sample=1536,
|
||||
outer_sample_offset=16,
|
||||
object_delay_samples=1473,
|
||||
)
|
||||
start = 1536 + 16 + 32 + 1473 + 64
|
||||
duration = 1536 - 64
|
||||
target = np.asarray(q_to_adm_xyz(target_q1, q_of(0, 62), q_of(0, 15)))
|
||||
np.testing.assert_allclose(timeline.positions_at(start - 1)[0], old)
|
||||
np.testing.assert_allclose(timeline.positions_at(start)[0], old)
|
||||
np.testing.assert_allclose(
|
||||
timeline.positions_at(start + duration // 2)[0],
|
||||
old + (target - old) * 0.5,
|
||||
)
|
||||
np.testing.assert_allclose(timeline.positions_at(start + duration)[0], target)
|
||||
|
||||
@unittest.skipUnless(TEST_MODEL.is_file(), "test model not present")
|
||||
def test_model_parser_and_direct_path_promote_to_double(self):
|
||||
model = load_personalized_headphone(TEST_MODEL)
|
||||
self.assertEqual(model.sample_rate, 48000)
|
||||
self.assertEqual(len(model.coefficients), 16033)
|
||||
self.assertEqual(
|
||||
model.coefficient_sha256,
|
||||
"2d4b40c27925ec8827556585d90c371384458d31bb84b20f94a946145558ea27",
|
||||
)
|
||||
result = direct_and_room_send(model, (0.0, 1.0, 0.0), 3)
|
||||
self.assertEqual(result.gains.dtype, np.complex128)
|
||||
with self.assertRaisesRegex(ValueError, "near, mid, or far"):
|
||||
direct_and_room_send(model, (0.0, 1.0, 0.0), 0)
|
||||
|
||||
@unittest.skipUnless(TEST_MODEL.is_file(), "test model not present")
|
||||
def test_renderer_keeps_float64_state_and_compensates_961_samples(self):
|
||||
renderer = RosellaBinauralRenderer(
|
||||
TEST_MODEL,
|
||||
chunk_frames=1,
|
||||
tail_seconds=0,
|
||||
room_impulse_slots=32,
|
||||
)
|
||||
source = np.zeros((1536, 16), dtype=np.float32)
|
||||
source[0, 1] = 1.0
|
||||
first = renderer.render_frame(source)
|
||||
tail = renderer.finish()
|
||||
self.assertEqual(first.dtype, np.float64)
|
||||
self.assertEqual(tail.dtype, np.float64)
|
||||
self.assertEqual(len(first), 1536 - 961)
|
||||
self.assertEqual(renderer.qmf_analysis.history.dtype, np.float64)
|
||||
self.assertEqual(renderer.hybrid_analysis.high_history.dtype, np.complex128)
|
||||
self.assertEqual(renderer.core.gains.dtype, np.complex128)
|
||||
self.assertEqual(renderer.core.room.tail.dtype, np.complex128)
|
||||
|
||||
@unittest.skipUnless(TEST_MODEL.is_file(), "test model not present")
|
||||
def test_native_backend_matches_python_backend(self):
|
||||
source = np.zeros((1536, 16), dtype=np.float32)
|
||||
source[0, 0] = 0.1
|
||||
source[64, 1] = -0.2
|
||||
python_renderer = RosellaBinauralRenderer(
|
||||
TEST_MODEL, backend="python", chunk_frames=1,
|
||||
tail_seconds=0, room_impulse_slots=64)
|
||||
try:
|
||||
native_renderer = RosellaBinauralRenderer(
|
||||
TEST_MODEL, backend="native",
|
||||
native_library=ROOT / "lib" / "eac3joc_core.dll",
|
||||
chunk_frames=1, tail_seconds=0)
|
||||
except (OSError, RuntimeError):
|
||||
python_renderer.close()
|
||||
self.skipTest("native binaural backend not built")
|
||||
try:
|
||||
self.assertEqual(native_renderer.dsp_backend, "native")
|
||||
python_output = np.concatenate((
|
||||
python_renderer.render_frame(source), python_renderer.finish()))
|
||||
native_output = np.concatenate((
|
||||
native_renderer.render_frame(source), native_renderer.finish()))
|
||||
np.testing.assert_allclose(
|
||||
native_output, python_output, rtol=0.0, atol=2.0e-14)
|
||||
finally:
|
||||
python_renderer.close()
|
||||
native_renderer.close()
|
||||
|
||||
def test_binaural_spool_preserves_float64_then_trims_tail(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
raw = Path(directory) / "binaural.raw"
|
||||
wav = Path(directory) / "binaural.wav"
|
||||
spool = BinauralPcmSpool(raw, 8, tail_threshold=1.0e-8)
|
||||
values = np.asarray([
|
||||
[0.0, 0.0],
|
||||
[0.25, -0.25],
|
||||
[1.0 + 2.0 ** -40, 0.0],
|
||||
[1.0e-9, 0.0],
|
||||
], dtype=np.float64)
|
||||
spool.write_frame(values)
|
||||
spool.finalize(minimum_samples=2)
|
||||
self.assertEqual(spool.values.dtype, np.float64)
|
||||
self.assertEqual(spool.sample_count, 3)
|
||||
self.assertEqual(spool.clipped_values, 1)
|
||||
info = write_pcm_wav(wav, spool.values, "float32")
|
||||
self.assertEqual(info["sample_count"], 3)
|
||||
self.assertEqual(info["channel_count"], 2)
|
||||
spool.close()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
Reference in New Issue
Block a user