Compare commits
14 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| afef6d2c52 | |||
| ab7e815a4d | |||
| b634326b5d | |||
| cf50beafcb | |||
| c8676d7aa3 | |||
| b619dee523 | |||
| 7da3a5eb95 | |||
| 128286153e | |||
| 85105f21d4 | |||
| 329445ed25 | |||
| 887b3e317f | |||
| 35536b0bd4 | |||
| 9b41fedb14 | |||
| 65a08e544d |
@@ -1,11 +1,14 @@
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
.pytest_cache/
|
||||
|
||||
.venv/
|
||||
venv/
|
||||
|
||||
build/
|
||||
output/
|
||||
tests/
|
||||
lib/
|
||||
metadata_cache/
|
||||
|
||||
*.metadata.json
|
||||
@@ -13,3 +16,8 @@ metadata_cache/
|
||||
*.variant-error.json
|
||||
*.objects16.f32le
|
||||
|
||||
HRTF/
|
||||
|
||||
*.sofa
|
||||
*.personalized_headphone
|
||||
*.jochrtf
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
MIT License
|
||||
|
||||
Copyright (c) 2026 TheM14
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
in the Software without restriction, including without limitation the rights
|
||||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||
copies of the Software, and to permit persons to whom the Software is
|
||||
furnished to do so, subject to the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be included in all
|
||||
copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||
SOFTWARE.
|
||||
+148
-12
@@ -4,9 +4,9 @@
|
||||
|
||||
> JustOneCacophony is an experimental/test implementation of E-AC-3 JOC for studying JOC parsing, reconstruction, rendering, and the associated mathematics.
|
||||
|
||||
The project can extract and parse EMDF, ID14 JOC parameters, and ID11 OAMD metadata from common E-AC-3 JOC streams. It combines those data with the core 5.1 PCM decoded by FFmpeg, reconstructs LFE plus 15 object channels, and writes either ADM BWF or a WAV file for a selected speaker layout.
|
||||
The project can extract and parse EMDF, ID14 JOC parameters, and ID11 OAMD metadata from common E-AC-3 JOC streams. It combines those data with the core 5.1 PCM decoded by FFmpeg, reconstructs LFE plus 15 object channels, and writes ADM BWF, a WAV file for a selected speaker layout, or direct binaural stereo using a standard SOFA HRTF.
|
||||
|
||||
This is research code, not a complete, standards-compliant, or production-grade Dolby JOC decoder. It covers only the stream forms currently implemented. Unknown variants fail explicitly—because when the math goes wrong, all that may remain is the cacophony.
|
||||
This is research code, not a complete, standards-compliant, or production-grade JOC decoder. It covers only the stream forms currently implemented. Unknown variants fail explicitly—because when the math goes wrong, all that may remain is the cacophony.
|
||||
|
||||
## Current features
|
||||
|
||||
@@ -16,7 +16,9 @@ This is research code, not a complete, standards-compliant, or production-grade
|
||||
- Reconstruct LFE plus 15 object channels through analysis QMF, parameter interpolation, the object matrix, and inverse QMF.
|
||||
- Write a 25-channel ADM BWF: a 10-channel 7.1.2 bed (silent except for LFE) plus 15 objects.
|
||||
- Render directly to `2.0`, `3.1`, `5.1`, `7.1`, `5.1.2`, `5.1.4`, `7.1.2`, `7.1.4`, `9.1.4`, or `9.1.6`.
|
||||
- Write float32 or PCM24 WAV and require an explicit policy when PCM24 would clip.
|
||||
- Run public SOFA binaural rendering directly from `pcm16 + ID11/OAMD`, without a temporary ADM BWF.
|
||||
- Keep the binaural DSP in float64/complex128, including 961-sample latency compensation, cross-frame state, and the room tail.
|
||||
- Use a shared float32/PCM24 WAV writer and explicit PCM24 clipping policy for direct outputs.
|
||||
- Use the NumPy backend or an optional C++20 core through `ctypes`; `auto` falls back to Python when the native library is unavailable.
|
||||
- Read or write metadata sidecars and produce metadata, timing, and output reports.
|
||||
|
||||
@@ -30,16 +32,19 @@ M4A / E-AC-3
|
||||
├─ ID11 OAMD → object positions and timing
|
||||
└─ LFE + 15 objects
|
||||
├─ 25ch ADM BWF
|
||||
└─ speaker WAV for the selected layout
|
||||
├─ speaker WAV for the selected layout
|
||||
└─ direct ID11 timeline + SOFA HRTF → binaural WAV
|
||||
```
|
||||
|
||||
The Python and C++ backends follow the same documented mathematics. The native core handles the state-heavy DSP and speaker rendering; high-level bitstream parsing, ADM assembly, and CLI behavior remain in Python.
|
||||
The Python and C++ backends follow the same mathematics for JOC object reconstruction and speaker rendering. The public SOFA binaural backend currently runs in Python; bitstream parsing, the OAMD timeline, and CLI behavior also remain in Python.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Python 3.10+
|
||||
- NumPy 1.24+
|
||||
- A standalone FFmpeg executable; `ffmpeg-python` is not required. FFmpeg is discovered through `PATH` by default or selected with `--ffmpeg`
|
||||
- h5py 3.8+
|
||||
- SciPy 1.10+
|
||||
- A standalone FFmpeg executable; `ffmpeg-python` is not required. FFmpeg is discovered through `PATH` by default or selected with `--ffmpeg`. On startup the decoder options are probed with `ffmpeg -h decoder=eac3`: a missing E-AC-3 decoder or `-drc_scale` is a hard error, while a missing `-target_level` only fails when `--eac3-target-level` is used
|
||||
- Optional: CMake and a C++20 toolchain to build the native core
|
||||
|
||||
Install the Python dependency in a project-specific environment:
|
||||
@@ -78,12 +83,58 @@ python main.py input.m4a --speaker-layout 5.1 --speaker-format int24
|
||||
python main.py input.m4a --speaker-layout 7.1.2 --speaker-output output.7.1.2.wav
|
||||
```
|
||||
|
||||
When PCM24 may clip in a non-interactive environment, select a policy explicitly:
|
||||
Write binaural stereo directly (ordinary objects are Near/Mid/Far only; Mid is
|
||||
the default). The HRTF input accepts three sources:
|
||||
|
||||
```powershell
|
||||
# 1) SOFA (defaults to HRTF/binaural.sofa, or an explicit path)
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone (defaults to HRTF/binaural.personalized_headphone)
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf compiled cache
|
||||
python main.py input.m4a --binaural --compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# Common options
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--binaural-mode near
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
python main.py input.m4a --binaural --binaural-output output.binaural.wav
|
||||
```
|
||||
|
||||
With none of the three specified, resolution tries, in order:
|
||||
`HRTF/binaural.sofa`, the unique `.jochrtf` under `output/hrtf-cache`, then
|
||||
`HRTF/binaural.personalized_headphone`; if none exist, an error asks for an
|
||||
explicit path.
|
||||
|
||||
- `.sofa` is the portable source of truth; it can hold self-scanned or any
|
||||
generic HRTF data.
|
||||
- `.personalized_headphone` is a model produced by Dolby's official
|
||||
personalization scan; its JSON parsing is implemented by this project
|
||||
(`src/rosella_model.py`) and does not invoke any Dolby software.
|
||||
- `.jochrtf` is a project-internal cache compiled from SOFA; it is disposable,
|
||||
rebuildable, and written to `output/hrtf-cache` by default.
|
||||
|
||||
HRTF data lives under `HRTF/` (git-ignored): the default SOFA
|
||||
`HRTF/binaural.sofa` and the default model
|
||||
`HRTF/binaural.personalized_headphone`. Because the cache contains transformed
|
||||
HRTF data, its use and redistribution remain subject to the source dataset's
|
||||
terms. See [Binaural Rendering](docs/binaural.en.md) and
|
||||
[Third-party notices](THIRD_PARTY_NOTICES.md) for format boundaries, formulas,
|
||||
state, timing, and distribution considerations.
|
||||
|
||||
Speaker and binaural output share peak analysis, the WAV writer, and clipping policy. When PCM24 may clip in a non-interactive environment, select a policy explicitly:
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action abort
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action float32
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action continue
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--binaural-format int24 --clip-action abort
|
||||
```
|
||||
|
||||
Metadata and diagnostics:
|
||||
@@ -95,6 +146,85 @@ python main.py input.m4a --metadata-cache metadata_cache
|
||||
python main.py input.m4a --metadata-dir metadata_cache
|
||||
```
|
||||
|
||||
### E-AC-3 decode-side dynamic range and level
|
||||
|
||||
By default FFmpeg applies the stream `dynrng` dynamic range compression when
|
||||
decoding E-AC-3 (`-drc_scale 1`). The core 5.1 PCM is the input of JOC object
|
||||
reconstruction, and `dynrng` is playback-time gain, so it is inherited linearly
|
||||
by every object and every output (ADM, speaker, binaural). This tool therefore
|
||||
decodes at **full dynamic range** by default:
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a # default: -drc_scale 0, full range
|
||||
python main.py input.m4a --eac3-drc-scale 1 # reproduce consumer playback
|
||||
python main.py input.m4a --eac3-drc-scale 0.5 # apply half of it
|
||||
python main.py input.m4a --eac3-target-level -27 # dialnorm-referenced level
|
||||
```
|
||||
|
||||
- `--eac3-drc-scale` (`0`–`6`, default `0`) maps to FFmpeg `-drc_scale`: the gain
|
||||
of each E-AC-3 block is `dynrng factor ^ value`. `0` disables DRC, `1` is the
|
||||
author's intent, and `>1` is asymmetric (loud parts fully compressed, quiet
|
||||
parts enhanced).
|
||||
- `--eac3-target-level` (`-31`–`0`, default `0` = off) maps to FFmpeg
|
||||
`-target_level`: a static per-frame gain of about `target_level - dialnorm` dB,
|
||||
independent of and stackable with `--eac3-drc-scale`. dialnorm is a per-stream
|
||||
property (measured Apple Music Atmos streams are about `-18` to `-19` dB, so
|
||||
`-27` is roughly `8`–`9` dB of attenuation).
|
||||
- The level change is expected: compared with the FFmpeg default, measured
|
||||
tracks move by `0` to `-2.15` dB peak and `0` to `-1.69` dB RMS (direction
|
||||
depends on the stream `dynrng`), so `output_clip.peak` and the PCM24 clipping
|
||||
decision in `.report.json` change accordingly.
|
||||
- `--gain-db` is a static gain applied **after** reconstruction (float64 on the
|
||||
binaural path) and is not the same thing as decode-side DRC, which is
|
||||
block-varying; do not use `--gain-db` to cancel it.
|
||||
- The `ffmpeg` field of `.report.json` records the FFmpeg version and the decode
|
||||
options that were actually passed (`version`, `eac3_decode_options`).
|
||||
|
||||
### Binaural render mode
|
||||
|
||||
`--binaural-mode off|near|mid|far` selects the binaural render mode; the default
|
||||
is `mid`, and both outputs share this single option:
|
||||
|
||||
- **Direct binaural rendering** (`--binaural`): `off` is rejected (error);
|
||||
near/mid/far apply, defaulting to `mid`;
|
||||
- **ADM BWF**: the low 3 binaural-render-mode bits of the last 15 JOC object
|
||||
entries in DBMD segment 10 carry `off=0/near=1/far=2/mid=3`, leaving the first
|
||||
10 bed entries unchanged; the default is `mid`, and `off` explicitly disables
|
||||
the binaural metadata hint.
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a --binaural-mode mid
|
||||
python main.py input.m4a --binaural-mode off # ADM BWF only: disable the DBMD hint
|
||||
```
|
||||
|
||||
**The default `mid` is a human-specified rendering hint**; it is not original
|
||||
binaural metadata extracted or recovered from the input E-AC-3 JOC bitstream,
|
||||
nor does it represent the original mix's per-object binaural settings. The hint
|
||||
does not change PCM, object trajectories, or direct speaker rendering. The
|
||||
adjacent `.report.json` records `binaural_mode` (the mode name) and
|
||||
`binaural_mode_value` (the ADM code; `null` for direct binaural output).
|
||||
|
||||
### OAMD time alignment
|
||||
|
||||
Object trajectories and direct speaker rendering both default to a metadata delay of `1473 samples`. This value describes the theoretical mapping between decoder-output PCM and OAMD updates. The speaker renderer retains its existing 32-sample control block, so the default update lands on effective block boundary `1472`:
|
||||
|
||||
```text
|
||||
align32(1473) = 1472
|
||||
```
|
||||
|
||||
Override the two paths with `--object-delay-samples` and `--speaker-metadata-offset`, respectively. The 1473-sample timing offset is distinct from the 640-value inverse-QMF filter/window state; 640 is a QMF state length, not a metadata delay.
|
||||
|
||||
The direct binaural path uses `--object-delay-samples`. Each ID11/OAMD event is
|
||||
placed on an absolute sample timeline from its frame start, outer-subpayload
|
||||
offset, and block offset, then shifted by that delay. Each 1536-sample input
|
||||
frame is processed as three consecutive 512-sample blocks; the interpolated
|
||||
position, direction, and profile are updated at each block's absolute starting
|
||||
sample.
|
||||
|
||||
### Binaural calculation
|
||||
|
||||
See [Binaural Rendering Mathematics](docs/binaural.en.md) for QMF, hybrid processing, direction fields, distance, ITD, room processing, the 512-sample parameter updates above, and 961-sample latency compensation.
|
||||
|
||||
For all options:
|
||||
|
||||
```powershell
|
||||
@@ -128,8 +258,10 @@ JustOneCacophony/
|
||||
├─ main.py command-line entry point
|
||||
├─ src/ Python implementation modules
|
||||
├─ native/ C/C++ acceleration core, C ABI, and required table data
|
||||
├─ data/ runtime table data for Python
|
||||
├─ data/ Python runtime table data
|
||||
├─ lib/ native runtime drop-in directory (create as needed)
|
||||
├─ HRTF/ user HRTF data directory (create as needed, git-ignored)
|
||||
├─ output/ output directory (create as needed; the .jochrtf cache defaults to its hrtf-cache subdirectory)
|
||||
├─ docs/ math and native-core notes in both languages
|
||||
├─ requirements.txt Python dependency
|
||||
├─ README.md Chinese documentation
|
||||
@@ -148,20 +280,24 @@ The main documented stages are:
|
||||
- OAMD Q15 coordinate conversion;
|
||||
- equal-power panning over target-layout regions;
|
||||
- layout-dependent position compensation and sample-wise gain ramps;
|
||||
- float32 and PCM24 output quantization.
|
||||
- float32 and PCM24 output quantization;
|
||||
- SOFA canonical import, 64-QMF/77-hybrid projection, `36×2×77` fifth-order fields, exactly-once delay/phase, project early/late room behavior, and special LFE.
|
||||
|
||||
See the [mathematical notes](docs/math.en.md) for the equations used by the decoding and rendering process.
|
||||
|
||||
## Known limitations
|
||||
|
||||
- Only the common contiguous EMDF transport is covered. Fragmented transport across multiple audio-block skip fields is not covered.
|
||||
- Dense JOC is the main path. The Sparse JOC branch should not be treated as supported.
|
||||
- The speaker path currently covers ordinary point objects; extent, spread, divergence, and similar modes are outside the supported scope.
|
||||
- The speaker and SOFA binaural paths currently cover ordinary point objects; extent, spread, diffuse, divergence, channel lock, and similar controls are outside the supported scope.
|
||||
- OAMD trim elements are boundary-checked and skipped; warp, balance, and trim parameters are not applied to raw object trajectories or speaker rendering.
|
||||
- Multi-data-point streams, uncommon band configurations, and unusual OAMD scheduling have less coverage than common 12-band, single-data-point material.
|
||||
- A speaker limiter is outside the current primary formula.
|
||||
- ADM output, native binaries, and speaker layouts still need broader interoperability checks across platforms, players, and real material.
|
||||
- The SOFA importer currently supports the strict `SimpleFreeFieldHRIR` FIR subset; other SOFA conventions require explicit adapters.
|
||||
- The binaural runtime is fixed at 48 kHz, fifth order, and one measurement-radius shell at a time; the public binaural backend defaults to the native accelerator and falls back to Python when the native library is unavailable.
|
||||
- ADM output, native binaries, speaker layouts, and binaural models still need broader interoperability checks across platforms, players, and real material.
|
||||
|
||||
## Documentation
|
||||
|
||||
- [Mathematical notes](docs/math.en.md) · [中文](docs/math.md)
|
||||
- [Native-core notes](docs/native.en.md) · [中文](docs/native.md)
|
||||
- [Binaural rendering](docs/binaural.en.md) · [中文](docs/binaural.md)
|
||||
|
||||
@@ -4,9 +4,9 @@
|
||||
|
||||
> JustOneCacophony 是一个 E-AC-3 JOC 的实验性 / 测试实现,用于研究 JOC 的解析、重建、渲染以及相关数学过程。
|
||||
|
||||
项目可以从常见 E-AC-3 JOC 码流中提取并解析 EMDF、ID14 JOC 参数和 ID11 OAMD 元数据,结合 FFmpeg 解码出的核心 5.1 PCM 重建 LFE 与 15 路对象 PCM,并输出 ADM BWF 或指定扬声器布局的 WAV。
|
||||
项目可以从常见 E-AC-3 JOC 码流中提取并解析 EMDF、ID14 JOC 参数和 ID11 OAMD 元数据,结合 FFmpeg 解码出的核心 5.1 PCM 重建 LFE 与 15 路对象 PCM,并输出 ADM BWF、指定扬声器布局的 WAV,或使用标准 SOFA HRTF 直接输出双耳 WAV。
|
||||
|
||||
这是研究代码,不是完整、标准兼容或生产级的 Dolby JOC 解码器。它只覆盖当前已实现的码流形态;遇到未知变体时会明确报错,而不是假装一切都很和谐——如果哪里算错了,它可能就真的只剩 cacophony 了。
|
||||
这是研究代码,不是完整、标准兼容或生产级的 JOC 解码器。它只覆盖当前已实现的码流形态;遇到未知变体时会明确报错,而不是假装一切都很和谐——如果哪里算错了,它可能就真的只剩 cacophony 了。
|
||||
|
||||
## 当前功能
|
||||
|
||||
@@ -16,7 +16,9 @@
|
||||
- 通过 analysis QMF、参数插值、对象矩阵和 inverse QMF 重建 LFE + 15 路对象 PCM;
|
||||
- 输出 25 声道 ADM BWF:10 声道 7.1.2 bed(除 LFE 外静音)+ 15 个对象;
|
||||
- 直接渲染 `2.0`、`3.1`、`5.1`、`7.1`、`5.1.2`、`5.1.4`、`7.1.2`、`7.1.4`、`9.1.4`、`9.1.6`;
|
||||
- 输出 float32 或 PCM24 WAV,并在 PCM24 削波前提供明确处理策略;
|
||||
- 从 `pcm16 + ID11/OAMD` 直接运行公开 SOFA 双耳渲染,不生成临时 ADM BWF;
|
||||
- 双耳 DSP 全程使用 float64/complex128,并保留 961-sample latency compensation、跨帧状态和 room 尾声;
|
||||
- 直接输出统一支持 float32 或 PCM24 WAV,并在 PCM24 削波前提供明确处理策略;
|
||||
- 使用 NumPy 后端,或通过 `ctypes` 调用可选的 C++20 原生核;`auto` 模式在原生库不可用时回退到 Python;
|
||||
- 读取或写入 metadata sidecar,并生成元数据、运行时间和输出摘要。
|
||||
|
||||
@@ -30,16 +32,21 @@ M4A / E-AC-3
|
||||
├─ ID11 OAMD → 对象位置与时间轨迹
|
||||
└─ LFE + 15 objects
|
||||
├─ 25ch ADM BWF
|
||||
└─ 指定布局的扬声器 WAV
|
||||
├─ 指定布局的扬声器 WAV
|
||||
└─ ID11 直接时间轴 + SOFA HRTF → 双耳 WAV
|
||||
```
|
||||
|
||||
Python 与 C++ 后端使用同一组已记录的数学过程。原生核只处理状态密集的 DSP 和扬声器渲染,高层位流解析、ADM 组装与命令行逻辑仍在 Python 中。
|
||||
Python 与 C++ 后端在 JOC 对象重建、扬声器渲染和公开 SOFA 双耳渲染中使用同一组
|
||||
数学过程;native 双耳后端与 Python 参考实现逐值一致(差异 < 1e-9)。位流解析、
|
||||
OAMD 时间轴和命令行逻辑在 Python 中。
|
||||
|
||||
## 环境
|
||||
|
||||
- Python 3.10+
|
||||
- NumPy 1.24+
|
||||
- 独立的 FFmpeg 可执行程序;不需要 `ffmpeg-python`。默认从 `PATH` 查找,也可通过 `--ffmpeg` 指定可执行文件路径
|
||||
- h5py 3.8+
|
||||
- SciPy 1.10+
|
||||
- 独立的 FFmpeg 可执行程序;不需要 `ffmpeg-python`。默认从 `PATH` 查找,也可通过 `--ffmpeg` 指定可执行文件路径。启动时会探测 `ffmpeg -h decoder=eac3`:缺 E-AC-3 解码器或 `-drc_scale` 直接报错,缺 `-target_level` 只在使用 `--eac3-target-level` 时报错
|
||||
- 可选:支持 C++20 的 CMake 工具链,用于自行构建原生核
|
||||
|
||||
建议在项目专用虚拟环境中安装依赖:
|
||||
@@ -78,12 +85,50 @@ python main.py input.m4a --speaker-layout 5.1 --speaker-format int24
|
||||
python main.py input.m4a --speaker-layout 7.1.2 --speaker-output output.7.1.2.wav
|
||||
```
|
||||
|
||||
在非交互环境请求 PCM24 且可能削波时,需要显式选择处理方式:
|
||||
直接输出双耳渲染 WAV(普通对象仅 Near/Mid/Far,默认 Mid)。HRTF 输入支持三种来源:
|
||||
|
||||
```powershell
|
||||
# 1) SOFA(缺省取 HRTF/binaural.sofa,也可显式指定)
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone(缺省取 HRTF/binaural.personalized_headphone)
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf 编译缓存
|
||||
python main.py input.m4a --binaural --compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# 常用选项
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--binaural-mode near
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
python main.py input.m4a --binaural --binaural-output output.binaural.wav
|
||||
```
|
||||
|
||||
三者都不指定时的自动选择顺序:`HRTF/binaural.sofa` → `output/hrtf-cache` 下唯一的
|
||||
`.jochrtf` → `HRTF/binaural.personalized_headphone`;都没有则报错并提示显式指定。
|
||||
|
||||
- `.sofa` 是可移植的 source of truth;可以是自行扫描或任何来源的通用 HRTF 数据。
|
||||
- `.personalized_headphone` 是杜比官方软件个性化扫描得到的模型,其 JSON 解析由
|
||||
本项目自行实现(`src/rosella_model.py`),不调用杜比软件。
|
||||
- `.jochrtf` 是从 SOFA 编译出的项目内部 cache,可删除、可从 SOFA 重建,默认写在
|
||||
`output/hrtf-cache`。
|
||||
|
||||
HRTF 数据统一放在 `HRTF/`(git 忽略):默认 SOFA `HRTF/binaural.sofa`、默认模型
|
||||
`HRTF/binaural.personalized_headphone`。cache 含有源 HRTF 的变换数据,使用与再分发
|
||||
仍受源数据许可约束;格式边界、计算公式、状态、时间轴及发布注意事项见
|
||||
[双耳渲染](docs/binaural.md) 和 [第三方通知](THIRD_PARTY_NOTICES.md)。
|
||||
|
||||
扬声器和双耳输出共享峰值检查、writer 与削波策略。在非交互环境请求 PCM24 且可能削波时,需要显式选择处理方式:
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action abort
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action float32
|
||||
python main.py input.m4a --speaker-layout 5.1 --speaker-format int24 --clip-action continue
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--binaural-format int24 --clip-action abort
|
||||
```
|
||||
|
||||
元数据与诊断:
|
||||
@@ -95,6 +140,61 @@ python main.py input.m4a --metadata-cache metadata_cache
|
||||
python main.py input.m4a --metadata-dir metadata_cache
|
||||
```
|
||||
|
||||
### E-AC-3 解码级动态范围与电平
|
||||
|
||||
FFmpeg 解码 E-AC-3 时默认施加码流 `dynrng` 动态范围压缩(`-drc_scale 1`)。核心 5.1 PCM 是 JOC 对象重建的输入,而 `dynrng` 属于回放期增益,会被线性继承到全部对象与成品(ADM/扬声器/双耳),因此本工具默认按**全动态范围**解码:
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a # 默认:-drc_scale 0,全动态范围
|
||||
python main.py input.m4a --eac3-drc-scale 1 # 复现消费者回放(码流作者意图)
|
||||
python main.py input.m4a --eac3-drc-scale 0.5 # 施加一半
|
||||
python main.py input.m4a --eac3-target-level -27 # 按码流 dialnorm 归一化电平
|
||||
```
|
||||
|
||||
- `--eac3-drc-scale`(`0`~`6`,默认 `0`)对应 FFmpeg 的 `-drc_scale`:每个 E-AC-3 block 的增益为 `dynrng 因子 ^ 该值`。`0` 关闭 DRC;`1` 为码流作者意图;`>1` 非对称(响处全压、轻处增强)。
|
||||
- `--eac3-target-level`(`-31`~`0`,默认 `0` 不施加)对应 FFmpeg 的 `-target_level`:按每帧 dialnorm 施加静态增益,约 `target_level - dialnorm` dB,与 `--eac3-drc-scale` 相互独立、可叠加。dialnorm 是逐码流属性(实测 Apple Music Atmos 流约 `-18`~`-19` dB,故 `-27` 约等于衰减 `8`~`9` dB)。
|
||||
- 电平变化是预期的:与 FFmpeg 默认值相比,实测曲目峰值变化 `0`~`-2.15` dB、RMS `0`~`-1.69` dB(方向取决于码流 `dynrng`),`.report.json` 的 `output_clip.peak` 与 int24 削波判定会随之变化。
|
||||
- `--gain-db` 是**重建之后**的静态增益(双耳路径 float64),与解码级 DRC 不是一回事;解码级 DRC 是按 block 时变的,不要用 `--gain-db` 去抵消它。
|
||||
- `.report.json` 的 `ffmpeg` 字段记录 FFmpeg 版本与实际下发的解码选项(`version`、`eac3_decode_options`)。
|
||||
|
||||
### 双耳渲染模式
|
||||
|
||||
`--binaural-mode off|near|mid|far` 选择双耳渲染模式,默认 `mid`,两种输出共用这一个选项:
|
||||
|
||||
- **直接双耳渲染**(`--binaural`):`off` 不可用(报错),near/mid/far 生效,默认 `mid`;
|
||||
- **ADM BWF**:DBMD segment 10 中后 15 个 JOC 对象的 binaural render mode 写
|
||||
`off=0/near=1/far=2/mid=3`,前 10 个 bed 保持不变,默认 `mid`;`off` 用于显式
|
||||
关闭双耳元数据提示。
|
||||
|
||||
```powershell
|
||||
python main.py input.m4a --binaural-mode mid
|
||||
python main.py input.m4a --binaural-mode off # 仅 ADM BWF:关闭 DBMD 双耳提示
|
||||
```
|
||||
|
||||
**默认 `mid` 是本工具人为指定的渲染提示**,不是从输入 E-AC-3 JOC 码流中提取或
|
||||
还原的原始双耳元数据,也不代表原始混音中各对象的双耳设置。该提示不改变 PCM、
|
||||
对象轨迹或直接扬声器渲染。输出旁的 `.report.json` 用 `binaural_mode`(模式名)
|
||||
和 `binaural_mode_value`(ADM 编码值,直接双耳输出时为 `null`)记录。
|
||||
|
||||
### OAMD 时间对齐
|
||||
|
||||
对象轨迹和直接扬声器渲染的 metadata delay 默认均为 `1473 samples`。该值描述 decoder 输出 PCM 与 OAMD 更新之间的理论时间映射;扬声器 renderer 仍使用现有的 32-sample control block,因此默认更新的实际 block boundary 为 `1472`:
|
||||
|
||||
```text
|
||||
align32(1473) = 1472
|
||||
```
|
||||
|
||||
可分别用 `--object-delay-samples` 和 `--speaker-metadata-offset` 覆盖默认值。这里的 1473 不应与 inverse-QMF 的 640 项 filter/window state 混淆;后者是 QMF 状态长度,不是 metadata delay。
|
||||
|
||||
直接双耳路径使用 `--object-delay-samples`。每个 ID11/OAMD event 先按 frame start、
|
||||
outer subpayload offset 与 block offset 落到绝对 sample timeline,再加该 delay;每个
|
||||
1536-sample 输入帧按三个连续 512-sample block 处理,并在每块的绝对起始 sample
|
||||
查询插值后的位置、更新方向和 profile。
|
||||
|
||||
### 双耳计算
|
||||
|
||||
双耳路径的 QMF、hybrid、方向场、距离、ITD、room、上述 512-sample 参数更新和 961-sample 延迟补偿见[双耳渲染数学](docs/binaural.md)。
|
||||
|
||||
更多参数可查看:
|
||||
|
||||
```powershell
|
||||
@@ -130,6 +230,8 @@ JustOneCacophony/
|
||||
├─ native/ C/C++ 加速核、C ABI 与必要表数据
|
||||
├─ data/ Python 运行时表数据
|
||||
├─ lib/ 原生运行库投放目录(按需创建)
|
||||
├─ HRTF/ 用户 HRTF 数据目录(按需创建,git 忽略)
|
||||
├─ output/ 输出目录(按需创建;.jochrtf 缓存默认在其 hrtf-cache 子目录)
|
||||
├─ docs/ 数学与原生核文档(中英文)
|
||||
├─ requirements.txt Python 依赖
|
||||
├─ README.md 中文说明
|
||||
@@ -148,20 +250,24 @@ JustOneCacophony/
|
||||
- OAMD Q15 坐标转换;
|
||||
- 基于目标布局 region 的等功率声像;
|
||||
- 布局位置补偿与逐样本增益斜坡;
|
||||
- float32 与 PCM24 输出量化。
|
||||
- float32 与 PCM24 输出量化;
|
||||
- SOFA canonical importer、64-QMF/77-hybrid 投影、`36×2×77` 五阶方向 field、exactly-once delay/phase、项目 early/late room 与 special LFE。
|
||||
|
||||
解码与渲染过程使用的公式见[数学说明](docs/math.md)。
|
||||
|
||||
## 已知限制
|
||||
|
||||
- 当前只覆盖常见 continuous EMDF transport;跨多个 audio-block skip field 的碎片化 transport 尚未覆盖。
|
||||
- Dense JOC 是当前主要路径;Sparse JOC 分支不应视为受支持能力。
|
||||
- 扬声器路径当前只覆盖普通点对象;extent、spread、divergence 等对象模式不在支持范围内。
|
||||
- 扬声器与 SOFA 双耳路径当前只覆盖普通点对象;extent、spread、diffuse、divergence、channel lock 等对象控制不在支持范围内。
|
||||
- OAMD trim element 会按声明边界校验并跳过;warp、balance 和 trim 参数不应用于当前原始对象轨迹或扬声器渲染。
|
||||
- 多数据点、少见参数带配置和特殊 OAMD 调度的覆盖度低于常见 12-band、单数据点素材。
|
||||
- 扬声器 limiter 不属于当前实现的主公式。
|
||||
- ADM 输出、原生库和扬声器布局仍需在更多平台、播放器与真实素材上确认互操作性。
|
||||
- SOFA importer 当前严格支持 `SimpleFreeFieldHRIR` FIR;其它 SOFA convention 需要显式 adapter。
|
||||
- 双耳 runtime 固定 48 kHz、五阶和一次选择一个 measurement-radius shell;公开双耳默认走 native 加速,原生库不可用时自动回退 Python。
|
||||
- ADM 输出、原生库、扬声器布局和双耳模型仍需在更多平台、播放器与真实素材上确认互操作性。
|
||||
|
||||
## 文档
|
||||
|
||||
- [数学说明](docs/math.md) · [English](docs/math.en.md)
|
||||
- [原生核说明](docs/native.md) · [English](docs/native.en.md)
|
||||
- [双耳渲染](docs/binaural.md) · [English](docs/binaural.en.md)
|
||||
|
||||
@@ -0,0 +1,41 @@
|
||||
# Third-party notices / 第三方通知
|
||||
|
||||
本文件记录 `data/rosella_kernels.npz`(`src/public_filterbank.py` 使用的滤波器组表)
|
||||
的公开标准来源,以及 HRTF 数据与专利的边界说明。
|
||||
|
||||
## 公开标准来源
|
||||
|
||||
64-QMF → 77-hybrid 结构与 13-tap 低带 prototype 定义于
|
||||
[3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
第 5.2.2 节(Table 1 的 $Q=8$/
|
||||
$Q=4$ 系数,delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr)$$
|
||||
|
||||
64-band QMF analysis 即 ISO/IEC 14496-3/AMD1:2003 第 4.B.18.2 节的 MPEG-4
|
||||
AAC/SBR 64 complex QMF bank;打包的 $64\times10$ 表是公开 640-tap prototype 的
|
||||
多相重排:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t}$$
|
||||
|
||||
QMF synthesis 表为 analysis 多相矩阵 $\mathbf{A}$ 的因果左逆
|
||||
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$(
|
||||
$\mathbf{P}$ 为 577-sample 延迟置换;
|
||||
全链 $961 = 577 + 6\times64$),rank-4 分解存储:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
hybrid synthesis 表为 77→64 重组:高频带恒等 $Y_{3+b}=X_{16+b}$,低频带:
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\mathrm{Re}X_q + j\,s_q\,\mathrm{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
相同数值可在 FFmpeg(`aacps_tablegen.h`、`aacsbrdata.h`)等公开实现中查到。
|
||||
|
||||
## HRTF 数据与 `.jochrtf`
|
||||
|
||||
`.jochrtf` 含有特定源 SOFA/HRTF 数据集的变换系数与 delay;其使用、复制与再分发
|
||||
仍受源数据集许可约束,权限不明确时应作为私有 cache 保存。
|
||||
|
||||
## 专利说明
|
||||
|
||||
标准可公开获取不等于获准实施相关专利。
|
||||
+64
-1
@@ -2,7 +2,9 @@
|
||||
|
||||
[中文](README.md)
|
||||
|
||||
`tables.npz` contains the static table data used by the Python path:
|
||||
This directory contains static production tables. It does not contain user HRTFs.
|
||||
|
||||
`tables.npz` contains the JOC core decoding tables:
|
||||
|
||||
```text
|
||||
analysis_window float64[10,64]
|
||||
@@ -18,3 +20,64 @@ joc_huff_code_7ch_pos_index_sparse int64[6,2]
|
||||
`src/joc_qmf.py` loads the QMF tables, while `src/joc_decode.py` loads the JOC Huffman trees. Python does not read C/C++ headers under `native/`.
|
||||
|
||||
The corresponding native data are stored in `native/src/qmf_tables.h` and `native/src/joc_huffman_tables.h`. Changes on either side should update the other and be checked for value-by-value agreement.
|
||||
|
||||
## Binaural rendering tables
|
||||
|
||||
`rosella_kernels.npz` contains the fixed 64-QMF/77-hybrid tables used by the
|
||||
public SOFA binaural path:
|
||||
|
||||
```text
|
||||
format_version little-endian int32[1]
|
||||
qmf_analysis_coefficients float32[64,10]
|
||||
hybrid_analysis_low_kernel float32[3,2,13,16,2]
|
||||
hybrid_synthesis_indices int16[154,4]
|
||||
hybrid_synthesis_values float32[154]
|
||||
qmf_synthesis_basis float64[64,4,128]
|
||||
qmf_synthesis_taps float64[64,10,4]
|
||||
```
|
||||
|
||||
The float32 table values are promoted to float64 when loaded.
|
||||
`src/public_filterbank.py` verifies the archive and every array by SHA-256.
|
||||
Those hashes, the table version, and the 77 reference band-center values all
|
||||
participate in the `.jochrtf` cache key. The full analysis/synthesis latency is
|
||||
961 samples.
|
||||
|
||||
The packaged tables implement publicly standardized filter banks, computable
|
||||
from the following formulas.
|
||||
|
||||
The 64-QMF → 77-hybrid structure, the 13-tap low-band prototypes, and their
|
||||
half-bin complex modulation are defined in
|
||||
[3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf),
|
||||
Section 5.2.2 (Table 1 $Q=8$/$Q=4$ coefficients, delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\!\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
The 64-band QMF analysis is the MPEG-4 AAC/SBR 64 complex QMF analysis bank of
|
||||
ISO/IEC 14496-3/AMD1:2003, subclause 4.B.18.2; the packaged $64\times10$ table
|
||||
is the polyphase reordering of the public 640-tap prototype $c_0,\dots,c_{639}$:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
The QMF synthesis table is the causal left inverse of the analysis polyphase
|
||||
matrix $\mathbf{A}$, i.e. the solution of $\mathbf{A}\,\mathbf{W}=\mathbf{P}$
|
||||
($\mathbf{P}$ is the 577-sample delay permutation; total latency
|
||||
$961 = 577 + 6\times64$), stored as a rank-4 factorization:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
The hybrid synthesis table is the 77→64 recombination: identity for the high
|
||||
bands, $Y_{3+b}=X_{16+b}$, and for the low bands ($C_p$ is the $8+4+4$ child
|
||||
partition):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\operatorname{Re}X_q + j\,s_q\,\operatorname{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
The same values also appear in other public implementations of these standards
|
||||
(for example FFmpeg's `aacps_tablegen.h` and `aacsbrdata.h`).
|
||||
|
||||
Public availability of a standard does not by itself grant permission to
|
||||
practice related patent claims.
|
||||
|
||||
SOFA is the user-visible source of truth. A `.jochrtf` file is a disposable JOC
|
||||
compiled HRTF cache that can be rebuilt from SOFA. The cache contains
|
||||
transformed source-HRTF data and remains subject to the source SOFA/HRTF
|
||||
dataset's licence and redistribution restrictions.
|
||||
|
||||
+54
-1
@@ -2,7 +2,9 @@
|
||||
|
||||
[English](README.en.md)
|
||||
|
||||
`tables.npz` 集中保存 Python 路径使用的静态表数据:
|
||||
本目录保存 Python 生产路径使用的静态表数据,不保存用户 HRTF。
|
||||
|
||||
`tables.npz` 保存 JOC 核心解码表:
|
||||
|
||||
```text
|
||||
analysis_window float64[10,64]
|
||||
@@ -18,3 +20,54 @@ joc_huff_code_7ch_pos_index_sparse int64[6,2]
|
||||
`src/joc_qmf.py` 读取 QMF 表,`src/joc_decode.py` 读取 JOC Huffman 树。Python 不读取 `native/` 下的 C/C++ 头文件。
|
||||
|
||||
原生侧对应数据分别位于 `native/src/qmf_tables.h` 与 `native/src/joc_huffman_tables.h`。修改任何一侧时,应同步更新另一侧并进行逐值一致性检查。
|
||||
|
||||
## 双耳渲染表
|
||||
|
||||
`rosella_kernels.npz` 保存公开 SOFA 双耳路径使用的 64-QMF/77-hybrid 固定表:
|
||||
|
||||
```text
|
||||
format_version little-endian int32[1]
|
||||
qmf_analysis_coefficients float32[64,10]
|
||||
hybrid_analysis_low_kernel float32[3,2,13,16,2]
|
||||
hybrid_synthesis_indices int16[154,4]
|
||||
hybrid_synthesis_values float32[154]
|
||||
qmf_synthesis_basis float64[64,4,128]
|
||||
qmf_synthesis_taps float64[64,10,4]
|
||||
```
|
||||
|
||||
float32 表值载入后提升为 float64。`src/public_filterbank.py` 在读取时校验 archive
|
||||
及每个数组的 SHA-256;这些 hash、table version 和 77 个 band-center 参考值共同进入
|
||||
`.jochrtf` cache key。analysis/synthesis 全链 latency 为 961 samples。
|
||||
|
||||
打包表实现的是公开标准化的滤波器组,各表可由如下公式计算。
|
||||
|
||||
64-QMF → 77-hybrid 结构、13-tap 低带 prototype 与半 bin 复调制定义于
|
||||
[3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
第 5.2.2 节(Table 1 的 $Q=8$/$Q=4$ 系数,delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\!\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
64-band QMF analysis 即 ISO/IEC 14496-3/AMD1:2003 第 4.B.18.2 节的 MPEG-4
|
||||
AAC/SBR 64 complex QMF bank;打包的 $64\times10$ 表是公开 640-tap prototype
|
||||
$c_0,\dots,c_{639}$ 的多相重排:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
QMF synthesis 表为上述 analysis 多相矩阵 $\mathbf{A}$ 的因果左逆,即求解
|
||||
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$($\mathbf{P}$ 为 577-sample 延迟置换;
|
||||
全链 $961 = 577 + 6\times64$),以 rank-4 分解形式存储:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
hybrid synthesis 表为 77→64 重组:高频带恒等 $Y_{3+b}=X_{16+b}$;低频带
|
||||
($C_p$ 为 $8+4+4$ 子带划分):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\operatorname{Re}X_q + j\,s_q\,\operatorname{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
相同数值可在 FFmpeg(`aacps_tablegen.h`、`aacsbrdata.h`)等公开实现中查到。
|
||||
|
||||
标准可公开获取不等于获准实施相关专利。
|
||||
|
||||
`.sofa` 是用户可见的 source of truth;`.jochrtf` 是可删除、可从 SOFA 重建的
|
||||
JOC compiled HRTF cache。cache 含有源 HRTF 的变换数据,仍受源 SOFA/HRTF
|
||||
数据集的许可与再分发限制约束。
|
||||
|
||||
Binary file not shown.
@@ -0,0 +1,284 @@
|
||||
# Binaural rendering
|
||||
|
||||
[中文](binaural.md) · [Back to README](../README.en.md)
|
||||
|
||||
JustOneCacophony's binaural backend supports three HRTF sources:
|
||||
`SimpleFreeFieldHRIR` SOFA, the Rosella `.personalized_headphone` model exported
|
||||
by Dolby's official personalization scan (its JSON parsing is implemented by
|
||||
this project and invokes no Dolby software), and the `.jochrtf` cache compiled
|
||||
from SOFA. SOFA is compiled into an in-memory directional field when the model
|
||||
is loaded. A `.jochrtf` file is only a disposable, reproducible JOC compiled
|
||||
HRTF cache; it is neither an interchange format nor a prerequisite for using
|
||||
SOFA.
|
||||
|
||||
```text
|
||||
SOFA FIR
|
||||
-> CanonicalHrtf
|
||||
-> 48 kHz / one radius shell / delay-phase policy
|
||||
-> 64-QMF / 77-hybrid projection
|
||||
-> fifth-order ACN/N3D real-SH field
|
||||
-> per-object direct + early reflections
|
||||
-> shared unitary-FDN late room
|
||||
-> float64 stereo
|
||||
```
|
||||
|
||||
## Inputs
|
||||
|
||||
The CLI has three mutually exclusive HRTF input sources; with none given, a
|
||||
default rule resolves the input:
|
||||
|
||||
```powershell
|
||||
# 1) SOFA: defaults to HRTF/binaural.sofa, or an explicit path
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone: defaults to HRTF/binaural.personalized_headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf: explicitly load a compiled cache
|
||||
python main.py input.m4a --binaural `
|
||||
--compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# Optional: create/reuse a transparent disk cache for SOFA
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
```
|
||||
|
||||
The default order is `HRTF/binaural.sofa`, then the unique `.jochrtf` under
|
||||
`output/hrtf-cache`, then `HRTF/binaural.personalized_headphone`; if none of
|
||||
the three exist, an error asks for an explicit path. Multiple `.jochrtf` files
|
||||
under `output/hrtf-cache` are also an error requiring an explicit choice.
|
||||
|
||||
The `.personalized_headphone` JSON parsing is implemented by this project
|
||||
(`src/rosella_model.py`) and does not invoke any Dolby software.
|
||||
|
||||
`--hrtf-cache-policy` accepts `none`, `memory`, or `disk`. The default is
|
||||
`memory`; neither `none` nor `memory` creates a file. `disk` writes to
|
||||
`output/hrtf-cache` by default, or to `--hrtf-cache-dir`. `--hrtf-radius-m`
|
||||
selects the nearest measurement-radius shell.
|
||||
|
||||
The Python API also uses explicit factories:
|
||||
|
||||
```python
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
|
||||
renderer = SofaBinauralBackend.from_sofa(
|
||||
"subject.sofa",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
cache_policy="memory",
|
||||
)
|
||||
|
||||
cached = SofaBinauralBackend.from_compiled_cache(
|
||||
"subject.jochrtf",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
)
|
||||
```
|
||||
|
||||
The factories never guess a format from an unknown suffix: SOFA and `.jochrtf`
|
||||
always use distinct loaders.
|
||||
|
||||
## Binaural render mode
|
||||
|
||||
`--binaural-mode off|near|mid|far` (default `mid`) is a **human-specified
|
||||
rendering hint**, not original binaural metadata extracted or recovered from the
|
||||
input E-AC-3 JOC bitstream:
|
||||
|
||||
- Direct binaural rendering (`--binaural`): near/mid/far apply, default `mid`;
|
||||
`off` is an error;
|
||||
- ADM BWF: the low 3 binaural-render-mode bits of the last 15 JOC object entries
|
||||
in DBMD segment 10 carry `off=0/near=1/far=2/mid=3`, leaving the first 10 bed
|
||||
entries unchanged; the default is `mid`, and `off` explicitly disables the
|
||||
binaural metadata hint.
|
||||
|
||||
## Canonical SOFA contract
|
||||
|
||||
The strict importer currently accepts:
|
||||
|
||||
- `Conventions=SOFA`;
|
||||
- `SOFAConventions=SimpleFreeFieldHRIR`, version `0.4`, `1.0`, or `1.1`;
|
||||
- `DataType=FIR` and `Data.IR[M,2,N]`;
|
||||
- one positive finite `Data.SamplingRate` in hertz/Hz;
|
||||
- spherical or Cartesian `SourcePosition`;
|
||||
- singleton or per-measurement `ListenerPosition/View/Up`;
|
||||
- two receivers whose listener-local lateral geometry uniquely identifies L/R;
|
||||
- one zero-offset emitter;
|
||||
- causal `Data.Delay[I,2]` or `[M,2]`;
|
||||
- an explicitly free-field/anechoic `RoomType`.
|
||||
|
||||
Receiver order comes from geometry, never from the receiver array index. SOFA
|
||||
listener coordinates are $+X$ front,
|
||||
$+Y$ left,
|
||||
$+Z$ up; ADM coordinates are
|
||||
$+X$ right,
|
||||
$+Y$ front,
|
||||
$+Z$ up:
|
||||
|
||||
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
|
||||
|
||||
`CanonicalHrtf` keeps `Data.IR` and `Data.Delay` separate. Only a time-domain
|
||||
baseline calls `materialized_measurement()` to apply delay once; the runtime SH
|
||||
path never materializes and then restores the delay. Non-48-kHz HRIRs are
|
||||
normalized with float64 `scipy.signal.resample_poly`, and delay samples scale by
|
||||
the same ratio.
|
||||
|
||||
GeneralFIR, BRIR, TF, multiple emitters, ambiguous receivers, and non-free-field
|
||||
data require convention-specific adapters. They cannot enter the core importer
|
||||
through a reshape.
|
||||
|
||||
## Exactly-once delay and phase
|
||||
|
||||
The compiler recognizes three mutually exclusive representations:
|
||||
|
||||
1. Nonzero `Data.Delay` is external to `Data.IR`; the FIR is not de-rotated and
|
||||
runtime applies the delay once.
|
||||
2. With `Data.Delay=0` and an ordinary positive-onset HRIR, each ear's main peak
|
||||
supplies arrival time. Compilation separates it and runtime restores it once.
|
||||
The current threshold is a peak index greater than two samples.
|
||||
3. With `Data.Delay=0` and both FIRs at a shared sample-zero origin, no external
|
||||
delay is invented. The authored complex phase stays in the fifth-order field.
|
||||
|
||||
No path may add a second ear delay or phase-group delay.
|
||||
|
||||
## Public filterbank and directional field
|
||||
|
||||
The runtime is fixed at:
|
||||
|
||||
- 48 kHz;
|
||||
- a 64-sample QMF hop;
|
||||
- 64-QMF / 77 hybrid bands;
|
||||
- 961 samples of analysis/synthesis latency;
|
||||
- fifth order, 36 terms, ACN/N3D real spherical harmonics;
|
||||
- float64 PCM, delay, SH, and room state; complex128 band transfers and spectra.
|
||||
|
||||
Real and imaginary unit gains for every hybrid band pass through the same
|
||||
analysis/synthesis chain to form a 154-real-parameter impulse dictionary. The
|
||||
compiler does not sample 77 FFT bins. Defaults are `1e-3` projection ridge and
|
||||
`1e-5` SH ridge. Coincident directions are merged before a spherical-Voronoi
|
||||
weighted ridge fit.
|
||||
|
||||
The fixed resource is `data/rosella_kernels.npz`, which implements publicly
|
||||
standardized filter banks, computable from the following formulas.
|
||||
|
||||
The hybrid analysis kernels are defined in [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf),
|
||||
Section 5.2.2 (Table 1 $Q=8$/
|
||||
$Q=4$ coefficients, delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
The QMF analysis table is the MPEG-4 AAC/SBR 64 complex QMF bank of
|
||||
ISO/IEC 14496-3/AMD1:2003, subclause 4.B.18.2, stored as the polyphase
|
||||
reordering of the public 640-tap prototype $c_0,\dots,c_{639}$:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
The QMF synthesis table is the causal left inverse of the analysis polyphase
|
||||
matrix $\mathbf{A}$, i.e. the solution of
|
||||
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$
|
||||
($\mathbf{P}$ is the 577-sample delay permutation; total latency
|
||||
$961 = 577 + 6\times64$), stored as a rank-4 factorization:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
The hybrid synthesis table is the 77→64 recombination: identity for the high
|
||||
bands, $Y_{3+b}=X_{16+b}$, and for the low bands(
|
||||
$C_p$ is the
|
||||
$8+4+4$ child partition):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\mathrm{Re}X_q + j\,s_q\,\mathrm{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
The loader verifies the archive and every array by SHA-256; the table version,
|
||||
all array hashes, and the 77 reference band-center values are part of the cache
|
||||
key. Public availability of a standard does not by itself grant permission to
|
||||
practice related patent claims. See
|
||||
[`data/README.en.md`](../data/README.en.md) and
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the sources and the
|
||||
rights boundary.
|
||||
|
||||
## `.jochrtf`
|
||||
|
||||
A `.jochrtf` file is a pickle-free compressed NumPy archive with an exact member set:
|
||||
|
||||
| key | dtype / shape |
|
||||
|---|---|
|
||||
| `metadata_json` | NumPy Unicode scalar containing JSON text (`dtype.kind == "U"`) |
|
||||
| `band_center_frequencies_hz` | little-endian `float64[77]` |
|
||||
| `coefficients` | little-endian `complex128[36,2,77]` |
|
||||
| `delay_coefficients` | little-endian `float64[36,2]` |
|
||||
| `delay_bounds` | little-endian `float64[2,2]` |
|
||||
|
||||
Metadata uses the `JOC-HRTF-CACHE` magic and records the schema, compiler and
|
||||
phase-policy versions, ACN/N3D convention, filterbank hashes, SOFA content
|
||||
SHA-256, sample rate, radius, order, both ridge values, payload hash, and fit
|
||||
report. Every setting that changes compilation participates in the cache key.
|
||||
Metadata never persists an absolute local `source_path`; it may keep a display
|
||||
name only.
|
||||
|
||||
Before constructing a field, the loader uses `allow_pickle=False` and validates
|
||||
ZIP members and expanded sizes, shapes, dtypes, byte order, contiguous layout,
|
||||
finite values, delay bounds, band centers, payload hash, and cache key. The
|
||||
writer uses a same-directory temporary file, `fsync`, a process-held OS file
|
||||
lock, and atomic `os.replace`. Its hidden `.lock` sidecar may remain and does not
|
||||
mean that a writer still owns the lock. Outdated, damaged, or mismatched
|
||||
caches cannot hit. SOFA input rebuilds an invalid cache; an explicitly selected
|
||||
cache reports the error.
|
||||
|
||||
Deleting a disk cache must not change the field or render produced from the same
|
||||
SOFA and compiler configuration.
|
||||
|
||||
A `.jochrtf` file contains directional-field coefficients and delay data
|
||||
transformed from the source HRIRs. Its reproducibility therefore does not make
|
||||
it licence-free. Creating a cache does not enlarge the rights granted by the
|
||||
source SOFA/HRTF dataset: use, copying, and redistribution remain subject to
|
||||
that dataset's terms. If those terms are unclear, keep `.jochrtf` as a private
|
||||
local cache and do not ship it with the program or another build artifact.
|
||||
`source_sha256` is only a content-integrity identifier, not proof of provenance
|
||||
or permission.
|
||||
|
||||
## JOC objects and room behavior
|
||||
|
||||
The production adapter retains the existing JOC schedule:
|
||||
|
||||
- `[1536,16]` input per frame;
|
||||
- channel 0 is special LFE and channels 1..15 are JOC objects;
|
||||
- ID11/OAMD positions use a sample-timed timeline;
|
||||
- source parameters update every 512 samples;
|
||||
- every object owns independent direct/early history while one late FDN is shared;
|
||||
- `finish()` drains early/late tails; output gain is explicit, with no implicit
|
||||
limiter or programme loudness normalization.
|
||||
|
||||
Near/Mid/Far, equal-power direct level, six first-order shoebox image sources,
|
||||
late sends, the unitary FDN, the 120–180 Hz cosine-squared LFE low-pass, and room
|
||||
calibration are JOC project-defined behavior, not constants published by SOFA or
|
||||
Dolby.
|
||||
|
||||
The public SOFA binaural renderer defaults to the C++20 native core under
|
||||
`--backend auto/native` (`ejoc_sofa_binaural_*` in `lib/eac3joc_core.dll`): the
|
||||
filterbank, the SH direction-field evaluation, the per-object direct/early
|
||||
histories and the shared FDN all run natively, while Python only compiles the
|
||||
SOFA source and issues the per-512-sample metadata updates. When the native
|
||||
library is unavailable the renderer falls back to the Python/NumPy reference
|
||||
implementation; the two agree to better than 1e-9. `--backend python` forces
|
||||
the Python backend.
|
||||
`--backend` still selects native/Python JOC reconstruction and speaker rendering;
|
||||
native acceleration for the public binaural DSP is outside the current API.
|
||||
|
||||
## Technical references and rights boundary
|
||||
|
||||
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
|
||||
- [3GPP TS 26.405 / ETSI TS 126 405 (64-QMF/77-hybrid definition)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
- [Dolby binaural render-mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
|
||||
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en), used only as
|
||||
architectural background for direct/early/late, subbands, and FDNs; it does
|
||||
not establish that any product uses a particular embodiment.
|
||||
|
||||
Public availability of a specification, source file, or patent document does
|
||||
not by itself authorize copying its contents, redistribution of derivatives,
|
||||
or practice of patent claims. These technical references grant no patent
|
||||
licence and make no non-infringement representation. Anyone preparing a release
|
||||
or product integration must assess the applicable data and software licences,
|
||||
patent permissions, and freedom to operate. See
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the public-standard
|
||||
provenance and rights boundary.
|
||||
@@ -0,0 +1,245 @@
|
||||
# 双耳渲染
|
||||
|
||||
[English](binaural.en.md) · [返回 README](../README.md)
|
||||
|
||||
JustOneCacophony 的双耳后端支持三种 HRTF 来源:`SimpleFreeFieldHRIR` SOFA、
|
||||
杜比官方软件个性化扫描导出的 Rosella `.personalized_headphone`(JSON 解析由本项目
|
||||
自行实现,不调用杜比软件),以及从 SOFA 编译出的 `.jochrtf` 缓存。SOFA 在模型加载
|
||||
时编译成内存方向场;`.jochrtf` 只是可删除、可重建的 JOC compiled HRTF cache,
|
||||
不是交换格式,也不是使用 SOFA 的前置步骤。
|
||||
|
||||
```text
|
||||
SOFA FIR
|
||||
-> CanonicalHrtf
|
||||
-> 48 kHz / 单 radius shell / delay-phase policy
|
||||
-> 64-QMF / 77-hybrid projection
|
||||
-> 五阶 ACN/N3D 实球谐场
|
||||
-> 逐对象 direct + early reflections
|
||||
-> shared unitary-FDN late room
|
||||
-> stereo float64
|
||||
```
|
||||
|
||||
## 输入接口
|
||||
|
||||
CLI 有三个互斥的 HRTF 输入来源;都不指定时按默认规则自动选择:
|
||||
|
||||
```powershell
|
||||
# 1) SOFA:缺省取 HRTF/binaural.sofa,也可显式指定
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone:缺省取 HRTF/binaural.personalized_headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf:显式读取预编译 cache
|
||||
python main.py input.m4a --binaural `
|
||||
--compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# 可选:SOFA 透明生成/复用磁盘 cache
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
```
|
||||
|
||||
默认选择顺序:`HRTF/binaural.sofa` → `output/hrtf-cache` 下唯一的 `.jochrtf` →
|
||||
`HRTF/binaural.personalized_headphone`;三者都没有时报错并提示显式指定。
|
||||
`output/hrtf-cache` 下有多个 `.jochrtf` 时同样报错,要求显式选择。
|
||||
|
||||
`.personalized_headphone` 的 JSON 解析由本项目自行实现(`src/rosella_model.py`),
|
||||
不调用任何杜比软件。
|
||||
|
||||
`--hrtf-cache-policy` 可取 `none`、`memory`、`disk`。默认是 `memory`;`none` 和
|
||||
`memory` 都不会创建磁盘文件。`disk` 默认写入 `output/hrtf-cache`,也可用
|
||||
`--hrtf-cache-dir` 指定。`--hrtf-radius-m` 选择距离目标最近的 measurement shell。
|
||||
|
||||
Python API 使用显式 factory:
|
||||
|
||||
```python
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
|
||||
renderer = SofaBinauralBackend.from_sofa(
|
||||
"subject.sofa",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
cache_policy="memory",
|
||||
)
|
||||
|
||||
cached = SofaBinauralBackend.from_compiled_cache(
|
||||
"subject.jochrtf",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
)
|
||||
```
|
||||
|
||||
文件工厂不会按“未知后缀”猜格式:SOFA 和 `.jochrtf` 始终走不同 loader。
|
||||
|
||||
## 双耳渲染模式
|
||||
|
||||
`--binaural-mode off|near|mid|far`(默认 `mid`)是**人为指定的渲染提示**,不是
|
||||
从输入 E-AC-3 JOC 码流提取或还原的原始双耳元数据:
|
||||
|
||||
- 直接双耳渲染(`--binaural`):near/mid/far 生效,默认 `mid`;`off` 报错;
|
||||
- ADM BWF:DBMD segment 10 中后 15 个 JOC 对象的 binaural render mode 写
|
||||
`off=0/near=1/far=2/mid=3`,前 10 个 bed 保持不变,默认 `mid`;`off` 用于显式
|
||||
关闭双耳元数据提示。
|
||||
|
||||
## Canonical SOFA 契约
|
||||
|
||||
当前 strict importer 接受:
|
||||
|
||||
- `Conventions=SOFA`;
|
||||
- `SOFAConventions=SimpleFreeFieldHRIR`,version `0.4`、`1.0` 或 `1.1`;
|
||||
- `DataType=FIR`,`Data.IR[M,2,N]`;
|
||||
- 单一正有限 `Data.SamplingRate`,单位为 hertz/Hz;
|
||||
- spherical 或 Cartesian `SourcePosition`;
|
||||
- 单值或 per-measurement 的 `ListenerPosition/View/Up`;
|
||||
- 两个能由 listener-local lateral 坐标唯一识别左右的 receiver;
|
||||
- 单一且零偏移的 emitter;
|
||||
- causal `Data.Delay[I,2]` 或 `[M,2]`;
|
||||
- 明确的 free-field/anechoic `RoomType`。
|
||||
|
||||
receiver 左右顺序由几何决定,不能假定 `Data.IR` 的 receiver index。SOFA listener
|
||||
坐标为 $+X$ front、
|
||||
$+Y$ left、
|
||||
$+Z$ up;ADM 坐标为
|
||||
$+X$ right、
|
||||
$+Y$ front、
|
||||
$+Z$ up,转换为:
|
||||
|
||||
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
|
||||
|
||||
`CanonicalHrtf` 将 `Data.IR` 与 `Data.Delay` 分开保存。只有时域 baseline 才调用
|
||||
`materialized_measurement()` 将 delay 应用一次;运行时 SH 路径不先 materialize。
|
||||
非 48 kHz HRIR 使用 float64 `scipy.signal.resample_poly` 规范化,delay samples 按
|
||||
相同比例缩放。
|
||||
|
||||
GeneralFIR、BRIR、TF、多 emitter、多义 receiver 或非 free-field 数据需要单独的
|
||||
convention adapter,不能只通过 reshape 进入核心 importer。
|
||||
|
||||
## Delay/phase:exactly once
|
||||
|
||||
编译器只允许三种互斥语义:
|
||||
|
||||
1. 非零 `Data.Delay` 是 `Data.IR` 外部 delay;FIR 不去旋,运行时应用一次。
|
||||
2. `Data.Delay=0` 且 HRIR 有普通正 onset:以每耳 main peak 分离 arrival,拟合后
|
||||
在运行时恢复一次;当前阈值为 peak index 大于 2 samples。
|
||||
3. `Data.Delay=0` 且双耳 FIR 共享 sample-0 起点:不发明外部 delay,原 complex
|
||||
phase 直接进入五阶场。
|
||||
|
||||
任何路径都不能再叠加第二套 ear delay 或 phase-group delay。
|
||||
|
||||
## 公开 filterbank 与方向场
|
||||
|
||||
运行时固定为:
|
||||
|
||||
- 48 kHz;
|
||||
- 64-sample QMF hop;
|
||||
- 64-QMF / 77-hybrid;
|
||||
- analysis/synthesis latency 961 samples;
|
||||
- 五阶、36 项、ACN/N3D real spherical harmonics;
|
||||
- PCM、delay、SH、room state 为 float64;频带传递和频域状态为 complex128。
|
||||
|
||||
每个 hybrid band 的 real/imaginary 单位增益都通过同一套 analysis/synthesis 链生成
|
||||
脉冲字典,共 154 个实参数;编译不是直接读取 77 个 FFT bin。默认 projection
|
||||
ridge 为 `1e-3`,SH ridge 为 `1e-5`。同方向 measurement 先合并,再用球面 Voronoi
|
||||
面积权重做 ridge fit。
|
||||
|
||||
固定表位于 `data/rosella_kernels.npz`,实现公开标准化的滤波器组,各表可由如下
|
||||
公式计算。
|
||||
|
||||
hybrid 分析核定义于 [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
第 5.2.2 节(Table 1 的 $Q=8$/
|
||||
$Q=4$ 系数,delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
QMF analysis 表即 MPEG-4 AAC/SBR(ISO/IEC 14496-3/AMD1:2003 第 4.B.18.2 节)
|
||||
的 64 complex QMF bank;打包的 $64\times10$ 表是公开 640-tap prototype
|
||||
$c_0,\dots,c_{639}$ 的多相重排:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
QMF synthesis 表为上述 analysis 多相矩阵 $\mathbf{A}$ 的因果左逆,即求解
|
||||
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$(
|
||||
$\mathbf{P}$ 为 577-sample 延迟置换;
|
||||
全链 $961 = 577 + 6\times64$),以 rank-4 分解形式存储:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
hybrid synthesis 表为 77→64 重组:高频带恒等 $Y_{3+b}=X_{16+b}$;低频带(
|
||||
$C_p$ 为
|
||||
$8+4+4$ 子带划分):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\mathrm{Re}X_q + j\,s_q\,\mathrm{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
loader 校验 archive 和每个数组的 SHA-256;table version、所有数组 hash 与
|
||||
77 个 band-center 参考值都属于 cache key。标准可公开获取不等于获准实施相关
|
||||
专利;更多来源信息见 [`data/README.md`](../data/README.md) 与
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
|
||||
|
||||
## `.jochrtf`
|
||||
|
||||
`.jochrtf` 是无 pickle 的压缩 NumPy archive,固定包含:
|
||||
|
||||
| key | dtype / shape |
|
||||
|---|---|
|
||||
| `metadata_json` | 含 JSON 文本的 NumPy Unicode scalar(`dtype.kind == "U"`) |
|
||||
| `band_center_frequencies_hz` | little-endian `float64[77]` |
|
||||
| `coefficients` | little-endian `complex128[36,2,77]` |
|
||||
| `delay_coefficients` | little-endian `float64[36,2]` |
|
||||
| `delay_bounds` | little-endian `float64[2,2]` |
|
||||
|
||||
metadata magic 固定为 `JOC-HRTF-CACHE`,并记录 schema/compiler/phase-policy、
|
||||
ACN/N3D、filterbank table hashes、SOFA content SHA-256、采样率、radius、order、
|
||||
两个 ridge、payload hash 和 fit report。cache key 覆盖所有会改变编译结果的字段。
|
||||
metadata 不保存本机绝对 `source_path`,仅可保存 source display name。
|
||||
|
||||
loader 使用 `allow_pickle=False`,并在构造对象前检查 ZIP 成员集、解压大小、shape、
|
||||
dtype、端序、连续布局、有限值、delay bounds、band centers、payload hash 和 cache
|
||||
key。writer 使用同目录临时文件、`fsync`、进程持有的 OS 文件锁和原子
|
||||
`os.replace`;对应的隐藏 `.lock` sidecar 可保留,但不代表仍有 writer 持锁。
|
||||
旧版本、损坏或配置不匹配的 cache 不能命中;从 SOFA 启动时会重建,显式 cache
|
||||
入口则直接报错。
|
||||
|
||||
删除磁盘 cache 后,从同一 SOFA 和同一编译配置得到的场与渲染结果不得改变。
|
||||
|
||||
`.jochrtf` 包含由源 HRIR 变换得到的方向场系数与 delay 数据,因此“可以重建”不表示
|
||||
它不受数据许可约束。生成 cache 不会扩大源 SOFA/HRTF 数据集授予的权利;cache 的
|
||||
使用、复制和再分发仍须遵守源数据集条款。不能确认条款时,应把 `.jochrtf` 作为本地
|
||||
私有 cache,不随程序或构建产物发布。`source_sha256` 只用于内容一致性校验,不是许可
|
||||
或来源证明。
|
||||
|
||||
## JOC 对象与房间
|
||||
|
||||
生产适配器继续使用现有 JOC 调度:
|
||||
|
||||
- 每帧输入 `[1536,16]`;
|
||||
- channel 0 是 special LFE,channel 1..15 是 JOC objects;
|
||||
- ID11/OAMD position 使用 sample-timed timeline;
|
||||
- 每 512 samples 更新方向/profile;
|
||||
- 每个对象拥有独立 direct/early history,late FDN 全局共享;
|
||||
- `finish()` 排空 early/late tail;输出增益显式应用,不隐含 limiter 或节目响度归一化。
|
||||
|
||||
Near/Mid/Far、equal-power direct level、六面 shoebox 一阶 image source、late send、
|
||||
unitary FDN、LFE 120–180 Hz cosine-squared 低通及 room calibration 都是 JOC
|
||||
项目定义行为,不是 SOFA 或 Dolby 公布常数。
|
||||
|
||||
公开 SOFA 双耳渲染在 `--backend auto/native` 下默认走 C++20 原生核
|
||||
(`lib/eac3joc_core.dll` 的 `ejoc_sofa_binaural_*` 接口:filterbank、SH 方向场求值、
|
||||
逐对象 early/direct 历史与共享 FDN 全部在原生侧执行,Python 只做 SOFA 编译与每
|
||||
512-sample 的元数据更新);原生库不可用时自动回退 Python/NumPy 参考实现,两者
|
||||
逐值一致(差异 < 1e-9)。`--backend python` 强制使用 Python 后端。
|
||||
|
||||
## 技术引用与权利边界
|
||||
|
||||
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
|
||||
- [3GPP TS 26.405 / ETSI TS 126 405(64-QMF/77-hybrid 定义)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
- [Dolby binaural render mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
|
||||
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en),仅作 direct/early/late、
|
||||
subband 与 FDN 架构背景,不证明某个产品使用特定实施例。
|
||||
|
||||
规范、源码或专利文献可公开获取,不等于获准复制其内容、再分发派生产物或实施其中的
|
||||
专利权利要求。本项目的技术引用本身不授予专利许可,也不作不侵权保证;准备发布或集成
|
||||
到产品的一方应自行审查适用的数据许可、软件许可、专利许可及 freedom-to-operate。
|
||||
公开标准来源与权利边界见
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
|
||||
+107
-78
@@ -4,7 +4,7 @@
|
||||
|
||||
This document covers only the signal model and formulas used in the JustOneCacophony research path: how JOC parameters combine with core PCM to reconstruct object signals, and how OAMD coordinates become speaker gains.
|
||||
|
||||
The formulas describe the dense-JOC and ordinary point-object paths studied by the project. They are not a complete definition of every E-AC-3 JOC variant.
|
||||
The formulas describe the JOC matrix parameters (both the dense and the sparse differential syntax) and the ordinary point-object paths studied by the project. They are not a complete definition of every E-AC-3 JOC variant.
|
||||
|
||||
## 1. Overall path and notation
|
||||
|
||||
@@ -52,9 +52,11 @@ $$
|
||||
N_f=1536=24\times64.
|
||||
$$
|
||||
|
||||
## 2. Dense-JOC matrix parameters
|
||||
## 2. JOC matrix parameters
|
||||
|
||||
### 2.1 Differential reconstruction
|
||||
For every object and data point, the quantized matrix `joc_mix_mtx_q` is defined on $N_q$ quantization levels. The `b_joc_sparse` flag selects one of two differential syntaxes: dense sends one MTX difference per core channel, while sparse sends one active channel plus one coefficient difference per parameter band.
|
||||
|
||||
### 2.1 Dense differential reconstruction
|
||||
|
||||
Let `quant_idx` be $q_i\in\{0,1\}$. The number of quantization levels is
|
||||
|
||||
@@ -75,38 +77,76 @@ $$
|
||||
For object $o$, data point $d$, core channel $c$, and parameter band $p$, the coded difference $\Delta_{o,d,c,p}$ reconstructs to
|
||||
|
||||
$$
|
||||
Q_{o,d,c,0}
|
||||
=
|
||||
Q_{o,d,c,0}=
|
||||
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}
|
||||
=
|
||||
Q_{o,d,c,p}=
|
||||
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
|
||||
\qquad p>0.
|
||||
$$
|
||||
|
||||
### 2.2 Dequantization
|
||||
### 2.2 Sparse differential reconstruction
|
||||
|
||||
Let $I_{o,d,p}$ be the `joc_channel_idx` symbol (IDX), $V_{o,d,p}$ the `joc_vec` symbol (VEC), and $N_c\in\{5,7\}$ the number of core channels. Each parameter band has exactly one active channel:
|
||||
|
||||
$$
|
||||
A_{o,d,p}=
|
||||
\begin{cases}
|
||||
I_{o,d,0}, & p=0,\\[2pt]
|
||||
\left(A_{o,d,p-1}+I_{o,d,p}\right)\bmod N_c, & p>0,
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
where $I_{o,d,0}$ is a 3-bit absolute channel index and every later IDX symbol is an increment relative to the previous **active channel**. The coefficient is a single accumulator running across parameter bands:
|
||||
|
||||
$$
|
||||
\kappa_{o,d,-1}=O^{(s)}_q,\qquad
|
||||
\kappa_{o,d,p}=
|
||||
\left(\kappa_{o,d,p-1}+V_{o,d,p}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
with a sparse starting point two quantization levels above the dense center offset:
|
||||
|
||||
$$
|
||||
O^{(s)}_q=
|
||||
\begin{cases}
|
||||
50, & q_i=0,\\
|
||||
100, & q_i=1.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
The accumulator is **not** reset when the active channel changes. The complete matrix is
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}=
|
||||
\begin{cases}
|
||||
\kappa_{o,d,p}, & c=A_{o,d,p},\\[2pt]
|
||||
\dfrac{N_q}{2}, & c\neq A_{o,d,p}.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
Non-active entries take $N_q/2$, which dequantizes to exactly 0.
|
||||
|
||||
### 2.3 Dequantization
|
||||
|
||||
The dequantized matrix coefficient is
|
||||
|
||||
$$
|
||||
D_{o,d,c,p}
|
||||
=
|
||||
D_{o,d,c,p}=
|
||||
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
|
||||
\frac{820}{4096(1+q_i)}.
|
||||
$$
|
||||
|
||||
The effective denominator is therefore 4096 in coarse mode and 8192 in fine mode.
|
||||
|
||||
### 2.3 JOC clipgain
|
||||
### 2.4 JOC clipgain
|
||||
|
||||
If the clipgain field consists of integer $x$ and mantissa $y$, then
|
||||
|
||||
$$
|
||||
G_{\mathrm{clip}}
|
||||
=
|
||||
G_{\mathrm{clip}}=
|
||||
1+\frac{y}{32}2^{x-4}.
|
||||
$$
|
||||
|
||||
@@ -152,8 +192,7 @@ $$
|
||||
$$
|
||||
|
||||
$$
|
||||
M_{o,c,b,t}
|
||||
=
|
||||
M_{o,c,b,t}=
|
||||
(1-\alpha_t)P_{o,c,b}
|
||||
+\alpha_tD_{o,c,p(b)}.
|
||||
$$
|
||||
@@ -181,9 +220,8 @@ $$
|
||||
Let $\mathcal A_b$ denote the 64-band analysis-QMF operator with polyphase history state. Then
|
||||
|
||||
$$
|
||||
X_{c,b,t}
|
||||
=
|
||||
\mathcal A_b\!\left(
|
||||
X_{c,b,t}=
|
||||
\mathcal A_b\left(
|
||||
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
|
||||
\mathbf s^{\mathrm A}_{c,t}
|
||||
\right).
|
||||
@@ -210,8 +248,7 @@ $$
|
||||
Band 0 of each surround channel additionally passes through a 21-tap complex FIR:
|
||||
|
||||
$$
|
||||
\widehat X_{c,0,t}
|
||||
=
|
||||
\widehat X_{c,0,t}=
|
||||
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
|
||||
$$
|
||||
|
||||
@@ -222,8 +259,7 @@ These delays and filter histories are decoder state and cannot be reset independ
|
||||
For each object $o$, subband $b$, and slot $t$, the object's frequency-domain value is a linear combination of the five core channels:
|
||||
|
||||
$$
|
||||
Z_{o,b,t}
|
||||
=
|
||||
Z_{o,b,t}=
|
||||
\sum_{c=0}^{4}
|
||||
M_{o,c,b,t}\widehat X_{c,b,t}.
|
||||
$$
|
||||
@@ -238,21 +274,20 @@ Write the 64 complex subbands as 128 interleaved real values in `src`. For $k=0\
|
||||
|
||||
$$
|
||||
\begin{aligned}
|
||||
\operatorname{zone}[2k] &= \operatorname{src}[4k],\\
|
||||
\operatorname{zone}[2k+1] &= -\operatorname{src}[4k+1],\\
|
||||
\operatorname{zone}[126-2k] &= \operatorname{src}[4k+2],\\
|
||||
\operatorname{zone}[127-2k] &= \operatorname{src}[4k+3].
|
||||
\mathrm{zone}[2k] &= \mathrm{src}[4k],\\
|
||||
\mathrm{zone}[2k+1] &= -\mathrm{src}[4k+1],\\
|
||||
\mathrm{zone}[126-2k] &= \mathrm{src}[4k+2],\\
|
||||
\mathrm{zone}[127-2k] &= \mathrm{src}[4k+3].
|
||||
\end{aligned}
|
||||
$$
|
||||
|
||||
Treat `zone` as 64 complex values and apply an unnormalized 64-point FFT:
|
||||
|
||||
$$
|
||||
F_k
|
||||
=
|
||||
F_k=
|
||||
\sum_{n=0}^{63}
|
||||
\operatorname{zone}_n
|
||||
\exp\!\left(-j\frac{2\pi kn}{64}\right).
|
||||
\mathrm{zone}_n
|
||||
\exp\left(-j\frac{2\pi kn}{64}\right).
|
||||
$$
|
||||
|
||||
### 7.2 Modulation and synthesis
|
||||
@@ -260,8 +295,7 @@ $$
|
||||
Define the rotation coefficient
|
||||
|
||||
$$
|
||||
r_k
|
||||
=
|
||||
r_k=
|
||||
\frac12\left(
|
||||
\sin\frac{\pi k}{128}
|
||||
+j\cos\frac{\pi k}{128}
|
||||
@@ -277,9 +311,8 @@ $$
|
||||
Let $\mathcal S$ denote polyphase synthesis with a 640-value synthesis window and cross-slot state:
|
||||
|
||||
$$
|
||||
\mathbf y_{o,t}
|
||||
=
|
||||
\mathcal S\!\left(
|
||||
\mathbf y_{o,t}=
|
||||
\mathcal S\left(
|
||||
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
|
||||
\right).
|
||||
$$
|
||||
@@ -287,9 +320,8 @@ $$
|
||||
Object output is
|
||||
|
||||
$$
|
||||
y_o[64t+r]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
y_o[64t+r]=
|
||||
\mathrm{clip}\left(
|
||||
16\,\mathbf y_{o,t}[r],-1,1
|
||||
\right)G_{\mathrm{clip}},
|
||||
$$
|
||||
@@ -301,9 +333,8 @@ where $r=0\ldots63$. Synthesis state must advance continuously by slot.
|
||||
LFE bypasses the object matrix and inverse QMF and uses a 1217-sample delay. After the input and output scale factors cancel:
|
||||
|
||||
$$
|
||||
y_{\mathrm{LFE}}[n]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
y_{\mathrm{LFE}}[n]=
|
||||
\mathrm{clip}\left(
|
||||
x_{\mathrm{LFE,core}}[n-1217],-1,1
|
||||
\right).
|
||||
$$
|
||||
@@ -313,9 +344,8 @@ $$
|
||||
The lateral and longitudinal grids use $N=62$; the height grid uses $N=15$. The quantizer is
|
||||
|
||||
$$
|
||||
q_N(k)
|
||||
=
|
||||
\min\!\left(
|
||||
q_N(k)=
|
||||
\min\left(
|
||||
32767,
|
||||
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
|
||||
\right).
|
||||
@@ -336,11 +366,11 @@ Their maximum runtime value is $32767/32768$, not exactly 1.
|
||||
For conversion to the ADM grid:
|
||||
|
||||
$$
|
||||
k_1=\operatorname{round}\!\left(\frac{62q_1}{32767}\right),
|
||||
k_1=\mathrm{round}\left(\frac{62q_1}{32767}\right),
|
||||
\quad
|
||||
k_2=\operatorname{round}\!\left(\frac{62q_2}{32767}\right),
|
||||
k_2=\mathrm{round}\left(\frac{62q_2}{32767}\right),
|
||||
\quad
|
||||
k_3=\operatorname{round}\!\left(\frac{15q_3}{32767}\right),
|
||||
k_3=\mathrm{round}\left(\frac{15q_3}{32767}\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
@@ -404,17 +434,15 @@ $$
|
||||
The two-dimensional point gain is
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{2D}}(u,v)
|
||||
=
|
||||
\mathbf G_{\mathrm{2D}}(u,v)=
|
||||
\mathbf h(u)\odot\mathbf v(v).
|
||||
$$
|
||||
|
||||
For 5.1-family layouts with one horizontal surround pair rather than separate side and rear pairs, the longitudinal coordinate is
|
||||
|
||||
$$
|
||||
v_{\mathrm{floor}}
|
||||
=
|
||||
\operatorname{clamp}(2v,0,1).
|
||||
v_{\mathrm{floor}}=
|
||||
\mathrm{clamp}(2v,0,1).
|
||||
$$
|
||||
|
||||
Other layouts use $v_{\mathrm{floor}}=v$.
|
||||
@@ -424,8 +452,7 @@ Other layouts use $v_{\mathrm{floor}}=v$.
|
||||
Three-dimensional layouts compute floor gain $\mathbf G_f$ and height gain $\mathbf G_h$ separately:
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)
|
||||
=
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)=
|
||||
\cos\left(\frac\pi2w\right)\mathbf G_f
|
||||
+
|
||||
\sin\left(\frac\pi2w\right)\mathbf G_h.
|
||||
@@ -450,8 +477,7 @@ $$
|
||||
Maximum position compensation is
|
||||
|
||||
$$
|
||||
A_{\max}
|
||||
=
|
||||
A_{\max}=
|
||||
-\max\left(4.5-1.5H-3F,0\right)
|
||||
\quad\text{dB}.
|
||||
$$
|
||||
@@ -459,15 +485,15 @@ $$
|
||||
Longitudinal and height weights are
|
||||
|
||||
$$
|
||||
p_v=\operatorname{clamp}\left(\frac v{0.6},0,1\right),
|
||||
p_v=\mathrm{clamp}\left(\frac v{0.6},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p_w=\operatorname{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
p_w=\mathrm{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p=\operatorname{clamp}(p_v+p_w,0,1).
|
||||
p=\mathrm{clamp}(p_v+p_w,0,1).
|
||||
$$
|
||||
|
||||
The linear compensation gain is
|
||||
@@ -479,8 +505,7 @@ $$
|
||||
The object's target-gain vector is
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{target}}
|
||||
=
|
||||
\mathbf G_{\mathrm{target}}=
|
||||
G_{\mathrm{object}}
|
||||
G_{\mathrm{pos}}
|
||||
\mathbf G_{\mathrm{point}}.
|
||||
@@ -491,29 +516,36 @@ $$
|
||||
The coded position of an OAMD update is
|
||||
|
||||
$$
|
||||
s_{\mathrm{coded}}
|
||||
=
|
||||
s_{\mathrm{coded}}=
|
||||
s_{\mathrm{frame}}
|
||||
+s_{\mathrm{outer}}
|
||||
+s_{\mathrm{OAMD}}
|
||||
+32f_{\mathrm{block}}.
|
||||
$$
|
||||
|
||||
For processing-block length $B=32$, the aligned update point is
|
||||
The theoretical update position on the decoder-output PCM timeline is
|
||||
|
||||
$$
|
||||
\widehat s
|
||||
=
|
||||
s_{\mathrm{theoretical}}=
|
||||
s_{\mathrm{coded}}+d_{\mathrm{decoder}},
|
||||
\qquad d_{\mathrm{decoder}}=1473.
|
||||
$$
|
||||
|
||||
The speaker renderer retains the existing processing-block length $B=32$, so the aligned update point is
|
||||
|
||||
$$
|
||||
\widehat s=
|
||||
B\left\lfloor
|
||||
\frac{s_{\mathrm{coded}}+B/2-1}{B}
|
||||
\frac{s_{\mathrm{theoretical}}+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
$$
|
||||
|
||||
Thus, for frame-aligned updates, `align32(1473)=1472`. The 1473 value is the theoretical decoder delay at the metadata interface; 1472 is its effective boundary in the current 32-sample control block. The 640-value inverse-QMF window/state is not part of this metadata-timing formula.
|
||||
|
||||
For ramp duration $D$, the number of blocks is
|
||||
|
||||
$$
|
||||
K
|
||||
=
|
||||
K=
|
||||
\left\lfloor
|
||||
\frac{D+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
@@ -540,8 +572,7 @@ If no new metadata update intervenes, this is equivalent to a sample-wise linear
|
||||
For target output channel $c$:
|
||||
|
||||
$$
|
||||
y_c[n]
|
||||
=
|
||||
y_c[n]=
|
||||
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
|
||||
+
|
||||
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
|
||||
@@ -550,8 +581,7 @@ $$
|
||||
Here
|
||||
|
||||
$$
|
||||
\delta_{c,\mathrm{LFE}}
|
||||
=
|
||||
\delta_{c,\mathrm{LFE}}=
|
||||
\begin{cases}
|
||||
1, & c\text{ is the target layout's LFE channel},\\
|
||||
0, & \text{otherwise}.
|
||||
@@ -563,16 +593,15 @@ A layout without LFE output does not mix input LFE into other channels. After ob
|
||||
For PCM24 output, quantization is
|
||||
|
||||
$$
|
||||
y_{24}[n]
|
||||
=
|
||||
\operatorname{trunc}\left(
|
||||
8388607\,\operatorname{clip}(y[n],-1,1)
|
||||
y_{24}[n]=
|
||||
\mathrm{trunc}\left(
|
||||
8388607\,\mathrm{clip}(y[n],-1,1)
|
||||
\right).
|
||||
$$
|
||||
|
||||
## 14. Scope of the formulas
|
||||
|
||||
- The JOC matrix section describes dense JOC; Sparse JOC uses a different sparse coefficient/index path.
|
||||
- The JOC matrix section covers both the dense MTX and the sparse IDX/VEC differential syntax.
|
||||
- The speaker-panning section describes ordinary point objects; extent, spread, divergence, and similar modes require additional models.
|
||||
- Multiple OAMD position blocks must be scheduled in time order.
|
||||
- A limiter is separate post-processing and is not included in the mixing equations above.
|
||||
|
||||
+107
-78
@@ -4,7 +4,7 @@
|
||||
|
||||
本文只说明 JustOneCacophony 研究路径中使用的信号模型和公式:JOC 参数如何与核心 PCM 结合并重建对象信号,以及 OAMD 坐标如何转换为扬声器增益。
|
||||
|
||||
这些公式描述项目当前研究的 dense JOC 与普通点对象路径,不代表对所有 E-AC-3 JOC 变体的完整定义。
|
||||
这些公式描述项目当前研究的 JOC 矩阵参数(dense 与 sparse 两条差分语法)与普通点对象路径,不代表对所有 E-AC-3 JOC 变体的完整定义。
|
||||
|
||||
## 1. 总体路径与记号
|
||||
|
||||
@@ -52,9 +52,11 @@ $$
|
||||
N_f=1536=24\times64.
|
||||
$$
|
||||
|
||||
## 2. Dense JOC 矩阵参数
|
||||
## 2. JOC 矩阵参数
|
||||
|
||||
### 2.1 差分还原
|
||||
每个对象、每个数据点的量化矩阵 `joc_mix_mtx_q` 都定义在 $N_q$ 个量化级上。标志位 `b_joc_sparse` 选择两条差分语法之一:dense 为每个核心声道各送一路 MTX 差分,sparse 每参数带只送一个 active 声道与一路系数差分。
|
||||
|
||||
### 2.1 Dense 差分还原
|
||||
|
||||
令 `quant_idx` 为 $q_i\in\{0,1\}$,量化级数为
|
||||
|
||||
@@ -75,38 +77,76 @@ $$
|
||||
对对象 $o$、数据点 $d$、核心声道 $c$ 和参数带 $p$,编码差分 $\Delta_{o,d,c,p}$ 还原为
|
||||
|
||||
$$
|
||||
Q_{o,d,c,0}
|
||||
=
|
||||
Q_{o,d,c,0}=
|
||||
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}
|
||||
=
|
||||
Q_{o,d,c,p}=
|
||||
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
|
||||
\qquad p>0.
|
||||
$$
|
||||
|
||||
### 2.2 去量化
|
||||
### 2.2 Sparse 差分还原
|
||||
|
||||
令 $I_{o,d,p}$ 为 `joc_channel_idx` 符号(IDX),$V_{o,d,p}$ 为 `joc_vec` 符号(VEC),$N_c\in\{5,7\}$ 为核心声道数。每参数带只有一个 active 声道
|
||||
|
||||
$$
|
||||
A_{o,d,p}=
|
||||
\begin{cases}
|
||||
I_{o,d,0}, & p=0,\\[2pt]
|
||||
\left(A_{o,d,p-1}+I_{o,d,p}\right)\bmod N_c, & p>0,
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
其中 $I_{o,d,0}$ 是 3 bit 绝对声道号,其余 IDX 符号是相对上一个 **active 声道**的增量。系数是一个跨参数带连续的单累加器
|
||||
|
||||
$$
|
||||
\kappa_{o,d,-1}=O^{(s)}_q,\qquad
|
||||
\kappa_{o,d,p}=
|
||||
\left(\kappa_{o,d,p-1}+V_{o,d,p}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
sparse 起点比 dense 的中心偏移高两个量化级:
|
||||
|
||||
$$
|
||||
O^{(s)}_q=
|
||||
\begin{cases}
|
||||
50, & q_i=0,\\
|
||||
100, & q_i=1.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
active 声道切换时累加器**不**重置。完整矩阵为
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}=
|
||||
\begin{cases}
|
||||
\kappa_{o,d,p}, & c=A_{o,d,p},\\[2pt]
|
||||
\dfrac{N_q}{2}, & c\neq A_{o,d,p}.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
非 active 项取 $N_q/2$,即去量化后恰为 0。
|
||||
|
||||
### 2.3 去量化
|
||||
|
||||
矩阵系数的去量化值为
|
||||
|
||||
$$
|
||||
D_{o,d,c,p}
|
||||
=
|
||||
D_{o,d,c,p}=
|
||||
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
|
||||
\frac{820}{4096(1+q_i)}.
|
||||
$$
|
||||
|
||||
因此 coarse 模式的有效分母为 4096,fine 模式为 8192。
|
||||
|
||||
### 2.3 JOC clipgain
|
||||
### 2.4 JOC clipgain
|
||||
|
||||
若 clipgain 字段由整数 $x$ 和尾数 $y$ 组成,则
|
||||
|
||||
$$
|
||||
G_{\mathrm{clip}}
|
||||
=
|
||||
G_{\mathrm{clip}}=
|
||||
1+\frac{y}{32}2^{x-4}.
|
||||
$$
|
||||
|
||||
@@ -152,8 +192,7 @@ $$
|
||||
$$
|
||||
|
||||
$$
|
||||
M_{o,c,b,t}
|
||||
=
|
||||
M_{o,c,b,t}=
|
||||
(1-\alpha_t)P_{o,c,b}
|
||||
+\alpha_tD_{o,c,p(b)}.
|
||||
$$
|
||||
@@ -181,9 +220,8 @@ $$
|
||||
令 $\mathcal A_b$ 表示带 polyphase 历史状态的 64-band analysis-QMF 算子,则
|
||||
|
||||
$$
|
||||
X_{c,b,t}
|
||||
=
|
||||
\mathcal A_b\!\left(
|
||||
X_{c,b,t}=
|
||||
\mathcal A_b\left(
|
||||
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
|
||||
\mathbf s^{\mathrm A}_{c,t}
|
||||
\right).
|
||||
@@ -210,8 +248,7 @@ $$
|
||||
环绕声道的 band 0 还经过 21-tap 复 FIR:
|
||||
|
||||
$$
|
||||
\widehat X_{c,0,t}
|
||||
=
|
||||
\widehat X_{c,0,t}=
|
||||
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
|
||||
$$
|
||||
|
||||
@@ -222,8 +259,7 @@ $$
|
||||
对每个对象 $o$、子带 $b$ 和时槽 $t$,对象频域值为五个核心声道的线性组合:
|
||||
|
||||
$$
|
||||
Z_{o,b,t}
|
||||
=
|
||||
Z_{o,b,t}=
|
||||
\sum_{c=0}^{4}
|
||||
M_{o,c,b,t}\widehat X_{c,b,t}.
|
||||
$$
|
||||
@@ -238,21 +274,20 @@ analysis 输入的 $1/16$ 缩放会在 inverse QMF 输出端由 $\times16$ 抵
|
||||
|
||||
$$
|
||||
\begin{aligned}
|
||||
\operatorname{zone}[2k] &= \operatorname{src}[4k],\\
|
||||
\operatorname{zone}[2k+1] &= -\operatorname{src}[4k+1],\\
|
||||
\operatorname{zone}[126-2k] &= \operatorname{src}[4k+2],\\
|
||||
\operatorname{zone}[127-2k] &= \operatorname{src}[4k+3].
|
||||
\mathrm{zone}[2k] &= \mathrm{src}[4k],\\
|
||||
\mathrm{zone}[2k+1] &= -\mathrm{src}[4k+1],\\
|
||||
\mathrm{zone}[126-2k] &= \mathrm{src}[4k+2],\\
|
||||
\mathrm{zone}[127-2k] &= \mathrm{src}[4k+3].
|
||||
\end{aligned}
|
||||
$$
|
||||
|
||||
把 `zone` 重新视为 64 个复数后执行未归一化 64 点 FFT:
|
||||
|
||||
$$
|
||||
F_k
|
||||
=
|
||||
F_k=
|
||||
\sum_{n=0}^{63}
|
||||
\operatorname{zone}_n
|
||||
\exp\!\left(-j\frac{2\pi kn}{64}\right).
|
||||
\mathrm{zone}_n
|
||||
\exp\left(-j\frac{2\pi kn}{64}\right).
|
||||
$$
|
||||
|
||||
### 7.2 调制与合成
|
||||
@@ -260,8 +295,7 @@ $$
|
||||
定义旋转系数
|
||||
|
||||
$$
|
||||
r_k
|
||||
=
|
||||
r_k=
|
||||
\frac12\left(
|
||||
\sin\frac{\pi k}{128}
|
||||
+j\cos\frac{\pi k}{128}
|
||||
@@ -277,9 +311,8 @@ $$
|
||||
令 $\mathcal S$ 表示带 640 项 synthesis window 和跨时槽状态的 polyphase 合成算子:
|
||||
|
||||
$$
|
||||
\mathbf y_{o,t}
|
||||
=
|
||||
\mathcal S\!\left(
|
||||
\mathbf y_{o,t}=
|
||||
\mathcal S\left(
|
||||
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
|
||||
\right).
|
||||
$$
|
||||
@@ -287,9 +320,8 @@ $$
|
||||
对象输出为
|
||||
|
||||
$$
|
||||
y_o[64t+r]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
y_o[64t+r]=
|
||||
\mathrm{clip}\left(
|
||||
16\,\mathbf y_{o,t}[r],-1,1
|
||||
\right)G_{\mathrm{clip}},
|
||||
$$
|
||||
@@ -301,9 +333,8 @@ $$
|
||||
LFE 不经过对象矩阵或 inverse QMF,而是使用 1217-sample 延迟。输入与输出端的比例因子抵消后:
|
||||
|
||||
$$
|
||||
y_{\mathrm{LFE}}[n]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
y_{\mathrm{LFE}}[n]=
|
||||
\mathrm{clip}\left(
|
||||
x_{\mathrm{LFE,core}}[n-1217],-1,1
|
||||
\right).
|
||||
$$
|
||||
@@ -313,9 +344,8 @@ $$
|
||||
横向和纵向网格使用 $N=62$,高度网格使用 $N=15$。量化函数为
|
||||
|
||||
$$
|
||||
q_N(k)
|
||||
=
|
||||
\min\!\left(
|
||||
q_N(k)=
|
||||
\min\left(
|
||||
32767,
|
||||
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
|
||||
\right).
|
||||
@@ -336,11 +366,11 @@ $$
|
||||
转换为 ADM 网格时:
|
||||
|
||||
$$
|
||||
k_1=\operatorname{round}\!\left(\frac{62q_1}{32767}\right),
|
||||
k_1=\mathrm{round}\left(\frac{62q_1}{32767}\right),
|
||||
\quad
|
||||
k_2=\operatorname{round}\!\left(\frac{62q_2}{32767}\right),
|
||||
k_2=\mathrm{round}\left(\frac{62q_2}{32767}\right),
|
||||
\quad
|
||||
k_3=\operatorname{round}\!\left(\frac{15q_3}{32767}\right),
|
||||
k_3=\mathrm{round}\left(\frac{15q_3}{32767}\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
@@ -404,17 +434,15 @@ $$
|
||||
二维点增益为
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{2D}}(u,v)
|
||||
=
|
||||
\mathbf G_{\mathrm{2D}}(u,v)=
|
||||
\mathbf h(u)\odot\mathbf v(v).
|
||||
$$
|
||||
|
||||
对于只有一对水平环绕、没有独立 side/rear 两对的 5.1 系列布局,纵向坐标使用
|
||||
|
||||
$$
|
||||
v_{\mathrm{floor}}
|
||||
=
|
||||
\operatorname{clamp}(2v,0,1).
|
||||
v_{\mathrm{floor}}=
|
||||
\mathrm{clamp}(2v,0,1).
|
||||
$$
|
||||
|
||||
其他布局使用 $v_{\mathrm{floor}}=v$。
|
||||
@@ -424,8 +452,7 @@ $$
|
||||
三维布局分别计算地面层增益 $\mathbf G_f$ 和高度层增益 $\mathbf G_h$:
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)
|
||||
=
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)=
|
||||
\cos\left(\frac\pi2w\right)\mathbf G_f
|
||||
+
|
||||
\sin\left(\frac\pi2w\right)\mathbf G_h.
|
||||
@@ -450,8 +477,7 @@ $$
|
||||
最大位置补偿为
|
||||
|
||||
$$
|
||||
A_{\max}
|
||||
=
|
||||
A_{\max}=
|
||||
-\max\left(4.5-1.5H-3F,0\right)
|
||||
\quad\text{dB}.
|
||||
$$
|
||||
@@ -459,15 +485,15 @@ $$
|
||||
前后与高度位置权重为
|
||||
|
||||
$$
|
||||
p_v=\operatorname{clamp}\left(\frac v{0.6},0,1\right),
|
||||
p_v=\mathrm{clamp}\left(\frac v{0.6},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p_w=\operatorname{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
p_w=\mathrm{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p=\operatorname{clamp}(p_v+p_w,0,1).
|
||||
p=\mathrm{clamp}(p_v+p_w,0,1).
|
||||
$$
|
||||
|
||||
线性补偿增益为
|
||||
@@ -479,8 +505,7 @@ $$
|
||||
对象的目标增益向量为
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{target}}
|
||||
=
|
||||
\mathbf G_{\mathrm{target}}=
|
||||
G_{\mathrm{object}}
|
||||
G_{\mathrm{pos}}
|
||||
\mathbf G_{\mathrm{point}}.
|
||||
@@ -491,29 +516,36 @@ $$
|
||||
OAMD 更新的编码位置为
|
||||
|
||||
$$
|
||||
s_{\mathrm{coded}}
|
||||
=
|
||||
s_{\mathrm{coded}}=
|
||||
s_{\mathrm{frame}}
|
||||
+s_{\mathrm{outer}}
|
||||
+s_{\mathrm{OAMD}}
|
||||
+32f_{\mathrm{block}}.
|
||||
$$
|
||||
|
||||
对处理块长度 $B=32$,更新点对齐为
|
||||
decoder 输出 PCM timeline 上的理论更新位置为
|
||||
|
||||
$$
|
||||
\widehat s
|
||||
=
|
||||
s_{\mathrm{theoretical}}=
|
||||
s_{\mathrm{coded}}+d_{\mathrm{decoder}},
|
||||
\qquad d_{\mathrm{decoder}}=1473.
|
||||
$$
|
||||
|
||||
扬声器 renderer 保留现有的处理块长度 $B=32$,更新点对齐为
|
||||
|
||||
$$
|
||||
\widehat s=
|
||||
B\left\lfloor
|
||||
\frac{s_{\mathrm{coded}}+B/2-1}{B}
|
||||
\frac{s_{\mathrm{theoretical}}+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
$$
|
||||
|
||||
因此,对 frame-aligned 更新有 `align32(1473)=1472`。1473 是 metadata interface 的理论 decoder delay;1472 是当前 32-sample control block 中的有效边界。inverse-QMF 使用的 640 项 window/state 不属于这条 metadata timing 公式。
|
||||
|
||||
给定 ramp duration $D$,block 数为
|
||||
|
||||
$$
|
||||
K
|
||||
=
|
||||
K=
|
||||
\left\lfloor
|
||||
\frac{D+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
@@ -540,8 +572,7 @@ $$
|
||||
对目标输出声道 $c$:
|
||||
|
||||
$$
|
||||
y_c[n]
|
||||
=
|
||||
y_c[n]=
|
||||
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
|
||||
+
|
||||
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
|
||||
@@ -550,8 +581,7 @@ $$
|
||||
其中
|
||||
|
||||
$$
|
||||
\delta_{c,\mathrm{LFE}}
|
||||
=
|
||||
\delta_{c,\mathrm{LFE}}=
|
||||
\begin{cases}
|
||||
1, & c\text{ 为目标布局的 LFE},\\
|
||||
0, & \text{其他声道}.
|
||||
@@ -563,16 +593,15 @@ $$
|
||||
若输出 PCM24,量化关系为
|
||||
|
||||
$$
|
||||
y_{24}[n]
|
||||
=
|
||||
\operatorname{trunc}\left(
|
||||
8388607\,\operatorname{clip}(y[n],-1,1)
|
||||
y_{24}[n]=
|
||||
\mathrm{trunc}\left(
|
||||
8388607\,\mathrm{clip}(y[n],-1,1)
|
||||
\right).
|
||||
$$
|
||||
|
||||
## 14. 公式适用范围
|
||||
|
||||
- JOC 矩阵部分描述 dense JOC;Sparse JOC 使用不同的稀疏系数/索引路径。
|
||||
- JOC 矩阵部分同时描述 dense MTX 与 sparse IDX/VEC 两条差分语法。
|
||||
- 扬声器声像部分描述普通点对象;extent、spread、divergence 等模式需要额外模型。
|
||||
- 多个 OAMD position block 必须按其时间顺序调度。
|
||||
- limiter 属于独立后处理,不包含在上述混音公式中。
|
||||
|
||||
+23
-4
@@ -42,7 +42,7 @@ int ejoc_renderer_process(
|
||||
float* output16_planar); /* [16][1536] */
|
||||
```
|
||||
|
||||
Python performs dense-JOC Huffman decoding, differential reconstruction, and dequantization before the call. Sparse JOC is not silently passed to the dense native path.
|
||||
Python performs JOC Huffman decoding, differential reconstruction, and dequantization before the call, with the dense and sparse syntaxes sharing one entry point. The native core consumes the already dequantized `dq` in double precision, and both syntaxes have the same layout at that ABI.
|
||||
|
||||
Thread control is exposed as:
|
||||
|
||||
@@ -116,7 +116,26 @@ Supported layouts:
|
||||
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
|
||||
```
|
||||
|
||||
## 6. Building
|
||||
## 6. Binaural-rendering ABI
|
||||
|
||||
The shared library provides a 512-sample float64 binaural DSP interface:
|
||||
|
||||
```c
|
||||
ejoc_binaural_renderer_handle ejoc_binaural_renderer_create(void);
|
||||
int ejoc_binaural_renderer_configure_kernels(...);
|
||||
int ejoc_binaural_renderer_configure_room(...);
|
||||
int ejoc_binaural_renderer_process(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
const double* input16_interleaved, /* [512][16] */
|
||||
const double* gains_complex, /* [16][2][77][2] */
|
||||
const double* room_sends, /* [16] */
|
||||
double output_gain,
|
||||
double* output_stereo_interleaved); /* [512][2] */
|
||||
```
|
||||
|
||||
Python parses the model, evaluates the OAMD timeline, and supplies complex gains and room sends every 512 samples. The C++ handle owns QMF, hybrid, recursive-room, and QMF-synthesis state. Inputs, state, accumulation, and output are double/complex double.
|
||||
|
||||
## 7. Building
|
||||
|
||||
The CMake definition is `native/CMakeLists.txt`. Run from the repository root:
|
||||
|
||||
@@ -138,7 +157,7 @@ The MSVC configuration uses the static CRT. Other runtime dependencies depend on
|
||||
|
||||
The repository does not include native binaries by default. A prebuilt Release runtime or a locally built runtime can be placed directly under `lib/`.
|
||||
|
||||
## 7. Runtime lookup and fallback
|
||||
## 8. Runtime lookup and fallback
|
||||
|
||||
Lookup order:
|
||||
|
||||
@@ -148,7 +167,7 @@ Lookup order:
|
||||
|
||||
`--backend auto` falls back to NumPy when loading fails, and `--backend python` skips native discovery. The current CLI also prints the failure and falls back for `--backend native`; this existing behavior should not be read as successful native execution.
|
||||
|
||||
## 8. Implementation boundaries
|
||||
## 9. Implementation boundaries
|
||||
|
||||
- The native layer accepts only dense-JOC data already parsed by Python.
|
||||
- The ABI fixes a 1536-sample JOC frame, at most 15 objects, at most 23 parameter bands, and at most 2 data points.
|
||||
|
||||
+36
-4
@@ -42,7 +42,7 @@ int ejoc_renderer_process(
|
||||
float* output16_planar); /* [16][1536] */
|
||||
```
|
||||
|
||||
Dense JOC 的 Huffman 解码、差分还原和去量化先在 Python 中完成。Sparse JOC 不会被静默送入 dense 原生路径。
|
||||
JOC 的 Huffman 解码、差分还原和去量化先在 Python 中完成,dense 与 sparse 两条语法共用同一条入口。原生核心消费已去量化的 `dq`(double),两条语法在该 ABI 上布局一致。
|
||||
|
||||
线程接口为:
|
||||
|
||||
@@ -116,7 +116,39 @@ int ejoc_speaker_renderer_process(
|
||||
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
|
||||
```
|
||||
|
||||
## 6. 构建
|
||||
## 6. 双耳渲染 ABI
|
||||
|
||||
共享库提供 512-sample float64 双耳 DSP:
|
||||
|
||||
```c
|
||||
ejoc_binaural_renderer_handle ejoc_binaural_renderer_create(void);
|
||||
int ejoc_binaural_renderer_configure_kernels(...);
|
||||
int ejoc_binaural_renderer_configure_room(...);
|
||||
int ejoc_binaural_renderer_process(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
const double* input16_interleaved, /* [512][16] */
|
||||
const double* gains_complex, /* [16][2][77][2] */
|
||||
const double* room_sends, /* [16] */
|
||||
double output_gain,
|
||||
double* output_stereo_interleaved); /* [512][2] */
|
||||
```
|
||||
|
||||
Python 负责模型解析、OAMD 时间轴和每 512 samples 的 complex gains/room sends。C++ handle 保存 QMF、hybrid、递归 room 和 QMF synthesis 状态。全部输入、状态、乘加和输出均为 double/complex double。
|
||||
|
||||
## 6.1 公开 SOFA 双耳渲染 ABI
|
||||
|
||||
共享库同时提供完整的原生 SOFA 双耳渲染器(`ejoc_sofa_binaural_*`),它镜像
|
||||
Python `SofaBinauralBackend` 的全部数学:64-QMF/77-hybrid analysis/synthesis、
|
||||
五阶 ACN/N3D 实球谐方向场求值、whole-QMF-slot 逐对象 delay 历史、六面一阶
|
||||
image-source early reflections、共享 unitary FDN late room、LFE 120–180 Hz
|
||||
低通与 961-sample latency 语义。kernel 表、编译好的 HRTF 场与房间常数通过
|
||||
`configure_kernels/configure_field/configure_room` 一次上传;每 512-sample
|
||||
block 先 `set_source` 更新 16 个 source,再 `process` 输入 PCM;`process` 返回
|
||||
裁剪后的 stereo 样本数(首个 961 samples 被丢弃)。`finish` 以 64-sample 对齐的
|
||||
块排空尾音。Python 桥位于 `src/sofa_native_backend.py`,与 Python 参考实现逐值
|
||||
一致(差异 < 1e-9);原生库缺失时 `main.py` 自动回退 Python。
|
||||
|
||||
## 7. 构建
|
||||
|
||||
CMake 定义位于 `native/CMakeLists.txt`。从仓库根目录运行:
|
||||
|
||||
@@ -138,7 +170,7 @@ MSVC 配置使用静态 CRT。其他运行时依赖由平台和工具链决定
|
||||
|
||||
仓库默认不附带原生二进制。预构建的 Release 运行库或自行构建的运行库均可直接放入 `lib/`。
|
||||
|
||||
## 7. 运行时查找与回退
|
||||
## 8. 运行时查找与回退
|
||||
|
||||
查找顺序为:
|
||||
|
||||
@@ -148,7 +180,7 @@ MSVC 配置使用静态 CRT。其他运行时依赖由平台和工具链决定
|
||||
|
||||
`--backend auto` 在加载失败时回退到 NumPy;`--backend python` 跳过原生探测。`--backend native` 当前也会打印失败原因后回退,这是现有 CLI 行为,不应理解为原生库已成功使用。
|
||||
|
||||
## 8. 实现边界
|
||||
## 9. 实现边界
|
||||
|
||||
- 原生层只接收 Python 已解析的 dense JOC 数据。
|
||||
- ABI 固定了 1536-sample JOC 帧、最多 15 个对象、最多 23 个参数带和最多 2 个数据点。
|
||||
|
||||
@@ -6,6 +6,7 @@ import math
|
||||
import os
|
||||
from pathlib import Path
|
||||
import platform
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
@@ -20,35 +21,146 @@ if str(SOURCE_DIR) not in sys.path:
|
||||
import numpy as np
|
||||
|
||||
import adm_assemble
|
||||
import adm_atmos
|
||||
from adm_validate import validate
|
||||
from metadata import DirectPayloadIndex, PayloadIndex, write_summary
|
||||
import oamd_tracks
|
||||
from renderer import JocRenderer
|
||||
from native_renderer import NativeBackendUnavailable, NativeJocRenderer
|
||||
from binaural_renderer import (
|
||||
DEFAULT_SOFA_HRTF,
|
||||
SofaBinauralRenderer,
|
||||
resolve_compiled_hrtf_cache,
|
||||
resolve_sofa_hrtf,
|
||||
)
|
||||
from rosella_binaural_renderer import (
|
||||
DEFAULT_PERSONALIZED_HEADPHONE,
|
||||
ROSSELLA_BLOCK_SAMPLES,
|
||||
ROSSELLA_LATENCY_SAMPLES,
|
||||
RosellaBinauralRenderer,
|
||||
resolve_personalized_headphone,
|
||||
)
|
||||
from sofa_hrtf_field import DEFAULT_HRTF_CACHE_DIR
|
||||
from speaker_backend import create_speaker_renderer
|
||||
from speaker_layouts import (SPEAKER_LAYOUT_CHOICES, get_speaker_layout,
|
||||
speaker_layout_display_name)
|
||||
from speaker_wav import SpeakerPcmSpool, write_speaker_wav
|
||||
from speaker_wav import BinauralPcmSpool, SpeakerPcmSpool, write_pcm_wav
|
||||
from variant_error import UnsupportedVariantError, write_variant_report
|
||||
|
||||
|
||||
RATE = 48000
|
||||
FRAME_SAMPLES = 1536
|
||||
DEFAULT_OUTPUT_DIR = PROJECT_DIR / "output"
|
||||
EAC3_DRC_SCALE_MAX = 6.0
|
||||
EAC3_TARGET_LEVEL_RANGE = (-31, 0)
|
||||
EAC3_DECODER_OPTION_RE = re.compile(r"(?m)^\s*-([A-Za-z0-9_]+)\s+<")
|
||||
|
||||
|
||||
def resolve_output(source, requested=None, speaker_layout=None):
|
||||
def resolve_output(source, requested=None, speaker_layout=None, *, binaural=False):
|
||||
"""解析成品路径;未指定时使用项目内的 ``output`` 目录。"""
|
||||
source = Path(source)
|
||||
if requested is not None:
|
||||
target = Path(requested)
|
||||
elif speaker_layout is not None:
|
||||
target = DEFAULT_OUTPUT_DIR / f"{source.stem}.{speaker_layout}.wav"
|
||||
elif binaural:
|
||||
target = DEFAULT_OUTPUT_DIR / f"{source.stem}.binaural.wav"
|
||||
else:
|
||||
target = DEFAULT_OUTPUT_DIR / (source.stem + ".adm.wav")
|
||||
return target.expanduser().resolve()
|
||||
|
||||
|
||||
def _find_default_compiled_hrtf_cache():
|
||||
"""在默认 cache 目录寻找唯一的 .jochrtf;无文件返回 None,多个则报错。"""
|
||||
directory = DEFAULT_HRTF_CACHE_DIR
|
||||
if not directory.is_dir():
|
||||
return None
|
||||
candidates = sorted(directory.glob("*.jochrtf"))
|
||||
if not candidates:
|
||||
return None
|
||||
if len(candidates) > 1:
|
||||
listing = ", ".join(path.name for path in candidates[:8])
|
||||
raise ValueError(
|
||||
f"{directory} 下有多个 .jochrtf 缓存({listing}…),无法自动选择;"
|
||||
"请用 --compiled-hrtf-cache PATH 或 --sofa-hrtf PATH 显式指定")
|
||||
return candidates[0]
|
||||
|
||||
|
||||
def resolve_binaural_hrtf_input(args, *, required):
|
||||
"""解析 binaural 的 HRTF 输入。
|
||||
|
||||
无显式输入时按顺序回退:默认 HRTF/binaural.sofa → 默认 cache 目录下唯一的
|
||||
.jochrtf → 默认 HRTF/binaural.personalized_headphone → 报错。
|
||||
只校验路径,不做编译。
|
||||
"""
|
||||
sofa = args.sofa_hrtf
|
||||
compiled = args.compiled_hrtf_cache
|
||||
private = args.personalized_headphone
|
||||
cache_policy = args.hrtf_cache_policy
|
||||
cache_dir = args.hrtf_cache_dir
|
||||
radius = args.hrtf_radius_m
|
||||
|
||||
if compiled is not None and cache_policy is not None:
|
||||
raise ValueError("显式 .jochrtf 输入不能再指定 --hrtf-cache-policy")
|
||||
if compiled is not None and radius != 1.0:
|
||||
raise ValueError("显式 .jochrtf 输入不能再选择 SOFA radius shell")
|
||||
if private is not None and (cache_policy is not None or cache_dir is not None
|
||||
or radius != 1.0):
|
||||
raise ValueError(
|
||||
"Rosella 模型输入不能使用 "
|
||||
"--hrtf-cache-policy/--hrtf-cache-dir/--hrtf-radius-m")
|
||||
|
||||
if required and sofa is None and compiled is None and private is None:
|
||||
if DEFAULT_SOFA_HRTF.is_file():
|
||||
sofa = DEFAULT_SOFA_HRTF
|
||||
else:
|
||||
compiled = _find_default_compiled_hrtf_cache()
|
||||
if compiled is None and DEFAULT_PERSONALIZED_HEADPHONE.is_file():
|
||||
private = DEFAULT_PERSONALIZED_HEADPHONE
|
||||
|
||||
if sofa is None and compiled is None and private is None:
|
||||
if cache_policy is not None or cache_dir is not None or radius != 1.0:
|
||||
raise ValueError("HRTF cache/radius 选项需要 --sofa-hrtf")
|
||||
if required:
|
||||
raise ValueError(
|
||||
"--binaural 未找到 HRTF 输入:默认 "
|
||||
f"{DEFAULT_SOFA_HRTF}、{DEFAULT_PERSONALIZED_HEADPHONE} 与 "
|
||||
f"{DEFAULT_HRTF_CACHE_DIR} 下的 .jochrtf 缓存都不存在;请用 "
|
||||
"--sofa-hrtf PATH、--compiled-hrtf-cache PATH 或 "
|
||||
"--personalized-headphone PATH 指定")
|
||||
return None
|
||||
|
||||
if sofa is None and (cache_policy is not None or cache_dir is not None
|
||||
or radius != 1.0):
|
||||
raise ValueError("HRTF cache/radius 选项需要 --sofa-hrtf")
|
||||
effective_policy = "memory" if cache_policy is None else cache_policy
|
||||
if cache_dir is not None and (sofa is None or effective_policy != "disk"):
|
||||
raise ValueError("--hrtf-cache-dir 仅与 SOFA 的 disk cache policy 一起使用")
|
||||
if sofa is not None:
|
||||
return {
|
||||
"kind": "sofa",
|
||||
"path": resolve_sofa_hrtf(sofa),
|
||||
"cache_policy": effective_policy,
|
||||
"cache_dir": (DEFAULT_HRTF_CACHE_DIR if cache_dir is None else
|
||||
cache_dir.expanduser().resolve()),
|
||||
}
|
||||
if compiled is not None:
|
||||
return {
|
||||
"kind": "compiled_cache",
|
||||
"path": resolve_compiled_hrtf_cache(compiled),
|
||||
"cache_policy": None,
|
||||
"cache_dir": None,
|
||||
}
|
||||
if private is not None:
|
||||
return {
|
||||
"kind": "rosella",
|
||||
"path": resolve_personalized_headphone(private),
|
||||
"cache_policy": None,
|
||||
"cache_dir": None,
|
||||
}
|
||||
return None
|
||||
|
||||
|
||||
def executable(value, name):
|
||||
path = shutil.which(value) if value else None
|
||||
if path is None and value and Path(value).is_file():
|
||||
@@ -75,6 +187,53 @@ def timed_call(timings, name, function, *args, **kwargs):
|
||||
timings[name] = time.perf_counter() - started
|
||||
|
||||
|
||||
def probe_eac3_decoder_options(ffmpeg):
|
||||
"""读取 ``ffmpeg -h decoder=eac3`` 暴露的 AVOption 名。"""
|
||||
result = subprocess.run(
|
||||
[ffmpeg, "-hide_banner", "-h", "decoder=eac3"],
|
||||
stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
|
||||
text=True, encoding="utf-8", errors="replace")
|
||||
options = frozenset(EAC3_DECODER_OPTION_RE.findall(result.stdout or ""))
|
||||
# decoder 名不存在时 ffmpeg 依然返回 0,因此以“解析不到任何选项”为失败。
|
||||
if not options:
|
||||
raise RuntimeError(
|
||||
"无法读取 FFmpeg 的 eac3 解码器选项(ffmpeg -h decoder=eac3);"
|
||||
"需要带 E-AC-3 解码器的构建")
|
||||
return options
|
||||
|
||||
|
||||
def ffmpeg_version(ffmpeg):
|
||||
"""FFmpeg 版本字符串;探测失败返回空串,不影响渲染。"""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
[ffmpeg, "-hide_banner", "-version"],
|
||||
stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
|
||||
text=True, encoding="utf-8", errors="replace")
|
||||
except OSError:
|
||||
return ""
|
||||
lines = (result.stdout or "").splitlines()
|
||||
line = lines[0].strip() if lines else ""
|
||||
prefix = "ffmpeg version "
|
||||
return line[len(prefix):].strip() if line.startswith(prefix) else line
|
||||
|
||||
|
||||
def eac3_decode_options(drc_scale, target_level, available):
|
||||
"""构造 ``-i`` 之前的 E-AC-3 解码选项,返回 ``(argv, report 片段)``。"""
|
||||
if "drc_scale" not in available:
|
||||
raise RuntimeError(
|
||||
"FFmpeg 的 eac3 解码器缺少 -drc_scale,无法关闭码流 DRC")
|
||||
# -drc_scale 始终显式下发:0(全动态范围)不是 ffmpeg 的默认值。
|
||||
argv = ["-drc_scale", format(float(drc_scale), ".10g")]
|
||||
if target_level:
|
||||
if "target_level" not in available:
|
||||
raise RuntimeError(
|
||||
"FFmpeg 的 eac3 解码器不支持 -target_level;请升级 FFmpeg "
|
||||
"或去掉 --eac3-target-level")
|
||||
argv += ["-target_level", str(int(target_level))]
|
||||
applied = {"drc_scale": float(drc_scale), "target_level": int(target_level)}
|
||||
return argv, applied
|
||||
|
||||
|
||||
def extract_eac3(ffmpeg, source, target):
|
||||
if source.suffix.lower() in (".eac3", ".ec3"):
|
||||
return source
|
||||
@@ -84,10 +243,10 @@ def extract_eac3(ffmpeg, source, target):
|
||||
return target
|
||||
|
||||
|
||||
def decode_core(ffmpeg, eac3, target, duration_sec=None):
|
||||
def decode_core(ffmpeg, eac3, target, duration_sec=None, *, options=()):
|
||||
# 5.1(side) 的 f32le 顺序为 FL FR FC LFE SL SR;JOC 使用其中 0,1,2,4,5。
|
||||
command = [ffmpeg, "-hide_banner", "-loglevel", "error", "-y", "-i", str(eac3),
|
||||
"-map", "0:a:0", "-vn"]
|
||||
command = [ffmpeg, "-hide_banner", "-loglevel", "error", "-y", *options,
|
||||
"-i", str(eac3), "-map", "0:a:0", "-vn"]
|
||||
if duration_sec is not None:
|
||||
command.extend(["-t", f"{duration_sec:.9f}"])
|
||||
command.extend(["-ac", "6", "-ar", str(RATE),
|
||||
@@ -104,8 +263,8 @@ def sha256(path):
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def choose_speaker_output_format(requested_format, clip_action, peak, clipped_values,
|
||||
*, input_func=input, interactive=None):
|
||||
def choose_pcm_output_format(requested_format, clip_action, peak, clipped_values,
|
||||
*, input_func=input, interactive=None):
|
||||
"""Resolve int24 clipping interactively or through an explicit policy."""
|
||||
if requested_format != "int24" or clipped_values == 0:
|
||||
return requested_format
|
||||
@@ -145,6 +304,10 @@ def choose_speaker_output_format(requested_format, clip_action, peak, clipped_va
|
||||
raise ValueError(f"未知 clip action: {action}")
|
||||
|
||||
|
||||
# Backward-compatible public name used by existing tests and callers.
|
||||
choose_speaker_output_format = choose_pcm_output_format
|
||||
|
||||
|
||||
def resolve_metadata(args, eac3, temp_dir):
|
||||
if args.metadata_dir:
|
||||
directory = Path(args.metadata_dir).resolve()
|
||||
@@ -197,7 +360,9 @@ def create_renderer(backend, gain, native_library=None, native_threads=None):
|
||||
|
||||
def render(index, bed_path, frame_count, raw_path, gain, progress_every,
|
||||
backend="auto", native_library=None, native_threads=None, frame_sink=None,
|
||||
speaker_renderer=None, speaker_sink=None, speaker_metadata_offset=1473):
|
||||
speaker_renderer=None, speaker_sink=None, speaker_metadata_offset=1473,
|
||||
binaural_renderer=None, binaural_sink=None, binaural_metadata_offset=1473,
|
||||
raw_scale=1.0):
|
||||
values = np.memmap(bed_path, dtype=np.float32, mode="r")
|
||||
frame_width = FRAME_SAMPLES * 6
|
||||
if values.size % frame_width:
|
||||
@@ -215,6 +380,9 @@ def render(index, bed_path, frame_count, raw_path, gain, progress_every,
|
||||
raw_write_seconds = 0.0
|
||||
speaker_render_seconds = 0.0
|
||||
speaker_write_seconds = 0.0
|
||||
binaural_render_seconds = 0.0
|
||||
binaural_write_seconds = 0.0
|
||||
elapsed = 0.0
|
||||
try:
|
||||
for frame_number, row in enumerate(index.rows[:frame_count]):
|
||||
bed6 = np.asarray(bed[frame_number], dtype=np.float32)
|
||||
@@ -225,7 +393,8 @@ def render(index, bed_path, frame_count, raw_path, gain, progress_every,
|
||||
dsp_seconds += time.perf_counter() - stage
|
||||
if output is not None:
|
||||
stage = time.perf_counter()
|
||||
output[frame_number] = pcm16.T
|
||||
output[frame_number] = np.multiply(
|
||||
pcm16.T, np.float32(raw_scale), dtype=np.float32)
|
||||
raw_write_seconds += time.perf_counter() - stage
|
||||
if frame_sink is not None:
|
||||
stage = time.perf_counter()
|
||||
@@ -239,6 +408,22 @@ def render(index, bed_path, frame_count, raw_path, gain, progress_every,
|
||||
stage = time.perf_counter()
|
||||
speaker_sink.write_frame(speaker_pcm)
|
||||
speaker_write_seconds += time.perf_counter() - stage
|
||||
if binaural_renderer is not None:
|
||||
payload = subs.get(11)
|
||||
outer_offset = (
|
||||
index.subpayload_sample_offset(row, 11)
|
||||
if payload is not None and hasattr(index, "subpayload_sample_offset")
|
||||
else 0
|
||||
)
|
||||
stage = time.perf_counter()
|
||||
binaural_pcm = binaural_renderer.render_frame(
|
||||
pcm16.T, payload, binaural_metadata_offset,
|
||||
outer_sample_offset=outer_offset)
|
||||
binaural_render_seconds += time.perf_counter() - stage
|
||||
if len(binaural_pcm):
|
||||
stage = time.perf_counter()
|
||||
binaural_sink.write_frame(binaural_pcm)
|
||||
binaural_write_seconds += time.perf_counter() - stage
|
||||
done = frame_number + 1
|
||||
if done % progress_every == 0 or done == frame_count:
|
||||
elapsed = time.perf_counter() - started
|
||||
@@ -246,6 +431,14 @@ def render(index, bed_path, frame_count, raw_path, gain, progress_every,
|
||||
eta = (frame_count - done) / max(speed, 1e-9)
|
||||
print(f"[JOC:{backend_info['name']}] {done}/{frame_count} "
|
||||
f"{speed:.1f} frame/s ETA {eta:.1f}s", flush=True)
|
||||
if binaural_renderer is not None:
|
||||
stage = time.perf_counter()
|
||||
binaural_tail = binaural_renderer.finish()
|
||||
binaural_render_seconds += time.perf_counter() - stage
|
||||
if len(binaural_tail):
|
||||
stage = time.perf_counter()
|
||||
binaural_sink.write_frame(binaural_tail)
|
||||
binaural_write_seconds += time.perf_counter() - stage
|
||||
if output is not None:
|
||||
output.flush()
|
||||
elapsed = time.perf_counter() - started
|
||||
@@ -256,6 +449,9 @@ def render(index, bed_path, frame_count, raw_path, gain, progress_every,
|
||||
close = getattr(speaker_renderer, "close", None)
|
||||
if close is not None:
|
||||
close()
|
||||
close = getattr(binaural_renderer, "close", None)
|
||||
if close is not None:
|
||||
close()
|
||||
breakdown = {
|
||||
"pipeline_wall_seconds": elapsed,
|
||||
"dsp_and_joc_parse_seconds": dsp_seconds,
|
||||
@@ -263,36 +459,86 @@ def render(index, bed_path, frame_count, raw_path, gain, progress_every,
|
||||
"raw_float_write_seconds": raw_write_seconds,
|
||||
"speaker_render_seconds": speaker_render_seconds,
|
||||
"speaker_spool_write_seconds": speaker_write_seconds,
|
||||
"binaural_render_seconds": binaural_render_seconds,
|
||||
"binaural_spool_write_seconds": binaural_write_seconds,
|
||||
}
|
||||
return dsp_seconds, backend_info, breakdown
|
||||
|
||||
|
||||
def build_parser():
|
||||
parser = argparse.ArgumentParser(
|
||||
description="JustOneCacophony (JOC):E-AC-3 JOC → 25ch ADM BWF 或扬声器 WAV")
|
||||
description=("JustOneCacophony (JOC):E-AC-3 JOC → 25ch ADM BWF、"
|
||||
"扬声器 WAV 或公开 SOFA 双耳 WAV"))
|
||||
parser.add_argument("input", type=Path, help="输入 .m4a/.eac3/.ec3")
|
||||
parser.add_argument("-o", "--output", type=Path, help="输出文件;默认按模式和布局命名")
|
||||
parser.add_argument("--speaker-output", type=Path,
|
||||
help="扬声器 WAV 路径;仅与 --speaker-layout 一起使用")
|
||||
parser.add_argument("--speaker-layout", choices=SPEAKER_LAYOUT_CHOICES,
|
||||
help="直接扬声器渲染布局,例如 2.0、5.1、7.1.2")
|
||||
parser.add_argument("--binaural-output", type=Path,
|
||||
help="双耳 WAV 路径;仅与 --binaural 一起使用")
|
||||
direct_mode = parser.add_mutually_exclusive_group()
|
||||
direct_mode.add_argument("--speaker-layout", choices=SPEAKER_LAYOUT_CHOICES,
|
||||
help="直接扬声器渲染布局,例如 2.0、5.1、7.1.2")
|
||||
direct_mode.add_argument("--binaural", action="store_true",
|
||||
help="直接 SOFA 双耳渲染;不生成临时 ADM BWF")
|
||||
parser.add_argument("--speaker-format", choices=("float32", "int24"), default="float32",
|
||||
help="扬声器 WAV 格式,默认 float32")
|
||||
parser.add_argument("--binaural-format", choices=("float32", "int24"), default="float32",
|
||||
help="双耳 WAV 格式,默认 float32")
|
||||
parser.add_argument("--clip-action", choices=("ask", "continue", "float32", "abort"),
|
||||
default="ask",
|
||||
help="int24 削波处理:交互询问、继续截断、改 float32 或中止")
|
||||
parser.add_argument("--speaker-metadata-offset", type=int, default=1473,
|
||||
help="扬声器渲染 metadata 相对帧偏移,默认 1473 samples")
|
||||
parser.add_argument("--binaural-mode", choices=("off", "near", "mid", "far"),
|
||||
default="mid",
|
||||
help="双耳渲染模式,默认 mid(人为指定的渲染提示,非码流 "
|
||||
"原始元数据);直接双耳渲染与 ADM BWF 的 DBMD 提示共用。"
|
||||
"off 仅用于 ADM BWF:关闭 DBMD 双耳提示(编码 0)")
|
||||
hrtf_input = parser.add_mutually_exclusive_group()
|
||||
hrtf_input.add_argument(
|
||||
"--sofa-hrtf", type=Path,
|
||||
help="SimpleFreeFieldHRIR SOFA;缺省时依次尝试 HRTF/binaural.sofa、"
|
||||
"output/hrtf-cache 下唯一的 .jochrtf、"
|
||||
"HRTF/binaural.personalized_headphone,均无则报错")
|
||||
hrtf_input.add_argument(
|
||||
"--compiled-hrtf-cache", type=Path,
|
||||
help="高级入口:显式读取 JOC .jochrtf compiled cache")
|
||||
hrtf_input.add_argument(
|
||||
"--personalized-headphone", type=Path, nargs="?",
|
||||
const=DEFAULT_PERSONALIZED_HEADPHONE,
|
||||
help="Rosella .personalized_headphone 模型;不带路径时默认 "
|
||||
"HRTF/binaural.personalized_headphone")
|
||||
parser.add_argument(
|
||||
"--hrtf-cache-policy", choices=("none", "memory", "disk"), default=None,
|
||||
help="SOFA 编译缓存;默认 memory,disk 写入可删除的 .jochrtf")
|
||||
parser.add_argument(
|
||||
"--hrtf-cache-dir", type=Path,
|
||||
help="disk cache 目录;默认 output/hrtf-cache")
|
||||
parser.add_argument(
|
||||
"--hrtf-radius-m", type=float, default=1.0,
|
||||
help="选择最近的 SOFA measurement-radius shell,默认 1.0 m")
|
||||
parser.add_argument("--binaural-tail-seconds", type=float, default=5.0,
|
||||
help="双耳 room/filterbank flush 上限,默认 5 秒")
|
||||
parser.add_argument("--binaural-tail-threshold", type=float, default=1.0e-8,
|
||||
help="双耳尾声裁切阈值,默认 1e-8;主体至少保留原时长")
|
||||
parser.add_argument("--binaural-chunk-frames", type=int, default=64,
|
||||
help="双耳内部批处理 E-AC-3 帧数,默认 64")
|
||||
parser.add_argument("--gain-db", type=float, default=0.0,
|
||||
help="成品增益 dB,默认 0(float32 系数 1.0)")
|
||||
help="成品增益 dB,默认 0;双耳路径以 float64 应用")
|
||||
parser.add_argument("--duration", type=float, help="只处理开头指定秒数")
|
||||
parser.add_argument("--object-delay-samples", type=int, default=640,
|
||||
help="可选的对象 PCM/OAMD 时间补偿,默认 640 samples")
|
||||
parser.add_argument("--object-delay-samples", type=int, default=1473,
|
||||
help="对象 PCM/OAMD 时间补偿;ADM 与双耳默认 1473 samples")
|
||||
parser.add_argument("--trajectory-mode", choices=("compact", "dense64"), default="compact",
|
||||
help="对象轨迹表示;compact 用长线性插值压缩 AXML,dense64 保留逐 64-sample 块")
|
||||
help="ADM 对象轨迹表示;直接双耳路径不序列化 AXML")
|
||||
parser.add_argument("--ffmpeg", default=os.environ.get("FFMPEG", "ffmpeg"))
|
||||
parser.add_argument("--eac3-drc-scale", type=float, default=0.0,
|
||||
help="E-AC-3 解码器 -drc_scale:0=关闭码流 dynrng(全动态范围),"
|
||||
"1=码流作者意图,>1 非对称;默认 0")
|
||||
parser.add_argument("--eac3-target-level", type=int, default=0,
|
||||
help="E-AC-3 解码器 -target_level:按码流 dialnorm 归一化电平,"
|
||||
"增益约 target_level - dialnorm dB;0=不施加,默认 0")
|
||||
parser.add_argument("--backend", choices=("auto", "native", "python"), default="auto",
|
||||
help="DSP 后端;auto 优先 C++,不可用时回退 Python")
|
||||
help="JOC/扬声器 DSP 后端;SOFA 双耳 DSP 当前使用 Python")
|
||||
parser.add_argument("--native-library", type=Path,
|
||||
help="显式指定原生库;默认从单层 lib 目录选择当前平台文件")
|
||||
parser.add_argument("--native-threads", type=int,
|
||||
@@ -326,25 +572,81 @@ def main(argv=None):
|
||||
if not source.is_file():
|
||||
raise FileNotFoundError(source)
|
||||
speaker_mode = args.speaker_layout is not None
|
||||
binaural_mode = bool(args.binaural)
|
||||
binaural_render_mode = args.binaural_mode
|
||||
if binaural_mode and binaural_render_mode == "off":
|
||||
raise ValueError(
|
||||
"--binaural-mode off 仅用于 ADM BWF 输出(关闭 DBMD 双耳提示);"
|
||||
"直接双耳渲染请使用 near/mid/far")
|
||||
if args.speaker_output is not None and not speaker_mode:
|
||||
raise ValueError("--speaker-output 必须与 --speaker-layout 一起使用")
|
||||
if args.output is not None and args.speaker_output is not None:
|
||||
raise ValueError("-o/--output 与 --speaker-output 不能同时使用")
|
||||
if args.binaural_output is not None and not binaural_mode:
|
||||
raise ValueError("--binaural-output 必须与 --binaural 一起使用")
|
||||
specific_outputs = [value for value in (args.speaker_output, args.binaural_output)
|
||||
if value is not None]
|
||||
if args.output is not None and specific_outputs:
|
||||
raise ValueError("-o/--output 与 --speaker-output/--binaural-output 不能同时使用")
|
||||
if len(specific_outputs) > 1:
|
||||
raise ValueError("--speaker-output 与 --binaural-output 不能同时使用")
|
||||
if args.speaker_metadata_offset < 0:
|
||||
raise ValueError("speaker-metadata-offset 不能为负数")
|
||||
hrtf_options_used = any((
|
||||
args.sofa_hrtf is not None,
|
||||
args.compiled_hrtf_cache is not None,
|
||||
args.personalized_headphone is not None,
|
||||
args.hrtf_cache_policy is not None,
|
||||
args.hrtf_cache_dir is not None,
|
||||
args.hrtf_radius_m != 1.0,
|
||||
))
|
||||
if hrtf_options_used and not binaural_mode:
|
||||
raise ValueError("SOFA/HRTF 选项仅与 --binaural 一起使用")
|
||||
if (not math.isfinite(args.binaural_tail_seconds)
|
||||
or args.binaural_tail_seconds < 0):
|
||||
raise ValueError("binaural-tail-seconds 必须是非负有限值")
|
||||
if (not math.isfinite(args.binaural_tail_threshold)
|
||||
or args.binaural_tail_threshold < 0):
|
||||
raise ValueError("binaural-tail-threshold 必须是非负有限值")
|
||||
if args.binaural_chunk_frames <= 0:
|
||||
raise ValueError("binaural-chunk-frames 必须大于 0")
|
||||
if not math.isfinite(args.hrtf_radius_m) or args.hrtf_radius_m <= 0.0:
|
||||
raise ValueError("hrtf-radius-m 必须是正有限值")
|
||||
requested_output = (args.speaker_output if args.speaker_output is not None
|
||||
else args.binaural_output if args.binaural_output is not None
|
||||
else args.output)
|
||||
output = resolve_output(
|
||||
source, requested_output, args.speaker_layout if speaker_mode else None)
|
||||
source, requested_output, args.speaker_layout if speaker_mode else None,
|
||||
binaural=binaural_mode)
|
||||
output.parent.mkdir(parents=True, exist_ok=True)
|
||||
if args.duration is not None and args.duration <= 0:
|
||||
raise ValueError("duration 必须大于 0")
|
||||
if args.object_delay_samples < 0:
|
||||
raise ValueError("object-delay-samples 不能为负数")
|
||||
gain = np.float32(10.0 ** (args.gain_db / 20.0))
|
||||
if not np.isfinite(gain):
|
||||
raise ValueError("gain-db 超出 float32 范围")
|
||||
gain_float64 = 10.0 ** (args.gain_db / 20.0)
|
||||
gain = np.float32(gain_float64)
|
||||
if not math.isfinite(gain_float64) or not np.isfinite(gain):
|
||||
raise ValueError("gain-db 超出支持范围")
|
||||
if (not math.isfinite(args.eac3_drc_scale)
|
||||
or not 0.0 <= args.eac3_drc_scale <= EAC3_DRC_SCALE_MAX):
|
||||
raise ValueError(f"eac3-drc-scale 必须在 0..{EAC3_DRC_SCALE_MAX:g} 之间")
|
||||
if not (EAC3_TARGET_LEVEL_RANGE[0] <= args.eac3_target_level
|
||||
<= EAC3_TARGET_LEVEL_RANGE[1]):
|
||||
raise ValueError("eac3-target-level 必须在 -31..0 之间")
|
||||
binaural_hrtf_input = resolve_binaural_hrtf_input(
|
||||
args, required=binaural_mode and not args.metadata_only)
|
||||
ffmpeg = executable(args.ffmpeg, "FFmpeg")
|
||||
decode_options = ()
|
||||
decode_option_info = None
|
||||
if not args.metadata_only:
|
||||
available = probe_eac3_decoder_options(ffmpeg)
|
||||
decode_options, applied = eac3_decode_options(
|
||||
args.eac3_drc_scale, args.eac3_target_level, available)
|
||||
decode_option_info = {
|
||||
"version": ffmpeg_version(ffmpeg),
|
||||
"eac3_decode_options": applied,
|
||||
}
|
||||
print(f"[decode] ffmpeg {decode_option_info['version']} "
|
||||
f"drc_scale={args.eac3_drc_scale:g} "
|
||||
f"target_level={args.eac3_target_level}", flush=True)
|
||||
|
||||
total_started = time.perf_counter()
|
||||
timings = {}
|
||||
@@ -380,7 +682,8 @@ def main(argv=None):
|
||||
|
||||
bed_path = timed_call(
|
||||
timings, "decode_core", decode_core,
|
||||
ffmpeg, eac3, temp_dir / "core51_f32le.raw", duration_sec)
|
||||
ffmpeg, eac3, temp_dir / "core51_f32le.raw", duration_sec,
|
||||
options=decode_options)
|
||||
raw_path = (output.with_name(output.name + ".objects16.f32le")
|
||||
if args.keep_raw else None)
|
||||
master = None
|
||||
@@ -388,7 +691,13 @@ def main(argv=None):
|
||||
speaker_wav_info = None
|
||||
speaker_clip_info = None
|
||||
speaker_actual_format = None
|
||||
binaural_backend_info = None
|
||||
binaural_hrtf_report = None
|
||||
binaural_wav_info = None
|
||||
binaural_clip_info = None
|
||||
binaural_actual_format = None
|
||||
if speaker_mode:
|
||||
timings["create_binaural_renderer"] = 0.0
|
||||
layout = get_speaker_layout(args.speaker_layout)
|
||||
speaker_name = speaker_layout_display_name(layout)
|
||||
speaker_decoder, speaker_backend_info = create_speaker_renderer(
|
||||
@@ -410,11 +719,11 @@ def main(argv=None):
|
||||
args.native_threads, None, speaker_decoder, spool,
|
||||
args.speaker_metadata_offset)
|
||||
spool.finalize()
|
||||
speaker_actual_format = choose_speaker_output_format(
|
||||
speaker_actual_format = choose_pcm_output_format(
|
||||
args.speaker_format, args.clip_action, spool.peak,
|
||||
spool.clipped_values)
|
||||
speaker_wav_info = timed_call(
|
||||
timings, "write_speaker_wav", write_speaker_wav,
|
||||
timings, "write_speaker_wav", write_pcm_wav,
|
||||
output, spool.values, speaker_actual_format, rate=RATE)
|
||||
speaker_clip_info = {
|
||||
"peak": spool.peak,
|
||||
@@ -430,8 +739,158 @@ def main(argv=None):
|
||||
timings["validate_adm"] = 0.0
|
||||
info = (f"speaker layout={speaker_name}, format={speaker_actual_format}, "
|
||||
f"peak={speaker_clip_info['peak']:.9g}")
|
||||
elif binaural_mode:
|
||||
hrtf_source = binaural_hrtf_input
|
||||
common_options = {
|
||||
"mode": binaural_render_mode,
|
||||
"object_delay_samples": args.object_delay_samples,
|
||||
"tail_seconds": args.binaural_tail_seconds,
|
||||
"output_gain": gain_float64,
|
||||
"chunk_frames": args.binaural_chunk_frames,
|
||||
}
|
||||
if hrtf_source["kind"] == "sofa":
|
||||
binaural_decoder = None
|
||||
if args.backend in ("auto", "native"):
|
||||
try:
|
||||
from sofa_native_backend import create_native_sofa_renderer
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
create_native_sofa_renderer,
|
||||
hrtf_source["path"],
|
||||
cache_policy=hrtf_source["cache_policy"],
|
||||
cache_dir=hrtf_source["cache_dir"],
|
||||
shell_radius_m=args.hrtf_radius_m,
|
||||
**common_options)
|
||||
except (ImportError, OSError, RuntimeError, ValueError) as exc:
|
||||
print(
|
||||
f"[binaural] native SOFA backend unavailable "
|
||||
f"({exc.__class__.__name__}: {exc}); "
|
||||
f"falling back to Python", flush=True)
|
||||
binaural_decoder = None
|
||||
if binaural_decoder is None:
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
SofaBinauralRenderer.from_sofa,
|
||||
hrtf_source["path"],
|
||||
cache_policy=hrtf_source["cache_policy"],
|
||||
cache_dir=hrtf_source["cache_dir"],
|
||||
shell_radius_m=args.hrtf_radius_m,
|
||||
**common_options)
|
||||
elif hrtf_source["kind"] == "rosella":
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
RosellaBinauralRenderer,
|
||||
hrtf_source["path"],
|
||||
mode=binaural_render_mode,
|
||||
object_delay_samples=args.object_delay_samples,
|
||||
tail_seconds=args.binaural_tail_seconds,
|
||||
output_gain=gain_float64,
|
||||
chunk_frames=args.binaural_chunk_frames,
|
||||
backend=args.backend,
|
||||
native_library=args.native_library)
|
||||
else:
|
||||
binaural_decoder = None
|
||||
if args.backend in ("auto", "native"):
|
||||
try:
|
||||
from sofa_native_backend import (
|
||||
create_native_compiled_cache_renderer)
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
create_native_compiled_cache_renderer,
|
||||
hrtf_source["path"],
|
||||
**common_options)
|
||||
except (ImportError, OSError, RuntimeError, ValueError) as exc:
|
||||
print(
|
||||
f"[binaural] native SOFA backend unavailable "
|
||||
f"({exc.__class__.__name__}: {exc}); "
|
||||
f"falling back to Python", flush=True)
|
||||
binaural_decoder = None
|
||||
if binaural_decoder is None:
|
||||
binaural_decoder = timed_call(
|
||||
timings, "create_binaural_renderer",
|
||||
SofaBinauralRenderer.from_compiled_cache,
|
||||
hrtf_source["path"],
|
||||
**common_options)
|
||||
print(
|
||||
f"[binaural] mode={binaural_render_mode} "
|
||||
f"backend={binaural_decoder.dsp_backend} "
|
||||
f"precision=float64/complex128 "
|
||||
f"hrtf={hrtf_source['kind']}:{hrtf_source['path']}", flush=True)
|
||||
if hrtf_source["kind"] == "rosella":
|
||||
flush_samples = math.ceil(
|
||||
(args.binaural_tail_seconds * RATE
|
||||
+ ROSSELLA_LATENCY_SAMPLES + ROSSELLA_BLOCK_SAMPLES)
|
||||
/ ROSSELLA_BLOCK_SAMPLES) * ROSSELLA_BLOCK_SAMPLES
|
||||
spool_capacity = frame_count * FRAME_SAMPLES + flush_samples
|
||||
else:
|
||||
spool_capacity = (
|
||||
frame_count * FRAME_SAMPLES
|
||||
+ binaural_decoder.finish_capacity_samples)
|
||||
spool = BinauralPcmSpool(
|
||||
temp_dir / "binaural_interleaved_f64.raw",
|
||||
spool_capacity,
|
||||
tail_threshold=args.binaural_tail_threshold)
|
||||
try:
|
||||
render_seconds, renderer_backend, render_breakdown = timed_call(
|
||||
timings, "render_and_stream", variant_call,
|
||||
output, source, render, index, bed_path, frame_count, raw_path,
|
||||
np.float32(1.0), max(1, args.progress_every),
|
||||
backend=args.backend, native_library=args.native_library,
|
||||
native_threads=args.native_threads,
|
||||
binaural_renderer=binaural_decoder, binaural_sink=spool,
|
||||
binaural_metadata_offset=args.object_delay_samples, raw_scale=gain)
|
||||
spool.finalize(minimum_samples=frame_count * FRAME_SAMPLES)
|
||||
binaural_actual_format = choose_pcm_output_format(
|
||||
args.binaural_format, args.clip_action, spool.peak,
|
||||
spool.clipped_values)
|
||||
binaural_wav_info = timed_call(
|
||||
timings, "write_binaural_wav", write_pcm_wav,
|
||||
output, spool.values, binaural_actual_format, rate=RATE)
|
||||
binaural_clip_info = {
|
||||
"peak": spool.peak,
|
||||
"over_unity_values": spool.clipped_values,
|
||||
"requested_format": args.binaural_format,
|
||||
"actual_format": binaural_actual_format,
|
||||
"clip_action": args.clip_action,
|
||||
"tail_threshold": args.binaural_tail_threshold,
|
||||
"source_samples": frame_count * FRAME_SAMPLES,
|
||||
"kept_samples": spool.sample_count,
|
||||
}
|
||||
binaural_backend_info = binaural_decoder.backend_info
|
||||
if hrtf_source["kind"] == "rosella":
|
||||
binaural_hrtf_report = {
|
||||
"input_kind": "rosella",
|
||||
"input_path": str(binaural_decoder.model_path.resolve()),
|
||||
"model_coefficient_sha256": (
|
||||
binaural_decoder.model.coefficient_sha256),
|
||||
"cache_policy": None,
|
||||
}
|
||||
else:
|
||||
binaural_hrtf_report = {
|
||||
"input_kind": binaural_backend_info["hrtf_input_kind"],
|
||||
"input_path": binaural_backend_info["hrtf_input_path"],
|
||||
"source_sha256": (
|
||||
binaural_backend_info["field"]["source_sha256"]),
|
||||
"cache_policy": binaural_backend_info["cache_policy"],
|
||||
"cache_key": (
|
||||
binaural_backend_info["field"]["cache_key"]),
|
||||
"format_version": (
|
||||
binaural_backend_info["field"]["format_version"]),
|
||||
}
|
||||
finally:
|
||||
spool.close()
|
||||
timings["build_adm_tracks"] = 0.0
|
||||
timings["finalize_adm"] = 0.0
|
||||
timings["validate_adm"] = 0.0
|
||||
info = (f"binaural mode={binaural_render_mode}, "
|
||||
f"format={binaural_actual_format}, "
|
||||
f"peak={binaural_clip_info['peak']:.9g}, "
|
||||
f"samples={binaural_clip_info['kept_samples']}")
|
||||
else:
|
||||
master = adm_assemble.StreamingMaster(output, duration_sec, rate=RATE)
|
||||
timings["create_binaural_renderer"] = 0.0
|
||||
master = adm_assemble.StreamingMaster(
|
||||
output, duration_sec, rate=RATE,
|
||||
joc_binaural_mode=adm_atmos.JOC_BINAURAL_MODES[args.binaural_mode])
|
||||
try:
|
||||
render_seconds, renderer_backend, render_breakdown = timed_call(
|
||||
timings, "render_and_stream", variant_call,
|
||||
@@ -461,10 +920,11 @@ def main(argv=None):
|
||||
else:
|
||||
output_sha = timed_call(timings, "sha256", sha256, output)
|
||||
total_seconds = time.perf_counter() - total_started
|
||||
mode_name = "speaker" if speaker_mode else "binaural" if binaural_mode else "adm"
|
||||
report = {
|
||||
"input": str(source),
|
||||
"output": str(output),
|
||||
"mode": "speaker" if speaker_mode else "adm",
|
||||
"mode": mode_name,
|
||||
"metadata": str(metadata_json) if metadata_json is not None else None,
|
||||
"metadata_backend": metadata_backend,
|
||||
"metadata_cache": str(metadata_cache_dir) if metadata_cache_dir is not None else None,
|
||||
@@ -472,8 +932,12 @@ def main(argv=None):
|
||||
"duration_sec": duration_sec,
|
||||
"gain_db": args.gain_db,
|
||||
"gain_float32": float(gain),
|
||||
"object_delay_samples": None if speaker_mode else args.object_delay_samples,
|
||||
"trajectory_mode": None if speaker_mode else args.trajectory_mode,
|
||||
"gain_float64": float(gain_float64),
|
||||
"object_delay_samples": (None if speaker_mode else args.object_delay_samples),
|
||||
"trajectory_mode": args.trajectory_mode if mode_name == "adm" else None,
|
||||
"binaural_mode_value": (
|
||||
adm_atmos.JOC_BINAURAL_MODES[args.binaural_mode]
|
||||
if mode_name == "adm" else None),
|
||||
"render_seconds": render_seconds,
|
||||
"render_breakdown": render_breakdown,
|
||||
"renderer_backend": renderer_backend,
|
||||
@@ -482,15 +946,25 @@ def main(argv=None):
|
||||
"speaker_metadata_offset": args.speaker_metadata_offset if speaker_mode else None,
|
||||
"speaker_clip": speaker_clip_info,
|
||||
"speaker_wav": speaker_wav_info,
|
||||
"streaming_adm": not speaker_mode,
|
||||
"binaural_renderer_backend": binaural_backend_info,
|
||||
"binaural_mode": (
|
||||
args.binaural_mode if (binaural_mode or mode_name == "adm") else None),
|
||||
"binaural_hrtf": (
|
||||
binaural_hrtf_report if binaural_backend_info else None),
|
||||
"binaural_clip": binaural_clip_info,
|
||||
"binaural_wav": binaural_wav_info,
|
||||
"output_clip": speaker_clip_info if speaker_mode else binaural_clip_info,
|
||||
"output_wav": speaker_wav_info if speaker_mode else binaural_wav_info,
|
||||
"streaming_adm": mode_name == "adm",
|
||||
"kept_raw": str(raw_path) if raw_path is not None else None,
|
||||
"timings": timings,
|
||||
"total_seconds": total_seconds,
|
||||
"adm_validation": None if speaker_mode else info,
|
||||
"adm_validation": info if mode_name == "adm" else None,
|
||||
"adm_metadata": getattr(master, "metadata_info", None) if master is not None else None,
|
||||
"sha256": output_sha,
|
||||
"python": platform.python_version(),
|
||||
"numpy": np.__version__,
|
||||
"ffmpeg": decode_option_info,
|
||||
}
|
||||
report_path = Path(str(output) + ".report.json")
|
||||
report_path.write_text(json.dumps(report, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
@@ -504,6 +978,11 @@ def main(argv=None):
|
||||
f"speaker={render_breakdown['speaker_render_seconds']:.2f}s "
|
||||
f"pipeline={render_breakdown['pipeline_wall_seconds']:.2f}s "
|
||||
f"total={report['total_seconds']:.2f}s")
|
||||
elif binaural_mode:
|
||||
print(f"[time] JOC-DSP={render_seconds:.2f}s ({renderer_backend['name']}) "
|
||||
f"binaural={render_breakdown['binaural_render_seconds']:.2f}s "
|
||||
f"pipeline={render_breakdown['pipeline_wall_seconds']:.2f}s "
|
||||
f"total={report['total_seconds']:.2f}s")
|
||||
else:
|
||||
print(f"[time] DSP={render_seconds:.2f}s ({renderer_backend['name']}) "
|
||||
f"render+ADM-stream={render_breakdown['pipeline_wall_seconds']:.2f}s "
|
||||
|
||||
@@ -8,6 +8,8 @@ find_package(Threads REQUIRED)
|
||||
add_library(eac3joc_core SHARED
|
||||
src/eac3joc_core.cpp
|
||||
src/speaker_renderer.cpp
|
||||
src/binaural_renderer.cpp
|
||||
src/sofa_binaural_renderer.cpp
|
||||
src/joc_huffman_tables.h
|
||||
src/qmf_tables.h
|
||||
src/speaker_layouts.h
|
||||
|
||||
@@ -29,11 +29,17 @@ enum {
|
||||
EJOC_MAX_DPOINTS = 2,
|
||||
EJOC_MAX_PARAMETER_BANDS = 23,
|
||||
EJOC_SPEAKER_BLOCK_SAMPLES = 32,
|
||||
EJOC_SPEAKER_COORDINATES = 3
|
||||
EJOC_SPEAKER_COORDINATES = 3,
|
||||
EJOC_BINAURAL_BLOCK_SAMPLES = 512,
|
||||
EJOC_BINAURAL_INPUT_CHANNELS = 16,
|
||||
EJOC_BINAURAL_OUTPUT_CHANNELS = 2,
|
||||
EJOC_BINAURAL_QMF_BANDS = 64,
|
||||
EJOC_BINAURAL_HYBRID_BANDS = 77
|
||||
};
|
||||
|
||||
typedef void* ejoc_renderer_handle;
|
||||
typedef void* ejoc_speaker_renderer_handle;
|
||||
typedef void* ejoc_binaural_renderer_handle;
|
||||
|
||||
/*
|
||||
Fixed array layouts used by ejoc_renderer_process():
|
||||
@@ -47,8 +53,9 @@ Fixed array layouts used by ejoc_renderer_process():
|
||||
output16 [16][1536]
|
||||
|
||||
Only objects selected by object_mask are read from the descriptor arrays.
|
||||
Sparse JOC must be rejected by the caller; this ABI accepts already dequantized
|
||||
dense matrix coefficients.
|
||||
dq carries already dequantized matrix coefficients in double precision; the
|
||||
caller performs the JOC bitstream differential decoding for both dense and
|
||||
sparse objects, so this ABI is identical for both syntaxes.
|
||||
*/
|
||||
|
||||
EJOC_API uint32_t EJOC_CALL ejoc_abi_version(void);
|
||||
@@ -114,6 +121,115 @@ EJOC_API int EJOC_CALL ejoc_speaker_renderer_process(
|
||||
const double* object_gains,
|
||||
double* output_interleaved);
|
||||
|
||||
EJOC_API ejoc_binaural_renderer_handle EJOC_CALL ejoc_binaural_renderer_create(void);
|
||||
EJOC_API void EJOC_CALL ejoc_binaural_renderer_destroy(ejoc_binaural_renderer_handle handle);
|
||||
EJOC_API int EJOC_CALL ejoc_binaural_renderer_reset(ejoc_binaural_renderer_handle handle);
|
||||
EJOC_API const char* EJOC_CALL ejoc_binaural_renderer_last_error(
|
||||
ejoc_binaural_renderer_handle handle);
|
||||
EJOC_API int EJOC_CALL ejoc_binaural_renderer_configure_kernels(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
const double* qmf_analysis,
|
||||
const double* hybrid_low,
|
||||
const int16_t* hybrid_indices,
|
||||
const double* hybrid_values,
|
||||
uint32_t hybrid_count,
|
||||
const double* qmf_basis,
|
||||
const double* qmf_taps);
|
||||
EJOC_API int EJOC_CALL ejoc_binaural_renderer_configure_room(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
uint32_t bands,
|
||||
uint32_t allpass_count,
|
||||
const uint32_t* allpass_delays,
|
||||
const double* allpass_gains,
|
||||
const uint32_t* fdn_delays,
|
||||
const double* fdn_matrix,
|
||||
uint32_t output_tap_delay,
|
||||
const double* feedback_complex,
|
||||
const double* output_taps,
|
||||
const double* output_complex,
|
||||
uint32_t extra_count,
|
||||
const uint32_t* extra_delays,
|
||||
const double* extra_fields_complex,
|
||||
const double* extra_matrices);
|
||||
EJOC_API int EJOC_CALL ejoc_binaural_renderer_process(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
const double* input16_interleaved,
|
||||
const double* gains_complex,
|
||||
const double* room_sends,
|
||||
double output_gain,
|
||||
double* output_stereo_interleaved);
|
||||
|
||||
/*
|
||||
Native SOFA binaural renderer.
|
||||
|
||||
The handle owns the complete runtime: 64-QMF/77-hybrid analysis and synthesis,
|
||||
fifth-order ACN/N3D real spherical-harmonic direction-field evaluation,
|
||||
per-object whole-QMF-slot delay histories, six first-order image-source early
|
||||
reflections, the shared unitary-FDN late room, the 120-180 Hz LFE low-pass and
|
||||
the 961-sample latency compensation. The caller configures the filterbank
|
||||
tables, the compiled HRTF field and the room constants once, then per 512-sample
|
||||
block updates every source with ejoc_sofa_binaural_set_source() and calls
|
||||
ejoc_sofa_binaural_process(). Process returns the number of trimmed stereo
|
||||
samples written; the first 961 processed samples across calls are discarded.
|
||||
*/
|
||||
typedef void* ejoc_sofa_binaural_handle;
|
||||
|
||||
EJOC_API ejoc_sofa_binaural_handle EJOC_CALL ejoc_sofa_binaural_create(void);
|
||||
EJOC_API void EJOC_CALL ejoc_sofa_binaural_destroy(ejoc_sofa_binaural_handle handle);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_reset(ejoc_sofa_binaural_handle handle);
|
||||
EJOC_API const char* EJOC_CALL ejoc_sofa_binaural_last_error(
|
||||
ejoc_sofa_binaural_handle handle);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_configure_kernels(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
const double* qmf_analysis,
|
||||
const double* hybrid_low,
|
||||
const int16_t* hybrid_indices,
|
||||
const double* hybrid_values,
|
||||
uint32_t hybrid_count,
|
||||
const double* qmf_basis,
|
||||
const double* qmf_taps);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_configure_field(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
const double* coefficients,
|
||||
const double* delay_coefficients,
|
||||
const double* delay_bounds,
|
||||
const double* band_centers,
|
||||
double measurement_radius_m);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_configure_room(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
const double* room_dims,
|
||||
const double* listener_pos,
|
||||
const double* wall_gains,
|
||||
double speed_of_sound,
|
||||
const uint32_t* fdn_delays,
|
||||
const double* fdn_feedback,
|
||||
double damping,
|
||||
double fdn_output_gain,
|
||||
const uint32_t* allpass_delays,
|
||||
const double* allpass_gains,
|
||||
uint32_t enable_early_reflections,
|
||||
uint32_t enable_late_room);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_set_source(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
uint32_t source,
|
||||
const double* position_adm,
|
||||
uint32_t profile,
|
||||
double gain,
|
||||
uint32_t enabled,
|
||||
uint32_t special_lfe,
|
||||
uint32_t fade);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_process(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
const double* input16_interleaved,
|
||||
uint32_t sample_count,
|
||||
double output_gain,
|
||||
double* output_stereo_interleaved);
|
||||
EJOC_API int EJOC_CALL ejoc_sofa_binaural_finish(
|
||||
ejoc_sofa_binaural_handle handle,
|
||||
uint32_t flush_samples,
|
||||
double* output_stereo_interleaved,
|
||||
uint32_t capacity);
|
||||
|
||||
#ifdef __cplusplus
|
||||
}
|
||||
#endif
|
||||
|
||||
@@ -0,0 +1,664 @@
|
||||
#define EJOC_BUILD_DLL
|
||||
#include "eac3joc_core.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <array>
|
||||
#include <cmath>
|
||||
#include <cstdint>
|
||||
#include <cstdio>
|
||||
#include <cstring>
|
||||
#include <new>
|
||||
#include <vector>
|
||||
|
||||
namespace ejoc::binaural {
|
||||
|
||||
struct Complex {
|
||||
double re;
|
||||
double im;
|
||||
};
|
||||
|
||||
inline Complex add(Complex a, Complex b) noexcept {
|
||||
return {a.re + b.re, a.im + b.im};
|
||||
}
|
||||
|
||||
inline Complex mul(Complex a, Complex b) noexcept {
|
||||
return {a.re * b.re - a.im * b.im, a.re * b.im + a.im * b.re};
|
||||
}
|
||||
|
||||
inline Complex scale(Complex value, double gain) noexcept {
|
||||
return {value.re * gain, value.im * gain};
|
||||
}
|
||||
|
||||
constexpr double kPi = 3.141592653589793238462643383279502884;
|
||||
constexpr int kChannels = EJOC_BINAURAL_INPUT_CHANNELS;
|
||||
constexpr int kEars = EJOC_BINAURAL_OUTPUT_CHANNELS;
|
||||
constexpr int kBlock = EJOC_BINAURAL_BLOCK_SAMPLES;
|
||||
constexpr int kSlots = kBlock / 64;
|
||||
constexpr int kQmf = EJOC_BINAURAL_QMF_BANDS;
|
||||
constexpr int kHybrid = EJOC_BINAURAL_HYBRID_BANDS;
|
||||
constexpr int kRank = 4;
|
||||
|
||||
class Renderer final {
|
||||
public:
|
||||
Renderer() noexcept {
|
||||
initialize_fft();
|
||||
reset();
|
||||
}
|
||||
|
||||
int configure_kernels(
|
||||
const double* qmf_analysis,
|
||||
const double* hybrid_low,
|
||||
const int16_t* hybrid_indices,
|
||||
const double* hybrid_values,
|
||||
uint32_t hybrid_count,
|
||||
const double* qmf_basis,
|
||||
const double* qmf_taps) noexcept {
|
||||
if (!qmf_analysis || !hybrid_low || !hybrid_indices || !hybrid_values ||
|
||||
!qmf_basis || !qmf_taps || hybrid_count == 0) {
|
||||
return fail("invalid binaural kernel configuration");
|
||||
}
|
||||
std::memcpy(qmf_analysis_.data(), qmf_analysis,
|
||||
qmf_analysis_.size() * sizeof(double));
|
||||
hybrid_low_.assign(hybrid_low, hybrid_low + 3 * 2 * 13 * 16 * 2);
|
||||
hybrid_indices_.assign(hybrid_indices, hybrid_indices + hybrid_count * 4);
|
||||
hybrid_values_.assign(hybrid_values, hybrid_values + hybrid_count);
|
||||
std::memcpy(qmf_basis_.data(), qmf_basis,
|
||||
qmf_basis_.size() * sizeof(double));
|
||||
std::memcpy(qmf_taps_.data(), qmf_taps,
|
||||
qmf_taps_.size() * sizeof(double));
|
||||
kernels_ready_ = true;
|
||||
reset();
|
||||
return 0;
|
||||
}
|
||||
|
||||
int configure_room(
|
||||
uint32_t bands,
|
||||
uint32_t allpass_count,
|
||||
const uint32_t* allpass_delays,
|
||||
const double* allpass_gains,
|
||||
const uint32_t* fdn_delays,
|
||||
const double* fdn_matrix,
|
||||
uint32_t output_tap_delay,
|
||||
const double* feedback_complex,
|
||||
const double* output_taps,
|
||||
const double* output_complex,
|
||||
uint32_t extra_count,
|
||||
const uint32_t* extra_delays,
|
||||
const double* extra_fields_complex,
|
||||
const double* extra_matrices) noexcept {
|
||||
if (bands != 64 || !fdn_delays || !fdn_matrix || !feedback_complex ||
|
||||
!output_taps || !output_complex ||
|
||||
(allpass_count && (!allpass_delays || !allpass_gains)) ||
|
||||
(extra_count && (!extra_delays || !extra_fields_complex || !extra_matrices))) {
|
||||
return fail("invalid binaural room configuration");
|
||||
}
|
||||
room_bands_ = bands;
|
||||
if (allpass_count) {
|
||||
allpass_delays_.assign(allpass_delays, allpass_delays + allpass_count);
|
||||
allpass_gains_.assign(allpass_gains, allpass_gains + allpass_count);
|
||||
} else {
|
||||
allpass_delays_.clear();
|
||||
allpass_gains_.clear();
|
||||
}
|
||||
allpass_offsets_.resize(allpass_count);
|
||||
allpass_positions_.assign(allpass_count, 0);
|
||||
size_t allpass_size = 0;
|
||||
for (uint32_t index = 0; index < allpass_count; ++index) {
|
||||
if (allpass_delays_[index] == 0) {
|
||||
return fail("binaural allpass delay must be positive");
|
||||
}
|
||||
allpass_offsets_[index] = allpass_size;
|
||||
allpass_size += static_cast<size_t>(allpass_delays_[index]) * bands;
|
||||
}
|
||||
allpass_memory_.assign(allpass_size, {});
|
||||
|
||||
room_capacity_ = 0;
|
||||
for (int branch = 0; branch < 4; ++branch) {
|
||||
fdn_delays_[branch] = fdn_delays[branch];
|
||||
room_capacity_ = std::max(room_capacity_, fdn_delays_[branch]);
|
||||
}
|
||||
if (room_capacity_ == 0) {
|
||||
return fail("binaural room delay must be positive");
|
||||
}
|
||||
std::copy(fdn_matrix, fdn_matrix + 16, fdn_matrix_.begin());
|
||||
output_tap_delay_ = output_tap_delay;
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
for (int branch = 0; branch < 4; ++branch) {
|
||||
const size_t complex_index = (static_cast<size_t>(band) * 4 + branch) * 2;
|
||||
feedback_[band][branch] = {
|
||||
feedback_complex[complex_index], feedback_complex[complex_index + 1]};
|
||||
output_taps_[band][branch] = output_taps[band * 4 + branch];
|
||||
for (int ear = 0; ear < 2; ++ear) {
|
||||
const size_t output_index =
|
||||
((static_cast<size_t>(ear) * 64 + band) * 4 + branch) * 2;
|
||||
output_matrix_[ear][band][branch] = {
|
||||
output_complex[output_index], output_complex[output_index + 1]};
|
||||
}
|
||||
}
|
||||
}
|
||||
room_memory_.assign(static_cast<size_t>(room_capacity_) * 64 * 4, {});
|
||||
if (extra_count) {
|
||||
extra_delays_.assign(extra_delays, extra_delays + extra_count);
|
||||
} else {
|
||||
extra_delays_.clear();
|
||||
}
|
||||
extra_fields_.resize(static_cast<size_t>(extra_count) * 64);
|
||||
extra_matrices_.resize(static_cast<size_t>(extra_count) * 16);
|
||||
for (uint32_t extra = 0; extra < extra_count; ++extra) {
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
const size_t source = (static_cast<size_t>(extra) * 64 + band) * 2;
|
||||
extra_fields_[static_cast<size_t>(extra) * 64 + band] = {
|
||||
extra_fields_complex[source], extra_fields_complex[source + 1]};
|
||||
}
|
||||
std::copy(extra_matrices + static_cast<size_t>(extra) * 16,
|
||||
extra_matrices + static_cast<size_t>(extra + 1) * 16,
|
||||
extra_matrices_.begin() + static_cast<size_t>(extra) * 16);
|
||||
}
|
||||
room_ready_ = true;
|
||||
reset();
|
||||
return 0;
|
||||
}
|
||||
|
||||
int reset() noexcept {
|
||||
qmf_history_.fill(0.0);
|
||||
hybrid_low_history_.fill({});
|
||||
hybrid_high_history_.fill({});
|
||||
synthesis_history_.fill(0.0);
|
||||
std::fill(allpass_memory_.begin(), allpass_memory_.end(), Complex{});
|
||||
std::fill(allpass_positions_.begin(), allpass_positions_.end(), 0u);
|
||||
std::fill(room_memory_.begin(), room_memory_.end(), Complex{});
|
||||
room_position_ = 0;
|
||||
error_[0] = '\0';
|
||||
return 0;
|
||||
}
|
||||
|
||||
const char* error() const noexcept {
|
||||
return error_[0] ? error_ : "";
|
||||
}
|
||||
|
||||
int process(
|
||||
const double* input,
|
||||
const double* gains,
|
||||
const double* room_sends,
|
||||
double output_gain,
|
||||
double* output) noexcept {
|
||||
if (!kernels_ready_ || !room_ready_) {
|
||||
return fail("binaural renderer is not configured");
|
||||
}
|
||||
if (!input || !gains || !room_sends || !output || !std::isfinite(output_gain)) {
|
||||
return fail("invalid binaural process arguments");
|
||||
}
|
||||
for (int slot = 0; slot < kSlots; ++slot) {
|
||||
std::array<Complex, kChannels * kQmf> qmf{};
|
||||
std::array<Complex, kChannels * kHybrid> hybrid{};
|
||||
analyze_qmf(input + static_cast<size_t>(slot) * 64 * kChannels, qmf);
|
||||
analyze_hybrid(qmf, hybrid);
|
||||
|
||||
std::array<Complex, kEars * kHybrid> rendered{};
|
||||
std::array<Complex, kHybrid> room_input{};
|
||||
for (int source = kChannels - 1; source >= 0; --source) {
|
||||
for (int band = 0; band < kHybrid; ++band) {
|
||||
const Complex value = hybrid[source * kHybrid + band];
|
||||
room_input[band] = add(room_input[band], scale(value, room_sends[source]));
|
||||
for (int ear = 0; ear < kEars; ++ear) {
|
||||
const size_t gain_index =
|
||||
(((static_cast<size_t>(source) * kEars + ear) * kHybrid + band) * 2);
|
||||
const Complex gain{gains[gain_index], gains[gain_index + 1]};
|
||||
rendered[ear * kHybrid + band] = add(
|
||||
rendered[ear * kHybrid + band], mul(value, gain));
|
||||
}
|
||||
}
|
||||
}
|
||||
const auto room = process_room(room_input);
|
||||
for (size_t index = 0; index < rendered.size(); ++index) {
|
||||
rendered[index] = add(rendered[index], room[index]);
|
||||
}
|
||||
|
||||
std::array<Complex, kEars * kQmf> qmf_output{};
|
||||
synthesize_hybrid(rendered, qmf_output);
|
||||
for (int ear = 0; ear < kEars; ++ear) {
|
||||
std::array<double, 64> samples{};
|
||||
synthesize_qmf(qmf_output.data() + ear * kQmf, ear, samples);
|
||||
for (int sample = 0; sample < 64; ++sample) {
|
||||
output[(static_cast<size_t>(slot) * 64 + sample) * 2 + ear] =
|
||||
samples[sample] * output_gain;
|
||||
}
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
private:
|
||||
int fail(const char* message) noexcept {
|
||||
std::snprintf(error_, sizeof(error_), "%s", message);
|
||||
return -1;
|
||||
}
|
||||
|
||||
void initialize_fft() noexcept {
|
||||
for (int index = 0; index < 128; ++index) {
|
||||
int value = index;
|
||||
int reversed = 0;
|
||||
for (int bit = 0; bit < 7; ++bit) {
|
||||
reversed = (reversed << 1) | (value & 1);
|
||||
value >>= 1;
|
||||
}
|
||||
bit_reverse_[index] = static_cast<uint8_t>(reversed);
|
||||
}
|
||||
for (int phase = 0; phase < 64; ++phase) {
|
||||
const double angle = -kPi * static_cast<double>(phase) / 128.0;
|
||||
premod_[phase] = {std::cos(angle), std::sin(angle)};
|
||||
const double post_angle =
|
||||
-3.0 * (static_cast<double>(phase) + 0.5) * kPi / 128.0;
|
||||
post_[phase] = {std::cos(post_angle), std::sin(post_angle)};
|
||||
even_post_[phase] = {0.0, (phase & 1) ? -1.0 : 1.0};
|
||||
}
|
||||
}
|
||||
|
||||
void fft128(std::array<Complex, 128>& values) const noexcept {
|
||||
for (int index = 0; index < 128; ++index) {
|
||||
const int reversed = bit_reverse_[index];
|
||||
if (reversed > index) {
|
||||
std::swap(values[index], values[reversed]);
|
||||
}
|
||||
}
|
||||
for (int length = 2; length <= 128; length <<= 1) {
|
||||
const double angle = -2.0 * kPi / static_cast<double>(length);
|
||||
const Complex step{std::cos(angle), std::sin(angle)};
|
||||
for (int start = 0; start < 128; start += length) {
|
||||
Complex rotation{1.0, 0.0};
|
||||
for (int offset = 0; offset < length / 2; ++offset) {
|
||||
const Complex even = values[start + offset];
|
||||
const Complex odd = mul(values[start + offset + length / 2], rotation);
|
||||
values[start + offset] = {even.re + odd.re, even.im + odd.im};
|
||||
values[start + offset + length / 2] = {
|
||||
even.re - odd.re, even.im - odd.im};
|
||||
rotation = mul(rotation, step);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void qmf_transform(const std::array<double, 64>& source,
|
||||
std::array<Complex, 64>& target) const noexcept {
|
||||
std::array<Complex, 128> work{};
|
||||
for (int phase = 0; phase < 64; ++phase) {
|
||||
work[phase] = scale(premod_[phase], source[phase]);
|
||||
}
|
||||
fft128(work);
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
target[band] = mul(work[band], post_[band]);
|
||||
}
|
||||
}
|
||||
|
||||
void analyze_qmf(const double* input,
|
||||
std::array<Complex, kChannels * kQmf>& output) noexcept {
|
||||
for (int channel = 0; channel < kChannels; ++channel) {
|
||||
for (int lag = 9; lag > 0; --lag) {
|
||||
for (int phase = 0; phase < 64; ++phase) {
|
||||
qmf_history_[qmf_history_index(lag, channel, phase)] =
|
||||
qmf_history_[qmf_history_index(lag - 1, channel, phase)];
|
||||
}
|
||||
}
|
||||
for (int phase = 0; phase < 64; ++phase) {
|
||||
qmf_history_[qmf_history_index(0, channel, phase)] =
|
||||
input[phase * kChannels + channel];
|
||||
}
|
||||
std::array<double, 64> even{};
|
||||
std::array<double, 64> odd{};
|
||||
for (int phase = 0; phase < 64; ++phase) {
|
||||
for (int lag = 0; lag < 10; ++lag) {
|
||||
const double value =
|
||||
qmf_history_[qmf_history_index(lag, channel, phase)] *
|
||||
qmf_analysis_[phase * 10 + lag];
|
||||
(lag & 1 ? odd[phase] : even[phase]) += value;
|
||||
}
|
||||
}
|
||||
std::array<Complex, 64> even_fft{};
|
||||
std::array<Complex, 64> odd_fft{};
|
||||
qmf_transform(even, even_fft);
|
||||
qmf_transform(odd, odd_fft);
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
output[channel * 64 + band] = add(
|
||||
odd_fft[band], mul(even_fft[band], even_post_[band]));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void analyze_hybrid(
|
||||
const std::array<Complex, kChannels * kQmf>& qmf,
|
||||
std::array<Complex, kChannels * kHybrid>& output) noexcept {
|
||||
for (int channel = 0; channel < kChannels; ++channel) {
|
||||
for (int lag = 12; lag > 0; --lag) {
|
||||
for (int band = 0; band < 3; ++band) {
|
||||
hybrid_low_history_[hybrid_low_history_index(lag, channel, band)] =
|
||||
hybrid_low_history_[hybrid_low_history_index(lag - 1, channel, band)];
|
||||
}
|
||||
}
|
||||
for (int band = 0; band < 3; ++band) {
|
||||
hybrid_low_history_[hybrid_low_history_index(0, channel, band)] =
|
||||
qmf[channel * 64 + band];
|
||||
}
|
||||
for (int output_band = 0; output_band < 16; ++output_band) {
|
||||
Complex value{};
|
||||
for (int lag = 0; lag < 13; ++lag) {
|
||||
for (int input_band = 0; input_band < 3; ++input_band) {
|
||||
const Complex source = hybrid_low_history_[
|
||||
hybrid_low_history_index(lag, channel, input_band)];
|
||||
const double components[2]{source.re, source.im};
|
||||
for (int input_component = 0; input_component < 2; ++input_component) {
|
||||
value.re += components[input_component] * hybrid_low_[
|
||||
hybrid_low_kernel_index(input_band, input_component, lag,
|
||||
output_band, 0)];
|
||||
value.im += components[input_component] * hybrid_low_[
|
||||
hybrid_low_kernel_index(input_band, input_component, lag,
|
||||
output_band, 1)];
|
||||
}
|
||||
}
|
||||
}
|
||||
output[channel * kHybrid + output_band] = value;
|
||||
}
|
||||
for (int band = 0; band < 61; ++band) {
|
||||
output[channel * kHybrid + 16 + band] =
|
||||
hybrid_high_history_[hybrid_high_history_index(0, channel, band)];
|
||||
for (int delay = 0; delay < 5; ++delay) {
|
||||
hybrid_high_history_[hybrid_high_history_index(delay, channel, band)] =
|
||||
hybrid_high_history_[hybrid_high_history_index(delay + 1, channel, band)];
|
||||
}
|
||||
hybrid_high_history_[hybrid_high_history_index(5, channel, band)] =
|
||||
qmf[channel * 64 + 3 + band];
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
std::array<Complex, kEars * kHybrid> process_room(
|
||||
const std::array<Complex, kHybrid>& input) noexcept {
|
||||
std::array<Complex, 64> filtered{};
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
filtered[band] = scale(input[band], 0.70710677);
|
||||
}
|
||||
for (size_t stage = 0; stage < allpass_delays_.size(); ++stage) {
|
||||
const uint32_t position = allpass_positions_[stage];
|
||||
const double gain = allpass_gains_[stage];
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
Complex& memory = allpass_memory_[
|
||||
allpass_offsets_[stage] + static_cast<size_t>(position) * 64 + band];
|
||||
const Complex residual = add(filtered[band], scale(memory, -gain));
|
||||
filtered[band] = add(scale(residual, gain), memory);
|
||||
memory = residual;
|
||||
}
|
||||
allpass_positions_[stage] = (position + 1) % allpass_delays_[stage];
|
||||
}
|
||||
|
||||
std::array<Complex, 64 * 4> branches{};
|
||||
std::array<Complex, 64 * 4> taps{};
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
for (int branch = 0; branch < 4; ++branch) {
|
||||
Complex value = filtered[band];
|
||||
for (int source = 0; source < 4; ++source) {
|
||||
const uint32_t position =
|
||||
(room_position_ + room_capacity_ - fdn_delays_[source]) % room_capacity_;
|
||||
value = add(value, scale(room_memory_[
|
||||
room_memory_index(position, band, source)],
|
||||
fdn_matrix_[branch * 4 + source]));
|
||||
}
|
||||
branches[band * 4 + branch] = value;
|
||||
const uint32_t tap_position =
|
||||
(room_position_ + room_capacity_ -
|
||||
(output_tap_delay_ % room_capacity_)) % room_capacity_;
|
||||
taps[band * 4 + branch] =
|
||||
room_memory_[room_memory_index(tap_position, band, branch)];
|
||||
}
|
||||
}
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
for (int branch = 0; branch < 4; ++branch) {
|
||||
room_memory_[room_memory_index(room_position_, band, branch)] =
|
||||
mul(branches[band * 4 + branch], feedback_[band][branch]);
|
||||
}
|
||||
}
|
||||
room_position_ = (room_position_ + 1) % room_capacity_;
|
||||
|
||||
std::array<Complex, 64 * 4> extra{};
|
||||
for (size_t index = 0; index < extra_delays_.size(); ++index) {
|
||||
const uint32_t position =
|
||||
(room_position_ + room_capacity_ -
|
||||
((extra_delays_[index] + 1) % room_capacity_)) % room_capacity_;
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
for (int target = 0; target < 4; ++target) {
|
||||
Complex mixed{};
|
||||
for (int source = 0; source < 4; ++source) {
|
||||
mixed = add(mixed, scale(room_memory_[
|
||||
room_memory_index(position, band, source)],
|
||||
extra_matrices_[index * 16 + target * 4 + source]));
|
||||
}
|
||||
extra[band * 4 + target] = add(
|
||||
extra[band * 4 + target],
|
||||
mul(mixed, extra_fields_[index * 64 + band]));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
std::array<Complex, kEars * kHybrid> output{};
|
||||
for (int ear = 0; ear < 2; ++ear) {
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
Complex value{};
|
||||
for (int branch = 0; branch < 4; ++branch) {
|
||||
const Complex signal = add(
|
||||
scale(taps[band * 4 + branch], output_taps_[band][branch]),
|
||||
extra[band * 4 + branch]);
|
||||
value = add(value, mul(
|
||||
signal, output_matrix_[ear][band][branch]));
|
||||
}
|
||||
output[ear * kHybrid + band] = value;
|
||||
}
|
||||
}
|
||||
return output;
|
||||
}
|
||||
|
||||
void synthesize_hybrid(
|
||||
const std::array<Complex, kEars * kHybrid>& input,
|
||||
std::array<Complex, kEars * kQmf>& output) const noexcept {
|
||||
for (size_t mapping = 0; mapping < hybrid_values_.size(); ++mapping) {
|
||||
const int16_t* index = hybrid_indices_.data() + mapping * 4;
|
||||
const int input_band = index[0];
|
||||
const int input_component = index[1];
|
||||
const int output_band = index[2];
|
||||
const int output_component = index[3];
|
||||
const double gain = hybrid_values_[mapping];
|
||||
for (int ear = 0; ear < 2; ++ear) {
|
||||
const Complex source = input[ear * kHybrid + input_band];
|
||||
Complex& target = output[ear * kQmf + output_band];
|
||||
const double component = input_component == 0 ? source.re : source.im;
|
||||
(output_component == 0 ? target.re : target.im) += component * gain;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void synthesize_qmf(const Complex* input, int ear,
|
||||
std::array<double, 64>& output) noexcept {
|
||||
std::array<double, 64 * kRank> features{};
|
||||
std::array<double, 128> flat{};
|
||||
for (int band = 0; band < 64; ++band) {
|
||||
flat[band * 2] = input[band].re;
|
||||
flat[band * 2 + 1] = input[band].im;
|
||||
}
|
||||
for (int phase = 0; phase < 64; ++phase) {
|
||||
for (int rank = 0; rank < kRank; ++rank) {
|
||||
double value = 0.0;
|
||||
const size_t base = (static_cast<size_t>(phase) * kRank + rank) * 128;
|
||||
for (int component = 0; component < 128; ++component) {
|
||||
value += flat[component] * qmf_basis_[base + component];
|
||||
}
|
||||
features[phase * kRank + rank] = value;
|
||||
}
|
||||
}
|
||||
for (int phase = 0; phase < 64; ++phase) {
|
||||
double value = 0.0;
|
||||
for (int lag = 0; lag < 10; ++lag) {
|
||||
for (int rank = 0; rank < kRank; ++rank) {
|
||||
const double feature = lag == 0
|
||||
? features[phase * kRank + rank]
|
||||
: synthesis_history_[synthesis_history_index(
|
||||
ear, lag - 1, phase, rank)];
|
||||
value += feature * qmf_taps_[
|
||||
((static_cast<size_t>(phase) * 10 + lag) * kRank + rank)];
|
||||
}
|
||||
}
|
||||
output[phase] = value;
|
||||
}
|
||||
for (int lag = 8; lag > 0; --lag) {
|
||||
for (int phase = 0; phase < 64; ++phase) {
|
||||
for (int rank = 0; rank < kRank; ++rank) {
|
||||
synthesis_history_[synthesis_history_index(ear, lag, phase, rank)] =
|
||||
synthesis_history_[synthesis_history_index(
|
||||
ear, lag - 1, phase, rank)];
|
||||
}
|
||||
}
|
||||
}
|
||||
for (int phase = 0; phase < 64; ++phase) {
|
||||
for (int rank = 0; rank < kRank; ++rank) {
|
||||
synthesis_history_[synthesis_history_index(ear, 0, phase, rank)] =
|
||||
features[phase * kRank + rank];
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
static size_t qmf_history_index(int lag, int channel, int phase) noexcept {
|
||||
return (static_cast<size_t>(lag) * kChannels + channel) * 64 + phase;
|
||||
}
|
||||
|
||||
static size_t hybrid_low_history_index(int lag, int channel, int band) noexcept {
|
||||
return (static_cast<size_t>(lag) * kChannels + channel) * 3 + band;
|
||||
}
|
||||
|
||||
static size_t hybrid_high_history_index(int delay, int channel, int band) noexcept {
|
||||
return (static_cast<size_t>(delay) * kChannels + channel) * 61 + band;
|
||||
}
|
||||
|
||||
static size_t hybrid_low_kernel_index(
|
||||
int input_band, int input_component, int lag,
|
||||
int output_band, int output_component) noexcept {
|
||||
return (((static_cast<size_t>(input_band) * 2 + input_component) * 13 + lag) *
|
||||
16 + output_band) * 2 + output_component;
|
||||
}
|
||||
|
||||
size_t room_memory_index(uint32_t position, int band, int branch) const noexcept {
|
||||
return (static_cast<size_t>(position) * 64 + band) * 4 + branch;
|
||||
}
|
||||
|
||||
static size_t synthesis_history_index(
|
||||
int ear, int lag, int phase, int rank) noexcept {
|
||||
return (((static_cast<size_t>(ear) * 9 + lag) * 64 + phase) * kRank + rank);
|
||||
}
|
||||
|
||||
bool kernels_ready_ = false;
|
||||
bool room_ready_ = false;
|
||||
std::array<double, 64 * 10> qmf_analysis_{};
|
||||
std::vector<double> hybrid_low_;
|
||||
std::vector<int16_t> hybrid_indices_;
|
||||
std::vector<double> hybrid_values_;
|
||||
std::array<double, 64 * kRank * 128> qmf_basis_{};
|
||||
std::array<double, 64 * 10 * kRank> qmf_taps_{};
|
||||
|
||||
std::array<double, 10 * kChannels * 64> qmf_history_{};
|
||||
std::array<Complex, 13 * kChannels * 3> hybrid_low_history_{};
|
||||
std::array<Complex, 6 * kChannels * 61> hybrid_high_history_{};
|
||||
std::array<double, kEars * 9 * 64 * kRank> synthesis_history_{};
|
||||
|
||||
uint32_t room_bands_ = 0;
|
||||
std::vector<uint32_t> allpass_delays_;
|
||||
std::vector<double> allpass_gains_;
|
||||
std::vector<size_t> allpass_offsets_;
|
||||
std::vector<uint32_t> allpass_positions_;
|
||||
std::vector<Complex> allpass_memory_;
|
||||
std::array<uint32_t, 4> fdn_delays_{};
|
||||
std::array<double, 16> fdn_matrix_{};
|
||||
uint32_t room_capacity_ = 0;
|
||||
uint32_t output_tap_delay_ = 0;
|
||||
std::array<std::array<Complex, 4>, 64> feedback_{};
|
||||
std::array<std::array<double, 4>, 64> output_taps_{};
|
||||
std::array<std::array<std::array<Complex, 4>, 64>, 2> output_matrix_{};
|
||||
std::vector<Complex> room_memory_;
|
||||
uint32_t room_position_ = 0;
|
||||
std::vector<uint32_t> extra_delays_;
|
||||
std::vector<Complex> extra_fields_;
|
||||
std::vector<double> extra_matrices_;
|
||||
|
||||
std::array<uint8_t, 128> bit_reverse_{};
|
||||
std::array<Complex, 64> premod_{};
|
||||
std::array<Complex, 64> post_{};
|
||||
std::array<Complex, 64> even_post_{};
|
||||
char error_[256]{};
|
||||
};
|
||||
|
||||
} // namespace ejoc::binaural
|
||||
|
||||
extern "C" {
|
||||
|
||||
ejoc_binaural_renderer_handle EJOC_CALL ejoc_binaural_renderer_create(void) {
|
||||
return new (std::nothrow) ejoc::binaural::Renderer();
|
||||
}
|
||||
|
||||
void EJOC_CALL ejoc_binaural_renderer_destroy(ejoc_binaural_renderer_handle handle) {
|
||||
delete static_cast<ejoc::binaural::Renderer*>(handle);
|
||||
}
|
||||
|
||||
int EJOC_CALL ejoc_binaural_renderer_reset(ejoc_binaural_renderer_handle handle) {
|
||||
return handle ? static_cast<ejoc::binaural::Renderer*>(handle)->reset() : -1;
|
||||
}
|
||||
|
||||
const char* EJOC_CALL ejoc_binaural_renderer_last_error(
|
||||
ejoc_binaural_renderer_handle handle) {
|
||||
return handle ? static_cast<ejoc::binaural::Renderer*>(handle)->error()
|
||||
: "null binaural renderer handle";
|
||||
}
|
||||
|
||||
int EJOC_CALL ejoc_binaural_renderer_configure_kernels(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
const double* qmf_analysis,
|
||||
const double* hybrid_low,
|
||||
const int16_t* hybrid_indices,
|
||||
const double* hybrid_values,
|
||||
uint32_t hybrid_count,
|
||||
const double* qmf_basis,
|
||||
const double* qmf_taps) {
|
||||
return handle ? static_cast<ejoc::binaural::Renderer*>(handle)->configure_kernels(
|
||||
qmf_analysis, hybrid_low, hybrid_indices, hybrid_values,
|
||||
hybrid_count, qmf_basis, qmf_taps) : -1;
|
||||
}
|
||||
|
||||
int EJOC_CALL ejoc_binaural_renderer_configure_room(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
uint32_t bands,
|
||||
uint32_t allpass_count,
|
||||
const uint32_t* allpass_delays,
|
||||
const double* allpass_gains,
|
||||
const uint32_t* fdn_delays,
|
||||
const double* fdn_matrix,
|
||||
uint32_t output_tap_delay,
|
||||
const double* feedback_complex,
|
||||
const double* output_taps,
|
||||
const double* output_complex,
|
||||
uint32_t extra_count,
|
||||
const uint32_t* extra_delays,
|
||||
const double* extra_fields_complex,
|
||||
const double* extra_matrices) {
|
||||
return handle ? static_cast<ejoc::binaural::Renderer*>(handle)->configure_room(
|
||||
bands, allpass_count, allpass_delays, allpass_gains,
|
||||
fdn_delays, fdn_matrix, output_tap_delay,
|
||||
feedback_complex, output_taps, output_complex,
|
||||
extra_count, extra_delays, extra_fields_complex, extra_matrices) : -1;
|
||||
}
|
||||
|
||||
int EJOC_CALL ejoc_binaural_renderer_process(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
const double* input16_interleaved,
|
||||
const double* gains_complex,
|
||||
const double* room_sends,
|
||||
double output_gain,
|
||||
double* output_stereo_interleaved) {
|
||||
return handle ? static_cast<ejoc::binaural::Renderer*>(handle)->process(
|
||||
input16_interleaved, gains_complex, room_sends,
|
||||
output_gain, output_stereo_interleaved) : -1;
|
||||
}
|
||||
|
||||
} // extern "C"
|
||||
@@ -664,13 +664,13 @@ uint32_t EJOC_CALL ejoc_abi_version(void) {
|
||||
|
||||
const char* EJOC_CALL ejoc_build_info(void) {
|
||||
#if defined(_MSC_VER)
|
||||
return "eac3joc-core abi=1 compiler=MSVC fft=fixed64 speaker=double crt=static-by-build";
|
||||
return "eac3joc-core abi=1 compiler=MSVC fft=fixed64 speaker=double binaural=double crt=static-by-build";
|
||||
#elif defined(__clang__)
|
||||
return "eac3joc-core abi=1 compiler=Clang fft=fixed64 speaker=double";
|
||||
return "eac3joc-core abi=1 compiler=Clang fft=fixed64 speaker=double binaural=double";
|
||||
#elif defined(__GNUC__)
|
||||
return "eac3joc-core abi=1 compiler=GCC fft=fixed64 speaker=double";
|
||||
return "eac3joc-core abi=1 compiler=GCC fft=fixed64 speaker=double binaural=double";
|
||||
#else
|
||||
return "eac3joc-core abi=1 compiler=unknown fft=fixed64 speaker=double";
|
||||
return "eac3joc-core abi=1 compiler=unknown fft=fixed64 speaker=double binaural=double";
|
||||
#endif
|
||||
}
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1 +1,3 @@
|
||||
numpy>=1.24
|
||||
scipy>=1.10
|
||||
h5py>=3.8
|
||||
|
||||
+8
-4
@@ -14,7 +14,7 @@ import adm_atmos
|
||||
|
||||
|
||||
def assemble_from_raw(raw16_path, out_path, scale=1.0, kf_tracks=None,
|
||||
duration_sec=None, rate=48000):
|
||||
duration_sec=None, rate=48000, joc_binaural_mode=4):
|
||||
"""16ch f32 交织 raw → 25ch ADM BWF(空 7.1.2 bed + LFE + 15 对象)。
|
||||
|
||||
raw16: (n, 16) 交织(ch0 = LFE,ch1-15 = 对象)。
|
||||
@@ -54,7 +54,8 @@ def assemble_from_raw(raw16_path, out_path, scale=1.0, kf_tracks=None,
|
||||
kf_tracks.append(("JOC_Object_%d" % (oi + 1),
|
||||
[(0.0, 0.0, 0.0, 0.0, max(duration_sec, 1e-6))]))
|
||||
adm_atmos.build_master(out_path, BedView(), ObjView(), kf_tracks,
|
||||
duration_sec, rate=rate)
|
||||
duration_sec, rate=rate,
|
||||
joc_binaural_mode=joc_binaural_mode)
|
||||
# 及时释放 Windows 文件句柄,允许 TemporaryDirectory 删除中间 raw。
|
||||
raw._mmap.close()
|
||||
return out_path
|
||||
@@ -69,12 +70,14 @@ class StreamingMaster:
|
||||
channels 10..24.
|
||||
"""
|
||||
|
||||
def __init__(self, out_path, duration_sec, rate=48000, block_samples=131072):
|
||||
def __init__(self, out_path, duration_sec, rate=48000, block_samples=131072,
|
||||
joc_binaural_mode=4):
|
||||
if block_samples < 1536:
|
||||
raise ValueError("block_samples must be at least one E-AC-3 frame")
|
||||
self.out_path = os.fspath(out_path)
|
||||
self.duration_sec = float(duration_sec)
|
||||
self.rate = int(rate)
|
||||
self.joc_binaural_mode = joc_binaural_mode
|
||||
self._sink = adm_atmos.Sink25(self.out_path, 25, self.rate)
|
||||
self._buffer = np.empty((int(block_samples), 25), dtype=np.float32)
|
||||
self._used = 0
|
||||
@@ -112,7 +115,8 @@ class StreamingMaster:
|
||||
import adm_serializer
|
||||
axml = adm_serializer.build_axml(kf_tracks, self.duration_sec)
|
||||
chna = adm_atmos.build_chna()
|
||||
dbmd = adm_atmos.build_dbmd(25)
|
||||
dbmd = adm_atmos.build_dbmd(
|
||||
25, joc_binaural_mode=self.joc_binaural_mode)
|
||||
trajectory_blocks = sum(len(track[1]) for track in kf_tracks)
|
||||
self.metadata_info = {
|
||||
"axml_bytes": len(axml),
|
||||
|
||||
+39
-12
@@ -2,6 +2,7 @@
|
||||
|
||||
输出由 10 声道 7.1.2 bed 和 15 路对象组成;RF64 尺寸字段在写入完成后回填。
|
||||
"""
|
||||
import operator
|
||||
import struct
|
||||
import numpy as np
|
||||
import xml.etree.ElementTree as ET
|
||||
@@ -21,6 +22,14 @@ BED_POS = [(-1.0, 1.0, 0.0), (1.0, 1.0, 0.0), (0.0, 1.0, 0.0),
|
||||
(-1.0, -1.0, 0.0), (1.0, -1.0, 0.0), (-1.0, 0.0, 1.0), (1.0, 0.0, 1.0)]
|
||||
|
||||
N_OBJ = 15
|
||||
JOC_BINAURAL_MODES = {
|
||||
"off": 0,
|
||||
"near": 1,
|
||||
"far": 2,
|
||||
"mid": 3,
|
||||
"unspecified": 4,
|
||||
}
|
||||
JOC_BINAURAL_MODE_DEFAULT = "unspecified"
|
||||
|
||||
def q_to_adm_xyz(q1, q2, q3):
|
||||
posX = min(1.0, round(q1 * 62 / 32767.0) / 62.0)
|
||||
@@ -186,7 +195,11 @@ def _checksum(seg):
|
||||
s += b
|
||||
return (~s + 1) & 0xFF
|
||||
|
||||
def build_dbmd(object_count=25):
|
||||
def build_dbmd(object_count=25, joc_binaural_mode=4):
|
||||
"""仅覆盖 segment 10 中 JOC object slots 10..24 的 mode 低 3 bit。"""
|
||||
mode = operator.index(joc_binaural_mode)
|
||||
if mode not in JOC_BINAURAL_MODES.values():
|
||||
raise ValueError(f"invalid JOC binaural render mode: {mode}")
|
||||
out = bytearray(struct.pack("<I", 0x01000006))
|
||||
dd = bytearray(96)
|
||||
dd[1] = 0x47
|
||||
@@ -209,17 +222,33 @@ def build_dbmd(object_count=25):
|
||||
ob[4] = object_count
|
||||
for i in range(5 + 262, len(ob)):
|
||||
ob[i] = 0x84
|
||||
# sync (4), count (2), reserved (1), nine 15-byte config trims,
|
||||
# then one trim-bypass byte per track before the headphone modes.
|
||||
# Preserve the existing template's bed fields and trailing bytes.
|
||||
object_modes = 4 + 2 + 1 + 9 * 15 + object_count
|
||||
for i in range(10, min(object_count, 10 + N_OBJ)):
|
||||
ob[object_modes + i] = (ob[object_modes + i] & 0xF8) | mode
|
||||
out.append(10); out += struct.pack("<H", len(ob)); out += bytes(ob)
|
||||
out.append(_checksum(ob))
|
||||
out += b"\x00\x00"
|
||||
return bytes(out)
|
||||
|
||||
class Sink25:
|
||||
"""RF64 ADM BWF writer.
|
||||
|
||||
Header layout is fixed so that sizes can be patched without rereading the
|
||||
file: RF64+size+WAVE (12) + ds64 chunk (8+28) + fmt chunk (8+16) + data
|
||||
chunk header (8). Sizes beyond 32 bits follow the RF64 convention: the
|
||||
chunk size field holds 0xFFFFFFFF and the true value lives in ds64.
|
||||
"""
|
||||
_DS64_BODY_OFFSET = 20
|
||||
_DATA_SIZE_OFFSET = 76
|
||||
|
||||
def __init__(self, path, channels, rate):
|
||||
self.ch = channels; self.rate = rate; self.frames = 0
|
||||
self.fp = open(path, "wb+")
|
||||
self.fp.write(b"RF64" + struct.pack("<I", 0xFFFFFFFF) + b"WAVE")
|
||||
self._chunk(b"ds64", b"\x00" * 64)
|
||||
self._chunk(b"ds64", b"\x00" * 28)
|
||||
self._chunk(b"fmt ", self._fmt())
|
||||
self._chunk(b"data", b"")
|
||||
def _chunk(self, cid, body):
|
||||
@@ -240,19 +269,17 @@ class Sink25:
|
||||
self._chunk(b"chna", chna_bytes)
|
||||
self._chunk(b"dbmd", dbmd_bytes)
|
||||
self.fp.seek(0, 2); total = self.fp.tell()
|
||||
self.fp.seek(0); head = self.fp.read()
|
||||
m = head.find(b"data")
|
||||
if m >= 0:
|
||||
self.fp.seek(m + 4); self.fp.write(struct.pack("<I", data_len))
|
||||
m = head.find(b"ds64")
|
||||
if m >= 0:
|
||||
self.fp.seek(m + 8)
|
||||
self.fp.write(struct.pack("<QQQI", total - 8, data_len, self.frames, 0))
|
||||
# RF64: 超过 32-bit 的 chunk size 字段写 0xFFFFFFFF,真实大小回填 ds64。
|
||||
self.fp.seek(self._DATA_SIZE_OFFSET)
|
||||
self.fp.write(struct.pack(
|
||||
"<I", data_len if data_len <= 0xFFFFFFFF else 0xFFFFFFFF))
|
||||
self.fp.seek(self._DS64_BODY_OFFSET)
|
||||
self.fp.write(struct.pack("<QQQI", total - 8, data_len, self.frames, 0))
|
||||
self.fp.flush()
|
||||
self.fp.close()
|
||||
|
||||
def build_master(out_path, bed_mm, obj_mm, kf_tracks, duration_sec, rate=48000,
|
||||
block=480000):
|
||||
block=480000, joc_binaural_mode=4):
|
||||
n = min(bed_mm.shape[0], obj_mm.shape[0])
|
||||
try:
|
||||
from . import adm_serializer
|
||||
@@ -261,7 +288,7 @@ def build_master(out_path, bed_mm, obj_mm, kf_tracks, duration_sec, rate=48000,
|
||||
serial_axml = adm_serializer.build_axml
|
||||
axml = serial_axml(kf_tracks, duration_sec)
|
||||
chna = build_chna()
|
||||
dbmd = build_dbmd(25)
|
||||
dbmd = build_dbmd(25, joc_binaural_mode=joc_binaural_mode)
|
||||
sink = Sink25(out_path, 25, rate)
|
||||
for st in range(0, n, block):
|
||||
en = min(n, st + block)
|
||||
|
||||
@@ -0,0 +1,188 @@
|
||||
"""Direct ID11/OAMD position scheduling for the binaural render path."""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
|
||||
import numpy as np
|
||||
|
||||
from adm_atmos import q_to_adm_xyz
|
||||
from oamd_bits import JocFieldState, frame_update
|
||||
from variant_error import UnsupportedVariantError
|
||||
|
||||
OAMD_UPDATE_QUANTUM_SAMPLES = 64
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PositionTransition:
|
||||
start_sample: int
|
||||
duration_samples: int
|
||||
origin: np.ndarray
|
||||
target: np.ndarray
|
||||
|
||||
@property
|
||||
def end_sample(self) -> int:
|
||||
return self.start_sample + self.duration_samples
|
||||
|
||||
|
||||
class _ObjectPositionTrack:
|
||||
def __init__(self):
|
||||
self.initial = np.zeros(3, dtype=np.float64)
|
||||
self.last_target = self.initial.copy()
|
||||
self.transitions: list[PositionTransition] = []
|
||||
self.cursor = 0
|
||||
self.last_query_sample = -1
|
||||
|
||||
def set_initial(self, position):
|
||||
target = np.asarray(position, dtype=np.float64)
|
||||
self.initial = target.copy()
|
||||
self.last_target = target.copy()
|
||||
|
||||
def append(self, start_sample: int, duration_samples: int, target,
|
||||
object_index: int):
|
||||
start = int(start_sample)
|
||||
duration = int(duration_samples)
|
||||
if start < 0 or duration < 0:
|
||||
raise ValueError("position transition timing must be non-negative")
|
||||
target = np.asarray(target, dtype=np.float64)
|
||||
if self.transitions:
|
||||
previous = self.transitions[-1]
|
||||
if start < previous.end_sample:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "overlapping_binaural_position_ramps",
|
||||
"同一对象的新位置更新在上一双耳 ramp 完成前到达",
|
||||
details={
|
||||
"object": object_index,
|
||||
"ramp_start_sample": previous.start_sample,
|
||||
"ramp_end_sample": previous.end_sample,
|
||||
"next_update_sample": start,
|
||||
})
|
||||
if start == previous.start_sample and previous.duration_samples == 0:
|
||||
self.transitions[-1] = PositionTransition(
|
||||
start, duration, previous.origin.copy(), target.copy())
|
||||
self.last_target = target.copy()
|
||||
return
|
||||
self.transitions.append(PositionTransition(
|
||||
start, duration, self.last_target.copy(), target.copy()))
|
||||
self.last_target = target.copy()
|
||||
|
||||
def position_at(self, sample: int) -> np.ndarray:
|
||||
sample = int(sample)
|
||||
if sample < self.last_query_sample:
|
||||
raise ValueError("binaural metadata positions must be queried monotonically")
|
||||
self.last_query_sample = sample
|
||||
while self.cursor < len(self.transitions):
|
||||
transition = self.transitions[self.cursor]
|
||||
if sample < transition.end_sample:
|
||||
break
|
||||
self.initial = transition.target.copy()
|
||||
self.cursor += 1
|
||||
if self.cursor >= len(self.transitions):
|
||||
return self.initial
|
||||
transition = self.transitions[self.cursor]
|
||||
if sample < transition.start_sample:
|
||||
return self.initial
|
||||
if transition.duration_samples == 0:
|
||||
return transition.target
|
||||
amount = (sample - transition.start_sample) / float(transition.duration_samples)
|
||||
return transition.origin + (transition.target - transition.origin) * amount
|
||||
|
||||
|
||||
class OamdPositionTimeline:
|
||||
"""Convert OAMD state updates into a sample-timed Cartesian trajectory."""
|
||||
|
||||
def __init__(self, object_count: int = 15):
|
||||
if object_count != 15:
|
||||
raise ValueError("JOC OAMD currently requires 15 object slots")
|
||||
self.object_count = int(object_count)
|
||||
self.state = JocFieldState()
|
||||
self.tracks = [_ObjectPositionTrack() for _ in range(self.object_count)]
|
||||
self.initialized = False
|
||||
self.previous_targets: list[tuple[float, float, float] | None] = [
|
||||
None] * self.object_count
|
||||
self.payload_count = 0
|
||||
self.transition_count = 0
|
||||
self.last_coded_event_sample = -1
|
||||
|
||||
def _targets(self) -> list[tuple[float, float, float]]:
|
||||
q = self.state.q
|
||||
return [
|
||||
q_to_adm_xyz(
|
||||
q[(object_index, "q1")],
|
||||
q[(object_index, "q2")],
|
||||
q[(object_index, "q3")],
|
||||
)
|
||||
for object_index in range(1, self.object_count + 1)
|
||||
]
|
||||
|
||||
def submit_update(self, update: dict, *, frame_start_sample: int,
|
||||
outer_sample_offset: int = 0,
|
||||
object_delay_samples: int = 1473,
|
||||
processed_sample: int = 0):
|
||||
"""Schedule one already-parsed :func:`oamd_bits.frame_update` result."""
|
||||
frame_start = int(frame_start_sample)
|
||||
outer_offset = int(outer_sample_offset)
|
||||
object_delay = int(object_delay_samples)
|
||||
if min(frame_start, outer_offset, object_delay) < 0:
|
||||
raise ValueError("OAMD frame, outer offset, and object delay must be non-negative")
|
||||
self.state.apply(update["values"])
|
||||
targets = self._targets()
|
||||
coded_event = (
|
||||
frame_start + outer_offset + int(update["block_offset_samples"]))
|
||||
if coded_event < self.last_coded_event_sample:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "non_monotonic_binaural_updates",
|
||||
"双耳 OAMD 更新时间倒退",
|
||||
details={
|
||||
"event_sample": coded_event,
|
||||
"previous_event_sample": self.last_coded_event_sample,
|
||||
})
|
||||
self.last_coded_event_sample = coded_event
|
||||
|
||||
if not self.initialized:
|
||||
if int(processed_sample) > 0:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "late_initial_binaural_state",
|
||||
"首个 OAMD 状态在双耳 PCM 已处理后才出现,无法回填 sample 0",
|
||||
details={
|
||||
"processed_sample": int(processed_sample),
|
||||
"first_event_sample": coded_event,
|
||||
})
|
||||
for index, target in enumerate(targets):
|
||||
self.tracks[index].set_initial(target)
|
||||
self.previous_targets[index] = target
|
||||
self.initialized = True
|
||||
self.payload_count += 1
|
||||
return
|
||||
|
||||
ramp_duration = int(update["ramp_duration_samples"])
|
||||
effective_ramp = max(0, ramp_duration - OAMD_UPDATE_QUANTUM_SAMPLES)
|
||||
transition_start = coded_event + object_delay
|
||||
if effective_ramp:
|
||||
transition_start += OAMD_UPDATE_QUANTUM_SAMPLES
|
||||
for index, target in enumerate(targets):
|
||||
if self.previous_targets[index] == target:
|
||||
continue
|
||||
self.tracks[index].append(
|
||||
transition_start, effective_ramp, target, index + 1)
|
||||
self.previous_targets[index] = target
|
||||
self.transition_count += 1
|
||||
self.payload_count += 1
|
||||
|
||||
def submit_payload(self, payload, *, frame_start_sample: int,
|
||||
outer_sample_offset: int = 0,
|
||||
object_delay_samples: int = 1473,
|
||||
processed_sample: int = 0):
|
||||
update = frame_update(payload)
|
||||
self.submit_update(
|
||||
update,
|
||||
frame_start_sample=frame_start_sample,
|
||||
outer_sample_offset=outer_sample_offset,
|
||||
object_delay_samples=object_delay_samples,
|
||||
processed_sample=processed_sample,
|
||||
)
|
||||
return update
|
||||
|
||||
def positions_at(self, sample: int) -> np.ndarray:
|
||||
return np.stack(
|
||||
[track.position_at(sample) for track in self.tracks], axis=0
|
||||
).astype(np.float64, copy=False)
|
||||
@@ -0,0 +1,214 @@
|
||||
"""ctypes bridge for the native float64 binaural DSP."""
|
||||
from __future__ import annotations
|
||||
|
||||
import ctypes
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
from native_renderer import ABI_VERSION, find_native_library
|
||||
from rosella_filterbank import DEFAULT_KERNEL_DATA, load_kernel_tables
|
||||
from rosella_model import RosellaModel
|
||||
|
||||
BLOCK_SAMPLES = 512
|
||||
INPUT_CHANNELS = 16
|
||||
OUTPUT_CHANNELS = 2
|
||||
HYBRID_BANDS = 77
|
||||
|
||||
|
||||
class NativeBinauralDsp:
|
||||
def __init__(self, model: RosellaModel, *, library_path=None,
|
||||
kernel_data: str | Path = DEFAULT_KERNEL_DATA):
|
||||
self.library_path = find_native_library(library_path)
|
||||
self._lib = ctypes.CDLL(str(self.library_path))
|
||||
self._bind()
|
||||
version = int(self._lib.ejoc_abi_version())
|
||||
if version != ABI_VERSION:
|
||||
raise RuntimeError(
|
||||
f"native ABI mismatch: expected {ABI_VERSION}, got {version}")
|
||||
self._handle = self._lib.ejoc_binaural_renderer_create()
|
||||
if not self._handle:
|
||||
raise RuntimeError("native binaural renderer creation failed")
|
||||
try:
|
||||
self._configure_kernels(kernel_data)
|
||||
self._configure_room(model)
|
||||
except Exception:
|
||||
self.close()
|
||||
raise
|
||||
|
||||
def _bind(self):
|
||||
void_p = ctypes.c_void_p
|
||||
f64_p = ctypes.POINTER(ctypes.c_double)
|
||||
i16_p = ctypes.POINTER(ctypes.c_int16)
|
||||
u32_p = ctypes.POINTER(ctypes.c_uint32)
|
||||
self._lib.ejoc_abi_version.argtypes = []
|
||||
self._lib.ejoc_abi_version.restype = ctypes.c_uint32
|
||||
self._lib.ejoc_binaural_renderer_create.argtypes = []
|
||||
self._lib.ejoc_binaural_renderer_create.restype = void_p
|
||||
self._lib.ejoc_binaural_renderer_destroy.argtypes = [void_p]
|
||||
self._lib.ejoc_binaural_renderer_destroy.restype = None
|
||||
self._lib.ejoc_binaural_renderer_reset.argtypes = [void_p]
|
||||
self._lib.ejoc_binaural_renderer_reset.restype = ctypes.c_int
|
||||
self._lib.ejoc_binaural_renderer_last_error.argtypes = [void_p]
|
||||
self._lib.ejoc_binaural_renderer_last_error.restype = ctypes.c_char_p
|
||||
self._lib.ejoc_binaural_renderer_configure_kernels.argtypes = [
|
||||
void_p, f64_p, f64_p, i16_p, f64_p, ctypes.c_uint32, f64_p, f64_p]
|
||||
self._lib.ejoc_binaural_renderer_configure_kernels.restype = ctypes.c_int
|
||||
self._lib.ejoc_binaural_renderer_configure_room.argtypes = [
|
||||
void_p, ctypes.c_uint32, ctypes.c_uint32, u32_p, f64_p,
|
||||
u32_p, f64_p, ctypes.c_uint32, f64_p, f64_p, f64_p,
|
||||
ctypes.c_uint32, u32_p, f64_p, f64_p]
|
||||
self._lib.ejoc_binaural_renderer_configure_room.restype = ctypes.c_int
|
||||
self._lib.ejoc_binaural_renderer_process.argtypes = [
|
||||
void_p, f64_p, f64_p, f64_p, ctypes.c_double, f64_p]
|
||||
self._lib.ejoc_binaural_renderer_process.restype = ctypes.c_int
|
||||
|
||||
def _raise(self, operation, status):
|
||||
message = self._lib.ejoc_binaural_renderer_last_error(self._handle)
|
||||
detail = (message or b"").decode("utf-8", "replace")
|
||||
raise RuntimeError(
|
||||
f"native binaural renderer {operation} failed ({status}): {detail}")
|
||||
|
||||
@staticmethod
|
||||
def _f64_pointer(values):
|
||||
return values.ctypes.data_as(ctypes.POINTER(ctypes.c_double))
|
||||
|
||||
def _configure_kernels(self, kernel_data):
|
||||
tables = load_kernel_tables(kernel_data)
|
||||
qmf_analysis = np.ascontiguousarray(
|
||||
tables["qmf_analysis_coefficients"], dtype=np.float64)
|
||||
hybrid_low = np.ascontiguousarray(
|
||||
tables["hybrid_analysis_low_kernel"], dtype=np.float64)
|
||||
hybrid_indices = np.ascontiguousarray(
|
||||
tables["hybrid_synthesis_indices"], dtype=np.int16)
|
||||
hybrid_values = np.ascontiguousarray(
|
||||
tables["hybrid_synthesis_values"], dtype=np.float64)
|
||||
qmf_basis = np.ascontiguousarray(
|
||||
tables["qmf_synthesis_basis"], dtype=np.float64)
|
||||
qmf_taps = np.ascontiguousarray(
|
||||
tables["qmf_synthesis_taps"], dtype=np.float64)
|
||||
status = self._lib.ejoc_binaural_renderer_configure_kernels(
|
||||
self._handle,
|
||||
self._f64_pointer(qmf_analysis),
|
||||
self._f64_pointer(hybrid_low),
|
||||
hybrid_indices.ctypes.data_as(ctypes.POINTER(ctypes.c_int16)),
|
||||
self._f64_pointer(hybrid_values),
|
||||
len(hybrid_values),
|
||||
self._f64_pointer(qmf_basis),
|
||||
self._f64_pointer(qmf_taps),
|
||||
)
|
||||
if status:
|
||||
self._raise("configure_kernels", status)
|
||||
|
||||
def _configure_room(self, model: RosellaModel):
|
||||
if float(model.table_a_scalar) >= 0.5:
|
||||
raise NotImplementedError("alternate table-A room mode")
|
||||
bands = min(64, model.table_a_dimension)
|
||||
allpass_delays = np.ascontiguousarray(
|
||||
model.table_a_option_ids, dtype=np.uint32)
|
||||
allpass_gains = np.ascontiguousarray(
|
||||
model.table_a_option_values, dtype=np.float64)
|
||||
fdn_delays = np.ascontiguousarray(
|
||||
model.table_a_four_integers, dtype=np.uint32)
|
||||
fdn_matrix = np.ascontiguousarray(
|
||||
np.asarray(model.table_a_vector16, dtype=np.float64).reshape(
|
||||
4, 4, order="F"))
|
||||
|
||||
filter8 = np.asarray(
|
||||
model.table_a_filter_8x64_padded, dtype=np.float64).reshape(20, 4, 2, 4)
|
||||
filter4 = np.asarray(
|
||||
model.table_a_filter_4x64_padded, dtype=np.float64).reshape(20, 4, 4)
|
||||
filter16 = np.asarray(
|
||||
model.table_a_filter_16x64_padded, dtype=np.float64).reshape(20, 4, 4, 4)
|
||||
feedback = np.empty((64, 4, 2), dtype=np.float64)
|
||||
output_taps = np.empty((64, 4), dtype=np.float64)
|
||||
output_matrix = np.empty((2, 64, 4, 2), dtype=np.float64)
|
||||
for band in range(64):
|
||||
group, lane = divmod(band, 4)
|
||||
feedback[band, :, 0] = filter8[group, :, 0, lane]
|
||||
feedback[band, :, 1] = filter8[group, :, 1, lane]
|
||||
output_taps[band] = filter4[group, :, lane]
|
||||
output_matrix[0, band, :, 0] = filter16[group, :, 0, lane]
|
||||
output_matrix[0, band, :, 1] = filter16[group, :, 1, lane]
|
||||
output_matrix[1, band, :, 0] = filter16[group, :, 2, lane]
|
||||
output_matrix[1, band, :, 1] = filter16[group, :, 3, lane]
|
||||
|
||||
extra_count = int(model.table_a_extra)
|
||||
extra_delays = np.ascontiguousarray(
|
||||
model.table_a_extra_indices, dtype=np.uint32)
|
||||
extra_fields = np.empty((extra_count, 64, 2), dtype=np.float64)
|
||||
extra_source = np.asarray(
|
||||
model.table_a_extra_fields_padded, dtype=np.float64).reshape(
|
||||
extra_count, 20, 2, 4)
|
||||
for extra in range(extra_count):
|
||||
for band in range(64):
|
||||
group, lane = divmod(band, 4)
|
||||
extra_fields[extra, band] = extra_source[extra, group, :, lane]
|
||||
extra_matrices = np.empty((extra_count, 4, 4), dtype=np.float64)
|
||||
for extra in range(extra_count):
|
||||
extra_matrices[extra] = np.asarray(
|
||||
model.table_a_extra_vectors[extra], dtype=np.float64).reshape(
|
||||
4, 4, order="F")
|
||||
|
||||
null_u32 = ctypes.POINTER(ctypes.c_uint32)()
|
||||
null_f64 = ctypes.POINTER(ctypes.c_double)()
|
||||
status = self._lib.ejoc_binaural_renderer_configure_room(
|
||||
self._handle,
|
||||
bands,
|
||||
len(allpass_delays),
|
||||
allpass_delays.ctypes.data_as(ctypes.POINTER(ctypes.c_uint32)),
|
||||
self._f64_pointer(allpass_gains),
|
||||
fdn_delays.ctypes.data_as(ctypes.POINTER(ctypes.c_uint32)),
|
||||
self._f64_pointer(fdn_matrix),
|
||||
int(model.table_a_integer),
|
||||
self._f64_pointer(feedback),
|
||||
self._f64_pointer(output_taps),
|
||||
self._f64_pointer(output_matrix),
|
||||
extra_count,
|
||||
(extra_delays.ctypes.data_as(ctypes.POINTER(ctypes.c_uint32))
|
||||
if extra_count else null_u32),
|
||||
self._f64_pointer(extra_fields) if extra_count else null_f64,
|
||||
self._f64_pointer(extra_matrices) if extra_count else null_f64,
|
||||
)
|
||||
if status:
|
||||
self._raise("configure_room", status)
|
||||
|
||||
def reset(self):
|
||||
if not self._handle:
|
||||
raise RuntimeError("native binaural renderer is closed")
|
||||
status = self._lib.ejoc_binaural_renderer_reset(self._handle)
|
||||
if status:
|
||||
self._raise("reset", status)
|
||||
|
||||
def process_block(self, pcm16, gains, room_sends, output_gain=1.0):
|
||||
if not self._handle:
|
||||
raise RuntimeError("native binaural renderer is closed")
|
||||
source = np.ascontiguousarray(pcm16, dtype=np.float64)
|
||||
gain_values = np.asarray(gains)
|
||||
sends = np.ascontiguousarray(room_sends, dtype=np.float64)
|
||||
if source.shape != (BLOCK_SAMPLES, INPUT_CHANNELS):
|
||||
raise ValueError(f"pcm16 block must be (512,16), got {source.shape}")
|
||||
if gain_values.shape != (INPUT_CHANNELS, OUTPUT_CHANNELS, HYBRID_BANDS):
|
||||
raise ValueError(f"gains must be (16,2,77), got {gain_values.shape}")
|
||||
direct = np.ascontiguousarray(
|
||||
gain_values, dtype=np.complex128).view(np.float64)
|
||||
if sends.shape != (INPUT_CHANNELS,):
|
||||
raise ValueError(f"room_sends must be (16,), got {sends.shape}")
|
||||
output = np.empty((BLOCK_SAMPLES, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
status = self._lib.ejoc_binaural_renderer_process(
|
||||
self._handle,
|
||||
self._f64_pointer(source),
|
||||
self._f64_pointer(direct),
|
||||
self._f64_pointer(sends),
|
||||
float(output_gain),
|
||||
self._f64_pointer(output),
|
||||
)
|
||||
if status:
|
||||
self._raise("process", status)
|
||||
return output
|
||||
|
||||
def close(self):
|
||||
handle = getattr(self, "_handle", None)
|
||||
if handle:
|
||||
self._lib.ejoc_binaural_renderer_destroy(handle)
|
||||
self._handle = None
|
||||
@@ -0,0 +1,272 @@
|
||||
"""JOC frame adapter for the public SOFA binaural backend."""
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
from binaural_metadata import OamdPositionTimeline
|
||||
from public_filterbank import ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
from sofa_hrtf_field import (
|
||||
DEFAULT_HRTF_CACHE_DIR,
|
||||
DEFAULT_PROJECTION_RIDGE,
|
||||
DEFAULT_SH_RIDGE,
|
||||
)
|
||||
|
||||
|
||||
SAMPLE_RATE = 48000
|
||||
FRAME_SAMPLES = 1536
|
||||
BINAURAL_BLOCK_SAMPLES = 512
|
||||
QMF_HOP_SAMPLES = 64
|
||||
BINAURAL_LATENCY_SAMPLES = ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
SOURCE_CHANNELS = 16
|
||||
OUTPUT_CHANNELS = 2
|
||||
PROJECT_DIR = Path(__file__).resolve().parent.parent
|
||||
DEFAULT_HRTF_DIR = PROJECT_DIR / "HRTF"
|
||||
DEFAULT_SOFA_HRTF = DEFAULT_HRTF_DIR / "binaural.sofa"
|
||||
|
||||
|
||||
def _resolve_hrtf_file(path: str | Path, suffix: str, label: str) -> Path:
|
||||
target = Path(path).expanduser().resolve()
|
||||
if target.suffix.lower() != suffix:
|
||||
raise ValueError(f"{label} must use the {suffix} extension: {target}")
|
||||
if not target.is_file():
|
||||
raise FileNotFoundError(f"{label} not found: {target}")
|
||||
return target
|
||||
|
||||
|
||||
def resolve_sofa_hrtf(path: str | Path) -> Path:
|
||||
"""Resolve an explicitly selected public SOFA source."""
|
||||
return _resolve_hrtf_file(path, ".sofa", "SOFA HRTF")
|
||||
|
||||
|
||||
def resolve_compiled_hrtf_cache(path: str | Path) -> Path:
|
||||
"""Resolve an explicitly selected JOC compiled HRTF cache."""
|
||||
return _resolve_hrtf_file(path, ".jochrtf", "compiled HRTF cache")
|
||||
|
||||
|
||||
class SofaBinauralRenderer:
|
||||
"""Render interleaved LFE plus fifteen JOC objects to stereo.
|
||||
|
||||
The adapter owns frame buffering and sample-timed OAMD updates. The
|
||||
backend owns the 64-QMF/77-hybrid state, the 961-sample latency policy,
|
||||
per-object direct/early state, and the shared late room.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self, backend, *,
|
||||
mode: str = "mid",
|
||||
object_delay_samples: int = 1473,
|
||||
tail_seconds: float = 5.0,
|
||||
chunk_frames: int = 64):
|
||||
required_interface = (
|
||||
"source_count", "default_profile", "set_source", "process",
|
||||
"finish", "finish_output_capacity", "info")
|
||||
missing = [name for name in required_interface if not hasattr(backend, name)]
|
||||
if missing:
|
||||
raise TypeError(
|
||||
f"backend must implement the binaural backend interface; "
|
||||
f"missing: {', '.join(missing)}")
|
||||
if backend.source_count != SOURCE_CHANNELS:
|
||||
raise ValueError(f"JOC binaural backend must have {SOURCE_CHANNELS} sources")
|
||||
if backend.default_profile != str(mode).lower():
|
||||
raise ValueError("backend default profile does not match renderer mode")
|
||||
if int(object_delay_samples) < 0:
|
||||
raise ValueError("object_delay_samples must be non-negative")
|
||||
if not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0:
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
if int(chunk_frames) <= 0:
|
||||
raise ValueError("chunk_frames must be positive")
|
||||
|
||||
self.backend = backend
|
||||
self.mode = str(mode).lower()
|
||||
self.object_delay_samples = int(object_delay_samples)
|
||||
self.tail_seconds = float(tail_seconds)
|
||||
self.chunk_frames = int(chunk_frames)
|
||||
self.chunk_samples = self.chunk_frames * FRAME_SAMPLES
|
||||
self.dsp_backend = getattr(backend, "dsp_backend", "python-sofa")
|
||||
self.timeline = OamdPositionTimeline(15)
|
||||
|
||||
self._input_buffer = np.empty(
|
||||
(self.chunk_samples, SOURCE_CHANNELS), dtype=np.float64)
|
||||
self._buffer_used = 0
|
||||
self.input_samples = 0
|
||||
self.processed_input_samples = 0
|
||||
self.output_samples = 0
|
||||
self.finished = False
|
||||
self.metadata_block_updates = 0
|
||||
|
||||
@classmethod
|
||||
def from_sofa(
|
||||
cls, sofa: str | Path, *,
|
||||
mode: str = "mid",
|
||||
cache_policy: str = "memory",
|
||||
cache_dir: str | Path | None = DEFAULT_HRTF_CACHE_DIR,
|
||||
shell_radius_m: float = 1.0,
|
||||
projection_ridge: float = DEFAULT_PROJECTION_RIDGE,
|
||||
sh_ridge: float = DEFAULT_SH_RIDGE,
|
||||
object_delay_samples: int = 1473,
|
||||
tail_seconds: float = 5.0,
|
||||
output_gain: float = 1.0,
|
||||
chunk_frames: int = 64) -> "SofaBinauralRenderer":
|
||||
source = resolve_sofa_hrtf(sofa)
|
||||
backend = SofaBinauralBackend.from_sofa(
|
||||
source,
|
||||
source_count=SOURCE_CHANNELS,
|
||||
default_profile=mode,
|
||||
output_gain=output_gain,
|
||||
cache_policy=cache_policy,
|
||||
cache_dir=cache_dir,
|
||||
shell_radius_m=shell_radius_m,
|
||||
projection_ridge=projection_ridge,
|
||||
sh_ridge=sh_ridge)
|
||||
return cls(
|
||||
backend,
|
||||
mode=mode,
|
||||
object_delay_samples=object_delay_samples,
|
||||
tail_seconds=tail_seconds,
|
||||
chunk_frames=chunk_frames)
|
||||
|
||||
@classmethod
|
||||
def from_compiled_cache(
|
||||
cls, cache: str | Path, *,
|
||||
mode: str = "mid",
|
||||
object_delay_samples: int = 1473,
|
||||
tail_seconds: float = 5.0,
|
||||
output_gain: float = 1.0,
|
||||
chunk_frames: int = 64) -> "SofaBinauralRenderer":
|
||||
source = resolve_compiled_hrtf_cache(cache)
|
||||
backend = SofaBinauralBackend.from_compiled_cache(
|
||||
source,
|
||||
source_count=SOURCE_CHANNELS,
|
||||
default_profile=mode,
|
||||
output_gain=output_gain)
|
||||
return cls(
|
||||
backend,
|
||||
mode=mode,
|
||||
object_delay_samples=object_delay_samples,
|
||||
tail_seconds=tail_seconds,
|
||||
chunk_frames=chunk_frames)
|
||||
|
||||
@property
|
||||
def finish_capacity_samples(self) -> int:
|
||||
return self.backend.finish_output_capacity(self.tail_seconds)
|
||||
|
||||
def _append_input(self, samples: np.ndarray) -> list[np.ndarray]:
|
||||
outputs = []
|
||||
source = np.asarray(samples, dtype=np.float64)
|
||||
position = 0
|
||||
while position < len(source):
|
||||
count = min(self.chunk_samples - self._buffer_used, len(source) - position)
|
||||
self._input_buffer[self._buffer_used:self._buffer_used + count] = (
|
||||
source[position:position + count])
|
||||
self._buffer_used += count
|
||||
position += count
|
||||
if self._buffer_used == self.chunk_samples:
|
||||
outputs.append(self._process_samples(self._input_buffer))
|
||||
self._buffer_used = 0
|
||||
return outputs
|
||||
|
||||
def render_frame(self, objects16, payload=None, metadata_offset=None,
|
||||
*, outer_sample_offset=0) -> np.ndarray:
|
||||
"""Submit one 1536-sample reconstructed frame and its ID11 payload."""
|
||||
if self.finished:
|
||||
raise RuntimeError("binaural renderer is already finished")
|
||||
source = np.asarray(objects16)
|
||||
if source.shape != (FRAME_SAMPLES, SOURCE_CHANNELS):
|
||||
raise ValueError(
|
||||
f"binaural frame must have shape ({FRAME_SAMPLES},{SOURCE_CHANNELS}), "
|
||||
f"got {source.shape}")
|
||||
frame_start = self.input_samples
|
||||
metadata_delay = (self.object_delay_samples if metadata_offset is None
|
||||
else int(metadata_offset))
|
||||
if metadata_delay < 0:
|
||||
raise ValueError("metadata_offset must be non-negative")
|
||||
if payload is not None:
|
||||
self.timeline.submit_payload(
|
||||
payload,
|
||||
frame_start_sample=frame_start,
|
||||
outer_sample_offset=int(outer_sample_offset),
|
||||
object_delay_samples=metadata_delay,
|
||||
processed_sample=self.processed_input_samples,
|
||||
)
|
||||
self.metadata_block_updates += 1
|
||||
self.input_samples += FRAME_SAMPLES
|
||||
chunks = self._append_input(source)
|
||||
if not chunks:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
return np.concatenate(chunks, axis=0) if len(chunks) > 1 else chunks[0]
|
||||
|
||||
def _set_block_parameters(self, sample: int) -> None:
|
||||
positions = self.timeline.positions_at(sample)
|
||||
self.backend.set_source(
|
||||
0, (0.0, 1.0, 0.0), profile=self.mode,
|
||||
special_lfe=True)
|
||||
for object_index in range(15):
|
||||
self.backend.set_source(
|
||||
object_index + 1,
|
||||
positions[object_index],
|
||||
profile=self.mode)
|
||||
|
||||
def _process_samples(self, source: np.ndarray) -> np.ndarray:
|
||||
values = np.asarray(source, dtype=np.float64)
|
||||
if values.ndim != 2 or values.shape[1] != SOURCE_CHANNELS:
|
||||
raise ValueError(f"expected [samples,{SOURCE_CHANNELS}], got {values.shape}")
|
||||
if len(values) % BINAURAL_BLOCK_SAMPLES:
|
||||
raise ValueError("binaural input must be divisible by 512 samples")
|
||||
outputs = []
|
||||
block_base = self.processed_input_samples
|
||||
for start in range(0, len(values), BINAURAL_BLOCK_SAMPLES):
|
||||
sample = block_base + start
|
||||
self._set_block_parameters(sample)
|
||||
outputs.append(self.backend.process(
|
||||
values[start:start + BINAURAL_BLOCK_SAMPLES]))
|
||||
self.processed_input_samples += len(values)
|
||||
nonempty = [value for value in outputs if len(value)]
|
||||
if not nonempty:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
output = np.concatenate(nonempty, axis=0)
|
||||
self.output_samples += len(output)
|
||||
return output
|
||||
|
||||
def finish(self) -> np.ndarray:
|
||||
"""Process pending source samples and drain early/late room state once."""
|
||||
if self.finished:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
outputs: list[np.ndarray] = []
|
||||
if self._buffer_used:
|
||||
outputs.append(self._process_samples(
|
||||
self._input_buffer[:self._buffer_used]))
|
||||
self._buffer_used = 0
|
||||
outputs.append(self.backend.finish(tail_seconds=self.tail_seconds))
|
||||
self.finished = True
|
||||
nonempty = [value for value in outputs if len(value)]
|
||||
if not nonempty:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
output = np.concatenate(nonempty, axis=0)
|
||||
self.output_samples += len(outputs[-1])
|
||||
return output
|
||||
|
||||
def close(self) -> None:
|
||||
self.finished = True
|
||||
|
||||
@property
|
||||
def backend_info(self) -> dict:
|
||||
info = self.backend.info()
|
||||
info.update({
|
||||
"adapter": "JOC 1536-frame / 512-sample metadata",
|
||||
"dsp_backend": self.dsp_backend,
|
||||
"mode": self.mode,
|
||||
"latency_compensated_samples": BINAURAL_LATENCY_SAMPLES,
|
||||
"object_delay_samples": self.object_delay_samples,
|
||||
"tail_seconds": self.tail_seconds,
|
||||
"metadata_payloads": self.timeline.payload_count,
|
||||
"metadata_position_transitions": self.timeline.transition_count,
|
||||
"input_samples": self.input_samples,
|
||||
"source_samples_processed": self.processed_input_samples,
|
||||
"output_samples_before_tail_trim": self.output_samples,
|
||||
"thread_safe": False,
|
||||
})
|
||||
return info
|
||||
+30
-5
@@ -163,7 +163,13 @@ def _marker_offsets(data):
|
||||
|
||||
|
||||
def find_joc_emdf(frame):
|
||||
"""返回同步帧中唯一、连续且包含 ID11/ID14 的 JOC EMDF 容器。"""
|
||||
"""返回同步帧中唯一、顶层连续且包含 ID11/ID14 的 JOC EMDF 容器。
|
||||
|
||||
EMDF payload 是不透明字节串,其中可能自然出现另一个 ``0x5838``。若从这个
|
||||
内嵌 marker 开始的后续随机位恰好也能通过容器语法探测,它仍不是一个独立的
|
||||
transport 容器。因此,候选的起点一旦落在较早 JOC 容器的声明范围内,就只把
|
||||
它记作内嵌伪候选,不参与“多个容器”的判定。
|
||||
"""
|
||||
matches = []
|
||||
offsets = _marker_offsets(frame)
|
||||
parsed_candidates = []
|
||||
@@ -193,13 +199,32 @@ def find_joc_emdf(frame):
|
||||
"candidate_parse_errors": parse_errors,
|
||||
"repair_hint": "检查 EMDF 是否跨多个 audio-block skip field 分片,或 payload config 是否变化",
|
||||
})
|
||||
if len(matches) != 1:
|
||||
starts = [item.start_bit for item in matches]
|
||||
top_level_matches = []
|
||||
nested_matches = []
|
||||
for container in sorted(matches, key=lambda item: item.start_bit):
|
||||
parent = next((candidate for candidate in top_level_matches
|
||||
if candidate.start_bit < container.start_bit <
|
||||
candidate.start_bit + len(candidate.raw) * 8), None)
|
||||
if parent is None:
|
||||
top_level_matches.append(container)
|
||||
else:
|
||||
nested_matches.append({
|
||||
"start_bit": container.start_bit,
|
||||
"end_bit": container.start_bit + len(container.raw) * 8,
|
||||
"parent_start_bit": parent.start_bit,
|
||||
"parent_end_bit": parent.start_bit + len(parent.raw) * 8,
|
||||
})
|
||||
if len(top_level_matches) != 1:
|
||||
starts = [item.start_bit for item in top_level_matches]
|
||||
raise UnsupportedVariantError(
|
||||
"emdf_transport", "multiple_joc_containers",
|
||||
"同步帧中存在多个可用 JOC EMDF,当前无法自动选择",
|
||||
details={"syncframe_bytes": len(frame), "joc_container_start_bits": starts})
|
||||
return matches[0]
|
||||
details={
|
||||
"syncframe_bytes": len(frame),
|
||||
"joc_container_start_bits": starts,
|
||||
"nested_joc_candidates": nested_matches,
|
||||
})
|
||||
return top_level_matches[0]
|
||||
|
||||
|
||||
def parse_container(data):
|
||||
|
||||
+96
-44
@@ -2,11 +2,11 @@
|
||||
|
||||
范围:
|
||||
- EMDF ID14 的 joc_header、joc_info 和 Huffman joc_data;
|
||||
- 差分解码得到 joc_mix_mtx_q;
|
||||
- 差分还原得到 joc_mix_mtx_q(dense MTX 与 sparse IDX/VEC 两条语法);
|
||||
- 去量化得到 joc_mix_mtx_dq;
|
||||
- 位流自洽验证(joc_data 后剩余 = padding_bits 0..7 + 可能 joc_ext_data)
|
||||
后续的时间插值、QMF/时域重建和 ``joc_clipgain`` 位于 ``renderer.py``。
|
||||
Sparse 分支仍缺少实际样本验证。
|
||||
公式与符号定义见 ``docs/math.md`` 第 2 节。
|
||||
"""
|
||||
from pathlib import Path
|
||||
|
||||
@@ -23,6 +23,7 @@ _HUFF_NAMES = (
|
||||
"joc_huff_code_7ch_pos_index_sparse",
|
||||
)
|
||||
|
||||
|
||||
def _load_huff_tables():
|
||||
with np.load(_TABLES_PATH) as tables:
|
||||
return {
|
||||
@@ -30,18 +31,24 @@ def _load_huff_tables():
|
||||
for name in _HUFF_NAMES
|
||||
}
|
||||
|
||||
|
||||
H = _load_huff_tables()
|
||||
|
||||
JOC_NUM_CHANNELS = {0: 5, 1: 7, 2: 7, 3: 5, 4: 7} # Table 33
|
||||
JOC_NUM_BANDS = {0: 1, 1: 3, 2: 5, 3: 7, 4: 9, 5: 12, 6: 15, 7: 23} # Table 35
|
||||
_PUBLIC_TABLE39_EXCERPT_UNUSED = { # 仅保留作表格差异说明;渲染映射在 joc_qmf.py。
|
||||
23: [0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22],
|
||||
}
|
||||
JOC_NUM_QUANT = {0: 96, 1: 192} # Table 51
|
||||
# dense 的量化零点就是 nquant/2;sparse 的递推起点比它高 2 个量化步。
|
||||
JOC_DENSE_OFFSET = {0: 48, 1: 96}
|
||||
JOC_SPARSE_OFFSET = {0: 50, 1: 100}
|
||||
|
||||
|
||||
class BR:
|
||||
"""MSB-first 位读取器;位置以载荷内的 bit offset 表示。"""
|
||||
|
||||
def __init__(self, data, pos=0):
|
||||
self.d = data
|
||||
self.p = pos
|
||||
|
||||
def bits(self, n):
|
||||
v = 0
|
||||
for _ in range(n):
|
||||
@@ -53,18 +60,19 @@ class BR:
|
||||
def huff_decode(tree, br):
|
||||
node = 0
|
||||
while node >= 0:
|
||||
b = br.bits(1)
|
||||
node = tree[node][b]
|
||||
node = tree[node][br.bits(1)]
|
||||
return -node - 1
|
||||
|
||||
|
||||
def get_huff_code(mode, typ, nch):
|
||||
if typ == "IDX":
|
||||
return H["joc_huff_code_5ch_pos_index_sparse" if nch == 5 else "joc_huff_code_7ch_pos_index_sparse"]
|
||||
return H["joc_huff_code_5ch_pos_index_sparse" if nch == 5
|
||||
else "joc_huff_code_7ch_pos_index_sparse"]
|
||||
if typ == "VEC":
|
||||
return H["joc_huff_code_coarse_coeff_sparse" if mode == 0 else "joc_huff_code_fine_coeff_sparse"]
|
||||
# MTX
|
||||
return H["joc_huff_code_coarse_generic" if mode == 0 else "joc_huff_code_fine_generic"]
|
||||
return H["joc_huff_code_coarse_coeff_sparse" if mode == 0
|
||||
else "joc_huff_code_fine_coeff_sparse"]
|
||||
return H["joc_huff_code_coarse_generic" if mode == 0
|
||||
else "joc_huff_code_fine_generic"] # MTX
|
||||
|
||||
|
||||
def parse_joc(payload):
|
||||
@@ -76,6 +84,8 @@ def parse_joc(payload):
|
||||
out["ext_config_idx"] = br.bits(3)
|
||||
n_objects = out["num_objects_bits"] + 1
|
||||
n_channels = JOC_NUM_CHANNELS.get(out["dmx_config_idx"])
|
||||
if n_channels is None:
|
||||
raise ValueError(f"未知的 JOC downmix 配置 {out['dmx_config_idx']}")
|
||||
out["n_objects"], out["n_channels"] = n_objects, n_channels
|
||||
out["clipgain_x_bits"] = br.bits(3)
|
||||
out["clipgain_y_bits"] = br.bits(5)
|
||||
@@ -98,78 +108,120 @@ def parse_joc(payload):
|
||||
o["offset_ts"] = [br.bits(5) + 1 for _ in range(o["n_dpoints"])]
|
||||
objs.append(o)
|
||||
out["objs"] = objs
|
||||
# joc_data(Huffman)
|
||||
# joc_data(Huffman):dense 逐声道逐带读 MTX;sparse 读 IDX 后读 VEC。
|
||||
for obj, o in enumerate(objs):
|
||||
if not o["present"]:
|
||||
continue
|
||||
nquant = 96 if o["quant_idx"] == 0 else 192
|
||||
o["channel_idx"] = []
|
||||
o["vec"] = []
|
||||
o["mtx"] = []
|
||||
for dp in range(o["n_dpoints"]):
|
||||
if o["sparse"] == 1:
|
||||
# Sparse JOC 使用 VEC/IDX Huffman 树;此分支尚无真实码流验证。
|
||||
ci0 = br.bits(3)
|
||||
tree = get_huff_code(n_channels, "IDX", n_channels)
|
||||
ci = [ci0] + [huff_decode(tree, br) for _ in range(o["n_bands"] - 1)]
|
||||
o["channel_idx"].append(ci)
|
||||
tree = get_huff_code(o["quant_idx"], "IDX", n_channels)
|
||||
idx = [br.bits(3)]
|
||||
idx += [huff_decode(tree, br) for _ in range(o["n_bands"] - 1)]
|
||||
tree = get_huff_code(o["quant_idx"], "VEC", n_channels)
|
||||
vec = [huff_decode(tree, br) for _ in range(o["n_bands"])]
|
||||
o["vec"].append(vec)
|
||||
o["channel_idx"].append(idx)
|
||||
o["vec"].append([huff_decode(tree, br) for _ in range(o["n_bands"])])
|
||||
o["mtx"].append(None)
|
||||
else:
|
||||
tree = get_huff_code(o["quant_idx"], "MTX", n_channels)
|
||||
mtx = [[huff_decode(tree, br) for _ in range(o["n_bands"])]
|
||||
for _ in range(n_channels)]
|
||||
o["mtx"].append(mtx)
|
||||
o["channel_idx"].append(None)
|
||||
o["vec"].append(None)
|
||||
out["data_end_bits"] = br.p
|
||||
out["remaining_bits"] = len(payload) * 8 - br.p
|
||||
out["tail_bytes"] = payload[br.p // 8:]
|
||||
return out
|
||||
|
||||
|
||||
def reconstruct_dense(o, dp, n_ch):
|
||||
"""Dense 差分还原 → 量化矩阵 ``[ch][pb]``。
|
||||
|
||||
每个核心声道各自从 ``nquant/2``(去量化 0)出发,沿参数带累加 MTX 符号。
|
||||
"""
|
||||
nquant = JOC_NUM_QUANT[o["quant_idx"]]
|
||||
offset = JOC_DENSE_OFFSET[o["quant_idx"]]
|
||||
mtx = o["mtx"][dp]
|
||||
if mtx is None:
|
||||
raise ValueError("Dense JOC 对象缺少 MTX 符号")
|
||||
q = np.zeros((n_ch, o["n_bands"]), dtype=np.int64)
|
||||
for ch in range(n_ch):
|
||||
q[ch][0] = (offset + mtx[ch][0]) % nquant
|
||||
for pb in range(1, o["n_bands"]):
|
||||
q[ch][pb] = (q[ch][pb - 1] + mtx[ch][pb]) % nquant
|
||||
return q
|
||||
|
||||
|
||||
def reconstruct_sparse(o, dp, n_ch):
|
||||
"""Sparse 差分还原 → 量化矩阵 ``[ch][pb]``。
|
||||
|
||||
每个参数带只有一个 active 声道:
|
||||
- ``active[0]`` 是 3 bit 绝对声道号,其后由 IDX 符号累加得到,
|
||||
因此递推锚点是**已重建**的 active 声道,而不是编码器送的符号本身;
|
||||
- 系数是一个跨参数带连续的单累加器(active 声道切换时**不**重置),
|
||||
起点为 sparse offset,增量为 VEC 符号;
|
||||
- 非 active 项取 ``nquant/2``,即去量化后的 0。
|
||||
|
||||
IDX 符号取值恒为 ``0..n_ch-1``(Huffman 叶数即声道数),
|
||||
故 ``(active + idx) % n_ch`` 与单次条件减等价。
|
||||
"""
|
||||
if n_ch not in (5, 7):
|
||||
raise ValueError(f"Sparse JOC 需要 5 或 7 个核心声道,实际 {n_ch}")
|
||||
nquant = JOC_NUM_QUANT[o["quant_idx"]]
|
||||
offset = JOC_SPARSE_OFFSET[o["quant_idx"]]
|
||||
n_bands = o["n_bands"]
|
||||
idx = o["channel_idx"][dp]
|
||||
vec = o["vec"][dp]
|
||||
if idx is None or vec is None:
|
||||
raise ValueError("Sparse JOC 对象缺少 channel_idx/vec 符号")
|
||||
if len(idx) != n_bands or len(vec) != n_bands:
|
||||
raise ValueError(
|
||||
f"Sparse JOC 维度不符: bands={n_bands}, idx={len(idx)}, vec={len(vec)}")
|
||||
if not 0 <= idx[0] < n_ch:
|
||||
raise ValueError(f"Sparse JOC 初始声道 {idx[0]} 超出 {n_ch} 个声道")
|
||||
|
||||
q = np.full((n_ch, n_bands), nquant // 2, dtype=np.int64)
|
||||
active = idx[0]
|
||||
coefficient = offset
|
||||
for pb in range(n_bands):
|
||||
if pb:
|
||||
active = (active + idx[pb]) % n_ch
|
||||
coefficient = (coefficient + vec[pb]) % nquant
|
||||
q[active][pb] = coefficient
|
||||
return q
|
||||
|
||||
|
||||
def diff_decode(out):
|
||||
"""6.6.2:差分解码 → joc_mix_mtx_q[obj][dp][ch][pb]。"""
|
||||
"""差分还原 → joc_mix_mtx_q[obj][dp][ch][pb]。"""
|
||||
mix_q = {}
|
||||
n_ch = out["n_channels"]
|
||||
for obj, o in enumerate(out["objs"]):
|
||||
if not o["present"]:
|
||||
continue
|
||||
nquant = 96 if o["quant_idx"] == 0 else 192
|
||||
q = np.zeros((o["n_dpoints"], n_ch, o["n_bands"]), dtype=np.int64)
|
||||
for dp in range(o["n_dpoints"]):
|
||||
if o["sparse"] == 1:
|
||||
# Sparse 差分路径尚无真实码流验证。
|
||||
offset = 50 if o["quant_idx"] == 0 else 100
|
||||
ci = o["channel_idx"][dp]
|
||||
vec = o["vec"][dp]
|
||||
for pb in range(o["n_bands"]):
|
||||
ci_mod = ci[0] if pb == 0 else (ci[pb - 1] + ci[pb]) % n_ch
|
||||
for ch in range(n_ch):
|
||||
if ch == ci_mod:
|
||||
if pb == 0:
|
||||
q[dp][ch][pb] = (offset + vec[pb]) % nquant
|
||||
else:
|
||||
q[dp][ch][pb] = (q[dp][ch][pb - 1] + vec[pb]) % nquant
|
||||
else:
|
||||
q[dp][ch][pb] = offset
|
||||
q[dp] = reconstruct_sparse(o, dp, n_ch)
|
||||
else:
|
||||
offset = 48 if o["quant_idx"] == 0 else 96
|
||||
mtx = o["mtx"][dp]
|
||||
for ch in range(n_ch):
|
||||
q[dp][ch][0] = (offset + mtx[ch][0]) % nquant
|
||||
for pb in range(1, o["n_bands"]):
|
||||
q[dp][ch][pb] = (q[dp][ch][pb - 1] + mtx[ch][pb]) % nquant
|
||||
q[dp] = reconstruct_dense(o, dp, n_ch)
|
||||
mix_q[obj] = q
|
||||
return mix_q
|
||||
|
||||
|
||||
def dequantize(out, mix_q):
|
||||
"""6.6.4:去量化 → joc_mix_mtx_dq。"""
|
||||
"""去量化 → joc_mix_mtx_dq。
|
||||
|
||||
Sparse 的非 active 项在 ``joc_mix_mtx_q`` 中取 ``nquant/2``,
|
||||
因此与 dense 共用同一条去量化公式即得到 0。
|
||||
"""
|
||||
mix_dq = {}
|
||||
for obj, o in enumerate(out["objs"]):
|
||||
if not o["present"]:
|
||||
continue
|
||||
nquant = 96 if o["quant_idx"] == 0 else 192
|
||||
nquant = JOC_NUM_QUANT[o["quant_idx"]]
|
||||
q = mix_q[obj]
|
||||
dq = (q.astype(np.float64) - nquant / 2) * 820 / (4096 * (1 + o["quant_idx"]))
|
||||
mix_dq[obj] = dq
|
||||
|
||||
@@ -246,20 +246,6 @@ def inspect(index, limit=None, print_frames=False):
|
||||
"parser_error": str(exc),
|
||||
"repair_hint": "检查 JOC header、对象数、参数带、Huffman 或扩展字段",
|
||||
}) from exc
|
||||
sparse = [i for i, obj in enumerate(parsed["objs"]) if obj["present"] and obj["sparse"]]
|
||||
if sparse:
|
||||
raise UnsupportedVariantError(
|
||||
"joc", "sparse_joc",
|
||||
"发现尚未验证的 Sparse JOC 帧",
|
||||
frame=frame_number,
|
||||
details={
|
||||
"sparse_objects": sparse,
|
||||
"downmix_config": parsed["dmx_config_idx"],
|
||||
"extension_config": parsed["ext_config_idx"],
|
||||
"objects": parsed["n_objects"],
|
||||
"payload": bytes_descriptor(subs[14]),
|
||||
"repair_hint": "需要 Sparse JOC 实际样本及对应输出建立回归后再启用",
|
||||
})
|
||||
if parsed["n_channels"] != 5 or parsed["n_objects"] > 15:
|
||||
raise UnsupportedVariantError(
|
||||
"joc", "unsupported_configuration",
|
||||
|
||||
@@ -208,8 +208,6 @@ class NativeJocRenderer:
|
||||
for object_index, info in enumerate(out["objs"]):
|
||||
if not info["present"]:
|
||||
continue
|
||||
if info["sparse"]:
|
||||
raise ValueError("native core does not accept unvalidated Sparse JOC")
|
||||
bands = int(info["n_bands"])
|
||||
points = int(info["n_dpoints"])
|
||||
if bands > MAX_BANDS or points > MAX_DPOINTS:
|
||||
|
||||
+503
-127
@@ -1,12 +1,14 @@
|
||||
"""OAMD 位载荷 → 16 个对象槽的 q1/q2/q3 增量状态。
|
||||
|
||||
槽 0 是 bed/LFE;槽 1..15 对应输出 ch1..15 的对象元数据。
|
||||
依据 ETSI TS 103 420 V1.2.1 clause 5。槽 0 为 bed/LFE,槽 1..15 为输出
|
||||
ch1..15;bed/ISF/未激活对象没有位置字段,保持上一帧位置。element 目录按声明长度
|
||||
驱动,未知 element 按边界跳过;个别编码器的 ``oa_element_size`` 比实际内容短时以
|
||||
结构解析为准,差异记入 ``diagnostics``,不算变体错误。
|
||||
"""
|
||||
import numpy as np
|
||||
|
||||
from variant_error import UnsupportedVariantError, bytes_descriptor
|
||||
|
||||
Q1_OFF = 192
|
||||
N_Q12 = 62
|
||||
N_Q3 = 15
|
||||
SAMPLE_OFFSET_INDEX = (8, 16, 18, 24)
|
||||
@@ -15,6 +17,21 @@ RAMP_DURATION_INDEX = (
|
||||
32, 64, 128, 256, 320, 480, 1000, 1001,
|
||||
1024, 1600, 1601, 1602, 1920, 2000, 2002, 2048,
|
||||
)
|
||||
OBJECT_ELEMENT_ID = 1
|
||||
# 5.6.0.11 Table 11b:ISF 类型 → 对象数。
|
||||
ISF_OBJECT_COUNTS = (4, 8, 10, 14, 15, 30)
|
||||
# 5.6.1.1.4 Table 12:10 bit 标准 bed 掩码按位对应的声道(LSB = RC_L/RC_R)。
|
||||
BED_CHANNEL_LABELS = (
|
||||
"RC_L/RC_R", "RC_C", "RC_LFE", "RC_LS/RC_RS", "RC_LB/RC_RB",
|
||||
"RC_TFL/RC_TFR", "RC_TSL/RC_TSR", "RC_TBL/RC_TBR", "RC_LW/RC_RW", "RC_LFE2",
|
||||
)
|
||||
# 5.6.1.1.5 Table 13:17 bit 非标准 bed 掩码按位对应的声道标签。
|
||||
NONSTD_BED_LABELS = (
|
||||
"RC_LFE2", "RC_RW", "RC_LW", "RC_TBR", "RC_TBL", "RC_TSR", "RC_TSL",
|
||||
"RC_TFR", "RC_TFL", "RC_RB", "RC_LB", "RC_RS", "RC_LS", "RC_LFE", "RC_C",
|
||||
"RC_R", "RC_L",
|
||||
)
|
||||
MAX_SLOTS = 16
|
||||
|
||||
|
||||
def q_of(k, n):
|
||||
@@ -24,36 +41,58 @@ def q_of(k, n):
|
||||
|
||||
def _payload_bits(bits_one):
|
||||
if isinstance(bits_one, (bytes, bytearray, memoryview)):
|
||||
src = np.frombuffer(bits_one, dtype=np.uint8)
|
||||
raw_payload = bytes(bits_one)
|
||||
bits = np.unpackbits(
|
||||
np.frombuffer(raw_payload, dtype=np.uint8), bitorder="big")
|
||||
else:
|
||||
src = np.asarray(bits_one, dtype=np.uint8)
|
||||
if src.ndim == 1 and src.shape[0] in (536, 552):
|
||||
bits = src
|
||||
elif src.size in (67, 69):
|
||||
bits = np.unpackbits(src.reshape(-1), bitorder="big")
|
||||
else:
|
||||
raw = (bytes(bits_one) if isinstance(bits_one, (bytes, bytearray, memoryview))
|
||||
else np.asarray(bits_one, dtype=np.uint8).tobytes())
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", f"payload_length_{len(raw)}B",
|
||||
f"发现未覆盖的 OAMD 载荷长度 {len(raw)}B",
|
||||
details={
|
||||
"supported_payload_bytes": [67, 69],
|
||||
"payload": bytes_descriptor(raw),
|
||||
"repair_hint": "检查 OAMD header、element 数量及可选字段造成的位偏移变化",
|
||||
})
|
||||
raw_payload = np.packbits(bits, bitorder="big").tobytes()
|
||||
src = np.asarray(bits_one)
|
||||
if src.ndim != 1:
|
||||
raw = np.asarray(bits_one, dtype=np.uint8).tobytes()
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "payload_shape",
|
||||
"OAMD 载荷必须是一维 byte 或 bit 序列",
|
||||
details={"shape": list(src.shape), "payload": bytes_descriptor(raw)})
|
||||
is_bit_vector = bool(src.size) and bool(np.all((src == 0) | (src == 1)))
|
||||
if is_bit_vector:
|
||||
if src.size % 8:
|
||||
raw = np.packbits(src.astype(np.uint8), bitorder="big").tobytes()
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "payload_bit_alignment",
|
||||
"OAMD bit 载荷没有按整字节对齐",
|
||||
details={
|
||||
"payload_bits": int(src.size),
|
||||
"payload": bytes_descriptor(raw),
|
||||
})
|
||||
bits = src.astype(np.uint8, copy=False)
|
||||
raw_payload = np.packbits(bits, bitorder="big").tobytes()
|
||||
else:
|
||||
try:
|
||||
values = src.astype(np.int64, copy=False)
|
||||
except (TypeError, ValueError, OverflowError) as exc:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "payload_type",
|
||||
"OAMD 载荷不能转换为 byte 序列",
|
||||
details={"dtype": str(src.dtype), "parser_error": str(exc)}) from exc
|
||||
if np.any(values < 0) or np.any(values > 255):
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "payload_byte_range",
|
||||
"OAMD byte 载荷包含 0..255 之外的值",
|
||||
details={"dtype": str(src.dtype), "payload_values": int(src.size)})
|
||||
raw_payload = values.astype(np.uint8).tobytes()
|
||||
bits = np.unpackbits(
|
||||
np.frombuffer(raw_payload, dtype=np.uint8), bitorder="big")
|
||||
return bits, raw_payload
|
||||
|
||||
|
||||
class _BitReader:
|
||||
def __init__(self, bits):
|
||||
def __init__(self, bits, position=0, limit=None):
|
||||
self.bits = bits
|
||||
self.position = 0
|
||||
self.position = int(position)
|
||||
self.limit = len(bits) if limit is None else int(limit)
|
||||
|
||||
def read(self, count):
|
||||
end = self.position + count
|
||||
if end > len(self.bits):
|
||||
if end > self.limit:
|
||||
raise ValueError(f"OAMD 位流越界: bit={self.position}, need={count}")
|
||||
value = 0
|
||||
for bit in self.bits[self.position:end]:
|
||||
@@ -65,130 +104,467 @@ class _BitReader:
|
||||
self.read(count)
|
||||
|
||||
|
||||
def _variable_bits(reader, width, max_groups=5):
|
||||
value = 0
|
||||
for _ in range(max_groups + 1):
|
||||
value += reader.read(width)
|
||||
if not reader.read(1):
|
||||
return value
|
||||
value = (value + 1) << width
|
||||
raise ValueError(f"OAMD variable_bits({width}) 延伸组过多")
|
||||
def _variable_bits_max(reader, width, max_groups):
|
||||
"""5.5.1 ``variable_bits_max(n, max_num_groups)``。"""
|
||||
value = reader.read(width)
|
||||
more = reader.read(1)
|
||||
num_group = 1
|
||||
if max_groups > num_group:
|
||||
if more:
|
||||
value = (value + 1) << width
|
||||
while more:
|
||||
value += reader.read(width)
|
||||
more = reader.read(1)
|
||||
if num_group >= max_groups:
|
||||
break
|
||||
if more:
|
||||
value = (value + 1) << width
|
||||
num_group += 1
|
||||
return value
|
||||
|
||||
|
||||
def _update_timing(bits, alternate_object_present, element_count, raw_payload):
|
||||
"""按 OAMD element/MDUpdateInfo 读取位置块的开始偏移和 ramp 时长。"""
|
||||
reader = _BitReader(bits)
|
||||
reader.position = 14
|
||||
for _ in range(element_count):
|
||||
element_index = reader.read(4)
|
||||
element_length = _variable_bits(reader, 4)
|
||||
element_end = reader.position + element_length + 1
|
||||
reader.skip(5 if alternate_object_present else 1)
|
||||
if element_index == 1:
|
||||
offset_code = reader.read(2)
|
||||
if offset_code == 0:
|
||||
sample_offset = 0
|
||||
elif offset_code == 1:
|
||||
sample_offset = SAMPLE_OFFSET_INDEX[reader.read(2)]
|
||||
elif offset_code == 2:
|
||||
sample_offset = reader.read(5)
|
||||
def _element_details(elements):
|
||||
return [{
|
||||
"ordinal": element["ordinal"],
|
||||
"element_id": element["element_id"],
|
||||
"size_bytes": element["size_bytes"],
|
||||
"header_start_bit": element["header_start_bit"],
|
||||
"body_start_bit": element["body_start_bit"],
|
||||
"body_end_bit": element["body_end_bit"],
|
||||
"parsed_end_bit": element.get("parsed_end_bit"),
|
||||
"discard_unknown": element["discard_unknown"],
|
||||
"alternate_data_id": element["alternate_data_id"],
|
||||
} for element in elements]
|
||||
|
||||
|
||||
def _syntax_error(variant, message, raw_payload, exc=None, **details):
|
||||
payload = {"payload": bytes_descriptor(raw_payload)}
|
||||
if exc is not None:
|
||||
payload["parser_error"] = str(exc)
|
||||
payload.update(details)
|
||||
return UnsupportedVariantError("oamd", variant, message, details=payload)
|
||||
|
||||
|
||||
def _parse_program_assignment(reader, raw_payload):
|
||||
"""5.5.3 ``program_assignment()``:bed / ISF / dynamic 对象数量。"""
|
||||
program = {
|
||||
"dynamic_object_only": bool(reader.read(1)),
|
||||
"lfe_present": False,
|
||||
"content_description": None,
|
||||
"bed_assignments": [],
|
||||
"num_bed_objects": 0,
|
||||
"isf_idx": None,
|
||||
"num_isf_objects": 0,
|
||||
"num_dynamic_objects": None,
|
||||
}
|
||||
if program["dynamic_object_only"]:
|
||||
program["lfe_present"] = bool(reader.read(1))
|
||||
# 5.6.4.8:对象顺序为 bed → ISF → dynamic,LFE 属于 bed,排在最前。
|
||||
program["num_bed_objects"] = 1 if program["lfe_present"] else 0
|
||||
return program
|
||||
|
||||
mask = reader.read(4)
|
||||
program["content_description"] = {
|
||||
"reserved": bool(mask & 0x8),
|
||||
"dynamic": bool(mask & 0x4),
|
||||
"isf": bool(mask & 0x2),
|
||||
"bed": bool(mask & 0x1),
|
||||
}
|
||||
if mask & 0x1:
|
||||
program["b_bed_chan_distribute"] = bool(reader.read(1))
|
||||
num_instances = (reader.read(3) + 2) if reader.read(1) else 1
|
||||
for instance in range(num_instances):
|
||||
if reader.read(1): # b_lfe_only
|
||||
program["bed_assignments"].append({
|
||||
"instance": instance, "lfe_only": True, "channels": ["RC_LFE"],
|
||||
})
|
||||
continue
|
||||
if reader.read(1): # b_standard_chan_assign
|
||||
bed_mask = reader.read(10)
|
||||
channels = []
|
||||
for bit in range(10):
|
||||
if (bed_mask >> bit) & 1:
|
||||
channels.extend(BED_CHANNEL_LABELS[bit].split("/"))
|
||||
program["bed_assignments"].append({
|
||||
"instance": instance, "lfe_only": False, "standard": True,
|
||||
"mask": bed_mask, "channels": channels,
|
||||
})
|
||||
else:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "md_sample_offset_mode",
|
||||
"OAMD 使用了当前未覆盖的 MD sample-offset 模式",
|
||||
details={"payload": bytes_descriptor(raw_payload)})
|
||||
bed_mask = reader.read(17)
|
||||
channels = [NONSTD_BED_LABELS[bit]
|
||||
for bit in range(17) if (bed_mask >> bit) & 1]
|
||||
program["bed_assignments"].append({
|
||||
"instance": instance, "lfe_only": False, "standard": False,
|
||||
"mask": bed_mask, "channels": channels,
|
||||
})
|
||||
program["num_bed_objects"] = sum(
|
||||
len(assignment["channels"]) for assignment in program["bed_assignments"])
|
||||
if mask & 0x2:
|
||||
isf_idx = reader.read(3)
|
||||
program["isf_idx"] = isf_idx
|
||||
program["num_isf_objects"] = (ISF_OBJECT_COUNTS[isf_idx]
|
||||
if isf_idx < len(ISF_OBJECT_COUNTS) else 0)
|
||||
if mask & 0x4:
|
||||
count = reader.read(5)
|
||||
if count == 0x1F:
|
||||
count += reader.read(7)
|
||||
program["num_dynamic_objects"] = count + 1
|
||||
if mask & 0x8:
|
||||
# 5.6.4.1:reserved_data_size = reserved_data_size_bits + 1(字节)。
|
||||
reader.skip((reader.read(4) + 1) * 8)
|
||||
return program
|
||||
|
||||
block_count = reader.read(3) + 1
|
||||
blocks = []
|
||||
for _block in range(block_count):
|
||||
block_offset_factor = reader.read(6)
|
||||
block_offset = sample_offset + block_offset_factor * 32
|
||||
ramp_code = reader.read(2)
|
||||
if ramp_code == 3:
|
||||
if reader.read(1):
|
||||
ramp_duration = RAMP_DURATION_INDEX[reader.read(4)]
|
||||
else:
|
||||
ramp_duration = reader.read(11)
|
||||
|
||||
def _parse_object_info_block(reader, object_index, in_bed_or_isf, raw_payload):
|
||||
"""5.5.9 ``object_info_block(0)``:一个对象的属性更新。"""
|
||||
info = {
|
||||
"object": object_index,
|
||||
"in_bed_or_isf": bool(in_bed_or_isf),
|
||||
"not_active": bool(reader.read(1)),
|
||||
}
|
||||
# blk == 0 且对象激活时 basic info 恒为 full update(status 0b01)。
|
||||
basic_status = 0 if info["not_active"] else 1
|
||||
info["basic_status"] = basic_status
|
||||
if basic_status in (1, 3):
|
||||
if basic_status == 1:
|
||||
# 5.5.10:object_basic_info[] = {true, true};
|
||||
# 5.6.4.12:array[1] = object_gain_idx,array[0] = b_default_object_priority。
|
||||
gain_present = priority_present = True
|
||||
else:
|
||||
flags = reader.read(2)
|
||||
gain_present = bool(flags & 1)
|
||||
priority_present = bool(flags & 2)
|
||||
if gain_present:
|
||||
gain_idx = reader.read(2)
|
||||
info["gain_idx"] = gain_idx
|
||||
if gain_idx == 2:
|
||||
info["gain_bits"] = reader.read(6)
|
||||
if priority_present:
|
||||
default_priority = bool(reader.read(1))
|
||||
info["default_priority"] = default_priority
|
||||
if not default_priority:
|
||||
info["priority_bits"] = reader.read(5)
|
||||
render_status = 0 if (info["not_active"] or info["in_bed_or_isf"]) else 1
|
||||
info["render_status"] = render_status
|
||||
if render_status in (1, 3):
|
||||
if render_status == 1:
|
||||
# 5.5.11:obj_render_info[] = {true, true, true, true}。
|
||||
position_present = zone_present = size_present = screen_present = True
|
||||
else:
|
||||
flags = reader.read(4)
|
||||
position_present = bool(flags & 0x1)
|
||||
zone_present = bool(flags & 0x2)
|
||||
size_present = bool(flags & 0x4)
|
||||
screen_present = bool(flags & 0x8)
|
||||
if position_present:
|
||||
# blk == 0 时 b_differential_position_specified 恒为 FALSE。
|
||||
x = reader.read(6)
|
||||
y = reader.read(6)
|
||||
z_sign = reader.read(1)
|
||||
z = reader.read(4)
|
||||
info["position"] = (x, y, z if z_sign else -z)
|
||||
if reader.read(1): # b_object_distance_specified
|
||||
if not reader.read(1): # b_object_at_infinity
|
||||
reader.read(4) # distance_factor_idx
|
||||
if zone_present:
|
||||
info["zone_constraints_idx"] = reader.read(3)
|
||||
info["enable_elevation"] = bool(reader.read(1))
|
||||
if size_present:
|
||||
size_idx = reader.read(2)
|
||||
info["object_size_idx"] = size_idx
|
||||
if size_idx == 1:
|
||||
reader.read(5)
|
||||
elif size_idx == 2:
|
||||
reader.read(15)
|
||||
if screen_present:
|
||||
if reader.read(1): # b_object_use_screen_ref
|
||||
reader.read(3) # screen_factor_bits
|
||||
reader.read(2) # depth_factor_idx
|
||||
info["snap"] = bool(reader.read(1))
|
||||
if reader.read(1): # b_additional_table_data_exists
|
||||
info["additional_table_bytes"] = reader.read(4) + 1
|
||||
reader.skip(info["additional_table_bytes"] * 8)
|
||||
return info
|
||||
|
||||
|
||||
def _parse_object_element(reader, element, raw_payload):
|
||||
"""5.5.5 ``object_element()``:时间信息 + 每个对象一条 object_info_block。"""
|
||||
object_count = element["object_count"]
|
||||
try:
|
||||
sample_offset_code = reader.read(2)
|
||||
if sample_offset_code == 0:
|
||||
sample_offset = 0
|
||||
elif sample_offset_code == 1:
|
||||
sample_offset = SAMPLE_OFFSET_INDEX[reader.read(2)]
|
||||
elif sample_offset_code == 2:
|
||||
sample_offset = reader.read(5)
|
||||
else:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "md_sample_offset_mode",
|
||||
"OAMD 使用了当前未覆盖的 MD sample-offset 模式",
|
||||
details={
|
||||
"sample_offset_code": sample_offset_code,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
block_count = reader.read(3) + 1
|
||||
blocks = []
|
||||
for _block in range(block_count):
|
||||
block_offset_factor = reader.read(6)
|
||||
ramp_code = reader.read(2)
|
||||
if ramp_code == 3:
|
||||
if reader.read(1):
|
||||
ramp_duration = RAMP_DURATION_INDEX[reader.read(4)]
|
||||
else:
|
||||
ramp_duration = RAMP_DURATIONS[ramp_code]
|
||||
blocks.append((block_offset, ramp_duration))
|
||||
if len(blocks) != 1:
|
||||
ramp_duration = reader.read(11)
|
||||
else:
|
||||
ramp_duration = RAMP_DURATIONS[ramp_code]
|
||||
blocks.append({
|
||||
"block_offset_samples": sample_offset + block_offset_factor * 32,
|
||||
"ramp_duration_samples": ramp_duration,
|
||||
})
|
||||
reserved_data_not_present = bool(reader.read(1))
|
||||
if not reserved_data_not_present:
|
||||
reader.read(5)
|
||||
objects = [
|
||||
_parse_object_info_block(
|
||||
reader, index, index < element["bed_isf_objects"], raw_payload)
|
||||
for index in range(object_count)
|
||||
]
|
||||
except UnsupportedVariantError:
|
||||
raise
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"object_element_syntax",
|
||||
"OAMD object element 字段越界或不完整",
|
||||
raw_payload, exc,
|
||||
element=_element_details([element])[0],
|
||||
object_count=object_count,
|
||||
) from exc
|
||||
return {
|
||||
"sample_offset": sample_offset,
|
||||
"blocks": blocks,
|
||||
"reserved_data_not_present": reserved_data_not_present,
|
||||
"objects": objects,
|
||||
}
|
||||
|
||||
|
||||
def _parse_elements(reader, alternate_present, object_count, program, raw_payload):
|
||||
"""5.5.4 ``oa_element_md()`` 目录;对象 element 按结构解析,其余按声明长度跳过。"""
|
||||
try:
|
||||
element_count = reader.read(4)
|
||||
if element_count == 0xF:
|
||||
element_count += reader.read(5)
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"element_count_truncated", "OAMD element 数量字段不完整",
|
||||
raw_payload, exc) from exc
|
||||
if element_count == 0:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "missing_object_element",
|
||||
"OAMD 没有声明任何 element",
|
||||
details={"payload": bytes_descriptor(raw_payload)})
|
||||
|
||||
bed_isf_objects = program["num_bed_objects"] + program["num_isf_objects"]
|
||||
elements = []
|
||||
object_element = None
|
||||
for ordinal in range(element_count):
|
||||
header_start = reader.position
|
||||
try:
|
||||
element_id = reader.read(4)
|
||||
size_bytes = _variable_bits_max(reader, 4, 4) + 1
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"element_header_truncated", "OAMD element header 不完整",
|
||||
raw_payload, exc,
|
||||
element_ordinal=ordinal, header_start_bit=header_start) from exc
|
||||
region_start = reader.position
|
||||
region_end = region_start + size_bytes * 8
|
||||
if region_end > reader.limit:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "element_bounds",
|
||||
"OAMD element 声明长度超过 payload 边界",
|
||||
details={
|
||||
"element_ordinal": ordinal,
|
||||
"element_id": element_id,
|
||||
"size_bytes": size_bytes,
|
||||
"body_start_bit": region_start,
|
||||
"body_end_bit": region_end,
|
||||
"payload_bits": reader.limit,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
try:
|
||||
alternate_data_id = (reader.read(4)
|
||||
if alternate_present else None)
|
||||
discard_unknown = bool(reader.read(1))
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"element_control_bounds", "OAMD element 太短,无法容纳控制字段",
|
||||
raw_payload, exc,
|
||||
element_ordinal=ordinal, element_id=element_id,
|
||||
size_bytes=size_bytes) from exc
|
||||
element = {
|
||||
"ordinal": ordinal,
|
||||
"element_id": element_id,
|
||||
"size_bytes": size_bytes,
|
||||
"header_start_bit": header_start,
|
||||
"body_start_bit": region_start,
|
||||
"body_end_bit": region_end,
|
||||
"data_start_bit": reader.position,
|
||||
"alternate_data_id": alternate_data_id,
|
||||
"discard_unknown": discard_unknown,
|
||||
"object_count": object_count,
|
||||
"bed_isf_objects": bed_isf_objects,
|
||||
}
|
||||
if element_id == OBJECT_ELEMENT_ID:
|
||||
if object_element is not None:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "multiple_position_blocks",
|
||||
"OAMD 一帧含多个对象位置更新块,固定位置窗口不能安全套用",
|
||||
"oamd", "multiple_object_elements",
|
||||
"OAMD 含多个 object element",
|
||||
details={
|
||||
"block_count": len(blocks),
|
||||
"blocks": blocks,
|
||||
"elements": _element_details(elements + [element]),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"repair_hint": "按 ObjectInfoBlock 顺序逐块解析坐标,再生成分段 ADM ramp",
|
||||
})
|
||||
return blocks[0]
|
||||
reader.position = element_end
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "missing_object_element",
|
||||
"OAMD 中没有 object element (element_index=1)",
|
||||
details={"payload": bytes_descriptor(raw_payload)})
|
||||
object_element = _parse_object_element(
|
||||
reader, element, raw_payload)
|
||||
element["parsed_end_bit"] = reader.position
|
||||
element["object_element"] = object_element
|
||||
# 个别编码器声明的 oa_element_size 比实际内容短;以结构解析结果为准。
|
||||
reader.position = max(region_end, reader.position)
|
||||
else:
|
||||
# trim / extended / 未知 element:按声明长度整体跳过(5.5.4)。
|
||||
element["parsed_end_bit"] = region_end
|
||||
reader.position = region_end
|
||||
elements.append(element)
|
||||
|
||||
if object_element is None:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "missing_object_element",
|
||||
"OAMD 缺少 object element",
|
||||
details={
|
||||
"elements": _element_details(elements),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
padding = reader.bits[reader.position:]
|
||||
if np.any(padding):
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "nonzero_padding",
|
||||
"OAMD payload 尾部 padding 含非零位",
|
||||
details={
|
||||
"elements": _element_details(elements),
|
||||
"padding_start_bit": reader.position,
|
||||
"padding_bits": "".join(str(int(bit)) for bit in padding[:64]),
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
return object_element, elements
|
||||
|
||||
|
||||
def frame_update(bits_one):
|
||||
"""单帧 OAMD → 位置字段增量及其 sample offset/ramp duration。"""
|
||||
bits, raw_payload = _payload_bits(bits_one)
|
||||
header = {
|
||||
"version": int((bits[0] << 1) | bits[1]),
|
||||
"objects_minus_one": int(sum(int(bits[2 + i]) << (4 - i) for i in range(5))),
|
||||
"dynamic_object_only": int(bits[7]),
|
||||
"lfe_present": int(bits[8]),
|
||||
"alternate_object_present": int(bits[9]),
|
||||
"element_count": int(sum(int(bits[10 + i]) << (3 - i) for i in range(4))),
|
||||
}
|
||||
if raw_payload[:2] != b"\x1f\x88":
|
||||
try:
|
||||
version = bits[0] << 1 | bits[1]
|
||||
except IndexError:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", f"header_signature_{raw_payload[:2].hex()}",
|
||||
"OAMD 长度已知,但 header 与当前位置字段布局不一致",
|
||||
"oamd", "header_truncated",
|
||||
"OAMD payload 不足以容纳 header",
|
||||
details={"payload_bits": len(bits),
|
||||
"payload": bytes_descriptor(raw_payload)}) from None
|
||||
reader = _BitReader(bits, 2)
|
||||
try:
|
||||
if version == 3:
|
||||
version += reader.read(3)
|
||||
object_count_bits = reader.read(5)
|
||||
if object_count_bits == 0x1F:
|
||||
object_count_bits += reader.read(7)
|
||||
object_count = object_count_bits + 1
|
||||
program = _parse_program_assignment(reader, raw_payload)
|
||||
alternate_present = bool(reader.read(1))
|
||||
object_element, elements = _parse_elements(
|
||||
reader, alternate_present, object_count, program, raw_payload)
|
||||
except UnsupportedVariantError:
|
||||
raise
|
||||
except ValueError as exc:
|
||||
raise _syntax_error(
|
||||
"header_truncated", "OAMD header 字段越界或不完整",
|
||||
raw_payload, exc, payload_bits=len(bits)) from exc
|
||||
|
||||
if version != 0:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "oamd_version",
|
||||
"OAMD 使用了当前未覆盖的 syntax version",
|
||||
details={
|
||||
"supported_header_prefix_hex": "1f88",
|
||||
"header_probe": header,
|
||||
"version": version,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"repair_hint": "按新 header 的 program assignment 和 element 布局重新定位对象位置字段",
|
||||
})
|
||||
block_offset, ramp_duration = _update_timing(
|
||||
bits, bool(header["alternate_object_present"]), header["element_count"], raw_payload)
|
||||
weights = 1 << np.arange(7, -1, -1)
|
||||
out = {}
|
||||
for obj in range(16):
|
||||
start = 112 + 31 * (obj - 3)
|
||||
wq1 = int((bits[start:start + 8] * weights).sum())
|
||||
wq2 = int((bits[start + 8:start + 16] * weights).sum())
|
||||
wq3 = int((bits[start + 16:start + 24] * weights).sum())
|
||||
if obj and (wq1 >> 6 != 3 or wq2 & 2 != 2 or wq3 & 0x1F != 1):
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "position_layout_signature",
|
||||
"OAMD 长度和 header 已知,但对象位置字段标记或位偏移发生变化",
|
||||
details={
|
||||
"object_slot": obj,
|
||||
"position_start_bit": start,
|
||||
"q1_window_hex": f"{wq1:02x}",
|
||||
"q2_window_hex": f"{wq2:02x}",
|
||||
"q3_window_hex": f"{wq3:02x}",
|
||||
"header_probe": header,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"repair_hint": "解析 OAMD element 可选字段并更新每个对象的位置窗口偏移",
|
||||
})
|
||||
k1 = wq1 - Q1_OFF
|
||||
if obj == 0:
|
||||
out[(0, "q1")] = None
|
||||
out[(0, "q2")] = None
|
||||
else:
|
||||
out[(obj, "q1")] = q_of(k1, N_Q12) if 0 <= k1 <= N_Q12 else None
|
||||
k2 = wq2 >> 2
|
||||
if obj != 0:
|
||||
out[(obj, "q2")] = q_of(k2, N_Q12) if 0 <= k2 <= N_Q12 else None
|
||||
k3 = (wq2 & 1) * 8 + (wq3 >> 5)
|
||||
out[(obj, "q3")] = q_of(k3, N_Q3) if 0 <= k3 <= N_Q3 else None
|
||||
if object_count > MAX_SLOTS:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "object_count",
|
||||
"OAMD 对象数超过当前 16 槽对象模型",
|
||||
details={
|
||||
"object_count": object_count,
|
||||
"max_slots": MAX_SLOTS,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
})
|
||||
blocks = object_element["blocks"]
|
||||
if len(blocks) != 1:
|
||||
raise UnsupportedVariantError(
|
||||
"oamd", "multiple_position_blocks",
|
||||
"OAMD 一帧含多个对象位置更新块,单次 frame_update 无法表达",
|
||||
details={
|
||||
"block_count": len(blocks),
|
||||
"blocks": blocks,
|
||||
"payload": bytes_descriptor(raw_payload),
|
||||
"repair_hint": "按 ObjectInfoBlock 顺序逐块解析坐标,再生成分段 ADM ramp",
|
||||
})
|
||||
|
||||
# 对外契约:始终给出 16 个槽;本帧没有覆盖到的槽保持上一帧位置。
|
||||
out = {(slot, field): None for slot in range(MAX_SLOTS)
|
||||
for field in ("q1", "q2", "q3")}
|
||||
for info in object_element["objects"]:
|
||||
slot = info["object"]
|
||||
position = info.get("position")
|
||||
if position is None:
|
||||
# bed/ISF 对象与未激活对象没有位置字段:保持上一帧位置。
|
||||
out[(slot, "q1")] = None
|
||||
out[(slot, "q2")] = None
|
||||
out[(slot, "q3")] = None
|
||||
continue
|
||||
x, y, z = position
|
||||
out[(slot, "q1")] = q_of(x, N_Q12) if 0 <= x <= N_Q12 else None
|
||||
out[(slot, "q2")] = q_of(y, N_Q12) if 0 <= y <= N_Q12 else None
|
||||
# 5.6.1.1.10/5.6.1.1.11:pos3D_Z 带符号;当前 16 槽模型只表达非负高度。
|
||||
out[(slot, "q3")] = q_of(z, N_Q3) if 0 <= z <= N_Q3 else (
|
||||
0 if z < 0 else None)
|
||||
|
||||
size_mismatch = [
|
||||
{
|
||||
"element_ordinal": element["ordinal"],
|
||||
"element_id": element["element_id"],
|
||||
"declared_size_bytes": element["size_bytes"],
|
||||
"declared_end_bit": element["body_end_bit"],
|
||||
"parsed_end_bit": element["parsed_end_bit"],
|
||||
}
|
||||
for element in elements
|
||||
if element["parsed_end_bit"] > element["body_end_bit"]
|
||||
]
|
||||
return {
|
||||
"values": out,
|
||||
"block_offset_samples": block_offset,
|
||||
"ramp_duration_samples": ramp_duration,
|
||||
"block_offset_samples": blocks[0]["block_offset_samples"],
|
||||
"ramp_duration_samples": blocks[0]["ramp_duration_samples"],
|
||||
"object_count": object_count,
|
||||
"program": {
|
||||
"dynamic_object_only": program["dynamic_object_only"],
|
||||
"lfe_present": program["lfe_present"],
|
||||
"content_description": program["content_description"],
|
||||
"num_bed_objects": program["num_bed_objects"],
|
||||
"num_isf_objects": program["num_isf_objects"],
|
||||
"num_dynamic_objects": program["num_dynamic_objects"],
|
||||
"bed_channels": [list(assignment["channels"])
|
||||
for assignment in program["bed_assignments"]],
|
||||
},
|
||||
"elements": _element_details(elements),
|
||||
"objects": object_element["objects"],
|
||||
"size_mismatch": size_mismatch,
|
||||
}
|
||||
|
||||
|
||||
|
||||
+2
-2
@@ -179,7 +179,7 @@ def _expand_events(events, total_samples, rate, update_quantum_samples,
|
||||
raise ValueError(f"未知 trajectory_mode: {trajectory_mode}")
|
||||
|
||||
def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536,
|
||||
update_quantum_samples=64, object_delay_samples=640,
|
||||
update_quantum_samples=64, object_delay_samples=1473,
|
||||
trajectory_mode="compact"):
|
||||
"""从统一 metadata index 构造 15 条 ADM 轨迹。
|
||||
|
||||
@@ -187,7 +187,7 @@ def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536,
|
||||
OAMD 的内外层 sample offset、block offset 和 ramp 均保留。
|
||||
``trajectory_mode="compact"`` 用一个长 ADM interpolation block 表示每条
|
||||
线性 ramp;``dense64`` 保留逐 64-sample 展开作为兼容回退。
|
||||
``object_delay_samples`` 将位置更新与对象逆 QMF 的输出时刻对齐。
|
||||
``object_delay_samples`` 将位置更新与对象 PCM 的 decoder 输出时刻对齐。
|
||||
slot1..15 与对象 PCM ch1..15 一一对应。
|
||||
"""
|
||||
frames = index.rows if frames is None else frames
|
||||
|
||||
@@ -0,0 +1,692 @@
|
||||
"""Public 64-QMF and 77-band hybrid filterbank for binaural rendering.
|
||||
|
||||
The fixed resource is ``data/rosella_kernels.npz``: the fixed 64-QMF /
|
||||
``3 -> 8+4+4`` 77-hybrid analysis tables and the causal synthesis tables
|
||||
computed from that analysis bank. The filter bank is publicly standardized:
|
||||
the 64-QMF → 77-hybrid structure, the 13-tap low-band prototypes and their
|
||||
half-bin complex modulation follow 3GPP TS 26.405 / ETSI TS 126 405 (Section
|
||||
5.2.2, Table 1, ``Q=8``/``Q=4``); the 64-band QMF analysis is the MPEG-4
|
||||
AAC/SBR 64 complex QMF analysis bank (ISO/IEC 14496-3/AMD1:2003, subclause
|
||||
4.B.18.2), stored here as the polyphase form
|
||||
``A[r,t] = ((-1)**t / 128) * c[63 - r + 64*t]`` of the public 640-tap SBR
|
||||
prototype. The QMF synthesis table is the causal left inverse of that
|
||||
analysis polyphase matrix (``A @ W = P`` with the 577-sample delay
|
||||
permutation; total latency ``961 = 577 + 6*64``), stored as the rank-4
|
||||
factorization ``W[b,l] = sum_r taps[b,l,r] * basis[b,r,:]``; the hybrid
|
||||
synthesis table is the 77->64 recombination (identity for the high bands,
|
||||
signed summation of each 8+4+4 child group for the low bands), stored as a
|
||||
154-entry sparse map. The archive and every array inside it are
|
||||
hash-validated before use, and those hashes participate in every
|
||||
compiled-HRTF cache key. Provenance and rights boundaries are documented in
|
||||
``data/README.md`` and ``THIRD_PARTY_NOTICES.md``.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from functools import lru_cache
|
||||
import hashlib
|
||||
import os
|
||||
from pathlib import Path
|
||||
import zipfile
|
||||
|
||||
import numpy as np
|
||||
|
||||
|
||||
PROJECT_DIR = Path(__file__).resolve().parent.parent
|
||||
DEFAULT_FILTERBANK_DATA = PROJECT_DIR / "data" / "rosella_kernels.npz"
|
||||
FILTERBANK_TABLE_VERSION = "joc-public-64qmf-77hybrid-v1"
|
||||
SAMPLE_RATE = 48000
|
||||
QMF_HOP = 64
|
||||
QMF_BANDS = 64
|
||||
HYBRID_BANDS = 77
|
||||
ANALYSIS_SYNTHESIS_LATENCY_SAMPLES = 961
|
||||
|
||||
_ARCHIVE_SHA256 = "C05BEF4D26E96ECBD4694E2572F05DA400255C777BA5047300B9D3B1F81081CD"
|
||||
_TABLE_SPECS = {
|
||||
"format_version": (np.dtype("<i4"), (1,),
|
||||
"67ABDD721024F0FF4E0B3F4C2FC13BC5BAD42D0B7851D456D88D203D15AAA450",
|
||||
False, False),
|
||||
"qmf_analysis_coefficients": (
|
||||
np.dtype("<f4"), (64, 10),
|
||||
"AEFF6C7117D41664B9C4BF03BBF563F5319EC1B8C551F171ADBB90CF19D9D306",
|
||||
False, False),
|
||||
"hybrid_analysis_low_kernel": (
|
||||
np.dtype("<f4"), (3, 2, 13, 16, 2),
|
||||
"D00D36133B81BA699A7630C4DF1BE203FA1B7E371E595EAAEBBE8957DB322627",
|
||||
False, False),
|
||||
"hybrid_synthesis_indices": (
|
||||
np.dtype("<i2"), (154, 4),
|
||||
"F5BEB3220E4530FCF28E7F4DA7F07E821074265D118C911D61A590E00753A573",
|
||||
True, False),
|
||||
"hybrid_synthesis_values": (
|
||||
np.dtype("<f4"), (154,),
|
||||
"99409FDD9D20D1D7C2BE16BBC1E2159C8042487227C72160850745164C9CEE7F",
|
||||
False, False),
|
||||
"qmf_synthesis_basis": (
|
||||
np.dtype("<f8"), (64, 4, 128),
|
||||
"A0C4A55385F6D6C7C92D7615C83AD5FBDA51046D9EF785CAC0B9AC9A760DC527",
|
||||
False, False),
|
||||
"qmf_synthesis_taps": (
|
||||
np.dtype("<f8"), (64, 10, 4),
|
||||
"CD7756D060D51FBF02F44C1CE53CB6225B221099505C94C3F58D3BEE6F428150",
|
||||
False, False),
|
||||
}
|
||||
|
||||
# Hybrid-band center frequencies measured from the public analysis bank at
|
||||
# 48 kHz (positive-frequency response peaks). They are part of the validated
|
||||
# reference behavior: the runtime uses them only for the fractional-delay band
|
||||
# phase and the project LFE low-pass, never as filterbank coefficients.
|
||||
_BAND_CENTER_FREQUENCIES_HZ = np.asarray([
|
||||
53.19564095937407,
|
||||
26.3876219849709,
|
||||
140.9074183269806,
|
||||
98.55238901464415,
|
||||
234.09258166297573,
|
||||
344.8745321543293,
|
||||
321.80435900819805,
|
||||
401.38762185365727,
|
||||
476.19239745597804,
|
||||
473.552388917001,
|
||||
648.8076026837931,
|
||||
719.8745321201852,
|
||||
780.1254679168173,
|
||||
851.1923973714038,
|
||||
1023.8076024849751,
|
||||
1155.125467695814,
|
||||
1293.0325067063661,
|
||||
1668.032511377253,
|
||||
2043.0325074138086,
|
||||
2456.967490229666,
|
||||
2831.96748495363,
|
||||
3206.967488172826,
|
||||
3581.9675052705525,
|
||||
3918.0325089666067,
|
||||
4331.967500201844,
|
||||
4706.967500350216,
|
||||
5043.032510771545,
|
||||
5456.9674914391635,
|
||||
5831.967490225267,
|
||||
6168.0324885741875,
|
||||
6543.032508972284,
|
||||
6956.96748844665,
|
||||
7293.032503259869,
|
||||
7668.032503551393,
|
||||
8043.032504806491,
|
||||
8418.032498852166,
|
||||
8793.032502508235,
|
||||
9206.967488589786,
|
||||
9543.032511549152,
|
||||
9956.967486913867,
|
||||
10293.032513008677,
|
||||
10668.032507835102,
|
||||
11043.032513641429,
|
||||
11456.967482937946,
|
||||
11793.032513984212,
|
||||
12206.967486015788,
|
||||
12543.032517062376,
|
||||
12956.96748635822,
|
||||
13331.967492165066,
|
||||
13706.967486991198,
|
||||
14043.032513086031,
|
||||
14456.967488451,
|
||||
14793.032511410214,
|
||||
15206.967497492202,
|
||||
15581.967501147887,
|
||||
15956.967495193188,
|
||||
16331.96749644876,
|
||||
16706.967496740173,
|
||||
17043.03251155318,
|
||||
17456.96749102773,
|
||||
17831.967511425748,
|
||||
18168.032509774734,
|
||||
18543.032508560515,
|
||||
18956.967489228293,
|
||||
19293.032499649784,
|
||||
19668.032499798002,
|
||||
20081.967491033392,
|
||||
20418.0324947296,
|
||||
20793.032511827063,
|
||||
21168.032515046092,
|
||||
21543.032509770488,
|
||||
21956.967492586176,
|
||||
22331.96748862257,
|
||||
22706.967493293465,
|
||||
23043.03251375972,
|
||||
23418.03251061199,
|
||||
23831.96749768645,
|
||||
], dtype=np.float64)
|
||||
|
||||
|
||||
def _sha256_bytes(values: bytes) -> str:
|
||||
return hashlib.sha256(values).hexdigest().upper()
|
||||
|
||||
|
||||
def _validate_npy_member_header(
|
||||
archive: zipfile.ZipFile, member: zipfile.ZipInfo,
|
||||
*, name: str, dtype: np.dtype, shape: tuple[int, ...],
|
||||
allow_fortran: bool) -> None:
|
||||
try:
|
||||
with archive.open(member, "r") as payload:
|
||||
version = np.lib.format.read_magic(payload)
|
||||
if version == (1, 0):
|
||||
actual_shape, actual_fortran_order, actual_dtype = (
|
||||
np.lib.format.read_array_header_1_0(
|
||||
payload, max_header_size=4096))
|
||||
elif version == (2, 0):
|
||||
actual_shape, actual_fortran_order, actual_dtype = (
|
||||
np.lib.format.read_array_header_2_0(
|
||||
payload, max_header_size=4096))
|
||||
else:
|
||||
raise ValueError(f"unsupported .npy version {version!r}")
|
||||
header_size = payload.tell()
|
||||
except (EOFError, OSError, ValueError) as exc:
|
||||
raise ValueError(
|
||||
f"invalid public filterbank table .npy header: {member.filename}: "
|
||||
f"{exc}") from exc
|
||||
|
||||
actual_shape = tuple(actual_shape)
|
||||
actual_dtype = np.dtype(actual_dtype)
|
||||
if (actual_shape != shape or actual_dtype != dtype
|
||||
or (bool(actual_fortran_order) and not allow_fortran)):
|
||||
expected_order = "C-order" if not allow_fortran else "C- or Fortran-order"
|
||||
actual_order = "Fortran-order" if actual_fortran_order else "C-order"
|
||||
raise ValueError(
|
||||
f"invalid public filterbank table .npy header for {name}: expected "
|
||||
f"{dtype}{shape} {expected_order}, got "
|
||||
f"{actual_dtype}{actual_shape} {actual_order}")
|
||||
expected_size = header_size + dtype.itemsize * int(np.prod(shape))
|
||||
if member.file_size != expected_size:
|
||||
raise ValueError(
|
||||
f"invalid public filterbank table .npy payload size for {name}: "
|
||||
f"expected {expected_size} bytes including the header, "
|
||||
f"got {member.file_size}")
|
||||
|
||||
|
||||
def _validate_table_members(stream) -> None:
|
||||
expected = {name + ".npy" for name in _TABLE_SPECS}
|
||||
limits = {
|
||||
name + ".npy": dtype.itemsize * int(np.prod(shape)) + 4096
|
||||
for name, (dtype, shape, _, _, _) in _TABLE_SPECS.items()
|
||||
}
|
||||
try:
|
||||
stream.seek(0)
|
||||
with zipfile.ZipFile(stream, "r") as archive:
|
||||
members = archive.infolist()
|
||||
if (len(members) != len(expected)
|
||||
or {member.filename for member in members} != expected):
|
||||
raise ValueError(
|
||||
"public filterbank table archive has an invalid member set")
|
||||
for member in members:
|
||||
if member.flag_bits & 0x1:
|
||||
raise ValueError("encrypted public filterbank tables are unsupported")
|
||||
if member.compress_type not in (
|
||||
zipfile.ZIP_STORED, zipfile.ZIP_DEFLATED):
|
||||
raise ValueError("unsupported public filterbank table compression")
|
||||
if member.file_size > limits[member.filename]:
|
||||
raise ValueError(
|
||||
f"public filterbank table member is unexpectedly large: "
|
||||
f"{member.filename}")
|
||||
if sum(member.file_size for member in members) > sum(limits.values()):
|
||||
raise ValueError("public filterbank tables expand beyond their size limit")
|
||||
for member in members:
|
||||
name = member.filename[:-4]
|
||||
dtype, shape, _, allow_fortran, _ = _TABLE_SPECS[name]
|
||||
_validate_npy_member_header(
|
||||
archive, member, name=name, dtype=dtype, shape=shape,
|
||||
allow_fortran=allow_fortran)
|
||||
except zipfile.BadZipFile as exc:
|
||||
raise ValueError(f"invalid public filterbank table archive: {exc}") from exc
|
||||
|
||||
|
||||
@lru_cache(maxsize=2)
|
||||
def _load_tables(path_string: str) -> dict[str, np.ndarray]:
|
||||
path = Path(path_string)
|
||||
if not path.is_file():
|
||||
raise FileNotFoundError(f"public filterbank table resource not found: {path}")
|
||||
with path.open("rb") as stream:
|
||||
archive_size = os.fstat(stream.fileno()).st_size
|
||||
if archive_size <= 0 or archive_size > 8 << 20:
|
||||
raise ValueError(
|
||||
f"public filterbank table resource is unexpectedly large: {path}")
|
||||
if path.resolve() == DEFAULT_FILTERBANK_DATA.resolve():
|
||||
digest = hashlib.sha256()
|
||||
for block in iter(lambda: stream.read(4 << 20), b""):
|
||||
digest.update(block)
|
||||
actual_archive_hash = digest.hexdigest().upper()
|
||||
if actual_archive_hash != _ARCHIVE_SHA256:
|
||||
raise ValueError(
|
||||
"public filterbank table archive hash mismatch: "
|
||||
f"expected {_ARCHIVE_SHA256}, got {actual_archive_hash}")
|
||||
_validate_table_members(stream)
|
||||
stream.seek(0)
|
||||
with np.load(stream, allow_pickle=False) as archive:
|
||||
if set(archive.files) != set(_TABLE_SPECS):
|
||||
raise ValueError("public filterbank table archive has an invalid key set")
|
||||
result: dict[str, np.ndarray] = {}
|
||||
for name, (dtype, shape, expected_hash, _, _) in _TABLE_SPECS.items():
|
||||
value = np.asarray(archive[name])
|
||||
if value.dtype != dtype or value.shape != shape:
|
||||
raise ValueError(
|
||||
f"invalid public filterbank table {name}: "
|
||||
f"expected {dtype}{shape}, got {value.dtype}{value.shape}")
|
||||
actual_hash = _sha256_bytes(value.tobytes(order="C"))
|
||||
if actual_hash != expected_hash:
|
||||
raise ValueError(f"public filterbank table hash mismatch: {name}")
|
||||
result[name] = np.ascontiguousarray(value)
|
||||
result[name].setflags(write=False)
|
||||
if int(result["format_version"][0]) != 1:
|
||||
raise ValueError("unsupported public filterbank table format version")
|
||||
return result
|
||||
|
||||
|
||||
def load_filterbank_tables(
|
||||
path: str | Path = DEFAULT_FILTERBANK_DATA) -> dict[str, np.ndarray]:
|
||||
"""Load the validated project resource used by the public filterbank."""
|
||||
cached = _load_tables(str(Path(path).expanduser().resolve()))
|
||||
result = {name: value.copy() for name, value in cached.items()}
|
||||
for value in result.values():
|
||||
value.setflags(write=False)
|
||||
return result
|
||||
|
||||
|
||||
def filterbank_fingerprint() -> dict:
|
||||
"""Return stable identifiers used in compiled-HRTF cache keys."""
|
||||
centers = np.ascontiguousarray(_BAND_CENTER_FREQUENCIES_HZ, dtype="<f8")
|
||||
return {
|
||||
"table_version": FILTERBANK_TABLE_VERSION,
|
||||
"archive_sha256": _ARCHIVE_SHA256,
|
||||
"band_centers_sha256": _sha256_bytes(centers.tobytes(order="C")),
|
||||
"array_sha256": {
|
||||
name: spec[2] for name, spec in _TABLE_SPECS.items()
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
class QmfAnalysis:
|
||||
"""Batchable 64-band analysis with float64 state and complex128 FFTs."""
|
||||
|
||||
def __init__(self, channels: int,
|
||||
table_data: str | Path = DEFAULT_FILTERBANK_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_filterbank_tables(table_data)
|
||||
self.coefficients = np.asarray(
|
||||
tables["qmf_analysis_coefficients"], dtype=np.float64)
|
||||
self.channels = int(channels)
|
||||
self.history = np.zeros((9, self.channels, 64), dtype=np.float64)
|
||||
phase = np.arange(64, dtype=np.float64)
|
||||
self.premod = np.exp(-1j * np.pi * phase / 128.0).astype(np.complex128)
|
||||
self.post = np.exp(
|
||||
-1j * 3.0 * (np.arange(64, dtype=np.float64) + 0.5) * np.pi / 128.0
|
||||
).astype(np.complex128)
|
||||
self.even_post = (
|
||||
1j * ((-1.0) ** np.arange(64, dtype=np.float64))
|
||||
).astype(np.complex128)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.history.fill(0.0)
|
||||
|
||||
def process_chunk(self, hops) -> np.ndarray:
|
||||
values = np.asarray(hops, dtype=np.float64)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 64):
|
||||
raise ValueError(f"expected [slots,{self.channels},64], got {values.shape}")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("QMF input contains non-finite values")
|
||||
count = values.shape[0]
|
||||
joined = np.concatenate((self.history, values), axis=0)
|
||||
even = np.zeros_like(values)
|
||||
odd = np.zeros_like(values)
|
||||
for lag in range(10):
|
||||
source = joined[9 - lag:9 - lag + count]
|
||||
target = even if lag % 2 == 0 else odd
|
||||
target += source * self.coefficients[:, lag][None, None, :]
|
||||
self.history[:] = joined[-9:]
|
||||
|
||||
def transform(block):
|
||||
prepared = block.astype(np.complex128, copy=False) * self.premod
|
||||
transformed = np.fft.fft(prepared, n=128, axis=-1)[..., :64]
|
||||
return transformed * self.post
|
||||
|
||||
return np.asarray(transform(odd) + transform(even) * self.even_post,
|
||||
dtype=np.complex128)
|
||||
|
||||
|
||||
class HybridAnalysis:
|
||||
"""Sparse 64-QMF to 77-hybrid analysis in float64/complex128."""
|
||||
|
||||
def __init__(self, channels: int,
|
||||
table_data: str | Path = DEFAULT_FILTERBANK_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_filterbank_tables(table_data)
|
||||
self.low_kernel = np.asarray(
|
||||
tables["hybrid_analysis_low_kernel"], dtype=np.float64)
|
||||
self.channels = int(channels)
|
||||
self.history = np.zeros((12, self.channels, 3, 2), dtype=np.float64)
|
||||
self.high_history = np.zeros(
|
||||
(6, self.channels, 61), dtype=np.complex128)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.history.fill(0.0)
|
||||
self.high_history.fill(0.0)
|
||||
|
||||
def process_chunk(self, qmf) -> np.ndarray:
|
||||
values = np.asarray(qmf, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 64):
|
||||
raise ValueError(f"expected [slots,{self.channels},64], got {values.shape}")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hybrid-analysis input contains non-finite values")
|
||||
count = values.shape[0]
|
||||
low = np.stack((values[:, :, :3].real, values[:, :, :3].imag), axis=-1)
|
||||
joined = np.concatenate((self.history, low), axis=0)
|
||||
output = np.zeros((count, self.channels, 77, 2), dtype=np.float64)
|
||||
for lag in range(13):
|
||||
source = joined[12 - lag:12 - lag + count]
|
||||
output[:, :, :16] += np.einsum(
|
||||
"tcpi,pibo->tcbo", source, self.low_kernel[:, :, lag],
|
||||
dtype=np.float64, optimize=False)
|
||||
self.history[:] = joined[-12:]
|
||||
|
||||
high_joined = np.concatenate((self.high_history, values[:, :, 3:]), axis=0)
|
||||
high = high_joined[:count]
|
||||
output[:, :, 16:, 0] = high.real
|
||||
output[:, :, 16:, 1] = high.imag
|
||||
self.high_history[:] = high_joined[-6:]
|
||||
return np.asarray(output[..., 0] + 1j * output[..., 1], dtype=np.complex128)
|
||||
|
||||
|
||||
class HybridSynthesis:
|
||||
"""Instantaneous sparse 77-hybrid to 64-QMF synthesis map."""
|
||||
|
||||
def __init__(self, channels: int,
|
||||
table_data: str | Path = DEFAULT_FILTERBANK_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_filterbank_tables(table_data)
|
||||
indices = np.asarray(tables["hybrid_synthesis_indices"], dtype=np.int64)
|
||||
values = np.asarray(tables["hybrid_synthesis_values"], dtype=np.float64)
|
||||
if indices.ndim != 2 or indices.shape[1] != 4 or len(indices) != len(values):
|
||||
raise ValueError("invalid hybrid synthesis sparse table")
|
||||
self.mapping = [
|
||||
(int(index[0]), int(index[1]), int(index[2]), int(index[3]), float(value))
|
||||
for index, value in zip(indices, values)
|
||||
]
|
||||
self.channels = int(channels)
|
||||
|
||||
def reset(self) -> None:
|
||||
return None
|
||||
|
||||
def process_chunk(self, hybrid) -> np.ndarray:
|
||||
values = np.asarray(hybrid, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 77):
|
||||
raise ValueError(f"expected [slots,{self.channels},77], got {values.shape}")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hybrid-synthesis input contains non-finite values")
|
||||
source = np.stack((values.real, values.imag), axis=-1)
|
||||
output = np.zeros((values.shape[0], self.channels, 64, 2), dtype=np.float64)
|
||||
for input_band, input_component, output_band, output_component, gain in self.mapping:
|
||||
output[:, :, output_band, output_component] += (
|
||||
source[:, :, input_band, input_component] * gain)
|
||||
return np.asarray(output[..., 0] + 1j * output[..., 1], dtype=np.complex128)
|
||||
|
||||
|
||||
class QmfSynthesis:
|
||||
"""Rank-4 64-band synthesis with float64 state and accumulation."""
|
||||
|
||||
def __init__(self, channels: int,
|
||||
table_data: str | Path = DEFAULT_FILTERBANK_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_filterbank_tables(table_data)
|
||||
self.basis = np.asarray(tables["qmf_synthesis_basis"], dtype=np.float64)
|
||||
self.taps = np.asarray(tables["qmf_synthesis_taps"], dtype=np.float64)
|
||||
if self.basis.shape != (64, 4, 128) or self.taps.shape != (64, 10, 4):
|
||||
raise ValueError("invalid QMF synthesis factorization")
|
||||
self.channels = int(channels)
|
||||
self.rank = 4
|
||||
self.history = np.zeros(
|
||||
(9, self.channels, 64, self.rank), dtype=np.float64)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.history.fill(0.0)
|
||||
|
||||
def process_chunk(self, qmf) -> np.ndarray:
|
||||
values = np.asarray(qmf, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 64):
|
||||
raise ValueError(f"expected [slots,{self.channels},64], got {values.shape}")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("QMF-synthesis input contains non-finite values")
|
||||
count = values.shape[0]
|
||||
flat = np.stack((values.real, values.imag), axis=-1).reshape(
|
||||
count * self.channels, 128)
|
||||
modulation = self.basis.reshape(64 * self.rank, 128)
|
||||
features = (flat @ modulation.T).reshape(
|
||||
count, self.channels, 64, self.rank)
|
||||
joined = np.concatenate((self.history, features), axis=0)
|
||||
output = np.zeros((count, self.channels, 64), dtype=np.float64)
|
||||
for lag in range(10):
|
||||
output += np.sum(
|
||||
joined[9 - lag:9 - lag + count]
|
||||
* self.taps[:, lag, :][None, None, :, :],
|
||||
axis=-1, dtype=np.float64)
|
||||
self.history[:] = joined[-9:]
|
||||
return output
|
||||
|
||||
|
||||
class PublicAnalysis77:
|
||||
"""Full-rate PCM to the public 77-band hybrid representation."""
|
||||
|
||||
def __init__(self, channels: int):
|
||||
self.channels = int(channels)
|
||||
if self.channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
self.qmf = QmfAnalysis(self.channels)
|
||||
self.hybrid = HybridAnalysis(self.channels)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.qmf.reset()
|
||||
self.hybrid.reset()
|
||||
|
||||
def process(self, samples) -> np.ndarray:
|
||||
values = np.asarray(samples, dtype=np.float64)
|
||||
if values.ndim == 1 and self.channels == 1:
|
||||
values = values[:, None]
|
||||
if values.ndim != 2 or values.shape[1] != self.channels:
|
||||
raise ValueError(f"samples must have shape [N,{self.channels}]")
|
||||
if len(values) % QMF_HOP:
|
||||
raise ValueError("sample count must be divisible by the 64-sample QMF hop")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("samples contain non-finite values")
|
||||
hops = values.reshape(-1, QMF_HOP, self.channels).transpose(0, 2, 1)
|
||||
return np.asarray(
|
||||
self.hybrid.process_chunk(self.qmf.process_chunk(hops)),
|
||||
dtype=np.complex128)
|
||||
|
||||
|
||||
class PublicSynthesis77:
|
||||
"""Public 77-band hybrid representation to full-rate PCM."""
|
||||
|
||||
def __init__(self, channels: int):
|
||||
self.channels = int(channels)
|
||||
if self.channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
self.hybrid = HybridSynthesis(self.channels)
|
||||
self.qmf = QmfSynthesis(self.channels)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.hybrid.reset()
|
||||
self.qmf.reset()
|
||||
|
||||
def process(self, hybrid) -> np.ndarray:
|
||||
values = np.asarray(hybrid, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, HYBRID_BANDS):
|
||||
raise ValueError(
|
||||
f"hybrid must have shape [slots,{self.channels},{HYBRID_BANDS}]")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hybrid input contains non-finite values")
|
||||
qmf = self.hybrid.process_chunk(values)
|
||||
time = self.qmf.process_chunk(qmf)
|
||||
return np.asarray(time.transpose(0, 2, 1).reshape(-1, self.channels),
|
||||
dtype=np.float64)
|
||||
|
||||
|
||||
def identity_impulse_response(sample_count: int = 4096) -> np.ndarray:
|
||||
sample_count = int(sample_count)
|
||||
if sample_count <= 0:
|
||||
raise ValueError("sample_count must be positive")
|
||||
total = ((sample_count + QMF_HOP - 1) // QMF_HOP) * QMF_HOP
|
||||
impulse = np.zeros((total, 1), dtype=np.float64)
|
||||
impulse[0, 0] = 1.0
|
||||
analysis = PublicAnalysis77(1)
|
||||
synthesis = PublicSynthesis77(1)
|
||||
return synthesis.process(analysis.process(impulse))[:, 0]
|
||||
|
||||
|
||||
@lru_cache(maxsize=2)
|
||||
def _hybrid_band_center_frequencies_hz_cached(rate: float) -> np.ndarray:
|
||||
centers = np.asarray(
|
||||
_BAND_CENTER_FREQUENCIES_HZ * (rate / SAMPLE_RATE), dtype=np.float64)
|
||||
centers.setflags(write=False)
|
||||
return centers
|
||||
|
||||
|
||||
def hybrid_band_center_frequencies_hz(
|
||||
sample_rate_hz: float = SAMPLE_RATE) -> np.ndarray:
|
||||
"""Return the 77 hybrid-band reference center frequencies."""
|
||||
rate = float(sample_rate_hz)
|
||||
if not np.isfinite(rate) or rate <= 0.0:
|
||||
raise ValueError("sample rate must be positive and finite")
|
||||
centers = _hybrid_band_center_frequencies_hz_cached(rate).copy()
|
||||
centers.setflags(write=False)
|
||||
return centers
|
||||
|
||||
|
||||
def table_info() -> dict:
|
||||
return {
|
||||
"resource": DEFAULT_FILTERBANK_DATA.name,
|
||||
"sample_rate_hz": SAMPLE_RATE,
|
||||
"qmf_bands": QMF_BANDS,
|
||||
"hybrid_bands": HYBRID_BANDS,
|
||||
"hop_samples": QMF_HOP,
|
||||
"analysis_synthesis_latency_samples": ANALYSIS_SYNTHESIS_LATENCY_SAMPLES,
|
||||
"precision": "float64/complex128",
|
||||
"fingerprint": filterbank_fingerprint(),
|
||||
"provenance": {
|
||||
"qmf": (
|
||||
"MPEG-4 AAC/SBR 64 complex QMF analysis (ISO/IEC "
|
||||
"14496-3/AMD1:2003 4.B.18.2), polyphase form of the public "
|
||||
"640-tap SBR prototype"),
|
||||
"hybrid": (
|
||||
"3GPP TS 26.405 / ETSI TS 126 405 5.2.2 Table 1 (Q=8/Q=4) "
|
||||
"with standard half-bin complex modulation"),
|
||||
"synthesis": (
|
||||
"causal left inverse of the public analysis bank "
|
||||
"(A·W = P, 577-sample QMF delay); 77→64 sparse recombination"),
|
||||
"resource": "data/rosella_kernels.npz",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
@lru_cache(maxsize=8)
|
||||
def _hybrid_gain_synthesis_dictionary_cached(count: int) -> np.ndarray:
|
||||
total = int(np.ceil(
|
||||
(ANALYSIS_SYNTHESIS_LATENCY_SAMPLES + count + 512) / QMF_HOP) * QMF_HOP)
|
||||
impulse = np.zeros((total, 1), dtype=np.float64)
|
||||
impulse[0, 0] = 1.0
|
||||
base = PublicAnalysis77(1).process(impulse)[:, 0, :]
|
||||
parameter_count = 2 * HYBRID_BANDS
|
||||
hybrid = np.zeros(
|
||||
(len(base), parameter_count, HYBRID_BANDS), dtype=np.complex128)
|
||||
for band in range(HYBRID_BANDS):
|
||||
hybrid[:, 2 * band, band] = base[:, band]
|
||||
hybrid[:, 2 * band + 1, band] = 1j * base[:, band]
|
||||
rendered = PublicSynthesis77(parameter_count).process(hybrid)
|
||||
start = ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
dictionary = np.asarray(
|
||||
rendered[start:start + count], dtype=np.float64).copy()
|
||||
dictionary.setflags(write=False)
|
||||
return dictionary
|
||||
|
||||
|
||||
def hybrid_gain_synthesis_dictionary(sample_count: int) -> np.ndarray:
|
||||
"""Return the 154-real-parameter analysis/gain/synthesis dictionary.
|
||||
|
||||
Each hybrid band contributes one real-gain and one imaginary-gain column.
|
||||
The common 961-sample filterbank latency is removed from every column.
|
||||
"""
|
||||
count = int(sample_count)
|
||||
if count <= 0:
|
||||
raise ValueError("sample_count must be positive")
|
||||
dictionary = _hybrid_gain_synthesis_dictionary_cached(count).copy()
|
||||
dictionary.setflags(write=False)
|
||||
return dictionary
|
||||
|
||||
|
||||
def project_hrir_to_hybrid_gains(
|
||||
hrir, *, embedded_delay_samples=None,
|
||||
sample_rate_hz: float = SAMPLE_RATE, ridge: float = 1.0e-3,
|
||||
) -> tuple[np.ndarray, dict]:
|
||||
"""Project FIRs and remove only a known embedded arrival delay.
|
||||
|
||||
Non-zero SOFA ``Data.Delay`` is external and must be passed as zero here.
|
||||
A positive onset separated from ``Data.IR`` is de-rotated once, then restored
|
||||
once by the runtime field. A zero-origin FIR keeps its authored complex
|
||||
phase and therefore also passes zero.
|
||||
"""
|
||||
values = np.asarray(hrir, dtype=np.float64)
|
||||
if values.ndim != 3 or values.shape[1] != 2 or values.shape[2] <= 0:
|
||||
raise ValueError("hrir must have shape [M,2,N]")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hrir contains non-finite values")
|
||||
regularization = float(ridge)
|
||||
if not np.isfinite(regularization) or regularization < 0.0:
|
||||
raise ValueError("projection ridge must be finite and non-negative")
|
||||
if embedded_delay_samples is None:
|
||||
delay = np.zeros(values.shape[:2], dtype=np.float64)
|
||||
else:
|
||||
delay = np.asarray(embedded_delay_samples, dtype=np.float64)
|
||||
if delay.shape != values.shape[:2] or not np.isfinite(delay).all():
|
||||
raise ValueError("embedded_delay_samples must have finite shape [M,2]")
|
||||
rate = float(sample_rate_hz)
|
||||
if not np.isfinite(rate) or rate <= 0.0:
|
||||
raise ValueError("sample_rate_hz must be positive and finite")
|
||||
|
||||
dictionary = _hybrid_gain_synthesis_dictionary_cached(values.shape[2])
|
||||
gram = dictionary.T @ dictionary
|
||||
scale = float(np.trace(gram)) / gram.shape[0]
|
||||
system = gram + regularization * scale * np.eye(gram.shape[0], dtype=np.float64)
|
||||
target = values.reshape(-1, values.shape[2]).T
|
||||
parameters = np.linalg.solve(system, dictionary.T @ target).T
|
||||
parts = parameters.reshape(values.shape[0], 2, 2 * HYBRID_BANDS)
|
||||
transfer = np.asarray(parts[..., 0::2] + 1j * parts[..., 1::2],
|
||||
dtype=np.complex128)
|
||||
centers = hybrid_band_center_frequencies_hz(rate)
|
||||
removal_phase = np.exp(
|
||||
2j * np.pi * delay[..., None] * centers[None, None, :] / rate)
|
||||
aligned = np.asarray(transfer * removal_phase, dtype=np.complex128)
|
||||
|
||||
reconstructed = dictionary @ parameters.T
|
||||
error = target - reconstructed
|
||||
reference_energy = np.sum(target * target, axis=0, dtype=np.float64)
|
||||
error_energy = np.sum(error * error, axis=0, dtype=np.float64)
|
||||
snr = 10.0 * np.log10(
|
||||
np.maximum(reference_energy, 1.0e-300)
|
||||
/ np.maximum(error_energy, 1.0e-300))
|
||||
report = {
|
||||
"method": "regularized public analysis/gain/synthesis dictionary",
|
||||
"dictionary_shape": list(dictionary.shape),
|
||||
"real_parameters": 2 * HYBRID_BANDS,
|
||||
"ridge": regularization,
|
||||
"embedded_delay_samples_min": float(np.min(delay)),
|
||||
"embedded_delay_samples_max": float(np.max(delay)),
|
||||
"fir_reconstruction_snr_db_median": float(np.median(snr)),
|
||||
"fir_reconstruction_snr_db_p05": float(np.percentile(snr, 5.0)),
|
||||
"fir_reconstruction_snr_db_min": float(np.min(snr)),
|
||||
"maximum_absolute_hybrid_gain": float(np.max(np.abs(aligned))),
|
||||
"precision": "float64/complex128",
|
||||
}
|
||||
return aligned, report
|
||||
|
||||
|
||||
def project_aligned_hrir_to_hybrid_gains(
|
||||
aligned_hrir, *, ridge: float = 1.0e-3) -> tuple[np.ndarray, dict]:
|
||||
return project_hrir_to_hybrid_gains(aligned_hrir, ridge=ridge)
|
||||
@@ -0,0 +1,253 @@
|
||||
"""Project-owned image-source early reflections and shared unitary FDN."""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
import math
|
||||
import numpy as np
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ShoeboxRoomConfig:
|
||||
dimensions_m: tuple[float, float, float] = (18.0, 18.0, 14.0)
|
||||
listener_position_m: tuple[float, float, float] = (9.0, 9.0, 7.0)
|
||||
wall_reflection_gain: tuple[float, float, float, float, float, float] = (
|
||||
0.62, 0.60, 0.58, 0.61, 0.52, 0.56)
|
||||
speed_of_sound_m_s: float = 343.3
|
||||
|
||||
def validate(self) -> None:
|
||||
dimensions = np.asarray(self.dimensions_m, dtype=np.float64)
|
||||
listener = np.asarray(self.listener_position_m, dtype=np.float64)
|
||||
gains = np.asarray(self.wall_reflection_gain, dtype=np.float64)
|
||||
if (dimensions.shape != (3,) or not np.isfinite(dimensions).all()
|
||||
or np.any(dimensions <= 0.0)):
|
||||
raise ValueError("room dimensions must be three positive finite values")
|
||||
if (listener.shape != (3,) or not np.isfinite(listener).all()
|
||||
or np.any(listener <= 0.0) or np.any(listener >= dimensions)):
|
||||
raise ValueError("listener must be strictly inside the shoebox")
|
||||
if (gains.shape != (6,) or not np.isfinite(gains).all()
|
||||
or np.any(np.abs(gains) >= 1.0)):
|
||||
raise ValueError("six finite wall gains must have magnitude below one")
|
||||
if not math.isfinite(self.speed_of_sound_m_s) or self.speed_of_sound_m_s <= 0.0:
|
||||
raise ValueError("speed of sound must be positive")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class EarlyReflection:
|
||||
wall: str
|
||||
direction_adm: np.ndarray
|
||||
path_distance_m: float
|
||||
extra_delay_samples: float
|
||||
reflection_gain: float
|
||||
|
||||
|
||||
_WALL_NAMES = ("left", "right", "back", "front", "floor", "ceiling")
|
||||
|
||||
|
||||
def first_order_image_sources(direction_adm, source_distance_m: float,
|
||||
sample_rate_hz: float,
|
||||
config: ShoeboxRoomConfig = ShoeboxRoomConfig()
|
||||
) -> tuple[EarlyReflection, ...]:
|
||||
"""Return six first-order image-source paths for one object."""
|
||||
config.validate()
|
||||
direction = np.asarray(direction_adm, dtype=np.float64)
|
||||
if direction.shape != (3,) or not np.isfinite(direction).all():
|
||||
raise ValueError("reflection direction must contain three finite ADM values")
|
||||
norm = float(np.linalg.norm(direction))
|
||||
if norm <= 1.0e-15:
|
||||
direction = np.asarray([0.0, 1.0, 0.0], dtype=np.float64)
|
||||
else:
|
||||
direction = direction / norm
|
||||
distance = float(source_distance_m)
|
||||
rate = float(sample_rate_hz)
|
||||
if not math.isfinite(distance) or distance <= 0.0 or not math.isfinite(rate) or rate <= 0.0:
|
||||
raise ValueError("source distance and sample rate must be positive")
|
||||
dimensions = np.asarray(config.dimensions_m, dtype=np.float64)
|
||||
listener = np.asarray(config.listener_position_m, dtype=np.float64)
|
||||
source = listener + direction * distance
|
||||
if np.any(source <= 0.0) or np.any(source >= dimensions):
|
||||
raise ValueError(
|
||||
"source lies outside the configured public shoebox; enlarge the room")
|
||||
images = []
|
||||
for axis in range(3):
|
||||
low = source.copy()
|
||||
low[axis] = -source[axis]
|
||||
high = source.copy()
|
||||
high[axis] = 2.0 * dimensions[axis] - source[axis]
|
||||
images.extend((low, high))
|
||||
result = []
|
||||
for wall, image, gain in zip(_WALL_NAMES, images, config.wall_reflection_gain):
|
||||
vector = image - listener
|
||||
path_distance = float(np.linalg.norm(vector))
|
||||
path_direction = vector / path_distance
|
||||
extra = max(0.0, (path_distance - distance)
|
||||
* rate / config.speed_of_sound_m_s)
|
||||
air = math.exp(-0.002 * max(path_distance - distance, 0.0))
|
||||
result.append(EarlyReflection(
|
||||
wall=wall,
|
||||
direction_adm=np.asarray(path_direction, dtype=np.float64),
|
||||
path_distance_m=path_distance,
|
||||
extra_delay_samples=extra,
|
||||
reflection_gain=float(gain) * air,
|
||||
))
|
||||
return tuple(result)
|
||||
|
||||
|
||||
def normalized_hadamard4() -> np.ndarray:
|
||||
return 0.5 * np.asarray([
|
||||
[1.0, 1.0, 1.0, 1.0],
|
||||
[1.0, -1.0, 1.0, -1.0],
|
||||
[1.0, 1.0, -1.0, -1.0],
|
||||
[1.0, -1.0, -1.0, 1.0],
|
||||
], dtype=np.float64)
|
||||
|
||||
|
||||
def _is_prime(value: int) -> bool:
|
||||
if value < 2:
|
||||
return False
|
||||
if value % 2 == 0:
|
||||
return value == 2
|
||||
limit = int(math.sqrt(value))
|
||||
return all(value % divisor for divisor in range(3, limit + 1, 2))
|
||||
|
||||
|
||||
def _next_prime(value: int) -> int:
|
||||
candidate = max(2, int(value))
|
||||
while not _is_prime(candidate):
|
||||
candidate += 1
|
||||
return candidate
|
||||
|
||||
|
||||
class SchroederAllpass:
|
||||
def __init__(self, delay_samples: int, gain: float):
|
||||
self.delay_samples = int(delay_samples)
|
||||
self.gain = float(gain)
|
||||
if self.delay_samples <= 0 or not 0.0 <= abs(self.gain) < 1.0:
|
||||
raise ValueError("all-pass delay must be positive and |gain| < 1")
|
||||
self.buffer = np.zeros(self.delay_samples, dtype=np.float64)
|
||||
self.position = 0
|
||||
|
||||
def reset(self) -> None:
|
||||
self.buffer.fill(0.0)
|
||||
self.position = 0
|
||||
|
||||
def process(self, values) -> np.ndarray:
|
||||
source = np.asarray(values, dtype=np.float64)
|
||||
output = np.empty_like(source)
|
||||
for index, value in enumerate(source):
|
||||
delayed = self.buffer[self.position]
|
||||
result = delayed - self.gain * value
|
||||
self.buffer[self.position] = value + self.gain * result
|
||||
self.position = (self.position + 1) % self.delay_samples
|
||||
output[index] = result
|
||||
return output
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class LateFdnConfig:
|
||||
sample_rate_hz: float = 48000.0
|
||||
rt60_seconds: float = 0.85
|
||||
damping: float = 0.32
|
||||
output_gain: float = 0.22
|
||||
delay_seconds: tuple[float, float, float, float] = (
|
||||
0.0297, 0.0371, 0.0411, 0.0437)
|
||||
allpass_seconds: tuple[float, float] = (0.0023, 0.0067)
|
||||
allpass_gain: tuple[float, float] = (0.63, 0.51)
|
||||
|
||||
|
||||
class SharedUnitaryFdn:
|
||||
"""One shared late room driven by the sum of all object room sends."""
|
||||
|
||||
def __init__(self, config: LateFdnConfig = LateFdnConfig()):
|
||||
self.config = config
|
||||
self.sample_rate_hz = float(config.sample_rate_hz)
|
||||
self.rt60_seconds = float(config.rt60_seconds)
|
||||
self.damping = float(config.damping)
|
||||
self.output_gain = float(config.output_gain)
|
||||
if (not math.isfinite(self.sample_rate_hz) or self.sample_rate_hz <= 0.0
|
||||
or not math.isfinite(self.rt60_seconds) or self.rt60_seconds <= 0.0):
|
||||
raise ValueError("FDN sample rate and RT60 must be positive and finite")
|
||||
if (not math.isfinite(self.damping) or not 0.0 <= self.damping < 1.0
|
||||
or not math.isfinite(self.output_gain)):
|
||||
raise ValueError("invalid FDN damping/output gain")
|
||||
delay_seconds = np.asarray(config.delay_seconds, dtype=np.float64)
|
||||
allpass_seconds = np.asarray(config.allpass_seconds, dtype=np.float64)
|
||||
allpass_gain = np.asarray(config.allpass_gain, dtype=np.float64)
|
||||
if (delay_seconds.shape != (4,) or not np.isfinite(delay_seconds).all()
|
||||
or np.any(delay_seconds <= 0.0)):
|
||||
raise ValueError("FDN requires four positive finite delay times")
|
||||
if (allpass_seconds.shape != (2,) or not np.isfinite(allpass_seconds).all()
|
||||
or np.any(allpass_seconds <= 0.0)):
|
||||
raise ValueError("FDN requires two positive finite all-pass delay times")
|
||||
if (allpass_gain.shape != (2,) or not np.isfinite(allpass_gain).all()
|
||||
or np.any(np.abs(allpass_gain) >= 1.0)):
|
||||
raise ValueError("FDN requires two finite all-pass gains with magnitude below one")
|
||||
self.matrix = normalized_hadamard4()
|
||||
self.delays = np.asarray([
|
||||
_next_prime(round(seconds * self.sample_rate_hz))
|
||||
for seconds in delay_seconds
|
||||
], dtype=np.int32)
|
||||
self.feedback_gain = np.power(
|
||||
10.0, -3.0 * self.delays / (self.rt60_seconds * self.sample_rate_hz)
|
||||
).astype(np.float64)
|
||||
self.buffers = [np.zeros(int(delay), dtype=np.float64) for delay in self.delays]
|
||||
self.positions = np.zeros(4, dtype=np.int32)
|
||||
self.damping_state = np.zeros(4, dtype=np.float64)
|
||||
self.input_vector = 0.5 * np.asarray([1.0, -1.0, 1.0, 1.0], dtype=np.float64)
|
||||
self.output_matrix = 0.5 * np.asarray([
|
||||
[1.0, 1.0, -1.0, -1.0],
|
||||
[1.0, -1.0, 1.0, -1.0],
|
||||
], dtype=np.float64)
|
||||
self.diffusers = [
|
||||
SchroederAllpass(
|
||||
_next_prime(round(seconds * self.sample_rate_hz)), gain)
|
||||
for seconds, gain in zip(allpass_seconds, allpass_gain)
|
||||
]
|
||||
|
||||
@property
|
||||
def tail_samples(self) -> int:
|
||||
return int(math.ceil(1.5 * self.rt60_seconds * self.sample_rate_hz))
|
||||
|
||||
def reset(self) -> None:
|
||||
for buffer in self.buffers:
|
||||
buffer.fill(0.0)
|
||||
self.positions.fill(0)
|
||||
self.damping_state.fill(0.0)
|
||||
for diffuser in self.diffusers:
|
||||
diffuser.reset()
|
||||
|
||||
def process(self, mono) -> np.ndarray:
|
||||
values = np.asarray(mono, dtype=np.float64)
|
||||
if values.ndim != 1 or not np.isfinite(values).all():
|
||||
raise ValueError("FDN input must be one finite mono vector")
|
||||
diffused = values
|
||||
for diffuser in self.diffusers:
|
||||
diffused = diffuser.process(diffused)
|
||||
output = np.empty((len(values), 2), dtype=np.float64)
|
||||
for sample, value in enumerate(diffused):
|
||||
delayed = np.asarray([
|
||||
self.buffers[line][int(self.positions[line])]
|
||||
for line in range(4)
|
||||
], dtype=np.float64)
|
||||
self.damping_state = (
|
||||
self.damping * self.damping_state + (1.0 - self.damping) * delayed)
|
||||
output[sample] = self.output_gain * (self.output_matrix @ self.damping_state)
|
||||
feedback = self.matrix @ (self.damping_state * self.feedback_gain)
|
||||
write = self.input_vector * value + feedback
|
||||
for line in range(4):
|
||||
position = int(self.positions[line])
|
||||
self.buffers[line][position] = write[line]
|
||||
self.positions[line] = (position + 1) % int(self.delays[line])
|
||||
return output
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"name": "SharedUnitaryFdn",
|
||||
"sample_rate_hz": self.sample_rate_hz,
|
||||
"rt60_seconds": self.rt60_seconds,
|
||||
"delay_samples": [int(value) for value in self.delays],
|
||||
"feedback_gain": [float(value) for value in self.feedback_gain],
|
||||
"matrix_unitarity_max_error": float(
|
||||
np.max(np.abs(self.matrix.T @ self.matrix - np.eye(4)))),
|
||||
"allpass_delay_samples": [value.delay_samples for value in self.diffusers],
|
||||
"precision": "float64",
|
||||
}
|
||||
@@ -0,0 +1,115 @@
|
||||
"""Project-owned distance policy for the public SOFA renderer."""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
import math
|
||||
import numpy as np
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class DistanceState:
|
||||
profile: str
|
||||
normalized_radius: float
|
||||
reference_distance_m: float
|
||||
physical_distance_m: float
|
||||
direction_adm: np.ndarray
|
||||
|
||||
|
||||
class ReferenceDistanceProfileV1:
|
||||
"""Reference behavior, not a claim about any external public standard."""
|
||||
|
||||
DISTANCE_M = {
|
||||
"near": 1.00000465,
|
||||
"mid": 2.19327927,
|
||||
"far": 6.40177584,
|
||||
}
|
||||
MINIMUM_DISTANCE_M = 0.10
|
||||
# Distance profiles in the reference renderer are presentation presets,
|
||||
# not an instruction to attenuate already-authored programme PCM by 1/r.
|
||||
# Use an energy-normalized dry/room crossfade instead. The coefficient is
|
||||
# an explicit project calibration target.
|
||||
ROOM_ENERGY_COUPLING_PER_M2 = 0.01318359375
|
||||
PUBLIC_ROOM_CALIBRATION_GAIN = 1.4
|
||||
# Public, project-owned room coupling; it is not a SOFA or Dolby constant.
|
||||
LATE_SEND = {
|
||||
"near": 0.06,
|
||||
"mid": 0.16,
|
||||
"far": 0.28,
|
||||
}
|
||||
|
||||
@classmethod
|
||||
def validate_profile(cls, profile: str) -> str:
|
||||
value = str(profile).strip().lower()
|
||||
if value not in cls.DISTANCE_M:
|
||||
raise ValueError("distance profile must be near, mid, or far")
|
||||
return value
|
||||
|
||||
@classmethod
|
||||
def map_adm_position(cls, position, profile: str) -> DistanceState:
|
||||
name = cls.validate_profile(profile)
|
||||
values = np.asarray(position, dtype=np.float64)
|
||||
if values.shape != (3,) or not np.isfinite(values).all():
|
||||
raise ValueError("ADM position must contain three finite Cartesian values")
|
||||
radius = float(np.linalg.norm(values))
|
||||
direction = (values / radius if radius > 1.0e-15
|
||||
else np.asarray([0.0, 1.0, 0.0], dtype=np.float64))
|
||||
reference = float(cls.DISTANCE_M[name])
|
||||
distance = max(float(cls.MINIMUM_DISTANCE_M), radius * reference)
|
||||
return DistanceState(
|
||||
profile=name,
|
||||
normalized_radius=radius,
|
||||
reference_distance_m=reference,
|
||||
physical_distance_m=distance,
|
||||
direction_adm=np.asarray(direction, dtype=np.float64),
|
||||
)
|
||||
|
||||
@staticmethod
|
||||
def inverse_distance_gain(measurement_radius_m: float,
|
||||
path_distance_m: float) -> float:
|
||||
radius = float(measurement_radius_m)
|
||||
distance = float(path_distance_m)
|
||||
if not (math.isfinite(radius) and math.isfinite(distance)):
|
||||
raise ValueError("measurement and path distances must be finite")
|
||||
if radius <= 0.0 or distance <= 0.0:
|
||||
raise ValueError("measurement and path distances must be positive")
|
||||
return radius / distance
|
||||
|
||||
@classmethod
|
||||
def direct_level_gain(cls, state: DistanceState) -> float:
|
||||
"""Programme-normalized direct level for a distance presentation.
|
||||
|
||||
Near is the SOFA reference response. Mid/Far use an equal-power dry
|
||||
coefficient rather than a physical free-field 1/r attenuation. Room
|
||||
distance still changes through image-path lengths and late send.
|
||||
"""
|
||||
if state.profile == "near":
|
||||
return 1.0
|
||||
distance = float(state.physical_distance_m)
|
||||
return 1.0 / math.sqrt(
|
||||
1.0 + cls.ROOM_ENERGY_COUPLING_PER_M2 * distance * distance)
|
||||
|
||||
@classmethod
|
||||
def room_calibration_gain(cls, state: DistanceState) -> float:
|
||||
del state
|
||||
return float(cls.PUBLIC_ROOM_CALIBRATION_GAIN)
|
||||
|
||||
@classmethod
|
||||
def late_send(cls, state: DistanceState) -> float:
|
||||
base = float(cls.LATE_SEND[state.profile])
|
||||
radial = math.sqrt(max(state.normalized_radius, 0.0))
|
||||
return base * min(max(radial, 0.25), 1.5)
|
||||
|
||||
@classmethod
|
||||
def info(cls) -> dict:
|
||||
return {
|
||||
"name": "ReferenceDistanceProfileV1",
|
||||
"reference_distance_m": dict(cls.DISTANCE_M),
|
||||
"minimum_distance_m": cls.MINIMUM_DISTANCE_M,
|
||||
"direct_level_policy": (
|
||||
"Near unity; Mid/Far equal-power dry coefficient, never raw 1/r "
|
||||
"programme attenuation"),
|
||||
"room_energy_coupling_per_m2": cls.ROOM_ENERGY_COUPLING_PER_M2,
|
||||
"public_room_calibration_gain": cls.PUBLIC_ROOM_CALIBRATION_GAIN,
|
||||
"late_send": dict(cls.LATE_SEND),
|
||||
"standard_claim": False,
|
||||
}
|
||||
@@ -0,0 +1,308 @@
|
||||
"""Rosella .personalized_headphone binaural renderer.
|
||||
|
||||
Rosella JSON 解析由本项目自行实现(src/rosella_model.py),不调用任何 Dolby
|
||||
软件;.personalized_headphone 是用户经官方软件个性化扫描得到的模型文件。
|
||||
该路径与 SOFA 路径各自独立完成 HRTF/room 参数求值,只在最外层的 JOC 调度
|
||||
(1536-sample 帧缓冲、sample-timed OAMD timeline、512-sample 参数更新、输出
|
||||
包装)处汇合。
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import math
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
from binaural_metadata import OamdPositionTimeline
|
||||
from binaural_native_renderer import NativeBinauralDsp
|
||||
from rosella_core import RosellaRenderer
|
||||
from rosella_direct import BINAURAL_PROFILE_NAMES
|
||||
from rosella_filterbank import (
|
||||
DEFAULT_KERNEL_DATA,
|
||||
HybridAnalysis,
|
||||
HybridSynthesis,
|
||||
QmfAnalysis,
|
||||
QmfSynthesis,
|
||||
)
|
||||
from rosella_model import RosellaModel, load_personalized_headphone
|
||||
|
||||
SAMPLE_RATE = 48000
|
||||
FRAME_SAMPLES = 1536
|
||||
ROSSELLA_BLOCK_SAMPLES = 512
|
||||
QMF_HOP_SAMPLES = 64
|
||||
ROSSELLA_LATENCY_SAMPLES = 961
|
||||
SOURCE_CHANNELS = 16
|
||||
OUTPUT_CHANNELS = 2
|
||||
PROJECT_DIR = Path(__file__).resolve().parent.parent
|
||||
DEFAULT_PERSONALIZED_HEADPHONE = (
|
||||
PROJECT_DIR / "HRTF" / "binaural.personalized_headphone")
|
||||
|
||||
|
||||
def _sha256_file(path: Path) -> str:
|
||||
digest = hashlib.sha256()
|
||||
with path.open("rb") as stream:
|
||||
for block in iter(lambda: stream.read(1 << 20), b""):
|
||||
digest.update(block)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def resolve_personalized_headphone(path: str | Path | None = None) -> Path:
|
||||
target = (DEFAULT_PERSONALIZED_HEADPHONE if path is None
|
||||
else Path(path).expanduser().resolve())
|
||||
if not target.is_file():
|
||||
raise FileNotFoundError(
|
||||
f"未找到双耳模型:{target}\n"
|
||||
"请将兼容模型保存为 HRTF/binaural.personalized_headphone,"
|
||||
"或通过参数指定文件。"
|
||||
)
|
||||
return target
|
||||
|
||||
|
||||
class RosellaBinauralRenderer:
|
||||
"""Render interleaved LFE plus fifteen objects to stereo."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
personalized_headphone: str | Path | RosellaModel,
|
||||
*,
|
||||
mode: str = "mid",
|
||||
kernel_data: str | Path = DEFAULT_KERNEL_DATA,
|
||||
object_delay_samples: int = 1473,
|
||||
tail_seconds: float = 5.0,
|
||||
output_gain: float = 1.0,
|
||||
chunk_frames: int = 64,
|
||||
room_impulse_slots: int = 4096,
|
||||
backend: str = "python",
|
||||
native_library=None):
|
||||
if mode not in BINAURAL_PROFILE_NAMES:
|
||||
raise ValueError("binaural mode must be near, mid, or far")
|
||||
if int(object_delay_samples) < 0:
|
||||
raise ValueError("object_delay_samples must be non-negative")
|
||||
if float(tail_seconds) < 0.0:
|
||||
raise ValueError("tail_seconds must be non-negative")
|
||||
if int(chunk_frames) <= 0:
|
||||
raise ValueError("chunk_frames must be positive")
|
||||
if not math.isfinite(float(output_gain)):
|
||||
raise ValueError("output_gain must be finite")
|
||||
if backend not in ("auto", "native", "python"):
|
||||
raise ValueError("backend must be auto, native, or python")
|
||||
|
||||
if isinstance(personalized_headphone, RosellaModel):
|
||||
self.model = personalized_headphone
|
||||
self.model_path = Path(self.model.source_path)
|
||||
else:
|
||||
self.model_path = resolve_personalized_headphone(personalized_headphone)
|
||||
self.model = load_personalized_headphone(self.model_path)
|
||||
if self.model.sample_rate != SAMPLE_RATE:
|
||||
raise ValueError(
|
||||
f"Rosella model sample rate must be {SAMPLE_RATE}, got {self.model.sample_rate}")
|
||||
|
||||
self.mode = mode
|
||||
self.profile_index = BINAURAL_PROFILE_NAMES[mode]
|
||||
self.kernel_data = Path(kernel_data).expanduser().resolve()
|
||||
self.kernel_data_sha256 = _sha256_file(self.kernel_data)
|
||||
self.object_delay_samples = int(object_delay_samples)
|
||||
self.tail_seconds = float(tail_seconds)
|
||||
self.output_gain = np.float64(output_gain)
|
||||
self.chunk_frames = int(chunk_frames)
|
||||
self.chunk_samples = self.chunk_frames * FRAME_SAMPLES
|
||||
|
||||
self.native_dsp = None
|
||||
self.backend_fallback = None
|
||||
if backend in ("auto", "native"):
|
||||
try:
|
||||
self.native_dsp = NativeBinauralDsp(
|
||||
self.model, library_path=native_library,
|
||||
kernel_data=self.kernel_data)
|
||||
except (AttributeError, OSError, RuntimeError) as exc:
|
||||
if backend == "native":
|
||||
raise RuntimeError(f"native binaural backend unavailable: {exc}") from exc
|
||||
self.backend_fallback = str(exc)
|
||||
if self.native_dsp is not None:
|
||||
self.dsp_backend = "native"
|
||||
self.qmf_analysis = None
|
||||
self.hybrid_analysis = None
|
||||
self.hybrid_synthesis = None
|
||||
self.qmf_synthesis = None
|
||||
self.core = RosellaRenderer(
|
||||
self.model, SOURCE_CHANNELS, create_room=False)
|
||||
else:
|
||||
self.dsp_backend = "python"
|
||||
self.qmf_analysis = QmfAnalysis(SOURCE_CHANNELS, self.kernel_data)
|
||||
self.hybrid_analysis = HybridAnalysis(SOURCE_CHANNELS, self.kernel_data)
|
||||
self.core = RosellaRenderer(
|
||||
self.model, SOURCE_CHANNELS,
|
||||
room_impulse_slots=room_impulse_slots)
|
||||
self.hybrid_synthesis = HybridSynthesis(OUTPUT_CHANNELS, self.kernel_data)
|
||||
self.qmf_synthesis = QmfSynthesis(OUTPUT_CHANNELS, self.kernel_data)
|
||||
self.timeline = OamdPositionTimeline(15)
|
||||
|
||||
self._input_buffer = np.empty(
|
||||
(self.chunk_samples, SOURCE_CHANNELS), dtype=np.float64)
|
||||
self._buffer_used = 0
|
||||
self.input_samples = 0
|
||||
self.processed_input_samples = 0
|
||||
self.raw_output_samples = 0
|
||||
self.output_samples = 0
|
||||
self.finished = False
|
||||
self.metadata_block_updates = 0
|
||||
|
||||
def _append_input(self, samples: np.ndarray) -> list[np.ndarray]:
|
||||
outputs = []
|
||||
source = np.asarray(samples, dtype=np.float64)
|
||||
position = 0
|
||||
while position < len(source):
|
||||
count = min(self.chunk_samples - self._buffer_used,
|
||||
len(source) - position)
|
||||
self._input_buffer[self._buffer_used:self._buffer_used + count] = (
|
||||
source[position:position + count])
|
||||
self._buffer_used += count
|
||||
position += count
|
||||
if self._buffer_used == self.chunk_samples:
|
||||
outputs.append(self._process_samples(self._input_buffer))
|
||||
self._buffer_used = 0
|
||||
return outputs
|
||||
|
||||
def render_frame(self, objects16, payload=None, metadata_offset=None,
|
||||
*, outer_sample_offset=0) -> np.ndarray:
|
||||
"""Submit one 1536-sample reconstructed frame and its ID11 payload."""
|
||||
if self.finished:
|
||||
raise RuntimeError("binaural renderer is already finished")
|
||||
source = np.asarray(objects16)
|
||||
if source.shape != (FRAME_SAMPLES, SOURCE_CHANNELS):
|
||||
raise ValueError(
|
||||
f"binaural frame must have shape ({FRAME_SAMPLES},{SOURCE_CHANNELS}), "
|
||||
f"got {source.shape}")
|
||||
frame_start = self.input_samples
|
||||
metadata_delay = (self.object_delay_samples if metadata_offset is None
|
||||
else int(metadata_offset))
|
||||
if metadata_delay < 0:
|
||||
raise ValueError("metadata_offset must be non-negative")
|
||||
if payload is not None:
|
||||
self.timeline.submit_payload(
|
||||
payload,
|
||||
frame_start_sample=frame_start,
|
||||
outer_sample_offset=int(outer_sample_offset),
|
||||
object_delay_samples=metadata_delay,
|
||||
processed_sample=self.processed_input_samples,
|
||||
)
|
||||
self.metadata_block_updates += 1
|
||||
self.input_samples += FRAME_SAMPLES
|
||||
chunks = self._append_input(source)
|
||||
if not chunks:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
return np.concatenate(chunks, axis=0) if len(chunks) > 1 else chunks[0]
|
||||
|
||||
def _set_block_parameters(self, sample: int):
|
||||
positions = self.timeline.positions_at(sample)
|
||||
self.core.set_source(0, (0.0, 1.0, 0.0), special_lfe=True)
|
||||
for object_index in range(15):
|
||||
self.core.set_source(
|
||||
object_index + 1, positions[object_index], self.profile_index)
|
||||
|
||||
def _process_samples(self, source: np.ndarray) -> np.ndarray:
|
||||
values = np.asarray(source, dtype=np.float64)
|
||||
if values.ndim != 2 or values.shape[1] != SOURCE_CHANNELS:
|
||||
raise ValueError(f"expected [samples,{SOURCE_CHANNELS}], got {values.shape}")
|
||||
if len(values) % ROSSELLA_BLOCK_SAMPLES:
|
||||
raise ValueError("binaural input must be divisible by 512 samples")
|
||||
blocks = len(values) // ROSSELLA_BLOCK_SAMPLES
|
||||
block_base = self.processed_input_samples
|
||||
|
||||
if self.native_dsp is not None:
|
||||
stereo = np.empty((len(values), OUTPUT_CHANNELS), dtype=np.float64)
|
||||
for block in range(blocks):
|
||||
sample = block_base + block * ROSSELLA_BLOCK_SAMPLES
|
||||
self._set_block_parameters(sample)
|
||||
start = block * ROSSELLA_BLOCK_SAMPLES
|
||||
stop = start + ROSSELLA_BLOCK_SAMPLES
|
||||
stereo[start:stop] = self.native_dsp.process_block(
|
||||
values[start:stop], self.core.gains, self.core.room_sends,
|
||||
self.output_gain)
|
||||
else:
|
||||
hops = values.reshape(
|
||||
blocks, ROSSELLA_BLOCK_SAMPLES // QMF_HOP_SAMPLES,
|
||||
QMF_HOP_SAMPLES, SOURCE_CHANNELS,
|
||||
).transpose(0, 1, 3, 2).reshape(
|
||||
blocks * (ROSSELLA_BLOCK_SAMPLES // QMF_HOP_SAMPLES),
|
||||
SOURCE_CHANNELS, QMF_HOP_SAMPLES)
|
||||
hybrid = self.hybrid_analysis.process_chunk(
|
||||
self.qmf_analysis.process_chunk(hops))
|
||||
direct = np.empty((blocks * 8, OUTPUT_CHANNELS, 77), dtype=np.complex128)
|
||||
room_send = np.empty((blocks * 8, 77), dtype=np.complex128)
|
||||
for block in range(blocks):
|
||||
sample = block_base + block * ROSSELLA_BLOCK_SAMPLES
|
||||
self._set_block_parameters(sample)
|
||||
start = block * 8
|
||||
stop = start + 8
|
||||
direct[start:stop], room_send[start:stop] = (
|
||||
self.core.direct_and_send_static(hybrid[start:stop]))
|
||||
rendered = direct + self.core.room.process_chunk(room_send)
|
||||
time_bands = self.qmf_synthesis.process_chunk(
|
||||
self.hybrid_synthesis.process_chunk(rendered))
|
||||
stereo = time_bands.transpose(0, 2, 1).reshape(
|
||||
blocks * ROSSELLA_BLOCK_SAMPLES, OUTPUT_CHANNELS)
|
||||
stereo *= self.output_gain
|
||||
|
||||
skip = max(0, min(
|
||||
len(stereo), ROSSELLA_LATENCY_SAMPLES - self.raw_output_samples))
|
||||
self.raw_output_samples += len(stereo)
|
||||
self.processed_input_samples += len(values)
|
||||
output = stereo[skip:]
|
||||
self.output_samples += len(output)
|
||||
return output
|
||||
|
||||
def finish(self) -> np.ndarray:
|
||||
"""Process pending source samples and preserve the configured room tail."""
|
||||
if self.finished:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
outputs: list[np.ndarray] = []
|
||||
if self._buffer_used:
|
||||
outputs.append(self._process_samples(
|
||||
self._input_buffer[:self._buffer_used]))
|
||||
self._buffer_used = 0
|
||||
flush_samples = math.ceil(
|
||||
(self.tail_seconds * SAMPLE_RATE
|
||||
+ ROSSELLA_LATENCY_SAMPLES + ROSSELLA_BLOCK_SAMPLES)
|
||||
/ ROSSELLA_BLOCK_SAMPLES) * ROSSELLA_BLOCK_SAMPLES
|
||||
while flush_samples:
|
||||
count = min(flush_samples, self.chunk_samples)
|
||||
zero = np.zeros((count, SOURCE_CHANNELS), dtype=np.float64)
|
||||
outputs.append(self._process_samples(zero))
|
||||
flush_samples -= count
|
||||
self.finished = True
|
||||
nonempty = [value for value in outputs if len(value)]
|
||||
if not nonempty:
|
||||
return np.empty((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
return np.concatenate(nonempty, axis=0)
|
||||
|
||||
def close(self):
|
||||
if self.native_dsp is not None:
|
||||
self.native_dsp.close()
|
||||
self.finished = True
|
||||
|
||||
@property
|
||||
def backend_info(self) -> dict:
|
||||
return {
|
||||
"name": self.dsp_backend,
|
||||
"precision": "float64/complex128",
|
||||
"fallback_reason": self.backend_fallback,
|
||||
"library": (str(self.native_dsp.library_path)
|
||||
if self.native_dsp is not None else None),
|
||||
"model": str(self.model_path.resolve()),
|
||||
"model_coefficients": int(len(self.model.coefficients)),
|
||||
"model_coefficient_sha256": self.model.coefficient_sha256,
|
||||
"model_version": self.model.coefficient_version,
|
||||
"kernel_data": str(self.kernel_data),
|
||||
"kernel_data_sha256": self.kernel_data_sha256,
|
||||
"mode": self.mode,
|
||||
"latency_compensated_samples": ROSSELLA_LATENCY_SAMPLES,
|
||||
"object_delay_samples": self.object_delay_samples,
|
||||
"tail_seconds": self.tail_seconds,
|
||||
"metadata_payloads": self.timeline.payload_count,
|
||||
"metadata_position_transitions": self.timeline.transition_count,
|
||||
"input_samples": self.input_samples,
|
||||
"processed_samples_including_flush": self.processed_input_samples,
|
||||
"output_samples_before_tail_trim": self.output_samples,
|
||||
}
|
||||
@@ -0,0 +1,81 @@
|
||||
"""Stateful float64/complex128 Rosella hybrid-band renderer."""
|
||||
from __future__ import annotations
|
||||
|
||||
import numpy as np
|
||||
|
||||
from rosella_direct import (
|
||||
PROFILE_MID,
|
||||
direct_and_room_send,
|
||||
special_lfe_direct,
|
||||
)
|
||||
from rosella_model import RosellaModel
|
||||
from rosella_room import RosellaRoomFir
|
||||
|
||||
|
||||
class RosellaRenderer:
|
||||
"""Hold per-source direct parameters and the cross-block room state."""
|
||||
|
||||
def __init__(self, model: RosellaModel, source_count: int,
|
||||
room_impulse_slots: int = 4096, *, create_room: bool = True):
|
||||
if source_count <= 0:
|
||||
raise ValueError("source_count must be positive")
|
||||
self.model = model
|
||||
self.source_count = int(source_count)
|
||||
self.room = (RosellaRoomFir(model, impulse_slots=room_impulse_slots)
|
||||
if create_room else None)
|
||||
self.positions = np.zeros((self.source_count, 3), dtype=np.float64)
|
||||
self.positions[:, 1] = 1.0
|
||||
self.profiles = np.full(self.source_count, PROFILE_MID, dtype=np.int32)
|
||||
self.special_lfe = np.zeros(self.source_count, dtype=bool)
|
||||
self.gains = np.empty(
|
||||
(self.source_count, 2, 77), dtype=np.complex128)
|
||||
self.room_sends = np.empty(self.source_count, dtype=np.float64)
|
||||
self._parameter_keys = [None] * self.source_count
|
||||
for source in range(self.source_count):
|
||||
self.set_source(source, self.positions[source], PROFILE_MID)
|
||||
|
||||
def reset(self):
|
||||
if self.room is not None:
|
||||
self.room.reset()
|
||||
|
||||
def set_source(self, source: int, position, profile: int = PROFILE_MID,
|
||||
*, special_lfe: bool = False):
|
||||
source = int(source)
|
||||
if not 0 <= source < self.source_count:
|
||||
raise IndexError(source)
|
||||
coordinates = np.asarray(position, dtype=np.float64)
|
||||
if coordinates.shape != (3,) or not np.all(np.isfinite(coordinates)):
|
||||
raise ValueError(f"source position must be three finite values, got {position!r}")
|
||||
effective_profile = 0 if special_lfe else int(profile)
|
||||
key = ((bool(special_lfe), effective_profile)
|
||||
+ tuple(float(value) for value in coordinates))
|
||||
if self._parameter_keys[source] == key:
|
||||
return
|
||||
self.positions[source] = coordinates
|
||||
self.profiles[source] = effective_profile
|
||||
self.special_lfe[source] = bool(special_lfe)
|
||||
parameters = (special_lfe_direct() if special_lfe else
|
||||
direct_and_room_send(self.model, coordinates, effective_profile))
|
||||
self.gains[source] = parameters.gains
|
||||
self.room_sends[source] = parameters.room_send
|
||||
self._parameter_keys[source] = key
|
||||
|
||||
def direct_and_send_static(self, sources):
|
||||
"""Mix one static-parameter slot chunk without advancing room state."""
|
||||
values = np.asarray(sources, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.source_count, 77):
|
||||
raise ValueError(
|
||||
f"expected [slots,{self.source_count},77], got {values.shape}")
|
||||
direct = np.zeros((values.shape[0], 2, 77), dtype=np.complex128)
|
||||
room_send = np.zeros((values.shape[0], 77), dtype=np.complex128)
|
||||
for source in range(self.source_count - 1, -1, -1):
|
||||
direct += values[:, source, None, :] * self.gains[source][None, :, :]
|
||||
room_send += values[:, source, :] * self.room_sends[source]
|
||||
return direct, room_send
|
||||
|
||||
def process_static_chunk(self, sources) -> np.ndarray:
|
||||
direct, room_send = self.direct_and_send_static(sources)
|
||||
if self.room is None:
|
||||
raise RuntimeError("room renderer is not configured")
|
||||
direct += self.room.process_chunk(room_send)
|
||||
return direct
|
||||
@@ -0,0 +1,302 @@
|
||||
"""Float64 Rosella direction, distance, HRTF, and room-send calculations."""
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
from dataclasses import dataclass
|
||||
|
||||
import numpy as np
|
||||
|
||||
from rosella_model import RosellaModel, direction_basis
|
||||
|
||||
PROFILE_NEAR = 1
|
||||
PROFILE_FAR = 2
|
||||
PROFILE_MID = 3
|
||||
BINAURAL_PROFILE_NAMES = {
|
||||
"near": PROFILE_NEAR,
|
||||
"far": PROFILE_FAR,
|
||||
"mid": PROFILE_MID,
|
||||
}
|
||||
|
||||
_SPECIAL_LFE_LOW_16 = np.asarray([
|
||||
0x402695EA, 0x3FE75979, 0x3F28CAAA, 0xBCE1FB2E,
|
||||
0xBDD8AF65, 0xBD8F426E, 0x3D996821, 0xBC16B3A0,
|
||||
0x3B64BAF1, 0xBC81ECFD, 0xBA3D892F, 0x3AF6A9F0,
|
||||
0xB9DD1C5F, 0x380A193F, 0x38052059, 0x351BCB34,
|
||||
], dtype=np.uint32).view(np.float32).astype(np.float64)
|
||||
_CENTRE_EQUAL = 0.9998489618301392
|
||||
_CENTRE_ALTERNATE = 0.7070000171661377
|
||||
|
||||
_FIELD_CACHE: dict[int, tuple[np.ndarray, np.ndarray]] = {}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class DirectResult:
|
||||
gains: np.ndarray # complex128 [ear=2, hybrid_band=77]
|
||||
room_send: np.float64
|
||||
physical_radius_m: np.float64
|
||||
normalized_radius: np.float64
|
||||
clamped_radius: np.float64
|
||||
delay_samples: np.float64
|
||||
delayed_ear: int | None
|
||||
|
||||
|
||||
def special_lfe_direct() -> DirectResult:
|
||||
"""Return the fixed 16-band low-pass used by a special/LFE source."""
|
||||
mono = np.zeros(77, dtype=np.complex128)
|
||||
mono[:16] = _SPECIAL_LFE_LOW_16
|
||||
return DirectResult(
|
||||
gains=np.repeat(mono[None, :], 2, axis=0),
|
||||
room_send=np.float64(0.0),
|
||||
physical_radius_m=np.float64(0.0),
|
||||
normalized_radius=np.float64(0.0),
|
||||
clamped_radius=np.float64(0.0),
|
||||
delay_samples=np.float64(0.0),
|
||||
delayed_ear=None,
|
||||
)
|
||||
|
||||
|
||||
def _round_away_from_zero(value: float) -> int:
|
||||
return math.floor(value + 0.5) if value >= 0.0 else math.ceil(value - 0.5)
|
||||
|
||||
|
||||
def _q15_position(position) -> np.ndarray:
|
||||
"""Quantize ADM Cartesian coordinates to the Rosella metadata grid.
|
||||
|
||||
Quantization is metadata decoding. The returned integer lanes are promoted
|
||||
to float64 before any geometry is evaluated.
|
||||
"""
|
||||
x, y, z = (float(value) for value in position)
|
||||
encoded = (
|
||||
min(max((x + 1.0) * 0.5, 0.0), 1.0),
|
||||
min(max((1.0 - y) * 0.5, 0.0), 1.0),
|
||||
min(max(z, -1.0), 1.0),
|
||||
)
|
||||
return np.asarray([
|
||||
min(_round_away_from_zero(value * 32768.0), 32767)
|
||||
for value in encoded
|
||||
], dtype=np.int32)
|
||||
|
||||
|
||||
def _profile_geometry(model: RosellaModel, position, profile_index: int):
|
||||
if profile_index not in (PROFILE_NEAR, PROFILE_FAR, PROFILE_MID):
|
||||
raise ValueError("binaural object profile must be near, mid, or far")
|
||||
profile = model.profiles[profile_index]
|
||||
encoded = _q15_position(position)
|
||||
q_front = 1.0 - 2.0 * float(encoded[1]) / 32768.0
|
||||
q_x = 2.0 * float(encoded[0]) / 32768.0 - 1.0
|
||||
q_vertical = float(encoded[2]) / 32768.0
|
||||
|
||||
if int(model.header_integer_fields[0]) != 0:
|
||||
if q_x == 0.0 and q_front == 0.0:
|
||||
mapped_front = 0.0
|
||||
mapped_lateral = 0.0
|
||||
mapped_vertical = q_vertical
|
||||
else:
|
||||
horizontal_max = max(abs(q_x), abs(q_front))
|
||||
horizontal_norm = ((q_x / horizontal_max) ** 2
|
||||
+ (q_front / horizontal_max) ** 2)
|
||||
if q_vertical == 0.0:
|
||||
vertical_norm = 1.0
|
||||
else:
|
||||
smaller = min(abs(q_vertical), horizontal_max)
|
||||
larger = max(abs(q_vertical), horizontal_max)
|
||||
vertical_norm = 1.0 + (smaller / larger) ** 2
|
||||
horizontal_factor = 1.0 / math.sqrt(horizontal_norm * vertical_norm)
|
||||
vertical_factor = 1.0 / math.sqrt(vertical_norm)
|
||||
mapped_front = q_front * horizontal_factor
|
||||
mapped_lateral = -q_x * horizontal_factor
|
||||
mapped_vertical = q_vertical * vertical_factor
|
||||
else:
|
||||
mapped_front = q_front
|
||||
mapped_lateral = -q_x
|
||||
mapped_vertical = q_vertical
|
||||
|
||||
scales = np.asarray(profile.axis_scales_internal, dtype=np.float64)
|
||||
scaled = np.asarray([
|
||||
mapped_front * scales[2],
|
||||
mapped_lateral * scales[0],
|
||||
mapped_vertical * scales[1],
|
||||
], dtype=np.float64)
|
||||
bounds = np.asarray(profile.bounds, dtype=np.float64)
|
||||
ray = 1.0
|
||||
for axis in range(3):
|
||||
value = scaled[axis]
|
||||
lower, upper = bounds[axis * 2:axis * 2 + 2]
|
||||
if value < lower:
|
||||
ray = min(ray, lower / value)
|
||||
elif value > upper:
|
||||
ray = min(ray, upper / value)
|
||||
if ray < 1.0:
|
||||
scaled *= ray
|
||||
|
||||
radius = float(np.linalg.norm(scaled))
|
||||
clamped = max(radius, float(profile.minimum_normalized_radius))
|
||||
alpha = radius / clamped
|
||||
direction = (scaled / radius if radius > 1.0e-30
|
||||
else np.asarray([1.0, 0.0, 0.0], dtype=np.float64))
|
||||
return profile, direction, radius, clamped, alpha
|
||||
|
||||
|
||||
def _logical_field(padded: np.ndarray) -> np.ndarray:
|
||||
result = np.empty((77, 36, 2), dtype=np.float64)
|
||||
source = np.asarray(padded, dtype=np.float64)
|
||||
for band in range(77):
|
||||
block, lane = divmod(band, 4)
|
||||
for term in range(36):
|
||||
for component in range(2):
|
||||
result[band, term, component] = source[
|
||||
lane + 4 * (term * 2 + component + 72 * block)]
|
||||
return result
|
||||
|
||||
|
||||
def _model_fields(model: RosellaModel) -> tuple[np.ndarray, np.ndarray]:
|
||||
key = id(model)
|
||||
fields = _FIELD_CACHE.get(key)
|
||||
if fields is None:
|
||||
fields = (_logical_field(model.field_left_padded),
|
||||
_logical_field(model.field_right_padded))
|
||||
_FIELD_CACHE[key] = fields
|
||||
return fields
|
||||
|
||||
|
||||
def _ear_geometry(model: RosellaModel, profile, direction, clamped: float,
|
||||
offset: float, correction: float):
|
||||
x, y, z = (float(value) for value in direction)
|
||||
inverse_distance = float(profile.inverse_distance_per_m)
|
||||
ear = float(offset) * inverse_distance / clamped
|
||||
y_minus = y - ear
|
||||
y_plus = y + ear
|
||||
common = x * x + z * z
|
||||
length_minus = math.sqrt(y_minus * y_minus + common)
|
||||
length_plus = math.sqrt(y_plus * y_plus + common)
|
||||
basis_minus = direction_basis(
|
||||
x / length_minus, y_minus / length_minus, z / length_minus,
|
||||
dtype=np.float64)
|
||||
basis_plus = direction_basis(
|
||||
x / length_plus, y_plus / length_plus, z / length_plus,
|
||||
dtype=np.float64)
|
||||
path_minus = length_minus * clamped
|
||||
path_plus = length_plus * clamped
|
||||
if correction != 0.0:
|
||||
multiplier = 2.0 * float(correction) * inverse_distance
|
||||
path_minus += max(float(np.dot(
|
||||
np.asarray(model.vector_left, dtype=np.float64), basis_minus)), 0.0) * multiplier
|
||||
path_plus += max(float(np.dot(
|
||||
np.asarray(model.vector_right, dtype=np.float64), basis_plus)), 0.0) * multiplier
|
||||
return basis_minus, basis_plus, path_minus, path_plus
|
||||
|
||||
|
||||
def _phase_groups(model: RosellaModel, delay_samples: float) -> np.ndarray:
|
||||
result = np.ones(77, dtype=np.complex128)
|
||||
current = 1.0 + 0.0j
|
||||
step = 1.0 + 0.0j
|
||||
value_index = 0
|
||||
for band, flag in enumerate(model.hybrid_flags):
|
||||
if flag != 2:
|
||||
if flag == 1:
|
||||
angle = float(model.hybrid_values[value_index]) * delay_samples
|
||||
value_index += 1
|
||||
step = complex(math.cos(angle), math.sin(angle))
|
||||
current *= step
|
||||
result[band] = current
|
||||
return result
|
||||
|
||||
|
||||
def direct_and_room_send(model: RosellaModel, position,
|
||||
profile_index: int) -> DirectResult:
|
||||
"""Evaluate one ordinary source using float64/complex128 throughout."""
|
||||
profile, direction, radius, clamped, alpha = _profile_geometry(
|
||||
model, position, profile_index)
|
||||
|
||||
_, _, path_minus, path_plus = _ear_geometry(
|
||||
model, profile, direction, clamped,
|
||||
float(model.model_scalars[1]), float(model.model_scalars[2]))
|
||||
delay = (abs(path_plus - path_minus)
|
||||
* float(profile.distance_scale_m)
|
||||
* (float(model.sample_rate) / 343.3) * alpha)
|
||||
delayed_ear = 0 if path_minus > path_plus else (
|
||||
1 if path_plus > path_minus else None)
|
||||
|
||||
_, _, weight_minus_path, weight_plus_path = _ear_geometry(
|
||||
model, profile, direction, clamped,
|
||||
float(model.model_scalars[3]), float(model.model_scalars[4]))
|
||||
weight_norm = math.sqrt(
|
||||
weight_minus_path * weight_minus_path
|
||||
+ weight_plus_path * weight_plus_path)
|
||||
weight_left = weight_plus_path / weight_norm
|
||||
weight_right = weight_minus_path / weight_norm
|
||||
|
||||
final_offset = float(model.model_scalars[0])
|
||||
if final_offset == 0.0:
|
||||
basis_minus = direction_basis(*direction, dtype=np.float64)
|
||||
basis_plus = basis_minus.copy()
|
||||
else:
|
||||
x, y, z = (float(value) for value in direction)
|
||||
ear = final_offset * float(profile.inverse_distance_per_m) / clamped
|
||||
y_minus = y - ear
|
||||
y_plus = y + ear
|
||||
common_length = x * x + z * z
|
||||
length_minus = math.sqrt(y_minus * y_minus + common_length)
|
||||
length_plus = math.sqrt(y_plus * y_plus + common_length)
|
||||
basis_minus = direction_basis(
|
||||
x / length_minus, y_minus / length_minus, z / length_minus,
|
||||
dtype=np.float64)
|
||||
basis_plus = direction_basis(
|
||||
x / length_plus, y_plus / length_plus, z / length_plus,
|
||||
dtype=np.float64)
|
||||
|
||||
field_left, field_right = _model_fields(model)
|
||||
left_components = np.einsum(
|
||||
"bjc,j->bc", field_left, basis_minus,
|
||||
dtype=np.float64, optimize=False)
|
||||
right_components = np.einsum(
|
||||
"bjc,j->bc", field_right, basis_plus,
|
||||
dtype=np.float64, optimize=False)
|
||||
left = left_components[:, 0] + 1j * left_components[:, 1]
|
||||
right = right_components[:, 0] + 1j * right_components[:, 1]
|
||||
if delayed_ear is not None:
|
||||
phase = _phase_groups(model, delay)
|
||||
if delayed_ear == 0:
|
||||
left *= phase
|
||||
else:
|
||||
right *= phase
|
||||
|
||||
effective_radius = (radius * float(model.header_float_scalars[0])
|
||||
* float(profile.distance_scale_m))
|
||||
if profile_index in (PROFILE_FAR, PROFILE_MID):
|
||||
common = 1.0 / math.sqrt(
|
||||
1.0 + float(model.header_float_scalars[1])
|
||||
* effective_radius * effective_radius)
|
||||
room_send = effective_radius * common
|
||||
else:
|
||||
common = 1.0
|
||||
room_send = 0.0
|
||||
|
||||
left_term0 = field_left[:, 0, 0] + 1j * field_left[:, 0, 1]
|
||||
right_term0 = field_right[:, 0, 0] + 1j * field_right[:, 0, 1]
|
||||
weights_are_default_equal = (
|
||||
float(model.model_scalars[3]) == 0.0
|
||||
and float(model.model_scalars[4]) == 0.0)
|
||||
if weights_are_default_equal:
|
||||
centre_left = weight_left * (1.0 - alpha) * _CENTRE_EQUAL
|
||||
centre_right = centre_left
|
||||
right_direction_weight = weight_left
|
||||
else:
|
||||
centre_left = (1.0 - alpha) * _CENTRE_ALTERNATE
|
||||
centre_right = centre_left
|
||||
right_direction_weight = weight_right
|
||||
|
||||
gains = np.empty((2, 77), dtype=np.complex128)
|
||||
gains[0] = common * (
|
||||
left * (weight_left * alpha) + left_term0 * centre_left)
|
||||
gains[1] = common * (
|
||||
right * (right_direction_weight * alpha) + right_term0 * centre_right)
|
||||
return DirectResult(
|
||||
gains=gains,
|
||||
room_send=np.float64(room_send),
|
||||
physical_radius_m=np.float64(float(profile.distance_scale_m) * radius),
|
||||
normalized_radius=np.float64(radius),
|
||||
clamped_radius=np.float64(clamped),
|
||||
delay_samples=np.float64(delay),
|
||||
delayed_ear=delayed_ear,
|
||||
)
|
||||
@@ -0,0 +1,190 @@
|
||||
"""Float64/complex128 Rosella QMF and hybrid filterbanks."""
|
||||
from __future__ import annotations
|
||||
|
||||
from functools import lru_cache
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
PROJECT_DIR = Path(__file__).resolve().parent.parent
|
||||
DEFAULT_KERNEL_DATA = PROJECT_DIR / "data" / "rosella_kernels.npz"
|
||||
|
||||
|
||||
@lru_cache(maxsize=4)
|
||||
def _load_tables(path_string: str) -> dict[str, np.ndarray]:
|
||||
path = Path(path_string)
|
||||
if not path.is_file():
|
||||
raise FileNotFoundError(f"Rosella kernel data not found: {path}")
|
||||
with np.load(path, allow_pickle=False) as archive:
|
||||
version = archive["format_version"]
|
||||
if version.shape != (1,) or int(version[0]) != 1:
|
||||
raise ValueError(f"unsupported Rosella kernel data version in {path}")
|
||||
return {name: archive[name].copy() for name in archive.files}
|
||||
|
||||
|
||||
def load_kernel_tables(path: str | Path = DEFAULT_KERNEL_DATA) -> dict[str, np.ndarray]:
|
||||
"""Load and cache the compact, production Rosella kernel tables."""
|
||||
return _load_tables(str(Path(path).expanduser().resolve()))
|
||||
|
||||
|
||||
class QmfAnalysis:
|
||||
"""Batchable 64-band analysis with float64 state and complex128 FFTs."""
|
||||
|
||||
def __init__(self, channels: int, kernel_data: str | Path = DEFAULT_KERNEL_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_kernel_tables(kernel_data)
|
||||
self.coefficients = np.asarray(
|
||||
tables["qmf_analysis_coefficients"], dtype=np.float64)
|
||||
if self.coefficients.shape != (64, 10):
|
||||
raise ValueError("invalid qmf_analysis_coefficients shape")
|
||||
self.channels = int(channels)
|
||||
self.history = np.zeros((9, self.channels, 64), dtype=np.float64)
|
||||
phase = np.arange(64, dtype=np.float64)
|
||||
self.premod = np.exp(-1j * np.pi * phase / 128.0).astype(np.complex128)
|
||||
self.post = np.exp(
|
||||
-1j * 3.0 * (np.arange(64, dtype=np.float64) + 0.5) * np.pi / 128.0
|
||||
).astype(np.complex128)
|
||||
self.even_post = (
|
||||
1j * ((-1.0) ** np.arange(64, dtype=np.float64))
|
||||
).astype(np.complex128)
|
||||
|
||||
def reset(self):
|
||||
self.history.fill(0.0)
|
||||
|
||||
def process_chunk(self, hops) -> np.ndarray:
|
||||
values = np.asarray(hops, dtype=np.float64)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 64):
|
||||
raise ValueError(f"expected [slots,{self.channels},64], got {values.shape}")
|
||||
count = values.shape[0]
|
||||
joined = np.concatenate((self.history, values), axis=0)
|
||||
even = np.zeros_like(values)
|
||||
odd = np.zeros_like(values)
|
||||
for lag in range(10):
|
||||
source = joined[9 - lag:9 - lag + count]
|
||||
target = even if lag % 2 == 0 else odd
|
||||
target += source * self.coefficients[:, lag][None, None, :]
|
||||
self.history[:] = joined[-9:]
|
||||
|
||||
def transform(block):
|
||||
prepared = block.astype(np.complex128, copy=False) * self.premod
|
||||
transformed = np.fft.fft(prepared, n=128, axis=-1)[..., :64]
|
||||
return transformed * self.post
|
||||
|
||||
return np.asarray(transform(odd) + transform(even) * self.even_post,
|
||||
dtype=np.complex128)
|
||||
|
||||
|
||||
class HybridAnalysis:
|
||||
"""Sparse 64-QMF to 77-hybrid analysis in float64/complex128."""
|
||||
|
||||
def __init__(self, channels: int, kernel_data: str | Path = DEFAULT_KERNEL_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_kernel_tables(kernel_data)
|
||||
self.low_kernel = np.asarray(
|
||||
tables["hybrid_analysis_low_kernel"], dtype=np.float64)
|
||||
if self.low_kernel.shape != (3, 2, 13, 16, 2):
|
||||
raise ValueError("invalid hybrid_analysis_low_kernel shape")
|
||||
self.channels = int(channels)
|
||||
self.history = np.zeros((12, self.channels, 3, 2), dtype=np.float64)
|
||||
self.high_history = np.zeros(
|
||||
(6, self.channels, 61), dtype=np.complex128)
|
||||
|
||||
def reset(self):
|
||||
self.history.fill(0.0)
|
||||
self.high_history.fill(0.0)
|
||||
|
||||
def process_chunk(self, qmf) -> np.ndarray:
|
||||
values = np.asarray(qmf, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 64):
|
||||
raise ValueError(f"expected [slots,{self.channels},64], got {values.shape}")
|
||||
count = values.shape[0]
|
||||
low = np.stack((values[:, :, :3].real, values[:, :, :3].imag), axis=-1)
|
||||
joined = np.concatenate((self.history, low), axis=0)
|
||||
output = np.zeros((count, self.channels, 77, 2), dtype=np.float64)
|
||||
for lag in range(13):
|
||||
source = joined[12 - lag:12 - lag + count]
|
||||
output[:, :, :16] += np.einsum(
|
||||
"tcpi,pibo->tcbo", source, self.low_kernel[:, :, lag],
|
||||
dtype=np.float64, optimize=False)
|
||||
self.history[:] = joined[-12:]
|
||||
|
||||
high_joined = np.concatenate((self.high_history, values[:, :, 3:]), axis=0)
|
||||
high = high_joined[:count]
|
||||
output[:, :, 16:, 0] = high.real
|
||||
output[:, :, 16:, 1] = high.imag
|
||||
self.high_history[:] = high_joined[-6:]
|
||||
return np.asarray(output[..., 0] + 1j * output[..., 1], dtype=np.complex128)
|
||||
|
||||
|
||||
class HybridSynthesis:
|
||||
"""Instantaneous sparse 77-hybrid to 64-QMF synthesis map."""
|
||||
|
||||
def __init__(self, channels: int, kernel_data: str | Path = DEFAULT_KERNEL_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_kernel_tables(kernel_data)
|
||||
indices = np.asarray(tables["hybrid_synthesis_indices"], dtype=np.int64)
|
||||
values = np.asarray(tables["hybrid_synthesis_values"], dtype=np.float64)
|
||||
if indices.ndim != 2 or indices.shape[1] != 4 or len(indices) != len(values):
|
||||
raise ValueError("invalid hybrid synthesis sparse table")
|
||||
self.mapping = [
|
||||
(int(index[0]), int(index[1]), int(index[2]), int(index[3]), float(value))
|
||||
for index, value in zip(indices, values)
|
||||
]
|
||||
self.channels = int(channels)
|
||||
|
||||
def reset(self):
|
||||
return None
|
||||
|
||||
def process_chunk(self, hybrid) -> np.ndarray:
|
||||
values = np.asarray(hybrid, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 77):
|
||||
raise ValueError(f"expected [slots,{self.channels},77], got {values.shape}")
|
||||
source = np.stack((values.real, values.imag), axis=-1)
|
||||
output = np.zeros((values.shape[0], self.channels, 64, 2), dtype=np.float64)
|
||||
for input_band, input_component, output_band, output_component, gain in self.mapping:
|
||||
output[:, :, output_band, output_component] += (
|
||||
source[:, :, input_band, input_component] * gain)
|
||||
return np.asarray(output[..., 0] + 1j * output[..., 1], dtype=np.complex128)
|
||||
|
||||
|
||||
class QmfSynthesis:
|
||||
"""Rank-4 64-band synthesis with float64 state and accumulation."""
|
||||
|
||||
def __init__(self, channels: int, kernel_data: str | Path = DEFAULT_KERNEL_DATA):
|
||||
if channels <= 0:
|
||||
raise ValueError("channels must be positive")
|
||||
tables = load_kernel_tables(kernel_data)
|
||||
self.basis = np.asarray(tables["qmf_synthesis_basis"], dtype=np.float64)
|
||||
self.taps = np.asarray(tables["qmf_synthesis_taps"], dtype=np.float64)
|
||||
if self.basis.shape != (64, 4, 128) or self.taps.shape != (64, 10, 4):
|
||||
raise ValueError("invalid QMF synthesis factorization")
|
||||
self.channels = int(channels)
|
||||
self.rank = 4
|
||||
self.history = np.zeros(
|
||||
(9, self.channels, 64, self.rank), dtype=np.float64)
|
||||
|
||||
def reset(self):
|
||||
self.history.fill(0.0)
|
||||
|
||||
def process_chunk(self, qmf) -> np.ndarray:
|
||||
values = np.asarray(qmf, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.channels, 64):
|
||||
raise ValueError(f"expected [slots,{self.channels},64], got {values.shape}")
|
||||
count = values.shape[0]
|
||||
flat = np.stack((values.real, values.imag), axis=-1).reshape(
|
||||
count * self.channels, 128)
|
||||
modulation = self.basis.reshape(64 * self.rank, 128)
|
||||
features = (flat @ modulation.T).reshape(
|
||||
count, self.channels, 64, self.rank)
|
||||
joined = np.concatenate((self.history, features), axis=0)
|
||||
output = np.zeros((count, self.channels, 64), dtype=np.float64)
|
||||
for lag in range(10):
|
||||
output += np.sum(
|
||||
joined[9 - lag:9 - lag + count]
|
||||
* self.taps[:, lag, :][None, None, :, :],
|
||||
axis=-1, dtype=np.float64)
|
||||
self.history[:] = joined[-9:]
|
||||
return output
|
||||
@@ -0,0 +1,517 @@
|
||||
"""Parser for ``.personalized_headphone`` and raw ``rp`` models."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import math
|
||||
from numbers import Real
|
||||
import struct
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
|
||||
Q15 = np.float32(1.0 / 32768.0)
|
||||
|
||||
|
||||
def _f32(value) -> np.float32:
|
||||
return np.float32(value)
|
||||
|
||||
|
||||
def _q15(value: int) -> np.float32:
|
||||
return _f32(_f32(value) * Q15)
|
||||
|
||||
|
||||
def _q15_exp(value: int, exponent: int) -> np.float32:
|
||||
return _f32(_q15(value) * _f32(np.ldexp(1.0, exponent)))
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class DistanceProfile:
|
||||
bounds: np.ndarray
|
||||
distance_scale_m: np.float32
|
||||
inverse_distance_per_m: np.float32
|
||||
axis_scales_internal: np.ndarray
|
||||
minimum_normalized_radius: np.float32
|
||||
|
||||
@property
|
||||
def floats(self) -> np.ndarray:
|
||||
return np.concatenate((
|
||||
self.bounds,
|
||||
np.asarray([self.distance_scale_m,
|
||||
self.inverse_distance_per_m], dtype=np.float32),
|
||||
self.axis_scales_internal,
|
||||
np.asarray([self.minimum_normalized_radius], dtype=np.float32),
|
||||
))
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RosellaModel:
|
||||
source_path: str
|
||||
coefficients: np.ndarray
|
||||
coefficient_sha256: str
|
||||
coefficient_version: str | None
|
||||
room_model: str | None
|
||||
table_a_dimension: int
|
||||
table_a_option: int
|
||||
table_a_extra: int
|
||||
table_a_header_field: int
|
||||
table_a_header_25: int
|
||||
table_a_control: int
|
||||
table_a_option_ids: np.ndarray
|
||||
table_a_option_values: np.ndarray
|
||||
table_a_scalar: np.float32
|
||||
table_a_filter_16x64_padded: np.ndarray
|
||||
table_a_four_integers: np.ndarray
|
||||
table_a_integer: int
|
||||
table_a_filter_8x64_padded: np.ndarray
|
||||
table_a_vector16: np.ndarray
|
||||
table_a_filter_4x64_padded: np.ndarray
|
||||
table_a_extra_indices: np.ndarray
|
||||
table_a_extra_fields_padded: np.ndarray
|
||||
table_a_extra_vectors: np.ndarray
|
||||
sample_rate: int
|
||||
matrix_exponent: int
|
||||
field_exponent: int
|
||||
matrix_left: np.ndarray
|
||||
matrix_right: np.ndarray
|
||||
vector_left: np.ndarray
|
||||
vector_right: np.ndarray
|
||||
field_left_padded: np.ndarray
|
||||
field_right_padded: np.ndarray
|
||||
field_left_odd_serialized_zero: bool
|
||||
hybrid_flags: np.ndarray
|
||||
hybrid_values: np.ndarray
|
||||
model_scalars: np.ndarray
|
||||
header_float_scalars: np.ndarray
|
||||
header_integer_fields: np.ndarray
|
||||
profiles: tuple[DistanceProfile, ...]
|
||||
profile_tail: np.ndarray
|
||||
post_fields: np.ndarray
|
||||
table_a_main_serialized: np.ndarray
|
||||
|
||||
|
||||
def _lane(data: bytes, index: int) -> int:
|
||||
if (index + 1) * 4 > len(data):
|
||||
raise ValueError(f"Rosella rp truncated before int32 lane {index}")
|
||||
return struct.unpack_from("<I", data, index * 4)[0]
|
||||
|
||||
|
||||
def inspect_rp(data: bytes) -> dict:
|
||||
"""Return the active-lane layout and checksum status for one raw rp image."""
|
||||
if len(data) < 20 or len(data) % 4:
|
||||
raise ValueError("Rosella rp must contain whole little-endian int32 lanes")
|
||||
if _lane(data, 0) != 0x7072:
|
||||
raise ValueError(f"bad Rosella rp magic: 0x{_lane(data, 0):08X}")
|
||||
|
||||
low16 = lambda value: value & 0xFFFF
|
||||
checksum = low16(_lane(data, 1))
|
||||
table_a_present = low16(_lane(data, 2))
|
||||
table_b_present = low16(_lane(data, 3))
|
||||
table_c_present = low16(_lane(data, 4))
|
||||
index = 5
|
||||
if table_a_present:
|
||||
table_a_dimension = low16(_lane(data, index))
|
||||
table_a_option = low16(_lane(data, index + 1))
|
||||
table_a_extra = low16(_lane(data, index + 2))
|
||||
index += 5
|
||||
else:
|
||||
table_a_dimension, table_a_option, table_a_extra = 77, 0, 0
|
||||
if table_b_present:
|
||||
if not table_a_present:
|
||||
raise ValueError("Rosella rp table B cannot be present without table A")
|
||||
table_b_dimension = low16(_lane(data, index))
|
||||
table_b_extra = low16(_lane(data, index + 1))
|
||||
table_b_groups = low16(_lane(data, index + 2))
|
||||
index += 3
|
||||
else:
|
||||
table_b_dimension = table_b_extra = table_b_groups = 0
|
||||
if table_c_present:
|
||||
table_c_dimension = low16(_lane(data, index))
|
||||
index += 1
|
||||
else:
|
||||
table_c_dimension = 0
|
||||
|
||||
payload_words = (
|
||||
(index - 2)
|
||||
+ table_b_present * (
|
||||
table_b_dimension + 380 * table_b_groups + table_b_extra + 79)
|
||||
+ table_a_present * (
|
||||
171 * table_a_extra + 79 + 2 * (table_a_option + 14 * table_a_dimension))
|
||||
+ 11
|
||||
+ table_c_present * (314 * table_c_dimension + 1)
|
||||
)
|
||||
total_lanes = 2 + payload_words
|
||||
if len(data) < total_lanes * 4:
|
||||
raise ValueError(
|
||||
f"Rosella rp truncated: need {total_lanes * 4} bytes, have {len(data)}")
|
||||
computed = 0xA569
|
||||
for lane_index in range(2, total_lanes):
|
||||
computed ^= low16(_lane(data, lane_index))
|
||||
computed &= 0xFFFF
|
||||
return {
|
||||
"stored_checksum": checksum,
|
||||
"computed_checksum": computed,
|
||||
"checksum_valid": computed == checksum,
|
||||
"table_a_present": table_a_present,
|
||||
"table_b_present": table_b_present,
|
||||
"table_c_present": table_c_present,
|
||||
"table_a_dimension": table_a_dimension,
|
||||
"table_a_option": table_a_option,
|
||||
"table_a_extra": table_a_extra,
|
||||
"table_b_dimension": table_b_dimension,
|
||||
"table_b_extra": table_b_extra,
|
||||
"table_b_groups": table_b_groups,
|
||||
"table_c_dimension": table_c_dimension,
|
||||
"active_int32_lanes": total_lanes,
|
||||
}
|
||||
|
||||
|
||||
def _json_int32(values) -> np.ndarray:
|
||||
if not isinstance(values, list):
|
||||
raise ValueError("rosella_coefficients must be a JSON array")
|
||||
result = np.empty(len(values), dtype=np.int32)
|
||||
for index, value in enumerate(values):
|
||||
if isinstance(value, bool) or not isinstance(value, Real):
|
||||
raise ValueError(f"rosella_coefficients[{index}] is not a number")
|
||||
numeric = float(value)
|
||||
if not math.isfinite(numeric) or numeric != math.trunc(numeric):
|
||||
raise ValueError(
|
||||
f"rosella_coefficients[{index}] is not an exact integer: {value!r}")
|
||||
integer = int(value)
|
||||
if integer < -(1 << 31) or integer > (1 << 31) - 1:
|
||||
raise ValueError(
|
||||
f"rosella_coefficients[{index}] is outside signed int32: {integer}")
|
||||
result[index] = integer
|
||||
return result
|
||||
|
||||
|
||||
def _load_coefficients(path: Path) -> tuple[np.ndarray, str | None, str | None]:
|
||||
if not path.is_file():
|
||||
raise FileNotFoundError(path)
|
||||
source = path.read_bytes()
|
||||
stripped = source.lstrip()
|
||||
if stripped.startswith(b"{"):
|
||||
try:
|
||||
document = json.loads(source.decode("utf-8"))
|
||||
virtualizer = document["personalized_hrtf"]["virtualizer_parameters"]
|
||||
coefficients = _json_int32(virtualizer["rosella_coefficients"])
|
||||
except (UnicodeDecodeError, json.JSONDecodeError, KeyError, TypeError) as exc:
|
||||
raise ValueError(f"invalid personalized_headphone JSON: {exc}") from exc
|
||||
version = virtualizer.get("rosella_coefficients_version")
|
||||
room = virtualizer.get("room_model")
|
||||
else:
|
||||
if len(source) % 4:
|
||||
raise ValueError("raw rp payload must contain complete int32 lanes")
|
||||
coefficients = np.frombuffer(source, dtype="<i4").copy()
|
||||
version = room = None
|
||||
return coefficients, version, room
|
||||
|
||||
|
||||
def _unpack_field(serialized: np.ndarray, directions: int,
|
||||
exponent: int) -> np.ndarray:
|
||||
expected = 154 * directions
|
||||
if serialized.size != expected:
|
||||
raise ValueError(f"expected {expected} field lanes, got {serialized.size}")
|
||||
padded = np.zeros(160 * directions, dtype=np.float32)
|
||||
stride8 = 8 * directions
|
||||
stride2 = 2 * directions
|
||||
for source_index, value in enumerate(serialized):
|
||||
group4 = (source_index % stride8) // stride2
|
||||
destination = ((group4 & 3) + 4 * (
|
||||
source_index % stride2 +
|
||||
2 * directions * (source_index // stride8 + (group4 >> 2))))
|
||||
padded[destination] = _q15_exp(int(value), exponent)
|
||||
return padded
|
||||
|
||||
|
||||
def _unpack_table_a_grid(serialized: np.ndarray, dimension: int,
|
||||
serialized_rows: int, padded_rows: int,
|
||||
lane_group: int) -> np.ndarray:
|
||||
if serialized.size != serialized_rows * dimension:
|
||||
raise ValueError("unexpected table-A grid size")
|
||||
padded = np.zeros(padded_rows * dimension, dtype=np.float32)
|
||||
group_width = lane_group * 4
|
||||
for source_index, value in enumerate(serialized):
|
||||
remainder = source_index % group_width
|
||||
destination = ((remainder // lane_group) + 4 * (
|
||||
remainder % lane_group +
|
||||
group_width // 4 * (source_index // group_width)))
|
||||
padded[destination] = _q15(int(value))
|
||||
return padded
|
||||
|
||||
|
||||
def _unpack_table_a_extra(serialized: np.ndarray) -> np.ndarray:
|
||||
if serialized.size != 154:
|
||||
raise ValueError("table-A extra field must contain 154 serialized values")
|
||||
padded = np.zeros(160, dtype=np.float32)
|
||||
for source_index, value in enumerate(serialized):
|
||||
remainder = source_index & 7
|
||||
destination = ((remainder >> 1) + 4 * (
|
||||
(source_index & 1) + 2 * (source_index >> 3)))
|
||||
padded[destination] = _q15(int(value))
|
||||
return padded
|
||||
|
||||
|
||||
def _parse_profile(values: np.ndarray, position: int) -> tuple[DistanceProfile, int]:
|
||||
bounds = np.asarray([_q15(int(value)) for value in values[position:position + 6]],
|
||||
dtype=np.float32)
|
||||
position += 6
|
||||
distance = _q15_exp(int(values[position]), int(values[position + 1]))
|
||||
position += 2
|
||||
remaining = np.asarray(
|
||||
[_q15(int(value)) for value in values[position:position + 5]],
|
||||
dtype=np.float32)
|
||||
position += 5
|
||||
return DistanceProfile(
|
||||
bounds=bounds,
|
||||
distance_scale_m=distance,
|
||||
inverse_distance_per_m=remaining[0],
|
||||
axis_scales_internal=remaining[1:4],
|
||||
minimum_normalized_radius=remaining[4],
|
||||
), position
|
||||
|
||||
|
||||
def load_personalized_headphone(path: str | Path) -> RosellaModel:
|
||||
path = Path(path).resolve()
|
||||
coefficients, version, room = _load_coefficients(path)
|
||||
raw = coefficients.astype("<i4", copy=False).tobytes()
|
||||
header = inspect_rp(raw)
|
||||
if not header["checksum_valid"] or header["active_int32_lanes"] != coefficients.size:
|
||||
raise ValueError("invalid or non-active Rosella rp coefficient sequence")
|
||||
if not header["table_a_present"] or not header["table_b_present"] or header["table_c_present"]:
|
||||
raise NotImplementedError("current local renderer requires table A+B and no table C")
|
||||
if header["table_a_dimension"] != 64 or header["table_a_option"] != 3:
|
||||
raise NotImplementedError("current local renderer requires the observed 64-channel HQMF layout")
|
||||
if header["table_b_dimension"] != 20 or header["table_b_groups"] != 36:
|
||||
raise NotImplementedError("current local renderer requires 20 hybrid groups and 36 direction terms")
|
||||
|
||||
values = coefficients
|
||||
extra = header["table_a_extra"]
|
||||
table_a_main_start = 13
|
||||
position = table_a_main_start
|
||||
table_a_control = int(values[position]) & 0xFFFF
|
||||
field_exponent = int(values[position])
|
||||
position += 1
|
||||
option_count = header["table_a_option"]
|
||||
option_ids = (values[position:position + option_count].astype(np.int64) &
|
||||
0xFFFF).astype(np.int32)
|
||||
position += option_count
|
||||
option_values = np.asarray(
|
||||
[_q15(int(value)) for value in values[position:position + option_count]],
|
||||
dtype=np.float32)
|
||||
position += option_count
|
||||
table_a_scalar = _q15(int(values[position]))
|
||||
position += 1
|
||||
dimension = header["table_a_dimension"]
|
||||
table_a_filter_16x64 = _unpack_table_a_grid(
|
||||
values[position:position + 16 * dimension], dimension, 16, 20, 16)
|
||||
position += 16 * dimension
|
||||
table_a_four_integers = (values[position:position + 4].astype(np.int64) &
|
||||
0xFFFF).astype(np.int32)
|
||||
position += 4
|
||||
table_a_integer = int(values[position]) & 0xFFFF
|
||||
position += 1
|
||||
table_a_filter_8x64 = _unpack_table_a_grid(
|
||||
values[position:position + 8 * dimension], dimension, 8, 10, 8)
|
||||
position += 8 * dimension
|
||||
table_a_vector16 = np.asarray(
|
||||
[_q15(int(value)) for value in values[position:position + 16]],
|
||||
dtype=np.float32)
|
||||
position += 16
|
||||
table_a_filter_4x64 = _unpack_table_a_grid(
|
||||
values[position:position + 4 * dimension], dimension, 4, 5, 4)
|
||||
position += 4 * dimension
|
||||
extra_indices = (values[position:position + extra].astype(np.int64) &
|
||||
0xFFFF).astype(np.int32)
|
||||
position += extra
|
||||
extra_fields = np.empty((extra, 160), dtype=np.float32)
|
||||
for index in range(extra):
|
||||
extra_fields[index] = _unpack_table_a_extra(values[position:position + 154])
|
||||
position += 154
|
||||
extra_vectors = np.empty((extra, 16), dtype=np.float32)
|
||||
for index in range(extra):
|
||||
extra_vectors[index] = np.asarray(
|
||||
[_q15(int(value)) for value in values[position:position + 16]],
|
||||
dtype=np.float32)
|
||||
position += 16
|
||||
table_b_start = position
|
||||
expected_table_b_start = table_a_main_start + 1821 + 171 * extra
|
||||
if table_b_start != expected_table_b_start:
|
||||
raise AssertionError(
|
||||
f"table-A parser ended at {table_b_start}, expected {expected_table_b_start}")
|
||||
table_a_main = values[table_a_main_start:table_b_start].copy()
|
||||
|
||||
sample_rate = 2 * (int(values[position]) & 0xFFFF)
|
||||
position += 1
|
||||
matrix_exponent = int(values[position])
|
||||
position += 1
|
||||
matrix_count = 36 * 36
|
||||
scale_matrix = lambda block: np.asarray(
|
||||
[_q15_exp(int(value), matrix_exponent) for value in block],
|
||||
dtype=np.float32).reshape(36, 36)
|
||||
matrix_left = scale_matrix(values[position:position + matrix_count])
|
||||
position += matrix_count
|
||||
matrix_right = scale_matrix(values[position:position + matrix_count])
|
||||
position += matrix_count
|
||||
vector_left = np.asarray(
|
||||
[_q15_exp(int(value), matrix_exponent)
|
||||
for value in values[position:position + 36]], dtype=np.float32)
|
||||
position += 36
|
||||
vector_right = np.asarray(
|
||||
[_q15_exp(int(value), matrix_exponent)
|
||||
for value in values[position:position + 36]], dtype=np.float32)
|
||||
position += 36
|
||||
|
||||
serialized_count = 154 * 36
|
||||
field_left_serialized = values[position:position + serialized_count]
|
||||
field_left = _unpack_field(field_left_serialized, 36, field_exponent)
|
||||
field_left_odd_zero = not np.any(
|
||||
np.abs(np.asarray([_q15_exp(int(value), field_exponent)
|
||||
for value in field_left_serialized[1::2]],
|
||||
dtype=np.float32)) > np.float32(1e-6))
|
||||
position += serialized_count
|
||||
field_right = _unpack_field(
|
||||
values[position:position + serialized_count], 36, field_exponent)
|
||||
position += serialized_count
|
||||
|
||||
hybrid_flags = (values[position:position + 20].astype(np.int64) & 0xFFFF).astype(np.int32)
|
||||
position += 20
|
||||
active_hybrid_values = int(np.count_nonzero(hybrid_flags == 1))
|
||||
if active_hybrid_values != header["table_b_extra"]:
|
||||
raise ValueError(
|
||||
f"hybrid value count {active_hybrid_values} != header {header['table_b_extra']}")
|
||||
hybrid_values = np.asarray(
|
||||
[_q15(int(value)) for value in values[position:position + active_hybrid_values]],
|
||||
dtype=np.float32)
|
||||
position += active_hybrid_values
|
||||
model_scalars = np.asarray(
|
||||
[_q15(int(value)) for value in values[position:position + 5]],
|
||||
dtype=np.float32)
|
||||
position += 5
|
||||
|
||||
expected_table_a_tail = table_b_start + (
|
||||
header["table_b_dimension"] +
|
||||
380 * header["table_b_groups"] +
|
||||
header["table_b_extra"] + 79)
|
||||
if position != expected_table_a_tail:
|
||||
raise AssertionError(f"table-B parser ended at {position}, expected {expected_table_a_tail}")
|
||||
|
||||
header_float_scalars = np.asarray([
|
||||
_q15(int(values[position])),
|
||||
_f32(_q15(int(values[position + 1])) * _f32(16.0)),
|
||||
], dtype=np.float32)
|
||||
header_integer_fields = np.asarray([
|
||||
int(values[position + 2]),
|
||||
int(values[position + 3]) & 0xFFFF,
|
||||
], dtype=np.int32)
|
||||
position += 4
|
||||
profiles = []
|
||||
for _ in range(4):
|
||||
profile, position = _parse_profile(values, position)
|
||||
profiles.append(profile)
|
||||
profile_tail = np.asarray(
|
||||
[_q15(int(value)) for value in values[position:position + 8]],
|
||||
dtype=np.float32)
|
||||
position += 8
|
||||
post_fields = values[position:position + 3].astype(np.int32, copy=True)
|
||||
position += 3
|
||||
if position != values.size:
|
||||
raise AssertionError(f"unparsed coefficient lanes: {values.size - position}")
|
||||
|
||||
return RosellaModel(
|
||||
source_path=str(path),
|
||||
coefficients=coefficients,
|
||||
coefficient_sha256=hashlib.sha256(raw).hexdigest(),
|
||||
coefficient_version=version,
|
||||
room_model=room,
|
||||
table_a_dimension=header["table_a_dimension"],
|
||||
table_a_option=header["table_a_option"],
|
||||
table_a_extra=extra,
|
||||
table_a_header_field=int(values[8]) & 0xFFFF,
|
||||
table_a_header_25=int(values[9]) & 0xFFFF,
|
||||
table_a_control=table_a_control,
|
||||
table_a_option_ids=option_ids,
|
||||
table_a_option_values=option_values,
|
||||
table_a_scalar=table_a_scalar,
|
||||
table_a_filter_16x64_padded=table_a_filter_16x64,
|
||||
table_a_four_integers=table_a_four_integers,
|
||||
table_a_integer=table_a_integer,
|
||||
table_a_filter_8x64_padded=table_a_filter_8x64,
|
||||
table_a_vector16=table_a_vector16,
|
||||
table_a_filter_4x64_padded=table_a_filter_4x64,
|
||||
table_a_extra_indices=extra_indices,
|
||||
table_a_extra_fields_padded=extra_fields,
|
||||
table_a_extra_vectors=extra_vectors,
|
||||
sample_rate=sample_rate,
|
||||
matrix_exponent=matrix_exponent,
|
||||
field_exponent=field_exponent,
|
||||
matrix_left=matrix_left,
|
||||
matrix_right=matrix_right,
|
||||
vector_left=vector_left,
|
||||
vector_right=vector_right,
|
||||
field_left_padded=field_left,
|
||||
field_right_padded=field_right,
|
||||
field_left_odd_serialized_zero=field_left_odd_zero,
|
||||
hybrid_flags=hybrid_flags,
|
||||
hybrid_values=hybrid_values,
|
||||
model_scalars=model_scalars,
|
||||
header_float_scalars=header_float_scalars,
|
||||
header_integer_fields=header_integer_fields,
|
||||
profiles=tuple(profiles),
|
||||
profile_tail=profile_tail,
|
||||
post_fields=post_fields,
|
||||
table_a_main_serialized=table_a_main,
|
||||
)
|
||||
|
||||
|
||||
def direction_basis(x: float, y: float, z: float,
|
||||
dtype=np.float64) -> np.ndarray:
|
||||
"""Return the observed 36-term Rosella direction basis."""
|
||||
f = dtype
|
||||
x, y, z = f(x), f(y), f(z)
|
||||
out = np.empty(36, dtype=dtype)
|
||||
yz = f(y * z)
|
||||
x2 = f(x * x)
|
||||
y2 = f(y * y)
|
||||
x2m02 = f(x2 - f(0.2))
|
||||
xy = f(x * y)
|
||||
out[0:4] = (f(1.0), x, y, z)
|
||||
out[4] = f(x2 - f(1.0 / 3.0))
|
||||
out[5] = xy
|
||||
out[6] = f(x * z)
|
||||
out[7] = f(y2 - f(1.0 / 3.0))
|
||||
out[8] = yz
|
||||
out[9] = f(f(x2 - f(0.6)) * x)
|
||||
out[10] = f(x2m02 * y)
|
||||
out[11] = f(x2m02 * z)
|
||||
out[12] = f(f(y2 - f(0.2)) * x)
|
||||
out[13] = f(yz * x)
|
||||
out[14] = f(f(y2 - f(0.6)) * y)
|
||||
out[15] = f(f(y2 - f(0.2)) * z)
|
||||
out[16] = f(f(x2 * x2) - f(0.2))
|
||||
out[17] = f(xy * x2)
|
||||
out[18] = f(f(x * z) * x2)
|
||||
out[19] = f(f(y2 * x2) - f(1.0 / 15.0))
|
||||
out[20] = f(yz * x2)
|
||||
out[21] = f(x * y2 * y)
|
||||
out[22] = f(x * y2 * z)
|
||||
out[23] = f(f(y2 * y2) - f(0.2))
|
||||
x4 = f(x2 * x2)
|
||||
x2y2 = f(y2 * x2)
|
||||
y4 = f(y2 * y2)
|
||||
out[24] = f(yz * y2)
|
||||
out[25] = f(f(x4 - f(3.0 / 7.0)) * x)
|
||||
out[26] = f(f(x4 - f(3.0 / 35.0)) * y)
|
||||
out[27] = f(f(x4 - f(3.0 / 35.0)) * z)
|
||||
out[28] = f(f(x2y2 - f(3.0 / 35.0)) * x)
|
||||
out[29] = f(f(x2 * z) * xy)
|
||||
out[30] = f(f(x2y2 - f(3.0 / 35.0)) * y)
|
||||
out[31] = f(f(x2y2 - f(1.0 / 35.0)) * z)
|
||||
out[32] = f(f(y4 - f(3.0 / 35.0)) * x)
|
||||
out[33] = f(f(y2 * z) * xy)
|
||||
out[34] = f(f(y4 - f(3.0 / 7.0)) * y)
|
||||
out[35] = f(f(y4 - f(3.0 / 35.0)) * z)
|
||||
return out
|
||||
@@ -0,0 +1,198 @@
|
||||
"""Float64 Rosella table-A room model and overlap-add realization."""
|
||||
from __future__ import annotations
|
||||
|
||||
import numpy as np
|
||||
|
||||
from rosella_model import RosellaModel
|
||||
|
||||
|
||||
class _RosellaRoomState:
|
||||
"""Recursive table-A state used to generate the stable FIR realization."""
|
||||
|
||||
def __init__(self, model: RosellaModel):
|
||||
self.model = model
|
||||
self.bands = min(64, model.table_a_dimension)
|
||||
self.delays = model.table_a_four_integers.astype(np.int32)
|
||||
self.capacity = int(np.max(self.delays))
|
||||
self.matrix = np.asarray(model.table_a_vector16, dtype=np.float64).reshape(
|
||||
4, 4, order="F")
|
||||
f8 = np.asarray(model.table_a_filter_8x64_padded, dtype=np.float64).reshape(
|
||||
20, 4, 2, 4)
|
||||
f4 = np.asarray(model.table_a_filter_4x64_padded, dtype=np.float64).reshape(
|
||||
20, 4, 4)
|
||||
f16 = np.asarray(model.table_a_filter_16x64_padded, dtype=np.float64).reshape(
|
||||
20, 4, 4, 4)
|
||||
self.feedback_real = np.empty((self.bands, 4), dtype=np.float64)
|
||||
self.feedback_imag = np.empty_like(self.feedback_real)
|
||||
self.output_tap = np.empty_like(self.feedback_real)
|
||||
self.left_real = np.empty_like(self.feedback_real)
|
||||
self.left_imag = np.empty_like(self.feedback_real)
|
||||
self.right_real = np.empty_like(self.feedback_real)
|
||||
self.right_imag = np.empty_like(self.feedback_real)
|
||||
for band in range(self.bands):
|
||||
group, lane = divmod(band, 4)
|
||||
self.feedback_real[band] = f8[group, :, 0, lane]
|
||||
self.feedback_imag[band] = f8[group, :, 1, lane]
|
||||
self.output_tap[band] = f4[group, :, lane]
|
||||
self.left_real[band] = f16[group, :, 0, lane]
|
||||
self.left_imag[band] = f16[group, :, 1, lane]
|
||||
self.right_real[band] = f16[group, :, 2, lane]
|
||||
self.right_imag[band] = f16[group, :, 3, lane]
|
||||
|
||||
self.allpass_gain = np.asarray(model.table_a_option_values, dtype=np.float64)
|
||||
self.allpass_delay = model.table_a_option_ids.astype(np.int32)
|
||||
self.allpass_real = [
|
||||
np.zeros((int(delay), self.bands), dtype=np.float64)
|
||||
for delay in self.allpass_delay
|
||||
]
|
||||
self.allpass_imag = [np.zeros_like(value) for value in self.allpass_real]
|
||||
self.allpass_position = np.zeros(len(self.allpass_real), dtype=np.int32)
|
||||
self.memory_real = np.zeros(
|
||||
(self.capacity, self.bands, 4), dtype=np.float64)
|
||||
self.memory_imag = np.zeros_like(self.memory_real)
|
||||
self.position = 0
|
||||
self.extra_fields = np.asarray(
|
||||
model.table_a_extra_fields_padded, dtype=np.float64).reshape(-1, 20, 2, 4)
|
||||
self.extra_matrices = [
|
||||
np.asarray(value, dtype=np.float64).reshape(4, 4, order="F")
|
||||
for value in model.table_a_extra_vectors
|
||||
]
|
||||
|
||||
def reset(self):
|
||||
for value in self.allpass_real + self.allpass_imag:
|
||||
value.fill(0.0)
|
||||
self.allpass_position.fill(0)
|
||||
self.memory_real.fill(0.0)
|
||||
self.memory_imag.fill(0.0)
|
||||
self.position = 0
|
||||
|
||||
def process_slot(self, room_send) -> np.ndarray:
|
||||
values = np.asarray(room_send, dtype=np.complex128)
|
||||
input_real = values[:self.bands].real * 0.70710677
|
||||
input_imag = values[:self.bands].imag * 0.70710677
|
||||
if float(self.model.table_a_scalar) >= 0.5:
|
||||
raise NotImplementedError("alternate Rosella table-A room mode")
|
||||
|
||||
for index, gain in enumerate(self.allpass_gain):
|
||||
position = int(self.allpass_position[index])
|
||||
previous_real = self.allpass_real[index][position].copy()
|
||||
previous_imag = self.allpass_imag[index][position].copy()
|
||||
residual_real = input_real - previous_real * gain
|
||||
residual_imag = input_imag - previous_imag * gain
|
||||
input_real = residual_real * gain + previous_real
|
||||
input_imag = residual_imag * gain + previous_imag
|
||||
self.allpass_real[index][position] = residual_real
|
||||
self.allpass_imag[index][position] = residual_imag
|
||||
self.allpass_position[index] = (
|
||||
position + 1) % len(self.allpass_real[index])
|
||||
|
||||
branch_real = np.repeat(input_real[:, None], 4, axis=1)
|
||||
branch_imag = np.repeat(input_imag[:, None], 4, axis=1)
|
||||
delayed_real = np.empty_like(branch_real)
|
||||
delayed_imag = np.empty_like(branch_imag)
|
||||
for branch, delay in enumerate(self.delays):
|
||||
delayed_real[:, branch] = self.memory_real[
|
||||
(self.position - int(delay)) % self.capacity, :, branch]
|
||||
delayed_imag[:, branch] = self.memory_imag[
|
||||
(self.position - int(delay)) % self.capacity, :, branch]
|
||||
branch_real += np.einsum(
|
||||
"bj,ij->bi", delayed_real, self.matrix,
|
||||
dtype=np.float64, optimize=False)
|
||||
branch_imag += np.einsum(
|
||||
"bj,ij->bi", delayed_imag, self.matrix,
|
||||
dtype=np.float64, optimize=False)
|
||||
|
||||
tap_index = (self.position - self.model.table_a_integer) % self.capacity
|
||||
tap_real = self.memory_real[tap_index].copy()
|
||||
tap_imag = self.memory_imag[tap_index].copy()
|
||||
next_real = branch_real * self.feedback_real - branch_imag * self.feedback_imag
|
||||
next_imag = branch_imag * self.feedback_real + branch_real * self.feedback_imag
|
||||
self.memory_real[self.position] = next_real
|
||||
self.memory_imag[self.position] = next_imag
|
||||
self.position = (self.position + 1) % self.capacity
|
||||
|
||||
extra_real = np.zeros_like(branch_real)
|
||||
extra_imag = np.zeros_like(branch_imag)
|
||||
for index, delay in enumerate(self.model.table_a_extra_indices):
|
||||
source_real = self.memory_real[
|
||||
(self.position - (int(delay) + 1)) % self.capacity]
|
||||
source_imag = self.memory_imag[
|
||||
(self.position - (int(delay) + 1)) % self.capacity]
|
||||
matrix = self.extra_matrices[index]
|
||||
mixed_real = np.einsum(
|
||||
"bj,ij->bi", source_real, matrix,
|
||||
dtype=np.float64, optimize=False)
|
||||
mixed_imag = np.einsum(
|
||||
"bj,ij->bi", source_imag, matrix,
|
||||
dtype=np.float64, optimize=False)
|
||||
coefficient_real = np.empty(self.bands, dtype=np.float64)
|
||||
coefficient_imag = np.empty(self.bands, dtype=np.float64)
|
||||
for band in range(self.bands):
|
||||
group, lane = divmod(band, 4)
|
||||
coefficient_real[band] = self.extra_fields[index, group, 0, lane]
|
||||
coefficient_imag[band] = self.extra_fields[index, group, 1, lane]
|
||||
extra_real += (mixed_real * coefficient_real[:, None]
|
||||
- mixed_imag * coefficient_imag[:, None])
|
||||
extra_imag += (mixed_imag * coefficient_real[:, None]
|
||||
+ mixed_real * coefficient_imag[:, None])
|
||||
|
||||
output_real = tap_real * self.output_tap + extra_real
|
||||
output_imag = tap_imag * self.output_tap + extra_imag
|
||||
left = np.sum(
|
||||
self.left_real * output_real - self.left_imag * output_imag,
|
||||
axis=1, dtype=np.float64)
|
||||
left_imag = np.sum(
|
||||
self.left_imag * output_real + self.left_real * output_imag,
|
||||
axis=1, dtype=np.float64)
|
||||
right = np.sum(
|
||||
self.right_real * output_real - self.right_imag * output_imag,
|
||||
axis=1, dtype=np.float64)
|
||||
right_imag = np.sum(
|
||||
self.right_imag * output_real + self.right_real * output_imag,
|
||||
axis=1, dtype=np.float64)
|
||||
result = np.zeros((2, 77), dtype=np.complex128)
|
||||
result[0, :self.bands] = left + 1j * left_imag
|
||||
result[1, :self.bands] = right + 1j * right_imag
|
||||
return result
|
||||
|
||||
|
||||
class RosellaRoomFir:
|
||||
"""Complex128 overlap-add room FIR generated locally from table-A."""
|
||||
|
||||
def __init__(self, model: RosellaModel, impulse_slots: int = 4096):
|
||||
if impulse_slots <= 0:
|
||||
raise ValueError("impulse_slots must be positive")
|
||||
reference = _RosellaRoomState(model)
|
||||
self.length = int(impulse_slots)
|
||||
self.kernel = np.empty((self.length, 2, 64), dtype=np.complex128)
|
||||
for slot in range(self.length):
|
||||
impulse = np.zeros(77, dtype=np.complex128)
|
||||
if slot == 0:
|
||||
impulse[:64] = 1.0
|
||||
self.kernel[slot] = reference.process_slot(impulse)[:, :64]
|
||||
self.tail = np.zeros((self.length - 1, 2, 64), dtype=np.complex128)
|
||||
self._fft_cache: dict[int, np.ndarray] = {}
|
||||
|
||||
def reset(self):
|
||||
self.tail.fill(0.0)
|
||||
|
||||
def process_chunk(self, room_send) -> np.ndarray:
|
||||
values = np.asarray(room_send, dtype=np.complex128)
|
||||
if values.ndim != 2 or values.shape[1] != 77:
|
||||
raise ValueError("room_send must have shape [slots,77]")
|
||||
count = len(values)
|
||||
if count == 0:
|
||||
return np.zeros((0, 2, 77), dtype=np.complex128)
|
||||
needed = count + self.length - 1
|
||||
fft_size = 1 << (needed - 1).bit_length()
|
||||
kernel_fft = self._fft_cache.get(fft_size)
|
||||
if kernel_fft is None:
|
||||
kernel_fft = np.fft.fft(self.kernel, fft_size, axis=0)
|
||||
self._fft_cache[fft_size] = kernel_fft
|
||||
input_fft = np.fft.fft(values[:, :64], fft_size, axis=0)
|
||||
block = np.fft.ifft(input_fft[:, None, :] * kernel_fft, axis=0)[:needed]
|
||||
block[:len(self.tail)] += self.tail
|
||||
result = np.zeros((count, 2, 77), dtype=np.complex128)
|
||||
result[:, :, :64] = block[:count]
|
||||
self.tail = block[count:count + self.length - 1].copy()
|
||||
return result
|
||||
@@ -0,0 +1,552 @@
|
||||
"""Stateful public SOFA binaural renderer.
|
||||
|
||||
The runtime topology mirrors the existing multi-object binaural path:
|
||||
64-QMF -> 77 hybrid -> per-object directional transfer -> stereo synthesis.
|
||||
The HRTF parameter source is a SOFA-derived fifth-order field. Early
|
||||
reflections and the late room use project-owned behavior.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
import math
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
from public_filterbank import (
|
||||
ANALYSIS_SYNTHESIS_LATENCY_SAMPLES,
|
||||
HYBRID_BANDS,
|
||||
QMF_HOP,
|
||||
PublicAnalysis77,
|
||||
PublicSynthesis77,
|
||||
table_info as filterbank_table_info,
|
||||
)
|
||||
from public_room import (
|
||||
LateFdnConfig,
|
||||
SharedUnitaryFdn,
|
||||
ShoeboxRoomConfig,
|
||||
first_order_image_sources,
|
||||
)
|
||||
from reference_distance import DistanceState, ReferenceDistanceProfileV1
|
||||
from sofa_canonical import CanonicalHrtf
|
||||
from sofa_hrtf_field import (
|
||||
DEFAULT_ORDER,
|
||||
DEFAULT_PROJECTION_RIDGE,
|
||||
DEFAULT_SH_RIDGE,
|
||||
SofaHrtfField,
|
||||
compile_sofa_hrtf,
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class HybridPath:
|
||||
label: str
|
||||
delay_slots: np.ndarray # whole-QMF delay per ear, [2]
|
||||
transfer: np.ndarray # [ear,77], includes residual delay and HRTF delay
|
||||
|
||||
def __post_init__(self):
|
||||
slots = np.asarray(self.delay_slots)
|
||||
if slots.shape == ():
|
||||
slots = np.repeat(slots, 2)
|
||||
if slots.shape != (2,) or slots.dtype.kind not in "iu":
|
||||
raise ValueError("hybrid path delay_slots must contain two integers")
|
||||
slots = np.asarray(slots, dtype=np.int64)
|
||||
transfer = np.asarray(self.transfer, dtype=np.complex128)
|
||||
if transfer.shape != (2, HYBRID_BANDS) or not np.isfinite(transfer).all():
|
||||
raise ValueError("hybrid path transfer must have finite shape [2,77]")
|
||||
if np.any(slots < 0):
|
||||
raise ValueError("hybrid path delay_slots must be non-negative")
|
||||
slots.setflags(write=False)
|
||||
transfer.setflags(write=False)
|
||||
object.__setattr__(self, "delay_slots", slots)
|
||||
object.__setattr__(self, "transfer", transfer)
|
||||
|
||||
|
||||
class HybridObjectPathRenderer:
|
||||
"""Per-object hybrid histories for direct and image-source paths."""
|
||||
|
||||
def __init__(self, source_count: int, *, history_slots: int = 256,
|
||||
transition_slots: int = 8):
|
||||
self.source_count = int(source_count)
|
||||
self.history_slots = int(history_slots)
|
||||
self.transition_slots = int(transition_slots)
|
||||
if min(self.source_count, self.history_slots) <= 0 or self.transition_slots < 0:
|
||||
raise ValueError("invalid hybrid path renderer dimensions")
|
||||
self.history = np.zeros(
|
||||
(self.source_count, self.history_slots, HYBRID_BANDS), dtype=np.complex128)
|
||||
self.position = 0
|
||||
self.current: list[tuple[HybridPath, ...]] = [tuple() for _ in range(self.source_count)]
|
||||
self.target: list[tuple[HybridPath, ...] | None] = [None] * self.source_count
|
||||
self.fade_position = np.zeros(self.source_count, dtype=np.int32)
|
||||
self.fade_total = np.zeros(self.source_count, dtype=np.int32)
|
||||
self.processed_slots = 0
|
||||
|
||||
def reset(self) -> None:
|
||||
self.history.fill(0.0)
|
||||
self.position = 0
|
||||
self.target = [None] * self.source_count
|
||||
self.fade_position.fill(0)
|
||||
self.fade_total.fill(0)
|
||||
self.processed_slots = 0
|
||||
|
||||
def set_paths(self, source: int, paths, *, fade_slots: int | None = None) -> None:
|
||||
source = int(source)
|
||||
if not 0 <= source < self.source_count:
|
||||
raise IndexError(source)
|
||||
values = tuple(paths)
|
||||
for path in values:
|
||||
if np.any(path.delay_slots >= self.history_slots):
|
||||
raise ValueError(
|
||||
f"path {path.label!r} needs {path.delay_slots.tolist()} slots, "
|
||||
f"history capacity is {self.history_slots}")
|
||||
fade = self.transition_slots if fade_slots is None else int(fade_slots)
|
||||
if fade < 0:
|
||||
raise ValueError("path fade must be non-negative")
|
||||
if self.target[source] is not None:
|
||||
# Normal 512-sample updates complete an 8-slot transition exactly.
|
||||
# If a caller updates faster, use the previous target as the new
|
||||
# stable side rather than resetting signal history.
|
||||
self.current[source] = self.target[source]
|
||||
self.target[source] = None
|
||||
if not self.processed_slots or fade == 0:
|
||||
self.current[source] = values
|
||||
self.target[source] = None
|
||||
self.fade_position[source] = 0
|
||||
self.fade_total[source] = 0
|
||||
else:
|
||||
self.target[source] = values
|
||||
self.fade_position[source] = 0
|
||||
self.fade_total[source] = fade
|
||||
|
||||
def _render_paths(self, source: int, paths: tuple[HybridPath, ...]) -> np.ndarray:
|
||||
result = np.zeros((2, HYBRID_BANDS), dtype=np.complex128)
|
||||
for path in paths:
|
||||
indices = (self.position - path.delay_slots) % self.history_slots
|
||||
delayed = self.history[source, indices, :]
|
||||
result += delayed * path.transfer
|
||||
return result
|
||||
|
||||
def process(self, hybrid) -> np.ndarray:
|
||||
values = np.asarray(hybrid, dtype=np.complex128)
|
||||
if values.ndim != 3 or values.shape[1:] != (self.source_count, HYBRID_BANDS):
|
||||
raise ValueError(
|
||||
f"hybrid input must have shape [slots,{self.source_count},77]")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("hybrid input contains non-finite values")
|
||||
output = np.zeros((len(values), 2, HYBRID_BANDS), dtype=np.complex128)
|
||||
for slot in range(len(values)):
|
||||
self.history[:, self.position, :] = values[slot]
|
||||
for source in range(self.source_count - 1, -1, -1):
|
||||
current = self._render_paths(source, self.current[source])
|
||||
target_paths = self.target[source]
|
||||
if target_paths is None:
|
||||
output[slot] += current
|
||||
continue
|
||||
target = self._render_paths(source, target_paths)
|
||||
self.fade_position[source] += 1
|
||||
amount = min(
|
||||
1.0, self.fade_position[source] / float(self.fade_total[source]))
|
||||
output[slot] += current * (1.0 - amount) + target * amount
|
||||
if self.fade_position[source] >= self.fade_total[source]:
|
||||
self.current[source] = target_paths
|
||||
self.target[source] = None
|
||||
self.fade_position[source] = 0
|
||||
self.fade_total[source] = 0
|
||||
self.position = (self.position + 1) % self.history_slots
|
||||
self.processed_slots += 1
|
||||
return output
|
||||
|
||||
|
||||
class _StereoDelay:
|
||||
def __init__(self, delay_samples: int):
|
||||
self.delay_samples = int(delay_samples)
|
||||
if self.delay_samples < 0:
|
||||
raise ValueError("delay must be non-negative")
|
||||
self.state = np.zeros((self.delay_samples, 2), dtype=np.float64)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.state.fill(0.0)
|
||||
|
||||
def process(self, values) -> np.ndarray:
|
||||
source = np.asarray(values, dtype=np.float64)
|
||||
if source.ndim != 2 or source.shape[1] != 2:
|
||||
raise ValueError("stereo delay input must have shape [samples,2]")
|
||||
if self.delay_samples == 0:
|
||||
return source.copy()
|
||||
joined = np.concatenate((self.state, source), axis=0)
|
||||
output = joined[:len(source)].copy()
|
||||
self.state = joined[len(source):len(source) + self.delay_samples].copy()
|
||||
return output
|
||||
|
||||
|
||||
class SofaBinauralBackend:
|
||||
"""SOFA-derived public 77-band/SH renderer with public room processing."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
field: SofaHrtfField,
|
||||
*,
|
||||
source_count: int = 16,
|
||||
sample_rate_hz: float = 48000.0,
|
||||
default_profile: str = "mid",
|
||||
enable_early_reflections: bool = True,
|
||||
enable_late_room: bool = True,
|
||||
room_config: ShoeboxRoomConfig = ShoeboxRoomConfig(),
|
||||
fdn_config: LateFdnConfig | None = None,
|
||||
transition_slots: int = 8,
|
||||
history_slots: int = 256,
|
||||
output_gain: float = 1.0):
|
||||
self.source_count = int(source_count)
|
||||
self.sample_rate_hz = float(sample_rate_hz)
|
||||
self.default_profile = ReferenceDistanceProfileV1.validate_profile(default_profile)
|
||||
self.enable_early_reflections = bool(enable_early_reflections)
|
||||
self.enable_late_room = bool(enable_late_room)
|
||||
self.room_config = room_config
|
||||
self.room_config.validate()
|
||||
self.output_gain = float(output_gain)
|
||||
if (self.source_count <= 0 or not math.isfinite(self.sample_rate_hz)
|
||||
or self.sample_rate_hz <= 0.0):
|
||||
raise ValueError("source_count and sample rate must be positive")
|
||||
if not math.isfinite(self.output_gain):
|
||||
raise ValueError("output gain must be finite")
|
||||
if abs(self.sample_rate_hz - 48000.0) > 1.0e-9:
|
||||
raise ValueError("the public binaural runtime requires 48 kHz")
|
||||
|
||||
if not isinstance(field, SofaHrtfField):
|
||||
raise TypeError(
|
||||
"field must be SofaHrtfField; use from_sofa() or "
|
||||
"from_compiled_cache() for file inputs")
|
||||
self.field = field
|
||||
self.hrtf_input_kind = "field"
|
||||
self.hrtf_input_path: str | None = None
|
||||
self.cache_policy: str | None = None
|
||||
if abs(self.field.sample_rate_hz - self.sample_rate_hz) > 1.0e-9:
|
||||
raise ValueError("HRTF field sample rate does not match the renderer")
|
||||
|
||||
self.early_history_slots = int(history_slots)
|
||||
if self.early_history_slots <= 0:
|
||||
raise ValueError("history_slots must be positive")
|
||||
maximum_hrtf_delay = float(np.max(self.field.delay_bounds[:, 1], initial=0.0))
|
||||
self.maximum_hrtf_delay_samples = maximum_hrtf_delay
|
||||
self.hrtf_history_slots = int(math.ceil(maximum_hrtf_delay / QMF_HOP))
|
||||
self.analysis = PublicAnalysis77(self.source_count)
|
||||
self.paths = HybridObjectPathRenderer(
|
||||
self.source_count,
|
||||
history_slots=self.early_history_slots + self.hrtf_history_slots,
|
||||
transition_slots=transition_slots)
|
||||
self.synthesis = PublicSynthesis77(2)
|
||||
actual_fdn_config = fdn_config or LateFdnConfig(sample_rate_hz=self.sample_rate_hz)
|
||||
if abs(actual_fdn_config.sample_rate_hz - self.sample_rate_hz) > 1.0e-9:
|
||||
raise ValueError("FDN sample rate does not match the renderer")
|
||||
self.fdn = SharedUnitaryFdn(actual_fdn_config)
|
||||
self.late_delay = _StereoDelay(ANALYSIS_SYNTHESIS_LATENCY_SAMPLES)
|
||||
|
||||
self.positions = np.zeros((self.source_count, 3), dtype=np.float64)
|
||||
self.positions[:, 1] = 1.0
|
||||
self.profiles = [self.default_profile] * self.source_count
|
||||
self.user_gain = np.ones(self.source_count, dtype=np.float64)
|
||||
self.special_lfe = np.zeros(self.source_count, dtype=bool)
|
||||
self.distance_state: list[DistanceState | None] = [None] * self.source_count
|
||||
self.late_current = np.zeros(self.source_count, dtype=np.float64)
|
||||
self.late_start = np.zeros(self.source_count, dtype=np.float64)
|
||||
self.late_target = np.zeros(self.source_count, dtype=np.float64)
|
||||
self.late_fade_position = np.zeros(self.source_count, dtype=np.int64)
|
||||
self.late_fade_total = np.zeros(self.source_count, dtype=np.int64)
|
||||
self.maximum_early_delay_samples = 0.0
|
||||
self.latency_to_discard = ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
self.processed_input_samples = 0
|
||||
self.output_samples = 0
|
||||
self.parameter_updates = 0
|
||||
self.finished = False
|
||||
for source in range(self.source_count):
|
||||
self.set_source(
|
||||
source, self.positions[source], profile=self.default_profile,
|
||||
fade=False)
|
||||
|
||||
@classmethod
|
||||
def from_sofa(
|
||||
cls, sofa: str | Path | CanonicalHrtf, *,
|
||||
cache_policy: str = "memory",
|
||||
cache_dir: str | Path | None = None,
|
||||
shell_radius_m: float = 1.0,
|
||||
order: int = DEFAULT_ORDER,
|
||||
projection_ridge: float = DEFAULT_PROJECTION_RIDGE,
|
||||
sh_ridge: float = DEFAULT_SH_RIDGE,
|
||||
**renderer_options) -> "SofaBinauralBackend":
|
||||
"""Compile a SOFA source once and construct the runtime renderer."""
|
||||
field = compile_sofa_hrtf(
|
||||
sofa,
|
||||
target_sample_rate_hz=float(
|
||||
renderer_options.get("sample_rate_hz", 48000.0)),
|
||||
shell_radius_m=shell_radius_m,
|
||||
order=order,
|
||||
projection_ridge=projection_ridge,
|
||||
sh_ridge=sh_ridge,
|
||||
cache_policy=cache_policy,
|
||||
cache_dir=cache_dir)
|
||||
result = cls(field, **renderer_options)
|
||||
result.hrtf_input_kind = "sofa"
|
||||
result.hrtf_input_path = (
|
||||
str(Path(sofa).expanduser().resolve())
|
||||
if not isinstance(sofa, CanonicalHrtf) else sofa.source_path)
|
||||
result.cache_policy = str(cache_policy).lower()
|
||||
return result
|
||||
|
||||
@classmethod
|
||||
def from_compiled_cache(
|
||||
cls, cache: str | Path, **renderer_options
|
||||
) -> "SofaBinauralBackend":
|
||||
"""Load an explicitly selected validated JOC compiled HRTF cache."""
|
||||
path = Path(cache).expanduser().resolve()
|
||||
field = SofaHrtfField.load(path)
|
||||
result = cls(field, **renderer_options)
|
||||
result.hrtf_input_kind = "compiled_cache"
|
||||
result.hrtf_input_path = str(path)
|
||||
result.cache_policy = None
|
||||
return result
|
||||
|
||||
def _make_path(self, label: str, direction_adm, path_distance_m: float,
|
||||
extra_delay_samples: float, amplitude: float) -> HybridPath:
|
||||
del path_distance_m
|
||||
evaluation = self.field.evaluate_adm(direction_adm)
|
||||
extra_delay = float(extra_delay_samples)
|
||||
if not math.isfinite(extra_delay) or extra_delay < 0.0:
|
||||
raise ValueError("path delay must be finite and non-negative")
|
||||
early_delay_slots = int(math.floor(extra_delay / QMF_HOP))
|
||||
if early_delay_slots >= self.early_history_slots:
|
||||
raise ValueError(
|
||||
f"path {label!r} needs {early_delay_slots} early-delay slots, "
|
||||
f"early history capacity is {self.early_history_slots}")
|
||||
total_delay = np.asarray(evaluation.delay_samples, dtype=np.float64) + extra_delay
|
||||
delay_slots = np.floor(total_delay / QMF_HOP).astype(np.int64)
|
||||
residual = total_delay - delay_slots * QMF_HOP
|
||||
propagation_phase = np.exp(
|
||||
-2j * np.pi * self.field.band_center_frequencies_hz[None, :]
|
||||
* residual[:, None]
|
||||
/ self.sample_rate_hz)
|
||||
transfer = np.asarray(
|
||||
evaluation.aligned_gains * propagation_phase * float(amplitude),
|
||||
dtype=np.complex128)
|
||||
return HybridPath(label, delay_slots, transfer)
|
||||
|
||||
def _ordinary_paths(self, state: DistanceState, gain: float) -> tuple[HybridPath, ...]:
|
||||
# Object PCM is programme-normalized. Physical distance controls room
|
||||
# geometry, while the project profile supplies the direct presentation
|
||||
# coefficient instead of applying a second free-field 1/r attenuation.
|
||||
direct_amplitude = gain * ReferenceDistanceProfileV1.direct_level_gain(state)
|
||||
room_gain = ReferenceDistanceProfileV1.room_calibration_gain(state)
|
||||
paths = [self._make_path(
|
||||
"direct", state.direction_adm, state.physical_distance_m, 0.0,
|
||||
direct_amplitude)]
|
||||
if self.enable_early_reflections:
|
||||
reflections = first_order_image_sources(
|
||||
state.direction_adm, state.physical_distance_m,
|
||||
self.sample_rate_hz, self.room_config)
|
||||
for reflection in reflections:
|
||||
amplitude = (
|
||||
gain * room_gain * reflection.reflection_gain
|
||||
* ReferenceDistanceProfileV1.inverse_distance_gain(
|
||||
self.field.measurement_radius_m,
|
||||
reflection.path_distance_m))
|
||||
paths.append(self._make_path(
|
||||
f"early:{reflection.wall}", reflection.direction_adm,
|
||||
reflection.path_distance_m, reflection.extra_delay_samples,
|
||||
amplitude))
|
||||
self.maximum_early_delay_samples = max(
|
||||
self.maximum_early_delay_samples,
|
||||
reflection.extra_delay_samples)
|
||||
return tuple(paths)
|
||||
|
||||
def _lfe_paths(self, gain: float) -> tuple[HybridPath, ...]:
|
||||
frequency = self.field.band_center_frequencies_hz
|
||||
lowpass = np.ones(HYBRID_BANDS, dtype=np.float64)
|
||||
lowpass[frequency >= 180.0] = 0.0
|
||||
transition = (frequency > 120.0) & (frequency < 180.0)
|
||||
amount = (frequency[transition] - 120.0) / 60.0
|
||||
lowpass[transition] = np.cos(0.5 * np.pi * amount) ** 2
|
||||
transfer = np.repeat(
|
||||
(gain * lowpass / math.sqrt(2.0))[None, :], 2, axis=0
|
||||
).astype(np.complex128)
|
||||
return (HybridPath(
|
||||
"public_lfe_lowpass", np.zeros(2, dtype=np.int64), transfer),)
|
||||
|
||||
def _set_late_target(self, source: int, value: float, fade: bool) -> None:
|
||||
value = float(value)
|
||||
fade_samples = (self.paths.transition_slots * QMF_HOP
|
||||
if fade and self.processed_input_samples else 0)
|
||||
if fade_samples == 0:
|
||||
self.late_current[source] = value
|
||||
self.late_start[source] = value
|
||||
self.late_target[source] = value
|
||||
self.late_fade_position[source] = 0
|
||||
self.late_fade_total[source] = 0
|
||||
else:
|
||||
self.late_start[source] = self.late_current[source]
|
||||
self.late_target[source] = value
|
||||
self.late_fade_position[source] = 0
|
||||
self.late_fade_total[source] = fade_samples
|
||||
|
||||
def set_source(self, source: int, position_adm, *, profile: str | None = None,
|
||||
gain: float = 1.0, enabled: bool = True,
|
||||
special_lfe: bool = False, fade: bool = True) -> None:
|
||||
if self.finished:
|
||||
raise RuntimeError("SOFA renderer is finished")
|
||||
source = int(source)
|
||||
if not 0 <= source < self.source_count:
|
||||
raise IndexError(source)
|
||||
gain = float(gain)
|
||||
if not math.isfinite(gain):
|
||||
raise ValueError("source gain must be finite")
|
||||
effective_gain = gain if enabled else 0.0
|
||||
name = self.default_profile if profile is None else profile
|
||||
state = ReferenceDistanceProfileV1.map_adm_position(position_adm, name)
|
||||
path_set = (self._lfe_paths(effective_gain) if special_lfe
|
||||
else self._ordinary_paths(state, effective_gain))
|
||||
self.paths.set_paths(
|
||||
source, path_set,
|
||||
fade_slots=(self.paths.transition_slots if fade else 0))
|
||||
late_send = (0.0 if special_lfe or not self.enable_late_room or not enabled
|
||||
else effective_gain
|
||||
* ReferenceDistanceProfileV1.room_calibration_gain(state)
|
||||
* ReferenceDistanceProfileV1.late_send(state))
|
||||
self._set_late_target(source, late_send, fade)
|
||||
self.positions[source] = np.asarray(position_adm, dtype=np.float64)
|
||||
self.profiles[source] = state.profile
|
||||
self.user_gain[source] = gain
|
||||
self.special_lfe[source] = bool(special_lfe)
|
||||
self.distance_state[source] = state
|
||||
self.parameter_updates += 1
|
||||
|
||||
def _late_send_envelope(self, sample_count: int) -> np.ndarray:
|
||||
envelope = np.empty((sample_count, self.source_count), dtype=np.float64)
|
||||
for source in range(self.source_count):
|
||||
total = int(self.late_fade_total[source])
|
||||
if total == 0:
|
||||
envelope[:, source] = self.late_current[source]
|
||||
continue
|
||||
start_position = int(self.late_fade_position[source])
|
||||
position = start_position + np.arange(1, sample_count + 1)
|
||||
amount = np.clip(position / float(total), 0.0, 1.0)
|
||||
envelope[:, source] = (
|
||||
self.late_start[source] * (1.0 - amount)
|
||||
+ self.late_target[source] * amount)
|
||||
new_position = start_position + sample_count
|
||||
if new_position >= total:
|
||||
self.late_current[source] = self.late_target[source]
|
||||
self.late_start[source] = self.late_target[source]
|
||||
self.late_fade_position[source] = 0
|
||||
self.late_fade_total[source] = 0
|
||||
else:
|
||||
self.late_current[source] = float(envelope[-1, source])
|
||||
self.late_fade_position[source] = new_position
|
||||
return envelope
|
||||
|
||||
def _process(self, sources) -> np.ndarray:
|
||||
values = np.asarray(sources, dtype=np.float64)
|
||||
if values.ndim != 2 or values.shape[1] != self.source_count:
|
||||
raise ValueError(f"sources must have shape [samples,{self.source_count}]")
|
||||
if len(values) % QMF_HOP:
|
||||
raise ValueError("SOFA backend input must be divisible by 64 samples")
|
||||
if not np.isfinite(values).all():
|
||||
raise ValueError("SOFA backend input contains non-finite values")
|
||||
hybrid = self.analysis.process(values)
|
||||
direct_and_early = self.paths.process(hybrid)
|
||||
direct_pcm = self.synthesis.process(direct_and_early)
|
||||
if self.enable_late_room:
|
||||
sends = self._late_send_envelope(len(values))
|
||||
mono = np.sum(values * sends, axis=1, dtype=np.float64)
|
||||
late_pcm = self.late_delay.process(self.fdn.process(mono))
|
||||
else:
|
||||
# Still advance any pending send fade deterministically.
|
||||
self._late_send_envelope(len(values))
|
||||
late_pcm = np.zeros_like(direct_pcm)
|
||||
mixed = np.asarray((direct_pcm + late_pcm) * self.output_gain, dtype=np.float64)
|
||||
skip = min(self.latency_to_discard, len(mixed))
|
||||
self.latency_to_discard -= skip
|
||||
self.processed_input_samples += len(values)
|
||||
output = mixed[skip:]
|
||||
self.output_samples += len(output)
|
||||
return output
|
||||
|
||||
def process(self, sources) -> np.ndarray:
|
||||
if self.finished:
|
||||
raise RuntimeError("SOFA renderer is finished")
|
||||
return self._process(sources)
|
||||
|
||||
def finish(self, *, tail_seconds: float | None = None) -> np.ndarray:
|
||||
if self.finished:
|
||||
return np.zeros((0, 2), dtype=np.float64)
|
||||
if tail_seconds is not None and (
|
||||
not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0):
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
requested = (self.fdn.tail_samples if tail_seconds is None
|
||||
else int(math.ceil(float(tail_seconds) * self.sample_rate_hz)))
|
||||
drain = max(
|
||||
requested if self.enable_late_room else 0,
|
||||
int(math.ceil(
|
||||
self.maximum_hrtf_delay_samples
|
||||
+ self.maximum_early_delay_samples)) + 2048,
|
||||
) + ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
drain = int(math.ceil(drain / QMF_HOP) * QMF_HOP)
|
||||
output = self._process(np.zeros((drain, self.source_count), dtype=np.float64))
|
||||
self.finished = True
|
||||
return output
|
||||
|
||||
def finish_output_capacity(self, tail_seconds: float | None = None) -> int:
|
||||
"""Return a conservative bound for one future :meth:`finish` output."""
|
||||
if tail_seconds is not None and (
|
||||
not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0):
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
requested = (self.fdn.tail_samples if tail_seconds is None
|
||||
else int(math.ceil(float(tail_seconds) * self.sample_rate_hz)))
|
||||
hrtf_bound = self.hrtf_history_slots * QMF_HOP
|
||||
early_bound = hrtf_bound + 2048
|
||||
if self.enable_early_reflections:
|
||||
early_bound += self.early_history_slots * QMF_HOP
|
||||
drain = max(requested if self.enable_late_room else 0, early_bound)
|
||||
drain += ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
return int(math.ceil(drain / QMF_HOP) * QMF_HOP)
|
||||
|
||||
def reset(self) -> None:
|
||||
self.analysis.reset()
|
||||
self.paths.reset()
|
||||
self.synthesis.reset()
|
||||
self.fdn.reset()
|
||||
self.late_delay.reset()
|
||||
self.latency_to_discard = ANALYSIS_SYNTHESIS_LATENCY_SAMPLES
|
||||
self.processed_input_samples = 0
|
||||
self.output_samples = 0
|
||||
self.finished = False
|
||||
for source in range(self.source_count):
|
||||
self.set_source(
|
||||
source, self.positions[source], profile=self.profiles[source],
|
||||
gain=float(self.user_gain[source]),
|
||||
special_lfe=bool(self.special_lfe[source]), fade=False)
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"name": "SofaBinauralBackend",
|
||||
"source_count": self.source_count,
|
||||
"sample_rate_hz": self.sample_rate_hz,
|
||||
"precision": "float64/complex128",
|
||||
"signal_path": (
|
||||
"public 64-QMF -> public 77-hybrid -> SOFA order-5 real-SH "
|
||||
"direct/early -> public synthesis + shared unitary FDN"),
|
||||
"hrtf_input_kind": self.hrtf_input_kind,
|
||||
"hrtf_input_path": self.hrtf_input_path,
|
||||
"cache_policy": self.cache_policy,
|
||||
"latency_compensated_samples": ANALYSIS_SYNTHESIS_LATENCY_SAMPLES,
|
||||
"enable_early_reflections": self.enable_early_reflections,
|
||||
"enable_late_room": self.enable_late_room,
|
||||
"early_history_slots": self.early_history_slots,
|
||||
"hrtf_history_slots": self.hrtf_history_slots,
|
||||
"maximum_hrtf_delay_samples": self.maximum_hrtf_delay_samples,
|
||||
"maximum_early_delay_samples": self.maximum_early_delay_samples,
|
||||
"parameter_updates": self.parameter_updates,
|
||||
"processed_input_samples_including_flush": self.processed_input_samples,
|
||||
"output_samples_before_trim": self.output_samples,
|
||||
"distance": ReferenceDistanceProfileV1.info(),
|
||||
"filterbank": filterbank_table_info(),
|
||||
"field": self.field.info(),
|
||||
"late_room": self.fdn.info(),
|
||||
}
|
||||
@@ -0,0 +1,637 @@
|
||||
"""Strict SimpleFreeFieldHRIR to canonical HRTF import.
|
||||
|
||||
The canonical representation keeps ``Data.IR`` and ``Data.Delay`` separate.
|
||||
No importer operation silently bakes the SOFA delay into the stored FIRs. A
|
||||
caller must explicitly request :meth:`CanonicalHrtf.materialized_measurement`
|
||||
when a time-domain FIR with ``Data.Delay`` applied exactly once is required.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from contextlib import contextmanager
|
||||
from dataclasses import dataclass, replace
|
||||
from fractions import Fraction
|
||||
import hashlib
|
||||
import math
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import h5py
|
||||
import numpy as np
|
||||
from scipy import signal
|
||||
from scipy.fft import next_fast_len
|
||||
|
||||
|
||||
_SUPPORTED_VERSIONS = {"0.4", "1.0", "1.1"}
|
||||
_FREE_FIELD_ROOM_TYPES = {"free field", "free-field", "anechoic", "hemi-anechoic"}
|
||||
_LENGTH_UNITS = {
|
||||
"m": 1.0,
|
||||
"metre": 1.0,
|
||||
"metres": 1.0,
|
||||
"meter": 1.0,
|
||||
"meters": 1.0,
|
||||
"cm": 1.0e-2,
|
||||
"centimetre": 1.0e-2,
|
||||
"centimetres": 1.0e-2,
|
||||
"centimeter": 1.0e-2,
|
||||
"centimeters": 1.0e-2,
|
||||
"mm": 1.0e-3,
|
||||
"millimetre": 1.0e-3,
|
||||
"millimetres": 1.0e-3,
|
||||
"millimeter": 1.0e-3,
|
||||
"millimeters": 1.0e-3,
|
||||
}
|
||||
_ANGLE_UNITS = {
|
||||
"degree": np.deg2rad,
|
||||
"degrees": np.deg2rad,
|
||||
"radian": lambda value: np.asarray(value, dtype=np.float64),
|
||||
"radians": lambda value: np.asarray(value, dtype=np.float64),
|
||||
}
|
||||
|
||||
|
||||
class SofaImportError(ValueError):
|
||||
"""The file is outside the deliberately narrow public SOFA contract."""
|
||||
|
||||
|
||||
def _text(value: Any) -> str:
|
||||
if isinstance(value, np.ndarray) and value.shape == ():
|
||||
value = value.item()
|
||||
if isinstance(value, (bytes, np.bytes_)):
|
||||
return value.decode("utf-8", "strict")
|
||||
return str(value)
|
||||
|
||||
|
||||
def _sha256_stream(stream) -> str:
|
||||
digest = hashlib.sha256()
|
||||
stream.seek(0)
|
||||
for block in iter(lambda: stream.read(4 << 20), b""):
|
||||
digest.update(block)
|
||||
stream.seek(0)
|
||||
return digest.hexdigest().upper()
|
||||
|
||||
|
||||
@contextmanager
|
||||
def _stable_hdf5_source(path: Path):
|
||||
"""Read arrays and content identity from one stable open-file snapshot."""
|
||||
with path.open("rb") as stream:
|
||||
before = _sha256_stream(stream)
|
||||
with h5py.File(stream, "r") as file:
|
||||
yield file, before
|
||||
after = _sha256_stream(stream)
|
||||
if after != before:
|
||||
raise SofaImportError("SOFA file changed while it was being imported")
|
||||
|
||||
|
||||
def _tokens(units: str) -> list[str]:
|
||||
return [token.strip().lower() for token in units.split(",") if token.strip()]
|
||||
|
||||
|
||||
def _coordinate_attributes(dataset: h5py.Dataset, *, inherit=None) -> tuple[str, str]:
|
||||
source = dataset.attrs
|
||||
if "Type" not in source or "Units" not in source:
|
||||
if inherit is None or "Type" not in inherit.attrs or "Units" not in inherit.attrs:
|
||||
raise SofaImportError(f"{dataset.name} must declare Type and Units")
|
||||
source = inherit.attrs
|
||||
return _text(source["Type"]).strip().lower(), _text(source["Units"]).strip()
|
||||
|
||||
|
||||
def coordinates_to_cartesian_m(values, coordinate_type: str, units: str,
|
||||
*, variable: str) -> np.ndarray:
|
||||
"""Convert a SOFA coordinate array to Cartesian metres without reshaping it."""
|
||||
data = np.asarray(values, dtype=np.float64)
|
||||
if data.shape[-1] != 3 or not np.isfinite(data).all():
|
||||
raise SofaImportError(f"{variable} must contain finite C=3 coordinates")
|
||||
kind = coordinate_type.strip().lower()
|
||||
unit_tokens = _tokens(units)
|
||||
if kind == "cartesian":
|
||||
if len(unit_tokens) == 1:
|
||||
factors = [_LENGTH_UNITS.get(unit_tokens[0])] * 3
|
||||
elif len(unit_tokens) == 3:
|
||||
factors = [_LENGTH_UNITS.get(token) for token in unit_tokens]
|
||||
else:
|
||||
factors = []
|
||||
if len(factors) != 3 or any(value is None for value in factors):
|
||||
raise SofaImportError(f"unsupported Cartesian units for {variable}: {units!r}")
|
||||
return data * np.asarray(factors, dtype=np.float64)
|
||||
if kind != "spherical" or len(unit_tokens) != 3:
|
||||
raise SofaImportError(
|
||||
f"unsupported coordinates for {variable}: Type={coordinate_type!r}, Units={units!r}")
|
||||
if unit_tokens[0] not in _ANGLE_UNITS or unit_tokens[1] not in _ANGLE_UNITS:
|
||||
raise SofaImportError(f"unsupported spherical angle units for {variable}: {units!r}")
|
||||
radius_factor = _LENGTH_UNITS.get(unit_tokens[2])
|
||||
if radius_factor is None:
|
||||
raise SofaImportError(f"unsupported spherical radius unit for {variable}: {units!r}")
|
||||
azimuth = _ANGLE_UNITS[unit_tokens[0]](data[..., 0])
|
||||
elevation = _ANGLE_UNITS[unit_tokens[1]](data[..., 1])
|
||||
radius = data[..., 2] * radius_factor
|
||||
if np.any(radius < 0.0):
|
||||
raise SofaImportError(f"{variable} contains a negative spherical radius")
|
||||
horizontal = np.cos(elevation)
|
||||
return np.stack(
|
||||
(radius * horizontal * np.cos(azimuth),
|
||||
radius * horizontal * np.sin(azimuth),
|
||||
radius * np.sin(elevation)),
|
||||
axis=-1,
|
||||
).astype(np.float64, copy=False)
|
||||
|
||||
|
||||
def _rows(file: h5py.File, name: str, measurements: int, *, inherit=None) -> np.ndarray:
|
||||
if name not in file:
|
||||
raise SofaImportError(f"missing required SOFA variable {name}")
|
||||
dataset = file[name]
|
||||
kind, units = _coordinate_attributes(dataset, inherit=inherit)
|
||||
result = coordinates_to_cartesian_m(dataset[...], kind, units, variable=name)
|
||||
if result.shape not in ((1, 3), (measurements, 3)):
|
||||
raise SofaImportError(
|
||||
f"{name} must have shape [I,C] or [M,C], got {result.shape}")
|
||||
return np.broadcast_to(result, (measurements, 3)).astype(np.float64, copy=True)
|
||||
|
||||
|
||||
def _receiver_rows(file: h5py.File, measurements: int) -> np.ndarray:
|
||||
if "ReceiverPosition" not in file:
|
||||
raise SofaImportError("missing required SOFA variable ReceiverPosition")
|
||||
dataset = file["ReceiverPosition"]
|
||||
raw = np.asarray(dataset[...], dtype=np.float64)
|
||||
if raw.shape not in ((2, 3, 1), (2, 3, measurements)):
|
||||
raise SofaImportError(
|
||||
"ReceiverPosition must have shape [R=2,C=3,I=1 or M]")
|
||||
values = np.moveaxis(raw, 1, -1) # [R,I/M,C]
|
||||
kind, units = _coordinate_attributes(dataset)
|
||||
cartesian = coordinates_to_cartesian_m(
|
||||
values, kind, units, variable="ReceiverPosition")
|
||||
cartesian = np.moveaxis(cartesian, 0, 1) # [I/M,R,C]
|
||||
return np.broadcast_to(cartesian, (measurements, 2, 3)).astype(
|
||||
np.float64, copy=True)
|
||||
|
||||
|
||||
def _emitter_is_origin(file: h5py.File, measurements: int) -> None:
|
||||
if "EmitterPosition" not in file:
|
||||
raise SofaImportError("missing required SOFA variable EmitterPosition")
|
||||
dataset = file["EmitterPosition"]
|
||||
raw = np.asarray(dataset[...], dtype=np.float64)
|
||||
if raw.shape not in ((1, 3, 1), (1, 3, measurements)):
|
||||
raise SofaImportError("SimpleFreeFieldHRIR v1 requires E=1 EmitterPosition[E,C,I/M]")
|
||||
values = np.moveaxis(raw, 1, -1)
|
||||
kind, units = _coordinate_attributes(dataset)
|
||||
cartesian = coordinates_to_cartesian_m(
|
||||
values, kind, units, variable="EmitterPosition")
|
||||
if np.max(np.abs(cartesian), initial=0.0) > 1.0e-9:
|
||||
raise SofaImportError("non-zero EmitterPosition needs a separate source-pose adapter")
|
||||
|
||||
|
||||
def _sampling_rate(file: h5py.File) -> float:
|
||||
if "Data.SamplingRate" not in file:
|
||||
raise SofaImportError("missing Data.SamplingRate")
|
||||
dataset = file["Data.SamplingRate"]
|
||||
values = np.asarray(dataset[...], dtype=np.float64).reshape(-1)
|
||||
if values.size != 1 or not math.isfinite(float(values[0])) or values[0] <= 0.0:
|
||||
raise SofaImportError("Data.SamplingRate must contain one positive finite value")
|
||||
units = _text(dataset.attrs.get("Units", "")).strip().lower()
|
||||
if units not in {"hertz", "hz"}:
|
||||
raise SofaImportError(f"Data.SamplingRate Units must be hertz, got {units!r}")
|
||||
return float(values[0])
|
||||
|
||||
|
||||
def _processing_label(file: h5py.File) -> str:
|
||||
parts = []
|
||||
for key in ("DatabaseName", "Title", "ListenerShortName", "Comment"):
|
||||
value = _text(file.attrs.get(key, "")).strip()
|
||||
if value and value not in parts:
|
||||
parts.append(value)
|
||||
return " | ".join(parts)
|
||||
|
||||
|
||||
def _read_delay(file: h5py.File, measurements: int) -> np.ndarray:
|
||||
if "Data.Delay" not in file:
|
||||
raise SofaImportError("missing Data.Delay")
|
||||
delay = np.asarray(file["Data.Delay"][...], dtype=np.float64)
|
||||
if delay.shape not in ((1, 2), (measurements, 2)) or not np.isfinite(delay).all():
|
||||
raise SofaImportError("Data.Delay must have finite shape [I=1,R=2] or [M,R=2]")
|
||||
delay = np.broadcast_to(delay, (measurements, 2)).astype(np.float64, copy=True)
|
||||
if np.min(delay, initial=0.0) < -1.0e-9:
|
||||
raise SofaImportError("negative Data.Delay is outside the supported causal contract")
|
||||
delay[delay < 0.0] = 0.0
|
||||
return delay
|
||||
|
||||
|
||||
def _readonly(array, dtype) -> np.ndarray:
|
||||
result = np.asarray(array, dtype=dtype)
|
||||
result.setflags(write=False)
|
||||
return result
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class CanonicalHrtf:
|
||||
source_path: str
|
||||
source_sha256: str
|
||||
convention: str
|
||||
convention_version: str
|
||||
sofa_version: str
|
||||
source_sample_rate_hz: float
|
||||
sample_rate_hz: float
|
||||
source_position_cartesian_m: np.ndarray # listener-local [M,3]
|
||||
listener_view: np.ndarray # world, normalized [M,3]
|
||||
listener_up: np.ndarray # world, orthonormal [M,3]
|
||||
receiver_position_cartesian_m: np.ndarray # listener-local, L/R [M,2,3]
|
||||
left_receiver_index: int
|
||||
right_receiver_index: int
|
||||
hrir: np.ndarray # canonical L/R [M,2,N]
|
||||
delay_samples: np.ndarray # canonical L/R [M,2], not applied
|
||||
measurement_radius_m: np.ndarray # [M]
|
||||
processing_label: str
|
||||
resampling_label: str = "none"
|
||||
|
||||
def __post_init__(self):
|
||||
object.__setattr__(self, "source_position_cartesian_m", _readonly(
|
||||
self.source_position_cartesian_m, np.float64))
|
||||
object.__setattr__(self, "listener_view", _readonly(self.listener_view, np.float64))
|
||||
object.__setattr__(self, "listener_up", _readonly(self.listener_up, np.float64))
|
||||
object.__setattr__(self, "receiver_position_cartesian_m", _readonly(
|
||||
self.receiver_position_cartesian_m, np.float64))
|
||||
object.__setattr__(self, "hrir", _readonly(self.hrir, np.float64))
|
||||
object.__setattr__(self, "delay_samples", _readonly(self.delay_samples, np.float64))
|
||||
object.__setattr__(self, "measurement_radius_m", _readonly(
|
||||
self.measurement_radius_m, np.float64))
|
||||
|
||||
@property
|
||||
def measurements(self) -> int:
|
||||
return int(self.hrir.shape[0])
|
||||
|
||||
@property
|
||||
def taps(self) -> int:
|
||||
return int(self.hrir.shape[2])
|
||||
|
||||
@property
|
||||
def unit_directions(self) -> np.ndarray:
|
||||
return self.source_position_cartesian_m / self.measurement_radius_m[:, None]
|
||||
|
||||
@property
|
||||
def shells_m(self) -> np.ndarray:
|
||||
return np.unique(np.round(self.measurement_radius_m, 9))
|
||||
|
||||
def shell_indices(self, radius_m: float) -> np.ndarray:
|
||||
shell = float(self.shells_m[np.argmin(np.abs(self.shells_m - float(radius_m)))])
|
||||
return np.flatnonzero(np.isclose(
|
||||
self.measurement_radius_m, shell, atol=5.0e-7, rtol=0.0))
|
||||
|
||||
def nearest_index(self, direction_sofa, radius_m: float = 1.0) -> tuple[int, float]:
|
||||
direction = np.asarray(direction_sofa, dtype=np.float64)
|
||||
if direction.shape != (3,) or not np.isfinite(direction).all():
|
||||
raise ValueError("direction must contain three finite SOFA Cartesian values")
|
||||
norm = float(np.linalg.norm(direction))
|
||||
if norm <= 1.0e-15:
|
||||
raise ValueError("direction must be non-zero")
|
||||
direction = direction / norm
|
||||
indices = self.shell_indices(radius_m)
|
||||
dots = self.unit_directions[indices] @ direction
|
||||
local = int(np.argmax(dots))
|
||||
error = math.degrees(math.acos(float(np.clip(dots[local], -1.0, 1.0))))
|
||||
return int(indices[local]), float(error)
|
||||
|
||||
def resampled(self, target_sample_rate_hz: float) -> "CanonicalHrtf":
|
||||
target = float(target_sample_rate_hz)
|
||||
if not math.isfinite(target) or target <= 0.0:
|
||||
raise ValueError("target sample rate must be positive and finite")
|
||||
if abs(target - self.sample_rate_hz) <= 1.0e-9:
|
||||
return self
|
||||
ratio = target / self.sample_rate_hz
|
||||
fraction = Fraction(ratio).limit_denominator(100000)
|
||||
if abs(float(fraction) - ratio) > 1.0e-10:
|
||||
raise ValueError("sample-rate ratio cannot be represented safely")
|
||||
converted = signal.resample_poly(
|
||||
np.asarray(self.hrir, dtype=np.float64), fraction.numerator,
|
||||
fraction.denominator, axis=-1, window=("kaiser", 8.6), padtype="constant")
|
||||
converted = np.asarray(converted, dtype=np.float64)
|
||||
return replace(
|
||||
self,
|
||||
sample_rate_hz=target,
|
||||
hrir=converted,
|
||||
delay_samples=np.asarray(self.delay_samples * ratio, dtype=np.float64),
|
||||
resampling_label=(
|
||||
f"scipy.signal.resample_poly {self.sample_rate_hz:g}->{target:g} Hz "
|
||||
f"({fraction.numerator}/{fraction.denominator}, Kaiser beta=8.6)"),
|
||||
)
|
||||
|
||||
def materialized_measurement(self, index: int, *, fractional_half_length: int = 48
|
||||
) -> np.ndarray:
|
||||
"""Return [L/R,taps] with SOFA Data.Delay applied exactly once."""
|
||||
index = int(index)
|
||||
if not 0 <= index < self.measurements:
|
||||
raise IndexError(index)
|
||||
ears = []
|
||||
for ear in range(2):
|
||||
ears.append(apply_fractional_delay(
|
||||
self.hrir[index, ear], float(self.delay_samples[index, ear]),
|
||||
half_length=fractional_half_length))
|
||||
length = max(map(len, ears))
|
||||
result = np.zeros((2, length), dtype=np.float64)
|
||||
for ear, value in enumerate(ears):
|
||||
result[ear, :len(value)] = value
|
||||
return result
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"source_path": self.source_path,
|
||||
"source_sha256": self.source_sha256,
|
||||
"convention": self.convention,
|
||||
"convention_version": self.convention_version,
|
||||
"sofa_version": self.sofa_version,
|
||||
"source_sample_rate_hz": self.source_sample_rate_hz,
|
||||
"sample_rate_hz": self.sample_rate_hz,
|
||||
"measurements": self.measurements,
|
||||
"taps": self.taps,
|
||||
"shells_m": [float(value) for value in self.shells_m],
|
||||
"source_receiver_order": [self.left_receiver_index, self.right_receiver_index],
|
||||
"canonical_ear_order": ["left", "right"],
|
||||
"data_delay_samples_min": float(np.min(self.delay_samples)),
|
||||
"data_delay_samples_max": float(np.max(self.delay_samples)),
|
||||
"data_delay_applied": False,
|
||||
"processing_label": self.processing_label,
|
||||
"resampling": self.resampling_label,
|
||||
"precision": "float64",
|
||||
}
|
||||
|
||||
|
||||
def load_simple_free_field_hrir(path, *, target_sample_rate_hz: float | None = None
|
||||
) -> CanonicalHrtf:
|
||||
"""Strictly import the supported SimpleFreeFieldHRIR subset."""
|
||||
source = Path(path).expanduser().resolve()
|
||||
if not source.is_file():
|
||||
raise FileNotFoundError(source)
|
||||
with _stable_hdf5_source(source) as (file, source_sha256):
|
||||
if _text(file.attrs.get("Conventions", "")) != "SOFA":
|
||||
raise SofaImportError("Conventions must be SOFA")
|
||||
convention = _text(file.attrs.get("SOFAConventions", ""))
|
||||
if convention != "SimpleFreeFieldHRIR":
|
||||
raise SofaImportError(
|
||||
f"unsupported SOFAConventions={convention!r}; convert explicitly first")
|
||||
convention_version = _text(file.attrs.get("SOFAConventionsVersion", ""))
|
||||
if convention_version not in _SUPPORTED_VERSIONS:
|
||||
raise SofaImportError(
|
||||
f"unsupported SimpleFreeFieldHRIR version {convention_version!r}; "
|
||||
f"supported={sorted(_SUPPORTED_VERSIONS)}")
|
||||
if _text(file.attrs.get("DataType", "")) != "FIR":
|
||||
raise SofaImportError("DataType must be FIR")
|
||||
room_type = _text(file.attrs.get("RoomType", "")).strip().lower()
|
||||
if room_type not in _FREE_FIELD_ROOM_TYPES:
|
||||
raise SofaImportError(f"RoomType must explicitly be free-field, got {room_type!r}")
|
||||
if "Data.IR" not in file:
|
||||
raise SofaImportError("missing Data.IR")
|
||||
hrir_source = np.asarray(file["Data.IR"][...], dtype=np.float64)
|
||||
if hrir_source.ndim != 3 or hrir_source.shape[1] != 2 or min(hrir_source.shape) <= 0:
|
||||
raise SofaImportError("Data.IR must have shape [M,R=2,N]")
|
||||
if not np.isfinite(hrir_source).all():
|
||||
raise SofaImportError("Data.IR contains non-finite values")
|
||||
measurements = int(hrir_source.shape[0])
|
||||
sample_rate = _sampling_rate(file)
|
||||
delay_source = _read_delay(file, measurements)
|
||||
_emitter_is_origin(file, measurements)
|
||||
|
||||
listener_position = _rows(file, "ListenerPosition", measurements)
|
||||
listener_view = _rows(file, "ListenerView", measurements)
|
||||
listener_up_raw = _rows(
|
||||
file, "ListenerUp", measurements,
|
||||
inherit=file["ListenerView"] if "ListenerView" in file else None)
|
||||
forward_norm = np.linalg.norm(listener_view, axis=1)
|
||||
if np.any(forward_norm <= 1.0e-12):
|
||||
raise SofaImportError("ListenerView must be non-zero")
|
||||
forward = listener_view / forward_norm[:, None]
|
||||
left = np.cross(listener_up_raw, forward)
|
||||
left_norm = np.linalg.norm(left, axis=1)
|
||||
if np.any(left_norm <= 1.0e-12):
|
||||
raise SofaImportError("ListenerUp must not be parallel to ListenerView")
|
||||
left /= left_norm[:, None]
|
||||
up = np.cross(forward, left)
|
||||
|
||||
if "SourcePosition" not in file:
|
||||
raise SofaImportError("missing required SOFA variable SourcePosition")
|
||||
source_dataset = file["SourcePosition"]
|
||||
source_type, source_units = _coordinate_attributes(source_dataset)
|
||||
source_world = coordinates_to_cartesian_m(
|
||||
source_dataset[...], source_type, source_units, variable="SourcePosition")
|
||||
if source_world.shape != (measurements, 3):
|
||||
raise SofaImportError("SourcePosition must have shape [M,C=3]")
|
||||
relative = source_world - listener_position
|
||||
source_local = np.stack(
|
||||
(np.sum(relative * forward, axis=1),
|
||||
np.sum(relative * left, axis=1),
|
||||
np.sum(relative * up, axis=1)), axis=1)
|
||||
radii = np.linalg.norm(source_local, axis=1)
|
||||
if np.any(radii <= 1.0e-8) or not np.isfinite(radii).all():
|
||||
raise SofaImportError("every source measurement must have a positive radius")
|
||||
|
||||
receiver = _receiver_rows(file, measurements)
|
||||
lateral_difference = receiver[:, 0, 1] - receiver[:, 1, 1]
|
||||
if np.all(lateral_difference > 1.0e-5):
|
||||
left_index, right_index = 0, 1
|
||||
elif np.all(lateral_difference < -1.0e-5):
|
||||
left_index, right_index = 1, 0
|
||||
else:
|
||||
raise SofaImportError(
|
||||
"ReceiverPosition does not identify one consistently-left and one "
|
||||
"consistently-right receiver")
|
||||
ear_order = [left_index, right_index]
|
||||
hrir = hrir_source[:, ear_order, :]
|
||||
delay = delay_source[:, ear_order]
|
||||
receiver = receiver[:, ear_order, :]
|
||||
processing_label = _processing_label(file)
|
||||
sofa_version = _text(file.attrs.get("Version", ""))
|
||||
|
||||
canonical = CanonicalHrtf(
|
||||
source_path=str(source),
|
||||
source_sha256=source_sha256,
|
||||
convention=convention,
|
||||
convention_version=convention_version,
|
||||
sofa_version=sofa_version,
|
||||
source_sample_rate_hz=sample_rate,
|
||||
sample_rate_hz=sample_rate,
|
||||
source_position_cartesian_m=source_local,
|
||||
listener_view=forward,
|
||||
listener_up=up,
|
||||
receiver_position_cartesian_m=receiver,
|
||||
left_receiver_index=left_index,
|
||||
right_receiver_index=right_index,
|
||||
hrir=hrir,
|
||||
delay_samples=delay,
|
||||
measurement_radius_m=radii,
|
||||
processing_label=processing_label,
|
||||
)
|
||||
return (canonical if target_sample_rate_hz is None
|
||||
else canonical.resampled(target_sample_rate_hz))
|
||||
|
||||
|
||||
def apply_fractional_delay(values, delay_samples: float, *, half_length: int = 48
|
||||
) -> np.ndarray:
|
||||
"""Apply one causal non-negative delay to a real FIR using windowed sinc."""
|
||||
source = np.asarray(values, dtype=np.float64)
|
||||
delay = float(delay_samples)
|
||||
if source.ndim != 1 or not np.isfinite(source).all():
|
||||
raise ValueError("fractional delay input must be a finite real vector")
|
||||
if not math.isfinite(delay) or delay < -1.0e-12:
|
||||
raise ValueError("fractional delay must be finite and non-negative")
|
||||
if delay < 1.0e-12:
|
||||
return source.copy()
|
||||
integer = int(math.floor(delay))
|
||||
fraction = delay - integer
|
||||
if fraction < 1.0e-12:
|
||||
return np.pad(source, (integer, 0)).astype(np.float64, copy=False)
|
||||
half = int(half_length)
|
||||
if half < 8:
|
||||
raise ValueError("fractional delay half_length must be at least 8")
|
||||
index = np.arange(-half, half + 1, dtype=np.float64)
|
||||
kernel = np.sinc(index - fraction) * np.kaiser(2 * half + 1, 8.6)
|
||||
kernel /= np.sum(kernel, dtype=np.float64)
|
||||
full = signal.fftconvolve(source, kernel, mode="full")
|
||||
causal = np.asarray(full[half:], dtype=np.float64)
|
||||
return np.pad(causal, (integer, 0)).astype(np.float64, copy=False)
|
||||
|
||||
|
||||
def shift_signal_fft(values, shift_samples: float) -> np.ndarray:
|
||||
"""Band-limited linear shift; positive is delay and negative is advance."""
|
||||
source = np.asarray(values, dtype=np.float64)
|
||||
shift = float(shift_samples)
|
||||
if source.ndim != 1 or not np.isfinite(source).all() or not math.isfinite(shift):
|
||||
raise ValueError("shift input and amount must be finite")
|
||||
if abs(shift) < 1.0e-12:
|
||||
return source.copy()
|
||||
guard = max(128, int(math.ceil(abs(shift))) + 64)
|
||||
needed = len(source) + 2 * guard
|
||||
fft_size = next_fast_len(needed)
|
||||
padded = np.zeros(fft_size, dtype=np.float64)
|
||||
padded[guard:guard + len(source)] = source
|
||||
bins = np.arange(fft_size // 2 + 1, dtype=np.float64)
|
||||
spectrum = np.fft.rfft(padded)
|
||||
spectrum *= np.exp(-2j * np.pi * bins * shift / fft_size)
|
||||
shifted = np.fft.irfft(spectrum, fft_size)
|
||||
return np.asarray(shifted[guard:guard + len(source)], dtype=np.float64)
|
||||
|
||||
|
||||
def estimate_interaural_delay_samples(left, right, sample_rate_hz: float,
|
||||
*, low_hz: float = 200.0,
|
||||
high_hz: float = 1500.0) -> float:
|
||||
"""Estimate L-minus-R delay by low-frequency circular phase coherence.
|
||||
|
||||
A coarse-to-fine delay search avoids the phase-unwrapping branch failures
|
||||
that ordinary straight-line regression can exhibit for strongly filtered
|
||||
Far responses.
|
||||
"""
|
||||
left = np.asarray(left, dtype=np.float64)
|
||||
right = np.asarray(right, dtype=np.float64)
|
||||
if left.shape != right.shape or left.ndim != 1:
|
||||
raise ValueError("ITD inputs must be equal-length vectors")
|
||||
fft_size = next_fast_len(max(4096, 4 * len(left)))
|
||||
left_spectrum = np.fft.rfft(left, fft_size)
|
||||
right_spectrum = np.fft.rfft(right, fft_size)
|
||||
frequency = np.fft.rfftfreq(fft_size, 1.0 / float(sample_rate_hz))
|
||||
selected = (frequency >= low_hz) & (frequency <= high_hz)
|
||||
if np.count_nonzero(selected) < 8:
|
||||
return 0.0
|
||||
cross = left_spectrum[selected] * np.conj(right_spectrum[selected])
|
||||
magnitude = np.abs(cross)
|
||||
maximum = float(np.max(magnitude, initial=0.0))
|
||||
if maximum <= 1.0e-20:
|
||||
return 0.0
|
||||
weighted_unit = cross / np.maximum(magnitude, 1.0e-30)
|
||||
weight = np.sqrt(magnitude / maximum)
|
||||
weighted_unit *= weight
|
||||
omega = 2.0 * np.pi * frequency[selected] / float(sample_rate_hz)
|
||||
limit = 0.0012 * float(sample_rate_hz)
|
||||
|
||||
def best(candidates: np.ndarray) -> float:
|
||||
steering = np.exp(1j * omega[:, None] * candidates[None, :])
|
||||
score = np.abs(weighted_unit @ steering)
|
||||
return float(candidates[int(np.argmax(score))])
|
||||
|
||||
coarse = np.arange(-limit, limit + 0.25, 0.5, dtype=np.float64)
|
||||
estimate = best(coarse)
|
||||
fine = np.arange(estimate - 0.6, estimate + 0.6001, 0.02, dtype=np.float64)
|
||||
return float(np.clip(best(fine), -limit, limit))
|
||||
|
||||
def _subsample_peak(values: np.ndarray) -> float:
|
||||
magnitude = np.abs(np.asarray(values, dtype=np.float64))
|
||||
index = int(np.argmax(magnitude))
|
||||
if index == 0 or index + 1 >= len(magnitude):
|
||||
return float(index)
|
||||
y0, y1, y2 = (float(magnitude[index - 1]), float(magnitude[index]),
|
||||
float(magnitude[index + 1]))
|
||||
denominator = y0 - 2.0 * y1 + y2
|
||||
correction = 0.0 if abs(denominator) < 1.0e-30 else 0.5 * (y0 - y2) / denominator
|
||||
return float(index + np.clip(correction, -0.5, 0.5))
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class TimeAlignedHrtf:
|
||||
canonical: CanonicalHrtf
|
||||
aligned_hrir: np.ndarray
|
||||
runtime_delay_samples: np.ndarray
|
||||
embedded_delay_removed_samples: np.ndarray
|
||||
delay_source: str
|
||||
|
||||
def __post_init__(self):
|
||||
object.__setattr__(self, "aligned_hrir", _readonly(self.aligned_hrir, np.float64))
|
||||
object.__setattr__(self, "runtime_delay_samples", _readonly(
|
||||
self.runtime_delay_samples, np.float64))
|
||||
object.__setattr__(self, "embedded_delay_removed_samples", _readonly(
|
||||
self.embedded_delay_removed_samples, np.float64))
|
||||
|
||||
|
||||
def time_align_hrtf(canonical: CanonicalHrtf) -> TimeAlignedHrtf:
|
||||
"""Separate one delay representation before directional interpolation.
|
||||
|
||||
Trusted non-zero ``Data.Delay`` is external to ``Data.IR`` and is therefore
|
||||
retained without de-rotating the FIR. When ``Data.Delay`` is identically
|
||||
zero, ordinary measured HRIRs with a positive onset use their per-ear main
|
||||
peaks. A zero-origin effective FIR is already expressed at one common
|
||||
time origin; its interaural phase is therefore retained in ``Data.IR``.
|
||||
|
||||
These representations are mutually exclusive. Runtime rendering must
|
||||
restore exactly the delay separated here and must not add any second ear
|
||||
delay or phase-group delay.
|
||||
"""
|
||||
hrir = np.asarray(canonical.hrir, dtype=np.float64)
|
||||
if np.max(np.abs(canonical.delay_samples), initial=0.0) > 1.0e-12:
|
||||
return TimeAlignedHrtf(
|
||||
canonical=canonical,
|
||||
aligned_hrir=hrir.copy(),
|
||||
runtime_delay_samples=np.asarray(canonical.delay_samples, dtype=np.float64),
|
||||
embedded_delay_removed_samples=np.zeros_like(canonical.delay_samples),
|
||||
delay_source="Data.Delay (external; applied once at render time)",
|
||||
)
|
||||
|
||||
measurements = canonical.measurements
|
||||
runtime = np.zeros((measurements, 2), dtype=np.float64)
|
||||
removed = np.zeros_like(runtime)
|
||||
used_peak = 0
|
||||
retained_embedded_phase = 0
|
||||
for measurement in range(measurements):
|
||||
peaks = np.asarray([
|
||||
_subsample_peak(hrir[measurement, 0]),
|
||||
_subsample_peak(hrir[measurement, 1]),
|
||||
], dtype=np.float64)
|
||||
if float(np.max(peaks)) > 2.0:
|
||||
delays = peaks
|
||||
used_peak += 1
|
||||
else:
|
||||
# An effective response can have both ear FIRs beginning at sample
|
||||
# zero while still carrying the correct ITD in complex phase. Do
|
||||
# not invent an external delay which SOFA did not author.
|
||||
delays = np.zeros(2, dtype=np.float64)
|
||||
retained_embedded_phase += 1
|
||||
runtime[measurement] = delays
|
||||
removed[measurement] = delays
|
||||
|
||||
aligned = np.empty_like(hrir)
|
||||
for measurement in range(measurements):
|
||||
for ear in range(2):
|
||||
aligned[measurement, ear] = shift_signal_fft(
|
||||
hrir[measurement, ear], -float(removed[measurement, ear]))
|
||||
source = (
|
||||
f"embedded Data.IR arrival separation: peak={used_peak}, "
|
||||
f"zero-origin embedded phase retained={retained_embedded_phase}; "
|
||||
"positive onset restored once at render time")
|
||||
return TimeAlignedHrtf(
|
||||
canonical=canonical,
|
||||
aligned_hrir=aligned,
|
||||
runtime_delay_samples=runtime,
|
||||
embedded_delay_removed_samples=removed,
|
||||
delay_source=source,
|
||||
)
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,398 @@
|
||||
"""Native float64 SOFA binaural DSP (ctypes bridge to eac3joc_core).
|
||||
|
||||
The C++ side mirrors the Python :class:`sofa_binaural_backend.SofaBinauralBackend`
|
||||
mathematics: 64-QMF/77-hybrid analysis and synthesis, fifth-order ACN/N3D real
|
||||
spherical-harmonic field evaluation, whole-QMF-slot per-object delay histories,
|
||||
six image-source early reflections, the shared unitary FDN late room, the LFE
|
||||
low-pass and the 961-sample latency policy. The compiled HRTF field, the
|
||||
filterbank tables and the room constants are uploaded once; per 512-sample
|
||||
block the adapter updates every source and streams PCM through the DLL.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import ctypes
|
||||
import math
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
from native_renderer import ABI_VERSION, find_native_library
|
||||
from public_filterbank import DEFAULT_FILTERBANK_DATA, load_filterbank_tables
|
||||
from public_room import LateFdnConfig, SharedUnitaryFdn, ShoeboxRoomConfig
|
||||
from reference_distance import ReferenceDistanceProfileV1
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
from sofa_hrtf_field import (
|
||||
DEFAULT_HRTF_CACHE_DIR,
|
||||
SofaHrtfField,
|
||||
compile_sofa_hrtf,
|
||||
)
|
||||
|
||||
BLOCK_SAMPLES = 512
|
||||
INPUT_CHANNELS = 16
|
||||
OUTPUT_CHANNELS = 2
|
||||
QMF_HOP = 64
|
||||
LATENCY_SAMPLES = 961
|
||||
|
||||
_PROFILE_INDEX = {"near": 0, "mid": 1, "far": 2}
|
||||
|
||||
|
||||
def _room_numbers(fdn_config: LateFdnConfig) -> dict:
|
||||
"""Derive the FDN delays/feedback with the same arithmetic as the Python room."""
|
||||
fdn = SharedUnitaryFdn(fdn_config)
|
||||
return {
|
||||
"fdn_delays": np.asarray(fdn.delays, dtype=np.uint32),
|
||||
"fdn_feedback": np.asarray(fdn.feedback_gain, dtype=np.float64),
|
||||
"damping": fdn.damping,
|
||||
"output_gain": fdn.output_gain,
|
||||
"allpass_delays": np.asarray(
|
||||
[diffuser.delay_samples for diffuser in fdn.diffusers], dtype=np.uint32),
|
||||
"allpass_gains": np.asarray(fdn_config.allpass_gain, dtype=np.float64),
|
||||
"tail_samples": fdn.tail_samples,
|
||||
}
|
||||
|
||||
|
||||
class NativeSofaBinauralDsp:
|
||||
"""Duck-type compatible with SofaBinauralBackend for the JOC adapter."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
field: SofaHrtfField,
|
||||
*,
|
||||
source_count: int = INPUT_CHANNELS,
|
||||
default_profile: str = "mid",
|
||||
enable_early_reflections: bool = True,
|
||||
enable_late_room: bool = True,
|
||||
room_config: ShoeboxRoomConfig = ShoeboxRoomConfig(),
|
||||
fdn_config: LateFdnConfig | None = None,
|
||||
library_path: str | Path | None = None):
|
||||
if not isinstance(field, SofaHrtfField):
|
||||
raise TypeError("field must be SofaHrtfField")
|
||||
self.source_count = int(source_count)
|
||||
self.default_profile = ReferenceDistanceProfileV1.validate_profile(default_profile)
|
||||
self.enable_early_reflections = bool(enable_early_reflections)
|
||||
self.enable_late_room = bool(enable_late_room)
|
||||
self.field = field
|
||||
if self.source_count != INPUT_CHANNELS:
|
||||
raise ValueError(f"native SOFA backend requires {INPUT_CHANNELS} sources")
|
||||
if abs(self.field.sample_rate_hz - 48000.0) > 1.0e-9:
|
||||
raise ValueError("the native SOFA binaural runtime requires 48 kHz")
|
||||
self.room_config = room_config
|
||||
self.room_config.validate()
|
||||
self.dsp_backend = "native-sofa"
|
||||
self.hrtf_input_kind = "field"
|
||||
self.hrtf_input_path = None
|
||||
self.cache_policy = None
|
||||
|
||||
self.library_path = find_native_library(library_path)
|
||||
self._lib = ctypes.CDLL(str(self.library_path))
|
||||
self._bind()
|
||||
version = int(self._lib.ejoc_abi_version())
|
||||
if version != ABI_VERSION:
|
||||
raise RuntimeError(
|
||||
f"native ABI mismatch: expected {ABI_VERSION}, got {version}")
|
||||
self._handle = self._lib.ejoc_sofa_binaural_create()
|
||||
if not self._handle:
|
||||
raise RuntimeError("native SOFA binaural renderer creation failed")
|
||||
try:
|
||||
self._configure_kernels()
|
||||
self._configure_field()
|
||||
self._configure_room(fdn_config)
|
||||
except Exception:
|
||||
self.close()
|
||||
raise
|
||||
|
||||
self.positions = np.zeros((self.source_count, 3), dtype=np.float64)
|
||||
self.positions[:, 1] = 1.0
|
||||
self.profiles = [self.default_profile] * self.source_count
|
||||
self.user_gain = np.ones(self.source_count, dtype=np.float64)
|
||||
self.special_lfe = np.zeros(self.source_count, dtype=bool)
|
||||
self.parameter_updates = 0
|
||||
self.finished = False
|
||||
for source in range(self.source_count):
|
||||
self.set_source(
|
||||
source, self.positions[source], profile=self.default_profile,
|
||||
fade=False)
|
||||
|
||||
def _bind(self):
|
||||
void_p = ctypes.c_void_p
|
||||
f64_p = ctypes.POINTER(ctypes.c_double)
|
||||
i16_p = ctypes.POINTER(ctypes.c_int16)
|
||||
u32_p = ctypes.POINTER(ctypes.c_uint32)
|
||||
self._lib.ejoc_abi_version.argtypes = []
|
||||
self._lib.ejoc_abi_version.restype = ctypes.c_uint32
|
||||
self._lib.ejoc_sofa_binaural_create.argtypes = []
|
||||
self._lib.ejoc_sofa_binaural_create.restype = void_p
|
||||
self._lib.ejoc_sofa_binaural_destroy.argtypes = [void_p]
|
||||
self._lib.ejoc_sofa_binaural_destroy.restype = None
|
||||
self._lib.ejoc_sofa_binaural_reset.argtypes = [void_p]
|
||||
self._lib.ejoc_sofa_binaural_reset.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_last_error.argtypes = [void_p]
|
||||
self._lib.ejoc_sofa_binaural_last_error.restype = ctypes.c_char_p
|
||||
self._lib.ejoc_sofa_binaural_configure_kernels.argtypes = [
|
||||
void_p, f64_p, f64_p, i16_p, f64_p, ctypes.c_uint32, f64_p, f64_p]
|
||||
self._lib.ejoc_sofa_binaural_configure_kernels.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_configure_field.argtypes = [
|
||||
void_p, f64_p, f64_p, f64_p, f64_p, ctypes.c_double]
|
||||
self._lib.ejoc_sofa_binaural_configure_field.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_configure_room.argtypes = [
|
||||
void_p, f64_p, f64_p, f64_p, ctypes.c_double, u32_p, f64_p,
|
||||
ctypes.c_double, ctypes.c_double, u32_p, f64_p,
|
||||
ctypes.c_uint32, ctypes.c_uint32]
|
||||
self._lib.ejoc_sofa_binaural_configure_room.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_set_source.argtypes = [
|
||||
void_p, ctypes.c_uint32, f64_p, ctypes.c_uint32, ctypes.c_double,
|
||||
ctypes.c_uint32, ctypes.c_uint32, ctypes.c_uint32]
|
||||
self._lib.ejoc_sofa_binaural_set_source.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_process.argtypes = [
|
||||
void_p, f64_p, ctypes.c_uint32, ctypes.c_double, f64_p]
|
||||
self._lib.ejoc_sofa_binaural_process.restype = ctypes.c_int
|
||||
self._lib.ejoc_sofa_binaural_finish.argtypes = [
|
||||
void_p, ctypes.c_uint32, f64_p, ctypes.c_uint32]
|
||||
self._lib.ejoc_sofa_binaural_finish.restype = ctypes.c_int
|
||||
|
||||
def _raise(self, operation, status):
|
||||
message = self._lib.ejoc_sofa_binaural_last_error(self._handle)
|
||||
detail = (message or b"").decode("utf-8", "replace")
|
||||
raise RuntimeError(
|
||||
f"native SOFA binaural renderer {operation} failed ({status}): {detail}")
|
||||
|
||||
@staticmethod
|
||||
def _f64_pointer(values):
|
||||
return values.ctypes.data_as(ctypes.POINTER(ctypes.c_double))
|
||||
|
||||
def _configure_kernels(self):
|
||||
tables = load_filterbank_tables(DEFAULT_FILTERBANK_DATA)
|
||||
qmf_analysis = np.ascontiguousarray(
|
||||
tables["qmf_analysis_coefficients"], dtype=np.float64)
|
||||
hybrid_low = np.ascontiguousarray(
|
||||
tables["hybrid_analysis_low_kernel"], dtype=np.float64)
|
||||
hybrid_indices = np.ascontiguousarray(
|
||||
tables["hybrid_synthesis_indices"], dtype=np.int16)
|
||||
hybrid_values = np.ascontiguousarray(
|
||||
tables["hybrid_synthesis_values"], dtype=np.float64)
|
||||
qmf_basis = np.ascontiguousarray(
|
||||
tables["qmf_synthesis_basis"], dtype=np.float64)
|
||||
qmf_taps = np.ascontiguousarray(
|
||||
tables["qmf_synthesis_taps"], dtype=np.float64)
|
||||
status = self._lib.ejoc_sofa_binaural_configure_kernels(
|
||||
self._handle,
|
||||
self._f64_pointer(qmf_analysis),
|
||||
self._f64_pointer(hybrid_low),
|
||||
hybrid_indices.ctypes.data_as(ctypes.POINTER(ctypes.c_int16)),
|
||||
self._f64_pointer(hybrid_values),
|
||||
len(hybrid_indices),
|
||||
self._f64_pointer(qmf_basis),
|
||||
self._f64_pointer(qmf_taps))
|
||||
if status:
|
||||
self._raise("configure_kernels", status)
|
||||
self._keepalive = (qmf_analysis, hybrid_low, hybrid_indices,
|
||||
hybrid_values, qmf_basis, qmf_taps)
|
||||
|
||||
def _configure_field(self):
|
||||
coefficients = np.ascontiguousarray(
|
||||
self.field.coefficients, dtype=np.complex128).view(np.float64)
|
||||
delay_coefficients = np.ascontiguousarray(
|
||||
self.field.delay_coefficients, dtype=np.float64)
|
||||
delay_bounds = np.ascontiguousarray(
|
||||
self.field.delay_bounds, dtype=np.float64)
|
||||
centers = np.ascontiguousarray(
|
||||
self.field.band_center_frequencies_hz, dtype=np.float64)
|
||||
status = self._lib.ejoc_sofa_binaural_configure_field(
|
||||
self._handle,
|
||||
self._f64_pointer(coefficients),
|
||||
self._f64_pointer(delay_coefficients),
|
||||
self._f64_pointer(delay_bounds),
|
||||
self._f64_pointer(centers),
|
||||
float(self.field.measurement_radius_m))
|
||||
if status:
|
||||
self._raise("configure_field", status)
|
||||
|
||||
def _configure_room(self, fdn_config: LateFdnConfig | None):
|
||||
actual = fdn_config or LateFdnConfig(sample_rate_hz=48000.0)
|
||||
numbers = _room_numbers(actual)
|
||||
dims = np.asarray(self.room_config.dimensions_m, dtype=np.float64)
|
||||
listener = np.asarray(self.room_config.listener_position_m, dtype=np.float64)
|
||||
walls = np.asarray(self.room_config.wall_reflection_gain, dtype=np.float64)
|
||||
status = self._lib.ejoc_sofa_binaural_configure_room(
|
||||
self._handle,
|
||||
self._f64_pointer(dims),
|
||||
self._f64_pointer(listener),
|
||||
self._f64_pointer(walls),
|
||||
float(self.room_config.speed_of_sound_m_s),
|
||||
numbers["fdn_delays"].ctypes.data_as(ctypes.POINTER(ctypes.c_uint32)),
|
||||
self._f64_pointer(numbers["fdn_feedback"]),
|
||||
float(numbers["damping"]),
|
||||
float(numbers["output_gain"]),
|
||||
numbers["allpass_delays"].ctypes.data_as(ctypes.POINTER(ctypes.c_uint32)),
|
||||
self._f64_pointer(numbers["allpass_gains"]),
|
||||
1 if self.enable_early_reflections else 0,
|
||||
1 if self.enable_late_room else 0)
|
||||
if status:
|
||||
self._raise("configure_room", status)
|
||||
self._fdn = SharedUnitaryFdn(actual)
|
||||
self._fdn_tail_samples = numbers["tail_samples"]
|
||||
|
||||
def set_source(self, source: int, position_adm, *, profile: str | None = None,
|
||||
gain: float = 1.0, enabled: bool = True,
|
||||
special_lfe: bool = False, fade: bool = True) -> None:
|
||||
if self.finished:
|
||||
raise RuntimeError("SOFA renderer is finished")
|
||||
source = int(source)
|
||||
if not 0 <= source < self.source_count:
|
||||
raise IndexError(source)
|
||||
name = self.default_profile if profile is None else profile
|
||||
position = np.asarray(position_adm, dtype=np.float64)
|
||||
if position.shape != (3,):
|
||||
raise ValueError("ADM position must contain three Cartesian values")
|
||||
status = self._lib.ejoc_sofa_binaural_set_source(
|
||||
self._handle, source,
|
||||
self._f64_pointer(np.ascontiguousarray(position)),
|
||||
_PROFILE_INDEX[ReferenceDistanceProfileV1.validate_profile(name)],
|
||||
float(gain), 1 if enabled else 0, 1 if special_lfe else 0,
|
||||
1 if fade else 0)
|
||||
if status:
|
||||
self._raise("set_source", status)
|
||||
self.positions[source] = position
|
||||
self.profiles[source] = name
|
||||
self.user_gain[source] = float(gain)
|
||||
self.special_lfe[source] = bool(special_lfe)
|
||||
self.parameter_updates += 1
|
||||
|
||||
def process(self, sources) -> np.ndarray:
|
||||
if self.finished:
|
||||
raise RuntimeError("SOFA renderer is finished")
|
||||
values = np.ascontiguousarray(sources, dtype=np.float64)
|
||||
if values.ndim != 2 or values.shape[1] != self.source_count:
|
||||
raise ValueError(f"sources must have shape [samples,{self.source_count}]")
|
||||
if len(values) % QMF_HOP or len(values) > BLOCK_SAMPLES:
|
||||
raise ValueError("native SOFA backend input must be a 64-aligned block")
|
||||
output = np.empty((len(values), OUTPUT_CHANNELS), dtype=np.float64)
|
||||
count = self._lib.ejoc_sofa_binaural_process(
|
||||
self._handle, self._f64_pointer(values), len(values), 1.0,
|
||||
self._f64_pointer(output))
|
||||
if count < 0:
|
||||
self._raise("process", count)
|
||||
return output[:count]
|
||||
|
||||
def finish(self, *, tail_seconds: float | None = None) -> np.ndarray:
|
||||
if self.finished:
|
||||
return np.zeros((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
if tail_seconds is not None and (
|
||||
not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0):
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
flush = self.finish_output_capacity(tail_seconds)
|
||||
pieces = []
|
||||
remaining = flush
|
||||
while remaining > 0:
|
||||
chunk = min(BLOCK_SAMPLES, remaining)
|
||||
output = np.empty((chunk, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
count = self._lib.ejoc_sofa_binaural_finish(
|
||||
self._handle, chunk, self._f64_pointer(output), chunk)
|
||||
if count < 0:
|
||||
self._raise("finish", count)
|
||||
pieces.append(output[:count])
|
||||
remaining -= chunk
|
||||
self.finished = True
|
||||
nonempty = [piece for piece in pieces if len(piece)]
|
||||
if not nonempty:
|
||||
return np.zeros((0, OUTPUT_CHANNELS), dtype=np.float64)
|
||||
return np.concatenate(nonempty, axis=0)
|
||||
|
||||
def finish_output_capacity(self, tail_seconds: float | None = None) -> int:
|
||||
if tail_seconds is not None and (
|
||||
not math.isfinite(float(tail_seconds)) or float(tail_seconds) < 0.0):
|
||||
raise ValueError("tail_seconds must be finite and non-negative")
|
||||
requested = (self._fdn_tail_samples if tail_seconds is None
|
||||
else int(math.ceil(float(tail_seconds) * 48000.0)))
|
||||
maximum_hrtf = float(np.max(self.field.delay_bounds[:, 1], initial=0.0))
|
||||
hrtf_slots = int(math.ceil(maximum_hrtf / QMF_HOP))
|
||||
hrtf_bound = hrtf_slots * QMF_HOP
|
||||
early_bound = hrtf_bound + 2048
|
||||
if self.enable_early_reflections:
|
||||
early_bound += 256 * QMF_HOP
|
||||
drain = max(requested if self.enable_late_room else 0, early_bound)
|
||||
drain += LATENCY_SAMPLES
|
||||
return int(math.ceil(drain / QMF_HOP) * QMF_HOP)
|
||||
|
||||
def reset(self) -> None:
|
||||
if self._lib.ejoc_sofa_binaural_reset(self._handle):
|
||||
self._raise("reset", -1)
|
||||
self.finished = False
|
||||
for source in range(self.source_count):
|
||||
self.set_source(
|
||||
source, self.positions[source], profile=self.profiles[source],
|
||||
gain=float(self.user_gain[source]),
|
||||
special_lfe=bool(self.special_lfe[source]), fade=False)
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"name": "NativeSofaBinauralDsp",
|
||||
"source_count": self.source_count,
|
||||
"sample_rate_hz": 48000.0,
|
||||
"precision": "float64/complex128",
|
||||
"signal_path": (
|
||||
"native 64-QMF -> native 77-hybrid -> SOFA order-5 real-SH "
|
||||
"direct/early -> native synthesis + shared unitary FDN"),
|
||||
"hrtf_input_kind": self.hrtf_input_kind,
|
||||
"hrtf_input_path": self.hrtf_input_path,
|
||||
"cache_policy": self.cache_policy,
|
||||
"latency_compensated_samples": LATENCY_SAMPLES,
|
||||
"enable_early_reflections": self.enable_early_reflections,
|
||||
"enable_late_room": self.enable_late_room,
|
||||
"early_history_slots": 256,
|
||||
"hrtf_history_slots": int(math.ceil(
|
||||
float(np.max(self.field.delay_bounds[:, 1], initial=0.0)) / QMF_HOP)),
|
||||
"maximum_hrtf_delay_samples": float(
|
||||
np.max(self.field.delay_bounds[:, 1], initial=0.0)),
|
||||
"parameter_updates": self.parameter_updates,
|
||||
"distance": ReferenceDistanceProfileV1.info(),
|
||||
"field": self.field.info(),
|
||||
"library_path": str(self.library_path),
|
||||
"native_backend": True,
|
||||
}
|
||||
|
||||
def close(self) -> None:
|
||||
handle = getattr(self, "_handle", None)
|
||||
if handle:
|
||||
self._lib.ejoc_sofa_binaural_destroy(handle)
|
||||
self._handle = None
|
||||
self.finished = True
|
||||
|
||||
|
||||
def create_native_sofa_renderer(
|
||||
sofa, *, mode="mid", cache_policy="memory", cache_dir=None,
|
||||
shell_radius_m=1.0, object_delay_samples=1473, tail_seconds=5.0,
|
||||
output_gain=1.0, chunk_frames=64):
|
||||
"""Compile a SOFA source and build a JOC adapter over the native DSP."""
|
||||
from binaural_renderer import SofaBinauralRenderer, resolve_sofa_hrtf
|
||||
|
||||
source = resolve_sofa_hrtf(sofa)
|
||||
field = compile_sofa_hrtf(
|
||||
source,
|
||||
shell_radius_m=shell_radius_m,
|
||||
cache_policy=cache_policy,
|
||||
cache_dir=cache_dir)
|
||||
backend = NativeSofaBinauralDsp(field, default_profile=mode)
|
||||
backend.hrtf_input_kind = "sofa"
|
||||
backend.hrtf_input_path = str(source)
|
||||
backend.cache_policy = str(cache_policy).lower()
|
||||
return SofaBinauralRenderer(
|
||||
backend, mode=mode, object_delay_samples=object_delay_samples,
|
||||
tail_seconds=tail_seconds, chunk_frames=chunk_frames)
|
||||
|
||||
|
||||
def create_native_compiled_cache_renderer(
|
||||
cache, *, mode="mid", object_delay_samples=1473, tail_seconds=5.0,
|
||||
output_gain=1.0, chunk_frames=64):
|
||||
"""Load a compiled cache and build a JOC adapter over the native DSP."""
|
||||
from binaural_renderer import SofaBinauralRenderer, resolve_compiled_hrtf_cache
|
||||
|
||||
source = resolve_compiled_hrtf_cache(cache)
|
||||
field = SofaHrtfField.load(source)
|
||||
backend = NativeSofaBinauralDsp(field, default_profile=mode)
|
||||
backend.hrtf_input_kind = "compiled_cache"
|
||||
backend.hrtf_input_path = str(source)
|
||||
backend.cache_policy = None
|
||||
return SofaBinauralRenderer(
|
||||
backend, mode=mode, object_delay_samples=object_delay_samples,
|
||||
tail_seconds=tail_seconds, chunk_frames=chunk_frames)
|
||||
+103
-29
@@ -1,6 +1,7 @@
|
||||
"""Streaming spool and WAV writer for direct speaker-layout output."""
|
||||
"""Shared PCM spool, peak analysis, and WAV writer for direct outputs."""
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
import struct
|
||||
from pathlib import Path
|
||||
|
||||
@@ -13,48 +14,115 @@ _PCM_GUID = bytes.fromhex("0100000000001000800000aa00389b71")
|
||||
_FLOAT_GUID = bytes.fromhex("0300000000001000800000aa00389b71")
|
||||
|
||||
|
||||
class SpeakerPcmSpool:
|
||||
"""Temporary interleaved float32 store with float64 peak analysis."""
|
||||
class PcmSpool:
|
||||
"""Temporary interleaved PCM store with float64 peak and clipping analysis.
|
||||
|
||||
def __init__(self, path, sample_count, channel_count):
|
||||
``storage_dtype`` controls only the temporary representation. Speaker
|
||||
output keeps its historical float32 spool, while binaural uses float64 so
|
||||
precision is reduced only by the selected final WAV format.
|
||||
"""
|
||||
|
||||
def __init__(self, path, sample_capacity, channel_count, *,
|
||||
expected_samples=None, storage_dtype="<f4",
|
||||
tail_threshold=None):
|
||||
self.path = Path(path)
|
||||
self.sample_count = int(sample_count)
|
||||
self.sample_capacity = int(sample_capacity)
|
||||
self.sample_count = int(
|
||||
self.sample_capacity if expected_samples is None else expected_samples)
|
||||
self.expected_samples = (
|
||||
None if expected_samples is None else int(expected_samples))
|
||||
self.channel_count = int(channel_count)
|
||||
self.storage_dtype = np.dtype(storage_dtype)
|
||||
self.tail_threshold = (
|
||||
None if tail_threshold is None else float(tail_threshold))
|
||||
if self.sample_capacity < 0 or self.channel_count <= 0:
|
||||
raise ValueError("invalid PCM spool dimensions")
|
||||
if self.expected_samples is not None and not (
|
||||
0 <= self.expected_samples <= self.sample_capacity):
|
||||
raise ValueError("expected_samples exceeds sample_capacity")
|
||||
if self.tail_threshold is not None and (
|
||||
not math.isfinite(self.tail_threshold) or self.tail_threshold < 0.0):
|
||||
raise ValueError("tail_threshold must be finite and non-negative")
|
||||
self.position = 0
|
||||
self.peak = 0.0
|
||||
self.clipped_values = 0
|
||||
self.values = np.memmap(
|
||||
self.path, dtype="<f4", mode="w+",
|
||||
shape=(self.sample_count, self.channel_count),
|
||||
self.last_above_threshold = -1
|
||||
self.kept_samples = None
|
||||
self._values = np.memmap(
|
||||
self.path, dtype=self.storage_dtype, mode="w+",
|
||||
shape=(self.sample_capacity, self.channel_count),
|
||||
)
|
||||
|
||||
@property
|
||||
def values(self):
|
||||
if self._values is None:
|
||||
raise RuntimeError("PCM spool is closed")
|
||||
length = (self.position if self.kept_samples is None
|
||||
else self.kept_samples)
|
||||
return self._values[:length]
|
||||
|
||||
def write_frame(self, pcm):
|
||||
values = np.asarray(pcm, dtype=np.float64)
|
||||
if values.ndim != 2 or values.shape[1] != self.channel_count:
|
||||
raise ValueError(
|
||||
f"speaker frame must have shape [samples,{self.channel_count}], got {values.shape}")
|
||||
if self.position + len(values) > self.sample_count:
|
||||
raise ValueError("speaker spool received more samples than allocated")
|
||||
f"PCM frame must have shape [samples,{self.channel_count}], got {values.shape}")
|
||||
if self.position + len(values) > self.sample_capacity:
|
||||
raise ValueError("PCM spool received more samples than allocated")
|
||||
if not np.all(np.isfinite(values)):
|
||||
raise ValueError("speaker renderer produced NaN or infinity")
|
||||
raise ValueError("renderer produced NaN or infinity")
|
||||
absolute = np.abs(values)
|
||||
if absolute.size:
|
||||
self.peak = max(self.peak, float(np.max(absolute)))
|
||||
self.clipped_values += int(np.count_nonzero(absolute > 1.0))
|
||||
self.values[self.position:self.position + len(values)] = values.astype(np.float32)
|
||||
if self.tail_threshold is not None:
|
||||
per_sample = np.max(absolute, axis=1)
|
||||
above = np.flatnonzero(per_sample > self.tail_threshold)
|
||||
if above.size:
|
||||
self.last_above_threshold = self.position + int(above[-1])
|
||||
self._values[self.position:self.position + len(values)] = values.astype(
|
||||
self.storage_dtype, copy=False)
|
||||
self.position += len(values)
|
||||
|
||||
def finalize(self):
|
||||
if self.position != self.sample_count:
|
||||
def finalize(self, *, minimum_samples=0):
|
||||
if (self.expected_samples is not None
|
||||
and self.position != self.expected_samples):
|
||||
raise ValueError(
|
||||
f"speaker spool has {self.position} samples, expected {self.sample_count}")
|
||||
self.values.flush()
|
||||
f"PCM spool has {self.position} samples, expected {self.expected_samples}")
|
||||
keep = self.position
|
||||
if self.tail_threshold is not None:
|
||||
keep = min(
|
||||
self.position,
|
||||
max(int(minimum_samples), self.last_above_threshold + 1),
|
||||
)
|
||||
self.kept_samples = keep
|
||||
self.sample_count = keep
|
||||
self._values.flush()
|
||||
return self
|
||||
|
||||
def close(self):
|
||||
values = self.values
|
||||
self.values = None
|
||||
del values
|
||||
values = self._values
|
||||
self._values = None
|
||||
if values is not None:
|
||||
del values
|
||||
|
||||
|
||||
class SpeakerPcmSpool(PcmSpool):
|
||||
"""Backward-compatible fixed-length float32 speaker spool."""
|
||||
|
||||
def __init__(self, path, sample_count, channel_count):
|
||||
super().__init__(
|
||||
path, sample_count, channel_count,
|
||||
expected_samples=sample_count, storage_dtype="<f4")
|
||||
|
||||
|
||||
class BinauralPcmSpool(PcmSpool):
|
||||
"""Float64 variable-tail spool for the binaural renderer."""
|
||||
|
||||
def __init__(self, path, sample_capacity, *, tail_threshold=1.0e-8):
|
||||
super().__init__(
|
||||
path, sample_capacity, 2,
|
||||
expected_samples=None, storage_dtype="<f8",
|
||||
tail_threshold=tail_threshold)
|
||||
|
||||
|
||||
def _fmt_chunk(channel_count, rate, sample_format):
|
||||
@@ -69,7 +137,7 @@ def _fmt_chunk(channel_count, rate, sample_format):
|
||||
simple_tag = WAVE_FORMAT_PCM
|
||||
guid = _PCM_GUID
|
||||
else:
|
||||
raise ValueError(f"unsupported speaker WAV format: {sample_format}")
|
||||
raise ValueError(f"unsupported WAV format: {sample_format}")
|
||||
block_align = channel_count * bytes_per_sample
|
||||
byte_rate = rate * block_align
|
||||
if channel_count <= 2:
|
||||
@@ -94,13 +162,13 @@ def _write_header(stream, channel_count, sample_count, rate, sample_format):
|
||||
riff_file_size = 12 + 8 + len(fmt) + 8 + data_size
|
||||
use_rf64 = riff_file_size - 8 > 0xFFFFFFFF
|
||||
if use_rf64:
|
||||
# RF64 + ds64 + fmt + data.
|
||||
file_size = 12 + 36 + 8 + len(fmt) + 8 + data_size
|
||||
stream.write(b"RF64")
|
||||
stream.write(struct.pack("<I", 0xFFFFFFFF))
|
||||
stream.write(b"WAVE")
|
||||
stream.write(b"ds64")
|
||||
stream.write(struct.pack("<IQQQI", 28, file_size - 8, data_size, sample_count, 0))
|
||||
stream.write(struct.pack(
|
||||
"<IQQQI", 28, file_size - 8, data_size, sample_count, 0))
|
||||
else:
|
||||
stream.write(b"RIFF")
|
||||
stream.write(struct.pack("<I", riff_file_size - 8))
|
||||
@@ -121,7 +189,8 @@ def _write_header(stream, channel_count, sample_count, rate, sample_format):
|
||||
|
||||
|
||||
def _pack_int24(values):
|
||||
scaled = (np.clip(values, -1.0, 1.0) * np.float32(8388607.0)).astype(np.int32)
|
||||
source = np.asarray(values)
|
||||
scaled = (np.clip(source, -1.0, 1.0) * np.float32(8388607.0)).astype(np.int32)
|
||||
unsigned = scaled.reshape(-1).view(np.uint32)
|
||||
packed = np.empty((unsigned.size, 3), dtype=np.uint8)
|
||||
packed[:, 0] = unsigned & 0xFF
|
||||
@@ -130,21 +199,22 @@ def _pack_int24(values):
|
||||
return packed.tobytes()
|
||||
|
||||
|
||||
def write_speaker_wav(path, pcm, sample_format, *, rate=48000, chunk_samples=262144):
|
||||
"""Write an interleaved float32 array/memmap as float32 or PCM24 WAV."""
|
||||
def write_pcm_wav(path, pcm, sample_format, *, rate=48000,
|
||||
chunk_samples=262144):
|
||||
"""Write an interleaved array/memmap as float32 or PCM24 WAV."""
|
||||
target = Path(path)
|
||||
values = np.asarray(pcm)
|
||||
if values.ndim != 2:
|
||||
raise ValueError(f"speaker PCM must be 2D, got {values.shape}")
|
||||
raise ValueError(f"PCM must be 2D, got {values.shape}")
|
||||
sample_count, channel_count = values.shape
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
with target.open("wb") as stream:
|
||||
info = _write_header(
|
||||
stream, channel_count, sample_count, int(rate), sample_format)
|
||||
for start in range(0, sample_count, int(chunk_samples)):
|
||||
block = np.asarray(values[start:start + chunk_samples], dtype="<f4")
|
||||
block = np.asarray(values[start:start + chunk_samples])
|
||||
if sample_format == "float32":
|
||||
stream.write(block.tobytes(order="C"))
|
||||
stream.write(block.astype("<f4", copy=False).tobytes(order="C"))
|
||||
else:
|
||||
stream.write(_pack_int24(block))
|
||||
info.update({
|
||||
@@ -155,3 +225,7 @@ def write_speaker_wav(path, pcm, sample_format, *, rate=48000, chunk_samples=262
|
||||
"file_bytes": target.stat().st_size,
|
||||
})
|
||||
return info
|
||||
|
||||
|
||||
# Existing imports remain valid.
|
||||
write_speaker_wav = write_pcm_wav
|
||||
|
||||
@@ -0,0 +1,144 @@
|
||||
"""Orthonormal real spherical harmonics in ACN order, through fifth order."""
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
import numpy as np
|
||||
from scipy.spatial import SphericalVoronoi
|
||||
|
||||
|
||||
def _associated_legendre(order: int, degree: int, x: np.ndarray) -> np.ndarray:
|
||||
"""P_degree^order(x), including the Condon-Shortley phase."""
|
||||
m = int(order)
|
||||
l = int(degree)
|
||||
if not 0 <= m <= l:
|
||||
raise ValueError("associated Legendre indices require 0 <= m <= l")
|
||||
x = np.asarray(x, dtype=np.float64)
|
||||
p_mm = np.ones_like(x)
|
||||
if m:
|
||||
double_factorial = 1.0
|
||||
for value in range(1, 2 * m, 2):
|
||||
double_factorial *= value
|
||||
p_mm = ((-1.0) ** m) * double_factorial * np.power(
|
||||
np.maximum(0.0, 1.0 - x * x), 0.5 * m)
|
||||
if l == m:
|
||||
return p_mm
|
||||
p_m1 = x * (2 * m + 1) * p_mm
|
||||
if l == m + 1:
|
||||
return p_m1
|
||||
previous_previous = p_mm
|
||||
previous = p_m1
|
||||
for current_degree in range(m + 2, l + 1):
|
||||
current = (
|
||||
(2 * current_degree - 1) * x * previous
|
||||
- (current_degree + m - 1) * previous_previous
|
||||
) / float(current_degree - m)
|
||||
previous_previous, previous = previous, current
|
||||
return previous
|
||||
|
||||
|
||||
def real_spherical_harmonics(directions, order: int = 5) -> np.ndarray:
|
||||
"""Return [directions,(order+1)^2] ACN/N3D real harmonics.
|
||||
|
||||
Coordinates use SOFA listener axes: +X front, +Y left, +Z up. The basis is
|
||||
orthonormal over the sphere and includes the Condon-Shortley phase.
|
||||
"""
|
||||
maximum_order = int(order)
|
||||
if not 0 <= maximum_order <= 12:
|
||||
raise ValueError("supported spherical-harmonic orders are 0..12")
|
||||
vectors = np.asarray(directions, dtype=np.float64)
|
||||
one = vectors.ndim == 1
|
||||
if one:
|
||||
vectors = vectors[None, :]
|
||||
if vectors.ndim != 2 or vectors.shape[1] != 3 or not np.isfinite(vectors).all():
|
||||
raise ValueError("directions must have finite shape [M,3]")
|
||||
length = np.linalg.norm(vectors, axis=1)
|
||||
if np.any(length <= 1.0e-15):
|
||||
raise ValueError("spherical-harmonic directions must be non-zero")
|
||||
unit = vectors / length[:, None]
|
||||
azimuth = np.arctan2(unit[:, 1], unit[:, 0])
|
||||
cos_colatitude = np.clip(unit[:, 2], -1.0, 1.0)
|
||||
result = np.empty((len(unit), (maximum_order + 1) ** 2), dtype=np.float64)
|
||||
column = 0
|
||||
for degree in range(maximum_order + 1):
|
||||
for m in range(-degree, degree + 1):
|
||||
absolute = abs(m)
|
||||
normalization = math.sqrt(
|
||||
(2 * degree + 1) / (4.0 * math.pi)
|
||||
* math.factorial(degree - absolute)
|
||||
/ math.factorial(degree + absolute))
|
||||
legendre = _associated_legendre(absolute, degree, cos_colatitude)
|
||||
if m < 0:
|
||||
value = math.sqrt(2.0) * normalization * legendre * np.sin(
|
||||
absolute * azimuth)
|
||||
elif m > 0:
|
||||
value = math.sqrt(2.0) * normalization * legendre * np.cos(
|
||||
m * azimuth)
|
||||
else:
|
||||
value = normalization * legendre
|
||||
result[:, column] = value
|
||||
column += 1
|
||||
return result[0] if one else result
|
||||
|
||||
|
||||
def spherical_voronoi_weights(directions) -> np.ndarray:
|
||||
"""Area weights for an irregular full-sphere grid, with uniform fallback."""
|
||||
vectors = np.asarray(directions, dtype=np.float64)
|
||||
if vectors.ndim != 2 or vectors.shape[1] != 3:
|
||||
raise ValueError("directions must have shape [M,3]")
|
||||
unit = vectors / np.linalg.norm(vectors, axis=1)[:, None]
|
||||
if len(unit) < 4:
|
||||
return np.full(len(unit), 1.0 / len(unit), dtype=np.float64)
|
||||
try:
|
||||
voronoi = SphericalVoronoi(unit, radius=1.0, center=np.zeros(3))
|
||||
areas = np.asarray(voronoi.calculate_areas(), dtype=np.float64)
|
||||
if not np.isfinite(areas).all() or np.any(areas <= 0.0):
|
||||
raise ValueError("invalid spherical Voronoi areas")
|
||||
return areas / np.sum(areas, dtype=np.float64)
|
||||
except (ValueError, RuntimeError, np.linalg.LinAlgError):
|
||||
return np.full(len(unit), 1.0 / len(unit), dtype=np.float64)
|
||||
|
||||
|
||||
def fit_real_spherical_harmonics(directions, values, *, order: int = 5,
|
||||
ridge: float = 1.0e-6,
|
||||
weights=None) -> np.ndarray:
|
||||
"""Weighted ridge fit. Output shape is [terms,...value trailing axes]."""
|
||||
basis = real_spherical_harmonics(directions, order=order)
|
||||
target = np.asarray(values)
|
||||
if target.shape[0] != basis.shape[0]:
|
||||
raise ValueError("spherical-harmonic target count does not match directions")
|
||||
if target.dtype.kind == "c":
|
||||
target = np.asarray(target, dtype=np.complex128)
|
||||
solve_dtype = np.complex128
|
||||
else:
|
||||
target = np.asarray(target, dtype=np.float64)
|
||||
solve_dtype = np.float64
|
||||
if weights is None:
|
||||
weight = spherical_voronoi_weights(directions)
|
||||
else:
|
||||
weight = np.asarray(weights, dtype=np.float64)
|
||||
if weight.shape != (len(basis),) or np.any(weight < 0.0) or not np.isfinite(weight).all():
|
||||
raise ValueError("weights must be finite non-negative [M]")
|
||||
total = float(np.sum(weight))
|
||||
if total <= 0.0:
|
||||
raise ValueError("weights must have positive sum")
|
||||
weight = weight / total
|
||||
flat = target.reshape(len(target), -1)
|
||||
weighted_basis = basis * weight[:, None]
|
||||
gram = basis.T @ weighted_basis
|
||||
regularization = float(ridge)
|
||||
if not math.isfinite(regularization) or regularization < 0.0:
|
||||
raise ValueError("ridge must be finite and non-negative")
|
||||
scale = float(np.trace(gram)) / gram.shape[0]
|
||||
system = gram + np.eye(gram.shape[0], dtype=np.float64) * regularization * scale
|
||||
right = basis.T @ (weight[:, None] * flat)
|
||||
coefficients = np.linalg.solve(system.astype(solve_dtype), right.astype(solve_dtype))
|
||||
return coefficients.reshape((basis.shape[1],) + target.shape[1:])
|
||||
|
||||
|
||||
def evaluate_real_spherical_harmonics(coefficients, directions,
|
||||
*, order: int = 5) -> np.ndarray:
|
||||
basis = real_spherical_harmonics(directions, order=order)
|
||||
coeff = np.asarray(coefficients)
|
||||
if coeff.shape[0] != (int(order) + 1) ** 2:
|
||||
raise ValueError("coefficient term count does not match order")
|
||||
return np.tensordot(basis, coeff, axes=([-1], [0]))
|
||||
Reference in New Issue
Block a user