添加双耳渲染功能

This commit is contained in:
2026-09-06 19:14:24 +08:00
parent 329445ed25
commit 2a296099fa
38 changed files with 9660 additions and 214 deletions
+277
View File
@@ -0,0 +1,277 @@
# Binaural rendering
[中文](binaural.md) · [Back to README](../README.en.md)
JustOneCacophony's binaural backend supports three HRTF sources:
`SimpleFreeFieldHRIR` SOFA, the Rosella `.personalized_headphone` model exported
by Dolby's official personalization scan (its JSON parsing is implemented by
this project and invokes no Dolby software), and the `.jochrtf` cache compiled
from SOFA. SOFA is compiled into an in-memory directional field when the model
is loaded. A `.jochrtf` file is only a disposable, reproducible JOC compiled
HRTF cache; it is neither an interchange format nor a prerequisite for using
SOFA.
```text
SOFA FIR
-> CanonicalHrtf
-> 48 kHz / one radius shell / delay-phase policy
-> 64-QMF / 77-hybrid projection
-> fifth-order ACN/N3D real-SH field
-> per-object direct + early reflections
-> shared unitary-FDN late room
-> float64 stereo
```
## Inputs
The CLI has three mutually exclusive HRTF input sources; with none given, a
default rule resolves the input:
```powershell
# 1) SOFA: defaults to HRTF/binaural.sofa, or an explicit path
python main.py input.m4a --binaural
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
# 2) Rosella .personalized_headphone: defaults to HRTF/binaural.personalized_headphone
python main.py input.m4a --binaural --personalized-headphone
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
# 3) .jochrtf: explicitly load a compiled cache
python main.py input.m4a --binaural `
--compiled-hrtf-cache C:\HRTF\subject.jochrtf
# Optional: create/reuse a transparent disk cache for SOFA
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
--hrtf-cache-policy disk
```
The default order is `HRTF/binaural.sofa`, then the unique `.jochrtf` under
`output/hrtf-cache`, then `HRTF/binaural.personalized_headphone`; if none of
the three exist, an error asks for an explicit path. Multiple `.jochrtf` files
under `output/hrtf-cache` are also an error requiring an explicit choice.
The `.personalized_headphone` JSON parsing is implemented by this project
(`src/rosella_model.py`) and does not invoke any Dolby software.
`--hrtf-cache-policy` accepts `none`, `memory`, or `disk`. The default is
`memory`; neither `none` nor `memory` creates a file. `disk` writes to
`output/hrtf-cache` by default, or to `--hrtf-cache-dir`. `--hrtf-radius-m`
selects the nearest measurement-radius shell.
The Python API also uses explicit factories:
```python
from sofa_binaural_backend import SofaBinauralBackend
renderer = SofaBinauralBackend.from_sofa(
"subject.sofa",
source_count=16,
default_profile="mid",
cache_policy="memory",
)
cached = SofaBinauralBackend.from_compiled_cache(
"subject.jochrtf",
source_count=16,
default_profile="mid",
)
```
The factories never guess a format from an unknown suffix: SOFA and `.jochrtf`
always use distinct loaders.
## Binaural render mode
`--binaural-mode off|near|mid|far` (default `mid`) is a **human-specified
rendering hint**, not original binaural metadata extracted or recovered from the
input E-AC-3 JOC bitstream:
- Direct binaural rendering (`--binaural`): near/mid/far apply, default `mid`;
`off` is an error;
- ADM BWF: the low 3 binaural-render-mode bits of the last 15 JOC object entries
in DBMD segment 10 carry `off=0/near=1/far=2/mid=3`, leaving the first 10 bed
entries unchanged; the default is `mid`, and `off` explicitly disables the
binaural metadata hint.
## Canonical SOFA contract
The strict importer currently accepts:
- `Conventions=SOFA`;
- `SOFAConventions=SimpleFreeFieldHRIR`, version `0.4`, `1.0`, or `1.1`;
- `DataType=FIR` and `Data.IR[M,2,N]`;
- one positive finite `Data.SamplingRate` in hertz/Hz;
- spherical or Cartesian `SourcePosition`;
- singleton or per-measurement `ListenerPosition/View/Up`;
- two receivers whose listener-local lateral geometry uniquely identifies L/R;
- one zero-offset emitter;
- causal `Data.Delay[I,2]` or `[M,2]`;
- an explicitly free-field/anechoic `RoomType`.
Receiver order comes from geometry, never from the receiver array index. SOFA
listener coordinates are $+X$ front, $+Y$ left, $+Z$ up; ADM coordinates are
$+X$ right, $+Y$ front, $+Z$ up:
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
`CanonicalHrtf` keeps `Data.IR` and `Data.Delay` separate. Only a time-domain
baseline calls `materialized_measurement()` to apply delay once; the runtime SH
path never materializes and then restores the delay. Non-48-kHz HRIRs are
normalized with float64 `scipy.signal.resample_poly`, and delay samples scale by
the same ratio.
GeneralFIR, BRIR, TF, multiple emitters, ambiguous receivers, and non-free-field
data require convention-specific adapters. They cannot enter the core importer
through a reshape.
## Exactly-once delay and phase
The compiler recognizes three mutually exclusive representations:
1. Nonzero `Data.Delay` is external to `Data.IR`; the FIR is not de-rotated and
runtime applies the delay once.
2. With `Data.Delay=0` and an ordinary positive-onset HRIR, each ear's main peak
supplies arrival time. Compilation separates it and runtime restores it once.
The current threshold is a peak index greater than two samples.
3. With `Data.Delay=0` and both FIRs at a shared sample-zero origin, no external
delay is invented. The authored complex phase stays in the fifth-order field.
No path may add a second ear delay or phase-group delay.
## Public filterbank and directional field
The runtime is fixed at:
- 48 kHz;
- a 64-sample QMF hop;
- 64-QMF / 77 hybrid bands;
- 961 samples of analysis/synthesis latency;
- fifth order, 36 terms, ACN/N3D real spherical harmonics;
- float64 PCM, delay, SH, and room state; complex128 band transfers and spectra.
Real and imaginary unit gains for every hybrid band pass through the same
analysis/synthesis chain to form a 154-real-parameter impulse dictionary. The
compiler does not sample 77 FFT bins. Defaults are `1e-3` projection ridge and
`1e-5` SH ridge. Coincident directions are merged before a spherical-Voronoi
weighted ridge fit.
The fixed resource is `data/rosella_kernels.npz`, which implements publicly
standardized filter banks, computable from the following formulas.
The hybrid analysis kernels are defined in [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf),
Section 5.2.2 (Table 1 $Q=8$/$Q=4$ coefficients, delay 6):
$$G_q^p[n] = g^p[n]\cdot\exp\!\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
The QMF analysis table is the MPEG-4 AAC/SBR 64 complex QMF bank of
ISO/IEC 14496-3/AMD1:2003, subclause 4.B.18.2, stored as the polyphase
reordering of the public 640-tap prototype $c_0,\dots,c_{639}$:
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
The QMF synthesis table is the causal left inverse of the analysis polyphase
matrix $\mathbf{A}$, i.e. the solution of $\mathbf{A}\,\mathbf{W}=\mathbf{P}$
($\mathbf{P}$ is the 577-sample delay permutation; total latency
$961 = 577 + 6\times64$), stored as a rank-4 factorization:
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
The hybrid synthesis table is the 77→64 recombination: identity for the high
bands, $Y_{3+b}=X_{16+b}$, and for the low bands ($C_p$ is the $8+4+4$ child
partition):
$$Y_p = \sum_{q\in C_p}\Bigl(\operatorname{Re}X_q + j\,s_q\,\operatorname{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
The loader verifies the archive and every array by SHA-256; the table version,
all array hashes, and the 77 reference band-center values are part of the cache
key. Public availability of a standard does not by itself grant permission to
practice related patent claims. See
[`data/README.en.md`](../data/README.en.md) and
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the sources and the
rights boundary.
## `.jochrtf`
A `.jochrtf` file is a pickle-free compressed NumPy archive with an exact member set:
| key | dtype / shape |
|---|---|
| `metadata_json` | NumPy Unicode scalar containing JSON text (`dtype.kind == "U"`) |
| `band_center_frequencies_hz` | little-endian `float64[77]` |
| `coefficients` | little-endian `complex128[36,2,77]` |
| `delay_coefficients` | little-endian `float64[36,2]` |
| `delay_bounds` | little-endian `float64[2,2]` |
Metadata uses the `JOC-HRTF-CACHE` magic and records the schema, compiler and
phase-policy versions, ACN/N3D convention, filterbank hashes, SOFA content
SHA-256, sample rate, radius, order, both ridge values, payload hash, and fit
report. Every setting that changes compilation participates in the cache key.
Metadata never persists an absolute local `source_path`; it may keep a display
name only.
Before constructing a field, the loader uses `allow_pickle=False` and validates
ZIP members and expanded sizes, shapes, dtypes, byte order, contiguous layout,
finite values, delay bounds, band centers, payload hash, and cache key. The
writer uses a same-directory temporary file, `fsync`, a process-held OS file
lock, and atomic `os.replace`. Its hidden `.lock` sidecar may remain and does not
mean that a writer still owns the lock. Outdated, damaged, or mismatched
caches cannot hit. SOFA input rebuilds an invalid cache; an explicitly selected
cache reports the error.
Deleting a disk cache must not change the field or render produced from the same
SOFA and compiler configuration.
A `.jochrtf` file contains directional-field coefficients and delay data
transformed from the source HRIRs. Its reproducibility therefore does not make
it licence-free. Creating a cache does not enlarge the rights granted by the
source SOFA/HRTF dataset: use, copying, and redistribution remain subject to
that dataset's terms. If those terms are unclear, keep `.jochrtf` as a private
local cache and do not ship it with the program or another build artifact.
`source_sha256` is only a content-integrity identifier, not proof of provenance
or permission.
## JOC objects and room behavior
The production adapter retains the existing JOC schedule:
- `[1536,16]` input per frame;
- channel 0 is special LFE and channels 1..15 are JOC objects;
- ID11/OAMD positions use a sample-timed timeline;
- source parameters update every 512 samples;
- every object owns independent direct/early history while one late FDN is shared;
- `finish()` drains early/late tails; output gain is explicit, with no implicit
limiter or programme loudness normalization.
Near/Mid/Far, equal-power direct level, six first-order shoebox image sources,
late sends, the unitary FDN, the 120–180 Hz cosine-squared LFE low-pass, and room
calibration are JOC project-defined behavior, not constants published by SOFA or
Dolby.
The public SOFA binaural renderer defaults to the C++20 native core under
`--backend auto/native` (`ejoc_sofa_binaural_*` in `lib/eac3joc_core.dll`): the
filterbank, the SH direction-field evaluation, the per-object direct/early
histories and the shared FDN all run natively, while Python only compiles the
SOFA source and issues the per-512-sample metadata updates. When the native
library is unavailable the renderer falls back to the Python/NumPy reference
implementation; the two agree to better than 1e-9. `--backend python` forces
the Python backend.
`--backend` still selects native/Python JOC reconstruction and speaker rendering;
native acceleration for the public binaural DSP is outside the current API.
## Technical references and rights boundary
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
- [3GPP TS 26.405 / ETSI TS 126 405 (64-QMF/77-hybrid definition)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
- [Dolby binaural render-mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en), used only as
architectural background for direct/early/late, subbands, and FDNs; it does
not establish that any product uses a particular embodiment.
Public availability of a specification, source file, or patent document does
not by itself authorize copying its contents, redistribution of derivatives,
or practice of patent claims. These technical references grant no patent
licence and make no non-infringement representation. Anyone preparing a release
or product integration must assess the applicable data and software licences,
patent permissions, and freedom to operate. See
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the public-standard
provenance and rights boundary.
+238
View File
@@ -0,0 +1,238 @@
# 双耳渲染
[English](binaural.en.md) · [返回 README](../README.md)
JustOneCacophony 的双耳后端支持三种 HRTF 来源:`SimpleFreeFieldHRIR` SOFA、
杜比官方软件个性化扫描导出的 Rosella `.personalized_headphone`(JSON 解析由本项目
自行实现,不调用杜比软件),以及从 SOFA 编译出的 `.jochrtf` 缓存。SOFA 在模型加载
时编译成内存方向场;`.jochrtf` 只是可删除、可重建的 JOC compiled HRTF cache,
不是交换格式,也不是使用 SOFA 的前置步骤。
```text
SOFA FIR
-> CanonicalHrtf
-> 48 kHz / 单 radius shell / delay-phase policy
-> 64-QMF / 77-hybrid projection
-> 五阶 ACN/N3D 实球谐场
-> 逐对象 direct + early reflections
-> shared unitary-FDN late room
-> stereo float64
```
## 输入接口
CLI 有三个互斥的 HRTF 输入来源;都不指定时按默认规则自动选择:
```powershell
# 1) SOFA:缺省取 HRTF/binaural.sofa,也可显式指定
python main.py input.m4a --binaural
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
# 2) Rosella .personalized_headphone:缺省取 HRTF/binaural.personalized_headphone
python main.py input.m4a --binaural --personalized-headphone
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
# 3) .jochrtf:显式读取预编译 cache
python main.py input.m4a --binaural `
--compiled-hrtf-cache C:\HRTF\subject.jochrtf
# 可选:SOFA 透明生成/复用磁盘 cache
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
--hrtf-cache-policy disk
```
默认选择顺序:`HRTF/binaural.sofa` → `output/hrtf-cache` 下唯一的 `.jochrtf` →
`HRTF/binaural.personalized_headphone`;三者都没有时报错并提示显式指定。
`output/hrtf-cache` 下有多个 `.jochrtf` 时同样报错,要求显式选择。
`.personalized_headphone` 的 JSON 解析由本项目自行实现(`src/rosella_model.py`),
不调用任何杜比软件。
`--hrtf-cache-policy` 可取 `none`、`memory`、`disk`。默认是 `memory`;`none` 和
`memory` 都不会创建磁盘文件。`disk` 默认写入 `output/hrtf-cache`,也可用
`--hrtf-cache-dir` 指定。`--hrtf-radius-m` 选择距离目标最近的 measurement shell。
Python API 使用显式 factory:
```python
from sofa_binaural_backend import SofaBinauralBackend
renderer = SofaBinauralBackend.from_sofa(
"subject.sofa",
source_count=16,
default_profile="mid",
cache_policy="memory",
)
cached = SofaBinauralBackend.from_compiled_cache(
"subject.jochrtf",
source_count=16,
default_profile="mid",
)
```
文件工厂不会按“未知后缀”猜格式:SOFA 和 `.jochrtf` 始终走不同 loader。
## 双耳渲染模式
`--binaural-mode off|near|mid|far`(默认 `mid`)是**人为指定的渲染提示**,不是
从输入 E-AC-3 JOC 码流提取或还原的原始双耳元数据:
- 直接双耳渲染(`--binaural`):near/mid/far 生效,默认 `mid`;`off` 报错;
- ADM BWF:DBMD segment 10 中后 15 个 JOC 对象的 binaural render mode 写
`off=0/near=1/far=2/mid=3`,前 10 个 bed 保持不变,默认 `mid`;`off` 用于显式
关闭双耳元数据提示。
## Canonical SOFA 契约
当前 strict importer 接受:
- `Conventions=SOFA`;
- `SOFAConventions=SimpleFreeFieldHRIR`,version `0.4`、`1.0` 或 `1.1`;
- `DataType=FIR`,`Data.IR[M,2,N]`;
- 单一正有限 `Data.SamplingRate`,单位为 hertz/Hz;
- spherical 或 Cartesian `SourcePosition`;
- 单值或 per-measurement 的 `ListenerPosition/View/Up`;
- 两个能由 listener-local lateral 坐标唯一识别左右的 receiver;
- 单一且零偏移的 emitter;
- causal `Data.Delay[I,2]` 或 `[M,2]`;
- 明确的 free-field/anechoic `RoomType`。
receiver 左右顺序由几何决定,不能假定 `Data.IR` 的 receiver index。SOFA listener
坐标为 $+X$ front、$+Y$ left、$+Z$ up;ADM 坐标为 $+X$ right、$+Y$ front、
$+Z$ up,转换为:
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
`CanonicalHrtf` 将 `Data.IR` 与 `Data.Delay` 分开保存。只有时域 baseline 才调用
`materialized_measurement()` 将 delay 应用一次;运行时 SH 路径不先 materialize。
非 48 kHz HRIR 使用 float64 `scipy.signal.resample_poly` 规范化,delay samples 按
相同比例缩放。
GeneralFIR、BRIR、TF、多 emitter、多义 receiver 或非 free-field 数据需要单独的
convention adapter,不能只通过 reshape 进入核心 importer。
## Delay/phase:exactly once
编译器只允许三种互斥语义:
1. 非零 `Data.Delay` 是 `Data.IR` 外部 delay;FIR 不去旋,运行时应用一次。
2. `Data.Delay=0` 且 HRIR 有普通正 onset:以每耳 main peak 分离 arrival,拟合后
在运行时恢复一次;当前阈值为 peak index 大于 2 samples。
3. `Data.Delay=0` 且双耳 FIR 共享 sample-0 起点:不发明外部 delay,原 complex
phase 直接进入五阶场。
任何路径都不能再叠加第二套 ear delay 或 phase-group delay。
## 公开 filterbank 与方向场
运行时固定为:
- 48 kHz;
- 64-sample QMF hop;
- 64-QMF / 77-hybrid;
- analysis/synthesis latency 961 samples;
- 五阶、36 项、ACN/N3D real spherical harmonics;
- PCM、delay、SH、room state 为 float64;频带传递和频域状态为 complex128。
每个 hybrid band 的 real/imaginary 单位增益都通过同一套 analysis/synthesis 链生成
脉冲字典,共 154 个实参数;编译不是直接读取 77 个 FFT bin。默认 projection
ridge 为 `1e-3`,SH ridge 为 `1e-5`。同方向 measurement 先合并,再用球面 Voronoi
面积权重做 ridge fit。
固定表位于 `data/rosella_kernels.npz`,实现公开标准化的滤波器组,各表可由如下
公式计算。
hybrid 分析核定义于 [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
第 5.2.2 节(Table 1 的 $Q=8$/$Q=4$ 系数,delay 6):
$$G_q^p[n] = g^p[n]\cdot\exp\!\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
QMF analysis 表即 MPEG-4 AAC/SBR(ISO/IEC 14496-3/AMD1:2003 第 4.B.18.2 节)
的 64 complex QMF bank;打包的 $64\times10$ 表是公开 640-tap prototype
$c_0,\dots,c_{639}$ 的多相重排:
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
QMF synthesis 表为上述 analysis 多相矩阵 $\mathbf{A}$ 的因果左逆,即求解
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$($\mathbf{P}$ 为 577-sample 延迟置换;
全链 $961 = 577 + 6\times64$),以 rank-4 分解形式存储:
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
hybrid synthesis 表为 77→64 重组:高频带恒等 $Y_{3+b}=X_{16+b}$;低频带
($C_p$ 为 $8+4+4$ 子带划分):
$$Y_p = \sum_{q\in C_p}\Bigl(\operatorname{Re}X_q + j\,s_q\,\operatorname{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
loader 校验 archive 和每个数组的 SHA-256;table version、所有数组 hash 与
77 个 band-center 参考值都属于 cache key。标准可公开获取不等于获准实施相关
专利;更多来源信息见 [`data/README.md`](../data/README.md) 与
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
## `.jochrtf`
`.jochrtf` 是无 pickle 的压缩 NumPy archive,固定包含:
| key | dtype / shape |
|---|---|
| `metadata_json` | 含 JSON 文本的 NumPy Unicode scalar(`dtype.kind == "U"`) |
| `band_center_frequencies_hz` | little-endian `float64[77]` |
| `coefficients` | little-endian `complex128[36,2,77]` |
| `delay_coefficients` | little-endian `float64[36,2]` |
| `delay_bounds` | little-endian `float64[2,2]` |
metadata magic 固定为 `JOC-HRTF-CACHE`,并记录 schema/compiler/phase-policy、
ACN/N3D、filterbank table hashes、SOFA content SHA-256、采样率、radius、order、
两个 ridge、payload hash 和 fit report。cache key 覆盖所有会改变编译结果的字段。
metadata 不保存本机绝对 `source_path`,仅可保存 source display name。
loader 使用 `allow_pickle=False`,并在构造对象前检查 ZIP 成员集、解压大小、shape、
dtype、端序、连续布局、有限值、delay bounds、band centers、payload hash 和 cache
key。writer 使用同目录临时文件、`fsync`、进程持有的 OS 文件锁和原子
`os.replace`;对应的隐藏 `.lock` sidecar 可保留,但不代表仍有 writer 持锁。
旧版本、损坏或配置不匹配的 cache 不能命中;从 SOFA 启动时会重建,显式 cache
入口则直接报错。
删除磁盘 cache 后,从同一 SOFA 和同一编译配置得到的场与渲染结果不得改变。
`.jochrtf` 包含由源 HRIR 变换得到的方向场系数与 delay 数据,因此“可以重建”不表示
它不受数据许可约束。生成 cache 不会扩大源 SOFA/HRTF 数据集授予的权利;cache 的
使用、复制和再分发仍须遵守源数据集条款。不能确认条款时,应把 `.jochrtf` 作为本地
私有 cache,不随程序或构建产物发布。`source_sha256` 只用于内容一致性校验,不是许可
或来源证明。
## JOC 对象与房间
生产适配器继续使用现有 JOC 调度:
- 每帧输入 `[1536,16]`;
- channel 0 是 special LFE,channel 1..15 是 JOC objects;
- ID11/OAMD position 使用 sample-timed timeline;
- 每 512 samples 更新方向/profile;
- 每个对象拥有独立 direct/early history,late FDN 全局共享;
- `finish()` 排空 early/late tail;输出增益显式应用,不隐含 limiter 或节目响度归一化。
Near/Mid/Far、equal-power direct level、六面 shoebox 一阶 image source、late send、
unitary FDN、LFE 120–180 Hz cosine-squared 低通及 room calibration 都是 JOC
项目定义行为,不是 SOFA 或 Dolby 公布常数。
公开 SOFA 双耳渲染在 `--backend auto/native` 下默认走 C++20 原生核
(`lib/eac3joc_core.dll` 的 `ejoc_sofa_binaural_*` 接口:filterbank、SH 方向场求值、
逐对象 early/direct 历史与共享 FDN 全部在原生侧执行,Python 只做 SOFA 编译与每
512-sample 的元数据更新);原生库不可用时自动回退 Python/NumPy 参考实现,两者
逐值一致(差异 < 1e-9)。`--backend python` 强制使用 Python 后端。
## 技术引用与权利边界
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
- [3GPP TS 26.405 / ETSI TS 126 405(64-QMF/77-hybrid 定义)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
- [Dolby binaural render mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en),仅作 direct/early/late、
subband 与 FDN 架构背景,不证明某个产品使用特定实施例。
规范、源码或专利文献可公开获取,不等于获准复制其内容、再分发派生产物或实施其中的
专利权利要求。本项目的技术引用本身不授予专利许可,也不作不侵权保证;准备发布或集成
到产品的一方应自行审查适用的数据许可、软件许可、专利许可及 freedom-to-operate。
公开标准来源与权利边界见
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
+22 -3
View File
@@ -116,7 +116,26 @@ Supported layouts:
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
```
## 6. Building
## 6. Binaural-rendering ABI
The shared library provides a 512-sample float64 binaural DSP interface:
```c
ejoc_binaural_renderer_handle ejoc_binaural_renderer_create(void);
int ejoc_binaural_renderer_configure_kernels(...);
int ejoc_binaural_renderer_configure_room(...);
int ejoc_binaural_renderer_process(
ejoc_binaural_renderer_handle handle,
const double* input16_interleaved, /* [512][16] */
const double* gains_complex, /* [16][2][77][2] */
const double* room_sends, /* [16] */
double output_gain,
double* output_stereo_interleaved); /* [512][2] */
```
Python parses the model, evaluates the OAMD timeline, and supplies complex gains and room sends every 512 samples. The C++ handle owns QMF, hybrid, recursive-room, and QMF-synthesis state. Inputs, state, accumulation, and output are double/complex double.
## 7. Building
The CMake definition is `native/CMakeLists.txt`. Run from the repository root:
@@ -138,7 +157,7 @@ The MSVC configuration uses the static CRT. Other runtime dependencies depend on
The repository does not include native binaries by default. A prebuilt Release runtime or a locally built runtime can be placed directly under `lib/`.
## 7. Runtime lookup and fallback
## 8. Runtime lookup and fallback
Lookup order:
@@ -148,7 +167,7 @@ Lookup order:
`--backend auto` falls back to NumPy when loading fails, and `--backend python` skips native discovery. The current CLI also prints the failure and falls back for `--backend native`; this existing behavior should not be read as successful native execution.
## 8. Implementation boundaries
## 9. Implementation boundaries
- The native layer accepts only dense-JOC data already parsed by Python.
- The ABI fixes a 1536-sample JOC frame, at most 15 objects, at most 23 parameter bands, and at most 2 data points.
+35 -3
View File
@@ -116,7 +116,39 @@ int ejoc_speaker_renderer_process(
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
```
## 6. 构建
## 6. 双耳渲染 ABI
共享库提供 512-sample float64 双耳 DSP:
```c
ejoc_binaural_renderer_handle ejoc_binaural_renderer_create(void);
int ejoc_binaural_renderer_configure_kernels(...);
int ejoc_binaural_renderer_configure_room(...);
int ejoc_binaural_renderer_process(
ejoc_binaural_renderer_handle handle,
const double* input16_interleaved, /* [512][16] */
const double* gains_complex, /* [16][2][77][2] */
const double* room_sends, /* [16] */
double output_gain,
double* output_stereo_interleaved); /* [512][2] */
```
Python 负责模型解析、OAMD 时间轴和每 512 samples 的 complex gains/room sends。C++ handle 保存 QMF、hybrid、递归 room 和 QMF synthesis 状态。全部输入、状态、乘加和输出均为 double/complex double。
## 6.1 公开 SOFA 双耳渲染 ABI
共享库同时提供完整的原生 SOFA 双耳渲染器(`ejoc_sofa_binaural_*`),它镜像
Python `SofaBinauralBackend` 的全部数学:64-QMF/77-hybrid analysis/synthesis、
五阶 ACN/N3D 实球谐方向场求值、whole-QMF-slot 逐对象 delay 历史、六面一阶
image-source early reflections、共享 unitary FDN late room、LFE 120–180 Hz
低通与 961-sample latency 语义。kernel 表、编译好的 HRTF 场与房间常数通过
`configure_kernels/configure_field/configure_room` 一次上传;每 512-sample
block 先 `set_source` 更新 16 个 source,再 `process` 输入 PCM;`process` 返回
裁剪后的 stereo 样本数(首个 961 samples 被丢弃)。`finish` 以 64-sample 对齐的
块排空尾音。Python 桥位于 `src/sofa_native_backend.py`,与 Python 参考实现逐值
一致(差异 < 1e-9);原生库缺失时 `main.py` 自动回退 Python。
## 7. 构建
CMake 定义位于 `native/CMakeLists.txt`。从仓库根目录运行:
@@ -138,7 +170,7 @@ MSVC 配置使用静态 CRT。其他运行时依赖由平台和工具链决定
仓库默认不附带原生二进制。预构建的 Release 运行库或自行构建的运行库均可直接放入 `lib/`。
## 7. 运行时查找与回退
## 8. 运行时查找与回退
查找顺序为:
@@ -148,7 +180,7 @@ MSVC 配置使用静态 CRT。其他运行时依赖由平台和工具链决定
`--backend auto` 在加载失败时回退到 NumPy;`--backend python` 跳过原生探测。`--backend native` 当前也会打印失败原因后回退,这是现有 CLI 行为,不应理解为原生库已成功使用。
## 8. 实现边界
## 9. 实现边界
- 原生层只接收 Python 已解析的 dense JOC 数据。
- ABI 固定了 1536-sample JOC 帧、最多 15 个对象、最多 23 个参数带和最多 2 个数据点。