Add binaural rendering support.
Native builds / linux-x64 (push) Failing after 11s
Native builds / macos-arm64 (push) Has been cancelled
Native builds / macos-x64 (push) Has been cancelled
Native builds / windows-x64 (push) Has been cancelled
Native builds / Publish GitHub Release (push) Has been cancelled
Native builds / linux-x64 (push) Failing after 11s
Native builds / macos-arm64 (push) Has been cancelled
Native builds / macos-x64 (push) Has been cancelled
Native builds / windows-x64 (push) Has been cancelled
Native builds / Publish GitHub Release (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1,277 @@
|
||||
# Binaural rendering
|
||||
|
||||
[中文](binaural.md) · [Back to README](../README.en.md)
|
||||
|
||||
JustOneCacophony's binaural backend supports three HRTF sources:
|
||||
`SimpleFreeFieldHRIR` SOFA, the Rosella `.personalized_headphone` model exported
|
||||
by Dolby's official personalization scan (its JSON parsing is implemented by
|
||||
this project and invokes no Dolby software), and the `.jochrtf` cache compiled
|
||||
from SOFA. SOFA is compiled into an in-memory directional field when the model
|
||||
is loaded. A `.jochrtf` file is only a disposable, reproducible JOC compiled
|
||||
HRTF cache; it is neither an interchange format nor a prerequisite for using
|
||||
SOFA.
|
||||
|
||||
```text
|
||||
SOFA FIR
|
||||
-> CanonicalHrtf
|
||||
-> 48 kHz / one radius shell / delay-phase policy
|
||||
-> 64-QMF / 77-hybrid projection
|
||||
-> fifth-order ACN/N3D real-SH field
|
||||
-> per-object direct + early reflections
|
||||
-> shared unitary-FDN late room
|
||||
-> float64 stereo
|
||||
```
|
||||
|
||||
## Inputs
|
||||
|
||||
The CLI has three mutually exclusive HRTF input sources; with none given, a
|
||||
default rule resolves the input:
|
||||
|
||||
```powershell
|
||||
# 1) SOFA: defaults to HRTF/binaural.sofa, or an explicit path
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone: defaults to HRTF/binaural.personalized_headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf: explicitly load a compiled cache
|
||||
python main.py input.m4a --binaural `
|
||||
--compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# Optional: create/reuse a transparent disk cache for SOFA
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
```
|
||||
|
||||
The default order is `HRTF/binaural.sofa`, then the unique `.jochrtf` under
|
||||
`output/hrtf-cache`, then `HRTF/binaural.personalized_headphone`; if none of
|
||||
the three exist, an error asks for an explicit path. Multiple `.jochrtf` files
|
||||
under `output/hrtf-cache` are also an error requiring an explicit choice.
|
||||
|
||||
The `.personalized_headphone` JSON parsing is implemented by this project
|
||||
(`src/rosella_model.py`) and does not invoke any Dolby software.
|
||||
|
||||
`--hrtf-cache-policy` accepts `none`, `memory`, or `disk`. The default is
|
||||
`memory`; neither `none` nor `memory` creates a file. `disk` writes to
|
||||
`output/hrtf-cache` by default, or to `--hrtf-cache-dir`. `--hrtf-radius-m`
|
||||
selects the nearest measurement-radius shell.
|
||||
|
||||
The Python API also uses explicit factories:
|
||||
|
||||
```python
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
|
||||
renderer = SofaBinauralBackend.from_sofa(
|
||||
"subject.sofa",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
cache_policy="memory",
|
||||
)
|
||||
|
||||
cached = SofaBinauralBackend.from_compiled_cache(
|
||||
"subject.jochrtf",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
)
|
||||
```
|
||||
|
||||
The factories never guess a format from an unknown suffix: SOFA and `.jochrtf`
|
||||
always use distinct loaders.
|
||||
|
||||
## Binaural render mode
|
||||
|
||||
`--binaural-mode off|near|mid|far` (default `mid`) is a **human-specified
|
||||
rendering hint**, not original binaural metadata extracted or recovered from the
|
||||
input E-AC-3 JOC bitstream:
|
||||
|
||||
- Direct binaural rendering (`--binaural`): near/mid/far apply, default `mid`;
|
||||
`off` is an error;
|
||||
- ADM BWF: the low 3 binaural-render-mode bits of the last 15 JOC object entries
|
||||
in DBMD segment 10 carry `off=0/near=1/far=2/mid=3`, leaving the first 10 bed
|
||||
entries unchanged; the default is `mid`, and `off` explicitly disables the
|
||||
binaural metadata hint.
|
||||
|
||||
## Canonical SOFA contract
|
||||
|
||||
The strict importer currently accepts:
|
||||
|
||||
- `Conventions=SOFA`;
|
||||
- `SOFAConventions=SimpleFreeFieldHRIR`, version `0.4`, `1.0`, or `1.1`;
|
||||
- `DataType=FIR` and `Data.IR[M,2,N]`;
|
||||
- one positive finite `Data.SamplingRate` in hertz/Hz;
|
||||
- spherical or Cartesian `SourcePosition`;
|
||||
- singleton or per-measurement `ListenerPosition/View/Up`;
|
||||
- two receivers whose listener-local lateral geometry uniquely identifies L/R;
|
||||
- one zero-offset emitter;
|
||||
- causal `Data.Delay[I,2]` or `[M,2]`;
|
||||
- an explicitly free-field/anechoic `RoomType`.
|
||||
|
||||
Receiver order comes from geometry, never from the receiver array index. SOFA
|
||||
listener coordinates are $+X$ front, $+Y$ left, $+Z$ up; ADM coordinates are
|
||||
$+X$ right, $+Y$ front, $+Z$ up:
|
||||
|
||||
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
|
||||
|
||||
`CanonicalHrtf` keeps `Data.IR` and `Data.Delay` separate. Only a time-domain
|
||||
baseline calls `materialized_measurement()` to apply delay once; the runtime SH
|
||||
path never materializes and then restores the delay. Non-48-kHz HRIRs are
|
||||
normalized with float64 `scipy.signal.resample_poly`, and delay samples scale by
|
||||
the same ratio.
|
||||
|
||||
GeneralFIR, BRIR, TF, multiple emitters, ambiguous receivers, and non-free-field
|
||||
data require convention-specific adapters. They cannot enter the core importer
|
||||
through a reshape.
|
||||
|
||||
## Exactly-once delay and phase
|
||||
|
||||
The compiler recognizes three mutually exclusive representations:
|
||||
|
||||
1. Nonzero `Data.Delay` is external to `Data.IR`; the FIR is not de-rotated and
|
||||
runtime applies the delay once.
|
||||
2. With `Data.Delay=0` and an ordinary positive-onset HRIR, each ear's main peak
|
||||
supplies arrival time. Compilation separates it and runtime restores it once.
|
||||
The current threshold is a peak index greater than two samples.
|
||||
3. With `Data.Delay=0` and both FIRs at a shared sample-zero origin, no external
|
||||
delay is invented. The authored complex phase stays in the fifth-order field.
|
||||
|
||||
No path may add a second ear delay or phase-group delay.
|
||||
|
||||
## Public filterbank and directional field
|
||||
|
||||
The runtime is fixed at:
|
||||
|
||||
- 48 kHz;
|
||||
- a 64-sample QMF hop;
|
||||
- 64-QMF / 77 hybrid bands;
|
||||
- 961 samples of analysis/synthesis latency;
|
||||
- fifth order, 36 terms, ACN/N3D real spherical harmonics;
|
||||
- float64 PCM, delay, SH, and room state; complex128 band transfers and spectra.
|
||||
|
||||
Real and imaginary unit gains for every hybrid band pass through the same
|
||||
analysis/synthesis chain to form a 154-real-parameter impulse dictionary. The
|
||||
compiler does not sample 77 FFT bins. Defaults are `1e-3` projection ridge and
|
||||
`1e-5` SH ridge. Coincident directions are merged before a spherical-Voronoi
|
||||
weighted ridge fit.
|
||||
|
||||
The fixed resource is `data/rosella_kernels.npz`, which implements publicly
|
||||
standardized filter banks, computable from the following formulas.
|
||||
|
||||
The hybrid analysis kernels are defined in [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf),
|
||||
Section 5.2.2 (Table 1 $Q=8$/$Q=4$ coefficients, delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\!\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
The QMF analysis table is the MPEG-4 AAC/SBR 64 complex QMF bank of
|
||||
ISO/IEC 14496-3/AMD1:2003, subclause 4.B.18.2, stored as the polyphase
|
||||
reordering of the public 640-tap prototype $c_0,\dots,c_{639}$:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
The QMF synthesis table is the causal left inverse of the analysis polyphase
|
||||
matrix $\mathbf{A}$, i.e. the solution of $\mathbf{A}\,\mathbf{W}=\mathbf{P}$
|
||||
($\mathbf{P}$ is the 577-sample delay permutation; total latency
|
||||
$961 = 577 + 6\times64$), stored as a rank-4 factorization:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
The hybrid synthesis table is the 77→64 recombination: identity for the high
|
||||
bands, $Y_{3+b}=X_{16+b}$, and for the low bands ($C_p$ is the $8+4+4$ child
|
||||
partition):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\operatorname{Re}X_q + j\,s_q\,\operatorname{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
The loader verifies the archive and every array by SHA-256; the table version,
|
||||
all array hashes, and the 77 reference band-center values are part of the cache
|
||||
key. Public availability of a standard does not by itself grant permission to
|
||||
practice related patent claims. See
|
||||
[`data/README.en.md`](../data/README.en.md) and
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the sources and the
|
||||
rights boundary.
|
||||
|
||||
## `.jochrtf`
|
||||
|
||||
A `.jochrtf` file is a pickle-free compressed NumPy archive with an exact member set:
|
||||
|
||||
| key | dtype / shape |
|
||||
|---|---|
|
||||
| `metadata_json` | NumPy Unicode scalar containing JSON text (`dtype.kind == "U"`) |
|
||||
| `band_center_frequencies_hz` | little-endian `float64[77]` |
|
||||
| `coefficients` | little-endian `complex128[36,2,77]` |
|
||||
| `delay_coefficients` | little-endian `float64[36,2]` |
|
||||
| `delay_bounds` | little-endian `float64[2,2]` |
|
||||
|
||||
Metadata uses the `JOC-HRTF-CACHE` magic and records the schema, compiler and
|
||||
phase-policy versions, ACN/N3D convention, filterbank hashes, SOFA content
|
||||
SHA-256, sample rate, radius, order, both ridge values, payload hash, and fit
|
||||
report. Every setting that changes compilation participates in the cache key.
|
||||
Metadata never persists an absolute local `source_path`; it may keep a display
|
||||
name only.
|
||||
|
||||
Before constructing a field, the loader uses `allow_pickle=False` and validates
|
||||
ZIP members and expanded sizes, shapes, dtypes, byte order, contiguous layout,
|
||||
finite values, delay bounds, band centers, payload hash, and cache key. The
|
||||
writer uses a same-directory temporary file, `fsync`, a process-held OS file
|
||||
lock, and atomic `os.replace`. Its hidden `.lock` sidecar may remain and does not
|
||||
mean that a writer still owns the lock. Outdated, damaged, or mismatched
|
||||
caches cannot hit. SOFA input rebuilds an invalid cache; an explicitly selected
|
||||
cache reports the error.
|
||||
|
||||
Deleting a disk cache must not change the field or render produced from the same
|
||||
SOFA and compiler configuration.
|
||||
|
||||
A `.jochrtf` file contains directional-field coefficients and delay data
|
||||
transformed from the source HRIRs. Its reproducibility therefore does not make
|
||||
it licence-free. Creating a cache does not enlarge the rights granted by the
|
||||
source SOFA/HRTF dataset: use, copying, and redistribution remain subject to
|
||||
that dataset's terms. If those terms are unclear, keep `.jochrtf` as a private
|
||||
local cache and do not ship it with the program or another build artifact.
|
||||
`source_sha256` is only a content-integrity identifier, not proof of provenance
|
||||
or permission.
|
||||
|
||||
## JOC objects and room behavior
|
||||
|
||||
The production adapter retains the existing JOC schedule:
|
||||
|
||||
- `[1536,16]` input per frame;
|
||||
- channel 0 is special LFE and channels 1..15 are JOC objects;
|
||||
- ID11/OAMD positions use a sample-timed timeline;
|
||||
- source parameters update every 512 samples;
|
||||
- every object owns independent direct/early history while one late FDN is shared;
|
||||
- `finish()` drains early/late tails; output gain is explicit, with no implicit
|
||||
limiter or programme loudness normalization.
|
||||
|
||||
Near/Mid/Far, equal-power direct level, six first-order shoebox image sources,
|
||||
late sends, the unitary FDN, the 120–180 Hz cosine-squared LFE low-pass, and room
|
||||
calibration are JOC project-defined behavior, not constants published by SOFA or
|
||||
Dolby.
|
||||
|
||||
The public SOFA binaural renderer defaults to the C++20 native core under
|
||||
`--backend auto/native` (`ejoc_sofa_binaural_*` in `lib/eac3joc_core.dll`): the
|
||||
filterbank, the SH direction-field evaluation, the per-object direct/early
|
||||
histories and the shared FDN all run natively, while Python only compiles the
|
||||
SOFA source and issues the per-512-sample metadata updates. When the native
|
||||
library is unavailable the renderer falls back to the Python/NumPy reference
|
||||
implementation; the two agree to better than 1e-9. `--backend python` forces
|
||||
the Python backend.
|
||||
`--backend` still selects native/Python JOC reconstruction and speaker rendering;
|
||||
native acceleration for the public binaural DSP is outside the current API.
|
||||
|
||||
## Technical references and rights boundary
|
||||
|
||||
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
|
||||
- [3GPP TS 26.405 / ETSI TS 126 405 (64-QMF/77-hybrid definition)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
- [Dolby binaural render-mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
|
||||
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en), used only as
|
||||
architectural background for direct/early/late, subbands, and FDNs; it does
|
||||
not establish that any product uses a particular embodiment.
|
||||
|
||||
Public availability of a specification, source file, or patent document does
|
||||
not by itself authorize copying its contents, redistribution of derivatives,
|
||||
or practice of patent claims. These technical references grant no patent
|
||||
licence and make no non-infringement representation. Anyone preparing a release
|
||||
or product integration must assess the applicable data and software licences,
|
||||
patent permissions, and freedom to operate. See
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the public-standard
|
||||
provenance and rights boundary.
|
||||
@@ -0,0 +1,238 @@
|
||||
# 双耳渲染
|
||||
|
||||
[English](binaural.en.md) · [返回 README](../README.md)
|
||||
|
||||
JustOneCacophony 的双耳后端支持三种 HRTF 来源:`SimpleFreeFieldHRIR` SOFA、
|
||||
杜比官方软件个性化扫描导出的 Rosella `.personalized_headphone`(JSON 解析由本项目
|
||||
自行实现,不调用杜比软件),以及从 SOFA 编译出的 `.jochrtf` 缓存。SOFA 在模型加载
|
||||
时编译成内存方向场;`.jochrtf` 只是可删除、可重建的 JOC compiled HRTF cache,
|
||||
不是交换格式,也不是使用 SOFA 的前置步骤。
|
||||
|
||||
```text
|
||||
SOFA FIR
|
||||
-> CanonicalHrtf
|
||||
-> 48 kHz / 单 radius shell / delay-phase policy
|
||||
-> 64-QMF / 77-hybrid projection
|
||||
-> 五阶 ACN/N3D 实球谐场
|
||||
-> 逐对象 direct + early reflections
|
||||
-> shared unitary-FDN late room
|
||||
-> stereo float64
|
||||
```
|
||||
|
||||
## 输入接口
|
||||
|
||||
CLI 有三个互斥的 HRTF 输入来源;都不指定时按默认规则自动选择:
|
||||
|
||||
```powershell
|
||||
# 1) SOFA:缺省取 HRTF/binaural.sofa,也可显式指定
|
||||
python main.py input.m4a --binaural
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa
|
||||
|
||||
# 2) Rosella .personalized_headphone:缺省取 HRTF/binaural.personalized_headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone
|
||||
python main.py input.m4a --binaural --personalized-headphone C:\HRTF\subject.personalized_headphone
|
||||
|
||||
# 3) .jochrtf:显式读取预编译 cache
|
||||
python main.py input.m4a --binaural `
|
||||
--compiled-hrtf-cache C:\HRTF\subject.jochrtf
|
||||
|
||||
# 可选:SOFA 透明生成/复用磁盘 cache
|
||||
python main.py input.m4a --binaural --sofa-hrtf C:\HRTF\subject.sofa `
|
||||
--hrtf-cache-policy disk
|
||||
```
|
||||
|
||||
默认选择顺序:`HRTF/binaural.sofa` → `output/hrtf-cache` 下唯一的 `.jochrtf` →
|
||||
`HRTF/binaural.personalized_headphone`;三者都没有时报错并提示显式指定。
|
||||
`output/hrtf-cache` 下有多个 `.jochrtf` 时同样报错,要求显式选择。
|
||||
|
||||
`.personalized_headphone` 的 JSON 解析由本项目自行实现(`src/rosella_model.py`),
|
||||
不调用任何杜比软件。
|
||||
|
||||
`--hrtf-cache-policy` 可取 `none`、`memory`、`disk`。默认是 `memory`;`none` 和
|
||||
`memory` 都不会创建磁盘文件。`disk` 默认写入 `output/hrtf-cache`,也可用
|
||||
`--hrtf-cache-dir` 指定。`--hrtf-radius-m` 选择距离目标最近的 measurement shell。
|
||||
|
||||
Python API 使用显式 factory:
|
||||
|
||||
```python
|
||||
from sofa_binaural_backend import SofaBinauralBackend
|
||||
|
||||
renderer = SofaBinauralBackend.from_sofa(
|
||||
"subject.sofa",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
cache_policy="memory",
|
||||
)
|
||||
|
||||
cached = SofaBinauralBackend.from_compiled_cache(
|
||||
"subject.jochrtf",
|
||||
source_count=16,
|
||||
default_profile="mid",
|
||||
)
|
||||
```
|
||||
|
||||
文件工厂不会按“未知后缀”猜格式:SOFA 和 `.jochrtf` 始终走不同 loader。
|
||||
|
||||
## 双耳渲染模式
|
||||
|
||||
`--binaural-mode off|near|mid|far`(默认 `mid`)是**人为指定的渲染提示**,不是
|
||||
从输入 E-AC-3 JOC 码流提取或还原的原始双耳元数据:
|
||||
|
||||
- 直接双耳渲染(`--binaural`):near/mid/far 生效,默认 `mid`;`off` 报错;
|
||||
- ADM BWF:DBMD segment 10 中后 15 个 JOC 对象的 binaural render mode 写
|
||||
`off=0/near=1/far=2/mid=3`,前 10 个 bed 保持不变,默认 `mid`;`off` 用于显式
|
||||
关闭双耳元数据提示。
|
||||
|
||||
## Canonical SOFA 契约
|
||||
|
||||
当前 strict importer 接受:
|
||||
|
||||
- `Conventions=SOFA`;
|
||||
- `SOFAConventions=SimpleFreeFieldHRIR`,version `0.4`、`1.0` 或 `1.1`;
|
||||
- `DataType=FIR`,`Data.IR[M,2,N]`;
|
||||
- 单一正有限 `Data.SamplingRate`,单位为 hertz/Hz;
|
||||
- spherical 或 Cartesian `SourcePosition`;
|
||||
- 单值或 per-measurement 的 `ListenerPosition/View/Up`;
|
||||
- 两个能由 listener-local lateral 坐标唯一识别左右的 receiver;
|
||||
- 单一且零偏移的 emitter;
|
||||
- causal `Data.Delay[I,2]` 或 `[M,2]`;
|
||||
- 明确的 free-field/anechoic `RoomType`。
|
||||
|
||||
receiver 左右顺序由几何决定,不能假定 `Data.IR` 的 receiver index。SOFA listener
|
||||
坐标为 $+X$ front、$+Y$ left、$+Z$ up;ADM 坐标为 $+X$ right、$+Y$ front、
|
||||
$+Z$ up,转换为:
|
||||
|
||||
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
|
||||
|
||||
`CanonicalHrtf` 将 `Data.IR` 与 `Data.Delay` 分开保存。只有时域 baseline 才调用
|
||||
`materialized_measurement()` 将 delay 应用一次;运行时 SH 路径不先 materialize。
|
||||
非 48 kHz HRIR 使用 float64 `scipy.signal.resample_poly` 规范化,delay samples 按
|
||||
相同比例缩放。
|
||||
|
||||
GeneralFIR、BRIR、TF、多 emitter、多义 receiver 或非 free-field 数据需要单独的
|
||||
convention adapter,不能只通过 reshape 进入核心 importer。
|
||||
|
||||
## Delay/phase:exactly once
|
||||
|
||||
编译器只允许三种互斥语义:
|
||||
|
||||
1. 非零 `Data.Delay` 是 `Data.IR` 外部 delay;FIR 不去旋,运行时应用一次。
|
||||
2. `Data.Delay=0` 且 HRIR 有普通正 onset:以每耳 main peak 分离 arrival,拟合后
|
||||
在运行时恢复一次;当前阈值为 peak index 大于 2 samples。
|
||||
3. `Data.Delay=0` 且双耳 FIR 共享 sample-0 起点:不发明外部 delay,原 complex
|
||||
phase 直接进入五阶场。
|
||||
|
||||
任何路径都不能再叠加第二套 ear delay 或 phase-group delay。
|
||||
|
||||
## 公开 filterbank 与方向场
|
||||
|
||||
运行时固定为:
|
||||
|
||||
- 48 kHz;
|
||||
- 64-sample QMF hop;
|
||||
- 64-QMF / 77-hybrid;
|
||||
- analysis/synthesis latency 961 samples;
|
||||
- 五阶、36 项、ACN/N3D real spherical harmonics;
|
||||
- PCM、delay、SH、room state 为 float64;频带传递和频域状态为 complex128。
|
||||
|
||||
每个 hybrid band 的 real/imaginary 单位增益都通过同一套 analysis/synthesis 链生成
|
||||
脉冲字典,共 154 个实参数;编译不是直接读取 77 个 FFT bin。默认 projection
|
||||
ridge 为 `1e-3`,SH ridge 为 `1e-5`。同方向 measurement 先合并,再用球面 Voronoi
|
||||
面积权重做 ridge fit。
|
||||
|
||||
固定表位于 `data/rosella_kernels.npz`,实现公开标准化的滤波器组,各表可由如下
|
||||
公式计算。
|
||||
|
||||
hybrid 分析核定义于 [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
第 5.2.2 节(Table 1 的 $Q=8$/$Q=4$ 系数,delay 6):
|
||||
|
||||
$$G_q^p[n] = g^p[n]\cdot\exp\!\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
|
||||
|
||||
QMF analysis 表即 MPEG-4 AAC/SBR(ISO/IEC 14496-3/AMD1:2003 第 4.B.18.2 节)
|
||||
的 64 complex QMF bank;打包的 $64\times10$ 表是公开 640-tap prototype
|
||||
$c_0,\dots,c_{639}$ 的多相重排:
|
||||
|
||||
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
|
||||
|
||||
QMF synthesis 表为上述 analysis 多相矩阵 $\mathbf{A}$ 的因果左逆,即求解
|
||||
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$($\mathbf{P}$ 为 577-sample 延迟置换;
|
||||
全链 $961 = 577 + 6\times64$),以 rank-4 分解形式存储:
|
||||
|
||||
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
|
||||
|
||||
hybrid synthesis 表为 77→64 重组:高频带恒等 $Y_{3+b}=X_{16+b}$;低频带
|
||||
($C_p$ 为 $8+4+4$ 子带划分):
|
||||
|
||||
$$Y_p = \sum_{q\in C_p}\Bigl(\operatorname{Re}X_q + j\,s_q\,\operatorname{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
|
||||
|
||||
loader 校验 archive 和每个数组的 SHA-256;table version、所有数组 hash 与
|
||||
77 个 band-center 参考值都属于 cache key。标准可公开获取不等于获准实施相关
|
||||
专利;更多来源信息见 [`data/README.md`](../data/README.md) 与
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
|
||||
|
||||
## `.jochrtf`
|
||||
|
||||
`.jochrtf` 是无 pickle 的压缩 NumPy archive,固定包含:
|
||||
|
||||
| key | dtype / shape |
|
||||
|---|---|
|
||||
| `metadata_json` | 含 JSON 文本的 NumPy Unicode scalar(`dtype.kind == "U"`) |
|
||||
| `band_center_frequencies_hz` | little-endian `float64[77]` |
|
||||
| `coefficients` | little-endian `complex128[36,2,77]` |
|
||||
| `delay_coefficients` | little-endian `float64[36,2]` |
|
||||
| `delay_bounds` | little-endian `float64[2,2]` |
|
||||
|
||||
metadata magic 固定为 `JOC-HRTF-CACHE`,并记录 schema/compiler/phase-policy、
|
||||
ACN/N3D、filterbank table hashes、SOFA content SHA-256、采样率、radius、order、
|
||||
两个 ridge、payload hash 和 fit report。cache key 覆盖所有会改变编译结果的字段。
|
||||
metadata 不保存本机绝对 `source_path`,仅可保存 source display name。
|
||||
|
||||
loader 使用 `allow_pickle=False`,并在构造对象前检查 ZIP 成员集、解压大小、shape、
|
||||
dtype、端序、连续布局、有限值、delay bounds、band centers、payload hash 和 cache
|
||||
key。writer 使用同目录临时文件、`fsync`、进程持有的 OS 文件锁和原子
|
||||
`os.replace`;对应的隐藏 `.lock` sidecar 可保留,但不代表仍有 writer 持锁。
|
||||
旧版本、损坏或配置不匹配的 cache 不能命中;从 SOFA 启动时会重建,显式 cache
|
||||
入口则直接报错。
|
||||
|
||||
删除磁盘 cache 后,从同一 SOFA 和同一编译配置得到的场与渲染结果不得改变。
|
||||
|
||||
`.jochrtf` 包含由源 HRIR 变换得到的方向场系数与 delay 数据,因此“可以重建”不表示
|
||||
它不受数据许可约束。生成 cache 不会扩大源 SOFA/HRTF 数据集授予的权利;cache 的
|
||||
使用、复制和再分发仍须遵守源数据集条款。不能确认条款时,应把 `.jochrtf` 作为本地
|
||||
私有 cache,不随程序或构建产物发布。`source_sha256` 只用于内容一致性校验,不是许可
|
||||
或来源证明。
|
||||
|
||||
## JOC 对象与房间
|
||||
|
||||
生产适配器继续使用现有 JOC 调度:
|
||||
|
||||
- 每帧输入 `[1536,16]`;
|
||||
- channel 0 是 special LFE,channel 1..15 是 JOC objects;
|
||||
- ID11/OAMD position 使用 sample-timed timeline;
|
||||
- 每 512 samples 更新方向/profile;
|
||||
- 每个对象拥有独立 direct/early history,late FDN 全局共享;
|
||||
- `finish()` 排空 early/late tail;输出增益显式应用,不隐含 limiter 或节目响度归一化。
|
||||
|
||||
Near/Mid/Far、equal-power direct level、六面 shoebox 一阶 image source、late send、
|
||||
unitary FDN、LFE 120–180 Hz cosine-squared 低通及 room calibration 都是 JOC
|
||||
项目定义行为,不是 SOFA 或 Dolby 公布常数。
|
||||
|
||||
公开 SOFA 双耳渲染在 `--backend auto/native` 下默认走 C++20 原生核
|
||||
(`lib/eac3joc_core.dll` 的 `ejoc_sofa_binaural_*` 接口:filterbank、SH 方向场求值、
|
||||
逐对象 early/direct 历史与共享 FDN 全部在原生侧执行,Python 只做 SOFA 编译与每
|
||||
512-sample 的元数据更新);原生库不可用时自动回退 Python/NumPy 参考实现,两者
|
||||
逐值一致(差异 < 1e-9)。`--backend python` 强制使用 Python 后端。
|
||||
|
||||
## 技术引用与权利边界
|
||||
|
||||
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
|
||||
- [3GPP TS 26.405 / ETSI TS 126 405(64-QMF/77-hybrid 定义)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
|
||||
- [Dolby binaural render mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
|
||||
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en),仅作 direct/early/late、
|
||||
subband 与 FDN 架构背景,不证明某个产品使用特定实施例。
|
||||
|
||||
规范、源码或专利文献可公开获取,不等于获准复制其内容、再分发派生产物或实施其中的
|
||||
专利权利要求。本项目的技术引用本身不授予专利许可,也不作不侵权保证;准备发布或集成
|
||||
到产品的一方应自行审查适用的数据许可、软件许可、专利许可及 freedom-to-operate。
|
||||
公开标准来源与权利边界见
|
||||
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
|
||||
+22
-3
@@ -116,7 +116,26 @@ Supported layouts:
|
||||
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
|
||||
```
|
||||
|
||||
## 6. Building
|
||||
## 6. Binaural-rendering ABI
|
||||
|
||||
The shared library provides a 512-sample float64 binaural DSP interface:
|
||||
|
||||
```c
|
||||
ejoc_binaural_renderer_handle ejoc_binaural_renderer_create(void);
|
||||
int ejoc_binaural_renderer_configure_kernels(...);
|
||||
int ejoc_binaural_renderer_configure_room(...);
|
||||
int ejoc_binaural_renderer_process(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
const double* input16_interleaved, /* [512][16] */
|
||||
const double* gains_complex, /* [16][2][77][2] */
|
||||
const double* room_sends, /* [16] */
|
||||
double output_gain,
|
||||
double* output_stereo_interleaved); /* [512][2] */
|
||||
```
|
||||
|
||||
Python parses the model, evaluates the OAMD timeline, and supplies complex gains and room sends every 512 samples. The C++ handle owns QMF, hybrid, recursive-room, and QMF-synthesis state. Inputs, state, accumulation, and output are double/complex double.
|
||||
|
||||
## 7. Building
|
||||
|
||||
The CMake definition is `native/CMakeLists.txt`. Run from the repository root:
|
||||
|
||||
@@ -138,7 +157,7 @@ The MSVC configuration uses the static CRT. Other runtime dependencies depend on
|
||||
|
||||
The repository does not include native binaries by default. A prebuilt Release runtime or a locally built runtime can be placed directly under `lib/`.
|
||||
|
||||
## 7. Runtime lookup and fallback
|
||||
## 8. Runtime lookup and fallback
|
||||
|
||||
Lookup order:
|
||||
|
||||
@@ -148,7 +167,7 @@ Lookup order:
|
||||
|
||||
`--backend auto` falls back to NumPy when loading fails, and `--backend python` skips native discovery. The current CLI also prints the failure and falls back for `--backend native`; this existing behavior should not be read as successful native execution.
|
||||
|
||||
## 8. Implementation boundaries
|
||||
## 9. Implementation boundaries
|
||||
|
||||
- The native layer accepts only dense-JOC data already parsed by Python.
|
||||
- The ABI fixes a 1536-sample JOC frame, at most 15 objects, at most 23 parameter bands, and at most 2 data points.
|
||||
|
||||
+35
-3
@@ -116,7 +116,39 @@ int ejoc_speaker_renderer_process(
|
||||
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
|
||||
```
|
||||
|
||||
## 6. 构建
|
||||
## 6. 双耳渲染 ABI
|
||||
|
||||
共享库提供 512-sample float64 双耳 DSP:
|
||||
|
||||
```c
|
||||
ejoc_binaural_renderer_handle ejoc_binaural_renderer_create(void);
|
||||
int ejoc_binaural_renderer_configure_kernels(...);
|
||||
int ejoc_binaural_renderer_configure_room(...);
|
||||
int ejoc_binaural_renderer_process(
|
||||
ejoc_binaural_renderer_handle handle,
|
||||
const double* input16_interleaved, /* [512][16] */
|
||||
const double* gains_complex, /* [16][2][77][2] */
|
||||
const double* room_sends, /* [16] */
|
||||
double output_gain,
|
||||
double* output_stereo_interleaved); /* [512][2] */
|
||||
```
|
||||
|
||||
Python 负责模型解析、OAMD 时间轴和每 512 samples 的 complex gains/room sends。C++ handle 保存 QMF、hybrid、递归 room 和 QMF synthesis 状态。全部输入、状态、乘加和输出均为 double/complex double。
|
||||
|
||||
## 6.1 公开 SOFA 双耳渲染 ABI
|
||||
|
||||
共享库同时提供完整的原生 SOFA 双耳渲染器(`ejoc_sofa_binaural_*`),它镜像
|
||||
Python `SofaBinauralBackend` 的全部数学:64-QMF/77-hybrid analysis/synthesis、
|
||||
五阶 ACN/N3D 实球谐方向场求值、whole-QMF-slot 逐对象 delay 历史、六面一阶
|
||||
image-source early reflections、共享 unitary FDN late room、LFE 120–180 Hz
|
||||
低通与 961-sample latency 语义。kernel 表、编译好的 HRTF 场与房间常数通过
|
||||
`configure_kernels/configure_field/configure_room` 一次上传;每 512-sample
|
||||
block 先 `set_source` 更新 16 个 source,再 `process` 输入 PCM;`process` 返回
|
||||
裁剪后的 stereo 样本数(首个 961 samples 被丢弃)。`finish` 以 64-sample 对齐的
|
||||
块排空尾音。Python 桥位于 `src/sofa_native_backend.py`,与 Python 参考实现逐值
|
||||
一致(差异 < 1e-9);原生库缺失时 `main.py` 自动回退 Python。
|
||||
|
||||
## 7. 构建
|
||||
|
||||
CMake 定义位于 `native/CMakeLists.txt`。从仓库根目录运行:
|
||||
|
||||
@@ -138,7 +170,7 @@ MSVC 配置使用静态 CRT。其他运行时依赖由平台和工具链决定
|
||||
|
||||
仓库默认不附带原生二进制。预构建的 Release 运行库或自行构建的运行库均可直接放入 `lib/`。
|
||||
|
||||
## 7. 运行时查找与回退
|
||||
## 8. 运行时查找与回退
|
||||
|
||||
查找顺序为:
|
||||
|
||||
@@ -148,7 +180,7 @@ MSVC 配置使用静态 CRT。其他运行时依赖由平台和工具链决定
|
||||
|
||||
`--backend auto` 在加载失败时回退到 NumPy;`--backend python` 跳过原生探测。`--backend native` 当前也会打印失败原因后回退,这是现有 CLI 行为,不应理解为原生库已成功使用。
|
||||
|
||||
## 8. 实现边界
|
||||
## 9. 实现边界
|
||||
|
||||
- 原生层只接收 Python 已解析的 dense JOC 数据。
|
||||
- ABI 固定了 1536-sample JOC 帧、最多 15 个对象、最多 23 个参数带和最多 2 个数据点。
|
||||
|
||||
Reference in New Issue
Block a user