Initial pure C++ implementation
ci / ubuntu-latest (push) Failing after 13s
ci / macos-latest (push) Has been cancelled
ci / windows-latest (push) Has been cancelled

This commit is contained in:
2026-09-25 04:17:30 +08:00
commit 42d6da6302
110 changed files with 37510 additions and 0 deletions
+246
View File
@@ -0,0 +1,246 @@
# Binaural rendering
[中文](binaural.md) · [Back to README](../README.en.md)
JustOneCacophony's binaural backend supports three HRTF sources:
`SimpleFreeFieldHRIR` SOFA, the Rosella `.personalized_headphone` model exported
by Dolby's official personalization scan (its JSON parsing is implemented by
this project and invokes no Dolby software), and the `.jochrtf` cache compiled
from SOFA. SOFA is compiled into an in-memory directional field when the model
is loaded. A `.jochrtf` file is only a disposable, reproducible JOC compiled
HRTF cache; it is neither an interchange format nor a prerequisite for using
SOFA.
```text
SOFA FIR
-> CanonicalHrtf
-> 48 kHz / one radius shell / delay-phase policy
-> 64-QMF / 77-hybrid projection
-> fifth-order ACN/N3D real-SH field
-> per-object direct + early reflections
-> shared unitary-FDN late room
-> float64 stereo
```
## Inputs
The binaural backend consumes a compiled directional field (a JOC compiled HRTF cache,
`.jochrtf`) plus the shared filterbank tables. Both are read-only inputs: this library
performs no parsing or conversion of measurement data formats.
```powershell
joc_cli input.m4a --binaural `
--compiled-hrtf-cache path\to\subject.jochrtf `
--kernels data\rosella_kernels.npz `
--binaural-mode mid --binaural-tail-seconds 5.0
```
The library exposes the same fields: `joc_task_config` for a file task,
`joc_stream_config` for streaming, where `hrtf_path`, `kernels_path`,
`binaural_mode` and `binaural_tail_seconds` configure the binaural path.
The `.jochrtf` file is an **input**, not a product of this library: compiling it from
SOFA data or measurements belongs to the toolchain and is decoupled from this
repository. Loading validates the member set, dtypes and shapes, C-contiguity,
CRC-32 and a payload hash recomputed over the members (see `.jochrtf` below).
`binaural_mode` is `near`, `mid` or `far` and selects one of the backend's three
preset parameter sets; `binaural_tail_seconds` sets the room-tail length drained on
flush (default 5.0 s).
## Binaural render mode
`--binaural-mode off|near|mid|far` (default `mid`) is a **human-specified
rendering hint**, not original binaural metadata extracted or recovered from the
input E-AC-3 JOC bitstream:
- Direct binaural rendering (`--binaural`): near/mid/far apply, default `mid`;
`off` is an error;
- ADM BWF: the low 3 binaural-render-mode bits of the last 15 JOC object entries
in DBMD segment 10 carry `off=0/near=1/far=2/mid=3`, leaving the first 10 bed
entries unchanged; the default is `mid`, and `off` explicitly disables the
binaural metadata hint.
## Canonical SOFA contract
The strict importer currently accepts:
- `Conventions=SOFA`;
- `SOFAConventions=SimpleFreeFieldHRIR`, version `0.4`, `1.0`, or `1.1`;
- `DataType=FIR` and `Data.IR[M,2,N]`;
- one positive finite `Data.SamplingRate` in hertz/Hz;
- spherical or Cartesian `SourcePosition`;
- singleton or per-measurement `ListenerPosition/View/Up`;
- two receivers whose listener-local lateral geometry uniquely identifies L/R;
- one zero-offset emitter;
- causal `Data.Delay[I,2]` or `[M,2]`;
- an explicitly free-field/anechoic `RoomType`.
Receiver order comes from geometry, never from the receiver array index. SOFA
listener coordinates are $+X$ front,
$+Y$ left,
$+Z$ up; ADM coordinates are
$+X$ right,
$+Y$ front,
$+Z$ up:
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
`CanonicalHrtf` keeps `Data.IR` and `Data.Delay` separate. Only a time-domain
baseline calls `materialized_measurement()` to apply delay once; the runtime SH
path never materializes and then restores the delay. Non-48-kHz HRIRs are
normalized with float64 `scipy.signal.resample_poly`, and delay samples scale by
the same ratio.
GeneralFIR, BRIR, TF, multiple emitters, ambiguous receivers, and non-free-field
data require convention-specific adapters. They cannot enter the core importer
through a reshape.
## Exactly-once delay and phase
The compiler recognizes three mutually exclusive representations:
1. Nonzero `Data.Delay` is external to `Data.IR`; the FIR is not de-rotated and
runtime applies the delay once.
2. With `Data.Delay=0` and an ordinary positive-onset HRIR, each ear's main peak
supplies arrival time. Compilation separates it and runtime restores it once.
The current threshold is a peak index greater than two samples.
3. With `Data.Delay=0` and both FIRs at a shared sample-zero origin, no external
delay is invented. The authored complex phase stays in the fifth-order field.
No path may add a second ear delay or phase-group delay.
## Public filterbank and directional field
The runtime is fixed at:
- 48 kHz;
- a 64-sample QMF hop;
- 64-QMF / 77 hybrid bands;
- 961 samples of analysis/synthesis latency;
- fifth order, 36 terms, ACN/N3D real spherical harmonics;
- float64 PCM, delay, SH, and room state; complex128 band transfers and spectra.
Real and imaginary unit gains for every hybrid band pass through the same
analysis/synthesis chain to form a 154-real-parameter impulse dictionary. The
compiler does not sample 77 FFT bins. Defaults are `1e-3` projection ridge and
`1e-5` SH ridge. Coincident directions are merged before a spherical-Voronoi
weighted ridge fit.
The fixed resource is `data/rosella_kernels.npz`, which implements publicly
standardized filter banks, computable from the following formulas.
The hybrid analysis kernels are defined in [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf),
Section 5.2.2 (Table 1 $Q=8$/
$Q=4$ coefficients, delay 6):
$$G_q^p[n] = g^p[n]\cdot\exp\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
The QMF analysis table is the MPEG-4 AAC/SBR 64 complex QMF bank of
ISO/IEC 14496-3/AMD1:2003, subclause 4.B.18.2, stored as the polyphase
reordering of the public 640-tap prototype $c_0,\dots,c_{639}$:
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
The QMF synthesis table is the causal left inverse of the analysis polyphase
matrix $\mathbf{A}$, i.e. the solution of
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$
($\mathbf{P}$ is the 577-sample delay permutation; total latency
$961 = 577 + 6\times64$), stored as a rank-4 factorization:
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
The hybrid synthesis table is the 77→64 recombination: identity for the high
bands, $Y_{3+b}=X_{16+b}$, and for the low bands(
$C_p$ is the
$8+4+4$ child partition):
$$Y_p = \sum_{q\in C_p}\Bigl(\mathrm{Re}X_q + j\,s_q\,\mathrm{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
The loader verifies the archive and every array by SHA-256; the table version,
all array hashes, and the 77 reference band-center values are part of the cache
key. Public availability of a standard does not by itself grant permission to
practice related patent claims. See
[`data/README.en.md`](../data/README.en.md) and
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the sources and the
rights boundary.
## `.jochrtf`
A `.jochrtf` file is a pickle-free compressed NumPy archive with an exact member set:
| key | dtype / shape |
|---|---|
| `metadata_json` | NumPy Unicode scalar containing JSON text (`dtype.kind == "U"`) |
| `band_center_frequencies_hz` | little-endian `float64[77]` |
| `coefficients` | little-endian `complex128[36,2,77]` |
| `delay_coefficients` | little-endian `float64[36,2]` |
| `delay_bounds` | little-endian `float64[2,2]` |
Metadata uses the `JOC-HRTF-CACHE` magic and records the schema, compiler and
phase-policy versions, ACN/N3D convention, filterbank hashes, SOFA content
SHA-256, sample rate, radius, order, both ridge values, payload hash, and fit
report. Every setting that changes compilation participates in the cache key.
Metadata never persists an absolute local `source_path`; it may keep a display
name only.
Before constructing a field, the loader uses `allow_pickle=False` and validates
ZIP members and expanded sizes, shapes, dtypes, byte order, contiguous layout,
finite values, delay bounds, band centers, payload hash, and cache key. The
writer uses a same-directory temporary file, `fsync`, a process-held OS file
lock, and atomic `os.replace`. Its hidden `.lock` sidecar may remain and does not
mean that a writer still owns the lock. Outdated, damaged, or mismatched
caches cannot hit. SOFA input rebuilds an invalid cache; an explicitly selected
cache reports the error.
Deleting a disk cache must not change the field or render produced from the same
SOFA and compiler configuration.
A `.jochrtf` file contains directional-field coefficients and delay data
transformed from the source HRIRs. Its reproducibility therefore does not make
it licence-free. Creating a cache does not enlarge the rights granted by the
source SOFA/HRTF dataset: use, copying, and redistribution remain subject to
that dataset's terms. If those terms are unclear, keep `.jochrtf` as a private
local cache and do not ship it with the program or another build artifact.
`source_sha256` is only a content-integrity identifier, not proof of provenance
or permission.
## JOC objects and room behavior
The production adapter retains the existing JOC schedule:
- `[1536,16]` input per frame;
- channel 0 is special LFE and channels 1..15 are JOC objects;
- ID11/OAMD positions use a sample-timed timeline;
- source parameters update every 512 samples;
- every object owns independent direct/early history while one late FDN is shared;
- `finish()` drains early/late tails; output gain is explicit, with no implicit
limiter or programme loudness normalization.
Near/Mid/Far, equal-power direct level, six first-order shoebox image sources,
late sends, the unitary FDN, the 120–180 Hz cosine-squared LFE low-pass, and room
calibration are JOC project-defined behavior, not constants published by SOFA or
Dolby.
The binaural renderer is implemented inside this library: the filterbank, the SH
directional-field evaluation, the per-object early/direct histories and the shared
FDN all run in `joc_core` (the `ejoc_sofa_binaural_*` kernel), consuming a compiled
directional field and 512-sample metadata updates.
## Technical references and rights boundary
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
- [3GPP TS 26.405 / ETSI TS 126 405 (64-QMF/77-hybrid definition)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
- [Dolby binaural render-mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en), used only as
architectural background for direct/early/late, subbands, and FDNs; it does
not establish that any product uses a particular embodiment.
Public availability of a specification, source file, or patent document does
not by itself authorize copying its contents, redistribution of derivatives,
or practice of patent claims. These technical references grant no patent
licence and make no non-infringement representation. Anyone preparing a release
or product integration must assess the applicable data and software licences,
patent permissions, and freedom to operate. See
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md) for the public-standard
provenance and rights boundary.
+211
View File
@@ -0,0 +1,211 @@
# 双耳渲染
[English](binaural.en.md) · [返回 README](../README.md)
JustOneCacophony 的双耳后端支持三种 HRTF 来源:`SimpleFreeFieldHRIR` SOFA、
杜比官方软件个性化扫描导出的 Rosella `.personalized_headphone`(JSON 解析由本项目
自行实现,不调用杜比软件),以及从 SOFA 编译出的 `.jochrtf` 缓存。SOFA 在模型加载
时编译成内存方向场;`.jochrtf` 只是可删除、可重建的 JOC compiled HRTF cache,
不是交换格式,也不是使用 SOFA 的前置步骤。
```text
SOFA FIR
-> CanonicalHrtf
-> 48 kHz / 单 radius shell / delay-phase policy
-> 64-QMF / 77-hybrid projection
-> 五阶 ACN/N3D 实球谐场
-> 逐对象 direct + early reflections
-> shared unitary-FDN late room
-> stereo float64
```
## 输入接口
双耳后端消费一个已编译的方向场(JOC compiled HRTF cache,`.jochrtf`)与共享滤波器组表;
两者都是只读输入,本库不做任何测量数据格式的解析或转换:
```powershell
joc_cli input.m4a --binaural `
--compiled-hrtf-cache path\to\subject.jochrtf `
--kernels data\rosella_kernels.npz `
--binaural-mode mid --binaural-tail-seconds 5.0
```
库接口使用同一组字段:文件任务用 `joc_task_config`,流式用 `joc_stream_config`,
其中 `hrtf_path`、`kernels_path`、`binaural_mode`、`binaural_tail_seconds` 决定双耳通路。
`.jochrtf` 是**输入**而不是本库的产物:从 SOFA 或测量数据编译该缓存属于工具链的职责,
与本仓库解耦。缓存加载时会校验成员集合、dtype 与形状、C 连续性、CRC-32 以及按成员重算的
载荷哈希(见下文 `.jochrtf` 一节)。
`binaural_mode` 取 `near`、`mid`、`far`,选择后端的三组预置参数;`binaural_tail_seconds`
决定 flush 时排空的房间尾音长度(默认 5.0 s)。
## 双耳渲染模式
`--binaural-mode off|near|mid|far`(默认 `mid`)是**人为指定的渲染提示**,不是
从输入 E-AC-3 JOC 码流提取或还原的原始双耳元数据:
- 直接双耳渲染(`--binaural`):near/mid/far 生效,默认 `mid`;`off` 报错;
- ADM BWF:DBMD segment 10 中后 15 个 JOC 对象的 binaural render mode 写
`off=0/near=1/far=2/mid=3`,前 10 个 bed 保持不变,默认 `mid`;`off` 用于显式
关闭双耳元数据提示。
## Canonical SOFA 契约
当前 strict importer 接受:
- `Conventions=SOFA`;
- `SOFAConventions=SimpleFreeFieldHRIR`,version `0.4`、`1.0` 或 `1.1`;
- `DataType=FIR`,`Data.IR[M,2,N]`;
- 单一正有限 `Data.SamplingRate`,单位为 hertz/Hz;
- spherical 或 Cartesian `SourcePosition`;
- 单值或 per-measurement 的 `ListenerPosition/View/Up`;
- 两个能由 listener-local lateral 坐标唯一识别左右的 receiver;
- 单一且零偏移的 emitter;
- causal `Data.Delay[I,2]` 或 `[M,2]`;
- 明确的 free-field/anechoic `RoomType`。
receiver 左右顺序由几何决定,不能假定 `Data.IR` 的 receiver index。SOFA listener
坐标为 $+X$ front、
$+Y$ left、
$+Z$ up;ADM 坐标为
$+X$ right、
$+Y$ front、
$+Z$ up,转换为:
$$\bigl(x_{\mathrm{SOFA}},\ y_{\mathrm{SOFA}},\ z_{\mathrm{SOFA}}\bigr) = \bigl(y_{\mathrm{ADM}},\ -x_{\mathrm{ADM}},\ z_{\mathrm{ADM}}\bigr)$$
`CanonicalHrtf` 将 `Data.IR` 与 `Data.Delay` 分开保存。只有时域 baseline 才调用
`materialized_measurement()` 将 delay 应用一次;运行时 SH 路径不先 materialize。
非 48 kHz HRIR 使用 float64 `scipy.signal.resample_poly` 规范化,delay samples 按
相同比例缩放。
GeneralFIR、BRIR、TF、多 emitter、多义 receiver 或非 free-field 数据需要单独的
convention adapter,不能只通过 reshape 进入核心 importer。
## Delay/phase:exactly once
编译器只允许三种互斥语义:
1. 非零 `Data.Delay` 是 `Data.IR` 外部 delay;FIR 不去旋,运行时应用一次。
2. `Data.Delay=0` 且 HRIR 有普通正 onset:以每耳 main peak 分离 arrival,拟合后
在运行时恢复一次;当前阈值为 peak index 大于 2 samples。
3. `Data.Delay=0` 且双耳 FIR 共享 sample-0 起点:不发明外部 delay,原 complex
phase 直接进入五阶场。
任何路径都不能再叠加第二套 ear delay 或 phase-group delay。
## 公开 filterbank 与方向场
运行时固定为:
- 48 kHz;
- 64-sample QMF hop;
- 64-QMF / 77-hybrid;
- analysis/synthesis latency 961 samples;
- 五阶、36 项、ACN/N3D real spherical harmonics;
- PCM、delay、SH、room state 为 float64;频带传递和频域状态为 complex128。
每个 hybrid band 的 real/imaginary 单位增益都通过同一套 analysis/synthesis 链生成
脉冲字典,共 154 个实参数;编译不是直接读取 77 个 FFT bin。默认 projection
ridge 为 `1e-3`,SH ridge 为 `1e-5`。同方向 measurement 先合并,再用球面 Voronoi
面积权重做 ridge fit。
固定表位于 `data/rosella_kernels.npz`,实现公开标准化的滤波器组,各表可由如下
公式计算。
hybrid 分析核定义于 [3GPP TS 26.405 / ETSI TS 126 405](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
第 5.2.2 节(Table 1 的 $Q=8$/
$Q=4$ 系数,delay 6):
$$G_q^p[n] = g^p[n]\cdot\exp\Bigl(j\,\frac{2\pi}{Q^p}\bigl(q+\tfrac12\bigr)(n-6)\Bigr),\qquad n=0,\dots,12$$
QMF analysis 表即 MPEG-4 AAC/SBR(ISO/IEC 14496-3/AMD1:2003 第 4.B.18.2 节)
的 64 complex QMF bank;打包的 $64\times10$ 表是公开 640-tap prototype
$c_0,\dots,c_{639}$ 的多相重排:
$$A_{r,t} = \frac{(-1)^t}{128}\,c_{63-r+64t},\qquad r=0,\dots,63,\ t=0,\dots,9$$
QMF synthesis 表为上述 analysis 多相矩阵 $\mathbf{A}$ 的因果左逆,即求解
$\mathbf{A}\,\mathbf{W}=\mathbf{P}$(
$\mathbf{P}$ 为 577-sample 延迟置换;
全链 $961 = 577 + 6\times64$),以 rank-4 分解形式存储:
$$W_{b,l} = \sum_{r=1}^{4} t_{b,l,r}\,\mathbf{b}_{b,r}^{\top}$$
hybrid synthesis 表为 77→64 重组:高频带恒等 $Y_{3+b}=X_{16+b}$;低频带(
$C_p$ 为
$8+4+4$ 子带划分):
$$Y_p = \sum_{q\in C_p}\Bigl(\mathrm{Re}X_q + j\,s_q\,\mathrm{Im}X_q\Bigr),\qquad s_q\in\{\pm1\}$$
loader 校验 archive 和每个数组的 SHA-256;table version、所有数组 hash 与
77 个 band-center 参考值都属于 cache key。标准可公开获取不等于获准实施相关
专利;更多来源信息见 [`data/README.md`](../data/README.md) 与
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
## `.jochrtf`
`.jochrtf` 是无 pickle 的压缩 NumPy archive,固定包含:
| key | dtype / shape |
|---|---|
| `metadata_json` | 含 JSON 文本的 NumPy Unicode scalar(`dtype.kind == "U"`) |
| `band_center_frequencies_hz` | little-endian `float64[77]` |
| `coefficients` | little-endian `complex128[36,2,77]` |
| `delay_coefficients` | little-endian `float64[36,2]` |
| `delay_bounds` | little-endian `float64[2,2]` |
metadata magic 固定为 `JOC-HRTF-CACHE`,并记录 schema/compiler/phase-policy、
ACN/N3D、filterbank table hashes、SOFA content SHA-256、采样率、radius、order、
两个 ridge、payload hash 和 fit report。cache key 覆盖所有会改变编译结果的字段。
metadata 不保存本机绝对 `source_path`,仅可保存 source display name。
loader 使用 `allow_pickle=False`,并在构造对象前检查 ZIP 成员集、解压大小、shape、
dtype、端序、连续布局、有限值、delay bounds、band centers、payload hash 和 cache
key。writer 使用同目录临时文件、`fsync`、进程持有的 OS 文件锁和原子
`os.replace`;对应的隐藏 `.lock` sidecar 可保留,但不代表仍有 writer 持锁。
旧版本、损坏或配置不匹配的 cache 不能命中;从 SOFA 启动时会重建,显式 cache
入口则直接报错。
删除磁盘 cache 后,从同一 SOFA 和同一编译配置得到的场与渲染结果不得改变。
`.jochrtf` 包含由源 HRIR 变换得到的方向场系数与 delay 数据,因此“可以重建”不表示
它不受数据许可约束。生成 cache 不会扩大源 SOFA/HRTF 数据集授予的权利;cache 的
使用、复制和再分发仍须遵守源数据集条款。不能确认条款时,应把 `.jochrtf` 作为本地
私有 cache,不随程序或构建产物发布。`source_sha256` 只用于内容一致性校验,不是许可
或来源证明。
## JOC 对象与房间
生产适配器继续使用现有 JOC 调度:
- 每帧输入 `[1536,16]`;
- channel 0 是 special LFE,channel 1..15 是 JOC objects;
- ID11/OAMD position 使用 sample-timed timeline;
- 每 512 samples 更新方向/profile;
- 每个对象拥有独立 direct/early history,late FDN 全局共享;
- `finish()` 排空 early/late tail;输出增益显式应用,不隐含 limiter 或节目响度归一化。
Near/Mid/Far、equal-power direct level、六面 shoebox 一阶 image source、late send、
unitary FDN、LFE 120–180 Hz cosine-squared 低通及 room calibration 都是 JOC
项目定义行为,不是 SOFA 或 Dolby 公布常数。
双耳渲染在本库内实现:filterbank、SH 方向场求值、逐对象 early/direct 历史与共享 FDN
全部执行于 `joc_core`(`ejoc_sofa_binaural_*` 内核),对外只消费编译好的方向场与
512-sample 粒度的元数据更新。
## 技术引用与权利边界
- [SOFA SimpleFreeFieldHRIR convention](https://www.sofaconventions.org/mediawiki/index.php/SimpleFreeFieldHRIR)
- [3GPP TS 26.405 / ETSI TS 126 405(64-QMF/77-hybrid 定义)](https://www.etsi.org/deliver/etsi_ts/126400_126499/126405/06.00.00_60/ts_126405v060000p.pdf)
- [Dolby binaural render mode workflow](https://professionalsupport.dolby.com/s/article/What-is-Binaural-Render-Mode-and-how-do-the-settings-affect-my-mix)
- [EP3090576A1](https://patents.google.com/patent/EP3090576A1/en),仅作 direct/early/late、
subband 与 FDN 架构背景,不证明某个产品使用特定实施例。
规范、源码或专利文献可公开获取,不等于获准复制其内容、再分发派生产物或实施其中的
专利权利要求。本项目的技术引用本身不授予专利许可,也不作不侵权保证;准备发布或集成
到产品的一方应自行审查适用的数据许可、软件许可、专利许可及 freedom-to-operate。
公开标准来源与权利边界见
[`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)。
+607
View File
@@ -0,0 +1,607 @@
# JustOneCacophony — E-AC-3 JOC decoding and rendering mathematics
[中文](math.md) · [Back to README](../README.en.md)
This document covers only the signal model and formulas used in the JustOneCacophony research path: how JOC parameters combine with core PCM to reconstruct object signals, and how OAMD coordinates become speaker gains.
The formulas describe the JOC matrix parameters (both the dense and the sparse differential syntax) and the ordinary point-object paths studied by the project. They are not a complete definition of every E-AC-3 JOC variant.
## 1. Overall path and notation
Object reconstruction:
```text
E-AC-3 core 5.1 PCM
+ ID14 JOC matrix parameters
→ analysis QMF
→ parameter-band expansion and time interpolation
→ object matrix
→ inverse QMF
→ LFE + 15 object PCM channels
```
Speaker rendering:
```text
LFE + 15 object PCM channels
+ ID11 OAMD coordinates and update timing
→ target-layout region
→ equal-power panning
→ position compensation
→ sample-wise gain ramp
→ speaker PCM
```
Main notation:
| Symbol | Meaning |
|---|---|
| $c=0\ldots4$ | core channels L, R, C, Ls, Rs |
| $o=0\ldots14$ | 15 JOC objects |
| $b=0\ldots63$ | complex QMF subbands |
| $t=0\ldots23$ | 24 64-sample slots per frame |
| $p(b)$ | JOC parameter band corresponding to QMF subband $b$ |
| $X_{c,b,t}$ | analysis-QMF value for a core channel |
| $M_{o,c,b,t}$ | object-matrix coefficient |
| $Z_{o,b,t}$ | inverse-QMF input for an object |
| $y_o[n]$ | time-domain object PCM |
The number of samples in one frame is
$$
N_f=1536=24\times64.
$$
## 2. JOC matrix parameters
For every object and data point, the quantized matrix `joc_mix_mtx_q` is defined on $N_q$ quantization levels. The `b_joc_sparse` flag selects one of two differential syntaxes: dense sends one MTX difference per core channel, while sparse sends one active channel plus one coefficient difference per parameter band.
### 2.1 Dense differential reconstruction
Let `quant_idx` be $q_i\in\{0,1\}$. The number of quantization levels is
$$
N_q=
\begin{cases}
96, & q_i=0,\\
192, & q_i=1.
\end{cases}
$$
The center offset is
$$
O_q=\frac{N_q}{2}.
$$
For object $o$, data point $d$, core channel $c$, and parameter band $p$, the coded difference $\Delta_{o,d,c,p}$ reconstructs to
$$
Q_{o,d,c,0}=
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
$$
$$
Q_{o,d,c,p}=
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
\qquad p>0.
$$
### 2.2 Sparse differential reconstruction
Let $I_{o,d,p}$ be the `joc_channel_idx` symbol (IDX), $V_{o,d,p}$ the `joc_vec` symbol (VEC), and $N_c\in\lbrace5,7\rbrace$ the number of core channels. Each parameter band has exactly one active channel:
$$
A_{o,d,p}=
\begin{cases}
I_{o,d,0}, & p=0,\\
\left(A_{o,d,p-1}+I_{o,d,p}\right)\bmod N_c, & p>0,
\end{cases}
$$
where $I_{o,d,0}$ is a 3-bit absolute channel index and every later IDX symbol is an increment relative to the previous **active channel**. The coefficient is a single accumulator running across parameter bands:
$$
\kappa_{o,d,-1}=O^{(s)}_q,\qquad
\kappa_{o,d,p}=
\left(\kappa_{o,d,p-1}+V_{o,d,p}\right)\bmod N_q,
$$
with a sparse starting point two quantization levels above the dense center offset:
$$
O^{(s)}_q=
\begin{cases}
50, & q_i=0,\\
100, & q_i=1.
\end{cases}
$$
The accumulator is **not** reset when the active channel changes. The complete matrix is
$$
Q_{o,d,c,p}=
\begin{cases}
\kappa_{o,d,p}, & c=A_{o,d,p},\\
\dfrac{N_q}{2}, & c\neq A_{o,d,p}.
\end{cases}
$$
Non-active entries take $N_q/2$, which dequantizes to exactly 0.
### 2.3 Dequantization
The dequantized matrix coefficient is
$$
D_{o,d,c,p}=
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
\frac{820}{4096(1+q_i)}.
$$
The effective denominator is therefore 4096 in coarse mode and 8192 in fine mode.
### 2.4 JOC clipgain
If the clipgain field consists of integer $x$ and mantissa $y$, then
$$
G_{\mathrm{clip}}=
1+\frac{y}{32}2^{x-4}.
$$
It is applied to object PCM after inverse QMF and does not apply to LFE.
## 3. Parameter-band expansion and time interpolation
### 3.1 Parameter bands to QMF subbands
The JOC matrix is coded in parameter bands, while the QMF contains 64 subbands. Let $p(b)$ identify the parameter band containing subband $b$. A parameter-band coefficient expands as
$$
D_{o,d,c,b}=D_{o,d,c,p(b)}.
$$
The common 12-band mapping is
$$
\begin{aligned}
\mathcal B_0 &= \{0\}, &
\mathcal B_1 &= \{1\}, &
\mathcal B_2 &= \{2\}, &
\mathcal B_3 &= \{3\},\\
\mathcal B_4 &= \{4,5\}, &
\mathcal B_5 &= \{6,7\}, &
\mathcal B_6 &= \{8,9,10\}, &
\mathcal B_7 &= \{11,12,13\},\\
\mathcal B_8 &= \{14,15,16,17\}, &
\mathcal B_9 &= \{18,\ldots,22\},\\
\mathcal B_{10} &= \{23,\ldots,34\}, &
\mathcal B_{11} &= \{35,\ldots,63\}.
\end{aligned}
$$
Thus $p(b)=k$ if and only if $b\in\mathcal B_k$. Other parameter-band counts use their corresponding subband boundaries.
### 3.2 One-data-point interpolation
Let $P_{o,c,b}$ be the previous frame-end value and $D_{o,c,p(b)}$ the current target. For slot $t=0\ldots23$:
$$
\alpha_t=\frac{t+1}{24},
$$
$$
M_{o,c,b,t}=
(1-\alpha_t)P_{o,c,b}
+\alpha_tD_{o,c,p(b)}.
$$
The first slot has therefore advanced by $1/24$ of the ramp, while the last slot equals the current target:
$$
M_{o,c,b,23}=D_{o,c,p(b)}.
$$
This value then becomes the previous state for the next frame.
### 3.3 Multiple data points
When a frame contains two data points, `offset_ts` gives the segment boundary. Each segment uses the same linear relation between the previous and next targets; step mode switches targets at the designated slot.
## 4. Analysis QMF for core PCM
The matrix input uses core channels L, R, C, Ls, and Rs; LFE follows a separate path. Core PCM is first scaled as
$$
\widetilde x_c[n]=\frac{x_c[n]}{16}.
$$
Let $\mathcal A_b$ denote the 64-band analysis-QMF operator with polyphase history state. Then
$$
X_{c,b,t}=
\mathcal A_b\left(
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
\mathbf s^{\mathrm A}_{c,t}
\right).
$$
This consists of the analysis window/polyphase stage, modulation, a 64-point FFT, and subband reordering. History state advances continuously across slots and frames.
## 5. QMF-domain processing of core channels
L, R, and C are delayed by ten QMF slots before entering the object matrix:
$$
\widehat X_{c,b,t}=X_{c,b,t-10},
\qquad c\in\{L,R,C\}.
$$
Ls and Rs use the same ten-slot delay and a $-j$ rotation for $b>0$:
$$
\widehat X_{c,b,t}=-jX_{c,b,t-10},
\qquad c\in\{Ls,Rs\},\ b>0.
$$
Band 0 of each surround channel additionally passes through a 21-tap complex FIR:
$$
\widehat X_{c,0,t}=
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
$$
These delays and filter histories are decoder state and cannot be reset independently for every frame.
## 6. Object matrix
For each object $o$, subband $b$, and slot $t$, the object's frequency-domain value is a linear combination of the five core channels:
$$
Z_{o,b,t}=
\sum_{c=0}^{4}
M_{o,c,b,t}\widehat X_{c,b,t}.
$$
The $1/16$ analysis-input scale is canceled by the $\times16$ factor after inverse QMF, so the matrix itself needs no additional empirical gain.
## 7. Object inverse QMF
### 7.1 Subband reorder
Write the 64 complex subbands as 128 interleaved real values in `src`. For $k=0\ldots31$:
$$
\begin{aligned}
\mathrm{zone}[2k] &= \mathrm{src}[4k],\\
\mathrm{zone}[2k+1] &= -\mathrm{src}[4k+1],\\
\mathrm{zone}[126-2k] &= \mathrm{src}[4k+2],\\
\mathrm{zone}[127-2k] &= \mathrm{src}[4k+3].
\end{aligned}
$$
Treat `zone` as 64 complex values and apply an unnormalized 64-point FFT:
$$
F_k=
\sum_{n=0}^{63}
\mathrm{zone}_n
\exp\left(-j\frac{2\pi kn}{64}\right).
$$
### 7.2 Modulation and synthesis
Define the rotation coefficient
$$
r_k=
\frac12\left(
\sin\frac{\pi k}{128}
+j\cos\frac{\pi k}{128}
\right),
$$
and compute
$$
R_k=2F_kr_k.
$$
Let $\mathcal S$ denote polyphase synthesis with a 640-value synthesis window and cross-slot state:
$$
\mathbf y_{o,t}=
\mathcal S\left(
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
\right).
$$
Object output is
$$
y_o[64t+r]=
\mathrm{clip}\left(
16\,\mathbf y_{o,t}[r],-1,1
\right)G_{\mathrm{clip}},
$$
where $r=0\ldots63$. Synthesis state must advance continuously by slot.
## 8. LFE path
LFE bypasses the object matrix and inverse QMF and uses a 1217-sample delay. After the input and output scale factors cancel:
$$
y_{\mathrm{LFE}}[n]=
\mathrm{clip}\left(
x_{\mathrm{LFE,core}}[n-1217],-1,1
\right).
$$
## 9. OAMD coordinates
The lateral and longitudinal grids use $N=62$; the height grid uses $N=15$. The quantizer is
$$
q_N(k)=
\min\left(
32767,
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
\right).
$$
OAR coordinates are
$$
u=\frac{q_1}{32768},
\qquad
v=\frac{q_2}{32768},
\qquad
w=\frac{q_3}{32768}.
$$
Their maximum runtime value is $32767/32768$, not exactly 1.
For conversion to the ADM grid:
$$
k_1=\mathrm{round}\left(\frac{62q_1}{32767}\right),
\quad
k_2=\mathrm{round}\left(\frac{62q_2}{32767}\right),
\quad
k_3=\mathrm{round}\left(\frac{15q_3}{32767}\right),
$$
$$
X=2\frac{k_1}{62}-1,
\qquad
Y=1-2\frac{k_2}{62},
\qquad
Z=\frac{k_3}{15}.
$$
The continuous-coordinate relation is
$$
u=\frac{X+1}{2},
\qquad
v=\frac{1-Y}{2},
\qquad
w=Z.
$$
## 10. Equal-power speaker panning
### 10.1 One-dimensional interpolation
Let adjacent speaker coordinates be $a_0<a_1$ and object position be $a$. The normalized position is
$$
\tau=\frac{a-a_0}{a_1-a_0}.
$$
Gains inside the interval are
$$
g_0(\tau)=\cos\left(\frac\pi2\tau\right),
\qquad
g_1(\tau)=\sin\left(\frac\pi2\tau\right),
$$
and satisfy
$$
g_0^2(\tau)+g_1^2(\tau)=1.
$$
Object positions outside the interval are clamped to the nearest endpoint.
### 10.2 Two-dimensional regions
Each row first produces a horizontal gain vector $\mathbf h_r(u)$. If the object lies between adjacent rows $r_0,r_1$:
$$
\eta=\frac{v-v_{r_0}}{v_{r_1}-v_{r_0}},
$$
$$
a_0=\cos\left(\frac\pi2\eta\right),
\qquad
a_1=\sin\left(\frac\pi2\eta\right).
$$
The two-dimensional point gain is
$$
\mathbf G_{\mathrm{2D}}(u,v)=
\mathbf h(u)\odot\mathbf v(v).
$$
For 5.1-family layouts with one horizontal surround pair rather than separate side and rear pairs, the longitudinal coordinate is
$$
v_{\mathrm{floor}}=
\mathrm{clamp}(2v,0,1).
$$
Other layouts use $v_{\mathrm{floor}}=v$.
### 10.3 Height layer
Three-dimensional layouts compute floor gain $\mathbf G_f$ and height gain $\mathbf G_h$ separately:
$$
\mathbf G_{\mathrm{point}}(u,v,w)=
\cos\left(\frac\pi2w\right)\mathbf G_f
+
\sin\left(\frac\pi2w\right)\mathbf G_h.
$$
When the floor and height speaker sets do not overlap and each layer uses equal-power interpolation:
$$
\left\|\mathbf G_{\mathrm{point}}\right\|_2=1.
$$
## 11. Layout-dependent position compensation
Let $N_h$ be the number of relevant height speakers and $N_f$ the number of relevant additional horizontal speakers:
$$
H=\min\left(\frac{N_h}{4},1\right),
\qquad
F=\min\left(\frac{N_f}{4},1\right).
$$
Maximum position compensation is
$$
A_{\max}=
-\max\left(4.5-1.5H-3F,0\right)
\quad\text{dB}.
$$
Longitudinal and height weights are
$$
p_v=\mathrm{clamp}\left(\frac v{0.6},0,1\right),
$$
$$
p_w=\mathrm{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
$$
$$
p=\mathrm{clamp}(p_v+p_w,0,1).
$$
The linear compensation gain is
$$
G_{\mathrm{pos}}=10^{A_{\max}p/20}.
$$
The object's target-gain vector is
$$
\mathbf G_{\mathrm{target}}=
G_{\mathrm{object}}
G_{\mathrm{pos}}
\mathbf G_{\mathrm{point}}.
$$
## 12. OAMD time alignment and gain ramps
The coded position of an OAMD update is
$$
s_{\mathrm{coded}}=
s_{\mathrm{frame}}
+s_{\mathrm{outer}}
+s_{\mathrm{OAMD}}
+32f_{\mathrm{block}}.
$$
The theoretical update position on the decoder-output PCM timeline is
$$
s_{\mathrm{theoretical}}=
s_{\mathrm{coded}}+d_{\mathrm{decoder}},
\qquad d_{\mathrm{decoder}}=1473.
$$
The speaker renderer retains the existing processing-block length $B=32$, so the aligned update point is
$$
\widehat s=
B\left\lfloor
\frac{s_{\mathrm{theoretical}}+B/2-1}{B}
\right\rfloor.
$$
Thus, for frame-aligned updates, `align32(1473)=1472`. The 1473 value is the theoretical decoder delay at the metadata interface; 1472 is its effective boundary in the current 32-sample control block. The 640-value inverse-QMF window/state is not part of this metadata-timing formula.
For ramp duration $D$, the number of blocks is
$$
K=
\left\lfloor
\frac{D+B/2-1}{B}
\right\rfloor.
$$
If current gain is $g_0$ and target gain is $g_1$, the increment per block is
$$
\Delta g=\frac{g_1-g_0}{K}.
$$
Sample $r=0\ldots B-1$ of block $j$ uses
$$
g_{j,r}=g_j+\frac rB\Delta g,
\qquad
g_{j+1}=g_j+\Delta g.
$$
If no new metadata update intervenes, this is equivalent to a sample-wise linear ramp of total length $KB$.
## 13. Final speaker mix
For target output channel $c$:
$$
y_c[n]=
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
+
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
$$
Here
$$
\delta_{c,\mathrm{LFE}}=
\begin{cases}
1, & c\text{ is the target layout's LFE channel},\\
0, & \text{otherwise}.
\end{cases}
$$
A layout without LFE output does not mix input LFE into other channels. After object accumulation, output channels are ordered as required by the target format.
For PCM24 output, quantization is
$$
y_{24}[n]=
\mathrm{trunc}\left(
8388607\,\mathrm{clip}(y[n],-1,1)
\right).
$$
## 14. Scope of the formulas
- The JOC matrix section covers both the dense MTX and the sparse IDX/VEC differential syntax.
- The speaker-panning section describes ordinary point objects; extent, spread, divergence, and similar modes require additional models.
- Multiple OAMD position blocks must be scheduled in time order.
- A limiter is separate post-processing and is not included in the mixing equations above.
+607
View File
@@ -0,0 +1,607 @@
# JustOneCacophony — E-AC-3 JOC 解码与渲染数学
[English](math.en.md) · [返回 README](../README.md)
本文只说明 JustOneCacophony 研究路径中使用的信号模型和公式:JOC 参数如何与核心 PCM 结合并重建对象信号,以及 OAMD 坐标如何转换为扬声器增益。
这些公式描述项目当前研究的 JOC 矩阵参数(dense 与 sparse 两条差分语法)与普通点对象路径,不代表对所有 E-AC-3 JOC 变体的完整定义。
## 1. 总体路径与记号
对象重建路径:
```text
E-AC-3 核心 5.1 PCM
+ ID14 JOC 矩阵参数
→ analysis QMF
→ 参数带展开与时间插值
→ 对象矩阵
→ inverse QMF
→ LFE + 15 路对象 PCM
```
扬声器渲染路径:
```text
LFE + 15 路对象 PCM
+ ID11 OAMD 坐标与更新时间
→ 目标布局 region
→ 等功率声像
→ 位置补偿
→ 逐样本增益斜坡
→ 扬声器 PCM
```
主要记号:
| 符号 | 含义 |
|---|---|
| $c=0\ldots4$ | 核心声道 L、R、C、Ls、Rs |
| $o=0\ldots14$ | 15 个 JOC 对象 |
| $b=0\ldots63$ | 复 QMF 子带 |
| $t=0\ldots23$ | 每帧 24 个 64-sample 时槽 |
| $p(b)$ | QMF 子带 $b$ 对应的 JOC 参数带 |
| $X_{c,b,t}$ | 核心声道的 analysis-QMF 值 |
| $M_{o,c,b,t}$ | 对象矩阵系数 |
| $Z_{o,b,t}$ | 对象的 inverse-QMF 输入 |
| $y_o[n]$ | 对象时域 PCM |
一帧的采样数为
$$
N_f=1536=24\times64.
$$
## 2. JOC 矩阵参数
每个对象、每个数据点的量化矩阵 `joc_mix_mtx_q` 都定义在 $N_q$ 个量化级上。标志位 `b_joc_sparse` 选择两条差分语法之一:dense 为每个核心声道各送一路 MTX 差分,sparse 每参数带只送一个 active 声道与一路系数差分。
### 2.1 Dense 差分还原
令 `quant_idx` 为 $q_i\in\{0,1\}$,量化级数为
$$
N_q=
\begin{cases}
96, & q_i=0,\\
192, & q_i=1.
\end{cases}
$$
中心偏移为
$$
O_q=\frac{N_q}{2}.
$$
对对象 $o$、数据点 $d$、核心声道 $c$ 和参数带 $p$,编码差分 $\Delta_{o,d,c,p}$ 还原为
$$
Q_{o,d,c,0}=
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
$$
$$
Q_{o,d,c,p}=
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
\qquad p>0.
$$
### 2.2 Sparse 差分还原
令 $I_{o,d,p}$ 为 `joc_channel_idx` 符号(IDX), $V_{o,d,p}$ 为 `joc_vec` 符号(VEC), $N_c\in\lbrace5,7\rbrace$ 为核心声道数。每参数带只有一个 active 声道
$$
A_{o,d,p}=
\begin{cases}
I_{o,d,0}, & p=0,\\
\left(A_{o,d,p-1}+I_{o,d,p}\right)\bmod N_c, & p>0,
\end{cases}
$$
其中 $I_{o,d,0}$ 是 3 bit 绝对声道号,其余 IDX 符号是相对上一个 **active 声道**的增量。系数是一个跨参数带连续的单累加器
$$
\kappa_{o,d,-1}=O^{(s)}_q,\qquad
\kappa_{o,d,p}=
\left(\kappa_{o,d,p-1}+V_{o,d,p}\right)\bmod N_q,
$$
sparse 起点比 dense 的中心偏移高两个量化级:
$$
O^{(s)}_q=
\begin{cases}
50, & q_i=0,\\
100, & q_i=1.
\end{cases}
$$
active 声道切换时累加器**不**重置。完整矩阵为
$$
Q_{o,d,c,p}=
\begin{cases}
\kappa_{o,d,p}, & c=A_{o,d,p},\\
\dfrac{N_q}{2}, & c\neq A_{o,d,p}.
\end{cases}
$$
非 active 项取 $N_q/2$,即去量化后恰为 0。
### 2.3 去量化
矩阵系数的去量化值为
$$
D_{o,d,c,p}=
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
\frac{820}{4096(1+q_i)}.
$$
因此 coarse 模式的有效分母为 4096,fine 模式为 8192。
### 2.4 JOC clipgain
若 clipgain 字段由整数 $x$ 和尾数 $y$ 组成,则
$$
G_{\mathrm{clip}}=
1+\frac{y}{32}2^{x-4}.
$$
它在对象 inverse QMF 之后作用于对象 PCM,不作用于 LFE。
## 3. 参数带展开与时间插值
### 3.1 参数带到 QMF 子带
JOC 矩阵按参数带编码,而 QMF 使用 64 个子带。令 $p(b)$ 表示子带 $b$ 所属的参数带,则每个参数带系数展开为
$$
D_{o,d,c,b}=D_{o,d,c,p(b)}.
$$
常见的 12-band 映射为
$$
\begin{aligned}
\mathcal B_0 &= \{0\}, &
\mathcal B_1 &= \{1\}, &
\mathcal B_2 &= \{2\}, &
\mathcal B_3 &= \{3\},\\
\mathcal B_4 &= \{4,5\}, &
\mathcal B_5 &= \{6,7\}, &
\mathcal B_6 &= \{8,9,10\}, &
\mathcal B_7 &= \{11,12,13\},\\
\mathcal B_8 &= \{14,15,16,17\}, &
\mathcal B_9 &= \{18,\ldots,22\},\\
\mathcal B_{10} &= \{23,\ldots,34\}, &
\mathcal B_{11} &= \{35,\ldots,63\}.
\end{aligned}
$$
其中 $p(b)=k$ 当且仅当 $b\in\mathcal B_k$。其他参数带数使用各自的子带边界。
### 3.2 单数据点插值
令上一帧末值为 $P_{o,c,b}$,当前目标值为 $D_{o,c,p(b)}$。对时槽 $t=0\ldots23$:
$$
\alpha_t=\frac{t+1}{24},
$$
$$
M_{o,c,b,t}=
(1-\alpha_t)P_{o,c,b}
+\alpha_tD_{o,c,p(b)}.
$$
因此第一时槽已经推进 ramp 的 $1/24$,最后一时槽等于当前目标:
$$
M_{o,c,b,23}=D_{o,c,p(b)}.
$$
该值随后成为下一帧的 previous 状态。
### 3.3 多数据点
当一帧含两个数据点时,`offset_ts` 给出分段边界。每一段在上一目标和下一目标之间使用相同的线性关系;阶跃模式则在指定时槽直接切换目标。
## 4. 核心 PCM 的 analysis QMF
矩阵输入使用核心声道 L、R、C、Ls、Rs;LFE 走独立路径。核心 PCM 先缩放为
$$
\widetilde x_c[n]=\frac{x_c[n]}{16}.
$$
令 $\mathcal A_b$ 表示带 polyphase 历史状态的 64-band analysis-QMF 算子,则
$$
X_{c,b,t}=
\mathcal A_b\left(
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
\mathbf s^{\mathrm A}_{c,t}
\right).
$$
该过程依次包含 analysis window/polyphase、调制、64 点 FFT 和子带重排。历史状态跨时槽和帧连续推进。
## 5. 核心声道的 QMF 域处理
L、R、C 在进入对象矩阵前延迟 10 个 QMF 时槽:
$$
\widehat X_{c,b,t}=X_{c,b,t-10},
\qquad c\in\{L,R,C\}.
$$
Ls、Rs 同样延迟 10 个时槽,并在 $b>0$ 时作 $-j$ 旋转:
$$
\widehat X_{c,b,t}=-jX_{c,b,t-10},
\qquad c\in\{Ls,Rs\},\ b>0.
$$
环绕声道的 band 0 还经过 21-tap 复 FIR:
$$
\widehat X_{c,0,t}=
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
$$
这些延迟和滤波历史属于解码状态,不能按帧独立清零。
## 6. 对象矩阵
对每个对象 $o$、子带 $b$ 和时槽 $t$,对象频域值为五个核心声道的线性组合:
$$
Z_{o,b,t}=
\sum_{c=0}^{4}
M_{o,c,b,t}\widehat X_{c,b,t}.
$$
analysis 输入的 $1/16$ 缩放会在 inverse QMF 输出端由 $\times16$ 抵消,因此矩阵本身不需要额外经验增益。
## 7. 对象 inverse QMF
### 7.1 子带重排
将 64 个复子带写成 128 个交织实数 `src`。对 $k=0\ldots31$:
$$
\begin{aligned}
\mathrm{zone}[2k] &= \mathrm{src}[4k],\\
\mathrm{zone}[2k+1] &= -\mathrm{src}[4k+1],\\
\mathrm{zone}[126-2k] &= \mathrm{src}[4k+2],\\
\mathrm{zone}[127-2k] &= \mathrm{src}[4k+3].
\end{aligned}
$$
把 `zone` 重新视为 64 个复数后执行未归一化 64 点 FFT:
$$
F_k=
\sum_{n=0}^{63}
\mathrm{zone}_n
\exp\left(-j\frac{2\pi kn}{64}\right).
$$
### 7.2 调制与合成
定义旋转系数
$$
r_k=
\frac12\left(
\sin\frac{\pi k}{128}
+j\cos\frac{\pi k}{128}
\right),
$$
并计算
$$
R_k=2F_kr_k.
$$
令 $\mathcal S$ 表示带 640 项 synthesis window 和跨时槽状态的 polyphase 合成算子:
$$
\mathbf y_{o,t}=
\mathcal S\left(
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
\right).
$$
对象输出为
$$
y_o[64t+r]=
\mathrm{clip}\left(
16\,\mathbf y_{o,t}[r],-1,1
\right)G_{\mathrm{clip}},
$$
其中 $r=0\ldots63$。synthesis 状态必须按时槽连续推进。
## 8. LFE 路径
LFE 不经过对象矩阵或 inverse QMF,而是使用 1217-sample 延迟。输入与输出端的比例因子抵消后:
$$
y_{\mathrm{LFE}}[n]=
\mathrm{clip}\left(
x_{\mathrm{LFE,core}}[n-1217],-1,1
\right).
$$
## 9. OAMD 坐标
横向和纵向网格使用 $N=62$,高度网格使用 $N=15$。量化函数为
$$
q_N(k)=
\min\left(
32767,
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
\right).
$$
OAR 坐标为
$$
u=\frac{q_1}{32768},
\qquad
v=\frac{q_2}{32768},
\qquad
w=\frac{q_3}{32768}.
$$
其最大运行值为 $32767/32768$,不是精确的 1。
转换为 ADM 网格时:
$$
k_1=\mathrm{round}\left(\frac{62q_1}{32767}\right),
\quad
k_2=\mathrm{round}\left(\frac{62q_2}{32767}\right),
\quad
k_3=\mathrm{round}\left(\frac{15q_3}{32767}\right),
$$
$$
X=2\frac{k_1}{62}-1,
\qquad
Y=1-2\frac{k_2}{62},
\qquad
Z=\frac{k_3}{15}.
$$
连续坐标关系为
$$
u=\frac{X+1}{2},
\qquad
v=\frac{1-Y}{2},
\qquad
w=Z.
$$
## 10. 等功率扬声器声像
### 10.1 一维插值
相邻扬声器坐标为 $a_0<a_1$,对象位置为 $a$。归一化位置为
$$
\tau=\frac{a-a_0}{a_1-a_0}.
$$
区间内的增益为
$$
g_0(\tau)=\cos\left(\frac\pi2\tau\right),
\qquad
g_1(\tau)=\sin\left(\frac\pi2\tau\right),
$$
并满足
$$
g_0^2(\tau)+g_1^2(\tau)=1.
$$
区间外的对象位置夹到最近端点。
### 10.2 二维 region
每一行先沿 $u$ 得到横向增益向量 $\mathbf h_r(u)$。若对象位于相邻两行 $r_0,r_1$ 之间:
$$
\eta=\frac{v-v_{r_0}}{v_{r_1}-v_{r_0}},
$$
$$
a_0=\cos\left(\frac\pi2\eta\right),
\qquad
a_1=\sin\left(\frac\pi2\eta\right).
$$
二维点增益为
$$
\mathbf G_{\mathrm{2D}}(u,v)=
\mathbf h(u)\odot\mathbf v(v).
$$
对于只有一对水平环绕、没有独立 side/rear 两对的 5.1 系列布局,纵向坐标使用
$$
v_{\mathrm{floor}}=
\mathrm{clamp}(2v,0,1).
$$
其他布局使用 $v_{\mathrm{floor}}=v$。
### 10.3 高度层
三维布局分别计算地面层增益 $\mathbf G_f$ 和高度层增益 $\mathbf G_h$:
$$
\mathbf G_{\mathrm{point}}(u,v,w)=
\cos\left(\frac\pi2w\right)\mathbf G_f
+
\sin\left(\frac\pi2w\right)\mathbf G_h.
$$
当地面层与高度层扬声器集合不重叠、且各层内部使用等功率插值时:
$$
\left\|\mathbf G_{\mathrm{point}}\right\|_2=1.
$$
## 11. 布局位置补偿
令 $N_h$ 为相关高度扬声器数,$N_f$ 为相关附加水平扬声器数:
$$
H=\min\left(\frac{N_h}{4},1\right),
\qquad
F=\min\left(\frac{N_f}{4},1\right).
$$
最大位置补偿为
$$
A_{\max}=
-\max\left(4.5-1.5H-3F,0\right)
\quad\text{dB}.
$$
前后与高度位置权重为
$$
p_v=\mathrm{clamp}\left(\frac v{0.6},0,1\right),
$$
$$
p_w=\mathrm{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
$$
$$
p=\mathrm{clamp}(p_v+p_w,0,1).
$$
线性补偿增益为
$$
G_{\mathrm{pos}}=10^{A_{\max}p/20}.
$$
对象的目标增益向量为
$$
\mathbf G_{\mathrm{target}}=
G_{\mathrm{object}}
G_{\mathrm{pos}}
\mathbf G_{\mathrm{point}}.
$$
## 12. OAMD 时间对齐与增益斜坡
OAMD 更新的编码位置为
$$
s_{\mathrm{coded}}=
s_{\mathrm{frame}}
+s_{\mathrm{outer}}
+s_{\mathrm{OAMD}}
+32f_{\mathrm{block}}.
$$
decoder 输出 PCM timeline 上的理论更新位置为
$$
s_{\mathrm{theoretical}}=
s_{\mathrm{coded}}+d_{\mathrm{decoder}},
\qquad d_{\mathrm{decoder}}=1473.
$$
扬声器 renderer 保留现有的处理块长度 $B=32$,更新点对齐为
$$
\widehat s=
B\left\lfloor
\frac{s_{\mathrm{theoretical}}+B/2-1}{B}
\right\rfloor.
$$
因此,对 frame-aligned 更新有 `align32(1473)=1472`。1473 是 metadata interface 的理论 decoder delay;1472 是当前 32-sample control block 中的有效边界。inverse-QMF 使用的 640 项 window/state 不属于这条 metadata timing 公式。
给定 ramp duration $D$,block 数为
$$
K=
\left\lfloor
\frac{D+B/2-1}{B}
\right\rfloor.
$$
若当前增益为 $g_0$、目标为 $g_1$,则每 block 的增量为
$$
\Delta g=\frac{g_1-g_0}{K}.
$$
第 $j$ 个 block 内的样本 $r=0\ldots B-1$ 使用
$$
g_{j,r}=g_j+\frac rB\Delta g,
\qquad
g_{j+1}=g_j+\Delta g.
$$
如果中途没有新的 metadata 更新,该过程等价于总长度 $KB$ 的逐样本线性斜坡。
## 13. 最终扬声器混音
对目标输出声道 $c$:
$$
y_c[n]=
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
+
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
$$
其中
$$
\delta_{c,\mathrm{LFE}}=
\begin{cases}
1, & c\text{ 为目标布局的 LFE},\\
0, & \text{其他声道}.
\end{cases}
$$
没有 LFE 输出的布局不把输入 LFE 混入其他声道。对象完成累加后,再按目标格式要求排列输出声道。
若输出 PCM24,量化关系为
$$
y_{24}[n]=
\mathrm{trunc}\left(
8388607\,\mathrm{clip}(y[n],-1,1)
\right).
$$
## 14. 公式适用范围
- JOC 矩阵部分同时描述 dense MTX 与 sparse IDX/VEC 两条差分语法。
- 扬声器声像部分描述普通点对象;extent、spread、divergence 等模式需要额外模型。
- 多个 OAMD position block 必须按其时间顺序调度。
- limiter 属于独立后处理,不包含在上述混音公式中。
+128
View File
@@ -0,0 +1,128 @@
# SIMD and runtime dispatch
[中文](simd.md) · [Back to README](../README.en.md)
The heaviest loops in the binaural path (QMF analysis and synthesis, the hybrid
analysis low join, hybrid-domain path rendering, the ROOM FFT, the spherical
harmonic alignment) each have a runtime-dispatched vector implementation: one
binary carries several instruction-set variants, asks the CPU once at startup and
runs the widest one. **The output is byte-identical either way** — that is a hard
constraint, not a goal.
```text
JOC_SIMD=auto|scalar|sse2|avx2|avx512|neon pin one tier (used for verification)
JOC_SIMD_LOG=1 report the ISA each kernel actually got
```
## Why the split has to happen per translation unit
MSVC has no function-level attribute like `__attribute__((target("avx2")))`: one
`.cpp` file gets one `/arch`. So every ISA is its own translation unit with its own
`/arch:AVX2` or `/arch:AVX512` (GCC/Clang: `-mavx2` / `-mavx512f`), and
`dispatch.cpp` fills the function table at run time. The baseline units — the
dispatcher itself, the CPU probe and the scalar reference — carry **no** `/arch` at
all and stay on the SSE2 that x86-64 guarantees.
A trap from the history of this tree: `JOC_ENABLE_AVX2` used to be global, so
turning it on put AVX2 instructions into the very code paths that exist for older
CPUs. It now only selects whether the AVX2 unit is compiled in.
## Directory layout
One flat directory, **the instruction set in the file name and never in a
subdirectory** — that is how FLAC does it (`lpc.c` sits next to
`lpc_intrin_sse2.c`, `lpc_intrin_avx2.c` and `lpc_intrin_neon.c`, with the CPU
probe in its own `cpu.c`).
```text
src/simd/
simd.h the contract: Isa / Kernel / Kernels / dimensions
cpu_probe.{h,cpp} "can this machine run ISA X": CPUID+XGETBV / __builtin_cpu_supports / getauxval
dispatch.cpp policy: JOC_SIMD, the fallback ladder, the table, the log
kernels_scalar.cpp the reference (Isa::scalar; always built, always selectable)
kernels_intrin_avx2.cpp /arch:AVX2 -mavx2
kernels_intrin_avx512.cpp /arch:AVX512 -mavx512f
kernels_intrin_neon.cpp AArch64 default
```
The module lives at `src/simd/`, not `src/dsp/simd/`: `foundation/`, `hrtf/` and
`binaural/` all call into these kernels, so it is a cross-cutting layer rather than
a submodule of the DSP code (and `src/dsp/` held nothing else).
The header comment of `simd.h` carries the same module map; keep the two in sync
when the layout changes.
## Three rules
1. **Only `kernels_intrin_*.cpp` gets a wider flag.** `CMakeLists.txt` names those
files explicitly with `set_source_files_properties`; every other target stays on
the architecture's guaranteed ISA. A unit that goes wide without matching that
name fails `devtools/vec/isa_audit.ps1`.
2. **No dynamic initialisation inside an ISA unit.** Those objects are linked into
the same image as the baseline, so a global constructor would execute a wide
instruction before the dispatcher has looked at the CPU. Constant tables are
fine — they land in `.rdata`.
3. **Bit-exactness comes from the lane assignment, not from the ISA.** A lane may
only carry mutually independent outputs; the rounding sequence of a single
output, the separation of multiply and add (never an FMA) and the summation
order all stay exactly as `kernels_scalar.cpp` wrote them. Layout changes that
only reorder stored doubles (rank-minor basis tables, term-ordered tap tables,
stage-contiguous twiddle tables, band-major ROOM planes) are allowed.
## How the choice is made
`dispatch.cpp` parses `JOC_SIMD` first (forcing a tier this build or this machine
does not have prints a diagnostic and falls back, rather than pretending and
crashing), then walks `avx512 → avx2 → sse2 → neon` and picks, per kernel, the
widest implementation that is both compiled into this binary **and** runnable
here, falling back to the scalar reference. An AVX-512 unit therefore costs
nothing on a CPU without AVX-512; it is simply never selected.
On x86 the probe requires CPUID *and* XGETBV to agree: CPUID says the silicon can
do it, XCR0 says the OS saves the registers it needs. Either one alone is not
enough — using AVX without OS state support corrupts other threads across a
context switch. AArch64 needs no probe; ASIMD is the architectural baseline.
`sse2` is a selectable tier with **no unit of its own**, on purpose: a 128-bit SSE2
register is the register a scalar double already occupies, SSE2 cannot widen
double-precision arithmetic, and hand-written SSE2 would only add moves. The tier
resolves to the baseline unit.
## Effect
30-second reference cases, one binary with only `JOC_SIMD` switched (DSP stage,
`t_render_dsp`):
| Case | `scalar` | `auto` | DSP speed-up | End-to-end wall clock |
|---|---|---|---|---|
| Binaural Rosella | 1.256 s | **0.640 s** | **1.96×** | 1.690 → **0.941 s** |
| Binaural SOFA | 1.560 s | **0.654 s** | **2.39×** | 1.859 → **0.849 s** |
| Speaker 5.1 / 9.1.6 / ADM | — | — | 1.00× | no regression (these kernels are not on those paths) |
The vectorised loops themselves gain more: synthesis basis 5.89×, 13-tap low join
4.50×, SOFA QMF synthesis 4.27×, 128-point FFT 2.24×. The whole pipeline stops
short of 8× because a good part of the time goes to parameter setup, straight
copies and file writing — none of which has independent work items — and because
SOFA's 33 M sin/cos calls per sample cannot be vectorised under a byte-exactness
contract.
## Verifying a change
```powershell
$env:JOC_SIMD_LOG='1' # per-kernel ISA on this machine
$env:JOC_SIMD='scalar' # force the reference: hashes must not move
pwsh -NoProfile -File devtools\vec\isa_audit.ps1 # disassemble every .obj: 0 unguarded wide instructions
pwsh -NoProfile -File devtools\vec\sha_matrix.ps1 # 5 tiers x 2 renders against the reference digests
cmd /c devtools\vec\build_kernel_probe.bat # per-kernel byte digests (8 kernels)
```
## Adding an ISA
1. Write `kernels_intrin_<isa>.cpp`, implementing the slots you have and leaving the
rest `nullptr` — the dispatcher falls back per kernel (that is how
`qmf_synthesis_basis` is handled in the AVX-512 unit).
2. Add it to `JOC_SIMD_SOURCES` in `CMakeLists.txt` with `JOC_SIMD_HAVE_<ISA>=1` and
its flag, and extend `Isa`, `isa_rank`, `isa_compiled`, `isa_supported` and the
`JOC_SIMD` name table in `simd.h` / `dispatch.cpp`.
3. Verify: the kernel digests must match the scalar unit byte for byte, and the
reference renders must keep their SHA-256 on every tier.
+107
View File
@@ -0,0 +1,107 @@
# SIMD 与运行时派发
[English](simd.en.md) · [返回 README](../README.md)
双耳通路里最重的那几段循环(QMF 分析/合成、混合分析的低频拼接、混合域路径渲染、
ROOM 的 FFT、球谐对齐)都有一份运行时分派的向量实现:同一份二进制里装多套 ISA 代码,
启动时问一次 CPU,然后选最宽的那套跑。**输出逐字节不变**——这是硬约束,不是目标。
```text
JOC_SIMD=auto|scalar|sse2|avx2|avx512|neon 强制某一层(验收用)
JOC_SIMD_LOG=1 打印每个 kernel 实际生效的 ISA
```
## 为什么必须"按编译单元分 ISA"
MSVC 没有 `__attribute__((target("avx2")))` 这类函数级多版本能力,一个 .cpp 只能有
一个 `/arch`。所以每个 ISA 一个编译单元,各自带自己的 `/arch:AVX2` / `/arch:AVX512`
(GCC/Clang 是 `-mavx2` / `-mavx512f`),由 `dispatch.cpp` 在运行时填函数表。
基线单元(含派发器本身、CPU 探测、标量参考实现)**不带任何 `/arch`**,它们只使用
x86-64 架构保证的 SSE2。
历史坑:早先的 `JOC_ENABLE_AVX2` 是**全局**的,一旦打开,连"给老 CPU 用"的基线路径
都带 AVX2 指令。现在这个选项只决定是否把 AVX2 单元编进二进制。
## 目录布局
一个扁平目录,**ISA 写在文件名里,不写进子目录**——这是 FLAC 的做法
(`src/libFLAC/lpc.c` 旁边就是 `lpc_intrin_sse2.c` / `lpc_intrin_avx2.c` /
`lpc_intrin_neon.c`,CPU 探测单独放在 `cpu.c`)。
```text
src/simd/
simd.h 唯一契约头:Isa / Kernel / Kernels / 维度常量
cpu_probe.{h,cpp} "这台机器能不能跑 ISA X":CPUID+XGETBV / __builtin_cpu_supports / getauxval
dispatch.cpp 策略:JOC_SIMD 解析、回退阶梯、填函数表、日志
kernels_scalar.cpp 参考实现(Isa::scalar,永远编译、永远可选中)
kernels_intrin_avx2.cpp /arch:AVX2 -mavx2
kernels_intrin_avx512.cpp /arch:AVX512 -mavx512f
kernels_intrin_neon.cpp AArch64 默认 -march=armv8-a+simd
```
`simd.h` 的头注释里有一份同样的模块地图,改布局时两处一起改。
模块放在 `src/simd/` 而不是 `src/dsp/simd/`:这些 kernel 被 `foundation/`、`hrtf/`、
`binaural/` 三个模块共用,是横切的一层,不是 DSP 的子模块(更何况 `src/dsp/` 里除了
`simd/` 空无一物)。
## 三条规则
1. **只有 `kernels_intrin_*.cpp` 拿更宽的编译开关。** `CMakeLists.txt` 用
`set_source_files_properties` 逐个点名,其余目标一律留在架构保证的 ISA 上。
任何不属于这个命名却带了宽指令的单元都会被 `devtools/vec/isa_audit.ps1` 判失败。
2. **ISA 单元里不许有动态初始化。** 它们和基线代码链进同一个镜像,全局构造函数会在
派发器看 CPU 之前就跑宽指令。常量表没问题(落在 `.rdata`)。
3. **逐位一致靠的是 lane 的划分,不是 ISA。** lane 里只能放**互相独立**的输出;单个输出
的舍入序列、乘加分离(绝不用 FMA)、求和顺序都保持 `kernels_scalar.cpp` 原样。
只改变 double **存放顺序**的布局改造(基函数表转秩小序、抽头表按项序、蝶形因子表
按级连续化、ROOM 谱平面改频带主序)是允许的。
## 运行时怎么选
`dispatch.cpp` 先解析 `JOC_SIMD`(强制一个本机不支持的层会打印诊断并回退,而不是假装
选中然后崩),再走阶梯 `avx512 → avx2 → sse2 → neon`,每个 kernel 单独挑"已编进本
二进制 **且** 本机可跑"的最宽实现,挑不到就落到标量参考实现。所以 AVX-512 单元在
不支持它的 CPU 上只是不被选中,不影响启动。
x86 的探测要 CPUID 与 XGETBV **同时**成立:CPUID 说明硅片有这个能力,XCR0 说明操作
系统会保存对应寄存器状态,缺一个就不能用(否则上下文切换会踩坏别的线程)。AArch64
不需要探测,ASIMD 是架构基线。
`sse2` 是一个有意保留的档位但**没有单独的单元**:128 位 SSE2 寄存器就是标量 double
已经在用的寄存器,SSE2 加宽不了双精度运算,手写只会多出搬运指令,所以它选中的是基线
单元。
## 效果
30 s 参考用例,同一二进制只切 `JOC_SIMD`(DSP 阶段 `t_render_dsp`):
| 用例 | `scalar` | `auto` | DSP 加速 | 端到端墙钟 |
|---|---|---|---|---|
| 双耳 Rosella | 1.256 s | **0.640 s** | **1.96×** | 1.690 → **0.941 s** |
| 双耳 SOFA | 1.560 s | **0.654 s** | **2.39×** | 1.859 → **0.849 s** |
| 扬声器 5.1 / 9.1.6 / ADM | — | — | 1.00× | 0 回归(不走这些 kernel) |
单看被向量化的循环,红利更大:合成基函数 5.89×、13 抽头低频拼接 4.50×、
SOFA QMF 合成 4.27×、128 点 FFT 2.24×。整条流水线到不了 8×,是因为相当一部分时间在
参数设置、直通拷贝、写盘这些没有独立工作项的代码上,以及 SOFA 每次采样的 33 M 次
sin/cos 按逐位契约不能向量化。
## 验证
```powershell
$env:JOC_SIMD_LOG='1' # 本机每个 kernel 实际选中的 ISA
$env:JOC_SIMD='scalar' # 强制参考路径:哈希必须一动不动
pwsh -NoProfile -File devtools\vec\isa_audit.ps1 # 反汇编全部 .obj:0 个无守卫的宽指令
pwsh -NoProfile -File devtools\vec\sha_matrix.ps1 # 5 档 × 2 渲染,逐字节比对参考摘要
cmd /c devtools\vec\build_kernel_probe.bat # kernel 级逐字节摘要(8 个 kernel)
```
## 加一个新的 ISA
1. 写 `kernels_intrin_<isa>.cpp`,实现能实现的槽位,其余留 `nullptr`——派发器会逐
kernel 回退(AVX-512 单元里的 `qmf_synthesis_basis` 就是这么处理的)。
2. 在 `CMakeLists.txt` 里加进 `JOC_SIMD_SOURCES`、`JOC_SIMD_HAVE_<ISA>=1` 和它的编译
开关,并在 `simd.h` / `dispatch.cpp` 里补上 `Isa`、`isa_rank`、`isa_compiled`、
`isa_supported` 与 `JOC_SIMD` 名字表。
3. 验证:kernel 摘要必须与标量逐字节相同,参考渲染在每个档位上的 SHA-256 都不能变。