Initial public release of JustOneCacophony

This commit is contained in:
TheM14
2026-09-01 16:31:38 +08:00
commit 6536782a0f
38 changed files with 9073 additions and 0 deletions
+578
View File
@@ -0,0 +1,578 @@
# JustOneCacophony — E-AC-3 JOC decoding and rendering mathematics
[中文](math.md) · [Back to README](../README.en.md)
This document covers only the signal model and formulas used in the JustOneCacophony research path: how JOC parameters combine with core PCM to reconstruct object signals, and how OAMD coordinates become speaker gains.
The formulas describe the dense-JOC and ordinary point-object paths studied by the project. They are not a complete definition of every E-AC-3 JOC variant.
## 1. Overall path and notation
Object reconstruction:
```text
E-AC-3 core 5.1 PCM
+ ID14 JOC matrix parameters
→ analysis QMF
→ parameter-band expansion and time interpolation
→ object matrix
→ inverse QMF
→ LFE + 15 object PCM channels
```
Speaker rendering:
```text
LFE + 15 object PCM channels
+ ID11 OAMD coordinates and update timing
→ target-layout region
→ equal-power panning
→ position compensation
→ sample-wise gain ramp
→ speaker PCM
```
Main notation:
| Symbol | Meaning |
|---|---|
| $c=0\ldots4$ | core channels L, R, C, Ls, Rs |
| $o=0\ldots14$ | 15 JOC objects |
| $b=0\ldots63$ | complex QMF subbands |
| $t=0\ldots23$ | 24 64-sample slots per frame |
| $p(b)$ | JOC parameter band corresponding to QMF subband $b$ |
| $X_{c,b,t}$ | analysis-QMF value for a core channel |
| $M_{o,c,b,t}$ | object-matrix coefficient |
| $Z_{o,b,t}$ | inverse-QMF input for an object |
| $y_o[n]$ | time-domain object PCM |
The number of samples in one frame is
$$
N_f=1536=24\times64.
$$
## 2. Dense-JOC matrix parameters
### 2.1 Differential reconstruction
Let `quant_idx` be $q_i\in\{0,1\}$. The number of quantization levels is
$$
N_q=
\begin{cases}
96, & q_i=0,\\
192, & q_i=1.
\end{cases}
$$
The center offset is
$$
O_q=\frac{N_q}{2}.
$$
For object $o$, data point $d$, core channel $c$, and parameter band $p$, the coded difference $\Delta_{o,d,c,p}$ reconstructs to
$$
Q_{o,d,c,0}
=
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
$$
$$
Q_{o,d,c,p}
=
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
\qquad p>0.
$$
### 2.2 Dequantization
The dequantized matrix coefficient is
$$
D_{o,d,c,p}
=
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
\frac{820}{4096(1+q_i)}.
$$
The effective denominator is therefore 4096 in coarse mode and 8192 in fine mode.
### 2.3 JOC clipgain
If the clipgain field consists of integer $x$ and mantissa $y$, then
$$
G_{\mathrm{clip}}
=
1+\frac{y}{32}2^{x-4}.
$$
It is applied to object PCM after inverse QMF and does not apply to LFE.
## 3. Parameter-band expansion and time interpolation
### 3.1 Parameter bands to QMF subbands
The JOC matrix is coded in parameter bands, while the QMF contains 64 subbands. Let $p(b)$ identify the parameter band containing subband $b$. A parameter-band coefficient expands as
$$
D_{o,d,c,b}=D_{o,d,c,p(b)}.
$$
The common 12-band mapping is
$$
\begin{aligned}
\mathcal B_0 &= \{0\}, &
\mathcal B_1 &= \{1\}, &
\mathcal B_2 &= \{2\}, &
\mathcal B_3 &= \{3\},\\
\mathcal B_4 &= \{4,5\}, &
\mathcal B_5 &= \{6,7\}, &
\mathcal B_6 &= \{8,9,10\}, &
\mathcal B_7 &= \{11,12,13\},\\
\mathcal B_8 &= \{14,15,16,17\}, &
\mathcal B_9 &= \{18,\ldots,22\},\\
\mathcal B_{10} &= \{23,\ldots,34\}, &
\mathcal B_{11} &= \{35,\ldots,63\}.
\end{aligned}
$$
Thus $p(b)=k$ if and only if $b\in\mathcal B_k$. Other parameter-band counts use their corresponding subband boundaries.
### 3.2 One-data-point interpolation
Let $P_{o,c,b}$ be the previous frame-end value and $D_{o,c,p(b)}$ the current target. For slot $t=0\ldots23$:
$$
\alpha_t=\frac{t+1}{24},
$$
$$
M_{o,c,b,t}
=
(1-\alpha_t)P_{o,c,b}
+\alpha_tD_{o,c,p(b)}.
$$
The first slot has therefore advanced by $1/24$ of the ramp, while the last slot equals the current target:
$$
M_{o,c,b,23}=D_{o,c,p(b)}.
$$
This value then becomes the previous state for the next frame.
### 3.3 Multiple data points
When a frame contains two data points, `offset_ts` gives the segment boundary. Each segment uses the same linear relation between the previous and next targets; step mode switches targets at the designated slot.
## 4. Analysis QMF for core PCM
The matrix input uses core channels L, R, C, Ls, and Rs; LFE follows a separate path. Core PCM is first scaled as
$$
\widetilde x_c[n]=\frac{x_c[n]}{16}.
$$
Let $\mathcal A_b$ denote the 64-band analysis-QMF operator with polyphase history state. Then
$$
X_{c,b,t}
=
\mathcal A_b\!\left(
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
\mathbf s^{\mathrm A}_{c,t}
\right).
$$
This consists of the analysis window/polyphase stage, modulation, a 64-point FFT, and subband reordering. History state advances continuously across slots and frames.
## 5. QMF-domain processing of core channels
L, R, and C are delayed by ten QMF slots before entering the object matrix:
$$
\widehat X_{c,b,t}=X_{c,b,t-10},
\qquad c\in\{L,R,C\}.
$$
Ls and Rs use the same ten-slot delay and a $-j$ rotation for $b>0$:
$$
\widehat X_{c,b,t}=-jX_{c,b,t-10},
\qquad c\in\{Ls,Rs\},\ b>0.
$$
Band 0 of each surround channel additionally passes through a 21-tap complex FIR:
$$
\widehat X_{c,0,t}
=
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
$$
These delays and filter histories are decoder state and cannot be reset independently for every frame.
## 6. Object matrix
For each object $o$, subband $b$, and slot $t$, the object's frequency-domain value is a linear combination of the five core channels:
$$
Z_{o,b,t}
=
\sum_{c=0}^{4}
M_{o,c,b,t}\widehat X_{c,b,t}.
$$
The $1/16$ analysis-input scale is canceled by the $\times16$ factor after inverse QMF, so the matrix itself needs no additional empirical gain.
## 7. Object inverse QMF
### 7.1 Subband reorder
Write the 64 complex subbands as 128 interleaved real values in `src`. For $k=0\ldots31$:
$$
\begin{aligned}
\operatorname{zone}[2k] &= \operatorname{src}[4k],\\
\operatorname{zone}[2k+1] &= -\operatorname{src}[4k+1],\\
\operatorname{zone}[126-2k] &= \operatorname{src}[4k+2],\\
\operatorname{zone}[127-2k] &= \operatorname{src}[4k+3].
\end{aligned}
$$
Treat `zone` as 64 complex values and apply an unnormalized 64-point FFT:
$$
F_k
=
\sum_{n=0}^{63}
\operatorname{zone}_n
\exp\!\left(-j\frac{2\pi kn}{64}\right).
$$
### 7.2 Modulation and synthesis
Define the rotation coefficient
$$
r_k
=
\frac12\left(
\sin\frac{\pi k}{128}
+j\cos\frac{\pi k}{128}
\right),
$$
and compute
$$
R_k=2F_kr_k.
$$
Let $\mathcal S$ denote polyphase synthesis with a 640-value synthesis window and cross-slot state:
$$
\mathbf y_{o,t}
=
\mathcal S\!\left(
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
\right).
$$
Object output is
$$
y_o[64t+r]
=
\operatorname{clip}\!\left(
16\,\mathbf y_{o,t}[r],-1,1
\right)G_{\mathrm{clip}},
$$
where $r=0\ldots63$. Synthesis state must advance continuously by slot.
## 8. LFE path
LFE bypasses the object matrix and inverse QMF and uses a 1217-sample delay. After the input and output scale factors cancel:
$$
y_{\mathrm{LFE}}[n]
=
\operatorname{clip}\!\left(
x_{\mathrm{LFE,core}}[n-1217],-1,1
\right).
$$
## 9. OAMD coordinates
The lateral and longitudinal grids use $N=62$; the height grid uses $N=15$. The quantizer is
$$
q_N(k)
=
\min\!\left(
32767,
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
\right).
$$
OAR coordinates are
$$
u=\frac{q_1}{32768},
\qquad
v=\frac{q_2}{32768},
\qquad
w=\frac{q_3}{32768}.
$$
Their maximum runtime value is $32767/32768$, not exactly 1.
For conversion to the ADM grid:
$$
k_1=\operatorname{round}\!\left(\frac{62q_1}{32767}\right),
\quad
k_2=\operatorname{round}\!\left(\frac{62q_2}{32767}\right),
\quad
k_3=\operatorname{round}\!\left(\frac{15q_3}{32767}\right),
$$
$$
X=2\frac{k_1}{62}-1,
\qquad
Y=1-2\frac{k_2}{62},
\qquad
Z=\frac{k_3}{15}.
$$
The continuous-coordinate relation is
$$
u=\frac{X+1}{2},
\qquad
v=\frac{1-Y}{2},
\qquad
w=Z.
$$
## 10. Equal-power speaker panning
### 10.1 One-dimensional interpolation
Let adjacent speaker coordinates be $a_0<a_1$ and object position be $a$. The normalized position is
$$
\tau=\frac{a-a_0}{a_1-a_0}.
$$
Gains inside the interval are
$$
g_0(\tau)=\cos\left(\frac\pi2\tau\right),
\qquad
g_1(\tau)=\sin\left(\frac\pi2\tau\right),
$$
and satisfy
$$
g_0^2(\tau)+g_1^2(\tau)=1.
$$
Object positions outside the interval are clamped to the nearest endpoint.
### 10.2 Two-dimensional regions
Each row first produces a horizontal gain vector $\mathbf h_r(u)$. If the object lies between adjacent rows $r_0,r_1$:
$$
\eta=\frac{v-v_{r_0}}{v_{r_1}-v_{r_0}},
$$
$$
a_0=\cos\left(\frac\pi2\eta\right),
\qquad
a_1=\sin\left(\frac\pi2\eta\right).
$$
The two-dimensional point gain is
$$
\mathbf G_{\mathrm{2D}}(u,v)
=
\mathbf h(u)\odot\mathbf v(v).
$$
For 5.1-family layouts with one horizontal surround pair rather than separate side and rear pairs, the longitudinal coordinate is
$$
v_{\mathrm{floor}}
=
\operatorname{clamp}(2v,0,1).
$$
Other layouts use $v_{\mathrm{floor}}=v$.
### 10.3 Height layer
Three-dimensional layouts compute floor gain $\mathbf G_f$ and height gain $\mathbf G_h$ separately:
$$
\mathbf G_{\mathrm{point}}(u,v,w)
=
\cos\left(\frac\pi2w\right)\mathbf G_f
+
\sin\left(\frac\pi2w\right)\mathbf G_h.
$$
When the floor and height speaker sets do not overlap and each layer uses equal-power interpolation:
$$
\left\|\mathbf G_{\mathrm{point}}\right\|_2=1.
$$
## 11. Layout-dependent position compensation
Let $N_h$ be the number of relevant height speakers and $N_f$ the number of relevant additional horizontal speakers:
$$
H=\min\left(\frac{N_h}{4},1\right),
\qquad
F=\min\left(\frac{N_f}{4},1\right).
$$
Maximum position compensation is
$$
A_{\max}
=
-\max\left(4.5-1.5H-3F,0\right)
\quad\text{dB}.
$$
Longitudinal and height weights are
$$
p_v=\operatorname{clamp}\left(\frac v{0.6},0,1\right),
$$
$$
p_w=\operatorname{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
$$
$$
p=\operatorname{clamp}(p_v+p_w,0,1).
$$
The linear compensation gain is
$$
G_{\mathrm{pos}}=10^{A_{\max}p/20}.
$$
The object's target-gain vector is
$$
\mathbf G_{\mathrm{target}}
=
G_{\mathrm{object}}
G_{\mathrm{pos}}
\mathbf G_{\mathrm{point}}.
$$
## 12. OAMD time alignment and gain ramps
The coded position of an OAMD update is
$$
s_{\mathrm{coded}}
=
s_{\mathrm{frame}}
+s_{\mathrm{outer}}
+s_{\mathrm{OAMD}}
+32f_{\mathrm{block}}.
$$
For processing-block length $B=32$, the aligned update point is
$$
\widehat s
=
B\left\lfloor
\frac{s_{\mathrm{coded}}+B/2-1}{B}
\right\rfloor.
$$
For ramp duration $D$, the number of blocks is
$$
K
=
\left\lfloor
\frac{D+B/2-1}{B}
\right\rfloor.
$$
If current gain is $g_0$ and target gain is $g_1$, the increment per block is
$$
\Delta g=\frac{g_1-g_0}{K}.
$$
Sample $r=0\ldots B-1$ of block $j$ uses
$$
g_{j,r}=g_j+\frac rB\Delta g,
\qquad
g_{j+1}=g_j+\Delta g.
$$
If no new metadata update intervenes, this is equivalent to a sample-wise linear ramp of total length $KB$.
## 13. Final speaker mix
For target output channel $c$:
$$
y_c[n]
=
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
+
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
$$
Here
$$
\delta_{c,\mathrm{LFE}}
=
\begin{cases}
1, & c\text{ is the target layout's LFE channel},\\
0, & \text{otherwise}.
\end{cases}
$$
A layout without LFE output does not mix input LFE into other channels. After object accumulation, output channels are ordered as required by the target format.
For PCM24 output, quantization is
$$
y_{24}[n]
=
\operatorname{trunc}\left(
8388607\,\operatorname{clip}(y[n],-1,1)
\right).
$$
## 14. Scope of the formulas
- The JOC matrix section describes dense JOC; Sparse JOC uses a different sparse coefficient/index path.
- The speaker-panning section describes ordinary point objects; extent, spread, divergence, and similar modes require additional models.
- Multiple OAMD position blocks must be scheduled in time order.
- A limiter is separate post-processing and is not included in the mixing equations above.
+578
View File
@@ -0,0 +1,578 @@
# JustOneCacophony — E-AC-3 JOC 解码与渲染数学
[English](math.en.md) · [返回 README](../README.md)
本文只说明 JustOneCacophony 研究路径中使用的信号模型和公式:JOC 参数如何与核心 PCM 结合并重建对象信号,以及 OAMD 坐标如何转换为扬声器增益。
这些公式描述项目当前研究的 dense JOC 与普通点对象路径,不代表对所有 E-AC-3 JOC 变体的完整定义。
## 1. 总体路径与记号
对象重建路径:
```text
E-AC-3 核心 5.1 PCM
+ ID14 JOC 矩阵参数
→ analysis QMF
→ 参数带展开与时间插值
→ 对象矩阵
→ inverse QMF
→ LFE + 15 路对象 PCM
```
扬声器渲染路径:
```text
LFE + 15 路对象 PCM
+ ID11 OAMD 坐标与更新时间
→ 目标布局 region
→ 等功率声像
→ 位置补偿
→ 逐样本增益斜坡
→ 扬声器 PCM
```
主要记号:
| 符号 | 含义 |
|---|---|
| $c=0\ldots4$ | 核心声道 L、R、C、Ls、Rs |
| $o=0\ldots14$ | 15 个 JOC 对象 |
| $b=0\ldots63$ | 复 QMF 子带 |
| $t=0\ldots23$ | 每帧 24 个 64-sample 时槽 |
| $p(b)$ | QMF 子带 $b$ 对应的 JOC 参数带 |
| $X_{c,b,t}$ | 核心声道的 analysis-QMF 值 |
| $M_{o,c,b,t}$ | 对象矩阵系数 |
| $Z_{o,b,t}$ | 对象的 inverse-QMF 输入 |
| $y_o[n]$ | 对象时域 PCM |
一帧的采样数为
$$
N_f=1536=24\times64.
$$
## 2. Dense JOC 矩阵参数
### 2.1 差分还原
令 `quant_idx` 为 $q_i\in\{0,1\}$,量化级数为
$$
N_q=
\begin{cases}
96, & q_i=0,\\
192, & q_i=1.
\end{cases}
$$
中心偏移为
$$
O_q=\frac{N_q}{2}.
$$
对对象 $o$、数据点 $d$、核心声道 $c$ 和参数带 $p$,编码差分 $\Delta_{o,d,c,p}$ 还原为
$$
Q_{o,d,c,0}
=
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
$$
$$
Q_{o,d,c,p}
=
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
\qquad p>0.
$$
### 2.2 去量化
矩阵系数的去量化值为
$$
D_{o,d,c,p}
=
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
\frac{820}{4096(1+q_i)}.
$$
因此 coarse 模式的有效分母为 4096,fine 模式为 8192。
### 2.3 JOC clipgain
若 clipgain 字段由整数 $x$ 和尾数 $y$ 组成,则
$$
G_{\mathrm{clip}}
=
1+\frac{y}{32}2^{x-4}.
$$
它在对象 inverse QMF 之后作用于对象 PCM,不作用于 LFE。
## 3. 参数带展开与时间插值
### 3.1 参数带到 QMF 子带
JOC 矩阵按参数带编码,而 QMF 使用 64 个子带。令 $p(b)$ 表示子带 $b$ 所属的参数带,则每个参数带系数展开为
$$
D_{o,d,c,b}=D_{o,d,c,p(b)}.
$$
常见的 12-band 映射为
$$
\begin{aligned}
\mathcal B_0 &= \{0\}, &
\mathcal B_1 &= \{1\}, &
\mathcal B_2 &= \{2\}, &
\mathcal B_3 &= \{3\},\\
\mathcal B_4 &= \{4,5\}, &
\mathcal B_5 &= \{6,7\}, &
\mathcal B_6 &= \{8,9,10\}, &
\mathcal B_7 &= \{11,12,13\},\\
\mathcal B_8 &= \{14,15,16,17\}, &
\mathcal B_9 &= \{18,\ldots,22\},\\
\mathcal B_{10} &= \{23,\ldots,34\}, &
\mathcal B_{11} &= \{35,\ldots,63\}.
\end{aligned}
$$
其中 $p(b)=k$ 当且仅当 $b\in\mathcal B_k$。其他参数带数使用各自的子带边界。
### 3.2 单数据点插值
令上一帧末值为 $P_{o,c,b}$,当前目标值为 $D_{o,c,p(b)}$。对时槽 $t=0\ldots23$:
$$
\alpha_t=\frac{t+1}{24},
$$
$$
M_{o,c,b,t}
=
(1-\alpha_t)P_{o,c,b}
+\alpha_tD_{o,c,p(b)}.
$$
因此第一时槽已经推进 ramp 的 $1/24$,最后一时槽等于当前目标:
$$
M_{o,c,b,23}=D_{o,c,p(b)}.
$$
该值随后成为下一帧的 previous 状态。
### 3.3 多数据点
当一帧含两个数据点时,`offset_ts` 给出分段边界。每一段在上一目标和下一目标之间使用相同的线性关系;阶跃模式则在指定时槽直接切换目标。
## 4. 核心 PCM 的 analysis QMF
矩阵输入使用核心声道 L、R、C、Ls、Rs;LFE 走独立路径。核心 PCM 先缩放为
$$
\widetilde x_c[n]=\frac{x_c[n]}{16}.
$$
令 $\mathcal A_b$ 表示带 polyphase 历史状态的 64-band analysis-QMF 算子,则
$$
X_{c,b,t}
=
\mathcal A_b\!\left(
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
\mathbf s^{\mathrm A}_{c,t}
\right).
$$
该过程依次包含 analysis window/polyphase、调制、64 点 FFT 和子带重排。历史状态跨时槽和帧连续推进。
## 5. 核心声道的 QMF 域处理
L、R、C 在进入对象矩阵前延迟 10 个 QMF 时槽:
$$
\widehat X_{c,b,t}=X_{c,b,t-10},
\qquad c\in\{L,R,C\}.
$$
Ls、Rs 同样延迟 10 个时槽,并在 $b>0$ 时作 $-j$ 旋转:
$$
\widehat X_{c,b,t}=-jX_{c,b,t-10},
\qquad c\in\{Ls,Rs\},\ b>0.
$$
环绕声道的 band 0 还经过 21-tap 复 FIR:
$$
\widehat X_{c,0,t}
=
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
$$
这些延迟和滤波历史属于解码状态,不能按帧独立清零。
## 6. 对象矩阵
对每个对象 $o$、子带 $b$ 和时槽 $t$,对象频域值为五个核心声道的线性组合:
$$
Z_{o,b,t}
=
\sum_{c=0}^{4}
M_{o,c,b,t}\widehat X_{c,b,t}.
$$
analysis 输入的 $1/16$ 缩放会在 inverse QMF 输出端由 $\times16$ 抵消,因此矩阵本身不需要额外经验增益。
## 7. 对象 inverse QMF
### 7.1 子带重排
将 64 个复子带写成 128 个交织实数 `src`。对 $k=0\ldots31$:
$$
\begin{aligned}
\operatorname{zone}[2k] &= \operatorname{src}[4k],\\
\operatorname{zone}[2k+1] &= -\operatorname{src}[4k+1],\\
\operatorname{zone}[126-2k] &= \operatorname{src}[4k+2],\\
\operatorname{zone}[127-2k] &= \operatorname{src}[4k+3].
\end{aligned}
$$
把 `zone` 重新视为 64 个复数后执行未归一化 64 点 FFT:
$$
F_k
=
\sum_{n=0}^{63}
\operatorname{zone}_n
\exp\!\left(-j\frac{2\pi kn}{64}\right).
$$
### 7.2 调制与合成
定义旋转系数
$$
r_k
=
\frac12\left(
\sin\frac{\pi k}{128}
+j\cos\frac{\pi k}{128}
\right),
$$
并计算
$$
R_k=2F_kr_k.
$$
令 $\mathcal S$ 表示带 640 项 synthesis window 和跨时槽状态的 polyphase 合成算子:
$$
\mathbf y_{o,t}
=
\mathcal S\!\left(
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
\right).
$$
对象输出为
$$
y_o[64t+r]
=
\operatorname{clip}\!\left(
16\,\mathbf y_{o,t}[r],-1,1
\right)G_{\mathrm{clip}},
$$
其中 $r=0\ldots63$。synthesis 状态必须按时槽连续推进。
## 8. LFE 路径
LFE 不经过对象矩阵或 inverse QMF,而是使用 1217-sample 延迟。输入与输出端的比例因子抵消后:
$$
y_{\mathrm{LFE}}[n]
=
\operatorname{clip}\!\left(
x_{\mathrm{LFE,core}}[n-1217],-1,1
\right).
$$
## 9. OAMD 坐标
横向和纵向网格使用 $N=62$,高度网格使用 $N=15$。量化函数为
$$
q_N(k)
=
\min\!\left(
32767,
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
\right).
$$
OAR 坐标为
$$
u=\frac{q_1}{32768},
\qquad
v=\frac{q_2}{32768},
\qquad
w=\frac{q_3}{32768}.
$$
其最大运行值为 $32767/32768$,不是精确的 1。
转换为 ADM 网格时:
$$
k_1=\operatorname{round}\!\left(\frac{62q_1}{32767}\right),
\quad
k_2=\operatorname{round}\!\left(\frac{62q_2}{32767}\right),
\quad
k_3=\operatorname{round}\!\left(\frac{15q_3}{32767}\right),
$$
$$
X=2\frac{k_1}{62}-1,
\qquad
Y=1-2\frac{k_2}{62},
\qquad
Z=\frac{k_3}{15}.
$$
连续坐标关系为
$$
u=\frac{X+1}{2},
\qquad
v=\frac{1-Y}{2},
\qquad
w=Z.
$$
## 10. 等功率扬声器声像
### 10.1 一维插值
相邻扬声器坐标为 $a_0<a_1$,对象位置为 $a$。归一化位置为
$$
\tau=\frac{a-a_0}{a_1-a_0}.
$$
区间内的增益为
$$
g_0(\tau)=\cos\left(\frac\pi2\tau\right),
\qquad
g_1(\tau)=\sin\left(\frac\pi2\tau\right),
$$
并满足
$$
g_0^2(\tau)+g_1^2(\tau)=1.
$$
区间外的对象位置夹到最近端点。
### 10.2 二维 region
每一行先沿 $u$ 得到横向增益向量 $\mathbf h_r(u)$。若对象位于相邻两行 $r_0,r_1$ 之间:
$$
\eta=\frac{v-v_{r_0}}{v_{r_1}-v_{r_0}},
$$
$$
a_0=\cos\left(\frac\pi2\eta\right),
\qquad
a_1=\sin\left(\frac\pi2\eta\right).
$$
二维点增益为
$$
\mathbf G_{\mathrm{2D}}(u,v)
=
\mathbf h(u)\odot\mathbf v(v).
$$
对于只有一对水平环绕、没有独立 side/rear 两对的 5.1 系列布局,纵向坐标使用
$$
v_{\mathrm{floor}}
=
\operatorname{clamp}(2v,0,1).
$$
其他布局使用 $v_{\mathrm{floor}}=v$。
### 10.3 高度层
三维布局分别计算地面层增益 $\mathbf G_f$ 和高度层增益 $\mathbf G_h$:
$$
\mathbf G_{\mathrm{point}}(u,v,w)
=
\cos\left(\frac\pi2w\right)\mathbf G_f
+
\sin\left(\frac\pi2w\right)\mathbf G_h.
$$
当地面层与高度层扬声器集合不重叠、且各层内部使用等功率插值时:
$$
\left\|\mathbf G_{\mathrm{point}}\right\|_2=1.
$$
## 11. 布局位置补偿
令 $N_h$ 为相关高度扬声器数,$N_f$ 为相关附加水平扬声器数:
$$
H=\min\left(\frac{N_h}{4},1\right),
\qquad
F=\min\left(\frac{N_f}{4},1\right).
$$
最大位置补偿为
$$
A_{\max}
=
-\max\left(4.5-1.5H-3F,0\right)
\quad\text{dB}.
$$
前后与高度位置权重为
$$
p_v=\operatorname{clamp}\left(\frac v{0.6},0,1\right),
$$
$$
p_w=\operatorname{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
$$
$$
p=\operatorname{clamp}(p_v+p_w,0,1).
$$
线性补偿增益为
$$
G_{\mathrm{pos}}=10^{A_{\max}p/20}.
$$
对象的目标增益向量为
$$
\mathbf G_{\mathrm{target}}
=
G_{\mathrm{object}}
G_{\mathrm{pos}}
\mathbf G_{\mathrm{point}}.
$$
## 12. OAMD 时间对齐与增益斜坡
OAMD 更新的编码位置为
$$
s_{\mathrm{coded}}
=
s_{\mathrm{frame}}
+s_{\mathrm{outer}}
+s_{\mathrm{OAMD}}
+32f_{\mathrm{block}}.
$$
对处理块长度 $B=32$,更新点对齐为
$$
\widehat s
=
B\left\lfloor
\frac{s_{\mathrm{coded}}+B/2-1}{B}
\right\rfloor.
$$
给定 ramp duration $D$,block 数为
$$
K
=
\left\lfloor
\frac{D+B/2-1}{B}
\right\rfloor.
$$
若当前增益为 $g_0$、目标为 $g_1$,则每 block 的增量为
$$
\Delta g=\frac{g_1-g_0}{K}.
$$
第 $j$ 个 block 内的样本 $r=0\ldots B-1$ 使用
$$
g_{j,r}=g_j+\frac rB\Delta g,
\qquad
g_{j+1}=g_j+\Delta g.
$$
如果中途没有新的 metadata 更新,该过程等价于总长度 $KB$ 的逐样本线性斜坡。
## 13. 最终扬声器混音
对目标输出声道 $c$:
$$
y_c[n]
=
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
+
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
$$
其中
$$
\delta_{c,\mathrm{LFE}}
=
\begin{cases}
1, & c\text{ 为目标布局的 LFE},\\
0, & \text{其他声道}.
\end{cases}
$$
没有 LFE 输出的布局不把输入 LFE 混入其他声道。对象完成累加后,再按目标格式要求排列输出声道。
若输出 PCM24,量化关系为
$$
y_{24}[n]
=
\operatorname{trunc}\left(
8388607\,\operatorname{clip}(y[n],-1,1)
\right).
$$
## 14. 公式适用范围
- JOC 矩阵部分描述 dense JOC;Sparse JOC 使用不同的稀疏系数/索引路径。
- 扬声器声像部分描述普通点对象;extent、spread、divergence 等模式需要额外模型。
- 多个 OAMD position block 必须按其时间顺序调度。
- limiter 属于独立后处理,不包含在上述混音公式中。
+159
View File
@@ -0,0 +1,159 @@
# JustOneCacophony native-core notes
[中文](native.md) · [Back to README](../README.en.md)
## 1. Responsibility boundary
`native/` contains only the state-heavy, frequently called DSP and speaker-rendering kernels. High-level EMDF/JOC/OAMD parsing, error reporting, ADM assembly, and the CLI remain in Python.
Python calls a C ABI through the standard-library `ctypes` module. The native core does not use pybind11, Cython, FFTW, MKL, or OpenMP. It is an optional acceleration path and does not expand the set of supported stream variants.
Main files:
```text
native/include/eac3joc_core.h C ABI
native/src/eac3joc_core.cpp JOC/QMF object reconstruction
native/src/speaker_renderer.cpp object-to-speaker rendering
native/src/qmf_tables.h QMF tables
native/src/speaker_layouts.h layout tables
native/src/joc_huffman_tables.h JOC Huffman tables
src/native_renderer.py JOC ctypes bridge
src/speaker_native_renderer.py speaker ctypes bridge
```
## 2. JOC rendering ABI
An opaque renderer owns all cross-frame state. Its main call is:
```c
int ejoc_renderer_process(
ejoc_renderer_handle handle,
const float* bed5_planar, /* [5][1536] */
const float* lfe, /* [1536] or NULL */
uint32_t object_mask,
const uint8_t* n_bands, /* [15] */
const uint8_t* n_dpoints, /* [15] */
const uint8_t* slope_idx, /* [15] */
const uint8_t* offset_ts, /* [15][2] */
const double* dq, /* [15][2][5][23] */
double clipgain,
float phase_new,
float output_scale,
float* output16_planar); /* [16][1536] */
```
Python performs dense-JOC Huffman decoding, differential reconstruction, and dequantization before the call. Sparse JOC is not silently passed to the dense native path.
Thread control is exposed as:
```c
int ejoc_renderer_set_threads(ejoc_renderer_handle handle, uint32_t total_threads);
uint32_t ejoc_renderer_thread_count(ejoc_renderer_handle handle);
```
`total_threads` includes the calling thread. Frames must be submitted sequentially to one renderer instance; the instance may parallelize work across objects and analysis channels.
## 3. Cross-frame state
Each JOC renderer stores:
- analysis FIFO: `double[5][9][64]`;
- L/R/C analysis delay: `float[3][10][64]`;
- Ls/Rs QMF delay: `complex<double>[2][10][64]`;
- Ls/Rs band-0 FIR history: `complex<double>[2][20]`;
- previous matrix interpolation values: `double[15][5][64]`;
- inverse-QMF state: `double[15][640]`;
- LFE delay: `double[1217]`.
This state belongs to the renderer instance. Processing cannot be arbitrarily segmented or reordered without a corresponding state checkpoint.
## 4. FFT, QMF, and precision
The native core contains a fixed 64-point radix-2 complex FFT:
- analysis QMF uses a forward FFT followed by division by 64;
- inverse QMF uses the fixed reorder, rotation, and 640-value active-window state;
- no external FFT library is called.
The JOC path uses:
- float32 core-PCM input;
- double matrices, complex QMF, FFT, FIR, and cross-frame state;
- float32 phase and final gain;
- float32 16-channel object output.
## 5. Speaker-rendering ABI
The same shared library exports object-to-speaker rendering:
```c
uint32_t ejoc_speaker_layout_channel_count(uint32_t speaker_bitfield);
ejoc_speaker_renderer_handle
ejoc_speaker_renderer_create(uint32_t speaker_bitfield);
int ejoc_speaker_renderer_process(
ejoc_speaker_renderer_handle handle,
const float* objects16_interleaved,
uint32_t sample_count,
uint32_t metadata_count,
const uint32_t* metadata_offsets,
const uint32_t* ramp_durations,
const uint16_t* positions_q15,
const uint8_t* region_indices,
const uint8_t* height_enabled,
const double* object_gains,
double* output_interleaved);
```
Input channel 0 is LFE and channels 1–15 are objects. Each metadata entry is an object-state snapshot. `sample_count` must be a multiple of 32; unfinished gain ramps remain in the handle and continue across calls.
The speaker path uses float32 object input, double coordinates/gains/accumulation, and interleaved double output. Quantization to float32 or PCM24 happens when the WAV is written.
Supported layouts:
```text
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
```
## 6. Building
The CMake definition is `native/CMakeLists.txt`. Run from the repository root:
```powershell
cmake -S native -B build/cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX="$PWD/lib"
cmake --build build/cmake --config Release
cmake --install build/cmake --config Release
```
Platform runtime names:
```text
Windows lib/eac3joc_core.dll
Linux lib/libeac3joc_core.so
macOS lib/libeac3joc_core.dylib
```
The MSVC configuration uses the static CRT. Other runtime dependencies depend on the platform and toolchain and should be checked independently before publishing a prebuilt library.
The repository does not include native binaries by default. A prebuilt Release runtime or a locally built runtime can be placed directly under `lib/`.
## 7. Runtime lookup and fallback
Lookup order:
1. explicit `--native-library`;
2. `EAC3JOC_NATIVE_LIBRARY`;
3. the standard platform filename under `lib/`.
`--backend auto` falls back to NumPy when loading fails, and `--backend python` skips native discovery. The current CLI also prints the failure and falls back for `--backend native`; this existing behavior should not be read as successful native execution.
## 8. Implementation boundaries
- The native layer accepts only dense-JOC data already parsed by Python.
- The ABI fixes a 1536-sample JOC frame, at most 15 objects, at most 23 parameter bands, and at most 2 data points.
- The shared library and Python bridge must report the same ABI version.
- Only the ABI and data types are specified across platforms; bit-identical float64 results are not guaranteed.
- Private table headers under `native/src/` serve the native side only. The current repository does not include the scripts that generated those headers.
See the [mathematical notes](math.en.md) for the related formulas.
+159
View File
@@ -0,0 +1,159 @@
# JustOneCacophony 原生核说明
[English](native.en.md) · [返回 README](../README.md)
## 1. 职责边界
`native/` 只承载状态密集、调用频繁的 DSP 与扬声器渲染核。EMDF/JOC/OAMD 高层解析、错误报告、ADM 组装和 CLI 保留在 Python 中。
Python 通过标准库 `ctypes` 调用 C ABI;原生核不使用 pybind11、Cython、FFTW、MKL 或 OpenMP。它是可选加速路径,不扩大项目所支持的码流范围。
主要文件:
```text
native/include/eac3joc_core.h C ABI
native/src/eac3joc_core.cpp JOC/QMF 对象重建
native/src/speaker_renderer.cpp 对象到扬声器渲染
native/src/qmf_tables.h QMF 表
native/src/speaker_layouts.h 布局表
native/src/joc_huffman_tables.h JOC Huffman 表
src/native_renderer.py JOC ctypes 桥
src/speaker_native_renderer.py 扬声器 ctypes 桥
```
## 2. JOC 渲染 ABI
一个 opaque renderer 保存所有跨帧状态。主要调用为:
```c
int ejoc_renderer_process(
ejoc_renderer_handle handle,
const float* bed5_planar, /* [5][1536] */
const float* lfe, /* [1536] or NULL */
uint32_t object_mask,
const uint8_t* n_bands, /* [15] */
const uint8_t* n_dpoints, /* [15] */
const uint8_t* slope_idx, /* [15] */
const uint8_t* offset_ts, /* [15][2] */
const double* dq, /* [15][2][5][23] */
double clipgain,
float phase_new,
float output_scale,
float* output16_planar); /* [16][1536] */
```
Dense JOC 的 Huffman 解码、差分还原和去量化先在 Python 中完成。Sparse JOC 不会被静默送入 dense 原生路径。
线程接口为:
```c
int ejoc_renderer_set_threads(ejoc_renderer_handle handle, uint32_t total_threads);
uint32_t ejoc_renderer_thread_count(ejoc_renderer_handle handle);
```
`total_threads` 包含调用线程。单个 renderer 实例必须顺序提交帧;实例内部可以按对象和 analysis channel 并行。
## 3. 跨帧状态
每个 JOC renderer 独立保存:
- analysis FIFO:`double[5][9][64]`;
- L/R/C analysis delay:`float[3][10][64]`;
- Ls/Rs QMF delay:`complex<double>[2][10][64]`;
- Ls/Rs band-0 FIR history:`complex<double>[2][20]`;
- 矩阵插值 previous:`double[15][5][64]`;
- inverse-QMF state:`double[15][640]`;
- LFE delay:`double[1217]`。
这些状态属于 renderer 实例,不能在无 checkpoint 的情况下任意分段或乱序处理。
## 4. FFT、QMF 与精度
原生核包含固定 64 点 radix-2 complex FFT:
- analysis QMF 使用 forward FFT 后除以 64;
- inverse QMF 使用固定重排、旋转和 640 项有效窗状态;
- 不调用外部 FFT 库。
JOC 路径的数值类型为:
- 核心 PCM 输入:float32;
- 矩阵、复 QMF、FFT、FIR 和跨帧状态:double;
- phase 与最终 gain:float32;
- 16 声道对象输出:float32。
## 5. 扬声器渲染 ABI
同一个共享库还导出对象到扬声器布局的渲染接口:
```c
uint32_t ejoc_speaker_layout_channel_count(uint32_t speaker_bitfield);
ejoc_speaker_renderer_handle
ejoc_speaker_renderer_create(uint32_t speaker_bitfield);
int ejoc_speaker_renderer_process(
ejoc_speaker_renderer_handle handle,
const float* objects16_interleaved,
uint32_t sample_count,
uint32_t metadata_count,
const uint32_t* metadata_offsets,
const uint32_t* ramp_durations,
const uint16_t* positions_q15,
const uint8_t* region_indices,
const uint8_t* height_enabled,
const double* object_gains,
double* output_interleaved);
```
输入声道 0 为 LFE,1–15 为对象。每个 metadata entry 是一份对象状态快照。`sample_count` 必须是 32 的倍数;未完成的增益斜坡保存在 handle 中并跨调用继续。
扬声器路径使用 float32 对象输入、double 坐标/增益/累加与 interleaved double 输出;写 WAV 时才量化为 float32 或 PCM24。
支持的布局为:
```text
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
```
## 6. 构建
CMake 定义位于 `native/CMakeLists.txt`。从仓库根目录运行:
```powershell
cmake -S native -B build/cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX="$PWD/lib"
cmake --build build/cmake --config Release
cmake --install build/cmake --config Release
```
平台运行库文件名:
```text
Windows lib/eac3joc_core.dll
Linux lib/libeac3joc_core.so
macOS lib/libeac3joc_core.dylib
```
MSVC 配置使用静态 CRT。其他运行时依赖由平台和工具链决定,发布预构建库前应对产物独立检查。
仓库默认不附带原生二进制。预构建的 Release 运行库或自行构建的运行库均可直接放入 `lib/`。
## 7. 运行时查找与回退
查找顺序为:
1. 显式 `--native-library`;
2. `EAC3JOC_NATIVE_LIBRARY`;
3. `lib/` 下当前平台的标准文件名。
`--backend auto` 在加载失败时回退到 NumPy;`--backend python` 跳过原生探测。`--backend native` 当前也会打印失败原因后回退,这是现有 CLI 行为,不应理解为原生库已成功使用。
## 8. 实现边界
- 原生层只接收 Python 已解析的 dense JOC 数据。
- ABI 固定了 1536-sample JOC 帧、最多 15 个对象、最多 23 个参数带和最多 2 个数据点。
- 共享库与 Python 桥需要 ABI version 一致。
- 跨平台只约定 ABI 与数据类型,不保证 float64 结果逐位一致。
- `native/src/` 中的私有表头只服务于原生侧;当前仓库不包含重新生成这些头文件的脚本。
相关公式见[数学说明](math.md)。