Initial public release of JustOneCacophony
This commit is contained in:
+578
@@ -0,0 +1,578 @@
|
||||
# JustOneCacophony — E-AC-3 JOC decoding and rendering mathematics
|
||||
|
||||
[中文](math.md) · [Back to README](../README.en.md)
|
||||
|
||||
This document covers only the signal model and formulas used in the JustOneCacophony research path: how JOC parameters combine with core PCM to reconstruct object signals, and how OAMD coordinates become speaker gains.
|
||||
|
||||
The formulas describe the dense-JOC and ordinary point-object paths studied by the project. They are not a complete definition of every E-AC-3 JOC variant.
|
||||
|
||||
## 1. Overall path and notation
|
||||
|
||||
Object reconstruction:
|
||||
|
||||
```text
|
||||
E-AC-3 core 5.1 PCM
|
||||
+ ID14 JOC matrix parameters
|
||||
→ analysis QMF
|
||||
→ parameter-band expansion and time interpolation
|
||||
→ object matrix
|
||||
→ inverse QMF
|
||||
→ LFE + 15 object PCM channels
|
||||
```
|
||||
|
||||
Speaker rendering:
|
||||
|
||||
```text
|
||||
LFE + 15 object PCM channels
|
||||
+ ID11 OAMD coordinates and update timing
|
||||
→ target-layout region
|
||||
→ equal-power panning
|
||||
→ position compensation
|
||||
→ sample-wise gain ramp
|
||||
→ speaker PCM
|
||||
```
|
||||
|
||||
Main notation:
|
||||
|
||||
| Symbol | Meaning |
|
||||
|---|---|
|
||||
| $c=0\ldots4$ | core channels L, R, C, Ls, Rs |
|
||||
| $o=0\ldots14$ | 15 JOC objects |
|
||||
| $b=0\ldots63$ | complex QMF subbands |
|
||||
| $t=0\ldots23$ | 24 64-sample slots per frame |
|
||||
| $p(b)$ | JOC parameter band corresponding to QMF subband $b$ |
|
||||
| $X_{c,b,t}$ | analysis-QMF value for a core channel |
|
||||
| $M_{o,c,b,t}$ | object-matrix coefficient |
|
||||
| $Z_{o,b,t}$ | inverse-QMF input for an object |
|
||||
| $y_o[n]$ | time-domain object PCM |
|
||||
|
||||
The number of samples in one frame is
|
||||
|
||||
$$
|
||||
N_f=1536=24\times64.
|
||||
$$
|
||||
|
||||
## 2. Dense-JOC matrix parameters
|
||||
|
||||
### 2.1 Differential reconstruction
|
||||
|
||||
Let `quant_idx` be $q_i\in\{0,1\}$. The number of quantization levels is
|
||||
|
||||
$$
|
||||
N_q=
|
||||
\begin{cases}
|
||||
96, & q_i=0,\\
|
||||
192, & q_i=1.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
The center offset is
|
||||
|
||||
$$
|
||||
O_q=\frac{N_q}{2}.
|
||||
$$
|
||||
|
||||
For object $o$, data point $d$, core channel $c$, and parameter band $p$, the coded difference $\Delta_{o,d,c,p}$ reconstructs to
|
||||
|
||||
$$
|
||||
Q_{o,d,c,0}
|
||||
=
|
||||
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}
|
||||
=
|
||||
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
|
||||
\qquad p>0.
|
||||
$$
|
||||
|
||||
### 2.2 Dequantization
|
||||
|
||||
The dequantized matrix coefficient is
|
||||
|
||||
$$
|
||||
D_{o,d,c,p}
|
||||
=
|
||||
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
|
||||
\frac{820}{4096(1+q_i)}.
|
||||
$$
|
||||
|
||||
The effective denominator is therefore 4096 in coarse mode and 8192 in fine mode.
|
||||
|
||||
### 2.3 JOC clipgain
|
||||
|
||||
If the clipgain field consists of integer $x$ and mantissa $y$, then
|
||||
|
||||
$$
|
||||
G_{\mathrm{clip}}
|
||||
=
|
||||
1+\frac{y}{32}2^{x-4}.
|
||||
$$
|
||||
|
||||
It is applied to object PCM after inverse QMF and does not apply to LFE.
|
||||
|
||||
## 3. Parameter-band expansion and time interpolation
|
||||
|
||||
### 3.1 Parameter bands to QMF subbands
|
||||
|
||||
The JOC matrix is coded in parameter bands, while the QMF contains 64 subbands. Let $p(b)$ identify the parameter band containing subband $b$. A parameter-band coefficient expands as
|
||||
|
||||
$$
|
||||
D_{o,d,c,b}=D_{o,d,c,p(b)}.
|
||||
$$
|
||||
|
||||
The common 12-band mapping is
|
||||
|
||||
$$
|
||||
\begin{aligned}
|
||||
\mathcal B_0 &= \{0\}, &
|
||||
\mathcal B_1 &= \{1\}, &
|
||||
\mathcal B_2 &= \{2\}, &
|
||||
\mathcal B_3 &= \{3\},\\
|
||||
\mathcal B_4 &= \{4,5\}, &
|
||||
\mathcal B_5 &= \{6,7\}, &
|
||||
\mathcal B_6 &= \{8,9,10\}, &
|
||||
\mathcal B_7 &= \{11,12,13\},\\
|
||||
\mathcal B_8 &= \{14,15,16,17\}, &
|
||||
\mathcal B_9 &= \{18,\ldots,22\},\\
|
||||
\mathcal B_{10} &= \{23,\ldots,34\}, &
|
||||
\mathcal B_{11} &= \{35,\ldots,63\}.
|
||||
\end{aligned}
|
||||
$$
|
||||
|
||||
Thus $p(b)=k$ if and only if $b\in\mathcal B_k$. Other parameter-band counts use their corresponding subband boundaries.
|
||||
|
||||
### 3.2 One-data-point interpolation
|
||||
|
||||
Let $P_{o,c,b}$ be the previous frame-end value and $D_{o,c,p(b)}$ the current target. For slot $t=0\ldots23$:
|
||||
|
||||
$$
|
||||
\alpha_t=\frac{t+1}{24},
|
||||
$$
|
||||
|
||||
$$
|
||||
M_{o,c,b,t}
|
||||
=
|
||||
(1-\alpha_t)P_{o,c,b}
|
||||
+\alpha_tD_{o,c,p(b)}.
|
||||
$$
|
||||
|
||||
The first slot has therefore advanced by $1/24$ of the ramp, while the last slot equals the current target:
|
||||
|
||||
$$
|
||||
M_{o,c,b,23}=D_{o,c,p(b)}.
|
||||
$$
|
||||
|
||||
This value then becomes the previous state for the next frame.
|
||||
|
||||
### 3.3 Multiple data points
|
||||
|
||||
When a frame contains two data points, `offset_ts` gives the segment boundary. Each segment uses the same linear relation between the previous and next targets; step mode switches targets at the designated slot.
|
||||
|
||||
## 4. Analysis QMF for core PCM
|
||||
|
||||
The matrix input uses core channels L, R, C, Ls, and Rs; LFE follows a separate path. Core PCM is first scaled as
|
||||
|
||||
$$
|
||||
\widetilde x_c[n]=\frac{x_c[n]}{16}.
|
||||
$$
|
||||
|
||||
Let $\mathcal A_b$ denote the 64-band analysis-QMF operator with polyphase history state. Then
|
||||
|
||||
$$
|
||||
X_{c,b,t}
|
||||
=
|
||||
\mathcal A_b\!\left(
|
||||
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
|
||||
\mathbf s^{\mathrm A}_{c,t}
|
||||
\right).
|
||||
$$
|
||||
|
||||
This consists of the analysis window/polyphase stage, modulation, a 64-point FFT, and subband reordering. History state advances continuously across slots and frames.
|
||||
|
||||
## 5. QMF-domain processing of core channels
|
||||
|
||||
L, R, and C are delayed by ten QMF slots before entering the object matrix:
|
||||
|
||||
$$
|
||||
\widehat X_{c,b,t}=X_{c,b,t-10},
|
||||
\qquad c\in\{L,R,C\}.
|
||||
$$
|
||||
|
||||
Ls and Rs use the same ten-slot delay and a $-j$ rotation for $b>0$:
|
||||
|
||||
$$
|
||||
\widehat X_{c,b,t}=-jX_{c,b,t-10},
|
||||
\qquad c\in\{Ls,Rs\},\ b>0.
|
||||
$$
|
||||
|
||||
Band 0 of each surround channel additionally passes through a 21-tap complex FIR:
|
||||
|
||||
$$
|
||||
\widehat X_{c,0,t}
|
||||
=
|
||||
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
|
||||
$$
|
||||
|
||||
These delays and filter histories are decoder state and cannot be reset independently for every frame.
|
||||
|
||||
## 6. Object matrix
|
||||
|
||||
For each object $o$, subband $b$, and slot $t$, the object's frequency-domain value is a linear combination of the five core channels:
|
||||
|
||||
$$
|
||||
Z_{o,b,t}
|
||||
=
|
||||
\sum_{c=0}^{4}
|
||||
M_{o,c,b,t}\widehat X_{c,b,t}.
|
||||
$$
|
||||
|
||||
The $1/16$ analysis-input scale is canceled by the $\times16$ factor after inverse QMF, so the matrix itself needs no additional empirical gain.
|
||||
|
||||
## 7. Object inverse QMF
|
||||
|
||||
### 7.1 Subband reorder
|
||||
|
||||
Write the 64 complex subbands as 128 interleaved real values in `src`. For $k=0\ldots31$:
|
||||
|
||||
$$
|
||||
\begin{aligned}
|
||||
\operatorname{zone}[2k] &= \operatorname{src}[4k],\\
|
||||
\operatorname{zone}[2k+1] &= -\operatorname{src}[4k+1],\\
|
||||
\operatorname{zone}[126-2k] &= \operatorname{src}[4k+2],\\
|
||||
\operatorname{zone}[127-2k] &= \operatorname{src}[4k+3].
|
||||
\end{aligned}
|
||||
$$
|
||||
|
||||
Treat `zone` as 64 complex values and apply an unnormalized 64-point FFT:
|
||||
|
||||
$$
|
||||
F_k
|
||||
=
|
||||
\sum_{n=0}^{63}
|
||||
\operatorname{zone}_n
|
||||
\exp\!\left(-j\frac{2\pi kn}{64}\right).
|
||||
$$
|
||||
|
||||
### 7.2 Modulation and synthesis
|
||||
|
||||
Define the rotation coefficient
|
||||
|
||||
$$
|
||||
r_k
|
||||
=
|
||||
\frac12\left(
|
||||
\sin\frac{\pi k}{128}
|
||||
+j\cos\frac{\pi k}{128}
|
||||
\right),
|
||||
$$
|
||||
|
||||
and compute
|
||||
|
||||
$$
|
||||
R_k=2F_kr_k.
|
||||
$$
|
||||
|
||||
Let $\mathcal S$ denote polyphase synthesis with a 640-value synthesis window and cross-slot state:
|
||||
|
||||
$$
|
||||
\mathbf y_{o,t}
|
||||
=
|
||||
\mathcal S\!\left(
|
||||
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
|
||||
\right).
|
||||
$$
|
||||
|
||||
Object output is
|
||||
|
||||
$$
|
||||
y_o[64t+r]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
16\,\mathbf y_{o,t}[r],-1,1
|
||||
\right)G_{\mathrm{clip}},
|
||||
$$
|
||||
|
||||
where $r=0\ldots63$. Synthesis state must advance continuously by slot.
|
||||
|
||||
## 8. LFE path
|
||||
|
||||
LFE bypasses the object matrix and inverse QMF and uses a 1217-sample delay. After the input and output scale factors cancel:
|
||||
|
||||
$$
|
||||
y_{\mathrm{LFE}}[n]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
x_{\mathrm{LFE,core}}[n-1217],-1,1
|
||||
\right).
|
||||
$$
|
||||
|
||||
## 9. OAMD coordinates
|
||||
|
||||
The lateral and longitudinal grids use $N=62$; the height grid uses $N=15$. The quantizer is
|
||||
|
||||
$$
|
||||
q_N(k)
|
||||
=
|
||||
\min\!\left(
|
||||
32767,
|
||||
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
|
||||
\right).
|
||||
$$
|
||||
|
||||
OAR coordinates are
|
||||
|
||||
$$
|
||||
u=\frac{q_1}{32768},
|
||||
\qquad
|
||||
v=\frac{q_2}{32768},
|
||||
\qquad
|
||||
w=\frac{q_3}{32768}.
|
||||
$$
|
||||
|
||||
Their maximum runtime value is $32767/32768$, not exactly 1.
|
||||
|
||||
For conversion to the ADM grid:
|
||||
|
||||
$$
|
||||
k_1=\operatorname{round}\!\left(\frac{62q_1}{32767}\right),
|
||||
\quad
|
||||
k_2=\operatorname{round}\!\left(\frac{62q_2}{32767}\right),
|
||||
\quad
|
||||
k_3=\operatorname{round}\!\left(\frac{15q_3}{32767}\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
X=2\frac{k_1}{62}-1,
|
||||
\qquad
|
||||
Y=1-2\frac{k_2}{62},
|
||||
\qquad
|
||||
Z=\frac{k_3}{15}.
|
||||
$$
|
||||
|
||||
The continuous-coordinate relation is
|
||||
|
||||
$$
|
||||
u=\frac{X+1}{2},
|
||||
\qquad
|
||||
v=\frac{1-Y}{2},
|
||||
\qquad
|
||||
w=Z.
|
||||
$$
|
||||
|
||||
## 10. Equal-power speaker panning
|
||||
|
||||
### 10.1 One-dimensional interpolation
|
||||
|
||||
Let adjacent speaker coordinates be $a_0<a_1$ and object position be $a$. The normalized position is
|
||||
|
||||
$$
|
||||
\tau=\frac{a-a_0}{a_1-a_0}.
|
||||
$$
|
||||
|
||||
Gains inside the interval are
|
||||
|
||||
$$
|
||||
g_0(\tau)=\cos\left(\frac\pi2\tau\right),
|
||||
\qquad
|
||||
g_1(\tau)=\sin\left(\frac\pi2\tau\right),
|
||||
$$
|
||||
|
||||
and satisfy
|
||||
|
||||
$$
|
||||
g_0^2(\tau)+g_1^2(\tau)=1.
|
||||
$$
|
||||
|
||||
Object positions outside the interval are clamped to the nearest endpoint.
|
||||
|
||||
### 10.2 Two-dimensional regions
|
||||
|
||||
Each row first produces a horizontal gain vector $\mathbf h_r(u)$. If the object lies between adjacent rows $r_0,r_1$:
|
||||
|
||||
$$
|
||||
\eta=\frac{v-v_{r_0}}{v_{r_1}-v_{r_0}},
|
||||
$$
|
||||
|
||||
$$
|
||||
a_0=\cos\left(\frac\pi2\eta\right),
|
||||
\qquad
|
||||
a_1=\sin\left(\frac\pi2\eta\right).
|
||||
$$
|
||||
|
||||
The two-dimensional point gain is
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{2D}}(u,v)
|
||||
=
|
||||
\mathbf h(u)\odot\mathbf v(v).
|
||||
$$
|
||||
|
||||
For 5.1-family layouts with one horizontal surround pair rather than separate side and rear pairs, the longitudinal coordinate is
|
||||
|
||||
$$
|
||||
v_{\mathrm{floor}}
|
||||
=
|
||||
\operatorname{clamp}(2v,0,1).
|
||||
$$
|
||||
|
||||
Other layouts use $v_{\mathrm{floor}}=v$.
|
||||
|
||||
### 10.3 Height layer
|
||||
|
||||
Three-dimensional layouts compute floor gain $\mathbf G_f$ and height gain $\mathbf G_h$ separately:
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)
|
||||
=
|
||||
\cos\left(\frac\pi2w\right)\mathbf G_f
|
||||
+
|
||||
\sin\left(\frac\pi2w\right)\mathbf G_h.
|
||||
$$
|
||||
|
||||
When the floor and height speaker sets do not overlap and each layer uses equal-power interpolation:
|
||||
|
||||
$$
|
||||
\left\|\mathbf G_{\mathrm{point}}\right\|_2=1.
|
||||
$$
|
||||
|
||||
## 11. Layout-dependent position compensation
|
||||
|
||||
Let $N_h$ be the number of relevant height speakers and $N_f$ the number of relevant additional horizontal speakers:
|
||||
|
||||
$$
|
||||
H=\min\left(\frac{N_h}{4},1\right),
|
||||
\qquad
|
||||
F=\min\left(\frac{N_f}{4},1\right).
|
||||
$$
|
||||
|
||||
Maximum position compensation is
|
||||
|
||||
$$
|
||||
A_{\max}
|
||||
=
|
||||
-\max\left(4.5-1.5H-3F,0\right)
|
||||
\quad\text{dB}.
|
||||
$$
|
||||
|
||||
Longitudinal and height weights are
|
||||
|
||||
$$
|
||||
p_v=\operatorname{clamp}\left(\frac v{0.6},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p_w=\operatorname{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p=\operatorname{clamp}(p_v+p_w,0,1).
|
||||
$$
|
||||
|
||||
The linear compensation gain is
|
||||
|
||||
$$
|
||||
G_{\mathrm{pos}}=10^{A_{\max}p/20}.
|
||||
$$
|
||||
|
||||
The object's target-gain vector is
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{target}}
|
||||
=
|
||||
G_{\mathrm{object}}
|
||||
G_{\mathrm{pos}}
|
||||
\mathbf G_{\mathrm{point}}.
|
||||
$$
|
||||
|
||||
## 12. OAMD time alignment and gain ramps
|
||||
|
||||
The coded position of an OAMD update is
|
||||
|
||||
$$
|
||||
s_{\mathrm{coded}}
|
||||
=
|
||||
s_{\mathrm{frame}}
|
||||
+s_{\mathrm{outer}}
|
||||
+s_{\mathrm{OAMD}}
|
||||
+32f_{\mathrm{block}}.
|
||||
$$
|
||||
|
||||
For processing-block length $B=32$, the aligned update point is
|
||||
|
||||
$$
|
||||
\widehat s
|
||||
=
|
||||
B\left\lfloor
|
||||
\frac{s_{\mathrm{coded}}+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
$$
|
||||
|
||||
For ramp duration $D$, the number of blocks is
|
||||
|
||||
$$
|
||||
K
|
||||
=
|
||||
\left\lfloor
|
||||
\frac{D+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
$$
|
||||
|
||||
If current gain is $g_0$ and target gain is $g_1$, the increment per block is
|
||||
|
||||
$$
|
||||
\Delta g=\frac{g_1-g_0}{K}.
|
||||
$$
|
||||
|
||||
Sample $r=0\ldots B-1$ of block $j$ uses
|
||||
|
||||
$$
|
||||
g_{j,r}=g_j+\frac rB\Delta g,
|
||||
\qquad
|
||||
g_{j+1}=g_j+\Delta g.
|
||||
$$
|
||||
|
||||
If no new metadata update intervenes, this is equivalent to a sample-wise linear ramp of total length $KB$.
|
||||
|
||||
## 13. Final speaker mix
|
||||
|
||||
For target output channel $c$:
|
||||
|
||||
$$
|
||||
y_c[n]
|
||||
=
|
||||
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
|
||||
+
|
||||
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
|
||||
$$
|
||||
|
||||
Here
|
||||
|
||||
$$
|
||||
\delta_{c,\mathrm{LFE}}
|
||||
=
|
||||
\begin{cases}
|
||||
1, & c\text{ is the target layout's LFE channel},\\
|
||||
0, & \text{otherwise}.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
A layout without LFE output does not mix input LFE into other channels. After object accumulation, output channels are ordered as required by the target format.
|
||||
|
||||
For PCM24 output, quantization is
|
||||
|
||||
$$
|
||||
y_{24}[n]
|
||||
=
|
||||
\operatorname{trunc}\left(
|
||||
8388607\,\operatorname{clip}(y[n],-1,1)
|
||||
\right).
|
||||
$$
|
||||
|
||||
## 14. Scope of the formulas
|
||||
|
||||
- The JOC matrix section describes dense JOC; Sparse JOC uses a different sparse coefficient/index path.
|
||||
- The speaker-panning section describes ordinary point objects; extent, spread, divergence, and similar modes require additional models.
|
||||
- Multiple OAMD position blocks must be scheduled in time order.
|
||||
- A limiter is separate post-processing and is not included in the mixing equations above.
|
||||
+578
@@ -0,0 +1,578 @@
|
||||
# JustOneCacophony — E-AC-3 JOC 解码与渲染数学
|
||||
|
||||
[English](math.en.md) · [返回 README](../README.md)
|
||||
|
||||
本文只说明 JustOneCacophony 研究路径中使用的信号模型和公式:JOC 参数如何与核心 PCM 结合并重建对象信号,以及 OAMD 坐标如何转换为扬声器增益。
|
||||
|
||||
这些公式描述项目当前研究的 dense JOC 与普通点对象路径,不代表对所有 E-AC-3 JOC 变体的完整定义。
|
||||
|
||||
## 1. 总体路径与记号
|
||||
|
||||
对象重建路径:
|
||||
|
||||
```text
|
||||
E-AC-3 核心 5.1 PCM
|
||||
+ ID14 JOC 矩阵参数
|
||||
→ analysis QMF
|
||||
→ 参数带展开与时间插值
|
||||
→ 对象矩阵
|
||||
→ inverse QMF
|
||||
→ LFE + 15 路对象 PCM
|
||||
```
|
||||
|
||||
扬声器渲染路径:
|
||||
|
||||
```text
|
||||
LFE + 15 路对象 PCM
|
||||
+ ID11 OAMD 坐标与更新时间
|
||||
→ 目标布局 region
|
||||
→ 等功率声像
|
||||
→ 位置补偿
|
||||
→ 逐样本增益斜坡
|
||||
→ 扬声器 PCM
|
||||
```
|
||||
|
||||
主要记号:
|
||||
|
||||
| 符号 | 含义 |
|
||||
|---|---|
|
||||
| $c=0\ldots4$ | 核心声道 L、R、C、Ls、Rs |
|
||||
| $o=0\ldots14$ | 15 个 JOC 对象 |
|
||||
| $b=0\ldots63$ | 复 QMF 子带 |
|
||||
| $t=0\ldots23$ | 每帧 24 个 64-sample 时槽 |
|
||||
| $p(b)$ | QMF 子带 $b$ 对应的 JOC 参数带 |
|
||||
| $X_{c,b,t}$ | 核心声道的 analysis-QMF 值 |
|
||||
| $M_{o,c,b,t}$ | 对象矩阵系数 |
|
||||
| $Z_{o,b,t}$ | 对象的 inverse-QMF 输入 |
|
||||
| $y_o[n]$ | 对象时域 PCM |
|
||||
|
||||
一帧的采样数为
|
||||
|
||||
$$
|
||||
N_f=1536=24\times64.
|
||||
$$
|
||||
|
||||
## 2. Dense JOC 矩阵参数
|
||||
|
||||
### 2.1 差分还原
|
||||
|
||||
令 `quant_idx` 为 $q_i\in\{0,1\}$,量化级数为
|
||||
|
||||
$$
|
||||
N_q=
|
||||
\begin{cases}
|
||||
96, & q_i=0,\\
|
||||
192, & q_i=1.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
中心偏移为
|
||||
|
||||
$$
|
||||
O_q=\frac{N_q}{2}.
|
||||
$$
|
||||
|
||||
对对象 $o$、数据点 $d$、核心声道 $c$ 和参数带 $p$,编码差分 $\Delta_{o,d,c,p}$ 还原为
|
||||
|
||||
$$
|
||||
Q_{o,d,c,0}
|
||||
=
|
||||
\left(O_q+\Delta_{o,d,c,0}\right)\bmod N_q,
|
||||
$$
|
||||
|
||||
$$
|
||||
Q_{o,d,c,p}
|
||||
=
|
||||
\left(Q_{o,d,c,p-1}+\Delta_{o,d,c,p}\right)\bmod N_q,
|
||||
\qquad p>0.
|
||||
$$
|
||||
|
||||
### 2.2 去量化
|
||||
|
||||
矩阵系数的去量化值为
|
||||
|
||||
$$
|
||||
D_{o,d,c,p}
|
||||
=
|
||||
\left(Q_{o,d,c,p}-\frac{N_q}{2}\right)
|
||||
\frac{820}{4096(1+q_i)}.
|
||||
$$
|
||||
|
||||
因此 coarse 模式的有效分母为 4096,fine 模式为 8192。
|
||||
|
||||
### 2.3 JOC clipgain
|
||||
|
||||
若 clipgain 字段由整数 $x$ 和尾数 $y$ 组成,则
|
||||
|
||||
$$
|
||||
G_{\mathrm{clip}}
|
||||
=
|
||||
1+\frac{y}{32}2^{x-4}.
|
||||
$$
|
||||
|
||||
它在对象 inverse QMF 之后作用于对象 PCM,不作用于 LFE。
|
||||
|
||||
## 3. 参数带展开与时间插值
|
||||
|
||||
### 3.1 参数带到 QMF 子带
|
||||
|
||||
JOC 矩阵按参数带编码,而 QMF 使用 64 个子带。令 $p(b)$ 表示子带 $b$ 所属的参数带,则每个参数带系数展开为
|
||||
|
||||
$$
|
||||
D_{o,d,c,b}=D_{o,d,c,p(b)}.
|
||||
$$
|
||||
|
||||
常见的 12-band 映射为
|
||||
|
||||
$$
|
||||
\begin{aligned}
|
||||
\mathcal B_0 &= \{0\}, &
|
||||
\mathcal B_1 &= \{1\}, &
|
||||
\mathcal B_2 &= \{2\}, &
|
||||
\mathcal B_3 &= \{3\},\\
|
||||
\mathcal B_4 &= \{4,5\}, &
|
||||
\mathcal B_5 &= \{6,7\}, &
|
||||
\mathcal B_6 &= \{8,9,10\}, &
|
||||
\mathcal B_7 &= \{11,12,13\},\\
|
||||
\mathcal B_8 &= \{14,15,16,17\}, &
|
||||
\mathcal B_9 &= \{18,\ldots,22\},\\
|
||||
\mathcal B_{10} &= \{23,\ldots,34\}, &
|
||||
\mathcal B_{11} &= \{35,\ldots,63\}.
|
||||
\end{aligned}
|
||||
$$
|
||||
|
||||
其中 $p(b)=k$ 当且仅当 $b\in\mathcal B_k$。其他参数带数使用各自的子带边界。
|
||||
|
||||
### 3.2 单数据点插值
|
||||
|
||||
令上一帧末值为 $P_{o,c,b}$,当前目标值为 $D_{o,c,p(b)}$。对时槽 $t=0\ldots23$:
|
||||
|
||||
$$
|
||||
\alpha_t=\frac{t+1}{24},
|
||||
$$
|
||||
|
||||
$$
|
||||
M_{o,c,b,t}
|
||||
=
|
||||
(1-\alpha_t)P_{o,c,b}
|
||||
+\alpha_tD_{o,c,p(b)}.
|
||||
$$
|
||||
|
||||
因此第一时槽已经推进 ramp 的 $1/24$,最后一时槽等于当前目标:
|
||||
|
||||
$$
|
||||
M_{o,c,b,23}=D_{o,c,p(b)}.
|
||||
$$
|
||||
|
||||
该值随后成为下一帧的 previous 状态。
|
||||
|
||||
### 3.3 多数据点
|
||||
|
||||
当一帧含两个数据点时,`offset_ts` 给出分段边界。每一段在上一目标和下一目标之间使用相同的线性关系;阶跃模式则在指定时槽直接切换目标。
|
||||
|
||||
## 4. 核心 PCM 的 analysis QMF
|
||||
|
||||
矩阵输入使用核心声道 L、R、C、Ls、Rs;LFE 走独立路径。核心 PCM 先缩放为
|
||||
|
||||
$$
|
||||
\widetilde x_c[n]=\frac{x_c[n]}{16}.
|
||||
$$
|
||||
|
||||
令 $\mathcal A_b$ 表示带 polyphase 历史状态的 64-band analysis-QMF 算子,则
|
||||
|
||||
$$
|
||||
X_{c,b,t}
|
||||
=
|
||||
\mathcal A_b\!\left(
|
||||
\widetilde x_c[64t],\ldots,\widetilde x_c[64t+63];
|
||||
\mathbf s^{\mathrm A}_{c,t}
|
||||
\right).
|
||||
$$
|
||||
|
||||
该过程依次包含 analysis window/polyphase、调制、64 点 FFT 和子带重排。历史状态跨时槽和帧连续推进。
|
||||
|
||||
## 5. 核心声道的 QMF 域处理
|
||||
|
||||
L、R、C 在进入对象矩阵前延迟 10 个 QMF 时槽:
|
||||
|
||||
$$
|
||||
\widehat X_{c,b,t}=X_{c,b,t-10},
|
||||
\qquad c\in\{L,R,C\}.
|
||||
$$
|
||||
|
||||
Ls、Rs 同样延迟 10 个时槽,并在 $b>0$ 时作 $-j$ 旋转:
|
||||
|
||||
$$
|
||||
\widehat X_{c,b,t}=-jX_{c,b,t-10},
|
||||
\qquad c\in\{Ls,Rs\},\ b>0.
|
||||
$$
|
||||
|
||||
环绕声道的 band 0 还经过 21-tap 复 FIR:
|
||||
|
||||
$$
|
||||
\widehat X_{c,0,t}
|
||||
=
|
||||
\sum_{k=0}^{20}h_kX_{c,0,t-k}.
|
||||
$$
|
||||
|
||||
这些延迟和滤波历史属于解码状态,不能按帧独立清零。
|
||||
|
||||
## 6. 对象矩阵
|
||||
|
||||
对每个对象 $o$、子带 $b$ 和时槽 $t$,对象频域值为五个核心声道的线性组合:
|
||||
|
||||
$$
|
||||
Z_{o,b,t}
|
||||
=
|
||||
\sum_{c=0}^{4}
|
||||
M_{o,c,b,t}\widehat X_{c,b,t}.
|
||||
$$
|
||||
|
||||
analysis 输入的 $1/16$ 缩放会在 inverse QMF 输出端由 $\times16$ 抵消,因此矩阵本身不需要额外经验增益。
|
||||
|
||||
## 7. 对象 inverse QMF
|
||||
|
||||
### 7.1 子带重排
|
||||
|
||||
将 64 个复子带写成 128 个交织实数 `src`。对 $k=0\ldots31$:
|
||||
|
||||
$$
|
||||
\begin{aligned}
|
||||
\operatorname{zone}[2k] &= \operatorname{src}[4k],\\
|
||||
\operatorname{zone}[2k+1] &= -\operatorname{src}[4k+1],\\
|
||||
\operatorname{zone}[126-2k] &= \operatorname{src}[4k+2],\\
|
||||
\operatorname{zone}[127-2k] &= \operatorname{src}[4k+3].
|
||||
\end{aligned}
|
||||
$$
|
||||
|
||||
把 `zone` 重新视为 64 个复数后执行未归一化 64 点 FFT:
|
||||
|
||||
$$
|
||||
F_k
|
||||
=
|
||||
\sum_{n=0}^{63}
|
||||
\operatorname{zone}_n
|
||||
\exp\!\left(-j\frac{2\pi kn}{64}\right).
|
||||
$$
|
||||
|
||||
### 7.2 调制与合成
|
||||
|
||||
定义旋转系数
|
||||
|
||||
$$
|
||||
r_k
|
||||
=
|
||||
\frac12\left(
|
||||
\sin\frac{\pi k}{128}
|
||||
+j\cos\frac{\pi k}{128}
|
||||
\right),
|
||||
$$
|
||||
|
||||
并计算
|
||||
|
||||
$$
|
||||
R_k=2F_kr_k.
|
||||
$$
|
||||
|
||||
令 $\mathcal S$ 表示带 640 项 synthesis window 和跨时槽状态的 polyphase 合成算子:
|
||||
|
||||
$$
|
||||
\mathbf y_{o,t}
|
||||
=
|
||||
\mathcal S\!\left(
|
||||
\mathbf R_{o,t},W,\mathbf s^{\mathrm S}_{o,t}
|
||||
\right).
|
||||
$$
|
||||
|
||||
对象输出为
|
||||
|
||||
$$
|
||||
y_o[64t+r]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
16\,\mathbf y_{o,t}[r],-1,1
|
||||
\right)G_{\mathrm{clip}},
|
||||
$$
|
||||
|
||||
其中 $r=0\ldots63$。synthesis 状态必须按时槽连续推进。
|
||||
|
||||
## 8. LFE 路径
|
||||
|
||||
LFE 不经过对象矩阵或 inverse QMF,而是使用 1217-sample 延迟。输入与输出端的比例因子抵消后:
|
||||
|
||||
$$
|
||||
y_{\mathrm{LFE}}[n]
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
x_{\mathrm{LFE,core}}[n-1217],-1,1
|
||||
\right).
|
||||
$$
|
||||
|
||||
## 9. OAMD 坐标
|
||||
|
||||
横向和纵向网格使用 $N=62$,高度网格使用 $N=15$。量化函数为
|
||||
|
||||
$$
|
||||
q_N(k)
|
||||
=
|
||||
\min\!\left(
|
||||
32767,
|
||||
\left\lfloor\frac{32768k}{N}+\frac12\right\rfloor
|
||||
\right).
|
||||
$$
|
||||
|
||||
OAR 坐标为
|
||||
|
||||
$$
|
||||
u=\frac{q_1}{32768},
|
||||
\qquad
|
||||
v=\frac{q_2}{32768},
|
||||
\qquad
|
||||
w=\frac{q_3}{32768}.
|
||||
$$
|
||||
|
||||
其最大运行值为 $32767/32768$,不是精确的 1。
|
||||
|
||||
转换为 ADM 网格时:
|
||||
|
||||
$$
|
||||
k_1=\operatorname{round}\!\left(\frac{62q_1}{32767}\right),
|
||||
\quad
|
||||
k_2=\operatorname{round}\!\left(\frac{62q_2}{32767}\right),
|
||||
\quad
|
||||
k_3=\operatorname{round}\!\left(\frac{15q_3}{32767}\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
X=2\frac{k_1}{62}-1,
|
||||
\qquad
|
||||
Y=1-2\frac{k_2}{62},
|
||||
\qquad
|
||||
Z=\frac{k_3}{15}.
|
||||
$$
|
||||
|
||||
连续坐标关系为
|
||||
|
||||
$$
|
||||
u=\frac{X+1}{2},
|
||||
\qquad
|
||||
v=\frac{1-Y}{2},
|
||||
\qquad
|
||||
w=Z.
|
||||
$$
|
||||
|
||||
## 10. 等功率扬声器声像
|
||||
|
||||
### 10.1 一维插值
|
||||
|
||||
相邻扬声器坐标为 $a_0<a_1$,对象位置为 $a$。归一化位置为
|
||||
|
||||
$$
|
||||
\tau=\frac{a-a_0}{a_1-a_0}.
|
||||
$$
|
||||
|
||||
区间内的增益为
|
||||
|
||||
$$
|
||||
g_0(\tau)=\cos\left(\frac\pi2\tau\right),
|
||||
\qquad
|
||||
g_1(\tau)=\sin\left(\frac\pi2\tau\right),
|
||||
$$
|
||||
|
||||
并满足
|
||||
|
||||
$$
|
||||
g_0^2(\tau)+g_1^2(\tau)=1.
|
||||
$$
|
||||
|
||||
区间外的对象位置夹到最近端点。
|
||||
|
||||
### 10.2 二维 region
|
||||
|
||||
每一行先沿 $u$ 得到横向增益向量 $\mathbf h_r(u)$。若对象位于相邻两行 $r_0,r_1$ 之间:
|
||||
|
||||
$$
|
||||
\eta=\frac{v-v_{r_0}}{v_{r_1}-v_{r_0}},
|
||||
$$
|
||||
|
||||
$$
|
||||
a_0=\cos\left(\frac\pi2\eta\right),
|
||||
\qquad
|
||||
a_1=\sin\left(\frac\pi2\eta\right).
|
||||
$$
|
||||
|
||||
二维点增益为
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{2D}}(u,v)
|
||||
=
|
||||
\mathbf h(u)\odot\mathbf v(v).
|
||||
$$
|
||||
|
||||
对于只有一对水平环绕、没有独立 side/rear 两对的 5.1 系列布局,纵向坐标使用
|
||||
|
||||
$$
|
||||
v_{\mathrm{floor}}
|
||||
=
|
||||
\operatorname{clamp}(2v,0,1).
|
||||
$$
|
||||
|
||||
其他布局使用 $v_{\mathrm{floor}}=v$。
|
||||
|
||||
### 10.3 高度层
|
||||
|
||||
三维布局分别计算地面层增益 $\mathbf G_f$ 和高度层增益 $\mathbf G_h$:
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{point}}(u,v,w)
|
||||
=
|
||||
\cos\left(\frac\pi2w\right)\mathbf G_f
|
||||
+
|
||||
\sin\left(\frac\pi2w\right)\mathbf G_h.
|
||||
$$
|
||||
|
||||
当地面层与高度层扬声器集合不重叠、且各层内部使用等功率插值时:
|
||||
|
||||
$$
|
||||
\left\|\mathbf G_{\mathrm{point}}\right\|_2=1.
|
||||
$$
|
||||
|
||||
## 11. 布局位置补偿
|
||||
|
||||
令 $N_h$ 为相关高度扬声器数,$N_f$ 为相关附加水平扬声器数:
|
||||
|
||||
$$
|
||||
H=\min\left(\frac{N_h}{4},1\right),
|
||||
\qquad
|
||||
F=\min\left(\frac{N_f}{4},1\right).
|
||||
$$
|
||||
|
||||
最大位置补偿为
|
||||
|
||||
$$
|
||||
A_{\max}
|
||||
=
|
||||
-\max\left(4.5-1.5H-3F,0\right)
|
||||
\quad\text{dB}.
|
||||
$$
|
||||
|
||||
前后与高度位置权重为
|
||||
|
||||
$$
|
||||
p_v=\operatorname{clamp}\left(\frac v{0.6},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p_w=\operatorname{clamp}\left(\frac{w-0.2}{0.8},0,1\right),
|
||||
$$
|
||||
|
||||
$$
|
||||
p=\operatorname{clamp}(p_v+p_w,0,1).
|
||||
$$
|
||||
|
||||
线性补偿增益为
|
||||
|
||||
$$
|
||||
G_{\mathrm{pos}}=10^{A_{\max}p/20}.
|
||||
$$
|
||||
|
||||
对象的目标增益向量为
|
||||
|
||||
$$
|
||||
\mathbf G_{\mathrm{target}}
|
||||
=
|
||||
G_{\mathrm{object}}
|
||||
G_{\mathrm{pos}}
|
||||
\mathbf G_{\mathrm{point}}.
|
||||
$$
|
||||
|
||||
## 12. OAMD 时间对齐与增益斜坡
|
||||
|
||||
OAMD 更新的编码位置为
|
||||
|
||||
$$
|
||||
s_{\mathrm{coded}}
|
||||
=
|
||||
s_{\mathrm{frame}}
|
||||
+s_{\mathrm{outer}}
|
||||
+s_{\mathrm{OAMD}}
|
||||
+32f_{\mathrm{block}}.
|
||||
$$
|
||||
|
||||
对处理块长度 $B=32$,更新点对齐为
|
||||
|
||||
$$
|
||||
\widehat s
|
||||
=
|
||||
B\left\lfloor
|
||||
\frac{s_{\mathrm{coded}}+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
$$
|
||||
|
||||
给定 ramp duration $D$,block 数为
|
||||
|
||||
$$
|
||||
K
|
||||
=
|
||||
\left\lfloor
|
||||
\frac{D+B/2-1}{B}
|
||||
\right\rfloor.
|
||||
$$
|
||||
|
||||
若当前增益为 $g_0$、目标为 $g_1$,则每 block 的增量为
|
||||
|
||||
$$
|
||||
\Delta g=\frac{g_1-g_0}{K}.
|
||||
$$
|
||||
|
||||
第 $j$ 个 block 内的样本 $r=0\ldots B-1$ 使用
|
||||
|
||||
$$
|
||||
g_{j,r}=g_j+\frac rB\Delta g,
|
||||
\qquad
|
||||
g_{j+1}=g_j+\Delta g.
|
||||
$$
|
||||
|
||||
如果中途没有新的 metadata 更新,该过程等价于总长度 $KB$ 的逐样本线性斜坡。
|
||||
|
||||
## 13. 最终扬声器混音
|
||||
|
||||
对目标输出声道 $c$:
|
||||
|
||||
$$
|
||||
y_c[n]
|
||||
=
|
||||
\delta_{c,\mathrm{LFE}}x_{\mathrm{LFE}}[n]
|
||||
+
|
||||
\sum_{o=1}^{15}x_o[n]g_{o,c}[n].
|
||||
$$
|
||||
|
||||
其中
|
||||
|
||||
$$
|
||||
\delta_{c,\mathrm{LFE}}
|
||||
=
|
||||
\begin{cases}
|
||||
1, & c\text{ 为目标布局的 LFE},\\
|
||||
0, & \text{其他声道}.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
没有 LFE 输出的布局不把输入 LFE 混入其他声道。对象完成累加后,再按目标格式要求排列输出声道。
|
||||
|
||||
若输出 PCM24,量化关系为
|
||||
|
||||
$$
|
||||
y_{24}[n]
|
||||
=
|
||||
\operatorname{trunc}\left(
|
||||
8388607\,\operatorname{clip}(y[n],-1,1)
|
||||
\right).
|
||||
$$
|
||||
|
||||
## 14. 公式适用范围
|
||||
|
||||
- JOC 矩阵部分描述 dense JOC;Sparse JOC 使用不同的稀疏系数/索引路径。
|
||||
- 扬声器声像部分描述普通点对象;extent、spread、divergence 等模式需要额外模型。
|
||||
- 多个 OAMD position block 必须按其时间顺序调度。
|
||||
- limiter 属于独立后处理,不包含在上述混音公式中。
|
||||
@@ -0,0 +1,159 @@
|
||||
# JustOneCacophony native-core notes
|
||||
|
||||
[中文](native.md) · [Back to README](../README.en.md)
|
||||
|
||||
## 1. Responsibility boundary
|
||||
|
||||
`native/` contains only the state-heavy, frequently called DSP and speaker-rendering kernels. High-level EMDF/JOC/OAMD parsing, error reporting, ADM assembly, and the CLI remain in Python.
|
||||
|
||||
Python calls a C ABI through the standard-library `ctypes` module. The native core does not use pybind11, Cython, FFTW, MKL, or OpenMP. It is an optional acceleration path and does not expand the set of supported stream variants.
|
||||
|
||||
Main files:
|
||||
|
||||
```text
|
||||
native/include/eac3joc_core.h C ABI
|
||||
native/src/eac3joc_core.cpp JOC/QMF object reconstruction
|
||||
native/src/speaker_renderer.cpp object-to-speaker rendering
|
||||
native/src/qmf_tables.h QMF tables
|
||||
native/src/speaker_layouts.h layout tables
|
||||
native/src/joc_huffman_tables.h JOC Huffman tables
|
||||
src/native_renderer.py JOC ctypes bridge
|
||||
src/speaker_native_renderer.py speaker ctypes bridge
|
||||
```
|
||||
|
||||
## 2. JOC rendering ABI
|
||||
|
||||
An opaque renderer owns all cross-frame state. Its main call is:
|
||||
|
||||
```c
|
||||
int ejoc_renderer_process(
|
||||
ejoc_renderer_handle handle,
|
||||
const float* bed5_planar, /* [5][1536] */
|
||||
const float* lfe, /* [1536] or NULL */
|
||||
uint32_t object_mask,
|
||||
const uint8_t* n_bands, /* [15] */
|
||||
const uint8_t* n_dpoints, /* [15] */
|
||||
const uint8_t* slope_idx, /* [15] */
|
||||
const uint8_t* offset_ts, /* [15][2] */
|
||||
const double* dq, /* [15][2][5][23] */
|
||||
double clipgain,
|
||||
float phase_new,
|
||||
float output_scale,
|
||||
float* output16_planar); /* [16][1536] */
|
||||
```
|
||||
|
||||
Python performs dense-JOC Huffman decoding, differential reconstruction, and dequantization before the call. Sparse JOC is not silently passed to the dense native path.
|
||||
|
||||
Thread control is exposed as:
|
||||
|
||||
```c
|
||||
int ejoc_renderer_set_threads(ejoc_renderer_handle handle, uint32_t total_threads);
|
||||
uint32_t ejoc_renderer_thread_count(ejoc_renderer_handle handle);
|
||||
```
|
||||
|
||||
`total_threads` includes the calling thread. Frames must be submitted sequentially to one renderer instance; the instance may parallelize work across objects and analysis channels.
|
||||
|
||||
## 3. Cross-frame state
|
||||
|
||||
Each JOC renderer stores:
|
||||
|
||||
- analysis FIFO: `double[5][9][64]`;
|
||||
- L/R/C analysis delay: `float[3][10][64]`;
|
||||
- Ls/Rs QMF delay: `complex<double>[2][10][64]`;
|
||||
- Ls/Rs band-0 FIR history: `complex<double>[2][20]`;
|
||||
- previous matrix interpolation values: `double[15][5][64]`;
|
||||
- inverse-QMF state: `double[15][640]`;
|
||||
- LFE delay: `double[1217]`.
|
||||
|
||||
This state belongs to the renderer instance. Processing cannot be arbitrarily segmented or reordered without a corresponding state checkpoint.
|
||||
|
||||
## 4. FFT, QMF, and precision
|
||||
|
||||
The native core contains a fixed 64-point radix-2 complex FFT:
|
||||
|
||||
- analysis QMF uses a forward FFT followed by division by 64;
|
||||
- inverse QMF uses the fixed reorder, rotation, and 640-value active-window state;
|
||||
- no external FFT library is called.
|
||||
|
||||
The JOC path uses:
|
||||
|
||||
- float32 core-PCM input;
|
||||
- double matrices, complex QMF, FFT, FIR, and cross-frame state;
|
||||
- float32 phase and final gain;
|
||||
- float32 16-channel object output.
|
||||
|
||||
## 5. Speaker-rendering ABI
|
||||
|
||||
The same shared library exports object-to-speaker rendering:
|
||||
|
||||
```c
|
||||
uint32_t ejoc_speaker_layout_channel_count(uint32_t speaker_bitfield);
|
||||
|
||||
ejoc_speaker_renderer_handle
|
||||
ejoc_speaker_renderer_create(uint32_t speaker_bitfield);
|
||||
|
||||
int ejoc_speaker_renderer_process(
|
||||
ejoc_speaker_renderer_handle handle,
|
||||
const float* objects16_interleaved,
|
||||
uint32_t sample_count,
|
||||
uint32_t metadata_count,
|
||||
const uint32_t* metadata_offsets,
|
||||
const uint32_t* ramp_durations,
|
||||
const uint16_t* positions_q15,
|
||||
const uint8_t* region_indices,
|
||||
const uint8_t* height_enabled,
|
||||
const double* object_gains,
|
||||
double* output_interleaved);
|
||||
```
|
||||
|
||||
Input channel 0 is LFE and channels 1–15 are objects. Each metadata entry is an object-state snapshot. `sample_count` must be a multiple of 32; unfinished gain ramps remain in the handle and continue across calls.
|
||||
|
||||
The speaker path uses float32 object input, double coordinates/gains/accumulation, and interleaved double output. Quantization to float32 or PCM24 happens when the WAV is written.
|
||||
|
||||
Supported layouts:
|
||||
|
||||
```text
|
||||
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
|
||||
```
|
||||
|
||||
## 6. Building
|
||||
|
||||
The CMake definition is `native/CMakeLists.txt`. Run from the repository root:
|
||||
|
||||
```powershell
|
||||
cmake -S native -B build/cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX="$PWD/lib"
|
||||
cmake --build build/cmake --config Release
|
||||
cmake --install build/cmake --config Release
|
||||
```
|
||||
|
||||
Platform runtime names:
|
||||
|
||||
```text
|
||||
Windows lib/eac3joc_core.dll
|
||||
Linux lib/libeac3joc_core.so
|
||||
macOS lib/libeac3joc_core.dylib
|
||||
```
|
||||
|
||||
The MSVC configuration uses the static CRT. Other runtime dependencies depend on the platform and toolchain and should be checked independently before publishing a prebuilt library.
|
||||
|
||||
The repository does not include native binaries by default. A prebuilt Release runtime or a locally built runtime can be placed directly under `lib/`.
|
||||
|
||||
## 7. Runtime lookup and fallback
|
||||
|
||||
Lookup order:
|
||||
|
||||
1. explicit `--native-library`;
|
||||
2. `EAC3JOC_NATIVE_LIBRARY`;
|
||||
3. the standard platform filename under `lib/`.
|
||||
|
||||
`--backend auto` falls back to NumPy when loading fails, and `--backend python` skips native discovery. The current CLI also prints the failure and falls back for `--backend native`; this existing behavior should not be read as successful native execution.
|
||||
|
||||
## 8. Implementation boundaries
|
||||
|
||||
- The native layer accepts only dense-JOC data already parsed by Python.
|
||||
- The ABI fixes a 1536-sample JOC frame, at most 15 objects, at most 23 parameter bands, and at most 2 data points.
|
||||
- The shared library and Python bridge must report the same ABI version.
|
||||
- Only the ABI and data types are specified across platforms; bit-identical float64 results are not guaranteed.
|
||||
- Private table headers under `native/src/` serve the native side only. The current repository does not include the scripts that generated those headers.
|
||||
|
||||
See the [mathematical notes](math.en.md) for the related formulas.
|
||||
+159
@@ -0,0 +1,159 @@
|
||||
# JustOneCacophony 原生核说明
|
||||
|
||||
[English](native.en.md) · [返回 README](../README.md)
|
||||
|
||||
## 1. 职责边界
|
||||
|
||||
`native/` 只承载状态密集、调用频繁的 DSP 与扬声器渲染核。EMDF/JOC/OAMD 高层解析、错误报告、ADM 组装和 CLI 保留在 Python 中。
|
||||
|
||||
Python 通过标准库 `ctypes` 调用 C ABI;原生核不使用 pybind11、Cython、FFTW、MKL 或 OpenMP。它是可选加速路径,不扩大项目所支持的码流范围。
|
||||
|
||||
主要文件:
|
||||
|
||||
```text
|
||||
native/include/eac3joc_core.h C ABI
|
||||
native/src/eac3joc_core.cpp JOC/QMF 对象重建
|
||||
native/src/speaker_renderer.cpp 对象到扬声器渲染
|
||||
native/src/qmf_tables.h QMF 表
|
||||
native/src/speaker_layouts.h 布局表
|
||||
native/src/joc_huffman_tables.h JOC Huffman 表
|
||||
src/native_renderer.py JOC ctypes 桥
|
||||
src/speaker_native_renderer.py 扬声器 ctypes 桥
|
||||
```
|
||||
|
||||
## 2. JOC 渲染 ABI
|
||||
|
||||
一个 opaque renderer 保存所有跨帧状态。主要调用为:
|
||||
|
||||
```c
|
||||
int ejoc_renderer_process(
|
||||
ejoc_renderer_handle handle,
|
||||
const float* bed5_planar, /* [5][1536] */
|
||||
const float* lfe, /* [1536] or NULL */
|
||||
uint32_t object_mask,
|
||||
const uint8_t* n_bands, /* [15] */
|
||||
const uint8_t* n_dpoints, /* [15] */
|
||||
const uint8_t* slope_idx, /* [15] */
|
||||
const uint8_t* offset_ts, /* [15][2] */
|
||||
const double* dq, /* [15][2][5][23] */
|
||||
double clipgain,
|
||||
float phase_new,
|
||||
float output_scale,
|
||||
float* output16_planar); /* [16][1536] */
|
||||
```
|
||||
|
||||
Dense JOC 的 Huffman 解码、差分还原和去量化先在 Python 中完成。Sparse JOC 不会被静默送入 dense 原生路径。
|
||||
|
||||
线程接口为:
|
||||
|
||||
```c
|
||||
int ejoc_renderer_set_threads(ejoc_renderer_handle handle, uint32_t total_threads);
|
||||
uint32_t ejoc_renderer_thread_count(ejoc_renderer_handle handle);
|
||||
```
|
||||
|
||||
`total_threads` 包含调用线程。单个 renderer 实例必须顺序提交帧;实例内部可以按对象和 analysis channel 并行。
|
||||
|
||||
## 3. 跨帧状态
|
||||
|
||||
每个 JOC renderer 独立保存:
|
||||
|
||||
- analysis FIFO:`double[5][9][64]`;
|
||||
- L/R/C analysis delay:`float[3][10][64]`;
|
||||
- Ls/Rs QMF delay:`complex<double>[2][10][64]`;
|
||||
- Ls/Rs band-0 FIR history:`complex<double>[2][20]`;
|
||||
- 矩阵插值 previous:`double[15][5][64]`;
|
||||
- inverse-QMF state:`double[15][640]`;
|
||||
- LFE delay:`double[1217]`。
|
||||
|
||||
这些状态属于 renderer 实例,不能在无 checkpoint 的情况下任意分段或乱序处理。
|
||||
|
||||
## 4. FFT、QMF 与精度
|
||||
|
||||
原生核包含固定 64 点 radix-2 complex FFT:
|
||||
|
||||
- analysis QMF 使用 forward FFT 后除以 64;
|
||||
- inverse QMF 使用固定重排、旋转和 640 项有效窗状态;
|
||||
- 不调用外部 FFT 库。
|
||||
|
||||
JOC 路径的数值类型为:
|
||||
|
||||
- 核心 PCM 输入:float32;
|
||||
- 矩阵、复 QMF、FFT、FIR 和跨帧状态:double;
|
||||
- phase 与最终 gain:float32;
|
||||
- 16 声道对象输出:float32。
|
||||
|
||||
## 5. 扬声器渲染 ABI
|
||||
|
||||
同一个共享库还导出对象到扬声器布局的渲染接口:
|
||||
|
||||
```c
|
||||
uint32_t ejoc_speaker_layout_channel_count(uint32_t speaker_bitfield);
|
||||
|
||||
ejoc_speaker_renderer_handle
|
||||
ejoc_speaker_renderer_create(uint32_t speaker_bitfield);
|
||||
|
||||
int ejoc_speaker_renderer_process(
|
||||
ejoc_speaker_renderer_handle handle,
|
||||
const float* objects16_interleaved,
|
||||
uint32_t sample_count,
|
||||
uint32_t metadata_count,
|
||||
const uint32_t* metadata_offsets,
|
||||
const uint32_t* ramp_durations,
|
||||
const uint16_t* positions_q15,
|
||||
const uint8_t* region_indices,
|
||||
const uint8_t* height_enabled,
|
||||
const double* object_gains,
|
||||
double* output_interleaved);
|
||||
```
|
||||
|
||||
输入声道 0 为 LFE,1–15 为对象。每个 metadata entry 是一份对象状态快照。`sample_count` 必须是 32 的倍数;未完成的增益斜坡保存在 handle 中并跨调用继续。
|
||||
|
||||
扬声器路径使用 float32 对象输入、double 坐标/增益/累加与 interleaved double 输出;写 WAV 时才量化为 float32 或 PCM24。
|
||||
|
||||
支持的布局为:
|
||||
|
||||
```text
|
||||
2.0 3.1 5.1 7.1 5.1.2 5.1.4 7.1.2 7.1.4 9.1.4 9.1.6
|
||||
```
|
||||
|
||||
## 6. 构建
|
||||
|
||||
CMake 定义位于 `native/CMakeLists.txt`。从仓库根目录运行:
|
||||
|
||||
```powershell
|
||||
cmake -S native -B build/cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX="$PWD/lib"
|
||||
cmake --build build/cmake --config Release
|
||||
cmake --install build/cmake --config Release
|
||||
```
|
||||
|
||||
平台运行库文件名:
|
||||
|
||||
```text
|
||||
Windows lib/eac3joc_core.dll
|
||||
Linux lib/libeac3joc_core.so
|
||||
macOS lib/libeac3joc_core.dylib
|
||||
```
|
||||
|
||||
MSVC 配置使用静态 CRT。其他运行时依赖由平台和工具链决定,发布预构建库前应对产物独立检查。
|
||||
|
||||
仓库默认不附带原生二进制。预构建的 Release 运行库或自行构建的运行库均可直接放入 `lib/`。
|
||||
|
||||
## 7. 运行时查找与回退
|
||||
|
||||
查找顺序为:
|
||||
|
||||
1. 显式 `--native-library`;
|
||||
2. `EAC3JOC_NATIVE_LIBRARY`;
|
||||
3. `lib/` 下当前平台的标准文件名。
|
||||
|
||||
`--backend auto` 在加载失败时回退到 NumPy;`--backend python` 跳过原生探测。`--backend native` 当前也会打印失败原因后回退,这是现有 CLI 行为,不应理解为原生库已成功使用。
|
||||
|
||||
## 8. 实现边界
|
||||
|
||||
- 原生层只接收 Python 已解析的 dense JOC 数据。
|
||||
- ABI 固定了 1536-sample JOC 帧、最多 15 个对象、最多 23 个参数带和最多 2 个数据点。
|
||||
- 共享库与 Python 桥需要 ABI version 一致。
|
||||
- 跨平台只约定 ABI 与数据类型,不保证 float64 结果逐位一致。
|
||||
- `native/src/` 中的私有表头只服务于原生侧;当前仓库不包含重新生成这些头文件的脚本。
|
||||
|
||||
相关公式见[数学说明](math.md)。
|
||||
Reference in New Issue
Block a user