diff --git a/README.en.md b/README.en.md index afd854f..22d2a24 100644 --- a/README.en.md +++ b/README.en.md @@ -288,7 +288,6 @@ See the [mathematical notes](docs/math.en.md) for the equations used by the deco ## Known limitations - Only the common contiguous EMDF transport is covered. Fragmented transport across multiple audio-block skip fields is not covered. -- Dense JOC is the main path. The Sparse JOC branch should not be treated as supported. - The speaker and SOFA binaural paths currently cover ordinary point objects; extent, spread, diffuse, divergence, channel lock, and similar controls are outside the supported scope. - OAMD trim elements are boundary-checked and skipped; warp, balance, and trim parameters are not applied to raw object trajectories or speaker rendering. - Multi-data-point streams, uncommon band configurations, and unusual OAMD scheduling have less coverage than common 12-band, single-data-point material. diff --git a/README.md b/README.md index 69316fe..4cb8476 100644 --- a/README.md +++ b/README.md @@ -258,7 +258,6 @@ JustOneCacophony/ ## 已知限制 - 当前只覆盖常见 continuous EMDF transport;跨多个 audio-block skip field 的碎片化 transport 尚未覆盖。 -- Dense JOC 是当前主要路径;Sparse JOC 分支不应视为受支持能力。 - 扬声器与 SOFA 双耳路径当前只覆盖普通点对象;extent、spread、diffuse、divergence、channel lock 等对象控制不在支持范围内。 - OAMD trim element 会按声明边界校验并跳过;warp、balance 和 trim 参数不应用于当前原始对象轨迹或扬声器渲染。 - 多数据点、少见参数带配置和特殊 OAMD 调度的覆盖度低于常见 12-band、单数据点素材。 diff --git a/docs/math.en.md b/docs/math.en.md index f3a1d59..dfa9955 100644 --- a/docs/math.en.md +++ b/docs/math.en.md @@ -4,7 +4,7 @@ This document covers only the signal model and formulas used in the JustOneCacophony research path: how JOC parameters combine with core PCM to reconstruct object signals, and how OAMD coordinates become speaker gains. -The formulas describe the dense-JOC and ordinary point-object paths studied by the project. They are not a complete definition of every E-AC-3 JOC variant. +The formulas describe the JOC matrix parameters (both the dense and the sparse differential syntax) and the ordinary point-object paths studied by the project. They are not a complete definition of every E-AC-3 JOC variant. ## 1. Overall path and notation @@ -52,9 +52,11 @@ $$ N_f=1536=24\times64. $$ -## 2. Dense-JOC matrix parameters +## 2. JOC matrix parameters -### 2.1 Differential reconstruction +For every object and data point, the quantized matrix `joc_mix_mtx_q` is defined on $N_q$ quantization levels. The `b_joc_sparse` flag selects one of two differential syntaxes: dense sends one MTX difference per core channel, while sparse sends one active channel plus one coefficient difference per parameter band. + +### 2.1 Dense differential reconstruction Let `quant_idx` be $q_i\in\{0,1\}$. The number of quantization levels is @@ -85,7 +87,49 @@ Q_{o,d,c,p}= \qquad p>0. $$ -### 2.2 Dequantization +### 2.2 Sparse differential reconstruction + +Let $I_{o,d,p}$ be the `joc_channel_idx` symbol (IDX), $V_{o,d,p}$ the `joc_vec` symbol (VEC), and $N_c\in\{5,7\}$ the number of core channels. Each parameter band has exactly one active channel: + +$$ +A_{o,d,p}= +\begin{cases} +I_{o,d,0}, & p=0,\\[2pt] +\left(A_{o,d,p-1}+I_{o,d,p}\right)\bmod N_c, & p>0, +\end{cases} +$$ + +where $I_{o,d,0}$ is a 3-bit absolute channel index and every later IDX symbol is an increment relative to the previous **active channel**. The coefficient is a single accumulator running across parameter bands: + +$$ +\kappa_{o,d,-1}=O^{(s)}_q,\qquad +\kappa_{o,d,p}= +\left(\kappa_{o,d,p-1}+V_{o,d,p}\right)\bmod N_q, +$$ + +with a sparse starting point two quantization levels above the dense center offset: + +$$ +O^{(s)}_q= +\begin{cases} +50, & q_i=0,\\ +100, & q_i=1. +\end{cases} +$$ + +The accumulator is **not** reset when the active channel changes. The complete matrix is + +$$ +Q_{o,d,c,p}= +\begin{cases} +\kappa_{o,d,p}, & c=A_{o,d,p},\\[2pt] +\dfrac{N_q}{2}, & c\neq A_{o,d,p}. +\end{cases} +$$ + +Non-active entries take $N_q/2$, which dequantizes to exactly 0. + +### 2.3 Dequantization The dequantized matrix coefficient is @@ -97,7 +141,7 @@ $$ The effective denominator is therefore 4096 in coarse mode and 8192 in fine mode. -### 2.3 JOC clipgain +### 2.4 JOC clipgain If the clipgain field consists of integer $x$ and mantissa $y$, then @@ -557,7 +601,7 @@ $$ ## 14. Scope of the formulas -- The JOC matrix section describes dense JOC; Sparse JOC uses a different sparse coefficient/index path. +- The JOC matrix section covers both the dense MTX and the sparse IDX/VEC differential syntax. - The speaker-panning section describes ordinary point objects; extent, spread, divergence, and similar modes require additional models. - Multiple OAMD position blocks must be scheduled in time order. - A limiter is separate post-processing and is not included in the mixing equations above. diff --git a/docs/math.md b/docs/math.md index d4bd414..4c89279 100644 --- a/docs/math.md +++ b/docs/math.md @@ -4,7 +4,7 @@ 本文只说明 JustOneCacophony 研究路径中使用的信号模型和公式:JOC 参数如何与核心 PCM 结合并重建对象信号,以及 OAMD 坐标如何转换为扬声器增益。 -这些公式描述项目当前研究的 dense JOC 与普通点对象路径,不代表对所有 E-AC-3 JOC 变体的完整定义。 +这些公式描述项目当前研究的 JOC 矩阵参数(dense 与 sparse 两条差分语法)与普通点对象路径,不代表对所有 E-AC-3 JOC 变体的完整定义。 ## 1. 总体路径与记号 @@ -52,9 +52,11 @@ $$ N_f=1536=24\times64. $$ -## 2. Dense JOC 矩阵参数 +## 2. JOC 矩阵参数 -### 2.1 差分还原 +每个对象、每个数据点的量化矩阵 `joc_mix_mtx_q` 都定义在 $N_q$ 个量化级上。标志位 `b_joc_sparse` 选择两条差分语法之一:dense 为每个核心声道各送一路 MTX 差分,sparse 每参数带只送一个 active 声道与一路系数差分。 + +### 2.1 Dense 差分还原 令 `quant_idx` 为 $q_i\in\{0,1\}$,量化级数为 @@ -85,7 +87,49 @@ Q_{o,d,c,p}= \qquad p>0. $$ -### 2.2 去量化 +### 2.2 Sparse 差分还原 + +令 $I_{o,d,p}$ 为 `joc_channel_idx` 符号(IDX),$V_{o,d,p}$ 为 `joc_vec` 符号(VEC),$N_c\in\{5,7\}$ 为核心声道数。每参数带只有一个 active 声道 + +$$ +A_{o,d,p}= +\begin{cases} +I_{o,d,0}, & p=0,\\[2pt] +\left(A_{o,d,p-1}+I_{o,d,p}\right)\bmod N_c, & p>0, +\end{cases} +$$ + +其中 $I_{o,d,0}$ 是 3 bit 绝对声道号,其余 IDX 符号是相对上一个 **active 声道**的增量。系数是一个跨参数带连续的单累加器 + +$$ +\kappa_{o,d,-1}=O^{(s)}_q,\qquad +\kappa_{o,d,p}= +\left(\kappa_{o,d,p-1}+V_{o,d,p}\right)\bmod N_q, +$$ + +sparse 起点比 dense 的中心偏移高两个量化级: + +$$ +O^{(s)}_q= +\begin{cases} +50, & q_i=0,\\ +100, & q_i=1. +\end{cases} +$$ + +active 声道切换时累加器**不**重置。完整矩阵为 + +$$ +Q_{o,d,c,p}= +\begin{cases} +\kappa_{o,d,p}, & c=A_{o,d,p},\\[2pt] +\dfrac{N_q}{2}, & c\neq A_{o,d,p}. +\end{cases} +$$ + +非 active 项取 $N_q/2$,即去量化后恰为 0。 + +### 2.3 去量化 矩阵系数的去量化值为 @@ -97,7 +141,7 @@ $$ 因此 coarse 模式的有效分母为 4096,fine 模式为 8192。 -### 2.3 JOC clipgain +### 2.4 JOC clipgain 若 clipgain 字段由整数 $x$ 和尾数 $y$ 组成,则 @@ -557,7 +601,7 @@ $$ ## 14. 公式适用范围 -- JOC 矩阵部分描述 dense JOC;Sparse JOC 使用不同的稀疏系数/索引路径。 +- JOC 矩阵部分同时描述 dense MTX 与 sparse IDX/VEC 两条差分语法。 - 扬声器声像部分描述普通点对象;extent、spread、divergence 等模式需要额外模型。 - 多个 OAMD position block 必须按其时间顺序调度。 - limiter 属于独立后处理,不包含在上述混音公式中。 diff --git a/docs/native.en.md b/docs/native.en.md index 3f5a1d6..d50cad7 100644 --- a/docs/native.en.md +++ b/docs/native.en.md @@ -42,7 +42,7 @@ int ejoc_renderer_process( float* output16_planar); /* [16][1536] */ ``` -Python performs dense-JOC Huffman decoding, differential reconstruction, and dequantization before the call. Sparse JOC is not silently passed to the dense native path. +Python performs JOC Huffman decoding, differential reconstruction, and dequantization before the call, with the dense and sparse syntaxes sharing one entry point. The native core consumes the already dequantized `dq` in double precision, and both syntaxes have the same layout at that ABI. Thread control is exposed as: diff --git a/docs/native.md b/docs/native.md index 970211c..f5ef3f3 100644 --- a/docs/native.md +++ b/docs/native.md @@ -42,7 +42,7 @@ int ejoc_renderer_process( float* output16_planar); /* [16][1536] */ ``` -Dense JOC 的 Huffman 解码、差分还原和去量化先在 Python 中完成。Sparse JOC 不会被静默送入 dense 原生路径。 +JOC 的 Huffman 解码、差分还原和去量化先在 Python 中完成,dense 与 sparse 两条语法共用同一条入口。原生核心消费已去量化的 `dq`(double),两条语法在该 ABI 上布局一致。 线程接口为: diff --git a/native/include/eac3joc_core.h b/native/include/eac3joc_core.h index 4bf7953..53e8bd4 100644 --- a/native/include/eac3joc_core.h +++ b/native/include/eac3joc_core.h @@ -53,8 +53,9 @@ Fixed array layouts used by ejoc_renderer_process(): output16 [16][1536] Only objects selected by object_mask are read from the descriptor arrays. -Sparse JOC must be rejected by the caller; this ABI accepts already dequantized -dense matrix coefficients. +dq carries already dequantized matrix coefficients in double precision; the +caller performs the JOC bitstream differential decoding for both dense and +sparse objects, so this ABI is identical for both syntaxes. */ EJOC_API uint32_t EJOC_CALL ejoc_abi_version(void); diff --git a/src/joc_decode.py b/src/joc_decode.py index 6216f70..bc716d1 100644 --- a/src/joc_decode.py +++ b/src/joc_decode.py @@ -2,11 +2,11 @@ 范围: - EMDF ID14 的 joc_header、joc_info 和 Huffman joc_data; - - 差分解码得到 joc_mix_mtx_q; + - 差分还原得到 joc_mix_mtx_q(dense MTX 与 sparse IDX/VEC 两条语法); - 去量化得到 joc_mix_mtx_dq; - 位流自洽验证(joc_data 后剩余 = padding_bits 0..7 + 可能 joc_ext_data) 后续的时间插值、QMF/时域重建和 ``joc_clipgain`` 位于 ``renderer.py``。 -Sparse 分支仍缺少实际样本验证。 +公式与符号定义见 ``docs/math.md`` 第 2 节。 """ from pathlib import Path @@ -23,6 +23,7 @@ _HUFF_NAMES = ( "joc_huff_code_7ch_pos_index_sparse", ) + def _load_huff_tables(): with np.load(_TABLES_PATH) as tables: return { @@ -30,18 +31,24 @@ def _load_huff_tables(): for name in _HUFF_NAMES } + H = _load_huff_tables() JOC_NUM_CHANNELS = {0: 5, 1: 7, 2: 7, 3: 5, 4: 7} # Table 33 JOC_NUM_BANDS = {0: 1, 1: 3, 2: 5, 3: 7, 4: 9, 5: 12, 6: 15, 7: 23} # Table 35 -_PUBLIC_TABLE39_EXCERPT_UNUSED = { # 仅保留作表格差异说明;渲染映射在 joc_qmf.py。 - 23: [0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22,22], -} +JOC_NUM_QUANT = {0: 96, 1: 192} # Table 51 +# dense 的量化零点就是 nquant/2;sparse 的递推起点比它高 2 个量化步。 +JOC_DENSE_OFFSET = {0: 48, 1: 96} +JOC_SPARSE_OFFSET = {0: 50, 1: 100} + class BR: + """MSB-first 位读取器;位置以载荷内的 bit offset 表示。""" + def __init__(self, data, pos=0): self.d = data self.p = pos + def bits(self, n): v = 0 for _ in range(n): @@ -53,18 +60,19 @@ class BR: def huff_decode(tree, br): node = 0 while node >= 0: - b = br.bits(1) - node = tree[node][b] + node = tree[node][br.bits(1)] return -node - 1 def get_huff_code(mode, typ, nch): if typ == "IDX": - return H["joc_huff_code_5ch_pos_index_sparse" if nch == 5 else "joc_huff_code_7ch_pos_index_sparse"] + return H["joc_huff_code_5ch_pos_index_sparse" if nch == 5 + else "joc_huff_code_7ch_pos_index_sparse"] if typ == "VEC": - return H["joc_huff_code_coarse_coeff_sparse" if mode == 0 else "joc_huff_code_fine_coeff_sparse"] - # MTX - return H["joc_huff_code_coarse_generic" if mode == 0 else "joc_huff_code_fine_generic"] + return H["joc_huff_code_coarse_coeff_sparse" if mode == 0 + else "joc_huff_code_fine_coeff_sparse"] + return H["joc_huff_code_coarse_generic" if mode == 0 + else "joc_huff_code_fine_generic"] # MTX def parse_joc(payload): @@ -76,6 +84,8 @@ def parse_joc(payload): out["ext_config_idx"] = br.bits(3) n_objects = out["num_objects_bits"] + 1 n_channels = JOC_NUM_CHANNELS.get(out["dmx_config_idx"]) + if n_channels is None: + raise ValueError(f"未知的 JOC downmix 配置 {out['dmx_config_idx']}") out["n_objects"], out["n_channels"] = n_objects, n_channels out["clipgain_x_bits"] = br.bits(3) out["clipgain_y_bits"] = br.bits(5) @@ -98,78 +108,120 @@ def parse_joc(payload): o["offset_ts"] = [br.bits(5) + 1 for _ in range(o["n_dpoints"])] objs.append(o) out["objs"] = objs - # joc_data(Huffman) + # joc_data(Huffman):dense 逐声道逐带读 MTX;sparse 读 IDX 后读 VEC。 for obj, o in enumerate(objs): if not o["present"]: continue - nquant = 96 if o["quant_idx"] == 0 else 192 o["channel_idx"] = [] o["vec"] = [] o["mtx"] = [] for dp in range(o["n_dpoints"]): if o["sparse"] == 1: - # Sparse JOC 使用 VEC/IDX Huffman 树;此分支尚无真实码流验证。 - ci0 = br.bits(3) - tree = get_huff_code(n_channels, "IDX", n_channels) - ci = [ci0] + [huff_decode(tree, br) for _ in range(o["n_bands"] - 1)] - o["channel_idx"].append(ci) + tree = get_huff_code(o["quant_idx"], "IDX", n_channels) + idx = [br.bits(3)] + idx += [huff_decode(tree, br) for _ in range(o["n_bands"] - 1)] tree = get_huff_code(o["quant_idx"], "VEC", n_channels) - vec = [huff_decode(tree, br) for _ in range(o["n_bands"])] - o["vec"].append(vec) + o["channel_idx"].append(idx) + o["vec"].append([huff_decode(tree, br) for _ in range(o["n_bands"])]) + o["mtx"].append(None) else: tree = get_huff_code(o["quant_idx"], "MTX", n_channels) mtx = [[huff_decode(tree, br) for _ in range(o["n_bands"])] for _ in range(n_channels)] o["mtx"].append(mtx) + o["channel_idx"].append(None) + o["vec"].append(None) out["data_end_bits"] = br.p out["remaining_bits"] = len(payload) * 8 - br.p out["tail_bytes"] = payload[br.p // 8:] return out +def reconstruct_dense(o, dp, n_ch): + """Dense 差分还原 → 量化矩阵 ``[ch][pb]``。 + + 每个核心声道各自从 ``nquant/2``(去量化 0)出发,沿参数带累加 MTX 符号。 + """ + nquant = JOC_NUM_QUANT[o["quant_idx"]] + offset = JOC_DENSE_OFFSET[o["quant_idx"]] + mtx = o["mtx"][dp] + if mtx is None: + raise ValueError("Dense JOC 对象缺少 MTX 符号") + q = np.zeros((n_ch, o["n_bands"]), dtype=np.int64) + for ch in range(n_ch): + q[ch][0] = (offset + mtx[ch][0]) % nquant + for pb in range(1, o["n_bands"]): + q[ch][pb] = (q[ch][pb - 1] + mtx[ch][pb]) % nquant + return q + + +def reconstruct_sparse(o, dp, n_ch): + """Sparse 差分还原 → 量化矩阵 ``[ch][pb]``。 + + 每个参数带只有一个 active 声道: + - ``active[0]`` 是 3 bit 绝对声道号,其后由 IDX 符号累加得到, + 因此递推锚点是**已重建**的 active 声道,而不是编码器送的符号本身; + - 系数是一个跨参数带连续的单累加器(active 声道切换时**不**重置), + 起点为 sparse offset,增量为 VEC 符号; + - 非 active 项取 ``nquant/2``,即去量化后的 0。 + + IDX 符号取值恒为 ``0..n_ch-1``(Huffman 叶数即声道数), + 故 ``(active + idx) % n_ch`` 与单次条件减等价。 + """ + if n_ch not in (5, 7): + raise ValueError(f"Sparse JOC 需要 5 或 7 个核心声道,实际 {n_ch}") + nquant = JOC_NUM_QUANT[o["quant_idx"]] + offset = JOC_SPARSE_OFFSET[o["quant_idx"]] + n_bands = o["n_bands"] + idx = o["channel_idx"][dp] + vec = o["vec"][dp] + if idx is None or vec is None: + raise ValueError("Sparse JOC 对象缺少 channel_idx/vec 符号") + if len(idx) != n_bands or len(vec) != n_bands: + raise ValueError( + f"Sparse JOC 维度不符: bands={n_bands}, idx={len(idx)}, vec={len(vec)}") + if not 0 <= idx[0] < n_ch: + raise ValueError(f"Sparse JOC 初始声道 {idx[0]} 超出 {n_ch} 个声道") + + q = np.full((n_ch, n_bands), nquant // 2, dtype=np.int64) + active = idx[0] + coefficient = offset + for pb in range(n_bands): + if pb: + active = (active + idx[pb]) % n_ch + coefficient = (coefficient + vec[pb]) % nquant + q[active][pb] = coefficient + return q + + def diff_decode(out): - """6.6.2:差分解码 → joc_mix_mtx_q[obj][dp][ch][pb]。""" + """差分还原 → joc_mix_mtx_q[obj][dp][ch][pb]。""" mix_q = {} n_ch = out["n_channels"] for obj, o in enumerate(out["objs"]): if not o["present"]: continue - nquant = 96 if o["quant_idx"] == 0 else 192 q = np.zeros((o["n_dpoints"], n_ch, o["n_bands"]), dtype=np.int64) for dp in range(o["n_dpoints"]): if o["sparse"] == 1: - # Sparse 差分路径尚无真实码流验证。 - offset = 50 if o["quant_idx"] == 0 else 100 - ci = o["channel_idx"][dp] - vec = o["vec"][dp] - for pb in range(o["n_bands"]): - ci_mod = ci[0] if pb == 0 else (ci[pb - 1] + ci[pb]) % n_ch - for ch in range(n_ch): - if ch == ci_mod: - if pb == 0: - q[dp][ch][pb] = (offset + vec[pb]) % nquant - else: - q[dp][ch][pb] = (q[dp][ch][pb - 1] + vec[pb]) % nquant - else: - q[dp][ch][pb] = offset + q[dp] = reconstruct_sparse(o, dp, n_ch) else: - offset = 48 if o["quant_idx"] == 0 else 96 - mtx = o["mtx"][dp] - for ch in range(n_ch): - q[dp][ch][0] = (offset + mtx[ch][0]) % nquant - for pb in range(1, o["n_bands"]): - q[dp][ch][pb] = (q[dp][ch][pb - 1] + mtx[ch][pb]) % nquant + q[dp] = reconstruct_dense(o, dp, n_ch) mix_q[obj] = q return mix_q def dequantize(out, mix_q): - """6.6.4:去量化 → joc_mix_mtx_dq。""" + """去量化 → joc_mix_mtx_dq。 + + Sparse 的非 active 项在 ``joc_mix_mtx_q`` 中取 ``nquant/2``, + 因此与 dense 共用同一条去量化公式即得到 0。 + """ mix_dq = {} for obj, o in enumerate(out["objs"]): if not o["present"]: continue - nquant = 96 if o["quant_idx"] == 0 else 192 + nquant = JOC_NUM_QUANT[o["quant_idx"]] q = mix_q[obj] dq = (q.astype(np.float64) - nquant / 2) * 820 / (4096 * (1 + o["quant_idx"])) mix_dq[obj] = dq diff --git a/src/metadata.py b/src/metadata.py index 534f077..996af5d 100644 --- a/src/metadata.py +++ b/src/metadata.py @@ -246,20 +246,6 @@ def inspect(index, limit=None, print_frames=False): "parser_error": str(exc), "repair_hint": "检查 JOC header、对象数、参数带、Huffman 或扩展字段", }) from exc - sparse = [i for i, obj in enumerate(parsed["objs"]) if obj["present"] and obj["sparse"]] - if sparse: - raise UnsupportedVariantError( - "joc", "sparse_joc", - "发现尚未验证的 Sparse JOC 帧", - frame=frame_number, - details={ - "sparse_objects": sparse, - "downmix_config": parsed["dmx_config_idx"], - "extension_config": parsed["ext_config_idx"], - "objects": parsed["n_objects"], - "payload": bytes_descriptor(subs[14]), - "repair_hint": "需要 Sparse JOC 实际样本及对应输出建立回归后再启用", - }) if parsed["n_channels"] != 5 or parsed["n_objects"] > 15: raise UnsupportedVariantError( "joc", "unsupported_configuration", diff --git a/src/native_renderer.py b/src/native_renderer.py index c964831..1ba013b 100644 --- a/src/native_renderer.py +++ b/src/native_renderer.py @@ -208,8 +208,6 @@ class NativeJocRenderer: for object_index, info in enumerate(out["objs"]): if not info["present"]: continue - if info["sparse"]: - raise ValueError("native core does not accept unvalidated Sparse JOC") bands = int(info["n_bands"]) points = int(info["n_dpoints"]) if bands > MAX_BANDS or points > MAX_DPOINTS: