Align OAMD metadata delay with decoder output
Native builds / linux-x64 (push) Failing after 12s
Native builds / macos-arm64 (push) Has been cancelled
Native builds / macos-x64 (push) Has been cancelled
Native builds / windows-x64 (push) Has been cancelled
Native builds / Publish GitHub Release (push) Has been cancelled

This commit is contained in:
TheM14
2026-09-02 22:17:53 +08:00
parent 98b419a7dc
commit b115283d84
6 changed files with 48 additions and 8 deletions
+10
View File
@@ -95,6 +95,16 @@ python main.py input.m4a --metadata-cache metadata_cache
python main.py input.m4a --metadata-dir metadata_cache python main.py input.m4a --metadata-dir metadata_cache
``` ```
### OAMD time alignment
Object trajectories and direct speaker rendering both default to a metadata delay of `1473 samples`. This value describes the theoretical mapping between decoder-output PCM and OAMD updates. The speaker renderer retains its existing 32-sample control block, so the default update lands on effective block boundary `1472`:
```text
align32(1473) = 1472
```
Override the two paths with `--object-delay-samples` and `--speaker-metadata-offset`, respectively. The 1473-sample timing offset is distinct from the 640-value inverse-QMF filter/window state; 640 is a QMF state length, not a metadata delay.
For all options: For all options:
```powershell ```powershell
+10
View File
@@ -95,6 +95,16 @@ python main.py input.m4a --metadata-cache metadata_cache
python main.py input.m4a --metadata-dir metadata_cache python main.py input.m4a --metadata-dir metadata_cache
``` ```
### OAMD 时间对齐
对象轨迹和直接扬声器渲染的 metadata delay 默认均为 `1473 samples`。该值描述 decoder 输出 PCM 与 OAMD 更新之间的理论时间映射;扬声器 renderer 仍使用现有的 32-sample control block,因此默认更新的实际 block boundary 为 `1472`:
```text
align32(1473) = 1472
```
可分别用 `--object-delay-samples` 和 `--speaker-metadata-offset` 覆盖默认值。这里的 1473 不应与 inverse-QMF 的 640 项 filter/window state 混淆;后者是 QMF 状态长度,不是 metadata delay。
更多参数可查看: 更多参数可查看:
```powershell ```powershell
+12 -2
View File
@@ -499,16 +499,26 @@ s_{\mathrm{frame}}
+32f_{\mathrm{block}}. +32f_{\mathrm{block}}.
$$ $$
For processing-block length $B=32$, the aligned update point is The theoretical update position on the decoder-output PCM timeline is
$$
s_{\mathrm{theoretical}}
=s_{\mathrm{coded}}+d_{\mathrm{decoder}},
\qquad d_{\mathrm{decoder}}=1473.
$$
The speaker renderer retains the existing processing-block length $B=32$, so the aligned update point is
$$ $$
\widehat s \widehat s
= =
B\left\lfloor B\left\lfloor
\frac{s_{\mathrm{coded}}+B/2-1}{B} \frac{s_{\mathrm{theoretical}}+B/2-1}{B}
\right\rfloor. \right\rfloor.
$$ $$
Thus, for frame-aligned updates, `align32(1473)=1472`. The 1473 value is the theoretical decoder delay at the metadata interface; 1472 is its effective boundary in the current 32-sample control block. The 640-value inverse-QMF window/state is not part of this metadata-timing formula.
For ramp duration $D$, the number of blocks is For ramp duration $D$, the number of blocks is
$$ $$
+12 -2
View File
@@ -499,16 +499,26 @@ s_{\mathrm{frame}}
+32f_{\mathrm{block}}. +32f_{\mathrm{block}}.
$$ $$
对处理块长度 $B=32$,更新点对齐为 decoder 输出 PCM timeline 上的理论更新位置为
$$
s_{\mathrm{theoretical}}
=s_{\mathrm{coded}}+d_{\mathrm{decoder}},
\qquad d_{\mathrm{decoder}}=1473.
$$
扬声器 renderer 保留现有的处理块长度 $B=32$,更新点对齐为
$$ $$
\widehat s \widehat s
= =
B\left\lfloor B\left\lfloor
\frac{s_{\mathrm{coded}}+B/2-1}{B} \frac{s_{\mathrm{theoretical}}+B/2-1}{B}
\right\rfloor. \right\rfloor.
$$ $$
因此,对 frame-aligned 更新有 `align32(1473)=1472`。1473 是 metadata interface 的理论 decoder delay;1472 是当前 32-sample control block 中的有效边界。inverse-QMF 使用的 640 项 window/state 不属于这条 metadata timing 公式。
给定 ramp duration $D$,block 数为 给定 ramp duration $D$,block 数为
$$ $$
+2 -2
View File
@@ -286,8 +286,8 @@ def build_parser():
parser.add_argument("--gain-db", type=float, default=0.0, parser.add_argument("--gain-db", type=float, default=0.0,
help="成品增益 dB,默认 0(float32 系数 1.0)") help="成品增益 dB,默认 0(float32 系数 1.0)")
parser.add_argument("--duration", type=float, help="只处理开头指定秒数") parser.add_argument("--duration", type=float, help="只处理开头指定秒数")
parser.add_argument("--object-delay-samples", type=int, default=640, parser.add_argument("--object-delay-samples", type=int, default=1473,
help="可选的对象 PCM/OAMD 时间补偿,默认 640 samples") help="可选的对象 PCM/OAMD 时间补偿,默认 1473 samples")
parser.add_argument("--trajectory-mode", choices=("compact", "dense64"), default="compact", parser.add_argument("--trajectory-mode", choices=("compact", "dense64"), default="compact",
help="对象轨迹表示;compact 用长线性插值压缩 AXML,dense64 保留逐 64-sample 块") help="对象轨迹表示;compact 用长线性插值压缩 AXML,dense64 保留逐 64-sample 块")
parser.add_argument("--ffmpeg", default=os.environ.get("FFMPEG", "ffmpeg")) parser.add_argument("--ffmpeg", default=os.environ.get("FFMPEG", "ffmpeg"))
+2 -2
View File
@@ -179,7 +179,7 @@ def _expand_events(events, total_samples, rate, update_quantum_samples,
raise ValueError(f"未知 trajectory_mode: {trajectory_mode}") raise ValueError(f"未知 trajectory_mode: {trajectory_mode}")
def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536, def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536,
update_quantum_samples=64, object_delay_samples=640, update_quantum_samples=64, object_delay_samples=1473,
trajectory_mode="compact"): trajectory_mode="compact"):
"""从统一 metadata index 构造 15 条 ADM 轨迹。 """从统一 metadata index 构造 15 条 ADM 轨迹。
@@ -187,7 +187,7 @@ def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536,
OAMD 的内外层 sample offset、block offset 和 ramp 均保留。 OAMD 的内外层 sample offset、block offset 和 ramp 均保留。
``trajectory_mode="compact"`` 用一个长 ADM interpolation block 表示每条 ``trajectory_mode="compact"`` 用一个长 ADM interpolation block 表示每条
线性 ramp;``dense64`` 保留逐 64-sample 展开作为兼容回退。 线性 ramp;``dense64`` 保留逐 64-sample 展开作为兼容回退。
``object_delay_samples`` 将位置更新与对象逆 QMF 的输出时刻对齐。 ``object_delay_samples`` 将位置更新与对象 PCM 的 decoder 输出时刻对齐。
slot1..15 与对象 PCM ch1..15 一一对应。 slot1..15 与对象 PCM ch1..15 一一对应。
""" """
frames = index.rows if frames is None else frames frames = index.rows if frames is None else frames