Align OAMD metadata delay with decoder output
Native builds / linux-x64 (push) Failing after 15s
Native builds / macos-arm64 (push) Has been cancelled
Native builds / macos-x64 (push) Has been cancelled
Native builds / windows-x64 (push) Has been cancelled
Native builds / Publish GitHub Release (push) Has been cancelled
Native builds / linux-x64 (push) Failing after 15s
Native builds / macos-arm64 (push) Has been cancelled
Native builds / macos-x64 (push) Has been cancelled
Native builds / windows-x64 (push) Has been cancelled
Native builds / Publish GitHub Release (push) Has been cancelled
This commit is contained in:
@@ -95,6 +95,16 @@ python main.py input.m4a --metadata-cache metadata_cache
|
|||||||
python main.py input.m4a --metadata-dir metadata_cache
|
python main.py input.m4a --metadata-dir metadata_cache
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### OAMD time alignment
|
||||||
|
|
||||||
|
Object trajectories and direct speaker rendering both default to a metadata delay of `1473 samples`. This value describes the theoretical mapping between decoder-output PCM and OAMD updates. The speaker renderer retains its DRP-compatible 32-sample control block, so the default update lands on effective block boundary `1472`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
align32(1473) = 1472
|
||||||
|
```
|
||||||
|
|
||||||
|
Override the two paths with `--object-delay-samples` and `--speaker-metadata-offset`, respectively. The 1473-sample timing offset is distinct from the 640-value inverse-QMF filter/window state; 640 is a QMF state length, not a metadata delay.
|
||||||
|
|
||||||
For all options:
|
For all options:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
|
|||||||
@@ -95,6 +95,16 @@ python main.py input.m4a --metadata-cache metadata_cache
|
|||||||
python main.py input.m4a --metadata-dir metadata_cache
|
python main.py input.m4a --metadata-dir metadata_cache
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### OAMD 时间对齐
|
||||||
|
|
||||||
|
对象轨迹和直接扬声器渲染的 metadata delay 默认均为 `1473 samples`。该值描述 decoder 输出 PCM 与 OAMD 更新之间的理论时间映射;扬声器 renderer 仍使用 DRP-compatible 的 32-sample control block,因此默认更新的实际 block boundary 为 `1472`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
align32(1473) = 1472
|
||||||
|
```
|
||||||
|
|
||||||
|
可分别用 `--object-delay-samples` 和 `--speaker-metadata-offset` 覆盖默认值。这里的 1473 不应与 inverse-QMF 的 640 项 filter/window state 混淆;后者是 QMF 状态长度,不是 metadata delay。
|
||||||
|
|
||||||
更多参数可查看:
|
更多参数可查看:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
|
|||||||
+12
-2
@@ -499,16 +499,26 @@ s_{\mathrm{frame}}
|
|||||||
+32f_{\mathrm{block}}.
|
+32f_{\mathrm{block}}.
|
||||||
$$
|
$$
|
||||||
|
|
||||||
For processing-block length $B=32$, the aligned update point is
|
The theoretical update position on the decoder-output PCM timeline is
|
||||||
|
|
||||||
|
$$
|
||||||
|
s_{\mathrm{theoretical}}
|
||||||
|
=s_{\mathrm{coded}}+d_{\mathrm{decoder}},
|
||||||
|
\qquad d_{\mathrm{decoder}}=1473.
|
||||||
|
$$
|
||||||
|
|
||||||
|
The speaker renderer retains the DRP-compatible processing-block length $B=32$, so the aligned update point is
|
||||||
|
|
||||||
$$
|
$$
|
||||||
\widehat s
|
\widehat s
|
||||||
=
|
=
|
||||||
B\left\lfloor
|
B\left\lfloor
|
||||||
\frac{s_{\mathrm{coded}}+B/2-1}{B}
|
\frac{s_{\mathrm{theoretical}}+B/2-1}{B}
|
||||||
\right\rfloor.
|
\right\rfloor.
|
||||||
$$
|
$$
|
||||||
|
|
||||||
|
Thus, for frame-aligned updates, `align32(1473)=1472`. The 1473 value is the theoretical decoder delay at the metadata interface; 1472 is its effective boundary in the current 32-sample control block. The 640-value inverse-QMF window/state is not part of this metadata-timing formula.
|
||||||
|
|
||||||
For ramp duration $D$, the number of blocks is
|
For ramp duration $D$, the number of blocks is
|
||||||
|
|
||||||
$$
|
$$
|
||||||
|
|||||||
+12
-2
@@ -499,16 +499,26 @@ s_{\mathrm{frame}}
|
|||||||
+32f_{\mathrm{block}}.
|
+32f_{\mathrm{block}}.
|
||||||
$$
|
$$
|
||||||
|
|
||||||
对处理块长度 $B=32$,更新点对齐为
|
decoder 输出 PCM timeline 上的理论更新位置为
|
||||||
|
|
||||||
|
$$
|
||||||
|
s_{\mathrm{theoretical}}
|
||||||
|
=s_{\mathrm{coded}}+d_{\mathrm{decoder}},
|
||||||
|
\qquad d_{\mathrm{decoder}}=1473.
|
||||||
|
$$
|
||||||
|
|
||||||
|
扬声器 renderer 保留 DRP-compatible 的处理块长度 $B=32$,更新点对齐为
|
||||||
|
|
||||||
$$
|
$$
|
||||||
\widehat s
|
\widehat s
|
||||||
=
|
=
|
||||||
B\left\lfloor
|
B\left\lfloor
|
||||||
\frac{s_{\mathrm{coded}}+B/2-1}{B}
|
\frac{s_{\mathrm{theoretical}}+B/2-1}{B}
|
||||||
\right\rfloor.
|
\right\rfloor.
|
||||||
$$
|
$$
|
||||||
|
|
||||||
|
因此,对 frame-aligned 更新有 `align32(1473)=1472`。1473 是 metadata interface 的理论 decoder delay;1472 是当前 32-sample control block 中的有效边界。inverse-QMF 使用的 640 项 window/state 不属于这条 metadata timing 公式。
|
||||||
|
|
||||||
给定 ramp duration $D$,block 数为
|
给定 ramp duration $D$,block 数为
|
||||||
|
|
||||||
$$
|
$$
|
||||||
|
|||||||
@@ -286,8 +286,8 @@ def build_parser():
|
|||||||
parser.add_argument("--gain-db", type=float, default=0.0,
|
parser.add_argument("--gain-db", type=float, default=0.0,
|
||||||
help="成品增益 dB,默认 0(float32 系数 1.0)")
|
help="成品增益 dB,默认 0(float32 系数 1.0)")
|
||||||
parser.add_argument("--duration", type=float, help="只处理开头指定秒数")
|
parser.add_argument("--duration", type=float, help="只处理开头指定秒数")
|
||||||
parser.add_argument("--object-delay-samples", type=int, default=640,
|
parser.add_argument("--object-delay-samples", type=int, default=1473,
|
||||||
help="可选的对象 PCM/OAMD 时间补偿,默认 640 samples")
|
help="可选的对象 PCM/OAMD 时间补偿,默认 1473 samples")
|
||||||
parser.add_argument("--trajectory-mode", choices=("compact", "dense64"), default="compact",
|
parser.add_argument("--trajectory-mode", choices=("compact", "dense64"), default="compact",
|
||||||
help="对象轨迹表示;compact 用长线性插值压缩 AXML,dense64 保留逐 64-sample 块")
|
help="对象轨迹表示;compact 用长线性插值压缩 AXML,dense64 保留逐 64-sample 块")
|
||||||
parser.add_argument("--ffmpeg", default=os.environ.get("FFMPEG", "ffmpeg"))
|
parser.add_argument("--ffmpeg", default=os.environ.get("FFMPEG", "ffmpeg"))
|
||||||
|
|||||||
+3
-2
@@ -179,7 +179,7 @@ def _expand_events(events, total_samples, rate, update_quantum_samples,
|
|||||||
raise ValueError(f"未知 trajectory_mode: {trajectory_mode}")
|
raise ValueError(f"未知 trajectory_mode: {trajectory_mode}")
|
||||||
|
|
||||||
def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536,
|
def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536,
|
||||||
update_quantum_samples=64, object_delay_samples=640,
|
update_quantum_samples=64, object_delay_samples=1473,
|
||||||
trajectory_mode="compact"):
|
trajectory_mode="compact"):
|
||||||
"""从统一 metadata index 构造 15 条 ADM 轨迹。
|
"""从统一 metadata index 构造 15 条 ADM 轨迹。
|
||||||
|
|
||||||
@@ -187,7 +187,8 @@ def build_adm_tracks(index, frames=None, rate=48000, frame_samples=1536,
|
|||||||
OAMD 的内外层 sample offset、block offset 和 ramp 均保留。
|
OAMD 的内外层 sample offset、block offset 和 ramp 均保留。
|
||||||
``trajectory_mode="compact"`` 用一个长 ADM interpolation block 表示每条
|
``trajectory_mode="compact"`` 用一个长 ADM interpolation block 表示每条
|
||||||
线性 ramp;``dense64`` 保留逐 64-sample 展开作为兼容回退。
|
线性 ramp;``dense64`` 保留逐 64-sample 展开作为兼容回退。
|
||||||
``object_delay_samples`` 将位置更新与对象逆 QMF 的输出时刻对齐。
|
``object_delay_samples`` 将位置更新与对象 PCM 的 decoder 输出时刻对齐;
|
||||||
|
默认 1473 与 DRP decoder/renderer metadata interface 保持一致。
|
||||||
slot1..15 与对象 PCM ch1..15 一一对应。
|
slot1..15 与对象 PCM ch1..15 一一对应。
|
||||||
"""
|
"""
|
||||||
frames = index.rows if frames is None else frames
|
frames = index.rows if frames is None else frames
|
||||||
|
|||||||
Reference in New Issue
Block a user