diff --git a/.gitignore b/.gitignore index 85d4c19..7173bf1 100644 --- a/.gitignore +++ b/.gitignore @@ -29,3 +29,6 @@ rosella_kernels.npz # The component writes its own log next to the DLL joc_decoder.log + +# Measurement record: local working notes, kept on disk and never published. +VERIFICATION.md diff --git a/README.en.md b/README.en.md new file mode 100644 index 0000000..bdf2f1a --- /dev/null +++ b/README.en.md @@ -0,0 +1,121 @@ +# foo_input_joc — foobar2000 input component for E-AC-3 JOC + +[简体中文](README.md) · [Install](#install) · [Settings](#settings) · [Building](#building) · [Known limitations](#known-limitations) + +A foobar2000 input component for **E-AC-3 JOC (Dolby Atmos)** files: the JOC object and OAMD +metadata is taken from the E-AC-3 syncframes and paired with the 5.1 core PCM that ffmpeg +decodes, then rendered in real time to headphones (HRTF) or to a speaker layout up to 7.1. The +rendering core is compiled into the component, so there is nothing to install beside +`foo_input_joc.dll`. + +Input is a bare `.eac3` / `.ec3` stream, or an E-AC-3 JOC track inside a container (`.mp4` +`.m4a` `.m4b` `.m4p` `.m4r` `.mov` `.mkv` `.mka` `.webm`). A file without JOC is not claimed: +plain AC-3 / E-AC-3, a container whose audio track is AAC, and transport streams are all handed +back to foobar2000 and play through the decoder it would have used anyway. + +## Requirements + +* foobar2000 1.6 (32-bit) or 2.x (32-bit and 64-bit). +* An `ffmpeg` executable, found on `PATH` by default; the preferences page can name one. +* For binaural output, an HRTF data file. **This repository does not ship one** — see + [HRTF data](#hrtf-data). + +## Install + +Download the package for your architecture from [Releases](../../releases) — a pushed `v*` tag +makes CI attach both — or build it yourself as described under [Building](#building) and take the +result from `dist\`; either way the file is `foo_input_joc--.fb2k-component`. Drop +it onto foobar2000, or use Preferences → Components → Install, and restart. Take `-x86` for +foobar2000 1.6 and 2.x 32-bit, `-x64` for 2.x 64-bit. + +To install by hand, copy `foo_input_joc.dll` into the per-component subdirectory of the +`user-components` folder the running version reads: + +| foobar2000 | Folder | +|---|---| +| 1.6 | `\user-components\foo_input_joc\` | +| 2.x | `\user-components\foo_input_joc\` (also in portable mode) | + +The subdirectory is required — a DLL lying directly in `user-components\` is not scanned — and a +component placed in the folder the running version does not read, or built for the other +architecture, is **ignored without a message**. + +`tools/deploy.ps1 -TestBed ` is the scripted form of the manual steps, and +`tools/run.ps1 -TestBed -Play ` plays a file unattended and prints the log. Let +foobar2000 exit through `/exit`: a force-killed instance leaves a `\running` marker +behind, and the next start then refuses to load any user component. + +### Containers need one look at the decoder list + +foobar2000 asks the decoders in the order shown in Preferences → **Decoding**, and the built-in +container readers are in that list. When one of them is offered an MP4 or Matroska file first, it +takes the file and the JOC objects are lost — such a file then plays as plain E-AC-3. + +So move **JOC decoder (E-AC-3 JOC)** above **foobar2000 MP4 Demuxer** and **foobar2000 +Matroska/WebM Reader** in that list. Bare `.eac3` / `.ec3` files do not depend on the order. + +## Settings + +Preferences → Tools → **JOC decoder**: + +* **Output** — binaural, or a speaker layout from 2.0 to 7.1; +* **Binaural mode** (near / mid / far) and the room **tail** in seconds; +* **HRTF source** — a **SOFA** file or a **Rosella** `.personalized_headphone` model. An empty + path means the default location, `\HRTF\binaural.sofa` or + `binaural.personalized_headphone`; +* **Gain** — a switch and a value in dB. Binaural rendering can exceed full scale on material + that does not clip in the core mix, so attenuation belongs here; +* the **ffmpeg** executable to use. + +No control on the page is disabled, and the status line states what is in effect. + +### HRTF data + +A SOFA measurement set or a personalised headphone model is supplied by whoever runs the +component, and is listed in `.gitignore` so it cannot be committed by accident. Speaker layouts +need none. Binaural rendering without an HRTF fails with a message naming the file it looked for. + +## Building + +```powershell +pwsh -File tools/setup_sdk.ps1 # official SDK into SDK/, pinned to target 1.5/1.6 +pwsh -File tools/build.ps1 # Win32 -> build\Win32\foo_input_joc.dll +pwsh -File tools/build.ps1 -Platform x64 +pwsh -File tools/package.ps1 # both, packaged into dist\*.fb2k-component +``` + +The configuration is fixed at `Release-Static` (static CRT, `/MT`), and `/fp:precise` is what the +byte-for-byte acceptance rests on, so it must not be changed. `foo_input_joc.vcxproj` builds +`kernel\joc_kernel.vcxproj` first through a project reference; the rendering core sources in +`kernel/` are compiled with `JOC_STATIC` / `EJOC_STATIC`, so their entry points are neither +imported nor exported. `tests\` holds the offline tools (bitstream self-test and cross-check, +render comparison, preferences-page layout check, container-probe check) and `tools\` the build +and test-bed scripts. Settings also read `JOC_*` environment overrides for one run (development +only); the list and what each one does is in `src\settings.cpp`. + +To diagnose a problem, read `joc_decoder.log` beside the DLL: the component writes its own +version, the core version and the log path there at start-up. + +## Known limitations + +* ADM BWF output is not implemented. +* A container is only claimed when this component is ahead of the built-in container reader in + Preferences → Decoding (see [Install](#containers-need-one-look-at-the-decoder-list)); the core + does not let a decoder ask for a file another entry has already taken. Such a file's tags stay + that reader's as well. +* Tags of a bare `.eac3` / `.ec3` (an ID3v2 tag in front of the stream, or an APEv2/ID3v1 tag + behind it) can be **read but not written**: no component claims raw E-AC-3 for writing, and + rewriting the whole file to insert a tag is not this component's job. +* Transport streams (`.ts`, `.m2ts`) are not claimed. +* Playback length is exactly the file's duration. The binaural renderer still computes its room + tail, but it is not delivered as playback time the file does not have. +* A seek re-enters the bitstream at the frame holding the target instead of decoding everything + in front of it — which is why the cost of a seek does not depend on where it lands. The + position is exact and does not drift; the samples are the same waveform handed to a decoder + that started there, so they differ from a straight play-through by a low-level noise floor of + −59 dBFS or quieter. + +## Licence + +`LICENSE` is the upstream MIT licence, copied unchanged; `kernel/` is a copy of the upstream +rendering core sources and keeps their notices. See [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md). diff --git a/README.md b/README.md index 6fe48b5..3500825 100644 --- a/README.md +++ b/README.md @@ -1,143 +1,103 @@ -# foo_input_joc +# foo_input_joc — foobar2000 的 E-AC-3 JOC 输入组件 -foobar2000 input component for **E-AC-3 JOC (Dolby Atmos)** files: the JOC objects are -rendered to binaural (HRTF) or to a speaker layout up to 7.1, in real time. +[English](README.en.md) · [安装](#安装) · [设置](#设置) · [构建](#构建) · [已知限制](#已知限制) -Two files with the same name pay for the whole thing: `joc_core`'s C++ sources are copied -into [`kernel/`](kernel/) and compiled straight into the component, so there is nothing to -install beside `foo_input_joc.dll`. +foobar2000 输入组件,播放 **E-AC-3 JOC(Dolby Atmos)** 文件:从 E-AC-3 同步帧里取出 JOC +对象与 OAMD 元数据,与 ffmpeg 解出的 5.1 核心 PCM 配对后实时渲染,输出双耳(HRTF)或最多 +7.1 的扬声器布局。渲染内核直接编译进组件,除 `foo_input_joc.dll` 之外不需要安装任何东西。 -## What it does +输入是裸流 `.eac3` / `.ec3`,或容器里的 E-AC-3 JOC 轨道(`.mp4` `.m4a` `.m4b` `.m4p` +`.m4r` `.mov` `.mkv` `.mka` `.webm`)。文件里没有 JOC 就不接管:普通 AC-3 / E-AC-3、音频轨道 +是 AAC 的容器、传输流都交还 foobar2000,由它原本的解码器播放。 -1. Finds the audio: a bare `.eac3` / `.ec3` stream is read as it is, while a container - (`.mp4`, `.m4a`, `.m4b`, `.m4p`, `.m4r`, `.mov`, `.mkv`, `.mka`, `.webm`) is looked into - first — the container's own headers say whether an E-AC-3 track is present and which - audio track it is (MP4 sample entry `ec-3`, Matroska `CodecID A_EAC3`), and ffmpeg then - copies that track out of the file byte for byte. The header walk is bounded and cheap, so - an MP4 holding AAC is declined without starting anything. -2. Decides from the bitstream whether it really carries JOC (an EMDF container holding both - the OAMD and the JOC payload — container metadata only ever says "E-AC-3", and the JOC - flag inside it is frequently missing). -3. A file with no E-AC-3 track, or with one that carries no JOC, is handed back to - foobar2000 with `exception_io_unsupported_format`, so the built-in decoder plays it — this - component never decodes plain AC-3 or E-AC-3. -4. A JOC file is decoded as: the syncframes go to the renderer as metadata, the 5.1 core - PCM comes from ffmpeg, and the renderer pairs them (one syncframe : 1536 bed samples) - and produces the output PCM, which is handed back to foobar2000. +## 环境要求 -``` -.eac3 / .ec3 file container (.mp4 .mkv .m4a ...) - │ ├─ header walk (src/container_scan.cpp) ── no E-AC-3 ──▶ next decoder - │ └─ E-AC-3 track ── ffmpeg -c:a copy ──▶ syncframes - ├─ JOC check (src/eac3_scan.cpp) ─── no JOC ──▶ built-in E-AC-3 decoder - └─ JOC - ├─ syncframes ────────────────────▶ renderer metadata - └─ ffmpeg -ac 6 -c:a pcm_f32le ───▶ 5.1 core PCM ──▶ renderer bed - │ - ▼ - 2 ch or ≤7.1 PCM ──▶ foobar2000 -``` +* foobar2000 1.6(32 位)或 2.x(32 位与 64 位)。 +* 一个 `ffmpeg` 可执行文件,默认从 `PATH` 找,设置页可以指定路径。 +* 双耳输出需要一份 HRTF 数据文件。**本仓库不附带**,见 [HRTF 数据](#hrtf-数据)。 -## Repository layout +## 安装 -| Path | Contents | +从 [Releases](../../releases) 下载对应架构的包(`v*` 标签触发的 CI 会把两个架构都附上去), +或按[构建](#构建)自行打出 `dist\` 下的产物;文件名都是 +`foo_input_joc-<版本>-<架构>.fb2k-component`。把它拖到 foobar2000 上,或用 +Preferences → Components → Install 安装,然后重启。1.6 和 2.x 32 位用 `-x86`,2.x 64 位用 +`-x64`。 + +手工安装就把 `foo_input_joc.dll` 放进当前版本会读取的 `user-components` 子目录: + +| foobar2000 | 目录 | |---|---| -| `kernel/` | Copy of the `joc_core` C++ sources (`include/` + `src/`) and `joc_kernel.vcxproj`, the static library the component links | -| `src/eac3_scan.*` | Syncframe walk and the JOC bitstream test | -| `src/container_scan.*` | Bounded header walk of MP4/MOV and Matroska: is there an E-AC-3 track, which one, and how long is the file | -| `src/joc_decode.*` | Decode engine: starts ffmpeg, drives the renderer, handles the end of stream. No foobar2000 headers, so it also builds into the offline tools | -| `src/input_joc.cpp` | The foobar2000 input: format recognition, yielding, `get_info`, `initialize`, `run` | -| `src/settings.*` | Configuration values and their environment overrides (development only) | -| `src/prefs.cpp`, `src/prefs.rc` | The preferences page | -| `src/log.*` | Diagnostic log written next to the DLL | -| `tests/` | Offline tools: bitstream self-test and cross-check against the renderer, render harness, preferences-page layout check, container-probe check | -| `tools/` | SDK fetch, build, package, deploy, unattended test bed run | +| 1.6 | `\user-components\foo_input_joc\` | +| 2.x | `\user-components\foo_input_joc\`(portable 模式同样如此) | -## Build +子目录是必须的——DLL 直接躺在 `user-components\` 下不会被扫描;放进当前版本不读的目录,或 +架构不对,都会**静默忽略**,不会有任何提示。 + +`tools/deploy.ps1 -TestBed ` 是上面手工步骤的脚本版; +`tools/run.ps1 -TestBed <路径> -Play <文件>` 可以无人值守播放并把日志打出来。让 foobar2000 +通过 `/exit` 正常退出:被强杀的实例会在 `\running` 留下标记,下次启动会拒绝加载任何 +用户组件。 + +### 容器需要调一次解码器顺序 + +foobar2000 按 Preferences → **Decoding** 里的顺序询问解码器,内置的容器读取器也在那张表里。 +如果它排在前面接到 MP4 / Matroska 文件,文件就被它拿走,JOC 对象随之丢失——于是听起来只是 +普通 E-AC-3。 + +所以要把 **JOC decoder (E-AC-3 JOC)** 提到 **foobar2000 MP4 Demuxer** 与 **foobar2000 +Matroska/WebM Reader** 之前。裸流 `.eac3` / `.ec3` 不受这个顺序影响。 + +## 设置 + +Preferences → Tools → **JOC decoder**: + +* **Output** —— 双耳,或 2.0 到 7.1 的扬声器布局; +* **Binaural mode**(near / mid / far)与房间 **tail** 秒数; +* **HRTF source** —— **SOFA** 文件或 **Rosella** `.personalized_headphone` 模型。路径留空表示 + 用默认位置 `<组件目录>\HRTF\` 下的 `binaural.sofa` 或 `binaural.personalized_headphone`; +* **Gain** —— 开关加 dB 值。双耳渲染在核心混音不削顶的素材上也可能超过满刻度,衰减放在这里; +* **ffmpeg** 可执行文件路径。 + +页面上的控件都不禁用,状态行会说明当前生效的是什么。 + +### HRTF 数据 + +SOFA 测量集或个性化耳机模型由使用组件的人自己提供,并且写在 `.gitignore` 里,避免误提交。 +扬声器布局不需要 HRTF。双耳渲染缺少 HRTF 时会报错并指出它找的是哪个文件。 + +## 构建 ```powershell -pwsh -File tools/setup_sdk.ps1 # official SDK into SDK/, pinned to target 1.5/1.6 +pwsh -File tools/setup_sdk.ps1 # 官方 SDK 拉进 SDK/,固定到 target 1.5/1.6 pwsh -File tools/build.ps1 # Win32 -> build\Win32\foo_input_joc.dll pwsh -File tools/build.ps1 -Platform x64 -pwsh -File tools/package.ps1 # both, packaged into dist\*.fb2k-component +pwsh -File tools/package.ps1 # 两个架构,打包到 dist\*.fb2k-component ``` -`Release-Static` uses the static CRT (`/MT`); `/fp:precise` is required and must not be -changed. `foo_input_joc.vcxproj` builds `kernel\joc_kernel.vcxproj` first through a project -reference. The copied kernel sources are compiled with `JOC_STATIC` / `EJOC_STATIC` so their -entry points are neither imported nor exported. +配置固定为 `Release-Static`(静态 CRT,`/MT`);`/fp:precise` 是逐字节验收的前提,不要改。 +`foo_input_joc.vcxproj` 通过项目引用先构建 `kernel\joc_kernel.vcxproj`;`kernel/` 里的渲染内核 +源码以 `JOC_STATIC` / `EJOC_STATIC` 编译,入口既不导入也不导出。`tests\` 是离线工具(码流 +自检与交叉核对、渲染比对、设置页布局检查、容器探测检查),`tools\` 是构建与测试床脚本。设置项 +另有 `JOC_*` 环境变量覆盖(仅用于开发运行),清单与含义在 `src\settings.cpp`。 -## Install +排查问题看 DLL 旁边的 `joc_decoder.log`;组件启动时会把自己的版本、核心版本、日志路径写在 +里面。 -Either drop `dist\foo_input_joc--.fb2k-component` onto foobar2000 (or use -Preferences → Components → Install), or copy `foo_input_joc.dll` into -`\user-components\foo_input_joc\`. The per-component subdirectory is required: -a DLL lying directly in `user-components\` is not scanned. 1.6 is 32-bit, 2.x ships both, -and a DLL of the wrong architecture is silently ignored. +## 已知限制 -`tools/deploy.ps1 -TestBed ` does the manual variant, and -`tools/run.ps1 -TestBed -Play ` runs it unattended and prints the log. -Always let foobar2000 exit through `/exit`; a force-killed instance leaves a -`\running` marker behind and the next start then refuses to load any user -component. +* 不实现 ADM BWF 输出。 +* 容器只有在它排在内置容器读取器之前时才会被接管(见[安装](#容器需要调一次解码器顺序));核心 + 不允许某个解码器去要一个已经被别的条目拿走的文件。这类文件的标签也仍旧归那个读取器。 +* 裸流 `.eac3` / `.ec3` 的标签(流前面的 ID3v2,或后面的 APEv2/ID3v1)**能读不能写**:没有 + 组件声明可以写裸 E-AC-3,为插入标签重写整个文件也不是本组件该做的事。 +* 传输流(`.ts`、`.m2ts`)不接管。 +* 播放长度严格等于文件时长。双耳渲染器仍会算出房间尾音,但它不作为文件本身没有的播放时间交付。 +* 跳转会从包含目标位置的那个帧重新进入码流,而不是把前面的内容全部解码一遍——这是跳转代价与 + 目标位置无关的原因。位置精确、不漂移;样本是同一段波形交给了从该处开始的解码器,与从头播放 + 相比差一个很低的噪声底(−59 dBFS 或更低)。 -### Containers need one look at the decoder list +## 许可 -foobar2000 tries the decoders in the order shown in Preferences → **Decoding** (the -"list of available decoders", where entries can be moved up and down). The built-in -container readers are in that list too, and when one of them is offered an MP4 or Matroska -file before this component, it takes the file and the JOC objects are lost — the file plays -as plain E-AC-3. - -So, to play JOC from a container, move **JOC decoder (E-AC-3 JOC)** above **foobar2000 MP4 -Demuxer** and **foobar2000 Matroska/WebM Reader** in that list. Nothing else is needed, and -bare `.eac3` / `.ec3` files are unaffected by the order. This is the same thing every -third-party decoder (the FFmpeg wrapper, for one) asks for, which is why the component does -not try to work around it. If a container still plays as plain E-AC-3, that list is where to -look. - -## Settings - -Preferences → Tools → **JOC decoder**: - -* **Output** — binaural, or a speaker layout from 2.0 to 7.1; -* **Binaural mode** (near / mid / far) and the room **tail** in seconds; -* **HRTF source** — a **SOFA** file or a **Rosella** `.personalized_headphone` model. Leave - the path empty to use the default location `\HRTF\`: - `binaural.sofa` or `binaural.personalized_headphone`; -* **Gain** — a switch plus a value in dB. Binaural rendering can exceed full scale on - material that does not clip in the core mix, so attenuation belongs here; -* the **ffmpeg** executable to use. - -Nothing on the page is disabled; the status line states what is in effect. - -**HRTF data is not distributed with this repository.** A SOFA measurement set or a -personalised headphone model is supplied by whoever runs the component (and is listed in -`.gitignore` so it cannot be committed by accident). Speaker layouts and every offline test -except binaural rendering work without one; binaural rendering without an HRTF fails with a -message naming the file it looked for. - -## Environment overrides - -Development only: they override the stored settings for one run and every use is logged. -`JOC_OUTPUT`, `JOC_LAYOUT`, `JOC_HRTF`, `JOC_HRTF_SOURCE`, `JOC_BINAURAL_MODE`, `JOC_GAIN_DB`, -`JOC_GAIN_ENABLED`, `JOC_TAIL_SECONDS`, `JOC_OBJECT_DELAY`, `JOC_THREADS`, `JOC_FFMPEG`, -`JOC_LOG`. - -## Known limitations - -* ADM BWF output is not implemented. -* A container is only claimed when this component is ahead of the built-in container reader - in Preferences → Decoding, as described under [Install](#install); the core does not let a - decoder ask for a file another entry has already taken. -* Transport streams (`.ts`, `.m2ts`) are not claimed. -* The room tail is returned in full; the reference command-line renderer additionally trims - trailing samples below a threshold, so its output can be shorter. -* x86 and x64 do not produce bit-identical binaural output (last-bit differences): the - renderer's SIMD dispatch only applies to x86-64/ARM64, so 32-bit builds take the scalar - path. The speaker path is bit-identical on both. - -## Licence - -`LICENSE` is the upstream MIT licence, copied unchanged; `kernel/` is a copy of the upstream -renderer sources and keeps their notices. See [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md). +`LICENSE` 是上游 MIT 许可,原样复制;`kernel/` 是上游渲染内核源码的副本,保留其声明。见 +[THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md)。 diff --git a/VERIFICATION.md b/VERIFICATION.md deleted file mode 100644 index c7c18fe..0000000 --- a/VERIFICATION.md +++ /dev/null @@ -1,144 +0,0 @@ -# Verification - -Measured results only. Each entry names the command or the log line it came from, so it can -be reproduced. Environment: Windows x64 host, official **foobar2000 1.6.19 x86** portable -installation, official SDK 2026-09-17 pinned to `FOOBAR2000_TARGET_VERSION 80`, MSVC 14.44, -ffmpeg 8.0. - -Test material is supplied locally and is **not** part of this repository: the `testdata/` and -`vectors/` files of the upstream renderer project, and — for the binaural measurements — an -HRTF file. Everything except binaural rendering runs without any HRTF; binaural runs take the -file as an argument (`tests/render_harness.cpp --hrtf …`) or use the default location beside -the DLL. - -## Component and renderer - -| Check | Result | -|---|---| -| Sources compiled in, nothing loaded at run time | log: `core: in-process renderer 0.1.0-m1 (abi 3), component built against abi 3` | -| One artefact, no companion DLL | `dist\*.fb2k-component` holds `foo_input_joc.dll` and `README.md` only | -| Kernel sources untouched | the upstream working tree's file timestamps are unchanged; it is only ever read | -| Both architectures build | `build\Win32\foo_input_joc.dll`, `build\x64\foo_input_joc.dll` | - -## Playback in foobar2000 1.6.19 x86 - -| Case | Log evidence | -|---|---| -| Speaker 7.1, 5 s file | `stream created, 8 output channel(s), layout=7.1` … `frames_in=157 frames_out=157 samples_out=241152` — 157 × 1536 exactly | -| Binaural, SOFA | `stream created, 2 output channel(s)` … `end of stream after 481215 frames` (241152 source + 240063 tail) | -| Binaural, Rosella model | `stream created, 2 output channel(s)` … `end of stream after 481855 frames` | -| Binaural, HRTF path left empty | resolves to `\HRTF\binaural.sofa` and produces the same 481215 frames as naming that file explicitly | -| Full 238 s file, binaural SOFA | `eac3 frames queued=7436, bed frames pushed=7436` … `samples_out=11661759`; no ffmpeg process left behind | -| Installed from the `.fb2k-component` package | unpacked into `user-components\foo_input_joc\`, plays with the default-folder HRTF | - -## Bitstream recognition - -`tests/scan_crosscheck.cpp` compares, frame by frame, the syncframe lengths and the JOC -verdict this component computes against the renderer's own `joc_eac3_frame_bytes()` / -`joc_parse_eac3_frame()`: - -``` -testdata\gold_forever.eac3 frames=64 plugin_joc=64 kernel_joc=64 len_mismatch=0 verdict_mismatch=0 -vectors\valid.eac3 frames=7 plugin_joc=7 kernel_joc=7 len_mismatch=0 verdict_mismatch=0 -build\plain_eac3.eac3 frames=64 plugin_joc=0 kernel_joc=0 len_mismatch=0 verdict_mismatch=0 -crosscheck: 3 file(s), AGREES WITH CORE -``` - -Ten corrupt vectors were compared as well: no frame-length disagreement, verdicts agreed on -9 of 10. The one difference is `corrupt_truncated_huffman.eac3`: this component only tests -for the container, while the renderer also parses the payload and reports a truncated -bitstream later. - -## Plain E-AC-3 is handed back - -Playing a file the component's own encoder produced without JOC: - -``` -decoder: open "...plain_eac3.eac3" reason=1 bytes=144384 frames=8 with_joc=0 -decoder: yielding to the built-in decoder (at least one examined syncframe has no JOC EMDF container) -``` - -The file then plays through the built-in decoder; no decode log appears for it. Verified both -before and after the renderer was compiled in. - -## Output identical to the reference renderer - -`tests/render_harness.cpp` drives the component's own engine and writes a WAV in the same -format the reference command-line renderer writes, so the two files can be compared byte for -byte. 30 s of the reference file, speaker 5.1: - -| Product | Whole-file SHA-256 | -|---|---| -| reference renderer (`--speaker-layout 5.1 --bed … --duration 30`) | `99a8e3edbd1c047a3c0f547eaf85e56941f882af9e65468fbde0a9f18b6c5b6e` | -| this component's engine, x64 | `99a8e3edbd1c047a3c0f547eaf85e56941f882af9e65468fbde0a9f18b6c5b6e` | -| this component's engine, x86 | `99a8e3edbd1c047a3c0f547eaf85e56941f882af9e65468fbde0a9f18b6c5b6e` | - -34,578,500 bytes each. Binaural with a real SOFA file: with the same input frame count the -whole file is identical too (`588a15ce5526977f…baa88`, 12,158,780 bytes), and over the whole -238 s file the total sample count matches exactly (11,661,759 = 11,420,735 program + -241,024 tail) with identical peak (1.144561172) and an identical SHA-256 over the reference -renderer's entire payload. - -The one structural difference is the tail: the reference renderer trims trailing samples -below 1e-8 and this component returns the tail in full. - -## Gain - -Same file, speaker 5.1, rendered with and without attenuation: - -| Check | Result | -|---|---| -| Peak | 0.183568597 → 0.092002235, i.e. exactly −6.000000 dB | -| RMS over the whole signal | exactly −6.000000 dB | -| Per-sample, 1,446,912 floats | largest deviation from the ideal scaling is 2.1e-06 relative; the gain is applied in double precision before the DSP, so the remaining difference is float32 rounding | -| +6 dB | same check passes | -| The switch | switch on and −6 dB → log `gain=-6.00 dB`, delivered peak 0.000063; switch off with −6 dB still stored → `gain=0.00 dB`, peak 0.000126; switch on and 0 dB → peak 0.000126 | - -## Preferences page - -`tests/prefs_layout_check.cpp` builds the page from the component's own dialog resource and -asserts, per control, that it is enabled, lies inside the client area, and is not covered by -another interactive control (static text and group boxes are transparent to the mouse, as -they are for real clicks): - -``` -dialog client=495x367 non-client=0x0 - WS_CAPTION=no WS_BORDER=no WS_CHILD=yes WS_VISIBLE=yes -content extent=485x354 client=495x367 (everything fits) -controls=26 problems=0 -``` - -The page has no caption of its own — the host draws the frame — and no control is ever -disabled, which is what "the option is there but cannot be clicked" otherwise looks like. - -Not verified here: how the page and the `%joc_*%` fields look on screen; that needs a human -in front of the window. - -## Container support (mp4 / m4a / mov / mkv / mka / webm) - -Verified with the Win32 build, a 5.1 speaker layout, and the component installed in a -portable foobar2000 1.6.19 profile: - -- `ffmpeg -i -c:a copy` muxed into MP4 and Matroska, then `-map 0:a:0 -c:a copy - -f eac3` extracted again, is byte-identical to a direct `-t 30 -c:a copy` of the source - (same SHA-256) — the renderer therefore sees the stored syncframes, not a re-encode. -- The header probe (`tests/container_scan_test.cpp`, no foobar2000 involved) reports, for the - same material: `joc.mp4 -> mp4 eac3=1 audio#0 codec=ec-3 30.016 s`, - `joc.mkv -> matroska eac3=1 audio#0 codec=A_EAC3 30.016 s`, `plain_eac3.mp4 -> ec-3`, - `ac3.mp4 -> ac-3 (declined)`, `aac.mp4 -> mp4a (declined)`. -- End to end, with the component ordered ahead of the container reader in - Preferences -> Decoding: `open()` is called, the track is found, the JOC verdict is positive, - the file is claimed, and playback reaches `end of stream`. -- Negative case end to end: an MP4 holding E-AC-3 without JOC yields with - `E-AC-3 track 0 carries no JOC`, and the built-in decoder plays it. -- Bare `.eac3` / `.ec3` handling is unchanged. - -## Bed decode and gain - -* The 5.1 bed is decoded with -drc_scale 0 -target_level 0, so it is taken as stored and the - decoder does not apply the stream's dynrng or target-level metadata. The resulting bed is - byte-identical to the one the reference renderer uses (same SHA-256). -* The master gain is applied once, by the JOC kernel, on every output path. Measured at -6 dB - against 0 dB: speaker 5.1 and SOFA binaural give 0.50119 in peak and in per-channel RMS, - Rosella binaural gives 0.50119 as well, and a 0 dB render is unaffected by the renderer's - output gain. \ No newline at end of file diff --git a/src/input_joc.cpp b/src/input_joc.cpp index 2e52023..06fecc6 100644 --- a/src/input_joc.cpp +++ b/src/input_joc.cpp @@ -11,9 +11,11 @@ #include #include #include +#include #include #include #include +#include #include #include @@ -31,6 +33,18 @@ constexpr std::size_t kSniffBytes = 256u * 1024u; constexpr std::size_t kRunFrames = 4096u; constexpr unsigned kSampleRate = 48000; +// Largest magnitude in a block, for the delivery check in decode_run(). +template +double peak_of(const Sample* values, std::size_t count) { + double peak = 0.0; + for (std::size_t i = 0; i < count; ++i) { + const double value = + values[i] < Sample(0) ? -static_cast(values[i]) : static_cast(values[i]); + if (value > peak) peak = value; + } + return peak; +} + // Identity in the decoder priority table. const GUID g_decoder_guid = {0x9c3f1d58, 0x27ab, 0x4e64, {0xb0, 0x93, 0x5e, 0x1c, 0xd7, 0x48, 0x2f, 0xa6}}; @@ -52,13 +66,131 @@ std::string file_name_of(const std::string& path) { return slash == std::string::npos ? path : path.substr(slash + 1); } +// --------------------------------------------------------------------------- +// Tags of a file this component has taken over. +// +// An MP4/M4A keeps its tags in its own metadata box, and the component that knows +// how to read and write them is the container reader the core already ships. +// Claiming a file for decoding must not take it away from that reader, and the SDK +// has no "decode with me, ask someone else for tags" arrangement -- whichever +// entry answers open() answers for everything. So the information read and write +// paths are forwarded to whichever other entry claims the file, and only the tags +// of its answer are merged into ours: the technical information stays this +// component's own, which is what tells a user the file is JOC rather than plain +// E-AC-3. +// --------------------------------------------------------------------------- + +// Entries other than this one that claim the path, in the user's own decoding +// order. Ourselves is never in the list: an open forwarded back here would enter +// open() again, for ever. +void forwarding_candidates(const char* url, pfc::list_t& out) { + out.remove_all(); + input_manager_v3::ptr manager; + if (input_manager_v3::tryGet(manager)) { + manager->get_enabled_inputs(out); + } else { + input_entry::g_find_inputs_by_path_ex(out, url, + [](input_entry::ptr) { return true; }); + } + const char* dot = std::strrchr(url, '.'); + const char* extension = (dot != nullptr) ? dot + 1 : ""; + const GUID self = g_decoder_guid; + for (t_size index = out.get_count(); index-- > 0;) { + input_entry::ptr entry = out[index]; + if (entry->get_guid_() == self || !entry->is_our_path(url, extension)) { + out.remove_by_idx(index); + } + } +} + +// Opens the file again through another entry, for information reading or writing. +// The file is left unopened on our side, so the other entry can have it to itself. +template +bool open_forwarded(service_ptr_t& out, const GUID& what_for, const char* url, + abort_callback& abort, pfc::string8* name) { + out.release(); + pfc::list_t candidates; + forwarding_candidates(url, candidates); + if (candidates.get_count() == 0) return false; + try { + GUID used = pfc::guid_null; + service_ptr opened = input_entry::g_open_from_list(candidates, what_for, nullptr, url, + nullptr, abort, &used); + if (!opened.is_valid() || !opened->service_query_t(out)) return false; + if (name != nullptr) { + input_entry::ptr entry = input_entry::g_find_by_guid(used); + *name = entry.is_valid() ? entry->get_name_() : "another component"; + } + return true; + } catch (const pfc::exception& error) { + joc_log::line("decoder: no other component answers for this file's tags: %s", + error.what()); + return false; + } +} + +// Bytes in front of the E-AC-3 stream, which is where a tagging tool puts an +// ID3v2 tag. The renderer refuses a stream that does not begin on a syncword and +// never resynchronises, so the walk and the feed both have to start after it. +t_filesize leading_tag_bytes(file::ptr const& source, abort_callback& abort) { + if (!source.is_valid()) return 0; + try { + if (source->get_position(abort) != 0) source->seek(0, abort); + return tag_processor::skip_id3v2(source, abort); + } catch (const pfc::exception& error) { + joc_log::line("decoder: cannot inspect the area in front of the stream: %s", + error.what()); + return 0; + } +} + +// Tags read straight from the file, for a bare stream that no other component +// claims: an ID3v2 tag in front of the syncframes, or an APEv2/ID3v1 tag behind +// them. Neither is part of E-AC-3, so a tag that is there was written by a +// tagging tool and is worth showing. +void read_local_tags(file::ptr const& source, file_info& info, abort_callback& abort) { + if (!source.is_valid()) return; + bool found = false; + try { + source->seek(0, abort); + tag_processor::read_id3v2(source, info, abort); + found = true; + } catch (const pfc::exception&) { + // No leading tag; the trailing one is still worth a look. + } + try { + tag_processor::read_trailing(source, info, abort); + found = true; + } catch (const pfc::exception&) { + } + if (found) { + joc_log::line("decoder: %u tag field(s) read from the file itself", + static_cast(info.meta_get_count())); + } +} + class input_joc : public input_stubs { public: void open(service_ptr_t hint, const char* path, t_input_open_reason reason, abort_callback& abort) { - if (reason == input_open_info_write) throw exception_tagging_unsupported(); m_path = (path != nullptr) ? path : ""; + if (reason == input_open_info_write) { + // Writing tags belongs to whoever owns the file's format, and that is + // not this component: its inputs are two ffmpeg children and the JOC + // renderer, none of which writes anything. The file is deliberately + // left unopened here, because a write-mode handle of ours would make + // the writer that replaces it fail on a sharing violation. + m_write_only = true; + if (!open_forwarded(m_forward_writer, input_info_writer::class_guid, m_path.c_str(), + abort, &m_forward_name)) { + throw exception_tagging_unsupported(); + } + joc_log::line("decoder: open \"%s\" reason=2 tags: written by %s", m_path.c_str(), + m_forward_name.c_str()); + return; + } + service_ptr_t source = hint; input_open_file_helper(source, path, reason, abort); m_file = source; @@ -78,7 +210,31 @@ public: if (joc_container::is_container_extension(extension)) { open_container(extension); - return; + } else { + open_bare(abort); + } + + // Whatever else happens, the tags of this file are read by the component + // that owns its format; failing to find one is not fatal, the technical + // information below is still worth showing. + if (!open_forwarded(m_forward_reader, input_info_reader::class_guid, m_path.c_str(), abort, + &m_forward_name)) { + joc_log::line("decoder: no other component reads this file's tags"); + } else { + joc_log::line("decoder: tags for \"%s\" are read by %s", m_path.c_str(), + m_forward_name.c_str()); + } + } + + // A bare stream is either read from the file or handed back. One thing has to + // happen first: a tag area in front of the syncframes is not part of the + // stream, and treating it as one would hand the file to the built-in decoder, + // which plays it without the Atmos objects. + void open_bare(abort_callback& abort) { + m_stream_start = leading_tag_bytes(m_file, abort); + if (m_stream_start != 0) { + joc_log::line("decoder: %llu byte(s) of tags in front of the stream are skipped", + static_cast(m_stream_start)); } pfc::array_t buffer; @@ -86,9 +242,8 @@ public: const std::size_t got = m_file->read(buffer.get_ptr(), kSniffBytes, abort); const joc_eac3::ScanResult scan = joc_eac3::scan(buffer.get_ptr(), got, 8); - joc_log::line("decoder: open \"%s\" reason=%d bytes=%llu frames=%llu with_joc=%llu", - m_path.c_str(), static_cast(reason), - static_cast(got), + joc_log::line("decoder: open \"%s\" reason=1 bytes=%llu frames=%llu with_joc=%llu", + m_path.c_str(), static_cast(got), static_cast(scan.frames_examined), static_cast(scan.frames_with_joc)); @@ -137,8 +292,18 @@ public: } void get_info(file_info& info, abort_callback& abort) { - (void)abort; - const joc_decode::FileProbe probe = joc_decode::probe_file(m_native_path.get_ptr()); + if (m_write_only) { + // An instance opened to write tags is the writer's reader: what it + // reports is exactly what the caller has just written. + if (m_forward_writer.is_valid()) { + m_forward_writer->get_info(0, info, abort); + return; + } + throw exception_tagging_unsupported(); + } + + const joc_decode::FileProbe probe = + joc_decode::probe_file(m_native_path.get_ptr(), 0, m_stream_start); const joc_decode::Settings settings = joc_settings::current(); // A container knows its own duration even though the E-AC-3 syncframes are @@ -171,6 +336,10 @@ public: info.info_set_int("bitspersample", 32); info.info_set("bitspersample_extra", "floating-point"); info.set_length(duration); + m_length = duration; + // The renderer keeps its room tail, but the stream this component hands over + // ends where the file ends: the tail is rendering, not playback time. + m_engine.set_length(duration); if (duration > 0.0) { const t_filesize bytes = m_file.is_valid() ? m_file->get_size(abort) : filesize_invalid; if (bytes != filesize_invalid && bytes > 0) { @@ -194,6 +363,32 @@ public: m_container.audio_index); info.info_set("joc_container", text); } + + // The file's own reader supplies the tags; nothing above this line is one. + // Only the metadata is taken over -- its technical information (E-AC-3, + // 6 channels, the stream's own bitrate) would replace this component's, + // which is the part that says whether the file is JOC. + bool have_tags = false; + if (m_forward_reader.is_valid()) { + try { + file_info_impl tags; + m_forward_reader->get_info(0, tags, abort); + info.copy_meta(tags); + have_tags = tags.meta_get_count() != 0; + joc_log::line("decoder: %u tag field(s) from %s", + static_cast(tags.meta_get_count()), + m_forward_name.c_str()); + } catch (const pfc::exception& error) { + joc_log::line("decoder: reading this file's own tags failed: %s", error.what()); + } + } + // A reader that answers for the format but has nothing to say about a bare + // stream is common -- ffmpeg's AC-3 decoder reads no tags at all -- while + // the file may still carry an ID3v2 or APEv2 tag a tagging tool wrote. + if (!have_tags && m_input_kind == joc_decode::InputKind::kBare) { + read_local_tags(m_file, info, abort); + } + joc_log::line("decoder: get_info duration=%.3f s frames=%llu channels=%u render=%s%s", duration, static_cast(frames), channels, render.c_str(), container ? " (container)" : ""); @@ -201,6 +396,7 @@ public: t_filestats2 get_stats2(uint32_t flags, abort_callback& abort) { if (m_file.is_valid()) return m_file->get_stats2_(flags, abort); + if (m_forward_writer.is_valid()) return m_forward_writer->get_stats2_(nullptr, flags, abort); throw exception_io_unsupported_format(); } @@ -209,6 +405,8 @@ public: m_settings = joc_settings::current(); m_settings.input_kind = m_input_kind; m_settings.audio_index = m_audio_index; + m_settings.stream_start_bytes = m_stream_start; + m_settings.length_seconds = m_length; joc_log::line("decoder: initialize flags=0x%X settings: %s", flags, joc_settings::describe(m_settings).c_str()); @@ -229,6 +427,7 @@ public: m_buffer.resize(kRunFrames * m_channels); m_frames_delivered = 0; m_reported = false; + m_delivery_mismatches = 0; joc_log::line("decoder: engine ready, %u output channel(s), %u frames per read", m_channels, static_cast(kRunFrames)); } @@ -242,50 +441,98 @@ public: joc_log::line("decoder: read failed: %s", error.c_str()); throw exception_io_data(error.c_str()); } - joc_log::line("decoder: end of stream after %llu frames", - static_cast(m_frames_delivered)); + joc_log::line("decoder: end of stream after %llu frames%s", + static_cast(m_frames_delivered), + m_delivery_mismatches == 0 ? "" + : " (the delivery changed samples)"); return false; } - chunk.set_data_size(frames * m_channels); - chunk.set_channels(m_channels, audio_chunk::g_guess_channel_config(m_channels)); - chunk.set_sample_rate(kSampleRate); - chunk.set_sample_count(frames); - std::memcpy(chunk.get_data(), m_buffer.data(), - frames * m_channels * sizeof(audio_sample)); + // The renderer produces float32 and a chunk holds audio_sample, which is float on + // 32-bit builds and double on 64-bit ones (SDK audio_math.h): the samples are + // converted, not copied. set_data_32() is the SDK's conversion for a float32 + // source, and it sets the channel count, the sample rate and the sample count. + chunk.set_data_32(m_buffer.data(), frames, m_channels, kSampleRate); m_frames_delivered += frames; + // A delivery that mangled the samples would be heard as noise rather than reported + // as a failure, so every chunk is checked: the conversion is exact, and the peak of + // what the chunk holds has to equal the peak of what the renderer produced. + const double produced = peak_of(m_buffer.data(), frames * m_channels); + const double delivered = + peak_of(chunk.get_data(), chunk.get_sample_count() * chunk.get_channels()); + if (delivered > produced + 1e-6 + produced * 1e-6 || + delivered < produced - 1e-6 - produced * 1e-6) { + ++m_delivery_mismatches; + if (m_delivery_mismatches == 1) { + joc_log::line("decoder: delivery changed the samples: peak %.9f produced, " + "%.9f delivered", + produced, delivered); + } + } if (!m_reported) { m_reported = true; - float peak = 0.0f; - for (std::size_t i = 0; i < frames * m_channels; ++i) { - const float value = m_buffer[i] < 0.0f ? -m_buffer[i] : m_buffer[i]; - if (value > peak) peak = value; - } - joc_log::line("decoder: first %llu frames delivered (%u ch), peak %.6f", - static_cast(frames), m_channels, - static_cast(peak)); + joc_log::line("decoder: first %llu frames delivered (%u ch), peak %.6f, " + "delivered peak %.6f", + static_cast(frames), m_channels, produced, + delivered); } return true; } - void decode_seek(double, abort_callback&) { - // The renderer is stateful and has no seek; can_seek() says so. - throw exception_io_unsupported_format(); + void decode_seek(double seconds, abort_callback& abort) { + // Walking the syncframe index of a long file is the only part of a seek + // that can take a while, and it polls this. + m_engine.set_abort_check([&abort] { return !abort.is_aborting(); }); + std::string error; + const bool ok = m_engine.seek(seconds, m_length, &error); + m_engine.set_abort_check(nullptr); + if (!ok) { + // An aborted seek reports itself as an abort, not as a decode failure. + abort.check(); + joc_log::line("decoder: seek to %.6f s failed: %s", seconds, error.c_str()); + throw exception_io_data(error.c_str()); + } + // The position reporting and the first-read statistics belong to the run + // that starts here, not to the one that was interrupted. + m_frames_delivered = 0; + m_reported = false; + m_delivery_mismatches = 0; + joc_log::line("decoder: seek to %.6f s, the next read starts at the target", seconds); } - bool decode_can_seek() { return false; } + bool decode_can_seek() { return true; } size_t extended_param(const GUID& type, size_t arg1, void* arg2, size_t arg2size) { (void)arg1; (void)arg2; (void)arg2size; + // A seek restarts both ffmpeg children and replays the renderer's warm-up, + // so it is worth avoiding the ones the core would only make speculatively. if (type == input_params::seeking_expensive) return 1; return 0; } - void retag(const file_info&, abort_callback&) { throw exception_tagging_unsupported(); } - void remove_tags(abort_callback&) { throw exception_tagging_unsupported(); } + void retag(const file_info& info, abort_callback& abort) { + if (!m_forward_writer.is_valid()) throw exception_tagging_unsupported(); + // A single-track input has no commit() of its own -- the SDK wrapper + // implements it as a no-op -- so the writer's commit has to happen here or + // nothing reaches the file. + m_forward_writer->set_info(0, info, abort); + m_forward_writer->commit(abort); + joc_log::line("decoder: %u tag field(s) written through %s", + static_cast(info.meta_get_count()), m_forward_name.c_str()); + } + + void remove_tags(abort_callback& abort) { + if (!m_forward_writer.is_valid()) throw exception_tagging_unsupported(); + input_info_writer_v2::ptr v2; + if (m_forward_writer->service_query_t(v2)) { + v2->remove_tags(abort); + return; + } + m_forward_writer->remove_tags_fallback(abort); + } static bool g_is_our_path(const char* path, const char* extension) { (void)path; @@ -334,6 +581,19 @@ private: unsigned m_channels = 2; std::uint64_t m_frames_delivered = 0; bool m_reported = false; + // Chunks whose delivered samples did not match what the renderer produced. + std::uint64_t m_delivery_mismatches = 0; + // Duration get_info() last reported; a seek needs it to tell "past the end" + // from "inside the file" without decoding anything. + double m_length = 0.0; + // Bytes in front of a bare stream, which is where an ID3v2 tag sits. + std::uint64_t m_stream_start = 0; + // The other component that answers for this file's tags, and the one that + // writes them. Only the writer exists on an instance opened to retag. + service_ptr_t m_forward_reader; + service_ptr_t m_forward_writer; + pfc::string8 m_forward_name; + bool m_write_only = false; }; static input_singletrack_factory_t g_input_joc_factory; diff --git a/src/joc_decode.cpp b/src/joc_decode.cpp index 0f2df38..31b1211 100644 --- a/src/joc_decode.cpp +++ b/src/joc_decode.cpp @@ -2,8 +2,10 @@ #include +#include #include #include +#include #include // The kernel copy that is compiled into this component; see kernel/. @@ -42,6 +44,7 @@ struct CoreApi { std::uint32_t*) = joc_stream_push; joc_error(JOC_CALL* pull)(joc_stream*, joc_stream_buffer*, std::uint32_t*) = joc_stream_pull; joc_error(JOC_CALL* flush)(joc_stream*) = joc_stream_flush; + joc_error(JOC_CALL* reset)(joc_stream*) = joc_stream_reset; joc_error(JOC_CALL* status)(const joc_stream*, joc_stream_status_info*) = joc_stream_status; joc_error(JOC_CALL* destroy)(joc_stream*) = joc_stream_destroy; std::uint32_t(JOC_CALL* abi_version)() = joc_abi_version; @@ -199,6 +202,7 @@ public: FILE_ATTRIBUTE_NORMAL | FILE_FLAG_SEQUENTIAL_SCAN, nullptr); return handle_ != INVALID_HANDLE_VALUE; } + bool is_open() const { return handle_ != INVALID_HANDLE_VALUE; } std::size_t read(void* destination, std::size_t bytes) { if (handle_ == INVALID_HANDLE_VALUE) return 0; DWORD got = 0; @@ -207,6 +211,12 @@ public: } return got; } + bool seek(std::uint64_t offset) { + if (handle_ == INVALID_HANDLE_VALUE) return false; + LARGE_INTEGER value{}; + value.QuadPart = static_cast(offset); + return SetFilePointerEx(handle_, value, nullptr, FILE_BEGIN) != FALSE; + } std::uint64_t size() const { LARGE_INTEGER value{}; if (handle_ == INVALID_HANDLE_VALUE || GetFileSizeEx(handle_, &value) == FALSE) return 0; @@ -228,6 +238,41 @@ const char* const kLayouts[] = {"2.0", "3.0", "3.1", "4.0", "5.0", "5. "9.1.6", "22.2"}; const unsigned kLayoutChannels[] = {2, 3, 4, 4, 5, 6, 8, 10, 7, 7, 8, 10, 12, 14, 16, 24}; +// Every ffmpeg child writes its diagnostics next to the component rather than into +// whatever working directory the host process happens to have; separate files so +// no child can truncate another's. +std::wstring stderr_path_for(const wchar_t* name) { + HMODULE self = nullptr; + GetModuleHandleExW(GET_MODULE_HANDLE_EX_FLAG_FROM_ADDRESS | + GET_MODULE_HANDLE_EX_FLAG_UNCHANGED_REFCOUNT, + reinterpret_cast(&speaker_channels), &self); + wchar_t path[4096] = {}; + const DWORD length = GetModuleFileNameW(self, path, 4096); + if (length == 0) return std::wstring(name); + const std::wstring text(path, length); + const std::wstring::size_type slash = text.find_last_of(L"\\/"); + return slash == std::wstring::npos ? std::wstring(name) + : text.substr(0, slash) + L"\\" + name; +} + +// The time ffmpeg's -ss takes for a given source sample. A seek always starts on +// a syncframe boundary, so the sample is exactly representable in microseconds. +std::wstring seek_time(std::uint64_t source_sample) { + if (source_sample == 0) return {}; + wchar_t text[64] = {}; + std::swprintf(text, 64, L"%.6f", static_cast(source_sample) / 48000.0); + return text; +} + +// Source sample a time offset names, rounded the way the core rounds a position +// (pfc::rint64, i.e. llrint: to nearest, ties to even). +std::int64_t sample_of_seconds(double seconds) { + if (!(seconds > 0.0)) return 0; // also catches NaN + const double value = seconds * 48000.0; + if (value >= 9.0e18) return 9223372036854775807LL; + return static_cast(std::llrint(value)); +} + } // namespace const char* const* speaker_layouts(std::size_t* count) { @@ -277,7 +322,8 @@ std::string resolve_hrtf_file(const Settings& settings) { return directory + "\\HRTF\\" + name; } -FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) { +FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes, + std::uint64_t start_offset) { FileProbe probe; InputFile file; if (!file.open(path)) { @@ -285,6 +331,15 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) { return probe; } const std::uint64_t size = file.size(); + if (start_offset >= size) { + probe.detail = "file too small to be E-AC-3"; + return probe; + } + if (start_offset != 0 && !file.seek(start_offset)) { + probe.detail = "cannot skip the leading tag area"; + return probe; + } + const std::uint64_t stream_bytes = size - start_offset; // A Media Library scan calls this for every file, so the whole stream is // walked only when that is cheap; otherwise the first window is enough, @@ -293,7 +348,7 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) { const std::size_t window = (max_scan_bytes != 0) ? max_scan_bytes : 256u * 1024u; std::vector buffer(static_cast( - (size < kFullWalkLimit && size > 0) ? size : window)); + (stream_bytes < kFullWalkLimit && stream_bytes > 0) ? stream_bytes : window)); const std::size_t got = file.read(buffer.data(), buffer.size()); if (got < 8) { probe.detail = "file too small to be E-AC-3"; @@ -310,7 +365,7 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) { } std::uint64_t frames = 0; - if (buffer.size() == size) { + if (buffer.size() == stream_bytes) { std::size_t offset = 0; while (true) { const std::size_t bytes = joc_eac3::frame_bytes_at(buffer.data(), got, offset); @@ -320,7 +375,7 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) { } probe.detail = "frame count walked over the whole file"; } else if (scan.all_frames_same_size) { - frames = size / scan.first_frame_bytes; + frames = stream_bytes / scan.first_frame_bytes; probe.detail = "frame count extrapolated from a constant frame size"; } else { // Variable frame size: count in the window and scale by the byte ratio. @@ -333,7 +388,8 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) { ++seen; } frames = (offset != 0) ? static_cast( - (static_cast(size) / static_cast(offset)) * + (static_cast(stream_bytes) / + static_cast(offset)) * static_cast(seen)) : 0; probe.detail = "frame count estimated from a variable frame size"; @@ -443,6 +499,7 @@ struct Engine::Impl { bool eac3_from_pipe = false; FfmpegPipe bed; Settings settings; + std::string input_path; // The metadata stream comes either from the file itself or from ffmpeg. std::size_t read_eac3(void* destination, std::size_t bytes) { @@ -464,8 +521,192 @@ struct Engine::Impl { std::vector bed_buffer; std::size_t bed_staged_bytes = 0; // bytes staged at the front of bed_buffer std::vector pull_buffer; + + // Seek state. seek_skip counts output frames still to be discarded: a seek + // restarts the inputs on the syncframe before the target, and the samples the + // renderer produces for the part already behind the target are dropped here. + std::uint64_t seek_skip = 0; + // Delivered-stream accounting: where the stream currently being delivered starts + // on the source timeline, how much of it has been handed over, and where it has + // to stop. end_sample is the file's own end, so a binaural room tail cannot + // turn into playback time the file does not have. + std::uint64_t start_sample = 0; + std::uint64_t delivered = 0; + std::uint64_t end_sample = 0; // 0 = no limit + // Syncframes of a container's metadata stream still to be read and thrown away. + // ffmpeg's input seek on a copy stream lands on the frame whose timestamp is + // at or after the target, which is not reliably the frame the bed restarts on, + // so a container seek re-reads the stream and drops the frames here instead. + std::uint64_t metadata_skip = 0; + bool at_end = false; // seek landed at or past the end + std::function abort_check; + // Bare-stream syncframe index: offset of every syncframe within the stream + // (so index_base has to be added to get a file offset). Filled by walking the + // file, and only as far as a seek asks for. + std::vector frame_offsets; + std::uint64_t index_base = 0; // file offset the stream starts at + std::uint64_t index_file_offset = 0; // file offset the walk continues from + std::uint64_t index_stream_offset = 0; // stream offset the walk continues from + bool index_complete = false; + + void reset_feed_state() { + frames_queued = 0; + bed_frames_pushed = 0; + eac3_eof = false; + bed_eof = false; + flushed = false; + trace_count = 0; + read_calls = 0; + eac3_carry = 0; + bed_bytes_read = 0; + bed_staged_bytes = 0; + seek_skip = 0; + metadata_skip = 0; + start_sample = 0; + delivered = 0; + at_end = false; + } + + void stop_inputs() { + bed.stop(); + eac3_pipe.stop(); + eac3.close(); + } + + void reset_index() { + frame_offsets.clear(); + index_base = settings.stream_start_bytes; + index_file_offset = index_base; + index_stream_offset = 0; + index_complete = false; + } + + // Grows the syncframe index until it holds `wanted` entries or the stream ends. + bool grow_frame_index(std::uint64_t wanted, std::string* error); + // Points the file at the syncframe `frame`, or reports at_end when the stream + // has fewer frames than that. + bool position_bare(std::uint64_t frame, std::string* error); + bool start_bed(std::uint64_t source_sample, std::string* error); + bool start_metadata(std::string* error); }; +bool Engine::Impl::grow_frame_index(std::uint64_t wanted, std::string* error) { + constexpr std::size_t kWindow = 256u * 1024u; + std::vector buffer(kWindow); + while (frame_offsets.size() < wanted && !index_complete) { + if (abort_check != nullptr && !abort_check()) { + if (error != nullptr) *error = "the seek was aborted"; + return false; + } + if (!eac3.is_open() && !eac3.open(input_path)) { + if (error != nullptr) *error = "cannot reopen the input file"; + return false; + } + const std::uint64_t size = eac3.size(); + if (index_file_offset + 4u > size) { + index_complete = true; + break; + } + if (!eac3.seek(index_file_offset)) { + if (error != nullptr) *error = "cannot position the input file"; + return false; + } + const std::uint64_t remaining = size - index_file_offset; + const std::size_t want = static_cast( + remaining < kWindow ? remaining : kWindow); + const std::size_t got = eac3.read(buffer.data(), want); + if (got < 4u) { + index_complete = true; + break; + } + std::size_t consumed = 0; + while (frame_offsets.size() < wanted) { + const std::size_t bytes = joc_eac3::frame_bytes_at(buffer.data(), got, consumed); + if (bytes == 0 || consumed + bytes > got) break; + frame_offsets.push_back(index_stream_offset + consumed); + consumed += bytes; + } + if (consumed == 0) { + // Not even one whole syncframe in a full window: the stream stops here. + index_complete = true; + break; + } + index_file_offset += consumed; + index_stream_offset += consumed; + } + return true; +} + +bool Engine::Impl::position_bare(std::uint64_t frame, std::string* error) { + if (!grow_frame_index(frame + 1u, error)) return false; + if (frame >= frame_offsets.size()) { + // Past the last syncframe: the decoder contract asks for a successful seek + // that the next read() answers with end of stream. + at_end = true; + stop_inputs(); + return true; + } + // Position check first: frame_offsets.size() is what bounds the index. + const std::uint64_t file_offset = + index_base + frame_offsets[static_cast(frame)]; + if (!eac3.is_open() && !eac3.open(input_path)) { + if (error != nullptr) *error = "cannot reopen the input file"; + return false; + } + // The renderer rejects a stream that does not begin on a syncword and never + // resynchronises, so a mispositioned start has to fail loudly here rather than + // turn into silence at the end of the file. + std::uint8_t header[8] = {}; + if (!eac3.seek(file_offset) || eac3.read(header, sizeof(header)) < 4u || + joc_eac3::frame_bytes_at(header, sizeof(header), 0) == 0) { + if (error != nullptr) { + *error = "the syncframe index does not point at a syncframe (offset " + + std::to_string(file_offset) + ")"; + } + return false; + } + if (!eac3.seek(file_offset)) { + if (error != nullptr) *error = "cannot position the input file"; + return false; + } + return true; +} + +bool Engine::Impl::start_bed(std::uint64_t source_sample, std::string* error) { + // The 5.1 core PCM, exactly as the reference renderer's own core decode does + // it: 5.1 interleaved float32 at 48 kHz, the layout the renderer expects + // (L R C LFE Ls Rs). + // -drc_scale 0 -target_level 0: the bed is taken as stored, without the + // stream's dynrng or target-level metadata being applied by the decoder. + // + // -ss is an input option: the demuxer is positioned on the frame boundary and + // decoding starts there, so a seek costs the same wherever it lands. The price is + // that the bed is carried by a decoder that started at the seek point rather than + // at the start of the file, which is a low-level, noise-like difference from a + // play-through rather than a misalignment: the position stays exact. + std::wstring input_arguments = L"-drc_scale 0 -target_level 0"; + const std::wstring offset = seek_time(source_sample); + if (!offset.empty()) input_arguments += L" -ss " + offset; + std::wstring bed_arguments = L"-map 0:a:"; + bed_arguments += std::to_wstring(settings.audio_index); + bed_arguments += L" -vn -ac 6 -ar 48000 -c:a pcm_f32le -f f32le -"; + return bed.start(settings.ffmpeg_path, input_path, input_arguments, bed_arguments, "bed", + stderr_path_for(L"joc_ffmpeg_bed.log"), error); +} + +bool Engine::Impl::start_metadata(std::string* error) { + // Stream copy from the start of the track: the syncframes arrive byte for byte + // as they are stored, which is what the JOC metadata needs, and a seek then + // discards whole syncframes from the front (see metadata_skip) rather than + // asking ffmpeg to position the stream. + std::wstring stream_arguments = L"-map 0:a:"; + stream_arguments += std::to_wstring(settings.audio_index); + stream_arguments += L" -vn -c:a copy -f eac3 -"; + return eac3_pipe.start(settings.ffmpeg_path, input_path, std::wstring(), stream_arguments, + "metadata", stderr_path_for(L"joc_ffmpeg_stream.log"), error, + 1u << 20); +} + Engine::Engine() : impl_(new Impl()) {} Engine::~Engine() { @@ -475,21 +716,30 @@ Engine::~Engine() { unsigned Engine::channels() const { return impl_->channels; } +void Engine::set_abort_check(std::function check) { impl_->abort_check = std::move(check); } + +void Engine::set_length(double seconds) { impl_->end_sample = sample_of_seconds(seconds); } + void Engine::stop() { Impl& impl = *impl_; if (impl.stream != nullptr && impl.api.destroy != nullptr) { impl.api.destroy(impl.stream); impl.stream = nullptr; } - impl.bed.stop(); - impl.eac3_pipe.stop(); - impl.eac3.close(); - impl.eac3_from_pipe = false; + impl.stop_inputs(); + impl.reset_feed_state(); } bool Engine::start(const std::string& input_path, const Settings& settings, std::string* error) { Impl& impl = *impl_; + // initialize() may be called more than once on the same instance, and each call + // has the renderer and its inputs start from scratch. + stop(); impl.settings = settings; + impl.input_path = input_path; + impl.reset_index(); + impl.reset_feed_state(); + impl.end_sample = sample_of_seconds(settings.length_seconds); impl.eac3_buffer.resize(kEac3Chunk); impl.bed_buffer.resize(kBedFramesChunk * kBedChannels); @@ -513,6 +763,12 @@ bool Engine::start(const std::string& input_path, const Settings& settings, std: if (error != nullptr) *error = "cannot open the input file"; return false; } + // A tag area in front of the stream is skipped here rather than by the + // renderer, which rejects a stream that does not begin on a syncword. + if (impl.index_base != 0 && !impl.eac3.seek(impl.index_base)) { + if (error != nullptr) *error = "cannot skip the leading tag area"; + return false; + } } joc_stream_config config{}; @@ -590,51 +846,89 @@ bool Engine::start(const std::string& input_path, const Settings& settings, std: // ffmpeg's stderr lands next to the component rather than in whatever working // directory the host process happens to have. The two children write separate // files so neither can truncate the other's diagnostics. - const auto stderr_path_for = [](const wchar_t* name) { - HMODULE self = nullptr; - GetModuleHandleExW(GET_MODULE_HANDLE_EX_FLAG_FROM_ADDRESS | - GET_MODULE_HANDLE_EX_FLAG_UNCHANGED_REFCOUNT, - reinterpret_cast(&speaker_channels), &self); - wchar_t path[4096] = {}; - const DWORD length = GetModuleFileNameW(self, path, 4096); - if (length == 0) return std::wstring(name); - const std::wstring text(path, length); - const std::wstring::size_type slash = text.find_last_of(L"\\/"); - return slash == std::wstring::npos ? std::wstring(name) - : text.substr(0, slash) + L"\\" + name; - }; + if (!impl.start_bed(0, error)) return false; + if (impl.eac3_from_pipe && !impl.start_metadata(error)) return false; + return true; +} - // The 5.1 core PCM, exactly as the reference renderer's own core decode does it: - // 5.1 interleaved float32 at 48 kHz, the layout the renderer expects - // (L R C LFE Ls Rs). - // -drc_scale 0 -target_level 0: the bed is taken as stored, without the stream's - // dynrng or target-level metadata being applied by the decoder. - std::wstring bed_arguments = L"-map 0:a:"; - bed_arguments += std::to_wstring(settings.audio_index); - bed_arguments += L" -vn -ac 6 -ar 48000 -c:a pcm_f32le -f f32le -"; - if (!impl.bed.start(settings.ffmpeg_path, input_path, L"-drc_scale 0 -target_level 0", bed_arguments, "bed", - stderr_path_for(L"joc_ffmpeg_bed.log"), error)) { +bool Engine::seek(double seconds, double total_seconds, std::string* error) { + Impl& impl = *impl_; + if (impl.stream == nullptr || impl.api.reset == nullptr) { + if (error != nullptr) *error = "the renderer is not running"; return false; } + // Where the stream that is about to be delivered starts: + // + // * the frame the target sits in, minus a pre-roll, so the renderer's own state + // -- object timeline, matrix interpolation, room tail -- has converged by the + // time the target itself is delivered. Both inputs restart there, and the bed + // is positioned with an input seek, which is what keeps the cost of a seek + // independent of where it lands; + // * the samples in front of the target are then dropped from the output, which is + // what makes the delivery start exactly on the requested sample. + const std::int64_t target = sample_of_seconds(seconds); + std::int64_t wanted = target - static_cast(impl.settings.pipeline_delay_samples); + if (wanted < 0) wanted = 0; + const std::uint64_t target_frame = static_cast(wanted) / kFrameSamples; + const std::uint64_t preroll = impl.settings.seek_preroll_frames; + const std::uint64_t frame = (target_frame > preroll) ? (target_frame - preroll) : 0; + const std::uint64_t first_sample = frame * kFrameSamples; + const std::uint64_t skip = static_cast(wanted) - first_sample; + + impl.stop_inputs(); + impl.reset_feed_state(); + if (impl.eac3_from_pipe) { - // Stream copy: the syncframes arrive byte for byte as they are stored, which - // is what the JOC metadata needs. - std::wstring stream_arguments = L"-map 0:a:"; - stream_arguments += std::to_wstring(settings.audio_index); - stream_arguments += L" -vn -c:a copy -f eac3 -"; - if (!impl.eac3_pipe.start(settings.ffmpeg_path, input_path, L"", stream_arguments, "metadata", - stderr_path_for(L"joc_ffmpeg_stream.log"), error, - 1u << 20)) { - return false; + // A container's frame count is not known before it is decoded, so the + // caller's duration is what says whether this lands past the end. + if (total_seconds > 0.0 && seconds >= total_seconds) { + impl.at_end = true; + joc_log::line("engine: seek %.6f s is at or past the end (%.6f s)", seconds, + total_seconds); } + } else if (!impl.position_bare(frame, error)) { + return false; } + + if (!impl.at_end) { + if (!impl.start_bed(first_sample, error)) return false; + if (impl.eac3_from_pipe) { + if (!impl.start_metadata(error)) return false; + impl.metadata_skip = frame; + } + impl.seek_skip = skip; + impl.start_sample = static_cast(wanted); + impl.delivered = 0; + } + + // joc_stream_reset keeps the renderer and its HRTF field: it drops the whole + // timeline, gain ramps, room tail and filter-bank history, which is exactly + // what a restart at another position needs. + const joc_error reset = impl.api.reset(impl.stream); + if (reset != JOC_OK) { + if (error != nullptr) { + *error = std::string("cannot reset the render stream: ") + + error_text(impl.api, reset); + } + return false; + } + joc_log::line("engine: seek %.6f s -> source sample %lld, frames %llu.. from sample %llu, " + "drop %llu sample(s)%s", + seconds, static_cast(target), + static_cast(frame), + static_cast(first_sample), + static_cast(impl.seek_skip), + impl.at_end ? " (at end)" : ""); return true; } std::size_t Engine::read(float* destination, std::size_t frames, std::string* error) { Impl& impl = *impl_; if (impl.stream == nullptr || frames == 0) return 0; + // A seek at or past the end of the file succeeds and leaves the next read to + // report end of stream. + if (impl.at_end) return 0; const bool trace = impl.trace_count < 6; ++impl.read_calls; if ((impl.read_calls % 50u) == 0u) { @@ -672,10 +966,36 @@ std::size_t Engine::read(float* destination, std::size_t frames, std::string* er // frame of the file unrendered. const std::size_t got = impl.read_eac3(impl.eac3_buffer.data() + impl.eac3_carry, impl.eac3_buffer.size() - impl.eac3_carry); - const std::size_t total = impl.eac3_carry + got; + std::size_t total = impl.eac3_carry + got; if (total == 0) { impl.eac3_eof = true; } else { + // Syncframes in front of a seek target are thrown away before + // anything is interpreted. They are dropped a whole window at a + // time, so only as much of the stream is read as the skip needs. + if (impl.metadata_skip != 0) { + std::size_t dropped = 0; + while (impl.metadata_skip != 0) { + const std::size_t bytes = + joc_eac3::frame_bytes_at(impl.eac3_buffer.data(), total, dropped); + if (bytes == 0 || dropped + bytes > total) break; + dropped += bytes; + --impl.metadata_skip; + } + if (dropped != 0) { + total -= dropped; + std::memmove(impl.eac3_buffer.data(), + impl.eac3_buffer.data() + dropped, total); + } + if (impl.metadata_skip != 0) { + // The window ended inside the part to be skipped: keep what + // is left and come back with the next read. + impl.eac3_carry = total; + if (got == 0) impl.eac3_eof = true; + continue; + } + } + std::size_t offset = 0; std::uint64_t complete = 0; while (true) { @@ -804,12 +1124,37 @@ std::size_t Engine::read(float* destination, std::size_t frames, std::string* er return 0; } if (produced != 0u) { + std::size_t count = produced; + if (impl.seek_skip != 0) { + // The renderer had to be fed from before the seek target, so the + // samples it produces for that part are dropped before anything + // reaches the caller: the first sample handed over is the target. + const std::size_t drop = (impl.seek_skip < count) + ? static_cast(impl.seek_skip) + : count; + impl.seek_skip -= drop; + count -= drop; + if (count != 0) { + std::memmove(destination, destination + drop * impl.channels, + count * impl.channels * sizeof(float)); + } + if (count == 0) continue; // the whole pull was pre-roll + } + // The renderer's binaural room tail is not part of the file: the stream + // ends where the file ends, not later. + if (impl.end_sample != 0) { + const std::uint64_t from = impl.start_sample + impl.delivered; + if (from >= impl.end_sample) return 0; + const std::uint64_t left = impl.end_sample - from; + if (left < count) count = static_cast(left); + } + impl.delivered += count; if (trace) { ++impl.trace_count; - joc_log::line("engine: pulled %u frame(s) on the %u%s attempt", produced, + joc_log::line("engine: pulled %u frame(s) on the %u%s attempt", count, impl.trace_count, impl.trace_count == 1 ? "st" : "th"); } - return produced; + return count; } // Nothing more can arrive once the metadata stream is drained and the bed diff --git a/src/joc_decode.h b/src/joc_decode.h index 53b7fad..e49ab03 100644 --- a/src/joc_decode.h +++ b/src/joc_decode.h @@ -19,10 +19,29 @@ #include #include +#include #include namespace joc_decode { +// Delay of the rendering pipeline itself, in output samples: in a run that began +// at source sample 0, output sample k carries the rendering of source sample +// k - this value. A seek starts the inputs this much earlier and drops the +// samples before the target. Measured as 0: the renderer's own filter-bank +// latency and its metadata delay are inside the renderer, not delays of the +// delivered stream. It stays a named, overridable constant so the measurement can +// be repeated. +constexpr std::uint32_t kJocSeekPipelineDelaySamples = 0; + +// Syncframes fed before the frame the target sits in. The renderer's state is +// rebuilt from the frames it is given, so a restart needs a moment to converge and +// the delivered part has to be past that. Two things converge at different rates: +// the metadata state (object positions, matrix interpolation, gain ramps), which +// the measured 3008-sample window covers, and the binaural room tail, which is +// recursive and can only be approached -- two syncframes are enough for the speaker +// path, and the tail keeps improving with more, which is why this is 64. +constexpr std::uint32_t kJocSeekPrerollFrames = 64; + enum class Output { kBinaural = 0, // 2 channels, HRTF rendering kSpeaker = 1, // N channels, named layout @@ -57,9 +76,22 @@ struct Settings { std::uint32_t binaural_mode = 3; // JOC_BINAURAL_MID double gain_db = 0.0; double tail_seconds = 5.0; + // Duration of the file, in seconds. The renderer ends a binaural stream with a + // room tail, which would be played as time the file does not have; the delivered + // stream is cut at exactly this much audio instead. Zero means "no limit". + double length_seconds = 0.0; std::uint32_t object_delay_samples = 1473; std::uint32_t native_threads = 0; std::string ffmpeg_path = "ffmpeg"; + // Bytes in front of the first syncframe -- a leading ID3v2 tag. The renderer + // rejects a stream that does not begin on a syncword, so both the walk and the + // feed start here. + std::uint64_t stream_start_bytes = 0; + // Samples of pipeline delay a seek compensates for, and syncframes it pre-rolls + // before the target: the calibrated constants above, unless the comparison + // harness overrides them to measure those constants. + std::uint32_t pipeline_delay_samples = kJocSeekPipelineDelaySamples; + std::uint32_t seek_preroll_frames = kJocSeekPrerollFrames; // Stop feeding the renderer after this many input syncframes and flush, which // is what the reference CLI's --duration does. Zero means "the whole file". // Only the comparison harness sets it; playback leaves it at zero. @@ -94,8 +126,10 @@ struct FileProbe { }; // Walks the file's syncframes. Cheap enough for a Media Library scan: it only -// does pointer arithmetic over the stream, no decoding, no HRTF work. -FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes = 0); +// does pointer arithmetic over the stream, no decoding, no HRTF work. The walk +// starts at `start_offset`, which is where the stream begins behind a tag area. +FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes = 0, + std::uint64_t start_offset = 0); // Decides whether the E-AC-3 track of a container file carries JOC, without // rendering anything: ffmpeg copies a short prefix of that track out (stream copy, @@ -118,6 +152,21 @@ public: unsigned channels() const; void stop(); + // Duration of the file, in seconds, for a caller that learns it after start() + // (reported duration of the track); zero means "no limit". + void set_length(double seconds); + + // Repositions so that the next sample read() returns is the rendering of source + // sample round(seconds * 48000): the inputs restart at the syncframe holding + // that sample and the output before it is dropped. Seeking at or past + // `total_seconds` (0 when the duration is unknown) succeeds and makes the next + // read() return 0, as the decoder contract requires. start() must have run. + bool seek(double seconds, double total_seconds, std::string* error); + + // Polled while the syncframe index is being walked, so a long seek stays + // interruptible; returning false aborts the seek. + void set_abort_check(std::function check); + private: struct Impl; Impl* impl_; diff --git a/src/main.cpp b/src/main.cpp index 8fae479..3dc711f 100644 --- a/src/main.cpp +++ b/src/main.cpp @@ -12,7 +12,7 @@ // Kept in one place: the string reported to foobar2000 and written to the log // must not drift apart. -#define JOC_VERSION "0.1.0" +#define JOC_VERSION "0.2.0" DECLARE_COMPONENT_VERSION("JOC decoder (E-AC-3 JOC)", JOC_VERSION, "Plays E-AC-3 JOC (Dolby Atmos) files: the JOC objects are " diff --git a/tests/render_harness.cpp b/tests/render_harness.cpp index d3a305d..42c97ec 100644 --- a/tests/render_harness.cpp +++ b/tests/render_harness.cpp @@ -14,15 +14,33 @@ // --tail S binaural tail seconds (default 5) // --ffmpeg PATH default ffmpeg // --max-frames N stop after N output frames (0 = all) +// --max-input-frames N stop feeding the renderer after N syncframes +// --seek-to S seek to S seconds before reading (repeatable; +// only the last one before reading takes effect) +// --seek-at N:S seek to S seconds once N output frames have been +// written (repeatable; N = 0 seeks before reading) +// --total-seconds S the duration seek() is told about (0 = unknown) +// --length S cut the delivered stream at S seconds (the file's +// own duration; the renderer still renders its tail) +// --pipeline-delay N delay a seek compensates for, in samples +// --preroll-frames N syncframes fed before the target frame +// --cycles N render the file N times in this process, in sequence +// --parallel N render it N times at once, in separate threads +// +// The two constants a seek uses are overridable so they can be checked by measurement. +// +// --cycles and --parallel cover what one render per process cannot: the renderer and the +// compiled-HRTF cache live as long as the process does, and foobar2000 starts the next +// file while the current one is still being read. // // It exists because the plugin's engine (src/joc_decode.*) has no foobar2000 -// dependency: the very code that plays in foobar2000 can be run here and its -// output compared against the reference renderer, which is what the bit-exactness -// contract is about. The WAV header mirrors the renderer's own writer: a simple -// IEEE-float header up to two channels, WAVE_FORMAT_EXTENSIBLE above that with a -// channel mask of zero. +// dependency: the very code that plays in foobar2000 can be run here, against the +// reference renderer or against the same engine's continuous render. The WAV header +// mirrors the renderer's own writer: a simple IEEE-float header up to two channels, +// WAVE_FORMAT_EXTENSIBLE above that with a channel mask of zero. #include +#include #include #include #include @@ -44,6 +62,16 @@ struct Options { double tail_seconds = 5.0; std::uint64_t max_frames = 0; std::uint64_t max_input_frames = 0; + std::vector seeks_before; // --seek-to + std::vector> seeks_at; // --seek-at + double total_seconds = 0.0; + double length_seconds = 0.0; // --length: cut the delivered stream here + std::uint32_t pipeline_delay = joc_decode::kJocSeekPipelineDelaySamples; + std::uint32_t preroll_frames = joc_decode::kJocSeekPrerollFrames; + bool container = false; + unsigned audio_index = 0; + unsigned cycles = 0; // --cycles N: N renders in one process, in sequence + unsigned parallel = 0; // --parallel N: N renders at once, in one process }; void put_u16(std::string* out, unsigned value) { @@ -111,6 +139,26 @@ bool parse(int argc, char** argv, Options* options) { else if (arg == "--tail") options->tail_seconds = std::atof(next().c_str()); else if (arg == "--max-frames") options->max_frames = std::strtoull(next().c_str(), nullptr, 10); else if (arg == "--max-input-frames") options->max_input_frames = std::strtoull(next().c_str(), nullptr, 10); + else if (arg == "--seek-to") options->seeks_before.push_back(std::atof(next().c_str())); + else if (arg == "--total-seconds") options->total_seconds = std::atof(next().c_str()); + else if (arg == "--length") options->length_seconds = std::atof(next().c_str()); + else if (arg == "--pipeline-delay") options->pipeline_delay = static_cast(std::strtoul(next().c_str(), nullptr, 10)); + else if (arg == "--preroll-frames") options->preroll_frames = static_cast(std::strtoul(next().c_str(), nullptr, 10)); + else if (arg == "--container") options->container = true; + else if (arg == "--cycles") options->cycles = static_cast(std::strtoul(next().c_str(), nullptr, 10)); + else if (arg == "--parallel") options->parallel = static_cast(std::strtoul(next().c_str(), nullptr, 10)); + else if (arg == "--audio-index") options->audio_index = static_cast(std::strtoul(next().c_str(), nullptr, 10)); + else if (arg == "--seek-at") { + const std::string value = next(); + const std::string::size_type colon = value.find(':'); + if (colon == std::string::npos) { + std::fprintf(stderr, "render_harness: --seek-at wants :\n"); + return false; + } + options->seeks_at.emplace_back( + std::strtoull(value.substr(0, colon).c_str(), nullptr, 10), + std::atof(value.substr(colon + 1).c_str())); + } else if (arg == "--help" || arg == "-h") return false; else { std::fprintf(stderr, "render_harness: unknown argument %s\n", arg.c_str()); @@ -120,47 +168,26 @@ bool parse(int argc, char** argv, Options* options) { return !options->input.empty() && !options->output.empty(); } -} // namespace - -int main(int argc, char** argv) { - Options options; - if (!parse(argc, argv, &options)) { - std::fprintf(stderr, - "usage: render_harness --input --output [--mode binaural|speaker]\n" - " [--layout NAME] [--hrtf-source sofa|rosella] [--hrtf PATH]\n" - " [--gain-db X] [--tail S] [--ffmpeg path] [--max-frames N]\n" - " [--max-input-frames N]\n"); - return 2; - } - - joc_decode::Settings settings; - settings.output = (options.mode == "speaker") ? joc_decode::Output::kSpeaker - : joc_decode::Output::kBinaural; - settings.speaker_layout = options.layout; - settings.hrtf_source = (options.hrtf_source == "rosella") ? joc_decode::HrtfSource::kRosella - : joc_decode::HrtfSource::kSofa; - settings.hrtf_file = options.hrtf; - settings.gain_db = options.gain_db; - settings.tail_seconds = options.tail_seconds; - settings.ffmpeg_path = options.ffmpeg; - settings.input_frame_limit = options.max_input_frames; +// One render: everything the component does for one file from start to stop. +// Returns false with *error set when the engine reports a failure. +bool render_once(const joc_decode::Settings& settings, const Options& options, + const std::string& output, std::string* error) { joc_decode::Engine engine; - std::string error; - if (!engine.start(options.input, settings, &error)) { - std::fprintf(stderr, "render_harness: engine start failed: %s\n", error.c_str()); - return 1; - } + if (!engine.start(options.input, settings, error)) return false; const unsigned channels = engine.channels(); if (channels == 0) { - std::fprintf(stderr, "render_harness: engine reported zero channels\n"); - return 1; + *error = "engine reported zero channels"; + return false; + } + for (const double seconds : options.seeks_before) { + if (!engine.seek(seconds, options.total_seconds, error)) return false; } - std::FILE* file = std::fopen(options.output.c_str(), "wb"); + std::FILE* file = std::fopen(output.c_str(), "wb"); if (file == nullptr) { - std::fprintf(stderr, "render_harness: cannot write %s\n", options.output.c_str()); - return 1; + *error = "cannot write " + output; + return false; } // The header carries the length, so write a placeholder and come back to it. const std::string header = wav_header(channels, 48000, 0); @@ -170,19 +197,27 @@ int main(int argc, char** argv) { std::vector buffer(kChunk * channels); std::uint64_t frames_written = 0; double peak = 0.0; + std::vector applied(options.seeks_at.size(), 0); for (;;) { + for (std::size_t i = 0; i < options.seeks_at.size(); ++i) { + if (applied[i] != 0 || options.seeks_at[i].first > frames_written) continue; + applied[i] = 1; + if (!engine.seek(options.seeks_at[i].second, options.total_seconds, error)) { + std::fclose(file); + return false; + } + } std::size_t want = kChunk; if (options.max_frames != 0) { if (frames_written >= options.max_frames) break; const std::uint64_t left = options.max_frames - frames_written; if (left < want) want = static_cast(left); } - const std::size_t frames = engine.read(buffer.data(), want, &error); + const std::size_t frames = engine.read(buffer.data(), want, error); if (frames == 0) { - if (!error.empty()) { - std::fprintf(stderr, "render_harness: read failed: %s\n", error.c_str()); + if (!error->empty()) { std::fclose(file); - return 1; + return false; } break; } @@ -201,9 +236,83 @@ int main(int argc, char** argv) { std::fwrite(final_header.data(), 1, final_header.size(), file); std::fclose(file); - std::printf("render_harness: %s -> %s\n", options.input.c_str(), options.output.c_str()); - std::printf(" channels=%u frames=%llu samples_per_channel=%llu peak=%.9f\n", channels, - static_cast(frames_written), + std::printf("render_harness: %s -> %s\n", options.input.c_str(), output.c_str()); + std::printf(" channels=%u frames=%llu peak=%.9f\n", channels, static_cast(frames_written), peak); - return 0; + return true; } + +std::string numbered(const std::string& output, unsigned index) { + return output + "." + std::to_string(index) + ".wav"; +} + +} // namespace + +int main(int argc, char** argv) { + Options options; + if (!parse(argc, argv, &options)) { + std::fprintf(stderr, + "usage: render_harness --input --output [--mode binaural|speaker]\n" + " [--layout NAME] [--hrtf-source sofa|rosella] [--hrtf PATH]\n" + " [--gain-db X] [--tail S] [--ffmpeg path] [--max-frames N]\n" + " [--max-input-frames N] [--seek-to S] [--seek-at N:S]\n" + " [--total-seconds S] [--pipeline-delay N] [--preroll-frames N]\n" + " [--container] [--audio-index N] [--length S]\n" + " [--cycles N] [--parallel N]\n"); + return 2; + } + + joc_decode::Settings settings; + settings.output = (options.mode == "speaker") ? joc_decode::Output::kSpeaker + : joc_decode::Output::kBinaural; + settings.speaker_layout = options.layout; + settings.hrtf_source = (options.hrtf_source == "rosella") ? joc_decode::HrtfSource::kRosella + : joc_decode::HrtfSource::kSofa; + settings.hrtf_file = options.hrtf; + settings.gain_db = options.gain_db; + settings.tail_seconds = options.tail_seconds; + settings.ffmpeg_path = options.ffmpeg; + settings.input_frame_limit = options.max_input_frames; + settings.pipeline_delay_samples = options.pipeline_delay; + settings.seek_preroll_frames = options.preroll_frames; + settings.length_seconds = options.length_seconds; + settings.input_kind = options.container ? joc_decode::InputKind::kContainer + : joc_decode::InputKind::kBare; + settings.audio_index = options.audio_index; + + std::string error; + if (options.parallel != 0) { + // The same file rendered by several engines at once, which is what a track + // change looks like from the engine's side: foobar2000 starts the next file + // while the current one is still being read. + std::vector threads; + std::vector errors(options.parallel); + for (unsigned i = 0; i < options.parallel; ++i) { + threads.emplace_back([&, i] { + if (!render_once(settings, options, numbered(options.output, i), &errors[i])) { + std::fprintf(stderr, "render_harness: parallel %u: %s\n", i, + errors[i].c_str()); + } + }); + } + for (std::thread& thread : threads) thread.join(); + for (const std::string& text : errors) { + if (!text.empty()) return 1; + } + return 0; + } + + const unsigned cycles = (options.cycles == 0) ? 1 : options.cycles; + for (unsigned cycle = 0; cycle < cycles; ++cycle) { + // Every cycle is a fresh engine in the same process: the compiled-HRTF cache + // and everything else the renderer keeps per process is reused, exactly as it + // is when a second file is played in the same foobar2000 session. + error.clear(); + const std::string output = (cycles == 1) ? options.output : numbered(options.output, cycle); + if (!render_once(settings, options, output, &error)) { + std::fprintf(stderr, "render_harness: cycle %u: %s\n", cycle, error.c_str()); + return 1; + } + } + return 0; +} \ No newline at end of file diff --git a/tools/deploy.ps1 b/tools/deploy.ps1 index 60e63b5..fd47d5d 100644 --- a/tools/deploy.ps1 +++ b/tools/deploy.ps1 @@ -4,15 +4,16 @@ # pwsh -File tools/deploy.ps1 -TestBed D:\fb2k -Platform x64 # # Where the component has to go depends on the foobar2000 layout, and getting it -# wrong is silent: foobar2000 simply never calls LoadLibrary on the file. Two -# layouts are in the wild: +# wrong is silent: foobar2000 simply never calls LoadLibrary on the file. # -# profile-relative \profile\user-components\\.dll -# foobar2000 1.6.19 portable and 2.x -# app-relative \user-components\\.dll -# older / repacked 1.6 installs whose profile is \configuration +# 1.6 \profile\user-components\\.dll +# (portable mode; without it the profile is in %APPDATA%) +# 2.0 and up \user-components\\.dll +# even in portable mode, whose profile is \profile +# older 1.6 \user-components\\.dll +# repacked installs whose profile is \configuration # -# The per-component subdirectory is required in both: a DLL lying directly in +# The per-component subdirectory is required in every case: a DLL lying directly in # user-components\ is not scanned. Installing a .fb2k-component package through # foobar2000 itself ends up doing the same thing. [CmdletBinding()] @@ -31,11 +32,16 @@ if (-not (Test-Path -LiteralPath (Join-Path $TestBed 'foobar2000.exe'))) { throw "no foobar2000.exe in $TestBed" } +# The core version decides the layout, and its own profile records it. +$versionFile = Join-Path $TestBed 'profile\version.txt' +$version = if (Test-Path -LiteralPath $versionFile) { (Get-Content -LiteralPath $versionFile -Raw).Trim() } else { '' } + if (-not $Target) { $appRelative = Test-Path -LiteralPath (Join-Path $TestBed 'configuration') + if ($version -match 'v(\d+)\.') { $appRelative = ([int]$Matches[1] -ge 2) } $root = if ($appRelative) { $TestBed } else { Join-Path $TestBed 'profile' } $Target = Join-Path $root 'user-components' - Write-Host ("layout: {0} ({1})" -f $(if ($appRelative) { 'app-relative' } else { 'profile-relative' }), $root) + Write-Host ("layout: {0} ({1}){2}" -f $(if ($appRelative) { 'app-relative' } else { 'profile-relative' }), $root, $(if ($version) { ", $version" } else { '' })) } # Portable profile: config stays inside the test bed instead of %APPDATA%.