4 Commits

Author SHA1 Message Date
TheM14 4d8118684f Fix MP4 probe when moov exceeds the window
build / windows (push) Has been cancelled
build / release (push) Has been cancelled
2026-09-27 15:58:23 +08:00
TheM14 bb464b2c03 Fix bed decoded off the E-AC-3 frame grid
build / windows (push) Has been cancelled
build / release (push) Has been cancelled
2026-09-27 14:46:12 +08:00
TheM14 71bec4bda3 Fix band-0 DC filter gating and improve settings UI
build / windows (push) Has been cancelled
build / release (push) Has been cancelled
2026-09-27 01:24:53 +08:00
TheM14 65f56990aa Fix missing tags, broken seeking, and the x64 playback crash
build / windows (push) Has been cancelled
build / release (push) Has been cancelled
2026-09-26 03:14:02 +08:00
17 changed files with 1244 additions and 449 deletions
+3
View File
@@ -29,3 +29,6 @@ rosella_kernels.npz
# The component writes its own log next to the DLL
joc_decoder.log
# Measurement record: local working notes, kept on disk and never published.
VERIFICATION.md
+121
View File
@@ -0,0 +1,121 @@
# foo_input_joc — foobar2000 input component for E-AC-3 JOC
[简体中文](README.md) · [Install](#install) · [Settings](#settings) · [Building](#building) · [Known limitations](#known-limitations)
A foobar2000 input component for **E-AC-3 JOC (Dolby Atmos)** files: the JOC object and OAMD
metadata is taken from the E-AC-3 syncframes and paired with the 5.1 core PCM that ffmpeg
decodes, then rendered in real time to headphones (HRTF) or to a speaker layout up to 7.1. The
rendering core is compiled into the component, so there is nothing to install beside
`foo_input_joc.dll`.
Input is a bare `.eac3` / `.ec3` stream, or an E-AC-3 JOC track inside a container (`.mp4`
`.m4a` `.m4b` `.m4p` `.m4r` `.mov` `.mkv` `.mka` `.webm`). A file without JOC is not claimed:
plain AC-3 / E-AC-3, a container whose audio track is AAC, and transport streams are all handed
back to foobar2000 and play through the decoder it would have used anyway.
## Requirements
* foobar2000 1.6 (32-bit) or 2.x (32-bit and 64-bit).
* An `ffmpeg` executable, found on `PATH` by default; the preferences page can name one.
* For binaural output, an HRTF data file. **This repository does not ship one** — see
[HRTF data](#hrtf-data).
## Install
Download the package for your architecture from [Releases](../../releases) — a pushed `v*` tag
makes CI attach both — or build it yourself as described under [Building](#building) and take the
result from `dist\`; either way the file is `foo_input_joc-<version>-<arch>.fb2k-component`. Drop
it onto foobar2000, or use Preferences → Components → Install, and restart. Take `-x86` for
foobar2000 1.6 and 2.x 32-bit, `-x64` for 2.x 64-bit.
To install by hand, copy `foo_input_joc.dll` into the per-component subdirectory of the
`user-components` folder the running version reads:
| foobar2000 | Folder |
|---|---|
| 1.6 | `<profile>\user-components\foo_input_joc\` |
| 2.x | `<app>\user-components\foo_input_joc\` (also in portable mode) |
The subdirectory is required — a DLL lying directly in `user-components\` is not scanned — and a
component placed in the folder the running version does not read, or built for the other
architecture, is **ignored without a message**.
`tools/deploy.ps1 -TestBed <portable foobar2000>` is the scripted form of the manual steps, and
`tools/run.ps1 -TestBed <path> -Play <file>` plays a file unattended and prints the log. Let
foobar2000 exit through `/exit`: a force-killed instance leaves a `<profile>\running` marker
behind, and the next start then refuses to load any user component.
### Containers need one look at the decoder list
foobar2000 asks the decoders in the order shown in Preferences → **Decoding**, and the built-in
container readers are in that list. When one of them is offered an MP4 or Matroska file first, it
takes the file and the JOC objects are lost — such a file then plays as plain E-AC-3.
So move **JOC decoder (E-AC-3 JOC)** above **foobar2000 MP4 Demuxer** and **foobar2000
Matroska/WebM Reader** in that list. Bare `.eac3` / `.ec3` files do not depend on the order.
## Settings
Preferences → Tools → **JOC decoder**:
* **Output** — binaural, or a speaker layout from 2.0 to 7.1;
* **Binaural mode** (near / mid / far) and the room **tail** in seconds;
* **HRTF source** — a **SOFA** file or a **Rosella** `.personalized_headphone` model. An empty
path means the default location, `<component directory>\HRTF\binaural.sofa` or
`binaural.personalized_headphone`;
* **Gain** — a switch and a value in dB. Binaural rendering can exceed full scale on material
that does not clip in the core mix, so attenuation belongs here;
* the **ffmpeg** executable to use.
No control on the page is disabled, and the status line states what is in effect.
### HRTF data
A SOFA measurement set or a personalised headphone model is supplied by whoever runs the
component, and is listed in `.gitignore` so it cannot be committed by accident. Speaker layouts
need none. Binaural rendering without an HRTF fails with a message naming the file it looked for.
## Building
```powershell
pwsh -File tools/setup_sdk.ps1 # official SDK into SDK/, pinned to target 1.5/1.6
pwsh -File tools/build.ps1 # Win32 -> build\Win32\foo_input_joc.dll
pwsh -File tools/build.ps1 -Platform x64
pwsh -File tools/package.ps1 # both, packaged into dist\*.fb2k-component
```
The configuration is fixed at `Release-Static` (static CRT, `/MT`), and `/fp:precise` is what the
byte-for-byte acceptance rests on, so it must not be changed. `foo_input_joc.vcxproj` builds
`kernel\joc_kernel.vcxproj` first through a project reference; the rendering core sources in
`kernel/` are compiled with `JOC_STATIC` / `EJOC_STATIC`, so their entry points are neither
imported nor exported. `tests\` holds the offline tools (bitstream self-test and cross-check,
render comparison, preferences-page layout check, container-probe check) and `tools\` the build
and test-bed scripts. Settings also read `JOC_*` environment overrides for one run (development
only); the list and what each one does is in `src\settings.cpp`.
To diagnose a problem, read `joc_decoder.log` beside the DLL: the component writes its own
version, the core version and the log path there at start-up.
## Known limitations
* ADM BWF output is not implemented.
* A container is only claimed when this component is ahead of the built-in container reader in
Preferences → Decoding (see [Install](#containers-need-one-look-at-the-decoder-list)); the core
does not let a decoder ask for a file another entry has already taken. Such a file's tags stay
that reader's as well.
* Tags of a bare `.eac3` / `.ec3` (an ID3v2 tag in front of the stream, or an APEv2/ID3v1 tag
behind it) can be **read but not written**: no component claims raw E-AC-3 for writing, and
rewriting the whole file to insert a tag is not this component's job.
* Transport streams (`.ts`, `.m2ts`) are not claimed.
* Playback length is exactly the file's duration. The binaural renderer still computes its room
tail, but it is not delivered as playback time the file does not have.
* A seek re-enters the bitstream at the frame holding the target instead of decoding everything
in front of it — which is why the cost of a seek does not depend on where it lands. The
position is exact and does not drift; the samples are the same waveform handed to a decoder
that started there, so they differ from a straight play-through by a low-level noise floor of
−59 dBFS or quieter.
## Licence
`LICENSE` is the upstream MIT licence, copied unchanged; `kernel/` is a copy of the upstream
rendering core sources and keeps their notices. See [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md).
+83 -123
View File
@@ -1,143 +1,103 @@
# foo_input_joc
# foo_input_joc — foobar2000 的 E-AC-3 JOC 输入组件
foobar2000 input component for **E-AC-3 JOC (Dolby Atmos)** files: the JOC objects are
rendered to binaural (HRTF) or to a speaker layout up to 7.1, in real time.
[English](README.en.md) · [安装](#安装) · [设置](#设置) · [构建](#构建) · [已知限制](#已知限制)
Two files with the same name pay for the whole thing: `joc_core`'s C++ sources are copied
into [`kernel/`](kernel/) and compiled straight into the component, so there is nothing to
install beside `foo_input_joc.dll`.
foobar2000 输入组件,播放 **E-AC-3 JOC(Dolby Atmos)** 文件:从 E-AC-3 同步帧里取出 JOC
对象与 OAMD 元数据,与 ffmpeg 解出的 5.1 核心 PCM 配对后实时渲染,输出双耳(HRTF)或最多
7.1 的扬声器布局。渲染内核直接编译进组件,除 `foo_input_joc.dll` 之外不需要安装任何东西。
## What it does
输入是裸流 `.eac3` / `.ec3`,或容器里的 E-AC-3 JOC 轨道(`.mp4` `.m4a` `.m4b` `.m4p`
`.m4r` `.mov` `.mkv` `.mka` `.webm`)。文件里没有 JOC 就不接管:普通 AC-3 / E-AC-3、音频轨道
是 AAC 的容器、传输流都交还 foobar2000,由它原本的解码器播放。
1. Finds the audio: a bare `.eac3` / `.ec3` stream is read as it is, while a container
(`.mp4`, `.m4a`, `.m4b`, `.m4p`, `.m4r`, `.mov`, `.mkv`, `.mka`, `.webm`) is looked into
first — the container's own headers say whether an E-AC-3 track is present and which
audio track it is (MP4 sample entry `ec-3`, Matroska `CodecID A_EAC3`), and ffmpeg then
copies that track out of the file byte for byte. The header walk is bounded and cheap, so
an MP4 holding AAC is declined without starting anything.
2. Decides from the bitstream whether it really carries JOC (an EMDF container holding both
the OAMD and the JOC payload — container metadata only ever says "E-AC-3", and the JOC
flag inside it is frequently missing).
3. A file with no E-AC-3 track, or with one that carries no JOC, is handed back to
foobar2000 with `exception_io_unsupported_format`, so the built-in decoder plays it — this
component never decodes plain AC-3 or E-AC-3.
4. A JOC file is decoded as: the syncframes go to the renderer as metadata, the 5.1 core
PCM comes from ffmpeg, and the renderer pairs them (one syncframe : 1536 bed samples)
and produces the output PCM, which is handed back to foobar2000.
## 环境要求
```
.eac3 / .ec3 file container (.mp4 .mkv .m4a ...)
│ ├─ header walk (src/container_scan.cpp) ── no E-AC-3 ──▶ next decoder
│ └─ E-AC-3 track ── ffmpeg -c:a copy ──▶ syncframes
├─ JOC check (src/eac3_scan.cpp) ─── no JOC ──▶ built-in E-AC-3 decoder
└─ JOC
├─ syncframes ────────────────────▶ renderer metadata
└─ ffmpeg -ac 6 -c:a pcm_f32le ───▶ 5.1 core PCM ──▶ renderer bed
│
▼
2 ch or ≤7.1 PCM ──▶ foobar2000
```
* foobar2000 1.6(32 位)或 2.x(32 位与 64 位)。
* 一个 `ffmpeg` 可执行文件,默认从 `PATH` 找,设置页可以指定路径。
* 双耳输出需要一份 HRTF 数据文件。**本仓库不附带**,见 [HRTF 数据](#hrtf-数据)。
## Repository layout
## 安装
| Path | Contents |
从 [Releases](../../releases) 下载对应架构的包(`v*` 标签触发的 CI 会把两个架构都附上去),
或按[构建](#构建)自行打出 `dist\` 下的产物;文件名都是
`foo_input_joc-<版本>-<架构>.fb2k-component`。把它拖到 foobar2000 上,或用
Preferences → Components → Install 安装,然后重启。1.6 和 2.x 32 位用 `-x86`,2.x 64 位用
`-x64`。
手工安装就把 `foo_input_joc.dll` 放进当前版本会读取的 `user-components` 子目录:
| foobar2000 | 目录 |
|---|---|
| `kernel/` | Copy of the `joc_core` C++ sources (`include/` + `src/`) and `joc_kernel.vcxproj`, the static library the component links |
| `src/eac3_scan.*` | Syncframe walk and the JOC bitstream test |
| `src/container_scan.*` | Bounded header walk of MP4/MOV and Matroska: is there an E-AC-3 track, which one, and how long is the file |
| `src/joc_decode.*` | Decode engine: starts ffmpeg, drives the renderer, handles the end of stream. No foobar2000 headers, so it also builds into the offline tools |
| `src/input_joc.cpp` | The foobar2000 input: format recognition, yielding, `get_info`, `initialize`, `run` |
| `src/settings.*` | Configuration values and their environment overrides (development only) |
| `src/prefs.cpp`, `src/prefs.rc` | The preferences page |
| `src/log.*` | Diagnostic log written next to the DLL |
| `tests/` | Offline tools: bitstream self-test and cross-check against the renderer, render harness, preferences-page layout check, container-probe check |
| `tools/` | SDK fetch, build, package, deploy, unattended test bed run |
| 1.6 | `<profile>\user-components\foo_input_joc\` |
| 2.x | `<app>\user-components\foo_input_joc\`(portable 模式同样如此) |
## Build
子目录是必须的——DLL 直接躺在 `user-components\` 下不会被扫描;放进当前版本不读的目录,或
架构不对,都会**静默忽略**,不会有任何提示。
`tools/deploy.ps1 -TestBed <portable foobar2000>` 是上面手工步骤的脚本版;
`tools/run.ps1 -TestBed <路径> -Play <文件>` 可以无人值守播放并把日志打出来。让 foobar2000
通过 `/exit` 正常退出:被强杀的实例会在 `<profile>\running` 留下标记,下次启动会拒绝加载任何
用户组件。
### 容器需要调一次解码器顺序
foobar2000 按 Preferences → **Decoding** 里的顺序询问解码器,内置的容器读取器也在那张表里。
如果它排在前面接到 MP4 / Matroska 文件,文件就被它拿走,JOC 对象随之丢失——于是听起来只是
普通 E-AC-3。
所以要把 **JOC decoder (E-AC-3 JOC)** 提到 **foobar2000 MP4 Demuxer** 与 **foobar2000
Matroska/WebM Reader** 之前。裸流 `.eac3` / `.ec3` 不受这个顺序影响。
## 设置
Preferences → Tools → **JOC decoder**:
* **Output** —— 双耳,或 2.0 到 7.1 的扬声器布局;
* **Binaural mode**(near / mid / far)与房间 **tail** 秒数;
* **HRTF source** —— **SOFA** 文件或 **Rosella** `.personalized_headphone` 模型。路径留空表示
用默认位置 `<组件目录>\HRTF\` 下的 `binaural.sofa` 或 `binaural.personalized_headphone`;
* **Gain** —— 开关加 dB 值。双耳渲染在核心混音不削顶的素材上也可能超过满刻度,衰减放在这里;
* **ffmpeg** 可执行文件路径。
页面上的控件都不禁用,状态行会说明当前生效的是什么。
### HRTF 数据
SOFA 测量集或个性化耳机模型由使用组件的人自己提供,并且写在 `.gitignore` 里,避免误提交。
扬声器布局不需要 HRTF。双耳渲染缺少 HRTF 时会报错并指出它找的是哪个文件。
## 构建
```powershell
pwsh -File tools/setup_sdk.ps1 # official SDK into SDK/, pinned to target 1.5/1.6
pwsh -File tools/setup_sdk.ps1 # 官方 SDK 拉进 SDK/,固定到 target 1.5/1.6
pwsh -File tools/build.ps1 # Win32 -> build\Win32\foo_input_joc.dll
pwsh -File tools/build.ps1 -Platform x64
pwsh -File tools/package.ps1 # both, packaged into dist\*.fb2k-component
pwsh -File tools/package.ps1 # 两个架构,打包到 dist\*.fb2k-component
```
`Release-Static` uses the static CRT (`/MT`); `/fp:precise` is required and must not be
changed. `foo_input_joc.vcxproj` builds `kernel\joc_kernel.vcxproj` first through a project
reference. The copied kernel sources are compiled with `JOC_STATIC` / `EJOC_STATIC` so their
entry points are neither imported nor exported.
配置固定为 `Release-Static`(静态 CRT,`/MT`);`/fp:precise` 是逐字节验收的前提,不要改。
`foo_input_joc.vcxproj` 通过项目引用先构建 `kernel\joc_kernel.vcxproj`;`kernel/` 里的渲染内核
源码以 `JOC_STATIC` / `EJOC_STATIC` 编译,入口既不导入也不导出。`tests\` 是离线工具(码流
自检与交叉核对、渲染比对、设置页布局检查、容器探测检查),`tools\` 是构建与测试床脚本。设置项
另有 `JOC_*` 环境变量覆盖(仅用于开发运行),清单与含义在 `src\settings.cpp`。
## Install
排查问题看 DLL 旁边的 `joc_decoder.log`;组件启动时会把自己的版本、核心版本、日志路径写在
里面。
Either drop `dist\foo_input_joc-<version>-<arch>.fb2k-component` onto foobar2000 (or use
Preferences → Components → Install), or copy `foo_input_joc.dll` into
`<profile>\user-components\foo_input_joc\`. The per-component subdirectory is required:
a DLL lying directly in `user-components\` is not scanned. 1.6 is 32-bit, 2.x ships both,
and a DLL of the wrong architecture is silently ignored.
## 已知限制
`tools/deploy.ps1 -TestBed <path to portable foobar2000>` does the manual variant, and
`tools/run.ps1 -TestBed <path> -Play <file>` runs it unattended and prints the log.
Always let foobar2000 exit through `/exit`; a force-killed instance leaves a
`<profile>\running` marker behind and the next start then refuses to load any user
component.
* 不实现 ADM BWF 输出。
* 容器只有在它排在内置容器读取器之前时才会被接管(见[安装](#容器需要调一次解码器顺序));核心
不允许某个解码器去要一个已经被别的条目拿走的文件。这类文件的标签也仍旧归那个读取器。
* 裸流 `.eac3` / `.ec3` 的标签(流前面的 ID3v2,或后面的 APEv2/ID3v1)**能读不能写**:没有
组件声明可以写裸 E-AC-3,为插入标签重写整个文件也不是本组件该做的事。
* 传输流(`.ts`、`.m2ts`)不接管。
* 播放长度严格等于文件时长。双耳渲染器仍会算出房间尾音,但它不作为文件本身没有的播放时间交付。
* 跳转会从包含目标位置的那个帧重新进入码流,而不是把前面的内容全部解码一遍——这是跳转代价与
目标位置无关的原因。位置精确、不漂移;样本是同一段波形交给了从该处开始的解码器,与从头播放
相比差一个很低的噪声底(−59 dBFS 或更低)。
### Containers need one look at the decoder list
## 许可
foobar2000 tries the decoders in the order shown in Preferences → **Decoding** (the
"list of available decoders", where entries can be moved up and down). The built-in
container readers are in that list too, and when one of them is offered an MP4 or Matroska
file before this component, it takes the file and the JOC objects are lost — the file plays
as plain E-AC-3.
So, to play JOC from a container, move **JOC decoder (E-AC-3 JOC)** above **foobar2000 MP4
Demuxer** and **foobar2000 Matroska/WebM Reader** in that list. Nothing else is needed, and
bare `.eac3` / `.ec3` files are unaffected by the order. This is the same thing every
third-party decoder (the FFmpeg wrapper, for one) asks for, which is why the component does
not try to work around it. If a container still plays as plain E-AC-3, that list is where to
look.
## Settings
Preferences → Tools → **JOC decoder**:
* **Output** — binaural, or a speaker layout from 2.0 to 7.1;
* **Binaural mode** (near / mid / far) and the room **tail** in seconds;
* **HRTF source** — a **SOFA** file or a **Rosella** `.personalized_headphone` model. Leave
the path empty to use the default location `<component directory>\HRTF\`:
`binaural.sofa` or `binaural.personalized_headphone`;
* **Gain** — a switch plus a value in dB. Binaural rendering can exceed full scale on
material that does not clip in the core mix, so attenuation belongs here;
* the **ffmpeg** executable to use.
Nothing on the page is disabled; the status line states what is in effect.
**HRTF data is not distributed with this repository.** A SOFA measurement set or a
personalised headphone model is supplied by whoever runs the component (and is listed in
`.gitignore` so it cannot be committed by accident). Speaker layouts and every offline test
except binaural rendering work without one; binaural rendering without an HRTF fails with a
message naming the file it looked for.
## Environment overrides
Development only: they override the stored settings for one run and every use is logged.
`JOC_OUTPUT`, `JOC_LAYOUT`, `JOC_HRTF`, `JOC_HRTF_SOURCE`, `JOC_BINAURAL_MODE`, `JOC_GAIN_DB`,
`JOC_GAIN_ENABLED`, `JOC_TAIL_SECONDS`, `JOC_OBJECT_DELAY`, `JOC_THREADS`, `JOC_FFMPEG`,
`JOC_LOG`.
## Known limitations
* ADM BWF output is not implemented.
* A container is only claimed when this component is ahead of the built-in container reader
in Preferences → Decoding, as described under [Install](#install); the core does not let a
decoder ask for a file another entry has already taken.
* Transport streams (`.ts`, `.m2ts`) are not claimed.
* The room tail is returned in full; the reference command-line renderer additionally trims
trailing samples below a threshold, so its output can be shorter.
* x86 and x64 do not produce bit-identical binaural output (last-bit differences): the
renderer's SIMD dispatch only applies to x86-64/ARM64, so 32-bit builds take the scalar
path. The speaker path is bit-identical on both.
## Licence
`LICENSE` is the upstream MIT licence, copied unchanged; `kernel/` is a copy of the upstream
renderer sources and keeps their notices. See [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md).
`LICENSE` 是上游 MIT 许可,原样复制;`kernel/` 是上游渲染内核源码的副本,保留其声明。见
[THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md)。
-144
View File
@@ -1,144 +0,0 @@
# Verification
Measured results only. Each entry names the command or the log line it came from, so it can
be reproduced. Environment: Windows x64 host, official **foobar2000 1.6.19 x86** portable
installation, official SDK 2026-09-17 pinned to `FOOBAR2000_TARGET_VERSION 80`, MSVC 14.44,
ffmpeg 8.0.
Test material is supplied locally and is **not** part of this repository: the `testdata/` and
`vectors/` files of the upstream renderer project, and — for the binaural measurements — an
HRTF file. Everything except binaural rendering runs without any HRTF; binaural runs take the
file as an argument (`tests/render_harness.cpp --hrtf …`) or use the default location beside
the DLL.
## Component and renderer
| Check | Result |
|---|---|
| Sources compiled in, nothing loaded at run time | log: `core: in-process renderer 0.1.0-m1 (abi 3), component built against abi 3` |
| One artefact, no companion DLL | `dist\*.fb2k-component` holds `foo_input_joc.dll` and `README.md` only |
| Kernel sources untouched | the upstream working tree's file timestamps are unchanged; it is only ever read |
| Both architectures build | `build\Win32\foo_input_joc.dll`, `build\x64\foo_input_joc.dll` |
## Playback in foobar2000 1.6.19 x86
| Case | Log evidence |
|---|---|
| Speaker 7.1, 5 s file | `stream created, 8 output channel(s), layout=7.1` … `frames_in=157 frames_out=157 samples_out=241152` — 157 × 1536 exactly |
| Binaural, SOFA | `stream created, 2 output channel(s)` … `end of stream after 481215 frames` (241152 source + 240063 tail) |
| Binaural, Rosella model | `stream created, 2 output channel(s)` … `end of stream after 481855 frames` |
| Binaural, HRTF path left empty | resolves to `<component directory>\HRTF\binaural.sofa` and produces the same 481215 frames as naming that file explicitly |
| Full 238 s file, binaural SOFA | `eac3 frames queued=7436, bed frames pushed=7436` … `samples_out=11661759`; no ffmpeg process left behind |
| Installed from the `.fb2k-component` package | unpacked into `user-components\foo_input_joc\`, plays with the default-folder HRTF |
## Bitstream recognition
`tests/scan_crosscheck.cpp` compares, frame by frame, the syncframe lengths and the JOC
verdict this component computes against the renderer's own `joc_eac3_frame_bytes()` /
`joc_parse_eac3_frame()`:
```
testdata\gold_forever.eac3 frames=64 plugin_joc=64 kernel_joc=64 len_mismatch=0 verdict_mismatch=0
vectors\valid.eac3 frames=7 plugin_joc=7 kernel_joc=7 len_mismatch=0 verdict_mismatch=0
build\plain_eac3.eac3 frames=64 plugin_joc=0 kernel_joc=0 len_mismatch=0 verdict_mismatch=0
crosscheck: 3 file(s), AGREES WITH CORE
```
Ten corrupt vectors were compared as well: no frame-length disagreement, verdicts agreed on
9 of 10. The one difference is `corrupt_truncated_huffman.eac3`: this component only tests
+9
View File
@@ -68,6 +68,15 @@ EJOC_API ejoc_renderer_handle EJOC_CALL ejoc_renderer_create(void);
EJOC_API void EJOC_CALL ejoc_renderer_destroy(ejoc_renderer_handle handle);
EJOC_API int EJOC_CALL ejoc_renderer_reset(ejoc_renderer_handle handle);
EJOC_API int EJOC_CALL ejoc_renderer_set_threads(ejoc_renderer_handle handle, uint32_t total_threads);
/*
Enables or disables the Ls/Rs band-0 21-tap DC compensation. The caller derives
it from the JOC downmix configuration: only configurations 3 and 4 enable the
filter. When disabled, band 0 keeps the common per-band processing (surround
delay plus -j rotation) instead of being overwritten by the FIR. The delay line
and DC history advance either way, so the flag may change between frames.
Defaults to enabled when never called.
*/
EJOC_API int EJOC_CALL ejoc_renderer_set_dc_filter(ejoc_renderer_handle handle, uint32_t enabled);
EJOC_API uint32_t EJOC_CALL ejoc_renderer_thread_count(ejoc_renderer_handle handle);
EJOC_API const char* EJOC_CALL ejoc_renderer_last_error(ejoc_renderer_handle handle);
+25 -9
View File
@@ -125,6 +125,10 @@ public:
return error_[0] ? error_ : "";
}
void set_dc_filter(const bool enabled) noexcept {
dc_filter_enabled_ = enabled;
}
int process(
const float* bed5,
const float* lfe,
@@ -305,16 +309,18 @@ private:
for (int i = 0; i < 4; ++i) {
dc_buffer[20 + i] = current[i][0];
}
for (int slot = 0; slot < 4; ++slot) {
Complex sum{0.0, 0.0};
for (int tap = 0; tap < 21; ++tap) {
const Complex sample = dc_buffer[slot + tap];
const double cr = kDcB[tap];
const double ci = kDcA[tap];
sum.re += sample.re * cr - sample.im * ci;
sum.im += sample.re * ci + sample.im * cr;
if (dc_filter_enabled_) {
for (int slot = 0; slot < 4; ++slot) {
Complex sum{0.0, 0.0};
for (int tap = 0; tap < 21; ++tap) {
const Complex sample = dc_buffer[slot + tap];
const double cr = kDcB[tap];
const double ci = kDcA[tap];
sum.re += sample.re * cr - sample.im * ci;
sum.im += sample.re * ci + sample.im * cr;
}
x_[channel][0][group + slot] = {2.0 * sum.re, 2.0 * sum.im};
}
x_[channel][0][group + slot] = {2.0 * sum.re, 2.0 * sum.im};
}
for (int i = 0; i < 20; ++i) {
surround_history_[surround][i] = dc_buffer[i + 4];
@@ -641,6 +647,8 @@ private:
float analysis_phase_;
alignas(64) Complex surround_delay_[2][10][64];
alignas(64) Complex surround_history_[2][20];
// band-0 的 21-tap DC 补偿开关;仅 downmix 配置 3/4 由调用方置位。
bool dc_filter_enabled_ = true;
alignas(64) double lfe_delay_[kLfeDelay];
alignas(64) double matrix_previous_[15][5][64];
alignas(64) double synthesis_state_[15][640];
@@ -696,6 +704,14 @@ int EJOC_CALL ejoc_renderer_set_threads(ejoc_renderer_handle handle, uint32_t to
return static_cast<ejoc::Renderer*>(handle)->set_threads(total_threads);
}
int EJOC_CALL ejoc_renderer_set_dc_filter(ejoc_renderer_handle handle, uint32_t enabled) {
if (!handle) {
return -1;
}
static_cast<ejoc::Renderer*>(handle)->set_dc_filter(enabled != 0);
return 0;
}
uint32_t EJOC_CALL ejoc_renderer_thread_count(ejoc_renderer_handle handle) {
if (!handle) {
return 0;
+9
View File
@@ -52,6 +52,15 @@ Status rebuild_objects16(ejoc_renderer_handle handle, const joc_frame_params& pa
}
out16->assign(static_cast<std::size_t>(JOC_OUTPUT_CHANNELS) * JOC_FRAME_SAMPLES, 0.0f);
// band-0 的 21-tap DC 补偿只在 downmix 配置 3/4 下启用,其余配置 band 0
// 走与其他 band 相同的处理。
const bool dc_filter = params.dmx_config_idx == 3 || params.dmx_config_idx == 4;
if (ejoc_renderer_set_dc_filter(handle, dc_filter ? 1u : 0u) != 0) {
if (error != nullptr) {
*error = "ejoc_renderer_set_dc_filter failed";
}
return Status::fail(JOC_ERR_RENDER_FAILED, stage::kDsp, "ejoc_renderer_set_dc_filter failed");
}
const int result = ejoc_renderer_process(
handle, bed5_planar, lfe, params.present_mask, n_bands, n_dpoints, slope_idx, offset_ts,
dq.data(), params.clipgain, 0.0625f, gain, out16->data());
+17 -2
View File
@@ -200,15 +200,24 @@ Result scan_mp4(const Window& file, std::size_t max_bytes) {
Result result;
result.kind = Kind::kMp4;
// moov is usually at the start for streamed files and at the end otherwise.
// moov is usually at the start for streamed files and at the end otherwise. The
// walk reads box headers where they lie and steps over mdat in one go, so the
// whole file is walked from the top: a fixed prefix window cannot reach a moov
// that is larger than the window, which is what an MP4 with its cover art stored
// as a video track produces (a 4.9 MB moov against a 4 MiB window). The tail
// range stays as a fallback for a file whose leading boxes do not parse.
struct Range {
std::uint64_t begin;
std::uint64_t end;
};
std::vector<Range> ranges{{0, (std::min<std::uint64_t>)(file.size(), max_bytes)}};
std::vector<Range> ranges{{0, file.size()}};
if (file.size() > max_bytes) {
ranges.push_back({file.size() - max_bytes, file.size()});
}
// An oversized moov is a file this probe cannot describe, and walking it would be
// unbounded work; it is reported instead of entered.
constexpr std::uint64_t kMaxMoovBytes = 64ull * 1024ull * 1024ull;
bool oversized_moov = false;
unsigned audio_seen = 0;
bool found_any_audio = false;
@@ -218,6 +227,10 @@ Result scan_mp4(const Window& file, std::size_t max_bytes) {
[&](const std::string& type, std::uint64_t payload, std::uint64_t box_end,
std::size_t) {
if (type != "moov") return true;
if (box_end - payload > kMaxMoovBytes) {
oversized_moov = true;
return true;
}
walk_boxes(file, payload, box_end, 1,
[&](const std::string& inner, std::uint64_t inner_payload,
std::uint64_t inner_end, std::size_t) {
@@ -273,6 +286,8 @@ Result scan_mp4(const Window& file, std::size_t max_bytes) {
result.detail = "mp4: first audio track is " +
(first_audio_format.empty() ? std::string("unknown")
: first_audio_format);
} else if (oversized_moov) {
result.detail = "mp4: moov is larger than the scan limit";
} else {
result.detail = "mp4: no audio track found in the scanned window";
}
+289 -29
View File
@@ -11,9 +11,11 @@
#include <SDK/audio_chunk.h>
#include <SDK/exception_io.h>
#include <SDK/file_info.h>
#include <SDK/file_info_impl.h>
#include <SDK/input.h>
#include <SDK/input_file_type.h>
#include <SDK/input_impl.h>
#include <SDK/tag_processor.h>
#include <cstring>
#include <string>
@@ -31,6 +33,18 @@ constexpr std::size_t kSniffBytes = 256u * 1024u;
constexpr std::size_t kRunFrames = 4096u;
constexpr unsigned kSampleRate = 48000;
// Largest magnitude in a block, for the delivery check in decode_run().
template <typename Sample>
double peak_of(const Sample* values, std::size_t count) {
double peak = 0.0;
for (std::size_t i = 0; i < count; ++i) {
const double value =
values[i] < Sample(0) ? -static_cast<double>(values[i]) : static_cast<double>(values[i]);
if (value > peak) peak = value;
}
return peak;
}
// Identity in the decoder priority table.
const GUID g_decoder_guid = {0x9c3f1d58, 0x27ab, 0x4e64, {0xb0, 0x93, 0x5e, 0x1c, 0xd7, 0x48, 0x2f, 0xa6}};
@@ -52,13 +66,131 @@ std::string file_name_of(const std::string& path) {
return slash == std::string::npos ? path : path.substr(slash + 1);
}
// ---------------------------------------------------------------------------
// Tags of a file this component has taken over.
//
// An MP4/M4A keeps its tags in its own metadata box, and the component that knows
// how to read and write them is the container reader the core already ships.
// Claiming a file for decoding must not take it away from that reader, and the SDK
// has no "decode with me, ask someone else for tags" arrangement -- whichever
// entry answers open() answers for everything. So the information read and write
// paths are forwarded to whichever other entry claims the file, and only the tags
// of its answer are merged into ours: the technical information stays this
// component's own, which is what tells a user the file is JOC rather than plain
// E-AC-3.
// ---------------------------------------------------------------------------
// Entries other than this one that claim the path, in the user's own decoding
// order. Ourselves is never in the list: an open forwarded back here would enter
// open() again, for ever.
void forwarding_candidates(const char* url, pfc::list_t<input_entry::ptr>& out) {
out.remove_all();
input_manager_v3::ptr manager;
if (input_manager_v3::tryGet(manager)) {
manager->get_enabled_inputs(out);
} else {
input_entry::g_find_inputs_by_path_ex(out, url,
[](input_entry::ptr) { return true; });
}
const char* dot = std::strrchr(url, '.');
const char* extension = (dot != nullptr) ? dot + 1 : "";
const GUID self = g_decoder_guid;
for (t_size index = out.get_count(); index-- > 0;) {
input_entry::ptr entry = out[index];
if (entry->get_guid_() == self || !entry->is_our_path(url, extension)) {
out.remove_by_idx(index);
}
}
}
// Opens the file again through another entry, for information reading or writing.
// The file is left unopened on our side, so the other entry can have it to itself.
template <typename t_interface>
bool open_forwarded(service_ptr_t<t_interface>& out, const GUID& what_for, const char* url,
abort_callback& abort, pfc::string8* name) {
out.release();
pfc::list_t<input_entry::ptr> candidates;
forwarding_candidates(url, candidates);
if (candidates.get_count() == 0) return false;
try {
GUID used = pfc::guid_null;
service_ptr opened = input_entry::g_open_from_list(candidates, what_for, nullptr, url,
nullptr, abort, &used);
if (!opened.is_valid() || !opened->service_query_t(out)) return false;
if (name != nullptr) {
input_entry::ptr entry = input_entry::g_find_by_guid(used);
*name = entry.is_valid() ? entry->get_name_() : "another component";
}
return true;
} catch (const pfc::exception& error) {
joc_log::line("decoder: no other component answers for this file's tags: %s",
error.what());
return false;
}
}
// Bytes in front of the E-AC-3 stream, which is where a tagging tool puts an
// ID3v2 tag. The renderer refuses a stream that does not begin on a syncword and
// never resynchronises, so the walk and the feed both have to start after it.
t_filesize leading_tag_bytes(file::ptr const& source, abort_callback& abort) {
if (!source.is_valid()) return 0;
try {
if (source->get_position(abort) != 0) source->seek(0, abort);
return tag_processor::skip_id3v2(source, abort);
} catch (const pfc::exception& error) {
joc_log::line("decoder: cannot inspect the area in front of the stream: %s",
error.what());
return 0;
}
}
// Tags read straight from the file, for a bare stream that no other component
// claims: an ID3v2 tag in front of the syncframes, or an APEv2/ID3v1 tag behind
// them. Neither is part of E-AC-3, so a tag that is there was written by a
// tagging tool and is worth showing.
void read_local_tags(file::ptr const& source, file_info& info, abort_callback& abort) {
if (!source.is_valid()) return;
bool found = false;
try {
source->seek(0, abort);
tag_processor::read_id3v2(source, info, abort);
found = true;
} catch (const pfc::exception&) {
// No leading tag; the trailing one is still worth a look.
}
try {
tag_processor::read_trailing(source, info, abort);
found = true;
} catch (const pfc::exception&) {
}
if (found) {
joc_log::line("decoder: %u tag field(s) read from the file itself",
static_cast<unsigned>(info.meta_get_count()));
}
}
class input_joc : public input_stubs {
public:
void open(service_ptr_t<file> hint, const char* path, t_input_open_reason reason,
abort_callback& abort) {
if (reason == input_open_info_write) throw exception_tagging_unsupported();
m_path = (path != nullptr) ? path : "";
if (reason == input_open_info_write) {
// Writing tags belongs to whoever owns the file's format, and that is
// not this component: its inputs are two ffmpeg children and the JOC
// renderer, none of which writes anything. The file is deliberately
// left unopened here, because a write-mode handle of ours would make
// the writer that replaces it fail on a sharing violation.
m_write_only = true;
if (!open_forwarded(m_forward_writer, input_info_writer::class_guid, m_path.c_str(),
abort, &m_forward_name)) {
throw exception_tagging_unsupported();
}
joc_log::line("decoder: open \"%s\" reason=2 tags: written by %s", m_path.c_str(),
m_forward_name.c_str());
return;
}
service_ptr_t<file> source = hint;
input_open_file_helper(source, path, reason, abort);
m_file = source;
@@ -78,7 +210,31 @@ public:
if (joc_container::is_container_extension(extension)) {
open_container(extension);
return;
} else {
open_bare(abort);
}
// Whatever else happens, the tags of this file are read by the component
// that owns its format; failing to find one is not fatal, the technical
// information below is still worth showing.
if (!open_forwarded(m_forward_reader, input_info_reader::class_guid, m_path.c_str(), abort,
&m_forward_name)) {
joc_log::line("decoder: no other component reads this file's tags");
} else {
joc_log::line("decoder: tags for \"%s\" are read by %s", m_path.c_str(),
m_forward_name.c_str());
}
}
// A bare stream is either read from the file or handed back. One thing has to
// happen first: a tag area in front of the syncframes is not part of the
// stream, and treating it as one would hand the file to the built-in decoder,
// which plays it without the Atmos objects.
void open_bare(abort_callback& abort) {
m_stream_start = leading_tag_bytes(m_file, abort);
if (m_stream_start != 0) {
joc_log::line("decoder: %llu byte(s) of tags in front of the stream are skipped",
static_cast<unsigned long long>(m_stream_start));
}
pfc::array_t<t_uint8> buffer;
@@ -86,9 +242,8 @@ public:
const std::size_t got = m_file->read(buffer.get_ptr(), kSniffBytes, abort);
const joc_eac3::ScanResult scan = joc_eac3::scan(buffer.get_ptr(), got, 8);
joc_log::line("decoder: open \"%s\" reason=%d bytes=%llu frames=%llu with_joc=%llu",
m_path.c_str(), static_cast<int>(reason),
static_cast<unsigned long long>(got),
joc_log::line("decoder: open \"%s\" reason=1 bytes=%llu frames=%llu with_joc=%llu",
m_path.c_str(), static_cast<unsigned long long>(got),
static_cast<unsigned long long>(scan.frames_examined),
static_cast<unsigned long long>(scan.frames_with_joc));
@@ -137,8 +292,18 @@ public:
}
void get_info(file_info& info, abort_callback& abort) {
(void)abort;
const joc_decode::FileProbe probe = joc_decode::probe_file(m_native_path.get_ptr());
if (m_write_only) {
// An instance opened to write tags is the writer's reader: what it
// reports is exactly what the caller has just written.
if (m_forward_writer.is_valid()) {
m_forward_writer->get_info(0, info, abort);
return;
}
throw exception_tagging_unsupported();
}
const joc_decode::FileProbe probe =
joc_decode::probe_file(m_native_path.get_ptr(), 0, m_stream_start);
const joc_decode::Settings settings = joc_settings::current();
// A container knows its own duration even though the E-AC-3 syncframes are
@@ -171,6 +336,10 @@ public:
info.info_set_int("bitspersample", 32);
info.info_set("bitspersample_extra", "floating-point");
info.set_length(duration);
m_length = duration;
// The renderer keeps its room tail, but the stream this component hands over
// ends where the file ends: the tail is rendering, not playback time.
m_engine.set_length(duration);
if (duration > 0.0) {
const t_filesize bytes = m_file.is_valid() ? m_file->get_size(abort) : filesize_invalid;
if (bytes != filesize_invalid && bytes > 0) {
@@ -194,6 +363,32 @@ public:
m_container.audio_index);
info.info_set("joc_container", text);
}
// The file's own reader supplies the tags; nothing above this line is one.
// Only the metadata is taken over -- its technical information (E-AC-3,
// 6 channels, the stream's own bitrate) would replace this component's,
// which is the part that says whether the file is JOC.
bool have_tags = false;
if (m_forward_reader.is_valid()) {
try {
file_info_impl tags;
m_forward_reader->get_info(0, tags, abort);
info.copy_meta(tags);
have_tags = tags.meta_get_count() != 0;
joc_log::line("decoder: %u tag field(s) from %s",
static_cast<unsigned>(tags.meta_get_count()),
m_forward_name.c_str());
} catch (const pfc::exception& error) {
joc_log::line("decoder: reading this file's own tags failed: %s", error.what());
}
}
// A reader that answers for the format but has nothing to say about a bare
// stream is common -- ffmpeg's AC-3 decoder reads no tags at all -- while
// the file may still carry an ID3v2 or APEv2 tag a tagging tool wrote.
if (!have_tags && m_input_kind == joc_decode::InputKind::kBare) {
read_local_tags(m_file, info, abort);
}
joc_log::line("decoder: get_info duration=%.3f s frames=%llu channels=%u render=%s%s",
duration, static_cast<unsigned long long>(frames), channels, render.c_str(),
container ? " (container)" : "");
@@ -201,6 +396,7 @@ public:
t_filestats2 get_stats2(uint32_t flags, abort_callback& abort) {
if (m_file.is_valid()) return m_file->get_stats2_(flags, abort);
if (m_forward_writer.is_valid()) return m_forward_writer->get_stats2_(nullptr, flags, abort);
throw exception_io_unsupported_format();
}
@@ -209,6 +405,8 @@ public:
m_settings = joc_settings::current();
m_settings.input_kind = m_input_kind;
m_settings.audio_index = m_audio_index;
m_settings.stream_start_bytes = m_stream_start;
m_settings.length_seconds = m_length;
joc_log::line("decoder: initialize flags=0x%X settings: %s", flags,
joc_settings::describe(m_settings).c_str());
@@ -229,6 +427,7 @@ public:
m_buffer.resize(kRunFrames * m_channels);
m_frames_delivered = 0;
m_reported = false;
m_delivery_mismatches = 0;
joc_log::line("decoder: engine ready, %u output channel(s), %u frames per read",
m_channels, static_cast<unsigned>(kRunFrames));
}
@@ -242,50 +441,98 @@ public:
joc_log::line("decoder: read failed: %s", error.c_str());
throw exception_io_data(error.c_str());
}
joc_log::line("decoder: end of stream after %llu frames",
static_cast<unsigned long long>(m_frames_delivered));
joc_log::line("decoder: end of stream after %llu frames%s",
static_cast<unsigned long long>(m_frames_delivered),
m_delivery_mismatches == 0 ? ""
: " (the delivery changed samples)");
return false;
}
chunk.set_data_size(frames * m_channels);
chunk.set_channels(m_channels, audio_chunk::g_guess_channel_config(m_channels));
chunk.set_sample_rate(kSampleRate);
chunk.set_sample_count(frames);
std::memcpy(chunk.get_data(), m_buffer.data(),
frames * m_channels * sizeof(audio_sample));
// The renderer produces float32 and a chunk holds audio_sample, which is float on
// 32-bit builds and double on 64-bit ones (SDK audio_math.h): the samples are
// converted, not copied. set_data_32() is the SDK's conversion for a float32
// source, and it sets the channel count, the sample rate and the sample count.
chunk.set_data_32(m_buffer.data(), frames, m_channels, kSampleRate);
m_frames_delivered += frames;
// A delivery that mangled the samples would be heard as noise rather than reported
// as a failure, so every chunk is checked: the conversion is exact, and the peak of
// what the chunk holds has to equal the peak of what the renderer produced.
const double produced = peak_of(m_buffer.data(), frames * m_channels);
const double delivered =
peak_of(chunk.get_data(), chunk.get_sample_count() * chunk.get_channels());
if (delivered > produced + 1e-6 + produced * 1e-6 ||
delivered < produced - 1e-6 - produced * 1e-6) {
++m_delivery_mismatches;
if (m_delivery_mismatches == 1) {
joc_log::line("decoder: delivery changed the samples: peak %.9f produced, "
"%.9f delivered",
produced, delivered);
}
}
if (!m_reported) {
m_reported = true;
float peak = 0.0f;
for (std::size_t i = 0; i < frames * m_channels; ++i) {
const float value = m_buffer[i] < 0.0f ? -m_buffer[i] : m_buffer[i];
if (value > peak) peak = value;
}
joc_log::line("decoder: first %llu frames delivered (%u ch), peak %.6f",
static_cast<unsigned long long>(frames), m_channels,
static_cast<double>(peak));
joc_log::line("decoder: first %llu frames delivered (%u ch), peak %.6f, "
"delivered peak %.6f",
static_cast<unsigned long long>(frames), m_channels, produced,
delivered);
}
return true;
}
void decode_seek(double, abort_callback&) {
// The renderer is stateful and has no seek; can_seek() says so.
throw exception_io_unsupported_format();
void decode_seek(double seconds, abort_callback& abort) {
// Walking the syncframe index of a long file is the only part of a seek
// that can take a while, and it polls this.
m_engine.set_abort_check([&abort] { return !abort.is_aborting(); });
std::string error;
const bool ok = m_engine.seek(seconds, m_length, &error);
m_engine.set_abort_check(nullptr);
if (!ok) {
// An aborted seek reports itself as an abort, not as a decode failure.
abort.check();
joc_log::line("decoder: seek to %.6f s failed: %s", seconds, error.c_str());
throw exception_io_data(error.c_str());
}
// The position reporting and the first-read statistics belong to the run
// that starts here, not to the one that was interrupted.
m_frames_delivered = 0;
m_reported = false;
m_delivery_mismatches = 0;
joc_log::line("decoder: seek to %.6f s, the next read starts at the target", seconds);
}
bool decode_can_seek() { return false; }
bool decode_can_seek() { return true; }
size_t extended_param(const GUID& type, size_t arg1, void* arg2, size_t arg2size) {
(void)arg1;
(void)arg2;
(void)arg2size;
// A seek restarts both ffmpeg children and replays the renderer's warm-up,
// so it is worth avoiding the ones the core would only make speculatively.
if (type == input_params::seeking_expensive) return 1;
return 0;
}
void retag(const file_info&, abort_callback&) { throw exception_tagging_unsupported(); }
void remove_tags(abort_callback&) { throw exception_tagging_unsupported(); }
void retag(const file_info& info, abort_callback& abort) {
if (!m_forward_writer.is_valid()) throw exception_tagging_unsupported();
// A single-track input has no commit() of its own -- the SDK wrapper
// implements it as a no-op -- so the writer's commit has to happen here or
// nothing reaches the file.
m_forward_writer->set_info(0, info, abort);
m_forward_writer->commit(abort);
joc_log::line("decoder: %u tag field(s) written through %s",
static_cast<unsigned>(info.meta_get_count()), m_forward_name.c_str());
}
void remove_tags(abort_callback& abort) {
if (!m_forward_writer.is_valid()) throw exception_tagging_unsupported();
input_info_writer_v2::ptr v2;
if (m_forward_writer->service_query_t(v2)) {
v2->remove_tags(abort);
return;
}
m_forward_writer->remove_tags_fallback(abort);
}
static bool g_is_our_path(const char* path, const char* extension) {
(void)path;
@@ -334,6 +581,19 @@ private:
unsigned m_channels = 2;
std::uint64_t m_frames_delivered = 0;
bool m_reported = false;
// Chunks whose delivered samples did not match what the renderer produced.
std::uint64_t m_delivery_mismatches = 0;
// Duration get_info() last reported; a seek needs it to tell "past the end"
// from "inside the file" without decoding anything.
double m_length = 0.0;
// Bytes in front of a bare stream, which is where an ID3v2 tag sits.
std::uint64_t m_stream_start = 0;
// The other component that answers for this file's tags, and the one that
// writes them. Only the writer exists on an instance opened to retag.
service_ptr_t<input_info_reader> m_forward_reader;
service_ptr_t<input_info_writer> m_forward_writer;
pfc::string8 m_forward_name;
bool m_write_only = false;
};
static input_singletrack_factory_t<input_joc> g_input_joc_factory;
+416 -44
View File
@@ -2,8 +2,10 @@
#include <windows.h>
#include <cmath>
#include <cstdio>
#include <cstring>
#include <cwchar>
#include <vector>
// The kernel copy that is compiled into this component; see kernel/.
@@ -28,6 +30,32 @@ constexpr std::size_t kBedReadBytes = 48u * 1024u;
constexpr std::size_t kBedPipeBytes = 1u << 20;
constexpr std::size_t kFrameSamples = JOC_FRAME_SAMPLES;
// The decoded bed has to stay on the E-AC-3 frame grid that the JOC matrix and the
// OAMD are indexed by. An mp4/mov edit list trims the decoded audio instead, which
// takes the bed off that grid by however much the list removes -- a Dolby Atmos
// download loses 2432 samples (1.58 frames) -- and every frame's matrix would then be
// applied to audio tens of milliseconds away from it, so a subset of the objects comes
// out attenuated. The demuxer's own option drops the trim; the option exists only on
// the mov/mp4 demuxer, so it is passed for that family alone.
bool has_mov_timeline(const std::string& path) {
static const char* const kExtensions[] = {".mp4", ".m4a", ".m4b", ".m4v",
".mov", ".3gp", ".3g2", ".mj2"};
std::string lowered = path;
for (char& character : lowered) {
if (character >= 'A' && character <= 'Z') {
character = static_cast<char>(character - 'A' + 'a');
}
}
for (const char* extension : kExtensions) {
const std::size_t length = std::strlen(extension);
if (lowered.size() >= length &&
lowered.compare(lowered.size() - length, length, extension) == 0) {
return true;
}
}
return false;
}
// ---------------------------------------------------------------------------
// Kernel entry points.
//
@@ -42,6 +70,7 @@ struct CoreApi {
std::uint32_t*) = joc_stream_push;
joc_error(JOC_CALL* pull)(joc_stream*, joc_stream_buffer*, std::uint32_t*) = joc_stream_pull;
joc_error(JOC_CALL* flush)(joc_stream*) = joc_stream_flush;
joc_error(JOC_CALL* reset)(joc_stream*) = joc_stream_reset;
joc_error(JOC_CALL* status)(const joc_stream*, joc_stream_status_info*) = joc_stream_status;
joc_error(JOC_CALL* destroy)(joc_stream*) = joc_stream_destroy;
std::uint32_t(JOC_CALL* abi_version)() = joc_abi_version;
@@ -199,6 +228,7 @@ public:
FILE_ATTRIBUTE_NORMAL | FILE_FLAG_SEQUENTIAL_SCAN, nullptr);
return handle_ != INVALID_HANDLE_VALUE;
}
bool is_open() const { return handle_ != INVALID_HANDLE_VALUE; }
std::size_t read(void* destination, std::size_t bytes) {
if (handle_ == INVALID_HANDLE_VALUE) return 0;
DWORD got = 0;
@@ -207,6 +237,12 @@ public:
}
return got;
}
bool seek(std::uint64_t offset) {
if (handle_ == INVALID_HANDLE_VALUE) return false;
LARGE_INTEGER value{};
value.QuadPart = static_cast<LONGLONG>(offset);
return SetFilePointerEx(handle_, value, nullptr, FILE_BEGIN) != FALSE;
}
std::uint64_t size() const {
LARGE_INTEGER value{};
if (handle_ == INVALID_HANDLE_VALUE || GetFileSizeEx(handle_, &value) == FALSE) return 0;
@@ -228,6 +264,41 @@ const char* const kLayouts[] = {"2.0", "3.0", "3.1", "4.0", "5.0", "5.
"9.1.6", "22.2"};
const unsigned kLayoutChannels[] = {2, 3, 4, 4, 5, 6, 8, 10, 7, 7, 8, 10, 12, 14, 16, 24};
// Every ffmpeg child writes its diagnostics next to the component rather than into
// whatever working directory the host process happens to have; separate files so
// no child can truncate another's.
std::wstring stderr_path_for(const wchar_t* name) {
HMODULE self = nullptr;
GetModuleHandleExW(GET_MODULE_HANDLE_EX_FLAG_FROM_ADDRESS |
GET_MODULE_HANDLE_EX_FLAG_UNCHANGED_REFCOUNT,
reinterpret_cast<LPCWSTR>(&speaker_channels), &self);
wchar_t path[4096] = {};
const DWORD length = GetModuleFileNameW(self, path, 4096);
if (length == 0) return std::wstring(name);
const std::wstring text(path, length);
const std::wstring::size_type slash = text.find_last_of(L"\\/");
return slash == std::wstring::npos ? std::wstring(name)
: text.substr(0, slash) + L"\\" + name;
}
// The time ffmpeg's -ss takes for a given source sample. A seek always starts on
// a syncframe boundary, so the sample is exactly representable in microseconds.
std::wstring seek_time(std::uint64_t source_sample) {
if (source_sample == 0) return {};
wchar_t text[64] = {};
std::swprintf(text, 64, L"%.6f", static_cast<double>(source_sample) / 48000.0);
return text;
}
// Source sample a time offset names, rounded the way the core rounds a position
// (pfc::rint64, i.e. llrint: to nearest, ties to even).
std::int64_t sample_of_seconds(double seconds) {
if (!(seconds > 0.0)) return 0; // also catches NaN
const double value = seconds * 48000.0;
if (value >= 9.0e18) return 9223372036854775807LL;
return static_cast<std::int64_t>(std::llrint(value));
}
} // namespace
const char* const* speaker_layouts(std::size_t* count) {
@@ -277,7 +348,8 @@ std::string resolve_hrtf_file(const Settings& settings) {
return directory + "\\HRTF\\" + name;
}
FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) {
FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes,
std::uint64_t start_offset) {
FileProbe probe;
InputFile file;
if (!file.open(path)) {
@@ -285,6 +357,15 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) {
return probe;
}
const std::uint64_t size = file.size();
if (start_offset >= size) {
probe.detail = "file too small to be E-AC-3";
return probe;
}
if (start_offset != 0 && !file.seek(start_offset)) {
probe.detail = "cannot skip the leading tag area";
return probe;
}
const std::uint64_t stream_bytes = size - start_offset;
// A Media Library scan calls this for every file, so the whole stream is
// walked only when that is cheap; otherwise the first window is enough,
@@ -293,7 +374,7 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) {
const std::size_t window =
(max_scan_bytes != 0) ? max_scan_bytes : 256u * 1024u;
std::vector<std::uint8_t> buffer(static_cast<std::size_t>(
(size < kFullWalkLimit && size > 0) ? size : window));
(stream_bytes < kFullWalkLimit && stream_bytes > 0) ? stream_bytes : window));
const std::size_t got = file.read(buffer.data(), buffer.size());
if (got < 8) {
probe.detail = "file too small to be E-AC-3";
@@ -310,7 +391,7 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) {
}
std::uint64_t frames = 0;
if (buffer.size() == size) {
if (buffer.size() == stream_bytes) {
std::size_t offset = 0;
while (true) {
const std::size_t bytes = joc_eac3::frame_bytes_at(buffer.data(), got, offset);
@@ -320,7 +401,7 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) {
}
probe.detail = "frame count walked over the whole file";
} else if (scan.all_frames_same_size) {
frames = size / scan.first_frame_bytes;
frames = stream_bytes / scan.first_frame_bytes;
probe.detail = "frame count extrapolated from a constant frame size";
} else {
// Variable frame size: count in the window and scale by the byte ratio.
@@ -333,7 +414,8 @@ FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes) {
++seen;
}
frames = (offset != 0) ? static_cast<std::uint64_t>(
(static_cast<double>(size) / static_cast<double>(offset)) *
(static_cast<double>(stream_bytes) /
static_cast<double>(offset)) *
static_cast<double>(seen))
: 0;
probe.detail = "frame count estimated from a variable frame size";
@@ -443,6 +525,7 @@ struct Engine::Impl {
bool eac3_from_pipe = false;
FfmpegPipe bed;
Settings settings;
std::string input_path;
// The metadata stream comes either from the file itself or from ffmpeg.
std::size_t read_eac3(void* destination, std::size_t bytes) {
@@ -464,8 +547,193 @@ struct Engine::Impl {
std::vector<float> bed_buffer;
std::size_t bed_staged_bytes = 0; // bytes staged at the front of bed_buffer
std::vector<float> pull_buffer;
// Seek state. seek_skip counts output frames still to be discarded: a seek
// restarts the inputs on the syncframe before the target, and the samples the
// renderer produces for the part already behind the target are dropped here.
std::uint64_t seek_skip = 0;
// Delivered-stream accounting: where the stream currently being delivered starts
// on the source timeline, how much of it has been handed over, and where it has
// to stop. end_sample is the file's own end, so a binaural room tail cannot
// turn into playback time the file does not have.
std::uint64_t start_sample = 0;
std::uint64_t delivered = 0;
std::uint64_t end_sample = 0; // 0 = no limit
// Syncframes of a container's metadata stream still to be read and thrown away.
// ffmpeg's input seek on a copy stream lands on the frame whose timestamp is
// at or after the target, which is not reliably the frame the bed restarts on,
// so a container seek re-reads the stream and drops the frames here instead.
std::uint64_t metadata_skip = 0;
bool at_end = false; // seek landed at or past the end
std::function<bool()> abort_check;
// Bare-stream syncframe index: offset of every syncframe within the stream
// (so index_base has to be added to get a file offset). Filled by walking the
// file, and only as far as a seek asks for.
std::vector<std::uint64_t> frame_offsets;
std::uint64_t index_base = 0; // file offset the stream starts at
std::uint64_t index_file_offset = 0; // file offset the walk continues from
std::uint64_t index_stream_offset = 0; // stream offset the walk continues from
bool index_complete = false;
void reset_feed_state() {
frames_queued = 0;
bed_frames_pushed = 0;
eac3_eof = false;
bed_eof = false;
flushed = false;
trace_count = 0;
read_calls = 0;
eac3_carry = 0;
bed_bytes_read = 0;
bed_staged_bytes = 0;
seek_skip = 0;
metadata_skip = 0;
start_sample = 0;
delivered = 0;
at_end = false;
}
void stop_inputs() {
bed.stop();
eac3_pipe.stop();
eac3.close();
}
void reset_index() {
frame_offsets.clear();
index_base = settings.stream_start_bytes;
index_file_offset = index_base;
index_stream_offset = 0;
index_complete = false;
}
// Grows the syncframe index until it holds `wanted` entries or the stream ends.
bool grow_frame_index(std::uint64_t wanted, std::string* error);
// Points the file at the syncframe `frame`, or reports at_end when the stream
// has fewer frames than that.
bool position_bare(std::uint64_t frame, std::string* error);
bool start_bed(std::uint64_t source_sample, std::string* error);
bool start_metadata(std::string* error);
};
bool Engine::Impl::grow_frame_index(std::uint64_t wanted, std::string* error) {
constexpr std::size_t kWindow = 256u * 1024u;
std::vector<std::uint8_t> buffer(kWindow);
while (frame_offsets.size() < wanted && !index_complete) {
if (abort_check != nullptr && !abort_check()) {
if (error != nullptr) *error = "the seek was aborted";
return false;
}
if (!eac3.is_open() && !eac3.open(input_path)) {
if (error != nullptr) *error = "cannot reopen the input file";
return false;
}
const std::uint64_t size = eac3.size();
if (index_file_offset + 4u > size) {
index_complete = true;
break;
}
if (!eac3.seek(index_file_offset)) {
if (error != nullptr) *error = "cannot position the input file";
return false;
}
const std::uint64_t remaining = size - index_file_offset;
const std::size_t want = static_cast<std::size_t>(
remaining < kWindow ? remaining : kWindow);
const std::size_t got = eac3.read(buffer.data(), want);
if (got < 4u) {
index_complete = true;
break;
}
std::size_t consumed = 0;
while (frame_offsets.size() < wanted) {
const std::size_t bytes = joc_eac3::frame_bytes_at(buffer.data(), got, consumed);
if (bytes == 0 || consumed + bytes > got) break;
frame_offsets.push_back(index_stream_offset + consumed);
consumed += bytes;
}
if (consumed == 0) {
// Not even one whole syncframe in a full window: the stream stops here.
index_complete = true;
break;
}
index_file_offset += consumed;
index_stream_offset += consumed;
}
return true;
}
bool Engine::Impl::position_bare(std::uint64_t frame, std::string* error) {
if (!grow_frame_index(frame + 1u, error)) return false;
if (frame >= frame_offsets.size()) {
// Past the last syncframe: the decoder contract asks for a successful seek
// that the next read() answers with end of stream.
at_end = true;
stop_inputs();
return true;
}
// Position check first: frame_offsets.size() is what bounds the index.
const std::uint64_t file_offset =
index_base + frame_offsets[static_cast<std::size_t>(frame)];
if (!eac3.is_open() && !eac3.open(input_path)) {
if (error != nullptr) *error = "cannot reopen the input file";
return false;
}
// The renderer rejects a stream that does not begin on a syncword and never
// resynchronises, so a mispositioned start has to fail loudly here rather than
// turn into silence at the end of the file.
std::uint8_t header[8] = {};
if (!eac3.seek(file_offset) || eac3.read(header, sizeof(header)) < 4u ||
joc_eac3::frame_bytes_at(header, sizeof(header), 0) == 0) {
if (error != nullptr) {
*error = "the syncframe index does not point at a syncframe (offset " +
std::to_string(file_offset) + ")";
}
return false;
}
if (!eac3.seek(file_offset)) {
if (error != nullptr) *error = "cannot position the input file";
return false;
}
return true;
}
bool Engine::Impl::start_bed(std::uint64_t source_sample, std::string* error) {
// The 5.1 core PCM, exactly as the reference renderer's own core decode does
// it: 5.1 interleaved float32 at 48 kHz, the layout the renderer expects
// (L R C LFE Ls Rs).
// -drc_scale 0 -target_level 0: the bed is taken as stored, without the
// stream's dynrng or target-level metadata being applied by the decoder.
//
// -ss is an input option: the demuxer is positioned on the frame boundary and
// decoding starts there, so a seek costs the same wherever it lands. The price is
// that the bed is carried by a decoder that started at the seek point rather than
// at the start of the file, which is a low-level, noise-like difference from a
// play-through rather than a misalignment: the position stays exact.
std::wstring input_arguments = L"-drc_scale 0 -target_level 0";
if (has_mov_timeline(input_path)) input_arguments += L" -ignore_editlist 1";
const std::wstring offset = seek_time(source_sample);
if (!offset.empty()) input_arguments += L" -ss " + offset;
std::wstring bed_arguments = L"-map 0:a:";
bed_arguments += std::to_wstring(settings.audio_index);
bed_arguments += L" -vn -ac 6 -ar 48000 -c:a pcm_f32le -f f32le -";
return bed.start(settings.ffmpeg_path, input_path, input_arguments, bed_arguments, "bed",
stderr_path_for(L"joc_ffmpeg_bed.log"), error);
}
bool Engine::Impl::start_metadata(std::string* error) {
// Stream copy from the start of the track: the syncframes arrive byte for byte
// as they are stored, which is what the JOC metadata needs, and a seek then
// discards whole syncframes from the front (see metadata_skip) rather than
// asking ffmpeg to position the stream.
std::wstring stream_arguments = L"-map 0:a:";
stream_arguments += std::to_wstring(settings.audio_index);
stream_arguments += L" -vn -c:a copy -f eac3 -";
return eac3_pipe.start(settings.ffmpeg_path, input_path, std::wstring(), stream_arguments,
"metadata", stderr_path_for(L"joc_ffmpeg_stream.log"), error,
1u << 20);
}
Engine::Engine() : impl_(new Impl()) {}
Engine::~Engine() {
@@ -475,21 +743,30 @@ Engine::~Engine() {
unsigned Engine::channels() const { return impl_->channels; }
void Engine::set_abort_check(std::function<bool()> check) { impl_->abort_check = std::move(check); }
void Engine::set_length(double seconds) { impl_->end_sample = sample_of_seconds(seconds); }
void Engine::stop() {
Impl& impl = *impl_;
if (impl.stream != nullptr && impl.api.destroy != nullptr) {
impl.api.destroy(impl.stream);
impl.stream = nullptr;
}
impl.bed.stop();
impl.eac3_pipe.stop();
impl.eac3.close();
impl.eac3_from_pipe = false;
impl.stop_inputs();
impl.reset_feed_state();
}
bool Engine::start(const std::string& input_path, const Settings& settings, std::string* error) {
Impl& impl = *impl_;
// initialize() may be called more than once on the same instance, and each call
// has the renderer and its inputs start from scratch.
stop();
impl.settings = settings;
impl.input_path = input_path;
impl.reset_index();
impl.reset_feed_state();
impl.end_sample = sample_of_seconds(settings.length_seconds);
impl.eac3_buffer.resize(kEac3Chunk);
impl.bed_buffer.resize(kBedFramesChunk * kBedChannels);
@@ -513,6 +790,12 @@ bool Engine::start(const std::string& input_path, const Settings& settings, std:
if (error != nullptr) *error = "cannot open the input file";
return false;
}
// A tag area in front of the stream is skipped here rather than by the
// renderer, which rejects a stream that does not begin on a syncword.
if (impl.index_base != 0 && !impl.eac3.seek(impl.index_base)) {
if (error != nullptr) *error = "cannot skip the leading tag area";
return false;
}
}
joc_stream_config config{};
@@ -590,51 +873,89 @@ bool Engine::start(const std::string& input_path, const Settings& settings, std:
// ffmpeg's stderr lands next to the component rather than in whatever working
// directory the host process happens to have. The two children write separate
// files so neither can truncate the other's diagnostics.
const auto stderr_path_for = [](const wchar_t* name) {
HMODULE self = nullptr;
GetModuleHandleExW(GET_MODULE_HANDLE_EX_FLAG_FROM_ADDRESS |
GET_MODULE_HANDLE_EX_FLAG_UNCHANGED_REFCOUNT,
reinterpret_cast<LPCWSTR>(&speaker_channels), &self);
wchar_t path[4096] = {};
const DWORD length = GetModuleFileNameW(self, path, 4096);
if (length == 0) return std::wstring(name);
const std::wstring text(path, length);
const std::wstring::size_type slash = text.find_last_of(L"\\/");
return slash == std::wstring::npos ? std::wstring(name)
: text.substr(0, slash) + L"\\" + name;
};
if (!impl.start_bed(0, error)) return false;
if (impl.eac3_from_pipe && !impl.start_metadata(error)) return false;
return true;
}
// The 5.1 core PCM, exactly as the reference renderer's own core decode does it:
// 5.1 interleaved float32 at 48 kHz, the layout the renderer expects
// (L R C LFE Ls Rs).
// -drc_scale 0 -target_level 0: the bed is taken as stored, without the stream's
// dynrng or target-level metadata being applied by the decoder.
std::wstring bed_arguments = L"-map 0:a:";
bed_arguments += std::to_wstring(settings.audio_index);
bed_arguments += L" -vn -ac 6 -ar 48000 -c:a pcm_f32le -f f32le -";
if (!impl.bed.start(settings.ffmpeg_path, input_path, L"-drc_scale 0 -target_level 0", bed_arguments, "bed",
stderr_path_for(L"joc_ffmpeg_bed.log"), error)) {
bool Engine::seek(double seconds, double total_seconds, std::string* error) {
Impl& impl = *impl_;
if (impl.stream == nullptr || impl.api.reset == nullptr) {
if (error != nullptr) *error = "the renderer is not running";
return false;
}
// Where the stream that is about to be delivered starts:
//
// * the frame the target sits in, minus a pre-roll, so the renderer's own state
// -- object timeline, matrix interpolation, room tail -- has converged by the
// time the target itself is delivered. Both inputs restart there, and the bed
// is positioned with an input seek, which is what keeps the cost of a seek
// independent of where it lands;
// * the samples in front of the target are then dropped from the output, which is
// what makes the delivery start exactly on the requested sample.
const std::int64_t target = sample_of_seconds(seconds);
std::int64_t wanted = target - static_cast<std::int64_t>(impl.settings.pipeline_delay_samples);
if (wanted < 0) wanted = 0;
const std::uint64_t target_frame = static_cast<std::uint64_t>(wanted) / kFrameSamples;
const std::uint64_t preroll = impl.settings.seek_preroll_frames;
const std::uint64_t frame = (target_frame > preroll) ? (target_frame - preroll) : 0;
const std::uint64_t first_sample = frame * kFrameSamples;
const std::uint64_t skip = static_cast<std::uint64_t>(wanted) - first_sample;
impl.stop_inputs();
impl.reset_feed_state();
if (impl.eac3_from_pipe) {
// Stream copy: the syncframes arrive byte for byte as they are stored, which
// is what the JOC metadata needs.
std::wstring stream_arguments = L"-map 0:a:";
stream_arguments += std::to_wstring(settings.audio_index);
stream_arguments += L" -vn -c:a copy -f eac3 -";
if (!impl.eac3_pipe.start(settings.ffmpeg_path, input_path, L"", stream_arguments, "metadata",
stderr_path_for(L"joc_ffmpeg_stream.log"), error,
1u << 20)) {
return false;
// A container's frame count is not known before it is decoded, so the
// caller's duration is what says whether this lands past the end.
if (total_seconds > 0.0 && seconds >= total_seconds) {
impl.at_end = true;
joc_log::line("engine: seek %.6f s is at or past the end (%.6f s)", seconds,
total_seconds);
}
} else if (!impl.position_bare(frame, error)) {
return false;
}
if (!impl.at_end) {
if (!impl.start_bed(first_sample, error)) return false;
if (impl.eac3_from_pipe) {
if (!impl.start_metadata(error)) return false;
impl.metadata_skip = frame;
}
impl.seek_skip = skip;
impl.start_sample = static_cast<std::uint64_t>(wanted);
impl.delivered = 0;
}
// joc_stream_reset keeps the renderer and its HRTF field: it drops the whole
// timeline, gain ramps, room tail and filter-bank history, which is exactly
// what a restart at another position needs.
const joc_error reset = impl.api.reset(impl.stream);
if (reset != JOC_OK) {
if (error != nullptr) {
*error = std::string("cannot reset the render stream: ") +
error_text(impl.api, reset);
}
return false;
}
joc_log::line("engine: seek %.6f s -> source sample %lld, frames %llu.. from sample %llu, "
"drop %llu sample(s)%s",
seconds, static_cast<long long>(target),
static_cast<unsigned long long>(frame),
static_cast<unsigned long long>(first_sample),
static_cast<unsigned long long>(impl.seek_skip),
impl.at_end ? " (at end)" : "");
return true;
}
std::size_t Engine::read(float* destination, std::size_t frames, std::string* error) {
Impl& impl = *impl_;
if (impl.stream == nullptr || frames == 0) return 0;
// A seek at or past the end of the file succeeds and leaves the next read to
// report end of stream.
if (impl.at_end) return 0;
const bool trace = impl.trace_count < 6;
++impl.read_calls;
if ((impl.read_calls % 50u) == 0u) {
@@ -672,10 +993,36 @@ std::size_t Engine::read(float* destination, std::size_t frames, std::string* er
// frame of the file unrendered.
const std::size_t got = impl.read_eac3(impl.eac3_buffer.data() + impl.eac3_carry,
impl.eac3_buffer.size() - impl.eac3_carry);
const std::size_t total = impl.eac3_carry + got;
std::size_t total = impl.eac3_carry + got;
if (total == 0) {
impl.eac3_eof = true;
} else {
// Syncframes in front of a seek target are thrown away before
// anything is interpreted. They are dropped a whole window at a
// time, so only as much of the stream is read as the skip needs.
if (impl.metadata_skip != 0) {
std::size_t dropped = 0;
while (impl.metadata_skip != 0) {
const std::size_t bytes =
joc_eac3::frame_bytes_at(impl.eac3_buffer.data(), total, dropped);
if (bytes == 0 || dropped + bytes > total) break;
dropped += bytes;
--impl.metadata_skip;
}
if (dropped != 0) {
total -= dropped;
std::memmove(impl.eac3_buffer.data(),
impl.eac3_buffer.data() + dropped, total);
}
if (impl.metadata_skip != 0) {
// The window ended inside the part to be skipped: keep what
// is left and come back with the next read.
impl.eac3_carry = total;
if (got == 0) impl.eac3_eof = true;
continue;
}
}
std::size_t offset = 0;
std::uint64_t complete = 0;
while (true) {
@@ -804,12 +1151,37 @@ std::size_t Engine::read(float* destination, std::size_t frames, std::string* er
return 0;
}
if (produced != 0u) {
std::size_t count = produced;
if (impl.seek_skip != 0) {
// The renderer had to be fed from before the seek target, so the
// samples it produces for that part are dropped before anything
// reaches the caller: the first sample handed over is the target.
const std::size_t drop = (impl.seek_skip < count)
? static_cast<std::size_t>(impl.seek_skip)
: count;
impl.seek_skip -= drop;
count -= drop;
if (count != 0) {
std::memmove(destination, destination + drop * impl.channels,
count * impl.channels * sizeof(float));
}
if (count == 0) continue; // the whole pull was pre-roll
}
// The renderer's binaural room tail is not part of the file: the stream
// ends where the file ends, not later.
if (impl.end_sample != 0) {
const std::uint64_t from = impl.start_sample + impl.delivered;
if (from >= impl.end_sample) return 0;
const std::uint64_t left = impl.end_sample - from;
if (left < count) count = static_cast<std::size_t>(left);
}
impl.delivered += count;
if (trace) {
++impl.trace_count;
joc_log::line("engine: pulled %u frame(s) on the %u%s attempt", produced,
joc_log::line("engine: pulled %u frame(s) on the %u%s attempt", count,
impl.trace_count, impl.trace_count == 1 ? "st" : "th");
}
return produced;
return count;
}
// Nothing more can arrive once the metadata stream is drained and the bed
+51 -2
View File
@@ -19,10 +19,29 @@
#include <cstddef>
#include <cstdint>
#include <functional>
#include <string>
namespace joc_decode {
// Delay of the rendering pipeline itself, in output samples: in a run that began
// at source sample 0, output sample k carries the rendering of source sample
// k - this value. A seek starts the inputs this much earlier and drops the
// samples before the target. Measured as 0: the renderer's own filter-bank
// latency and its metadata delay are inside the renderer, not delays of the
// delivered stream. It stays a named, overridable constant so the measurement can
// be repeated.
constexpr std::uint32_t kJocSeekPipelineDelaySamples = 0;
// Syncframes fed before the frame the target sits in. The renderer's state is
// rebuilt from the frames it is given, so a restart needs a moment to converge and
// the delivered part has to be past that. Two things converge at different rates:
// the metadata state (object positions, matrix interpolation, gain ramps), which
// the measured 3008-sample window covers, and the binaural room tail, which is
// recursive and can only be approached -- two syncframes are enough for the speaker
// path, and the tail keeps improving with more, which is why this is 64.
constexpr std::uint32_t kJocSeekPrerollFrames = 64;
enum class Output {
kBinaural = 0, // 2 channels, HRTF rendering
kSpeaker = 1, // N channels, named layout
@@ -57,9 +76,22 @@ struct Settings {
std::uint32_t binaural_mode = 3; // JOC_BINAURAL_MID
double gain_db = 0.0;
double tail_seconds = 5.0;
// Duration of the file, in seconds. The renderer ends a binaural stream with a
// room tail, which would be played as time the file does not have; the delivered
// stream is cut at exactly this much audio instead. Zero means "no limit".
double length_seconds = 0.0;
std::uint32_t object_delay_samples = 1473;
std::uint32_t native_threads = 0;
std::string ffmpeg_path = "ffmpeg";
// Bytes in front of the first syncframe -- a leading ID3v2 tag. The renderer
// rejects a stream that does not begin on a syncword, so both the walk and the
// feed start here.
std::uint64_t stream_start_bytes = 0;
// Samples of pipeline delay a seek compensates for, and syncframes it pre-rolls
// before the target: the calibrated constants above, unless the comparison
// harness overrides them to measure those constants.
std::uint32_t pipeline_delay_samples = kJocSeekPipelineDelaySamples;
std::uint32_t seek_preroll_frames = kJocSeekPrerollFrames;
// Stop feeding the renderer after this many input syncframes and flush, which
// is what the reference CLI's --duration does. Zero means "the whole file".
// Only the comparison harness sets it; playback leaves it at zero.
@@ -94,8 +126,10 @@ struct FileProbe {
};
// Walks the file's syncframes. Cheap enough for a Media Library scan: it only
// does pointer arithmetic over the stream, no decoding, no HRTF work.
FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes = 0);
// does pointer arithmetic over the stream, no decoding, no HRTF work. The walk
// starts at `start_offset`, which is where the stream begins behind a tag area.
FileProbe probe_file(const std::string& path, std::size_t max_scan_bytes = 0,
std::uint64_t start_offset = 0);
// Decides whether the E-AC-3 track of a container file carries JOC, without
// rendering anything: ffmpeg copies a short prefix of that track out (stream copy,
@@ -118,6 +152,21 @@ public:
unsigned channels() const;
void stop();
// Duration of the file, in seconds, for a caller that learns it after start()
// (reported duration of the track); zero means "no limit".
void set_length(double seconds);
// Repositions so that the next sample read() returns is the rendering of source
// sample round(seconds * 48000): the inputs restart at the syncframe holding
// that sample and the output before it is dropped. Seeking at or past
// `total_seconds` (0 when the duration is unknown) succeeds and makes the next
// read() return 0, as the decoder contract requires. start() must have run.
bool seek(double seconds, double total_seconds, std::string* error);
// Polled while the syncframe index is being walked, so a long seek stays
// interruptible; returning false aborts the seek.
void set_abort_check(std::function<bool()> check);
private:
struct Impl;
Impl* impl_;
+1 -1
View File
@@ -12,7 +12,7 @@
// Kept in one place: the string reported to foobar2000 and written to the log
// must not drift apart.
#define JOC_VERSION "0.1.0"
#define JOC_VERSION "0.2.2"
DECLARE_COMPONENT_VERSION("JOC decoder (E-AC-3 JOC)", JOC_VERSION,
"Plays E-AC-3 JOC (Dolby Atmos) files: the JOC objects are "
+25 -19
View File
@@ -189,6 +189,7 @@ private:
}
if (pick_file(m_hwnd, path, title, filter)) {
set_text(GetDlgItem(m_hwnd, IDC_EDIT_HRTF), path);
update_status();
notify();
}
return TRUE;
@@ -219,6 +220,10 @@ private:
return TRUE;
}
if (code == EN_CHANGE || code == CBN_SELCHANGE) {
if (id == IDC_EDIT_HRTF || id == IDC_COMBO_LAYOUT ||
id == IDC_EDIT_GAIN) {
update_status();
}
notify();
return TRUE;
}
@@ -338,35 +343,36 @@ private:
void update_enabled_state() { update_status(); }
void update_status() {
const joc_settings::Values values = joc_settings::read();
std::string status;
const joc_settings::Values values = read_controls();
std::string output_status;
if (values.output != 0) {
char text[96] = {};
const unsigned channels = joc_decode::speaker_channels(values.speaker_layout);
std::snprintf(text, sizeof(text), "扬声器布局 %s(%u 声道)",
values.speaker_layout.c_str(), channels);
status = text;
} else {
// Show what will actually be read, so an empty box is not a mystery.
joc_decode::Settings effective = joc_settings::current();
effective.hrtf_source =
static_cast<joc_decode::HrtfSource>(values.hrtf_source);
effective.hrtf_file = values.hrtf_file;
const std::string file = joc_decode::resolve_hrtf_file(effective);
status = std::string(joc_settings::hrtf_source_name(values.hrtf_source)) + ":" +
(file.empty() ? std::string("无法确定默认路径")
: (values.hrtf_file.empty() ? "默认 " + file : file));
output_status = text;
}
// Say what is actually in effect, including the parts that do not apply
// to the selected output mode, so nothing has to be greyed out.
set_text(GetDlgItem(m_hwnd, IDC_LABEL_OUTPUT_STATUS), output_status);
joc_decode::Settings effective = joc_settings::current();
effective.hrtf_source =
static_cast<joc_decode::HrtfSource>(values.hrtf_source);
effective.hrtf_file = values.hrtf_file;
const std::string file = joc_decode::resolve_hrtf_file(effective);
std::string hrtf_status =
std::string(joc_settings::hrtf_source_name(values.hrtf_source)) + ":";
if (values.hrtf_file.empty()) hrtf_status += "默认";
hrtf_status += "\n" +
(file.empty() ? std::string("无法确定默认路径") : file);
set_text(GetDlgItem(m_hwnd, IDC_LABEL_STATUS), hrtf_status);
char gain[96] = {};
if (values.gain_enabled) {
std::snprintf(gain, sizeof(gain), "\n增益开 %.2f dB", values.gain_db);
std::snprintf(gain, sizeof(gain), "增益开 %.2f dB", values.gain_db);
} else {
std::snprintf(gain, sizeof(gain), "\n增益关(输出不衰减)");
std::snprintf(gain, sizeof(gain), "增益关(输出不衰减)");
}
status += gain;
set_text(GetDlgItem(m_hwnd, IDC_LABEL_STATUS), status);
set_text(GetDlgItem(m_hwnd, IDC_LABEL_GAIN_STATUS), gain);
}
bool changed() const {
+23 -21
View File
@@ -7,11 +7,11 @@
// sample page uses (SDK\foobar2000\foo_sample\foo_sample.rc). A CAPTION statement
// implies WS_CAPTION, which would draw a second frame inside the host's and shift
// every control down by the caption height -- visible but not clickable.
IDD_JOC_PREFS DIALOGEX 0, 0, 330, 226
IDD_JOC_PREFS DIALOGEX 0, 0, 330, 273
STYLE DS_SETFONT | WS_CHILD
FONT 8, "Microsoft Sans Serif", 400, 0, 0x0
BEGIN
GROUPBOX "输出", IDC_GROUP_OUTPUT, 7, 5, 316, 64
GROUPBOX "输出", IDC_GROUP_OUTPUT, 7, 5, 316, 61
CONTROL "双耳(HRTF)", IDC_RADIO_BINAURAL, "Button",
BS_AUTORADIOBUTTON | WS_GROUP | WS_TABSTOP, 16, 20, 72, 10
CONTROL "扬声器布局", IDC_RADIO_SPEAKER, "Button",
@@ -22,31 +22,33 @@ BEGIN
LTEXT "双耳模式", IDC_LABEL_MODE, 172, 22, 42, 10
COMBOBOX IDC_COMBO_MODE, 216, 19, 96, 90,
CBS_DROPDOWNLIST | WS_VSCROLL | WS_TABSTOP
LTEXT "房间尾音(秒)", IDC_LABEL_TAIL, 172, 41, 46, 10
EDITTEXT IDC_EDIT_TAIL, 220, 38, 40, 13, ES_AUTOHSCROLL
LTEXT "房间尾音(秒)", IDC_LABEL_TAIL, 172, 41, 62, 10
EDITTEXT IDC_EDIT_TAIL, 238, 38, 40, 13, ES_AUTOHSCROLL
LTEXT "", IDC_LABEL_OUTPUT_STATUS, 16, 52, 300, 10
GROUPBOX "HRTF", IDC_GROUP_HRTF, 7, 74, 316, 84
GROUPBOX "HRTF", IDC_GROUP_HRTF, 7, 71, 316, 100
CONTROL "SOFA 文件", IDC_RADIO_HRTF_SOFA, "Button",
BS_AUTORADIOBUTTON | WS_GROUP | WS_TABSTOP, 16, 89, 56, 10
BS_AUTORADIOBUTTON | WS_GROUP | WS_TABSTOP, 16, 86, 56, 10
CONTROL "Rosella 个性化模型", IDC_RADIO_HRTF_ROSELLA, "Button",
BS_AUTORADIOBUTTON | WS_TABSTOP, 78, 89, 78, 10
LTEXT "HRTF 文件(留空=默认位置)", IDC_LABEL_HRTF, 16, 106, 92, 10
EDITTEXT IDC_EDIT_HRTF, 110, 103, 164, 13, ES_AUTOHSCROLL
PUSHBUTTON "浏览…", IDC_BROWSE_HRTF, 279, 103, 38, 13
BS_AUTORADIOBUTTON | WS_TABSTOP, 78, 86, 78, 10
LTEXT "HRTF 文件(留空=默认位置)", IDC_LABEL_HRTF, 16, 105, 106, 10
EDITTEXT IDC_EDIT_HRTF, 126, 102, 148, 13, ES_AUTOHSCROLL
PUSHBUTTON "浏览…", IDC_BROWSE_HRTF, 279, 102, 38, 13
LTEXT "留空时读组件目录下 HRTF\\binaural.sofa 或 binaural.personalized_headphone;SOFA 由渲染器内部编译,滤波组表已编入库中。",
IDC_LABEL_HRTF_HINT, 16, 122, 300, 30
IDC_LABEL_HRTF_HINT, 16, 121, 300, 20
LTEXT "", IDC_LABEL_STATUS, 16, 145, 300, 20
GROUPBOX "增益", IDC_GROUP_GAIN, 7, 162, 152, 56
GROUPBOX "增益", IDC_GROUP_GAIN, 7, 176, 316, 48
CONTROL "应用增益", IDC_CHECK_GAIN, "Button",
BS_AUTOCHECKBOX | WS_TABSTOP, 16, 178, 58, 10
LTEXT "dB", IDC_LABEL_GAIN, 80, 178, 12, 10
EDITTEXT IDC_EDIT_GAIN, 96, 175, 40, 13, ES_AUTOHSCROLL
BS_AUTOCHECKBOX | WS_TABSTOP, 16, 192, 58, 10
LTEXT "dB", IDC_LABEL_GAIN, 80, 192, 12, 10
EDITTEXT IDC_EDIT_GAIN, 96, 189, 40, 13, ES_AUTOHSCROLL
LTEXT "", IDC_LABEL_GAIN_STATUS, 145, 192, 171, 10
LTEXT "双耳渲染后可能超过 0 dBFS,可用负值衰减。",
IDC_LABEL_GAIN_HINT, 16, 194, 134, 20
IDC_LABEL_GAIN_HINT, 16, 208, 300, 10
GROUPBOX "其它", IDC_GROUP_MISC, 165, 162, 158, 56
LTEXT "ffmpeg.exe", IDC_LABEL_FFMPEG, 173, 178, 44, 10
EDITTEXT IDC_EDIT_FFMPEG, 216, 175, 62, 13, ES_AUTOHSCROLL
PUSHBUTTON "浏览…", IDC_BROWSE_FFMPEG, 282, 175, 35, 13
LTEXT "", IDC_LABEL_STATUS, 173, 194, 144, 20
GROUPBOX "其它", IDC_GROUP_MISC, 7, 229, 316, 38
LTEXT "ffmpeg.exe", IDC_LABEL_FFMPEG, 16, 245, 44, 10
EDITTEXT IDC_EDIT_FFMPEG, 64, 242, 210, 13, ES_AUTOHSCROLL
PUSHBUTTON "浏览…", IDC_BROWSE_FFMPEG, 279, 242, 38, 13
END
+2
View File
@@ -12,6 +12,7 @@
#define IDC_COMBO_MODE 2016
#define IDC_LABEL_TAIL 2017
#define IDC_EDIT_TAIL 2018
#define IDC_LABEL_OUTPUT_STATUS 2019
#define IDC_GROUP_HRTF 2020
#define IDC_RADIO_HRTF_SOFA 2021
@@ -26,6 +27,7 @@
#define IDC_LABEL_GAIN 2032
#define IDC_EDIT_GAIN 2033
#define IDC_LABEL_GAIN_HINT 2034
#define IDC_LABEL_GAIN_STATUS 2035
#define IDC_GROUP_MISC 2040
#define IDC_LABEL_FFMPEG 2041
+156 -47
View File
@@ -14,15 +14,33 @@
// --tail S binaural tail seconds (default 5)
// --ffmpeg PATH default ffmpeg
// --max-frames N stop after N output frames (0 = all)
// --max-input-frames N stop feeding the renderer after N syncframes
// --seek-to S seek to S seconds before reading (repeatable;
// only the last one before reading takes effect)
// --seek-at N:S seek to S seconds once N output frames have been
// written (repeatable; N = 0 seeks before reading)
// --total-seconds S the duration seek() is told about (0 = unknown)
// --length S cut the delivered stream at S seconds (the file's
// own duration; the renderer still renders its tail)
// --pipeline-delay N delay a seek compensates for, in samples
// --preroll-frames N syncframes fed before the target frame
// --cycles N render the file N times in this process, in sequence
// --parallel N render it N times at once, in separate threads
//
// The two constants a seek uses are overridable so they can be checked by measurement.
//
// --cycles and --parallel cover what one render per process cannot: the renderer and the
// compiled-HRTF cache live as long as the process does, and foobar2000 starts the next
// file while the current one is still being read.
//
// It exists because the plugin's engine (src/joc_decode.*) has no foobar2000
// dependency: the very code that plays in foobar2000 can be run here and its
// output compared against the reference renderer, which is what the bit-exactness
// contract is about. The WAV header mirrors the renderer's own writer: a simple
// IEEE-float header up to two channels, WAVE_FORMAT_EXTENSIBLE above that with a
// channel mask of zero.
// dependency: the very code that plays in foobar2000 can be run here, against the
// reference renderer or against the same engine's continuous render. The WAV header
// mirrors the renderer's own writer: a simple IEEE-float header up to two channels,
// WAVE_FORMAT_EXTENSIBLE above that with a channel mask of zero.
#include <cstdio>
#include <thread>
#include <cstdlib>
#include <cstring>
#include <string>
@@ -44,6 +62,16 @@ struct Options {
double tail_seconds = 5.0;
std::uint64_t max_frames = 0;
std::uint64_t max_input_frames = 0;
std::vector<double> seeks_before; // --seek-to
std::vector<std::pair<std::uint64_t, double>> seeks_at; // --seek-at
double total_seconds = 0.0;
double length_seconds = 0.0; // --length: cut the delivered stream here
std::uint32_t pipeline_delay = joc_decode::kJocSeekPipelineDelaySamples;
std::uint32_t preroll_frames = joc_decode::kJocSeekPrerollFrames;
bool container = false;
unsigned audio_index = 0;
unsigned cycles = 0; // --cycles N: N renders in one process, in sequence
unsigned parallel = 0; // --parallel N: N renders at once, in one process
};
void put_u16(std::string* out, unsigned value) {
@@ -111,6 +139,26 @@ bool parse(int argc, char** argv, Options* options) {
else if (arg == "--tail") options->tail_seconds = std::atof(next().c_str());
else if (arg == "--max-frames") options->max_frames = std::strtoull(next().c_str(), nullptr, 10);
else if (arg == "--max-input-frames") options->max_input_frames = std::strtoull(next().c_str(), nullptr, 10);
else if (arg == "--seek-to") options->seeks_before.push_back(std::atof(next().c_str()));
else if (arg == "--total-seconds") options->total_seconds = std::atof(next().c_str());
else if (arg == "--length") options->length_seconds = std::atof(next().c_str());
else if (arg == "--pipeline-delay") options->pipeline_delay = static_cast<std::uint32_t>(std::strtoul(next().c_str(), nullptr, 10));
else if (arg == "--preroll-frames") options->preroll_frames = static_cast<std::uint32_t>(std::strtoul(next().c_str(), nullptr, 10));
else if (arg == "--container") options->container = true;
else if (arg == "--cycles") options->cycles = static_cast<unsigned>(std::strtoul(next().c_str(), nullptr, 10));
else if (arg == "--parallel") options->parallel = static_cast<unsigned>(std::strtoul(next().c_str(), nullptr, 10));
else if (arg == "--audio-index") options->audio_index = static_cast<unsigned>(std::strtoul(next().c_str(), nullptr, 10));
else if (arg == "--seek-at") {
const std::string value = next();
const std::string::size_type colon = value.find(':');
if (colon == std::string::npos) {
std::fprintf(stderr, "render_harness: --seek-at wants <frames>:<seconds>\n");
return false;
}
options->seeks_at.emplace_back(
std::strtoull(value.substr(0, colon).c_str(), nullptr, 10),
std::atof(value.substr(colon + 1).c_str()));
}
else if (arg == "--help" || arg == "-h") return false;
else {
std::fprintf(stderr, "render_harness: unknown argument %s\n", arg.c_str());
@@ -120,47 +168,26 @@ bool parse(int argc, char** argv, Options* options) {
return !options->input.empty() && !options->output.empty();
}
} // namespace
int main(int argc, char** argv) {
Options options;
if (!parse(argc, argv, &options)) {
std::fprintf(stderr,
"usage: render_harness --input <eac3> --output <wav> [--mode binaural|speaker]\n"
" [--layout NAME] [--hrtf-source sofa|rosella] [--hrtf PATH]\n"
" [--gain-db X] [--tail S] [--ffmpeg path] [--max-frames N]\n"
" [--max-input-frames N]\n");
return 2;
}
joc_decode::Settings settings;
settings.output = (options.mode == "speaker") ? joc_decode::Output::kSpeaker
: joc_decode::Output::kBinaural;
settings.speaker_layout = options.layout;
settings.hrtf_source = (options.hrtf_source == "rosella") ? joc_decode::HrtfSource::kRosella
: joc_decode::HrtfSource::kSofa;
settings.hrtf_file = options.hrtf;
settings.gain_db = options.gain_db;
settings.tail_seconds = options.tail_seconds;
settings.ffmpeg_path = options.ffmpeg;
settings.input_frame_limit = options.max_input_frames;
// One render: everything the component does for one file from start to stop.
// Returns false with *error set when the engine reports a failure.
bool render_once(const joc_decode::Settings& settings, const Options& options,
const std::string& output, std::string* error) {
joc_decode::Engine engine;
std::string error;
if (!engine.start(options.input, settings, &error)) {
std::fprintf(stderr, "render_harness: engine start failed: %s\n", error.c_str());
return 1;
}
if (!engine.start(options.input, settings, error)) return false;
const unsigned channels = engine.channels();
if (channels == 0) {
std::fprintf(stderr, "render_harness: engine reported zero channels\n");
return 1;
*error = "engine reported zero channels";
return false;
}
for (const double seconds : options.seeks_before) {
if (!engine.seek(seconds, options.total_seconds, error)) return false;
}
std::FILE* file = std::fopen(options.output.c_str(), "wb");
std::FILE* file = std::fopen(output.c_str(), "wb");
if (file == nullptr) {
std::fprintf(stderr, "render_harness: cannot write %s\n", options.output.c_str());
return 1;
*error = "cannot write " + output;
return false;
}
// The header carries the length, so write a placeholder and come back to it.
const std::string header = wav_header(channels, 48000, 0);
@@ -170,19 +197,27 @@ int main(int argc, char** argv) {
std::vector<float> buffer(kChunk * channels);
std::uint64_t frames_written = 0;
double peak = 0.0;
std::vector<char> applied(options.seeks_at.size(), 0);
for (;;) {
for (std::size_t i = 0; i < options.seeks_at.size(); ++i) {
if (applied[i] != 0 || options.seeks_at[i].first > frames_written) continue;
applied[i] = 1;
if (!engine.seek(options.seeks_at[i].second, options.total_seconds, error)) {
std::fclose(file);
return false;
}
}
std::size_t want = kChunk;
if (options.max_frames != 0) {
if (frames_written >= options.max_frames) break;
const std::uint64_t left = options.max_frames - frames_written;
if (left < want) want = static_cast<std::size_t>(left);
}
const std::size_t frames = engine.read(buffer.data(), want, &error);
const std::size_t frames = engine.read(buffer.data(), want, error);
if (frames == 0) {
if (!error.empty()) {
std::fprintf(stderr, "render_harness: read failed: %s\n", error.c_str());
if (!error->empty()) {
std::fclose(file);
return 1;
return false;
}
break;
}
@@ -201,9 +236,83 @@ int main(int argc, char** argv) {
std::fwrite(final_header.data(), 1, final_header.size(), file);
std::fclose(file);
std::printf("render_harness: %s -> %s\n", options.input.c_str(), options.output.c_str());
std::printf(" channels=%u frames=%llu samples_per_channel=%llu peak=%.9f\n", channels,
static_cast<unsigned long long>(frames_written),
std::printf("render_harness: %s -> %s\n", options.input.c_str(), output.c_str());
std::printf(" channels=%u frames=%llu peak=%.9f\n", channels,
static_cast<unsigned long long>(frames_written), peak);
return 0;
return true;
}
std::string numbered(const std::string& output, unsigned index) {
return output + "." + std::to_string(index) + ".wav";
}
} // namespace
int main(int argc, char** argv) {
Options options;
if (!parse(argc, argv, &options)) {
std::fprintf(stderr,
"usage: render_harness --input <eac3> --output <wav> [--mode binaural|speaker]\n"
" [--layout NAME] [--hrtf-source sofa|rosella] [--hrtf PATH]\n"
" [--gain-db X] [--tail S] [--ffmpeg path] [--max-frames N]\n"
" [--max-input-frames N] [--seek-to S] [--seek-at N:S]\n"
" [--total-seconds S] [--pipeline-delay N] [--preroll-frames N]\n"
" [--container] [--audio-index N] [--length S]\n"
" [--cycles N] [--parallel N]\n");
return 2;
}
joc_decode::Settings settings;
settings.output = (options.mode == "speaker") ? joc_decode::Output::kSpeaker
: joc_decode::Output::kBinaural;
settings.speaker_layout = options.layout;
settings.hrtf_source = (options.hrtf_source == "rosella") ? joc_decode::HrtfSource::kRosella
: joc_decode::HrtfSource::kSofa;
settings.hrtf_file = options.hrtf;
settings.gain_db = options.gain_db;
settings.tail_seconds = options.tail_seconds;
settings.ffmpeg_path = options.ffmpeg;
settings.input_frame_limit = options.max_input_frames;
settings.pipeline_delay_samples = options.pipeline_delay;
settings.seek_preroll_frames = options.preroll_frames;
settings.length_seconds = options.length_seconds;
settings.input_kind = options.container ? joc_decode::InputKind::kContainer
: joc_decode::InputKind::kBare;
settings.audio_index = options.audio_index;
std::string error;
if (options.parallel != 0) {
// The same file rendered by several engines at once, which is what a track
// change looks like from the engine's side: foobar2000 starts the next file
// while the current one is still being read.
std::vector<std::thread> threads;
std::vector<std::string> errors(options.parallel);
for (unsigned i = 0; i < options.parallel; ++i) {
threads.emplace_back([&, i] {
if (!render_once(settings, options, numbered(options.output, i), &errors[i])) {
std::fprintf(stderr, "render_harness: parallel %u: %s\n", i,
errors[i].c_str());
}
});
}
for (std::thread& thread : threads) thread.join();
for (const std::string& text : errors) {
if (!text.empty()) return 1;
}
return 0;
}
const unsigned cycles = (options.cycles == 0) ? 1 : options.cycles;
for (unsigned cycle = 0; cycle < cycles; ++cycle) {
// Every cycle is a fresh engine in the same process: the compiled-HRTF cache
// and everything else the renderer keeps per process is reused, exactly as it
// is when a second file is played in the same foobar2000 session.
error.clear();
const std::string output = (cycles == 1) ? options.output : numbered(options.output, cycle);
if (!render_once(settings, options, output, &error)) {
std::fprintf(stderr, "render_harness: cycle %u: %s\n", cycle, error.c_str());
return 1;
}
}
return 0;
}
+14 -8
View File
@@ -4,15 +4,16 @@
# pwsh -File tools/deploy.ps1 -TestBed D:\fb2k -Platform x64
#
# Where the component has to go depends on the foobar2000 layout, and getting it
# wrong is silent: foobar2000 simply never calls LoadLibrary on the file. Two
# layouts are in the wild:
# wrong is silent: foobar2000 simply never calls LoadLibrary on the file.
#
# profile-relative <app>\profile\user-components\<name>\<name>.dll
# foobar2000 1.6.19 portable and 2.x
# app-relative <app>\user-components\<name>\<name>.dll
# older / repacked 1.6 installs whose profile is <app>\configuration
# 1.6 <app>\profile\user-components\<name>\<name>.dll
# (portable mode; without it the profile is in %APPDATA%)
# 2.0 and up <app>\user-components\<name>\<name>.dll
# even in portable mode, whose profile is <app>\profile
# older 1.6 <app>\user-components\<name>\<name>.dll
# repacked installs whose profile is <app>\configuration
#
# The per-component subdirectory is required in both: a DLL lying directly in
# The per-component subdirectory is required in every case: a DLL lying directly in
# user-components\ is not scanned. Installing a .fb2k-component package through
# foobar2000 itself ends up doing the same thing.
[CmdletBinding()]
@@ -31,11 +32,16 @@ if (-not (Test-Path -LiteralPath (Join-Path $TestBed 'foobar2000.exe'))) {
throw "no foobar2000.exe in $TestBed"
}
# The core version decides the layout, and its own profile records it.
$versionFile = Join-Path $TestBed 'profile\version.txt'
$version = if (Test-Path -LiteralPath $versionFile) { (Get-Content -LiteralPath $versionFile -Raw).Trim() } else { '' }
if (-not $Target) {
$appRelative = Test-Path -LiteralPath (Join-Path $TestBed 'configuration')
if ($version -match 'v(\d+)\.') { $appRelative = ([int]$Matches[1] -ge 2) }
$root = if ($appRelative) { $TestBed } else { Join-Path $TestBed 'profile' }
$Target = Join-Path $root 'user-components'
Write-Host ("layout: {0} ({1})" -f $(if ($appRelative) { 'app-relative' } else { 'profile-relative' }), $root)
Write-Host ("layout: {0} ({1}){2}" -f $(if ($appRelative) { 'app-relative' } else { 'profile-relative' }), $root, $(if ($version) { ", $version" } else { '' }))
}
# Portable profile: config stays inside the test bed instead of %APPDATA%.