Files
JustOneCacophony/docs/binaural.en.md
T

7.5 KiB
Raw Blame History

JustOneCacophony — Binaural Rendering Mathematics

中文 · Back to README

This document defines the pcm16 + ID11/OAMD → stereo calculation. The path begins after object reconstruction and does not pass through ADM BWF or AXML.

1. Signal path and notation

LFE + 15 object PCM channels
  → 64-band QMF analysis
  → 77-band hybrid analysis
  → per-object geometry, transfer functions, and room send
  → direct accumulation + room network
  → hybrid synthesis
  → QMF synthesis
  → 961-sample latency compensation
  → stereo WAV
Symbol Meaning
s=0\ldots15 input source; source 0 is LFE
e\in\{L,R\} output ear
k=0\ldots63 QMF band
h=0\ldots76 hybrid band
j=0\ldots35 direction-basis term
m 64-sample QMF slot

A control block is

N_b=512=8\times64,

and an input frame is

N_f=1536=3N_b.

All filter and room state continues across frame boundaries.

2. QMF analysis

Let a_{p,\ell} be the fixed 64×10 polyphase coefficients and r_{s,\ell,p}[m] the current and previous nine phase vectors:


E_{s,p}[m]=\sum_{\ell\text{ even}}a_{p,\ell}r_{s,\ell,p}[m],

O_{s,p}[m]=\sum_{\ell\text{ odd}}a_{p,\ell}r_{s,\ell,p}[m].

Define


\mathcal Q(v)_k=
\operatorname{FFT}_{128}
\left([v[p]e^{-j\pi p/128}]_{p=0}^{63},0_{64}\right)_k
 e^{-j3\pi(k+1/2)/128}.

The complex QMF output is


X_{s,k}[m]=\mathcal Q(O_s)_k+j(-1)^k\mathcal Q(E_s)_k.

3. Hybrid analysis

The lowest three QMF bands are split into sixteen hybrid bands by a 13-slot FIR:


H_{s,h,o}[m]
=
\sum_{p=0}^{2}\sum_{i=0}^{1}\sum_{\ell=0}^{12}
X_{s,p,i}[m-\ell]K_{p,i,\ell,h,o},
\qquad h=0\ldots15.

The remaining bands are delayed QMF bands 3..63:


H_{s,16+q}[m]=X_{s,3+q}[m-6],
\qquad q=0\ldots60.

4. OAMD coordinates and time

The Q15 object fields are restored to their discrete grids:


u_1=\min\left(1,\frac{\operatorname{round}(62q_1/32767)}{62}\right),$$

u_2=\min\left(1,\frac{\operatorname{round}(62q_2/32767)}{62}\right),$$


u_3=\operatorname{clip}\left(
\frac{\operatorname{round}(15q_3/32767)}{15},-1,1\right),$$

(X,Y,Z)=(2u_1-1,\ 1-2u_2,\ u_3).



An update is coded at

n_{\mathrm{coded}} =n_{\mathrm{frame}}+n_{\mathrm{outer}}+n_{\mathrm{block}}.



The first valid state is the position at sample 0. Later updates add the object delay $D_o=1473$. For $R>64$:

n_{\mathrm{start}}=n_{\mathrm{coded}}+D_o+64,$$

R_{\mathrm{eff}}=R-64,

\mathbf p[n]=(1-\alpha)\mathbf p_0+\alpha\mathbf p_1,
\qquad
\alpha=\frac{n-n_{\mathrm{start}}}{R_{\mathrm{eff}}}.

The position is evaluated at each 512-sample block boundary.

5. Distance profile and direction

Each Near, Mid, or Far profile contains six bounds, distance scale D, inverse scale D^{-1}, three axis scales, and minimum radius \rho_{\min}.

After axis conversion and scale:


\mathbf s=(a_zq_f,a_xq_l,a_yq_v).

A single ray factor \lambda\le1 keeps the point inside the profile bounds:


\mathbf s'=\lambda\mathbf s.

Then


\rho=\|\mathbf s'\|_2,
\quad
\rho_c=\max(\rho,\rho_{\min}),
\quad
\alpha=\rho/\rho_c,

\mathbf d=\mathbf s'/\rho,
\qquad
R=D\rho.

6. Direction basis and ear paths

The direction is expanded into a fixed 36-term polynomial basis:


\mathbf b(\mathbf d)=
[1,x,y,z,x^2-\tfrac13,xy,xz,y^2-\tfrac13,yz,\ldots]^T.

For ear offset e:


\epsilon=\frac{eD^{-1}}{\rho_c},

\mathbf d_{\mp}=
\frac{(x,y\mp\epsilon,z)}{\|(x,y\mp\epsilon,z)\|_2}.

The normalized paths are


\ell_{\mp}=\rho_c\sqrt{x^2+(y\mp\epsilon)^2+z^2}.

A model direction vector may add a non-negative path correction:


\ell'_e=\ell_e+
\max(\mathbf v_e^T\mathbf b_e,0)\,2cD^{-1}.

The interaural delay is


\tau=|\ell'_+-\ell'_-|D\frac{48000}{343.3}\alpha.

The longer path receives the hybrid phase

P_h=e^{j\omega_h\tau}.

7. Direction fields and direct gains

Each ear has a 77×36 complex field:


C_{e,h}(\mathbf d_e)=
\sum_{j=0}^{35}F_{e,h,j}b_j(\mathbf d_e).

Path weights are


w_L=\frac{\ell_+}{\sqrt{\ell_-^2+\ell_+^2}},
\qquad
w_R=\frac{\ell_-}{\sqrt{\ell_-^2+\ell_+^2}}.

For effective distance R_e=\rho s_dD, Mid and Far use


g_c=\frac{1}{\sqrt{1+s_rR_e^2}},
\qquad
g_{\mathrm{room}}=R_eg_c.

Near uses g_c=1 and g_{\mathrm{room}}=0. With field term zero denoted by C^{(0)}:


G_{L,h}=g_c[C_{L,h}w_L\alpha+C_{L,h}^{(0)}c_L(1-\alpha)],

G_{R,h}=g_c[C_{R,h}w_R\alpha+C_{R,h}^{(0)}c_R(1-\alpha)].

8. LFE

LFE bypasses ordinary-object geometry:


G_{L,h}^{\mathrm{LFE}}=G_{R,h}^{\mathrm{LFE}}=
\begin{cases}
g_h,&0\le h<16,\\0,&16\le h<77.
\end{cases}
 2.60290003,  1.80741799,  0.659342408, -0.0275855921,
-0.105803289, -0.0699509233, 0.0749056414, -0.00919809937,
 0.00349014648,-0.0158600751,-0.000723021978,0.00188189559,
-0.000421735429,0.0000329252762,0.0000317397971,0.000000580376991

Its room send is zero.

9. Source accumulation and room network

Direct output and room input are


Y^{\mathrm{direct}}_{e,h}=
\sum_{s=0}^{15}H_{s,h}G_{s,e,h},

U_h=\sum_{s=1}^{15}H_{s,h}g_{\mathrm{room},s}.

The room input is scaled by 0.70710677. Each all-pass stage uses

r[n]=x[n]-ad[n], y[n]=ar[n]+d[n].

For the four-branch delay network:


\mathbf b_h[m]=U_h[m]\mathbf1+M\mathbf d_h[m],
m_{h,i}[m]=f_{h,i}b_{h,i}[m].

The main tap, optional extra taps, and ear output matrices produce


Y^{\mathrm{room}}_{e,h}[m]=
\sum_{i=0}^{3}O_{e,h,i}z_{h,i}[m].

The final hybrid signal is

Y_{e,h}=Y^{\mathrm{direct}}_{e,h}+Y^{\mathrm{room}}_{e,h}.

The Python backend uses a finite complex FIR/overlap-add realization. The C++ backend keeps the recursive room state directly.

10. Hybrid and QMF synthesis

Hybrid synthesis is a 154-entry sparse map. For an entry (h,i,k,o,w):

Q_{e,k,o}[m]\mathrel{+}=Y_{e,h,i}[m]w.

The complex QMF vector is flattened to


\mathbf q_e=[\Re Q_{e,0},\Im Q_{e,0},\ldots,\Re Q_{e,63},\Im Q_{e,63}]^T.

Rank-four features and ten-slot synthesis are

f_{e,p,r}[m]=\mathbf b_{p,r}^T\mathbf q_e[m],

y_e[64m+p]=
\sum_{\ell=0}^{9}\sum_{r=0}^{3}
t_{p,\ell,r}f_{e,p,r}[m-\ell].

11. Latency, tail, and precision

The filterbank latency is 961 samples and is removed once at the beginning of the continuous stream. Zero input is then processed to release filterbank and room state. Tail trimming keeps the final sample satisfying


\max(|y_L[n]|,|y_R[n]|)>10^{-8},

while never shortening the output below the source PCM length.

All internal state, geometry, field products, room processing, source accumulation, and tail processing use float64/complex128. Conversion to float32 or PCM24 occurs only in the final writer.

12. Backends and model path

Python and C++ use the same fixed tables, parsed model parameters, 512-sample control timeline, direct gains, room sends, latency compensation, and tail policy.

The C++ backend owns QMF, hybrid, recursive room, and synthesis state. Python supplies parsed parameters and per-block gains.

The default model path is

HRTF/binaural.personalized_headphone

Override it with --personalized-headphone PATH.

A SOFA FIR cannot be converted into this parameter model by array rearrangement alone. A conversion requires fitting the direction fields, ITD, distance profiles, ear geometry, and room parameters.

13. Scope

The current path covers fifteen point objects and one special LFE source. Extent, spread, diffuse, divergence, channel lock, and unsupported OAMD element variants are outside this model.