A causal flow-matching model renders streaming binaural speech from mono audio and speaker/listener poses, reaching a 42% confusion rate against real recordings in an AB test.
We initialize some variables, including t, ϕt(z), and δ
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models
A causal flow-matching model renders streaming binaural speech from mono audio and speaker/listener poses, reaching a 42% confusion rate against real recordings in an AB test.