Using Mimi neural codec features with label-delayed training reduces endpoint cutoff errors by 42.7% (single-stream) and 37.5% (two-stream) at 160 ms median latency.
Improved End- of-Query Detection for Streaming Speech Recognition,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
Using Mimi neural codec features with label-delayed training reduces endpoint cutoff errors by 42.7% (single-stream) and 37.5% (two-stream) at 160 ms median latency.