WhisperFlow combines a learned 'hush word', beam pruning, and CPU/GPU pipelining to cut streaming Whisper latency by 1.6x-4.7x on client devices with near-unchanged accuracy.
Accessed: 2024-11-3
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
WhisperFlow: speech foundation models in real time
WhisperFlow combines a learned 'hush word', beam pruning, and CPU/GPU pipelining to cut streaming Whisper latency by 1.6x-4.7x on client devices with near-unchanged accuracy.