A two-stage Whisper-plus-FlanT5 system with diversity-based hypothesis selection reduces dysarthric speech WER from 11.60% to 7.34% on the development set, while single-word recognition stays at 63.08% WER.
Our approach first uses ASR models to generate multiple transcription hypotheses, captur- ing different interpretations of the acoustic input
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Exploring Generative Error Correction for Dysarthric Speech Recognition
A two-stage Whisper-plus-FlanT5 system with diversity-based hypothesis selection reduces dysarthric speech WER from 11.60% to 7.34% on the development set, while single-word recognition stays at 63.08% WER.