Personalized LoRA fine-tuning of Whisper on both read and spontaneous stuttered speech from one speaker cuts word error rate roughly in half compared to a generalized model.
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven technologies inaccessible to people who stutter. The variability of disfluencies across speakers and contexts further complicates ASR training, compounded by limited annotated stuttered speech data. In this paper, we investigate fine-tuning ASRs for stuttered speech, comparing generalized models (trained across multiple speakers) to personalized models tailored to individual speech characteristics. Using a diverse range of voice-AI scenarios, including virtual assistants and video interviews, we evaluate how personalization affects transcription accuracy. Our findings show that personalized ASRs significantly reduce word error rates, especially in spontaneous speech, highlighting the potential of tailored models for more inclusive voice technologies.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
Personalized LoRA fine-tuning of Whisper on both read and spontaneous stuttered speech from one speaker cuts word error rate roughly in half compared to a generalized model.