Merging multiple fine-tuned Whisper models reduces word error rate on dysarthric speech by 12-16% relative to standard fine-tuning, with gains on long audio and low-data settings.
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Automatic Speech Recognition (ASR) has advanced with Speech Foundation Models (SFMs), yet performance degrades on dysarthric speech due to variability and limited data. This study as part of the submission to the Speech Accessibility challenge, explored model merging to improve ASR generalization using Whisper as the base SFM. We compared fine-tuning with single-trajectory merging, combining models from one fine-tuning path, and multi-run merging, merging independently trained models. Our best multi-run merging approach achieved a 12% relative decrease of WER over classic fine-tuning, and a 16.2% relative decrease on long-form audios, a major loss contributor in dysarthric ASR. Merging more and more models led to continuous gains, remained effective in low-data regimes, and generalized across model architectures. These results highlight model merging as an easily replicable adaptation method that consistently improves ASR without additional inference cost or hyperparameter tuning.
citation-role summary
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
Merging multiple fine-tuned Whisper models reduces word error rate on dysarthric speech by 12-16% relative to standard fine-tuning, with gains on long audio and low-data settings.