Merging multiple fine-tuned Whisper models reduces word error rate on dysarthric speech by 12-16% relative to standard fine-tuning, with gains on long audio and low-data settings.
As shown in Fig- ure 1, MAST outperformed standard fine-tuning, demonstrat- ing the benefits of weight averaging along a single optimization 2https://github.com/openai/whisper path
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
Merging multiple fine-tuned Whisper models reduces word error rate on dysarthric speech by 12-16% relative to standard fine-tuning, with gains on long audio and low-data settings.