REVIEW 6 cited by
Fine-tuning Whisper on Low-Resource Languages for Real-World Applications
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents a new approach to fine-tuning OpenAI's Whisper model for low-resource languages by introducing a novel data generation method that converts sentence-level data into a long-form corpus, using Swiss German as a case study. Non-sentence-level data, which could improve the performance of long-form audio, is difficult to obtain and often restricted by copyright laws. Our method bridges this gap by transforming more accessible sentence-level data into a format that preserves the model's ability to handle long-form audio and perform segmentation without requiring non-sentence-level data. Our data generation process improves performance in several real-world applications and leads to the development of a new state-of-the-art speech-to-text (STT) model for Swiss German. We compare our model with a non-fine-tuned Whisper and our previous state-of-the-art Swiss German STT models, where our new model achieves higher BLEU scores. Our results also indicate that the proposed method is adaptable to other low-resource languages, supported by written guidance and code that allows the creation of fine-tuned Whisper models, which keep segmentation capabilities and allow the transcription of longer audio files using only sentence-level data with high quality.
Forward citations
Cited by 6 Pith papers
-
End-to-End Spoken Grammatical Error Correction
End-to-end Whisper models, trained with 2,500 hours of pseudo-labeled speech, fluent prompts, aligned references, and confidence filtering, outperform cascaded systems on spoken grammatical error correction and feedback.
-
Advancing STT for Low-Resource Real-World Speech
Fine-tuning Whisper on a new 303-hour corpus of spontaneous Swiss German broadcast speech improves WER from 21% to 17% for large-v3 and outperforms models trained on sentence-level corpora.
-
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
Pseudo-labelling and prompting with fluent transcriptions improve end-to-end spoken grammatical error correction and feedback for Whisper-based models, but the benefits depend on model size.
-
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
Fine-tuning Whisper-large-v2 on 10,000 hours of synthesized Mandarin plus small real English/code-switching sets yields Twister, cutting mixed error rate by up to 56% on code-switching and 19% on Taiwanese Mandarin.
-
Voice Adaptation for Swiss German
Fine-tuning XTTS-v2 on about 5,000 hours of weakly labeled Swiss podcast audio produces a voice adaptation model that renders Standard German text in seven Swiss German dialect regions with near-reference quality in h...
-
Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models
Fine-tuning Whisper-Small on 3,520 Assamese clips from Common Voice cuts word error rate from 201% to 44% and character error rate from 191% to 13%.
Discussion (0). Sign in to comment.