Pith. sign in

REVIEW 6 cited by

Fine-tuning Whisper on Low-Resource Languages for Real-World Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.15726 v3 pith:LYVW4HWQ submitted 2024-12-20 cs.CL eess.AS

classification cs.CLeess.AS
keywords datamodelwhisperaudiogermanlanguageslong-formlow-resource
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a new approach to fine-tuning OpenAI's Whisper model for low-resource languages by introducing a novel data generation method that converts sentence-level data into a long-form corpus, using Swiss German as a case study. Non-sentence-level data, which could improve the performance of long-form audio, is difficult to obtain and often restricted by copyright laws. Our method bridges this gap by transforming more accessible sentence-level data into a format that preserves the model's ability to handle long-form audio and perform segmentation without requiring non-sentence-level data. Our data generation process improves performance in several real-world applications and leads to the development of a new state-of-the-art speech-to-text (STT) model for Swiss German. We compare our model with a non-fine-tuned Whisper and our previous state-of-the-art Swiss German STT models, where our new model achieves higher BLEU scores. Our results also indicate that the proposed method is adaptable to other low-resource languages, supported by written guidance and code that allows the creation of fine-tuned Whisper models, which keep segmentation capabilities and allow the transcription of longer audio files using only sentence-level data with high quality.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. End-to-End Spoken Grammatical Error Correction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    End-to-end Whisper models, trained with 2,500 hours of pseudo-labeled speech, fluent prompts, aligned references, and confidence filtering, outperform cascaded systems on spoken grammatical error correction and feedback.

  2. Advancing STT for Low-Resource Real-World Speech

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Fine-tuning Whisper on a new 303-hour corpus of spontaneous Swiss German broadcast speech improves WER from 21% to 17% for large-v3 and outperforms models trained on sentence-level corpora.

  3. Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Pseudo-labelling and prompting with fluent transcriptions improve end-to-end spoken grammatical error correction and feedback for Whisper-based models, but the benefits depend on model size.

  4. A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Fine-tuning Whisper-large-v2 on 10,000 hours of synthesized Mandarin plus small real English/code-switching sets yields Twister, cutting mixed error rate by up to 56% on code-switching and 19% on Taiwanese Mandarin.

  5. Voice Adaptation for Swiss German

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning XTTS-v2 on about 5,000 hours of weakly labeled Swiss podcast audio produces a voice adaptation model that renders Standard German text in seven Swiss German dialect regions with near-reference quality in h...

  6. Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Fine-tuning Whisper-Small on 3,520 Assamese clips from Common Voice cuts word error rate from 201% to 44% and character error rate from 191% to 13%.

Pith tools