Pith. sign in

REVIEW 2 cited by

Whispering in Amharic: Fine-tuning Whisper for Low-resource Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.18485 v2 pith:5OPCBCOH submitted 2025-03-24 cs.CL cs.LG

classification cs.CLcs.LG
keywords amharicdatamodelfine-tuningfleurslow-resourcewhisperdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work explores fine-tuning OpenAI's Whisper automatic speech recognition (ASR) model for Amharic, a low-resource language, to improve transcription accuracy. While the foundational Whisper model struggles with Amharic due to limited representation in its training data, we fine-tune it using datasets like Mozilla Common Voice, FLEURS, and the BDU-speech dataset. The best-performing model, Whispersmall-am, significantly improves when finetuned on a mix of existing FLEURS data and new, unseen Amharic datasets. Training solely on new data leads to poor performance, but combining it with FLEURS data reinforces the model, enabling better specialization in Amharic. We also demonstrate that normalizing Amharic homophones significantly enhances Word Error Rate (WER) and Bilingual Evaluation Understudy (BLEU) scores. This study underscores the importance of fine-tuning strategies and dataset composition for improving ASR in low-resource languages, providing insights for future Amharic speech recognition research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Assamese and Hindi, selected by acoustic and typological similarity to Warlpiri, cut Whisper WER/CER most; acoustic similarity best predicts fine-tuning gains, inventory/typology zero-shot.

  2. A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Fine-tuning Whisper-large-v2 on 10,000 hours of synthesized Mandarin plus small real English/code-switching sets yields Twister, cutting mixed error rate by up to 56% on code-switching and 19% on Taiwanese Mandarin.

Pith tools