Pith. sign in

REVIEW 1 cited by

Using fine-tuning and min lookahead beam search to improve Whisper

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.10299 v1 pith:W7BHIWJF submitted 2023-09-19 eess.AS cs.CLcs.LGcs.SD

classification eess.AScs.CLcs.LGcs.SD
keywords whisperbeamsearchalgorithmfine-tuninglanguageslookaheadcompared
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The performance of Whisper in low-resource languages is still far from perfect. In addition to a lack of training data on low-resource languages, we identify some limitations in the beam search algorithm used in Whisper. To address these issues, we fine-tune Whisper on additional data and propose an improved decoding algorithm. On the Vietnamese language, fine-tuning Whisper-Tiny with LoRA leads to an improvement of 38.49 in WER over the zero-shot Whisper-Tiny setting which is a further reduction of 1.45 compared to full-parameter fine-tuning. Additionally, by using Filter-Ends and Min Lookahead decoding algorithms, the WER reduces by 2.26 on average over a range of languages compared to standard beam search. These results generalise to larger Whisper model sizes. We also prove a theorem that Min Lookahead outperforms the standard beam search algorithm used in Whisper.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Advancing STT for Low-Resource Real-World Speech

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Fine-tuning Whisper on a new 303-hour corpus of spontaneous Swiss German broadcast speech improves WER from 21% to 17% for large-v3 and outperforms models trained on sentence-level corpora.

Pith tools