Pith. sign in

REVIEW 1 cited by

N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.02902 v1 pith:P7NM2H4F submitted 2023-06-05 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords whisperarabicspeechunderconditionsdatadialectsdiverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Whisper, the recently developed multilingual weakly supervised model, is reported to perform well on multiple speech recognition benchmarks in both monolingual and multilingual settings. However, it is not clear how Whisper would fare under diverse conditions even on languages it was evaluated on such as Arabic. In this work, we address this gap by comprehensively evaluating Whisper on several varieties of Arabic speech for the ASR task. Our evaluation covers most publicly available Arabic speech data and is performed under n-shot (zero-, few-, and full) finetuning. We also investigate the robustness of Whisper under completely novel conditions, such as in dialect-accented standard Arabic and in unseen dialects for which we develop evaluation data. Our experiments show that although Whisper zero-shot outperforms fully finetuned XLS-R models on all datasets, its performance deteriorates significantly in the zero-shot setting for five unseen dialects (i.e., Algeria, Jordan, Palestine, UAE, and Yemen).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Open Universal Arabic ASR Leaderboard

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A new Arabic ASR leaderboard ranks 14 open-source models on five multi-dialect datasets and analyzes robustness, speaker bias, and efficiency.

Pith tools