Pith. sign in

REVIEW 1 cited by

WER We Stand: Benchmarking Urdu ASR Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.11252 v3 pith:6YQO53WX submitted 2024-09-17 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechurdumodelsconversationaldatasetbenchmarkingerrorevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a comprehensive evaluation of Urdu Automatic Speech Recognition (ASR) models. We analyze the performance of three ASR model families: Whisper, MMS, and Seamless-M4T using Word Error Rate (WER), along with a detailed examination of the most frequent wrong words and error types including insertions, deletions, and substitutions. Our analysis is conducted using two types of datasets, read speech and conversational speech. Notably, we present the first conversational speech dataset designed for benchmarking Urdu ASR models. We find that seamless-large outperforms other ASR models on the read speech dataset, while whisper-large performs best on the conversational speech dataset. Furthermore, this evaluation highlights the complexities of assessing ASR models for low-resource languages like Urdu using quantitative metrics alone and emphasizes the need for a robust Urdu text normalization system. Our findings contribute valuable insights for developing robust ASR systems for low-resource languages like Urdu.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Assessing the Feasibility of Lightweight Whisper Models for Low-Resource Urdu Transcription

    cs.CL 2025-08 conditional novelty 3.0 of 10

    Whisper-Small achieves a 33.68% word error rate on a 36-sample Urdu dataset, outperforming Tiny (67.08%) and Base (53.67%) in zero-shot transcription.

Pith tools