Multilingual ASR models show 39.7-297% zero-shot WER on Pashto public data, Whisper models output correct script in under 0.8% of cases, and fine-tuned models degrade to 32.5-59% WER on out-of-domain sets.
Advocating character error rate for multilingual ASR evaluation
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2years
2026 2verdicts
CONDITIONAL 2representative citing papers
A 536-hour unscripted telephonic ASR benchmark covering 15 Indian languages and 139 regional clusters, with multi-reference transcripts for spelling variation and district-level performance analysis.
citing papers explorer
-
Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation
Multilingual ASR models show 39.7-297% zero-shot WER on Pashto public data, Whisper models output correct script in under 0.8% of cases, and fine-tuned models degrade to 32.5-59% WER on out-of-domain sets.
-
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
A 536-hour unscripted telephonic ASR benchmark covering 15 Indian languages and 139 regional clusters, with multi-reference transcripts for spelling variation and district-level performance analysis.