Pith. sign in

REVIEW 56 cited by

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.01037 v3 pith:GYDIN6CX submitted 2023-03-02 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords modellanguagesspeechmultilingualrecognitionacrossautomaticdataset
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages. This is achieved by pre-training the encoder of the model on a large unlabeled multilingual dataset of 12 million (M) hours spanning over 300 languages, and fine-tuning on a smaller labeled dataset. We use multilingual pre-training with random-projection quantization and speech-text modality matching to achieve state-of-the-art performance on downstream multilingual ASR and speech-to-text translation tasks. We also demonstrate that despite using a labeled training set 1/7-th the size of that used for the Whisper model, our model exhibits comparable or better performance on both in-domain and out-of-domain speech recognition tasks across many languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 56 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 112 citations worldwide. Full citation record

  1. OLMoASR: Open Models and Data for Training Robust Speech Recognition Models

    cs.SD 2025-08 conditional novelty 7.0 of 10

    An open 1M-hour English speech dataset plus Whisper-architecture models trained on it match Whisper's word error rates on short and long-form benchmarks.

  2. SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR

    eess.AS 2024-12 conditional novelty 7.0 of 10

    SQ-Whisper injects trainable speaker queries into Whisper's encoder and decoder, beating prior target-speaker ASR systems with new state-of-the-art WERs on Libri2Mix and WSJ0-2Mix.

  3. SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SraVaani-1.0 reports the broadest open ASR coverage for Indic languages to date, with competitive word error rates on 17 benchmark languages and the only transcription output for 44 low-resource languages, evaluated o...

  4. Gemma 4 Technical Report

    cs.CL 2026-07 accept novelty 6.0 of 10

    Gemma 4 open multimodal models (dense + MoE) with thinking mode, encoder-free 12B path, and KV/memory optimizations leap prior Gemma and rival larger open models on STEM, multimodal, long-context, and Arena benchmarks.

  5. UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations

    eess.AS 2026-04 unverdicted novelty 6.0 of 10

    UniPASE extends the PASE framework with DeWavLM-Omni to convert degraded speech into high-fidelity, low-hallucination audio across sampling rates via phonetic enhancement, acoustic adaptation, and multi-rate vocoding.

  6. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  7. The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties

    cs.CL 2025-09 conditional novelty 6.0 of 10

    The ML-SUPERB 2.0 Challenge adds accented and dialectal speech evaluation to multilingual ASR, and all five submitted systems beat the baselines.

  8. CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese

    cs.CL 2025-08 conditional novelty 6.0 of 10

    CAMOES is a new open benchmark and model collection for European Portuguese ASR, cutting word error rate by about 35% over the best zero-shot model.

  9. Geolocation-Aware Robust Spoken Language Identification

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Auxiliary geolocation prediction with injected conditioning signals improves dialect and accent robustness in SSL-based spoken language identification, reaching new SOTA on FLEURS and ML-SUPERB 2.0.

  10. LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness

    eess.AS 2025-08 conditional novelty 6.0 of 10

    LCS-CTC, a phoneme recognizer trained with similarity-aware LCS alignment masks constraining CTC, outperforms vanilla CTC on all reported PER, WPER, boundary-loss, and articulatory metrics.

  11. Identifying Hearing Difficulty Moments in Conversational Audio

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Prompted Gemini audio models detect hearing difficulty moments in conversation audio with F1 0.87, beating ASR hotword and Wav2Vec baselines.

  12. OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder

    cs.SD 2025-07 conditional novelty 6.0 of 10

    OpenBEATs releases the BEATs audio pretraining pipeline, trains 300M-parameter models on 20k hours of multi-domain audio, and reports strong results across 25 datasets, including bioacoustics and reasoning tasks.

  13. NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A real-world continual learning benchmark for multilingual ASR built from 3,250 hours of Indian language speech shows that no current CL method performs consistently across language- and domain-incremental scenarios.

  14. A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition

    cs.SD 2025-06 conditional novelty 6.0 of 10

    An iterative segmentation-based self-training method for Whisper improved long dysarthric speech recognition and achieved second place in both WER and SemScore at the SAP Challenge.

  15. Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A quality audit of Common Voice, FLEURS, and VoxPopuli finds serious micro- and macro-level data defects concentrated in less institutionalized languages, with a case study of Taiwanese Southern Min.

  16. Early Attentive Sparsification Accelerates Neural Speech Transcription

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Attention-based early audio-token sparsification at 40-60% sparsity accelerates Whisper ASR up to 1.6x with under 1% relative WER loss, across ten model variants, with no fine-tuning.

  17. OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Freezing OWSM v3.1 and adding dynamic-vocabulary biasing modules improves rare-word recognition and reduces real-time factor on LibriSpeech 100.

  18. GigaAM: Efficient Self-Supervised Learner for Speech Recognition

    eess.AS 2025-06 conditional novelty 6.0 of 10

    GigaAM, a Russian ASR model family pretrained with CTC-teacher cluster targets, reports roughly 50% lower WER than Whisper-large-v3 on three Russian benchmarks.

  19. NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.

  20. The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence

    cs.CL 2025-05 conditional novelty 6.0 of 10

    In 150k-hour speech-to-text training, a sub-exponential learning-rate warmup prevents divergence, while a faster warmup only speeds early convergence and does not improve the final model.

  21. LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Using CTC posterior scores from a speech encoder as soft weights over a language model's token embeddings lets a frozen encoder and a fine-tuned LLM work together, with a 49% average word error reduction on MLS.

  22. Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Miipher-2 restores degraded speech in known and unknown languages without conditioning, using a frozen 300-language USM encoder, parallel adapters, and a memory-efficient WaveFit vocoder.

  23. OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A new open suite of 13 multilingual speech models, up to 18B parameters, yields empirical scaling laws for ASR and speech translation performance.

  24. Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning

    eess.AS 2024-12 conditional novelty 6.0 of 10

    Reinforcement learning with word-error and meaning-preservation rewards adapts an LLM-based speech recognizer to disordered speech better than supervised fine-tuning in this study.

  25. How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A survey of 110 SimulST papers shows most systems rely on unrealistic human pre-segmented audio and inconsistent terminology, and it offers a taxonomy and recommendations to fix both.

  26. Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous Speech

    cs.CL 2024-12 conditional novelty 6.0 of 10

    GPT-4 ratings of five dementia-related language symptoms, added to 40 standard linguistic features, improve automatic Alzheimer's detection from spontaneous speech transcripts, reaching AUROC 0.931 on ADReSS.

  27. TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

    eess.AS 2024-12 conditional novelty 6.0 of 10

    TouchASP trains a single elastic mixture-of-experts ASR model on 1M hours of partly pseudo-labeled audio and reports SpeechIO CER dropping from 4.98% to 2.45% while adding multi-task perception.

  28. Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition

    eess.AS 2024-12 conditional novelty 6.0 of 10

    Transliterated zero-shot domain adaptation reduces ASR word error rate by 9.2% relative to wav2vec 2.0 by pre-training on transliterated pseudo-labels from a related source language.

  29. Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition

    eess.AS 2024-11 reject novelty 6.0 of 10

    A learned future-audio density ratio, FoCCE, is inserted into the streaming transducer forward recursion during training and modestly reduces word error rates.

  30. k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning

    eess.AS 2024-11 conditional novelty 6.0 of 10

    Zipformer-based self-supervised pretraining is up to 3.5x faster than HuBERT and yields lower ASR word error rates on LibriSpeech.

  31. A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition

    eess.AS 2026-03 conditional novelty 5.5 of 10

    On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.

  32. On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin

    eess.AS 2026-08 conditional novelty 5.0 of 10

    Staged continual pre-training, dialect fine-tuning, and on-policy self-distillation improves Chinese multi-dialect ASR while preserving Mandarin CER, outperforming continued teacher-forced fine-tuning.

  33. MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Grouping 495 languages into roughly 16 clusters and routing speech to group-specific LoRA experts improves multilingual ASR error rates over dense and random baselines.

  34. GigaAM Multilingual: Foundation Model for Underrepresented Languages

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Cluster-balanced HuBERT-style pre-training on 2M hours plus domain-aware fine-tuning yields a compact encoder that outperforms larger open multilingual ASR models on Kazakh, Kyrgyz and Uzbek.

  35. Towards Improved Speech Recognition through Optimized Synthetic Data Generation

    eess.AS 2025-08 conditional novelty 5.0 of 10

    A fine-tuned TTS with Whisper-based filtering generates synthetic Quebec French speech that trains ASR models well when combined with 10-60 hours of real audio, though a 13-14% WER gap to real-data training remains.

  36. Efficient Multilingual ASR Finetuning via LoRA Language Experts

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Monolingual LoRA language experts, combined by weighted merging (MoLE) or layer-wise knowledge distillation, improve Whisper-based multilingual ASR by about 10-15% relative WER over a plain multilingual LoRA baseline.

  37. A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Fine-tuning Whisper-large-v2 on 10,000 hours of synthesized Mandarin plus small real English/code-switching sets yields Twister, cutting mixed error rate by up to 56% on code-switching and 19% on Taiwanese Mandarin.

  38. OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    OWSM v4 models, trained on a cleaned 166k-hour multilingual YODAS subset, beat prior open OWSM models and are competitive with Whisper and MMS on several benchmarks.

  39. Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A Data2Vec2 speech encoder pre-trained on 300,000 hours of unlabeled Chinese dialect speech, connected to a small Qwen LLM via a linear projector and fine-tuned in four stages, sets a new state of the art on Chinese d...

  40. DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A single speech encoder trained via ASR-aware distillation with variable attention masking performs competitively in both streaming and full-context modes at 200M and 2B scale.

  41. VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A 68M-parameter Vietnamese ASR model, pretrained on 70,000 hours of unlabeled audio and fine-tuned on 50 hours of labels, reports average WER 8.31, beating Whisper Large-v3 and commercial systems.

  42. Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A spoken large language model learns to judge its own transcription difficulty and routes only hard speech to a stronger ASR model, cutting cost and improving word error rate.

  43. Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Granite-speech-3.3-2b and Granite-speech-3.3-8b achieve competitive English ASR word error rates, with the 8B model beating several larger proprietary models on multiple public benchmarks while remaining fully open-source.

  44. Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.

  45. Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A target-phrase-only edit representation offers the best accuracy-versus-output-length trade-off for LLM-based ASR post editing, closing 50-60% of the WER gap to full rewriting while losing only 10-20% of the length savings.

  46. BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR

    cs.CL 2025-01 conditional novelty 5.0 of 10

    BLR-MoE, which adds language-specific attention experts, expert pruning, and router fine-tuning to the LR-MoE architecture, reduces WER by 16.09% relative on a 10,000-hour multilingual ASR benchmark.

  47. AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A bias-list-conditioned tagger and decoder improves Vietnamese code-switching ASR text normalization, cutting word error rate by up to 56.2% and 36.8% on the authors' new synthetic test sets.

  48. Towards Early Prediction of Self-Supervised Speech Model Performance

    cs.SD 2025-01 conditional novelty 5.0 of 10

    Rank and clustering metrics on SSL speech embeddings predict downstream ASR and SV performance better than the pre-training loss, especially for in-domain ASR.

  49. UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition

    eess.AS 2024-12 conditional novelty 5.0 of 10

    UME upcycles pretrained dense ASR checkpoints into larger MoE models via weight copying, layer freezing, and expert balancing, yielding up to 11.9% relative CER reduction and 86.7% training time savings versus trainin...

  50. Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding

    cs.CL 2024-12 conditional novelty 5.0 of 10

    An LSTM-based encoder refiner plus language-aware dual adapters with a fusion module cuts Mandarin-English code-switching ASR errors on SEAME.

  51. Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A bilevel training method that jointly optimizes supervised and unsupervised losses outperforms pretraining-then-finetuning for ASR on LibriSpeech, Switchboard, and an in-house dataset.

  52. DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

    cs.CL 2026-07 reject novelty 4.0 of 10

    Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.

  53. Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach

    eess.AS 2025-05 conditional novelty 4.0 of 10

    Llama-SMoP-DEDR, a sparse mixture of projectors with modality-specific experts and routers, lowers word error rate for LLM-based AVSR on LRS3, mainly with smaller LLMs.

  54. Optimized Self-supervised Training with BEST-RQ for Speech Recognition

    cs.SD 2025-01 conditional novelty 4.0 of 10

    Adding multiple codebooks, a KL-divergence regularizer, and cluster-specific codebooks to BEST-RQ improves LibriSpeech ASR word error rates by up to 30.6% relative with faster convergence.

  55. Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Language-family prompt tuning and a BPE-token-extended tokenizer improve Whisper's WER and inference speed on eight Indian languages, with 250 added tokens per language as the best configuration.

  56. MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

    cs.CL 2024-12 conditional novelty 4.0 of 10

    An openly released 630M-parameter BEST-RQ speech encoder, pretrained on 200k hours, matches state-of-the-art encoders on several ASR benchmarks and holds its own on ten SUPERB tasks.

Pith tools