Pith. sign in

REVIEW 23 cited by

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.01037 v3 pith:GYDIN6CX submitted 2023-03-02 cs.CL cs.SDeess.AS

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

classification cs.CL cs.SDeess.AS
keywords modellanguagesspeechmultilingualrecognitionacrossautomaticdataset
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages. This is achieved by pre-training the encoder of the model on a large unlabeled multilingual dataset of 12 million (M) hours spanning over 300 languages, and fine-tuning on a smaller labeled dataset. We use multilingual pre-training with random-projection quantization and speech-text modality matching to achieve state-of-the-art performance on downstream multilingual ASR and speech-to-text translation tasks. We also demonstrate that despite using a labeled training set 1/7-th the size of that used for the Whisper model, our model exhibits comparable or better performance on both in-domain and out-of-domain speech recognition tasks across many languages.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions

    cs.CL 2026-05 conditional novelty 7.0

    FalAR is a new 5,800-hour EP parliamentary speech corpus with speaker annotations on 4,850 hours that yields up to 14% relative WER improvement in ASR pre-training experiments.

  2. Moshi: a speech-text foundation model for real-time dialogue

    eess.AS 2024-09 accept novelty 7.0

    Moshi is the first real-time full-duplex spoken large language model that casts dialogue as speech-to-speech generation using parallel audio streams and an inner monologue of time-aligned text tokens.

  3. Gemma 4 Technical Report

    cs.CL 2026-07 accept novelty 6.0

    Gemma 4 open multimodal models (dense + MoE) with thinking mode, encoder-free 12B path, and KV/memory optimizations leap prior Gemma and rival larger open models on STEM, multimodal, long-context, and Arena benchmarks.

  4. MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production

    cs.DC 2026-05 unverdicted novelty 6.0

    MegaScale-Omni delivers 1.27x-7.57x higher throughput for dynamic multimodal LLM training by decoupling encoder and LLM parallelism, using unified colocation, and applying adaptive workload balancing.

  5. BlasBench: An Open Benchmark for Irish Speech Recognition

    cs.CL 2026-04 conditional novelty 6.0

    BlasBench supplies an Irish-aware normalizer and scoring harness that enables reproducible ASR comparisons and exposes a 33-43 point generalization gap for fine-tuned models versus 7-10 points for massively multilingual ones.

  6. PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer

    cs.CV 2026-04 unverdicted novelty 6.0

    PoM is a new linear-complexity token mixer using learned polynomials that matches attention performance in transformers while enabling efficient long-sequence processing.

  7. BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation

    cs.CL 2025-10 unverdicted novelty 6.0

    BEARD adapts Whisper encoder for ATC domain via BEST-RQ and distillation on 5000h unlabeled speech then 2h labeled fine-tuning, delivering 12% relative WER gain over fine-tuned baseline.

  8. Gemini: A Family of Highly Capable Multimodal Models

    cs.CL 2023-12 conditional novelty 6.0

    Gemini Ultra reaches human-expert performance on MMLU for the first time and sets new state-of-the-art results on 30 of 32 benchmarks, including all 20 multimodal ones tested.

  9. AudioPaLM: A Large Language Model That Can Speak and Listen

    cs.CL 2023-06 unverdicted novelty 6.0

    AudioPaLM unifies PaLM-2 and AudioLM to outperform prior systems on speech translation while enabling zero-shot speech-to-text for many unseen language pairs and voice transfer from short prompts.

  10. A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition

    eess.AS 2026-03 conditional novelty 5.5

    On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.

  11. GigaAM Multilingual: Foundation Model for Underrepresented Languages

    eess.AS 2026-07 conditional novelty 5.0

    Cluster-balanced HuBERT-style pre-training on 2M hours plus domain-aware fine-tuning yields a compact encoder that outperforms larger open multilingual ASR models on Kazakh, Kyrgyz and Uzbek.

  12. Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization

    eess.AS 2026-05 unverdicted novelty 5.0

    Task-dependent simulation strategies for synthetic conversational data allow synthetic-only training to approach real-data baselines for multi-talker ASR and diarization, with mixing yielding further gains.

  13. UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations

    eess.AS 2026-04 unverdicted novelty 5.0

    UniPASE extends the PASE framework with DeWavLM-Omni to convert degraded speech into high-fidelity, low-hallucination audio across sampling rates via phonetic enhancement, acoustic adaptation, and multi-rate vocoding.

  14. A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models

    cs.SD 2026-01 unverdicted novelty 5.0

    Prioritizing longest utterances in SSL speech pre-training data outperforms random or diversity-based sampling for ASR performance while using half the data volume.

  15. Kimi-Audio Technical Report

    eess.AS 2025-04 unverdicted novelty 5.0

    Kimi-Audio is an open-source audio foundation model that achieves state-of-the-art results on speech recognition, audio understanding, question answering, and conversation after pre-training on more than 13 million ho...

  16. Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study

    cs.CL 2026-06 unverdicted novelty 4.0

    Case study finds that fine-tuned ASR models outperform human listeners on Dutch dysarthric continuous speech from one speaker, lowering WER from over 70% to over 23%.

  17. Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

    cs.SD 2026-06 unverdicted novelty 4.0

    Three modifications to BEST-RQ quantization (PCA projection, iterative codebook refinement, codebook distillation) reduce WER from 10.1% to 8.8% on LibriSpeech test-other.

  18. Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR

    cs.AI 2026-06 conditional novelty 4.0

    Multilingual encoder warm-start gives a data-scale-limited WER advantage in cross-lingual streaming ASR that decays with target data volume and is independent of latency tier.

  19. ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

    eess.AS 2026-06 unverdicted novelty 4.0

    ESPnet3 introduces a new modular architecture with DataOrganizer and sharding to cut training time and simplify model integration for speech research.

  20. Dolphin-CN-Dialect: Where Chinese Dialects Matter

    cs.CL 2026-05 unverdicted novelty 4.0

    Dolphin-CN-Dialect is a compact ASR model that boosts Chinese dialect accuracy through balanced sampling of rare dialects and character-level tokenization while staying smaller than recent open-source competitors.

  21. Online Predictive Coding for Dual-Mode Self-Supervised Speech Model

    cs.SD 2026-06 unverdicted novelty 3.0

    Proposes OPC and dual-mode LN to improve dual-mode SSL speech models, reducing WER gap at 160 ms latency on LibriSpeech from 3.65% to 3.40% (test-clean).

  22. On The Landscape of Spoken Language Models: A Comprehensive Survey

    cs.CL 2025-04 unverdicted novelty 3.0

    A literature survey that organizes spoken language models by architecture, training, and evaluation choices and identifies key challenges and future directions.

  23. Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)

    cs.CL 2025-01 unverdicted novelty 2.0

    A literature survey of Small Language Models (1-8B parameters) that can perform comparably or better than larger models, covering general-purpose and task-specific approaches plus creation techniques.