REVIEW 56 cited by
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages. This is achieved by pre-training the encoder of the model on a large unlabeled multilingual dataset of 12 million (M) hours spanning over 300 languages, and fine-tuning on a smaller labeled dataset. We use multilingual pre-training with random-projection quantization and speech-text modality matching to achieve state-of-the-art performance on downstream multilingual ASR and speech-to-text translation tasks. We also demonstrate that despite using a labeled training set 1/7-th the size of that used for the Whisper model, our model exhibits comparable or better performance on both in-domain and out-of-domain speech recognition tasks across many languages.
Forward citations
Cited by 56 Pith papers
-
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
An open 1M-hour English speech dataset plus Whisper-architecture models trained on it match Whisper's word error rates on short and long-form benchmarks.
-
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
SQ-Whisper injects trainable speaker queries into Whisper's encoder and decoder, beating prior target-speaker ASR systems with new state-of-the-art WERs on Libri2Mix and WSJ0-2Mix.
-
SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages
SraVaani-1.0 reports the broadest open ASR coverage for Indic languages to date, with competitive word error rates on 17 benchmark languages and the only transcription output for 44 low-resource languages, evaluated o...
-
Gemma 4 Technical Report
Gemma 4 open multimodal models (dense + MoE) with thinking mode, encoder-free 12B path, and KV/memory optimizations leap prior Gemma and rival larger open models on STEM, multimodal, long-context, and Arena benchmarks.
-
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
UniPASE extends the PASE framework with DeWavLM-Omni to convert degraded speech into high-fidelity, low-hallucination audio across sampling rates via phonetic enhancement, acoustic adaptation, and multi-rate vocoding.
-
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.
-
The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
The ML-SUPERB 2.0 Challenge adds accented and dialectal speech evaluation to multilingual ASR, and all five submitted systems beat the baselines.
-
CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
CAMOES is a new open benchmark and model collection for European Portuguese ASR, cutting word error rate by about 35% over the best zero-shot model.
-
Geolocation-Aware Robust Spoken Language Identification
Auxiliary geolocation prediction with injected conditioning signals improves dialect and accent robustness in SSL-based spoken language identification, reaching new SOTA on FLEURS and ML-SUPERB 2.0.
-
LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
LCS-CTC, a phoneme recognizer trained with similarity-aware LCS alignment masks constraining CTC, outperforms vanilla CTC on all reported PER, WPER, boundary-loss, and articulatory metrics.
-
Identifying Hearing Difficulty Moments in Conversational Audio
Prompted Gemini audio models detect hearing difficulty moments in conversation audio with F1 0.87, beating ASR hotword and Wav2Vec baselines.
-
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
OpenBEATs releases the BEATs audio pretraining pipeline, trains 300M-parameter models on 20k hours of multi-domain audio, and reports strong results across 25 datasets, including bioacoustics and reasoning tasks.
-
NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
A real-world continual learning benchmark for multilingual ASR built from 3,250 hours of Indian language speech shows that no current CL method performs consistently across language- and domain-incremental scenarios.
-
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
An iterative segmentation-based self-training method for Whisper improved long dysarthric speech recognition and achieved second place in both WER and SemScore at the SAP Challenge.
-
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
A quality audit of Common Voice, FLEURS, and VoxPopuli finds serious micro- and macro-level data defects concentrated in less institutionalized languages, with a case study of Taiwanese Southern Min.
-
Early Attentive Sparsification Accelerates Neural Speech Transcription
Attention-based early audio-token sparsification at 40-60% sparsity accelerates Whisper ASR up to 1.6x with under 1% relative WER loss, across ten model variants, with no fine-tuning.
-
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
Freezing OWSM v3.1 and adding dynamic-vocabulary biasing modules improves rare-word recognition and reduces real-time factor on LibriSpeech 100.
-
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
GigaAM, a Russian ASR model family pretrained with CTC-teacher cluster targets, reports roughly 50% lower WER than Whisper-large-v3 on three Russian benchmarks.
-
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.
-
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
In 150k-hour speech-to-text training, a sub-exponential learning-rate warmup prevents divergence, while a faster warmup only speeds early convergence and does not improve the final model.
-
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
Using CTC posterior scores from a speech encoder as soft weights over a language model's token embeddings lets a frozen encoder and a fine-tuned LLM work together, with a 49% average word error reduction on MLS.
-
Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration
Miipher-2 restores degraded speech in known and unknown languages without conditioning, using a frozen 300-language USM encoder, parallel adapters, and a memory-efficient WaveFit vocoder.
-
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
A new open suite of 13 multilingual speech models, up to 18B parameters, yields empirical scaling laws for ASR and speech translation performance.
-
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
Reinforcement learning with word-error and meaning-preservation rewards adapts an LLM-based speech recognizer to disordered speech better than supervised fine-tuning in this study.
-
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
A survey of 110 SimulST papers shows most systems rely on unrealistic human pre-segmented audio and inconsistent terminology, and it offers a taxonomy and recommendations to fix both.
-
Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous Speech
GPT-4 ratings of five dementia-related language symptoms, added to 40 standard linguistic features, improve automatic Alzheimer's detection from spontaneous speech transcripts, reaching AUROC 0.931 on ADReSS.
-
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
TouchASP trains a single elastic mixture-of-experts ASR model on 1M hours of partly pseudo-labeled audio and reports SpeechIO CER dropping from 4.98% to 2.45% while adding multi-task perception.
-
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
Transliterated zero-shot domain adaptation reduces ASR word error rate by 9.2% relative to wav2vec 2.0 by pre-training on transliterated pseudo-labels from a related source language.
-
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
A learned future-audio density ratio, FoCCE, is inserted into the streaming transducer forward recursion during training and modestly reduces word error rates.
-
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
Zipformer-based self-supervised pretraining is up to 3.5x faster than HuBERT and yields lower ASR word error rates on LibriSpeech.
-
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.
-
On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin
Staged continual pre-training, dialect fine-tuning, and on-policy self-distillation improves Chinese multi-dialect ASR while preserving Mandarin CER, outperforming continued teacher-forced fine-tuning.
-
MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition
Grouping 495 languages into roughly 16 clusters and routing speech to group-specific LoRA experts improves multilingual ASR error rates over dense and random baselines.
-
GigaAM Multilingual: Foundation Model for Underrepresented Languages
Cluster-balanced HuBERT-style pre-training on 2M hours plus domain-aware fine-tuning yields a compact encoder that outperforms larger open multilingual ASR models on Kazakh, Kyrgyz and Uzbek.
-
Towards Improved Speech Recognition through Optimized Synthetic Data Generation
A fine-tuned TTS with Whisper-based filtering generates synthetic Quebec French speech that trains ASR models well when combined with 10-60 hours of real audio, though a 13-14% WER gap to real-data training remains.
-
Efficient Multilingual ASR Finetuning via LoRA Language Experts
Monolingual LoRA language experts, combined by weighted merging (MoLE) or layer-wise knowledge distillation, improve Whisper-based multilingual ASR by about 10-15% relative WER over a plain multilingual LoRA baseline.
-
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
Fine-tuning Whisper-large-v2 on 10,000 hours of synthesized Mandarin plus small real English/code-switching sets yields Twister, cutting mixed error rate by up to 56% on code-switching and 19% on Taiwanese Mandarin.
-
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
OWSM v4 models, trained on a cleaned 166k-hour multilingual YODAS subset, beat prior open OWSM models and are competitive with Whisper and MMS on several benchmarks.
-
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
A Data2Vec2 speech encoder pre-trained on 300,000 hours of unlabeled Chinese dialect speech, connected to a small Qwen LLM via a linear projector and fine-tuned in four stages, sets a new state of the art on Chinese d...
-
DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation
A single speech encoder trained via ASR-aware distillation with variable attention masking performs competitively in both streaming and full-context modes at 200M and 2B scale.
-
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
A 68M-parameter Vietnamese ASR model, pretrained on 70,000 hours of unlabeled audio and fine-tuned on 50 hours of labels, reports average WER 8.31, beating Whisper Large-v3 and commercial systems.
-
Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
A spoken large language model learns to judge its own transcription difficulty and routes only hard speech to a stronger ASR model, cutting cost and improving word error rate.
-
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
Granite-speech-3.3-2b and Granite-speech-3.3-8b achieve competitive English ASR word error rates, with the 8B model beating several larger proprietary models on multiple public benchmarks while remaining fully open-source.
-
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.
-
Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing
A target-phrase-only edit representation offers the best accuracy-versus-output-length trade-off for LLM-based ASR post editing, closing 50-60% of the WER gap to full rewriting while losing only 10-20% of the length savings.
-
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
BLR-MoE, which adds language-specific attention experts, expert pruning, and router fine-tuning to the LR-MoE architecture, reduces WER by 16.09% relative on a 10,000-hour multilingual ASR benchmark.
-
AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR
A bias-list-conditioned tagger and decoder improves Vietnamese code-switching ASR text normalization, cutting word error rate by up to 56.2% and 36.8% on the authors' new synthetic test sets.
-
Towards Early Prediction of Self-Supervised Speech Model Performance
Rank and clustering metrics on SSL speech embeddings predict downstream ASR and SV performance better than the pre-training loss, especially for in-domain ASR.
-
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition
UME upcycles pretrained dense ASR checkpoints into larger MoE models via weight copying, layer freezing, and expert balancing, yielding up to 11.9% relative CER reduction and 86.7% training time savings versus trainin...
-
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
An LSTM-based encoder refiner plus language-aware dual adapters with a fusion module cuts Mandarin-English code-switching ASR errors on SEAME.
-
Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition
A bilevel training method that jointly optimizes supervised and unsupervised losses outperforms pretraining-then-finetuning for ASR on LibriSpeech, Switchboard, and an in-house dataset.
-
DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages
Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.
-
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
Llama-SMoP-DEDR, a sparse mixture of projectors with modality-specific experts and routers, lowers word error rate for LLM-based AVSR on LRS3, mainly with smaller LLMs.
-
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
Adding multiple codebooks, a KL-divergence regularizer, and cluster-specific codebooks to BEST-RQ improves LibriSpeech ASR word error rates by up to 30.6% relative with faster convergence.
-
Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization
Language-family prompt tuning and a BPE-token-extended tokenizer improve Whisper's WER and inference speed on eight Indian languages, with 250 added tokens per language as the best configuration.
-
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
An openly released 630M-parameter BEST-RQ speech encoder, pretrained on 200k hours, matches state-of-the-art encoders on several ASR benchmarks and holds its own on ten SUPERB tasks.
Discussion (0). Continue with ORCID to comment.