REVIEW 33 cited by
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
read the original abstract
This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0. We train models with up to 2B parameters on nearly half a million hours of publicly available speech audio in 128 languages, an order of magnitude more public data than the largest known prior work. Our evaluation covers a wide range of tasks, domains, data regimes and languages, both high and low-resource. On the CoVoST-2 speech translation benchmark, we improve the previous state of the art by an average of 7.4 BLEU over 21 translation directions into English. For speech recognition, XLS-R improves over the best known prior work on BABEL, MLS, CommonVoice as well as VoxPopuli, lowering error rates by 14-34% relative on average. XLS-R also sets a new state of the art on VoxLingua107 language identification. Moreover, we show that with sufficient model size, cross-lingual pretraining can outperform English-only pretraining when translating English speech into other languages, a setting which favors monolingual pretraining. We hope XLS-R can help to improve speech processing tasks for many more languages of the world.
Forward citations
Cited by 33 Pith papers
-
MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio
MixFake is a new benchmark for mixed-authenticity audio and a multi-stream prompt tuning method achieves 0.95% EER foreground and 7.72% absolute gain in complex background deepfake detection.
-
Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection
MARS is a transfer-based black-box attack that uses bi-level optimization on semantic and artifact anchors to escape the linearity trap and improve attack success rates on SSL-SVDD by up to 36%.
-
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
PVP models speaker-specific phoneme acoustic distributions with lightweight GMMs trained only on real speech to detect deepfakes of persons-of-interest, outperforming generic detectors and introducing a new Chinese PO...
-
Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation
Multilingual ASR models show 39.7-297% zero-shot WER on Pashto public data, Whisper models output correct script in under 0.8% of cases, and fine-tuned models degrade to 32.5-59% WER on out-of-domain sets.
-
A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection
Spoof-SUPERB benchmark shows large-scale discriminative SSL models such as XLS-R, UniSpeech-SAT, and WavLM Large outperform others in audio deepfake detection and maintain robustness under acoustic degradations.
-
Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry
A four-backbone SSL ensemble achieves near-perfect deepfake detection in the ImageCLEF 2026 track, while the same team's generated audio ranks first in the generation track, revealing an OR-versus-AND asymmetry betwee...
-
Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
Pre-adaptation on related auxiliary languages yields no practically meaningful WER gains for large multilingual ASR once even 1 hour of target-language data is available.
-
An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures
Controlled non-speech interventions cause the largest detection-cost spikes, confirming non-speech structure as the dominant confound-driven shortcut in ASVspoof-trained XLS-R + RawGAT-ST models.
-
Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study
Across six pooling heads and six frozen SSL backbones on English and Mandarin depression speech, a third of configurations collapse to single-class prediction, so backbone- and seed-robustness should be first-class ev...
-
Syntactic Belief Update as the Driver of Garden Path Processing Difficulty
Syntactic belief update via generalized Rényi divergence on syntactic trees predicts garden path reading times better than lexical surprisal.
-
What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations
Linear probes on frozen self-supervised representations closely approach bespoke deepfake detector performance on benchmarks, indicating benchmarks largely measure general modality understanding.
-
Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing
A gated fusion of XLSR-53 and CORES features with energy margin and diversity losses reaches 97.6% ID accuracy and reduces FPR95 by 83.5% relative to the Interspeech 2025 baseline on MLAAD.
-
SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
SONAR improves audio deepfake detection by explicitly aligning low- and high-frequency representations for real speech and repelling them for fakes, setting new benchmark EERs on ASVspoof 2021 and in-the-wild data.
-
Forensic Similarity for Speech Deepfakes
Introduces forensic similarity for speech deepfakes via a Siamese feature extractor and similarity network to verify shared forensic traces and source models between audio segments.
-
Towards Digital Preservation of Efik: TTS for a Low-Resource African Language
First end-to-end Efik TTS baseline: a 3-hour single-speaker corpus and four fine-tuned models, with MMS-TTS best at MOS 3.80±0.63 but residual tonal errors.
-
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
Adding dataset identity as an auxiliary task or adversarial label improves aggregate audio-deepfake detection EER on the 2025 Speech DeepFake Arena benchmark.
-
Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR
Language-balanced gradient projection plus experience replay yields near-zero average forgetting when adapting Whisper-large-v3 to low-resource languages while preserving target plasticity.
-
GigaAM Multilingual: Foundation Model for Underrepresented Languages
Cluster-balanced HuBERT-style pre-training on 2M hours plus domain-aware fine-tuning yields a compact encoder that outperforms larger open multilingual ASR models on Kazakh, Kyrgyz and Uzbek.
-
Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
Pre-adapting multilingual ASR models on linguistically related languages does not yield practically meaningful target-language gains once one hour of target data is used.
-
Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack
KEYAC dataset benchmarks speech models for keyboard acoustic side-channel attacks, with KAN fine-tuning setting new SOTA by addressing nonlinear feature interactions.
-
Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack
KEYAC dataset created; KAN fine-tuning achieves SOTA on acoustic side-channel keystroke recognition from speech representations under zero-shot and partial fine-tuning.
-
Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal
A framework using native-only trained discrete token surprisal and DTW alignment features improves pronunciation assessment PCC to 0.66 on SpeechOcean762, approaching supervised performance.
-
Teffic-Audio: Tell Fact from Fiction
A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.
-
DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages
Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.
-
Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin
Truncated SSL backbone with logistic classifier detects audio deepfakes on-device, claimed to outperform AASIST by 10% while running 40% faster, packaged as a browser plugin.
-
From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa
Fine-tuned MMS model reaches 9.48% WER on Fongbe benchmark while Whisper on Hausa videos yields 6,770 segments rated 57.4/100 quality, with Fongbe lower at 36.5/100.
-
Beyond Speaker Independence: Evaluating Cross-Lingual Acoustic-to-Articulatory Inversion Across Finnish and Russian
Benchmarks on the new FROST-EMA corpus show cross-language mismatch drops Pearson correlation by 0.10-0.20 while cross-gender mismatch drops it by 0.05-0.10.
-
Pretrained self-supervised speech models can recognize unseen consonants
Fine-tuned Wav2Vec2 and HuBERT models recognize click consonants more accurately than non-clicks in G|ui and West !Xoon data.
-
Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection
Cosine similarity in SupCon with a delayed negative queue on wav2vec2 XLS-R yields the lowest equal error rates for deepfake audio detection on in-the-wild and pooled evaluations.
-
Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus
A bilingual TTS system for the Peruvian Constitution in Quechua and Spanish is developed with XTTS v2, F5-TTS, and DiFlow-TTS, releasing checkpoints and audio to support low-resource speech synthesis.
-
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning
A survey that organizes audio SSL into five objective paradigms, relates their demands to architectural biases, and interprets downstream applications as tests of generalization.
-
Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings
Evaluation of open-source and commercial ASR models on narrow-band Hindi and Indian English shows poor zero-shot results and inconsistent fine-tuning benefits tied to pretraining exposure.
-
SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing
SpAArSIST sparsifies AASIST by swapping learned pooling for explicit magnitude-based scoring and mean aggregation, cutting compute 20.7% and improving In-the-Wild EER to 2.82%.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.