Pith. sign in

REVIEW 14 cited by

Unsupervised Cross-lingual Representation Learning for Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.13979 v2 pith:KJKTX77O submitted 2020-06-24 cs.CL cs.LGcs.SDeess.AS

classification cs.CLcs.LGcs.SDeess.AS
keywords speechlanguagesmodelcross-lingualpretrainingrepresentationsacrossapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages. We build on wav2vec 2.0 which is trained by solving a contrastive task over masked latent speech representations and jointly learns a quantization of the latents shared across languages. The resulting model is fine-tuned on labeled data and experiments show that cross-lingual pretraining significantly outperforms monolingual pretraining. On the CommonVoice benchmark, XLSR shows a relative phoneme error rate reduction of 72% compared to the best known results. On BABEL, our approach improves word error rate by 16% relative compared to a comparable system. Our approach enables a single multilingual speech recognition model which is competitive to strong individual models. Analysis shows that the latent discrete speech representations are shared across languages with increased sharing for related languages. We hope to catalyze research in low-resource speech understanding by releasing XLSR-53, a large model pretrained in 53 languages.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs

    cs.AR 2026-03 unverdicted novelty 7.0 of 10

    ENEC delivers 3.43X higher throughput than DietGPU and 1.12X better compression ratio than nvCOMP for lossless model weight compression on Ascend NPUs, yielding up to 6.3X end-to-end inference speedup.

  2. Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.

  3. Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Assamese and Hindi, selected by acoustic and typological similarity to Warlpiri, cut Whisper WER/CER most; acoustic similarity best predicts fine-tuning gains, inventory/typology zero-shot.

  4. Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy

    cs.AI 2026-05 conditional novelty 6.0 of 10

    VST combines deep-learning stuttering classification with multi-agent LLM reasoning to produce clinician-reviewed, evidence-based therapy plans for stuttering.

  5. FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec

    cs.SD 2025-10 unverdicted novelty 6.0 of 10

    FAC-FACodec is a controllable zero-shot foreign accent conversion framework using a factorized speech codec that adds an explicit parameter for adjusting pronunciation-level accent modification strength.

  6. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  7. Prominence-aware automatic speech recognition for conversational speech

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A wav2vec2 speech recognizer was extended with word-level prominence labels, yielding simultaneous transcription and prominence annotation of conversational Austrian German with accuracy near the detector-only system.

  8. Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models

    eess.AS 2025-08 conditional novelty 6.0 of 10

    A hybrid decoding scheme with a fast TDT draft decoder and selective transformer patches matches baseline word error rate while cutting decoder latency roughly threefold.

  9. Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Proposes Bayesian factorized adaptation for multilingual ASR to handle code-switching, reporting 32.87% fewer errors on switched words and 5.31% better overall WER while preserving monolingual accuracy with small synt...

  10. Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems

    cs.SD 2025-09 reject novelty 5.0 of 10

    A new evaluation protocol exhaustively pairs 164 speech synthesizers with nine bona fide speech types and reports max-pooled EERs, revealing larger failures than pooled averages show.

  11. Pretrained self-supervised speech models can recognize unseen consonants

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    Fine-tuned Wav2Vec2 and HuBERT models recognize click consonants more accurately than non-clicks in G|ui and West !Xoon data.

  12. Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    Proposes Weighted-Attention and Self-Conditioning to reduce decoder inconsistencies and WER in Whisper for Dravidian and low-resource languages.

  13. TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation

    cs.CL 2025-08 conditional novelty 4.0 of 10

    Adding task-specific learned vectors to acoustic embeddings lets a transducer ASR model train on partially labeled data, matching or beating the fully labeled TokenVerse baseline on most tasks.

  14. A study on the impact of region specific data on the performance of Indic ASR

    eess.AS 2026-06 unverdicted novelty 3.0 of 10

    Empirical study finds consistent positive correlation between inter-district geographic distance and ASR word error rate when models are finetuned on single-district Indic speech data.

Pith tools