REVIEW 14 cited by
Unsupervised Cross-lingual Representation Learning for Speech Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages. We build on wav2vec 2.0 which is trained by solving a contrastive task over masked latent speech representations and jointly learns a quantization of the latents shared across languages. The resulting model is fine-tuned on labeled data and experiments show that cross-lingual pretraining significantly outperforms monolingual pretraining. On the CommonVoice benchmark, XLSR shows a relative phoneme error rate reduction of 72% compared to the best known results. On BABEL, our approach improves word error rate by 16% relative compared to a comparable system. Our approach enables a single multilingual speech recognition model which is competitive to strong individual models. Analysis shows that the latent discrete speech representations are shared across languages with increased sharing for related languages. We hope to catalyze research in low-resource speech understanding by releasing XLSR-53, a large model pretrained in 53 languages.
Forward citations
Cited by 14 Pith papers
-
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
ENEC delivers 3.43X higher throughput than DietGPU and 1.12X better compression ratio than nvCOMP for lossless model weight compression on Ascend NPUs, yielding up to 6.3X end-to-end inference speedup.
-
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.
-
Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR
Assamese and Hindi, selected by acoustic and typological similarity to Warlpiri, cut Whisper WER/CER most; acoustic similarity best predicts fine-tuning gains, inventory/typology zero-shot.
-
Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy
VST combines deep-learning stuttering classification with multi-agent LLM reasoning to produce clinician-reviewed, evidence-based therapy plans for stuttering.
-
FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec
FAC-FACodec is a controllable zero-shot foreign accent conversion framework using a factorized speech codec that adds an explicit parameter for adjusting pronunciation-level accent modification strength.
-
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.
-
Prominence-aware automatic speech recognition for conversational speech
A wav2vec2 speech recognizer was extended with word-level prominence labels, yielding simultaneous transcription and prominence annotation of conversational Austrian German with accuracy near the detector-only system.
-
Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models
A hybrid decoding scheme with a fast TDT draft decoder and selective transformer patches matches baseline word error rate while cutting decoder latency roughly threefold.
-
Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR
Proposes Bayesian factorized adaptation for multilingual ASR to handle code-switching, reporting 32.87% fewer errors on switched words and 5.31% better overall WER while preserving monolingual accuracy with small synt...
-
Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems
A new evaluation protocol exhaustively pairs 164 speech synthesizers with nine bona fide speech types and reports max-pooled EERs, revealing larger failures than pooled averages show.
-
Pretrained self-supervised speech models can recognize unseen consonants
Fine-tuned Wav2Vec2 and HuBERT models recognize click consonants more accurately than non-clicks in G|ui and West !Xoon data.
-
Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages
Proposes Weighted-Attention and Self-Conditioning to reduce decoder inconsistencies and WER in Whisper for Dravidian and low-resource languages.
-
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
Adding task-specific learned vectors to acoustic embeddings lets a transducer ASR model train on partially labeled data, matching or beating the fully labeled TokenVerse baseline on most tasks.
-
A study on the impact of region specific data on the performance of Indic ASR
Empirical study finds consistent positive correlation between inter-district geographic distance and ASR word error rate when models are finetuned on single-district Indic speech data.
Discussion (0). Sign in to comment.