REVIEW 12 cited by
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce VoxPopuli, a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages. It is the largest open data to date for unsupervised representation learning as well as semi-supervised learning. VoxPopuli also contains 1.8K hours of transcribed speeches in 16 languages and their aligned oral interpretations into 5 other languages totaling 5.1K hours. We provide speech recognition baselines and validate the versatility of VoxPopuli unlabelled data in semi-supervised learning under challenging out-of-domain settings. We will release the corpus at https://github.com/facebookresearch/voxpopuli under an open license.
Forward citations
Cited by 12 Pith papers
-
SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.
-
ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
ParsVoice is an open ~1,800–2,200-hour Persian audiobook-derived speech-text corpus, much larger than prior open Persian TTS datasets.
-
From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model
Intermediate wav2vec 2.0 layers provide the best self-supervised speech representations for learning articulatory gestures and imitating speech across speakers.
-
On Barriers to Archival Audio Processing
On archival radio audio, Whisper V3 identifies languages with 91.3% accuracy, but speaker embeddings lose similarity across ages and languages, making speaker recognition unreliable for indexing.
-
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
A benchmark of eight post-training quantization methods on Whisper and Moonshine edge speech models across seven datasets, finding 8-bit is safe and 3-bit weights are viable for larger models with advanced methods like SpQR.
-
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.
-
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
A hierarchical Q-Former compresses speech to about 1.67 tokens/sec, enabling hour-long audio processing with near-linear memory scaling and competitive benchmark scores.
-
Group Relative Policy Optimization for Speech Recognition
Applying GRPO with rule-based rewards to LLM-based ASR improves WER by up to 18.4% relative and reduces hallucination errors on unseen acoustic conditions.
-
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings
A parameter-efficient adapter bridging Whisper and TinyLlama reports relative improvements in speech recognition, named entity recognition, and sentiment analysis on low-resource benchmarks.
-
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
The paper introduces AER, an LLM-judged question-answering metric for evaluating ASR output in LLM applications, and shows it does not correlate strongly with WER.
-
Unified Semi-Supervised Pipeline for Automatic Speech Recognition
A new semi-supervised ASR framework, TopIPL, combines a dynamic pseudo-label cache and top-N checkpoint teacher averaging to improve WER by up to 40 percent in low-resource settings.
-
Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning
The IT-IST IWSLT 2025 submission shows that a 1.5B language model with a speech encoder can do reasonable ASR after alignment, but struggles with ST and SQA.
Discussion (0). Sign in to comment.