Pith. sign in

REVIEW 17 cited by

VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.00390 v2 pith:SUXYLW2A submitted 2021-01-02 cs.CL eess.AS

classification cs.CLeess.AS
keywords learningvoxpopulicorpusdatahourslanguagessemi-supervisedspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce VoxPopuli, a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages. It is the largest open data to date for unsupervised representation learning as well as semi-supervised learning. VoxPopuli also contains 1.8K hours of transcribed speeches in 16 languages and their aligned oral interpretations into 5 other languages totaling 5.1K hours. We provide speech recognition baselines and validate the versatility of VoxPopuli unlabelled data in semi-supervised learning under challenging out-of-domain settings. We will release the corpus at https://github.com/facebookresearch/voxpopuli under an open license.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.

  2. ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis

    cs.SD 2025-10 conditional novelty 6.0 of 10

    ParsVoice is an open ~1,800–2,200-hour Persian audiobook-derived speech-text corpus, much larger than prior open Persian TTS datasets.

  3. From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model

    eess.AS 2025-09 conditional novelty 6.0 of 10

    Intermediate wav2vec 2.0 layers provide the best self-supervised speech representations for learning articulatory gestures and imitating speech across speakers.

  4. On Barriers to Archival Audio Processing

    cs.SD 2025-07 conditional novelty 6.0 of 10

    On archival radio audio, Whisper V3 identifies languages with 91.3% accuracy, but speaker embeddings lose similarity across ages and languages, making speaker recognition unreliable for indexing.

  5. Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A benchmark of eight post-training quantization methods on Whisper and Moonshine edge speech models across seven datasets, finding 8-bit is safe and 3-bit weights are viable for larger models with advanced methods like SpQR.

  6. MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MFLA adds finite look-ahead attention plus a CIF-based token counter to Whisper, enabling streaming recognition with a wait-k latency-quality trade-off.

  7. HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A 30-second counting task, embedded with speaker-identification models, predicts male sleep apnea (AUC 0.64) and shows gender- and condition-specific model rankings across a new 7,188-recording clinical speech benchmark.

  8. A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition

    eess.AS 2026-03 conditional novelty 5.5 of 10

    On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.

  9. FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation

    eess.AS 2026-01 conditional novelty 5.0 of 10

    A hierarchical Q-Former compresses speech to about 1.67 tokens/sec, enabling hour-long audio processing with near-linear memory scaling and competitive benchmark scores.

  10. Group Relative Policy Optimization for Speech Recognition

    eess.AS 2025-09 conditional novelty 5.0 of 10

    Applying GRPO with rule-based rewards to LLM-based ASR improves WER by up to 18.4% relative and reduces hallucination errors on unseen acoustic conditions.

  11. SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A parameter-efficient adapter bridging Whisper and TinyLlama reports relative improvements in speech recognition, named entity recognition, and sentiment analysis on low-resource benchmarks.

  12. An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications

    eess.AS 2025-07 reject novelty 5.0 of 10

    The paper introduces AER, an LLM-judged question-answering metric for evaluating ASR output in LLM applications, and shows it does not correlate strongly with WER.

  13. Unified Semi-Supervised Pipeline for Automatic Speech Recognition

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A new semi-supervised ASR framework, TopIPL, combines a dynamic pseudo-label cache and top-N checkpoint teacher averaging to improve WER by up to 40 percent in low-resource settings.

  14. DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A single speech encoder trained via ASR-aware distillation with variable attention masking performs competitively in both streaming and full-context modes at 200M and 2B scale.

  15. From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.

  16. Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning

    cs.CL 2025-06 conditional novelty 4.0 of 10

    The IT-IST IWSLT 2025 submission shows that a 1.5B language model with a speech encoder can do reasonable ASR after alignment, but struggles with ST and SQA.

  17. EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

    cs.SD 2025-05 reject novelty 4.0 of 10

    EZ-VC combines discrete units from a multilingual self-supervised encoder (Xeus) with an F5-TTS flow-matching decoder to achieve zero-shot any-to-any voice conversion, without text labels or multiple disentangling encoders.

Pith tools