Pith. sign in

REVIEW 73 cited by

Common Voice: A Massively-Multilingual Speech Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.06670 v2 pith:BXCKJI2R submitted 2019-12-13 cs.CL cs.LG

classification cs.CLcs.LG
keywords speechcommonlanguagesvoicerecognitioncorpusdataaudio
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identification). To achieve scale and sustainability, the Common Voice project employs crowdsourcing for both data collection and data validation. The most recent release includes 29 languages, and as of November 2019 there are a total of 38 languages collecting data. Over 50,000 individuals have participated so far, resulting in 2,500 hours of collected audio. To our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages. As an example use case for Common Voice, we present speech recognition experiments using Mozilla's DeepSpeech Speech-to-Text toolkit. By applying transfer learning from a source English model, we find an average Character Error Rate improvement of 5.99 +/- 5.48 for twelve target languages (German, French, Italian, Turkish, Catalan, Slovenian, Welsh, Irish, Breton, Tatar, Chuvash, and Kabyle). For most of these languages, these are the first ever published results on end-to-end Automatic Speech Recognition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 73 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 73 Pith citations

  1. REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

    cs.CL 2026-07 conditional novelty 7.0 of 10

    REDDIT corrects non-speech-induced timestamp drift in autoregressive ASR by editing timestamp targets under cached replay context while anchoring non-timestamp behavior to the frozen base distribution.

  2. Continual Learning with Embedding Layer Surgery and Task-wise Beam Search using Whisper

    cs.CL 2025-01 conditional novelty 7.0 of 10

    Embedding Layer Surgery and Task-wise Beam Search reduce catastrophic forgetting when adding new languages to Whisper, lowering old-language AWER from 14.2% to 11.9% versus Experience Replay.

  3. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.

  4. Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

    eess.AS 2025-11 conditional novelty 6.0 of 10

    A 10.7M-pair audio-caption corpus and systematic comparison show contrastive pretraining is more data-efficient while captioning scales better, and supervised initialization yields diminishing returns.

  5. ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis

    cs.SD 2025-10 conditional novelty 6.0 of 10

    ParsVoice is an open ~1,800–2,200-hour Persian audiobook-derived speech-text corpus, much larger than prior open Persian TTS datasets.

  6. AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A four-part benchmark plus a semantic/acoustic token taxonomy for comparing audio codecs, with correlation analysis across ten models.

  7. AHELM: A Holistic Evaluation of Audio-Language Models

    cs.AI 2025-08 conditional novelty 6.0 of 10

    AHELM standardizes evaluation of audio-language models across 10 aspects and shows simple ASR+LLM systems are competitive with multimodal models.

  8. LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.

  9. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  10. WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes

    cs.CR 2025-07 conditional novelty 6.0 of 10

    WaveVerify embeds audio watermarks with a FiLM-based generator and extracts them with a Mixture-of-Experts detector, reporting zero bit error and high localization under common distortions.

  11. DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.

  12. Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A benchmark of eight post-training quantization methods on Whisper and Moonshine edge speech models across seven datasets, finding 8-bit is safe and 3-bit weights are viable for larger models with advanced methods like SpQR.

  13. Word stress in self-supervised speech models: A cross-linguistic comparison

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Stress classifiers trained on Wav2vec 2.0 embeddings distinguish stressed and unstressed syllables in Dutch, English, German, Polish and Hungarian, with language-specific structure that separates fixed and variable st...

  14. Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A new public benchmark for Arabic mispronunciation detection using Quranic recitation, with baseline models reaching F1 scores below 30%.

  15. DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DeRAGEC explicitly denoises retrieved named-entity candidates with phonetic scores, definitions, and synthetic rationales, improving ASR error-correction WER and NE hit ratio without additional training.

  16. Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

    eess.AS 2025-06 conditional novelty 6.0 of 10

    An LLM conditioned on speaker embeddings and utterance time boundaries jointly transcribes and timestamps overlapping multi-speaker speech.

  17. NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.

  18. FeatureSense: Protecting Speaker Attributes in Always-On Audio Sensing System

    cs.SD 2025-05 conditional novelty 6.0 of 10

    FeatureSense exposes hand-picked audio features instead of raw audio and introduces the SILI metric, claiming 60.6% lower speaker attribute leakage while keeping sound classification accuracy.

  19. PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems

    eess.AS 2025-05 conditional novelty 6.0 of 10

    PSRB, a 10.4-hour Persian benchmark built from 3,372 clips and 756 speakers, evaluates ten ASR models and introduces SW-WER, showing that systems are far weaker on regional accents, children's speech, and informal aud...

  20. TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A new 143-indicator rubric applied to 114 human-voice datasets shows that documentation of consent, privacy, and harmful content is rare, and that scraping yields scale at the cost of documented ethical practices.

  21. Word Level Timestamp Generation for Automatic Speech Recognition and Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The paper teaches the Canary ASR and speech-translation model to output word-level start and end timestamps directly using forced-alignment teacher labels.

  22. Audio Large Language Models Can Be Descriptive Speech Quality Evaluators

    cs.SD 2025-01 conditional novelty 6.0 of 10

    Audio LLMs fine-tuned with token-level distillation against an LLM teacher can predict speech quality scores and generate natural-language descriptions, including A/B comparisons.

  23. Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Omni-Emotion combines face, audio, and video features in a large language model to achieve state-of-the-art scores on emotion recognition and emotion reasoning benchmarks.

  24. TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

    eess.AS 2024-12 conditional novelty 6.0 of 10

    TouchASP trains a single elastic mixture-of-experts ASR model on 1M hours of partly pseudo-labeled audio and reports SpeechIO CER dropping from 4.98% to 2.45% while adding multi-task perception.

  25. Bridging the Data Provenance Gap Across Text, Speech and Video

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...

  26. TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

    cs.SD 2024-12 conditional novelty 6.0 of 10

    TouchTTS reports a 51.6% data retention rate using a noise-robust tokenizer and two-ASR cross-validation, and a Qwen-backbone flow model that unifies streaming and non-streaming synthesis while matching CosyVoice on PER.

  27. Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models

    cs.SD 2024-11 conditional novelty 6.0 of 10

    The CAA benchmark applies content, emotional, explicit noise, and implicit noise attacks to six audio-language models and finds GPT-4o the most robust.

  28. A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition

    eess.AS 2026-03 conditional novelty 5.5 of 10

    On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.

  29. Leveraging Beam Search Information for Confidence Estimation in E2E ASR

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A 0.6k-parameter module that scores ASR tokens and words using only beam-search scores, ranks, context sums, and top-k alternatives substantially reduces calibration error, especially worst-case MCE.

  30. Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

    cs.SD 2026-07 conditional novelty 5.0 of 10

    Multi-ratio diffusion reconstruction residuals, added as a scalar-gated correction to a frozen WavLM anchor, lower ITW EER to 15.3% vs 18.3% for a separately optimized reference.

  31. Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Qwen-Audio-3.0-TTS claims state-of-the-art controllable multilingual text-to-speech across 16 languages and 20 Chinese dialects, using a 12.5 Hz tokenizer and multi-stage RL.

  32. ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

    eess.AS 2026-07 conditional novelty 5.0 of 10

    ProPS uses a mixture density network conditioned on SBERT text embeddings to generate Gaussian mixture models over speaker x-vectors from natural language profile descriptions.

  33. FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation

    eess.AS 2026-01 conditional novelty 5.0 of 10

    A hierarchical Q-Former compresses speech to about 1.67 tokens/sec, enabling hour-long audio processing with near-linear memory scaling and competitive benchmark scores.

  34. DarkStream: real-time speech anonymization with low latency

    eess.AS 2025-09 conditional novelty 5.0 of 10

    DarkStream combines a causal content encoder with limited lookahead, k-means quantization, and a GAN-based pseudo-speaker embedding to anonymize speech in real time with near-chance speaker-verification error rates.

  35. Group Relative Policy Optimization for Speech Recognition

    eess.AS 2025-09 conditional novelty 5.0 of 10

    Applying GRPO with rule-based rewards to LLM-based ASR improves WER by up to 18.4% relative and reduces hallucination errors on unseen acoustic conditions.

  36. Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal

    eess.AS 2025-09 conditional novelty 5.0 of 10

    Dharawal speech embeddings from a 107-language model rank Latin, Maori, Korean, Thai, and Welsh as the most similar high-resource languages, though confusion-based and geometry-based rankings differ.

  37. Large Language Model Data Generation for Enhanced Intent Recognition in German Speech

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    LLM-generated German text data improves intent recognition for elderly German speakers, and the smaller German-focused LeoLM outperforms the much larger ChatGPT as a data generator.

  38. PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy

    cs.CY 2025-07 conditional novelty 5.0 of 10

    Interviews with 20 voice actors reveal risks beyond consent, credit, and compensation, leading to a PRAC3 framework that adds privacy, reputation, and accountability.

  39. Unified Semi-Supervised Pipeline for Automatic Speech Recognition

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A new semi-supervised ASR framework, TopIPL, combines a dynamic pseudo-label cache and top-N checkpoint teacher averaging to improve WER by up to 40 percent in low-resource settings.

  40. WAKE: Watermarking Audio with Key Enrichment

    cs.SD 2025-06 conditional novelty 5.0 of 10

    WAKE embeds and decodes multiple 32-bit audio watermarks with separate 8-bit keys using an invertible neural network, avoiding the overwriting problem in existing systems.

  41. OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    OWSM v4 models, trained on a cleaned 166k-hour multilingual YODAS subset, beat prior open OWSM models and are competitive with Whisper and MMS on several benchmarks.

  42. Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use

    cs.CL 2025-05 conditional novelty 5.0 of 10

    The Loquacious Set is a curated 25,000-hour English ASR corpus combining six open datasets, with commercial-ready licenses and conformer baselines that reach 4.6% WER on LibriSpeech test-other.

  43. DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A single speech encoder trained via ASR-aware distillation with variable attention masking performs competitively in both streaming and full-context modes at 200M and 2B scale.

  44. CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning

    cs.SD 2025-05 reject novelty 5.0 of 10

    A universal adversarial perturbation framework claiming to protect speech against zero-shot voice cloning by degrading cloned outputs while preserving input naturalness.

  45. Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Multi-task behavior imitation with speech-text interleaving improves speech LLM generalization on prompts and zero-shot tasks using only paired speech and transcripts.

  46. From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.

  47. IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

    cs.SD 2025-02 conditional novelty 5.0 of 10

    IndexTTS combines character and pinyin modeling to make Chinese polyphone pronunciation controllable, and reports improved zero-shot voice cloning and naturalness over open-source TTS baselines.

  48. Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.

  49. MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A single aligned multimodal model, MinMo, achieves strong or state-of-the-art results on speech recognition, translation, emotion recognition, voice generation, and full-duplex dialogue while preserving the underlying...

  50. Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A systematic comparison of crowdsourcing, audiobooks, pseudo-labeling, and volunteer recording for Armenian and Georgian ASR, with open datasets and models achieving 9.9% and 5.73% WER respectively.

  51. VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

    eess.AS 2024-12 conditional novelty 5.0 of 10

    VoiceDiT generates speech and matching environmental sounds from text, audio, or image prompts, and reports better speech intelligibility than the VoiceLDM baseline.

  52. AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues

    cs.CV 2024-12 conditional novelty 5.0 of 10

    AV-EmoDialog uses speech and face encoders with a large language model to generate emotion-aware dialogue responses from audio-visual input, reporting better emotional alignment than the compared baselines.

  53. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

    eess.AS 2024-12 conditional novelty 5.0 of 10

    A 0.5B spoken dialogue model trained end-to-end in one stage, with grouped semantic tokens for faster generation and text-only history for multi-turn dialogue.

  54. Open Universal Arabic ASR Leaderboard

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A new Arabic ASR leaderboard ranks 14 open-source models on five multi-dialect datasets and analyzes robustness, speaker bias, and efficiency.

  55. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

    cs.CV 2024-12 conditional novelty 5.0 of 10

    The authors integrate streaming perception, compressed long-term memory, and a reasoning model into one open-source system, reporting SOTA open-source results on several video and audio benchmarks.

  56. Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Fine-tuning Whisper-Small on 3,520 Assamese clips from Common Voice cuts word error rate from 201% to 44% and character error rate from 191% to 13%.

  57. Interfaze: The Future of AI is built on Task-Specific Small Models

    cs.AI 2026-02 reject novelty 4.0 of 10

    Interfaze-Beta uses small specialist models and tools to build a compact context that a general-purpose LLM answers from, reporting competitive benchmark scores without reproducible evidence.

  58. Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages

    eess.AS 2025-07 conditional novelty 4.0 of 10

    In a five-language zero-shot TTS study, no single duration prediction strategy dominates: speaker-prompted durations help some languages, infilling durations help others, and results vary by metric.

  59. ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A three-stage iterative LoRA training recipe (Focus, Feed Back, Fix) is applied to Whisper-large-v3 and Qwen2-Audio, reporting WER reductions on a multilingual ASR benchmark, with the gains attributed to the iterative...

  60. Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit

    cs.CL 2025-06 conditional novelty 4.0 of 10

    LoRA fine-tuning plus custom text normalization reduces Whisper word error rate on cockpit pilot speech from 68.49% to 26.26%.

See all 73 Pith citations

Pith tools