REVIEW 73 cited by
Common Voice: A Massively-Multilingual Speech Corpus
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identification). To achieve scale and sustainability, the Common Voice project employs crowdsourcing for both data collection and data validation. The most recent release includes 29 languages, and as of November 2019 there are a total of 38 languages collecting data. Over 50,000 individuals have participated so far, resulting in 2,500 hours of collected audio. To our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages. As an example use case for Common Voice, we present speech recognition experiments using Mozilla's DeepSpeech Speech-to-Text toolkit. By applying transfer learning from a source English model, we find an average Character Error Rate improvement of 5.99 +/- 5.48 for twelve target languages (German, French, Italian, Turkish, Catalan, Slovenian, Welsh, Irish, Breton, Tatar, Chuvash, and Kabyle). For most of these languages, these are the first ever published results on end-to-end Automatic Speech Recognition.
Forward citations
Showing 60 of 73 Pith papers that cite this
-
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
REDDIT corrects non-speech-induced timestamp drift in autoregressive ASR by editing timestamp targets under cached replay context while anchoring non-timestamp behavior to the frozen base distribution.
-
Continual Learning with Embedding Layer Surgery and Task-wise Beam Search using Whisper
Embedding Layer Surgery and Task-wise Beam Search reduce catastrophic forgetting when adding new languages to Whisper, lowering old-language AWER from 14.2% to 11.9% versus Experience Replay.
-
SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.
-
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
A 10.7M-pair audio-caption corpus and systematic comparison show contrastive pretraining is more data-efficient while captioning scales better, and supervised initialization yields diminishing returns.
-
ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
ParsVoice is an open ~1,800–2,200-hour Persian audiobook-derived speech-text corpus, much larger than prior open Persian TTS datasets.
-
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
A four-part benchmark plus a semantic/acoustic token taxonomy for comparing audio codecs, with correlation analysis across ten models.
-
AHELM: A Holistic Evaluation of Audio-Language Models
AHELM standardizes evaluation of audio-language models across 10 aspects and shows simple ASR+LLM systems are competitive with multimodal models.
-
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.
-
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.
-
WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes
WaveVerify embeds audio watermarks with a FiLM-based generator and extracts them with a Mixture-of-Experts detector, reporting zero bit error and high localization under common distortions.
-
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.
-
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
A benchmark of eight post-training quantization methods on Whisper and Moonshine edge speech models across seven datasets, finding 8-bit is safe and 3-bit weights are viable for larger models with advanced methods like SpQR.
-
Word stress in self-supervised speech models: A cross-linguistic comparison
Stress classifiers trained on Wav2vec 2.0 embeddings distinguish stressed and unstressed syllables in Dutch, English, German, Polish and Hungarian, with language-specific structure that separates fixed and variable st...
-
Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study
A new public benchmark for Arabic mispronunciation detection using Quranic recitation, with baseline models reaching F1 scores below 30%.
-
DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction
DeRAGEC explicitly denoises retrieved named-entity candidates with phonetic scores, definitions, and synthetic rationales, improving ASR error-correction WER and NE hit ratio without additional training.
-
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
An LLM conditioned on speaker embeddings and utterance time boundaries jointly transcribes and timestamps overlapping multi-speaker speech.
-
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.
-
FeatureSense: Protecting Speaker Attributes in Always-On Audio Sensing System
FeatureSense exposes hand-picked audio features instead of raw audio and introduces the SILI metric, claiming 60.6% lower speaker attribute leakage while keeping sound classification accuracy.
-
PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
PSRB, a 10.4-hour Persian benchmark built from 3,372 clips and 756 speakers, evaluates ten ASR models and introduces SW-WER, showing that systems are far weaker on regional accents, children's speech, and informal aud...
-
TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
A new 143-indicator rubric applied to 114 human-voice datasets shows that documentation of consent, privacy, and harmful content is rare, and that scraping yields scale at the cost of documented ethical practices.
-
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
The paper teaches the Canary ASR and speech-translation model to output word-level start and end timestamps directly using forced-alignment teacher labels.
-
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
Audio LLMs fine-tuned with token-level distillation against an LLM teacher can predict speech quality scores and generate natural-language descriptions, including A/B comparisons.
-
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
Omni-Emotion combines face, audio, and video features in a large language model to achieve state-of-the-art scores on emotion recognition and emotion reasoning benchmarks.
-
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
TouchASP trains a single elastic mixture-of-experts ASR model on 1M hours of partly pseudo-labeled audio and reports SpeechIO CER dropping from 4.98% to 2.45% while adding multi-task perception.
-
Bridging the Data Provenance Gap Across Text, Speech and Video
A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...
-
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
TouchTTS reports a 51.6% data retention rate using a noise-robust tokenizer and two-ASR cross-validation, and a Qwen-backbone flow model that unifies streaming and non-streaming synthesis while matching CosyVoice on PER.
-
Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
The CAA benchmark applies content, emotional, explicit noise, and implicit noise attacks to six audio-language models and finds GPT-4o the most robust.
-
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.
-
Leveraging Beam Search Information for Confidence Estimation in E2E ASR
A 0.6k-parameter module that scores ASR tokens and words using only beam-search scores, ranks, context sums, and top-k alternatives substantially reduces calibration error, especially worst-case MCE.
-
Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection
Multi-ratio diffusion reconstruction residuals, added as a scalar-gated correction to a frozen WavLM anchor, lower ITW EER to 15.3% vs 18.3% for a separately optimized reference.
-
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
Qwen-Audio-3.0-TTS claims state-of-the-art controllable multilingual text-to-speech across 16 languages and 20 Chinese dialects, using a 12.5 Hz tokenizer and multi-stage RL.
-
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
ProPS uses a mixture density network conditioned on SBERT text embeddings to generate Gaussian mixture models over speaker x-vectors from natural language profile descriptions.
-
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
A hierarchical Q-Former compresses speech to about 1.67 tokens/sec, enabling hour-long audio processing with near-linear memory scaling and competitive benchmark scores.
-
DarkStream: real-time speech anonymization with low latency
DarkStream combines a causal content encoder with limited lookahead, k-means quantization, and a GAN-based pseudo-speaker embedding to anonymize speech in real time with near-chance speaker-verification error rates.
-
Group Relative Policy Optimization for Speech Recognition
Applying GRPO with rule-based rewards to LLM-based ASR improves WER by up to 18.4% relative and reduces hallucination errors on unseen acoustic conditions.
-
Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal
Dharawal speech embeddings from a 107-language model rank Latin, Maori, Korean, Thai, and Welsh as the most similar high-resource languages, though confusion-based and geometry-based rankings differ.
-
Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
LLM-generated German text data improves intent recognition for elderly German speakers, and the smaller German-focused LeoLM outperforms the much larger ChatGPT as a data generator.
-
PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy
Interviews with 20 voice actors reveal risks beyond consent, credit, and compensation, leading to a PRAC3 framework that adds privacy, reputation, and accountability.
-
Unified Semi-Supervised Pipeline for Automatic Speech Recognition
A new semi-supervised ASR framework, TopIPL, combines a dynamic pseudo-label cache and top-N checkpoint teacher averaging to improve WER by up to 40 percent in low-resource settings.
-
WAKE: Watermarking Audio with Key Enrichment
WAKE embeds and decodes multiple 32-bit audio watermarks with separate 8-bit keys using an invertible neural network, avoiding the overwriting problem in existing systems.
-
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
OWSM v4 models, trained on a cleaned 166k-hour multilingual YODAS subset, beat prior open OWSM models and are competitive with Whisper and MMS on several benchmarks.
-
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
The Loquacious Set is a curated 25,000-hour English ASR corpus combining six open datasets, with commercial-ready licenses and conformer baselines that reach 4.6% WER on LibriSpeech test-other.
-
DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation
A single speech encoder trained via ASR-aware distillation with variable attention masking performs competitively in both streaming and full-context modes at 200M and 2B scale.
-
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
A universal adversarial perturbation framework claiming to protect speech against zero-shot voice cloning by degrading cloned outputs while preserving input naturalness.
-
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
Multi-task behavior imitation with speech-text interleaving improves speech LLM generalization on prompts and zero-shot tasks using only paired speech and transcripts.
-
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.
-
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
IndexTTS combines character and pinyin modeling to make Chinese polyphone pronunciation controllable, and reports improved zero-shot voice cloning and naturalness over open-source TTS baselines.
-
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.
-
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
A single aligned multimodal model, MinMo, achieves strong or state-of-the-art results on speech recognition, translation, emotion recognition, voice generation, and full-duplex dialogue while preserving the underlying...
-
Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages
A systematic comparison of crowdsourcing, audiobooks, pseudo-labeling, and volunteer recording for Armenian and Georgian ASR, with open datasets and models achieving 9.9% and 5.73% WER respectively.
-
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
VoiceDiT generates speech and matching environmental sounds from text, audio, or image prompts, and reports better speech intelligibility than the VoiceLDM baseline.
-
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
AV-EmoDialog uses speech and face encoders with a large language model to generate emotion-aware dialogue responses from audio-visual input, reporting better emotional alignment than the compared baselines.
-
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
A 0.5B spoken dialogue model trained end-to-end in one stage, with grouped semantic tokens for faster generation and text-only history for multi-turn dialogue.
-
Open Universal Arabic ASR Leaderboard
A new Arabic ASR leaderboard ranks 14 open-source models on five multi-dialect datasets and analyzes robustness, speaker bias, and efficiency.
-
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
The authors integrate streaming perception, compressed long-term memory, and a reasoning model into one open-source system, reporting SOTA open-source results on several video and audio benchmarks.
-
Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models
Fine-tuning Whisper-Small on 3,520 Assamese clips from Common Voice cuts word error rate from 201% to 44% and character error rate from 191% to 13%.
-
Interfaze: The Future of AI is built on Task-Specific Small Models
Interfaze-Beta uses small specialist models and tools to build a compact context that a general-purpose LLM answers from, reporting competitive benchmark scores without reproducible evidence.
-
Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages
In a five-language zero-shot TTS study, no single duration prediction strategy dominates: speaker-prompted durations help some languages, infilling durations help others, and results vary by metric.
-
ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition
A three-stage iterative LoRA training recipe (Focus, Feed Back, Fix) is applied to Whisper-large-v3 and Qwen2-Audio, reporting WER reductions on a multilingual ASR benchmark, with the gains attributed to the iterative...
-
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
LoRA fine-tuning plus custom text normalization reduces Whisper word error rate on cockpit pilot speech from 68.49% to 26.26%.
Discussion (0). Continue with ORCID to comment.