Pith. sign in

REVIEW 16 cited by

Scaling Speech Technology to 1,000+ Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13516 v1 pith:HFJQH56S submitted 2023-05-22 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords languagesspeechmodelmultilingualtechnologyfractionlanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to about one hundred languages which is a small fraction of the over 7,000 languages spoken around the world. The Massively Multilingual Speech (MMS) project increases the number of supported languages by 10-40x, depending on the task. The main ingredients are a new dataset based on readings of publicly available religious texts and effectively leveraging self-supervised learning. We built pre-trained wav2vec 2.0 models covering 1,406 languages, a single multilingual automatic speech recognition model for 1,107 languages, speech synthesis models for the same number of languages, as well as a language identification model for 4,017 languages. Experiments show that our multilingual speech recognition model more than halves the word error rate of Whisper on 54 languages of the FLEURS benchmark while being trained on a small fraction of the labeled data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A framework for analyzing concept representations in neural models

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    A new framework shows concept subspaces are not unique, estimator choice affects containment and disentanglement, LEACE works well but generalizes poorly, and HuBERT encodes phone info as contained and disentangled fr...

  2. Tadabur: A Large-Scale Quran Audio Dataset

    cs.SD 2026-04 unverdicted novelty 7.0 of 10

    Tadabur is a large-scale Quran audio dataset with over 1400 hours from 600+ reciters to support speech research and benchmarks.

  3. Training-Free Cross-Lingual Dysarthria Severity Assessment via Phonological Subspace Analysis in Self-Supervised Speech Representations

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    A training-free method quantifies dysarthria severity via d-prime scores on phonological contrasts in HuBERT embeddings, correlating with clinical ratings across 5 languages and multiple conditions.

  4. Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation

    cs.CL 2026-04 conditional novelty 7.0 of 10

    Multilingual ASR models show 39.7-297% zero-shot WER on Pashto public data, Whisper models output correct script in under 0.8% of cases, and fine-tuned models degrade to 32.5-59% WER on out-of-domain sets.

  5. CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Multi-model ASR consensus (BEACON) curates 413 h of CHILDES with corrected timestamps; the 283 h ASR subset yields up to 19.5% relative WER reduction on four held-out child benchmarks.

  6. BlasBench: An Open Benchmark for Irish Speech Recognition

    cs.CL 2026-04 conditional novelty 6.0 of 10

    BlasBench supplies an Irish-aware normalizer and scoring harness that enables reproducible ASR comparisons and exposes a 33-43 point generalization gap for fine-tuned models versus 7-10 points for massively multilingual ones.

  7. Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative

    cs.SD 2026-03 accept novelty 6.0 of 10

    RuASD is a comprehensive Russian speech anti-spoofing dataset featuring 37 synthesis systems and a robustness evaluation pipeline for real-world channel distortions.

  8. Coherence in the brain unfolds across separable temporal regimes

    q-bio.NC 2025-12 conditional novelty 6.0 of 10

    Language coherence arises from slow contextual integration in default-mode cortex and rapid event-driven reconfiguration in auditory and language areas, captured by LLM-derived signals in single-subject fMRI.

  9. On Barriers to Archival Audio Processing

    cs.SD 2025-07 conditional novelty 6.0 of 10

    On archival radio audio, Whisper V3 identifies languages with 91.3% accuracy, but speaker embeddings lose similarity across ages and languages, making speaker recognition unreliable for indexing.

  10. MLAAD: The Multi-Language Audio Anti-Spoofing Dataset

    cs.SD 2024-01 unverdicted novelty 6.0 of 10

    MLAAD provides a large-scale multi-language synthetic audio dataset for training and evaluating audio anti-spoofing models, showing better training performance than InTheWild and FakeOrReal and alternating superiority...

  11. Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

    cs.CL 2026-07 conditional novelty 5.5 of 10

    First end-to-end Efik TTS baseline: a 3-hour single-speaker corpus and four fine-tuned models, with MMS-TTS best at MOS 3.80±0.63 but residual tonal errors.

  12. A Comparative Study of Pre-trained Speech Encoders and Training Objectives for Large-Scale Indic Spoken Language Identification

    eess.AS 2026-06 unverdicted novelty 5.0 of 10

    Frozen FastConformer with hierarchical softmax achieves over 90% macro accuracy on out-of-domain Indic LID benchmarks for 42 languages and outperforms Whisper and other objectives in cross-corpus settings.

  13. DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

    cs.CL 2026-07 reject novelty 4.0 of 10

    Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.

  14. Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    A shared LoRA adapter on VoxCPM2 trained on 26 hours of Khmer and Korean data improves Khmer MOS from 3.85 to 4.23 using under 3% of parameters, with no gain for Korean.

  15. Pretrained self-supervised speech models can recognize unseen consonants

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    Fine-tuned Wav2Vec2 and HuBERT models recognize click consonants more accurately than non-clicks in G|ui and West !Xoon data.

  16. A Hybrid Machine Learning Framework for Optimizing Crop Selection via Agronomic and Economic Forecasting

    cs.LG 2025-07 reject novelty 3.0 of 10

    A two-stage ML pipeline that recommends crops by predicted market price after filtering for agronomic suitability, delivered through a Kannada voice interface.

Pith tools