Pith. sign in

REVIEW 4 cited by

Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19674 v1 pith:RZFDABDC submitted 2024-06-28 cs.CL cs.LGcs.SDeess.AS

Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

classification cs.CL cs.LGcs.SDeess.AS
keywords dataspeechtranslationmodeltrainingdynamiclessmodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent advances in speech recognition and translation rely on hundreds of thousands of hours of Internet speech data. We argue that state-of-the art accuracy can be reached without relying on web-scale data. Canary - multilingual ASR and speech translation model, outperforms current state-of-the-art models - Whisper, OWSM, and Seamless-M4T on English, French, Spanish, and German languages, while being trained on an order of magnitude less data than these models. Three key factors enables such data-efficient model: (1) a FastConformer-based attention encoder-decoder architecture (2) training on synthetic data generated with machine translation and (3) advanced training techniques: data-balancing, dynamic data blending, dynamic bucketing and noise-robust fine-tuning. The model, weights, and training code will be open-sourced.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CTC-Seeded Token Edit Refinement for Non-Autoregressive Speech Recognition

    eess.AS 2026-06 unverdicted novelty 6.0

    CTC-seeded variable-length edit refinement with a diffusion-based Edit Flow decoder achieves WER reductions in non-autoregressive ASR using only two inference steps plus classifier-free guidance.

  2. BlasBench: An Open Benchmark for Irish Speech Recognition

    cs.CL 2026-04 conditional novelty 6.0

    BlasBench supplies an Irish-aware normalizer and scoring harness that enables reproducible ASR comparisons and exposes a 33-43 point generalization gap for fine-tuned models versus 7-10 points for massively multilingual ones.

  3. LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness

    eess.AS 2025-08 conditional novelty 6.0

    LCS-CTC, a phoneme recognizer trained with similarity-aware LCS alignment masks constraining CTC, outperforms vanilla CTC on all reported PER, WPER, boundary-loss, and articulatory metrics.

  4. Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech

    eess.AS 2026-05 unverdicted novelty 4.0

    Frame-aligned fusion of Canary and WavLM encoders, with WavLM temporally prepared via learnable strided convolution, outperforms other fusion strategies and reaches Eval RMSE 24.96 and Corr 0.796 on non-intrusive inte...