REVIEW 10 cited by
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advances in speech recognition and translation rely on hundreds of thousands of hours of Internet speech data. We argue that state-of-the art accuracy can be reached without relying on web-scale data. Canary - multilingual ASR and speech translation model, outperforms current state-of-the-art models - Whisper, OWSM, and Seamless-M4T on English, French, Spanish, and German languages, while being trained on an order of magnitude less data than these models. Three key factors enables such data-efficient model: (1) a FastConformer-based attention encoder-decoder architecture (2) training on synthetic data generated with machine translation and (3) advanced training techniques: data-balancing, dynamic data blending, dynamic bucketing and noise-robust fine-tuning. The model, weights, and training code will be open-sourced.
Forward citations
Cited by 10 Pith papers
-
Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
Six monolingual 27M-parameter ASR models are reported to outperform Whisper Tiny and Small, and sometimes Whisper Medium, but several evaluations use test sets that were included in training.
-
LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
LCS-CTC, a phoneme recognizer trained with similarity-aware LCS alignment masks constraining CTC, outperforms vanilla CTC on all reported PER, WPER, boundary-loss, and articulatory metrics.
-
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
The paper teaches the Canary ASR and speech-translation model to output word-level start and end timestamps directly using forced-alignment teacher labels.
-
Granary: Speech Recognition and Translation Dataset in 25 European Languages
Granary releases 643k hours of pseudo-labeled ASR data and 351k hours of X-to-English translation pairs for 25 European languages, with evidence that its cleaned English and Croatian subsets train ASR models comparabl...
-
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
A new open suite of 13 multilingual speech models, up to 18B parameters, yields empirical scaling laws for ASR and speech translation performance.
-
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
Conditioning an SOT multi-talker ASR decoder on EEND-EDA speaker embeddings and activity information lowers WER on Libri2Mix and Libri3Mix, provided the diarization branch is accurate.
-
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
The authors present SwitchLingua, a large multilingual code-switching text and audio dataset, and SAER, a semantic-aware error metric for code-switching ASR evaluation.
-
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
A 68M-parameter Vietnamese ASR model, pretrained on 70,000 hours of unlabeled audio and fine-tuned on 50 hours of labels, reports average WER 8.31, beating Whisper Large-v3 and commercial systems.
-
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.
-
Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People
For the Clarity speech intelligibility task, a single encoder layer plus a temporal transformer head and a three-model ensemble beats use-all-layers pooling.
Discussion (0). Continue with ORCID to comment.