The Loquacious Set is a curated 25,000-hour English ASR corpus combining six open datasets, with commercial-ready licenses and conformer baselines that reach 4.6% WER on LibriSpeech test-other.
The dataset is easy to reproduce thanks to the SpeechBrain re- leased source code and can be loaded in a single line of code
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
The Loquacious Set is a curated 25,000-hour English ASR corpus combining six open datasets, with commercial-ready licenses and conformer baselines that reach 4.6% WER on LibriSpeech test-other.