SimClass is a new 391-hour simulated classroom speech dataset with game-engine babble noise; ASR fine-tuning on it beats Librispeech and TEDLIUM on real classroom test sets.
My Science Tutor (MyST) -- A Large Corpus of Children's Conversational Speech
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This article describes the MyST corpus developed as part of the My Science Tutor project -- one of the largest collections of children's conversational speech comprising approximately 400 hours, spanning some 230K utterances across about 10.5K virtual tutor sessions by around 1.3K third, fourth and fifth grade students. 100K of all utterances have been transcribed thus far. The corpus is freely available (https://myst.cemantix.org) for non-commercial use using a creative commons license. It is also available for commercial use (https://boulderlearning.com/resources/myst-corpus/). To date, ten organizations have licensed the corpus for commercial use, and approximately 40 university and other not-for-profit research groups have downloaded the corpus. It is our hope that the corpus can be used to improve automatic speech recognition algorithms, build and evaluate conversational AI agents for education, and together help accelerate development of multimodal applications to improve children's excitement and learning about science, and help them learn remotely.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
SimClass is a new 391-hour simulated classroom speech dataset with game-engine babble noise; ASR fine-tuning on it beats Librispeech and TEDLIUM on real classroom test sets.