Using subtitles as prompts and pseudo transcripts as targets, with a Gini-based attention weighting, refines Whisper's transcripts on low-resource Flemish TV speech without verbatim labels.
Weakly Supervised Construction of ASR Systems with Massive Video Data
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Building Automatic Speech Recognition (ASR) systems from scratch is significantly challenging, mostly due to the time-consuming and financially-expensive process of annotating a large amount of audio data with transcripts. Although several unsupervised pre-training models have been proposed, applying such models directly might still be sub-optimal if more labeled, training data could be obtained without a large cost. In this paper, we present a weakly supervised framework for constructing ASR systems with massive video data. As videos often contain human-speech audios aligned with subtitles, we consider videos as an important knowledge source, and propose an effective approach to extract high-quality audios aligned with transcripts from videos based on Optical Character Recognition (OCR). The underlying ASR model can be fine-tuned to fit any domain-specific target training datasets after weakly supervised pre-training. Extensive experiments show that our framework can easily produce state-of-the-art results on six public datasets for Mandarin speech recognition.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Refining Transcripts With TV Subtitles by Prompt-Based Weakly Supervised Training of ASR
Using subtitles as prompts and pseudo transcripts as targets, with a Gini-based attention weighting, refines Whisper's transcripts on low-resource Flemish TV speech without verbatim labels.