A multitask classifier built on Whisper features flags multi-speaker, non-English, music, noisy, and synthetic speech clips in large in-the-wild corpora, and a new 21k-clip labeled dataset is released.
It was shown that the AITW dataset can be used to train classifiers on multispeaker, foreign language, background music, noise and synthetic speech labels
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
A multitask classifier built on Whisper features flags multi-speaker, non-English, music, noisy, and synthetic speech clips in large in-the-wild corpora, and a new 21k-clip labeled dataset is released.