SSL convolutional audio models pre-trained on speech, non-speech, or both perform nearly equally well across speech and non-speech downstream tasks, while domain-specific baselines struggle outside their domains.
Embeddings for up to 2000 random samples from the validation partition of select datasets representing speech, non-speech and VAD audio domains
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study
SSL convolutional audio models pre-trained on speech, non-speech, or both perform nearly equally well across speech and non-speech downstream tasks, while domain-specific baselines struggle outside their domains.