Newer, larger deep learning models show no consistent gains over older architectures for speech emotion recognition across two naturalistic benchmarks, with results sensitive to model selection and hyperparameters.
Computer Audition: From Task-Specific Machine Learning to Foundation Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Foundation models (FMs) are increasingly spearheading recent advances on a variety of tasks that fall under the purview of computer audition -- the use of machines to understand sounds. They feature several advantages over traditional pipelines: among others, the ability to consolidate multiple tasks in a single model, the option to leverage knowledge from other modalities, and the readily-available interaction with human users. Naturally, these promises have created substantial excitement in the audio community, and have led to a wave of early attempts to build new, general-purpose foundation models for audio. In the present contribution, we give an overview of computational audio analysis as it transitions from traditional pipelines towards auditory foundation models. Our work highlights the key operating principles that underpin those models, and showcases how they can accommodate multiple tasks that the audio community previously tackled separately.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
Newer, larger deep learning models show no consistent gains over older architectures for speech emotion recognition across two naturalistic benchmarks, with results sensitive to model selection and hyperparameters.