A simple multimodal system with soft labels and audio augmentation reaches top-3 in the IS25 emotion recognition challenge.
One is to study pre-trained speech models with emotional speech data like Emotion2Vec [21]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
A simple multimodal system with soft labels and audio augmentation reaches top-3 in the IS25 emotion recognition challenge.