A multimodal ensemble combining Whisper audio features, RoBERTa text, quantized F0 and spectral features reaches 39.79% Macro F1 on the INTERSPEECH 2025 naturalistic speech emotion recognition test set.
Our evaluation of unimodal models demonstrated the strong performance of Whisper and XEUS, highlighting their robustness for SER in spontaneous speech
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
A multimodal ensemble combining Whisper audio features, RoBERTa text, quantized F0 and spectral features reaches 39.79% Macro F1 on the INTERSPEECH 2025 naturalistic speech emotion recognition test set.