An ARI layer-mixing module plus co-attention fusion over gender, speaker, style, and ASR auxiliary tasks gives 76.64 to 77.74 percent unweighted accuracy on IEMOCAP, topping prior published results on three SSL encoders.
Temporal modeling matters: A novel temporal emotional modeling approach for speech emotion recognition,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
An ARI layer-mixing module plus co-attention fusion over gender, speaker, style, and ASR auxiliary tasks gives 76.64 to 77.74 percent unweighted accuracy on IEMOCAP, topping prior published results on three SSL encoders.