An ARI layer-mixing module plus co-attention fusion over gender, speaker, style, and ASR auxiliary tasks gives 76.64 to 77.74 percent unweighted accuracy on IEMOCAP, topping prior published results on three SSL encoders.
Speech emotion recognition using self-supervised features,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
An ARI layer-mixing module plus co-attention fusion over gender, speaker, style, and ASR auxiliary tasks gives 76.64 to 77.74 percent unweighted accuracy on IEMOCAP, topping prior published results on three SSL encoders.