HuMP-CAT fuses HuBERT, MFCC and prosodic features with a cross-attention transformer and fine-tunes on small target-language subsets to reach a stated 78.75% average accuracy across seven SER corpora.
A comprehensive review of speech emotion recognition systems,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition
HuMP-CAT fuses HuBERT, MFCC and prosodic features with a cross-attention transformer and fine-tunes on small target-language subsets to reach a stated 78.75% average accuracy across seven SER corpora.