An end-to-end LLM system that jointly predicts transcripts, emotion descriptors, and emotion labels from speech, with VAE-based disentanglement, beats its own multi-task baselines on IEMOCAP and MELD.
In contrast, prior research employs external ASR models to generate transcripts before performing SER [12–17], which may lead to low SER performance affected by ASR er- rors
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
An end-to-end LLM system that jointly predicts transcripts, emotion descriptors, and emotion labels from speech, with VAE-based disentanglement, beats its own multi-task baselines on IEMOCAP and MELD.