An end-to-end LLM system that jointly predicts transcripts, emotion descriptors, and emotion labels from speech, with VAE-based disentanglement, beats its own multi-task baselines on IEMOCAP and MELD.
14200220, 14200021, 14200324, Innovation Technology Fund grant No
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
An end-to-end LLM system that jointly predicts transcripts, emotion descriptors, and emotion labels from speech, with VAE-based disentanglement, beats its own multi-task baselines on IEMOCAP and MELD.