Timestamp alignment between ASR transcripts and speaker diarization is claimed to improve speech emotion recognition, but the experiment conflates alignment with fine-tuning of the feature extractors.
IEMOCAP: In- teractive emotional dyadic motion capture database,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
REJECT 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
Timestamp alignment between ASR transcripts and speaker diarization is claimed to improve speech emotion recognition, but the experiment conflates alignment with fine-tuning of the feature extractors.