REVIEW 3 cited by
DialogueTRM: Exploring the Intra- and Inter-Modal Emotional Behaviors in the Conversation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Emotion Recognition in Conversations (ERC) is essential for building empathetic human-machine systems. Existing studies on ERC primarily focus on summarizing the context information in a conversation, however, ignoring the differentiated emotional behaviors within and across different modalities. Designing appropriate strategies that fit the differentiated multi-modal emotional behaviors can produce more accurate emotional predictions. Thus, we propose the DialogueTransformer to explore the differentiated emotional behaviors from the intra- and inter-modal perspectives. For intra-modal, we construct a novel Hierarchical Transformer that can easily switch between sequential and feed-forward structures according to the differentiated context preference within each modality. For inter-modal, we constitute a novel Multi-Grained Interactive Fusion that applies both neuron- and vector-grained feature interactions to learn the differentiated contributions across all modalities. Experimental results show that DialogueTRM outperforms the state-of-the-art by a significant margin on three benchmark datasets.
Forward citations
Cited by 3 Pith papers
-
Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC
VEGA aligns multimodal emotion features with CLIP-derived visual emotion prototypes and reports SOTA on IEMOCAP and MELD.
-
TED: Turn Emphasis with Dialogue Feature Attention for Emotion Recognition in Conversation
TED adds dialogue-aware attention weighting (turn priority, speaker/listener factors) to a RoBERTa-based turn-averaging model and reports the best IEMOCAP score, though the gain is minimal.
-
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
WavFusion fuses wav2vec 2.0 audio features with text and visual features using gated cross-modal attention and a homogeneous-feature margin loss, reporting modest state-of-the-art gains on IEMOCAP and MELD.
Discussion (0). Continue with ORCID to comment.