REVIEW 4 cited by
Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Emotion and Intent Joint Understanding in Multimodal Conversation (MC-EIU) aims to decode the semantic information manifested in a multimodal conversational history, while inferring the emotions and intents simultaneously for the current utterance. MC-EIU is enabling technology for many human-computer interfaces. However, there is a lack of available datasets in terms of annotation, modality, language diversity, and accessibility. In this work, we propose an MC-EIU dataset, which features 7 emotion categories, 9 intent categories, 3 modalities, i.e., textual, acoustic, and visual content, and two languages, i.e., English and Mandarin. Furthermore, it is completely open-source for free access. To our knowledge, MC-EIU is the first comprehensive and rich emotion and intent joint understanding dataset for multimodal conversation. Together with the release of the dataset, we also develop an Emotion and Intent Interaction (EI$^2$) network as a reference system by modeling the deep correlation between emotion and intent in the multimodal conversation. With comparative experiments and ablation studies, we demonstrate the effectiveness of the proposed EI$^2$ method on the MC-EIU dataset. The dataset and codes will be made available at: https://github.com/MC-EIU/MC-EIU.
Forward citations
Cited by 4 Pith papers
-
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
EmotionTalk provides 19,250 utterances from 744 Chinese dyadic dialogues with emotion, sentiment, and speaking-style caption annotations.
-
OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing
Rationale-privileged on-policy self-distillation reaches 84.19 mean on MER-UniBench by scoring student rollouts with a local teacher that alone sees frontier-generated multimodal evidence.
-
Deep Learning Approaches for Multimodal Intent Recognition: A Survey
A survey of deep learning methods for intent recognition, tracing the field from unimodal text, audio, vision, and EEG approaches to multimodal fusion, alignment, knowledge-augmented, and multi-task models.
-
End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning
On the MC-EIU dataset, semi-supervised training with HuBERT and RoBERTa plus late fusion raises joint emotion and intent recognition scores above unimodal baselines.
Discussion (0). Continue with ORCID to comment.