REVIEW 1 cited by
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
read the original abstract
Emotion recognition and sentiment analysis are pivotal tasks in speech and language processing, particularly in real-world scenarios involving multi-party, conversational data. This paper presents a multimodal approach to tackle these challenges on a well-known dataset. We propose a system that integrates four key modalities/channels using pre-trained models: RoBERTa for text, Wav2Vec2 for speech, a proposed FacialNet for facial expressions, and a CNN+Transformer architecture trained from scratch for video analysis. Feature embeddings from each modality are concatenated to form a multimodal vector, which is then used to predict emotion and sentiment labels. The multimodal system demonstrates superior performance compared to unimodal approaches, achieving an accuracy of 66.36% for emotion recognition and 72.15% for sentiment analysis.
Forward citations
Cited by 1 Pith paper
-
EmoScene: A Dual-space Dataset for Controllable Affective Image Generation
EmoScene contributes 1.2M images annotated with discrete emotions, continuous VAD scores, perceptual attributes, and captions, plus a cross-attention modulation that shifts generated images toward requested affective targets.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.