Pith. sign in

REVIEW 1 cited by

Holistic Visual-Textual Sentiment Analysis with Prior Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12981 v2 pith:KGHKRBPK submitted 2022-11-23 cs.CV cs.MM

classification cs.CVcs.MM
keywords sentimentvisual-textualanalysisfeaturesbranchmethodvisualexpert
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Visual-textual sentiment analysis aims to predict sentiment with the input of a pair of image and text, which poses a challenge in learning effective features for diverse input images. To address this, we propose a holistic method that achieves robust visual-textual sentiment analysis by exploiting a rich set of powerful pre-trained visual and textual prior models. The proposed method consists of four parts: (1) a visual-textual branch to learn features directly from data for sentiment analysis, (2) a visual expert branch with a set of pre-trained "expert" encoders to extract selected semantic visual features, (3) a CLIP branch to implicitly model visual-textual correspondence, and (4) a multimodal feature fusion network based on BERT to fuse multimodal features and make sentiment predictions. Extensive experiments on three datasets show that our method produces better visual-textual sentiment analysis performance than existing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A structured prompt that makes LLaVA predict image, text, and multimodal sentiment labels together yields state-of-the-art accuracy and F1 on MVSA-Single.

Pith tools