A structured prompt that makes LLaVA predict image, text, and multimodal sentiment labels together yields state-of-the-art accuracy and F1 on MVSA-Single.
A Fair and Comprehensive Comparison of Multimodal Tweet Sentiment Analysis Methods
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Opinion and sentiment analysis is a vital task to characterize subjective information in social media posts. In this paper, we present a comprehensive experimental evaluation and comparison with six state-of-the-art methods, from which we have re-implemented one of them. In addition, we investigate different textual and visual feature embeddings that cover different aspects of the content, as well as the recently introduced multimodal CLIP embeddings. Experimental results are presented for two different publicly available benchmark datasets of tweets and corresponding images. In contrast to the evaluation methodology of previous work, we introduce a reproducible and fair evaluation scheme to make results comparable. Finally, we conduct an error analysis to outline the limitations of the methods and possibilities for the future work.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier
A structured prompt that makes LLaVA predict image, text, and multimodal sentiment labels together yields state-of-the-art accuracy and F1 on MVSA-Single.