A new controlled benchmark shows that four open-source vision-language models detect multimodal sarcasm largely from lexical, stylistic, and OCR surface cues, not from pragmatic image-text understanding.
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models
A new controlled benchmark shows that four open-source vision-language models detect multimodal sarcasm largely from lexical, stylistic, and OCR surface cues, not from pragmatic image-text understanding.