Vision-language models with web-sourced evidence images can verify cross-modal entity consistency in news, outperforming a CNN baseline for locations and events in a zero-shot setting.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Verifying Cross-modal Entity Consistency in News using Vision-language Models
Vision-language models with web-sourced evidence images can verify cross-modal entity consistency in news, outperforming a CNN baseline for locations and events in a zero-shot setting.