Vision-language models recall facts much better when the entity is named in text than when the same entity appears only in an image, and hidden-state probes can detect many of these recall failures.
Finally, the QA pairs are deduplicated using both exact match and Llama-3.1-8B
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can VLMs Recall Factual Associations From Visual References?
Vision-language models recall facts much better when the entity is named in text than when the same entity appears only in an image, and hidden-state probes can detect many of these recall failures.