Internal attention patterns in multimodal LLMs are used to define an attention accuracy metric and a benchmark for detecting cases where a model answers correctly while attending to the wrong image.
Claude 3 haiku: our fastest model yet
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
Internal attention patterns in multimodal LLMs are used to define an attention accuracy metric and a benchmark for detecting cases where a model answers correctly while attending to the wrong image.