For coarse visual attributes like color, count, and size, MLLMs encode counterfactual image evidence in their final layers yet fail to follow instructions about whether to trust it; a learned steering vector improves this controllability.
ROME : Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense
1 Pith paper cite this work, alongside 6 external citations. Polarity classification is still indexing.
1
Pith paper citing it
6
external citations · OpenAlex
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
For coarse visual attributes like color, count, and size, MLLMs encode counterfactual image evidence in their final layers yet fail to follow instructions about whether to trust it; a learned steering vector improves this controllability.