A new benchmark shows that multimodal LLMs, including GPT-4o, consistently fail to recognize objects when their colors are modified, and that larger language models can degrade the vision encoder's performance during fine-tuning.
Breaking common sense: WHOOPS! A vision- and-language benchmark of synthetic and compositional im- ages
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?
A new benchmark shows that multimodal LLMs, including GPT-4o, consistently fail to recognize objects when their colors are modified, and that larger language models can degrade the vision encoder's performance during fine-tuning.