Pith. sign in

Breaking common sense: WHOOPS! A vision- and-language benchmark of synthetic and compositional im- ages

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CV 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?

cs.CV · 2024-11-26 · conditional · novelty 6.0

A new benchmark shows that multimodal LLMs, including GPT-4o, consistently fail to recognize objects when their colors are modified, and that larger language models can degrade the vision encoder's performance during fine-tuning.

citing papers explorer

Showing 1 of 1 citing paper.

  • NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? cs.CV · 2024-11-26 · conditional · none · ref 3

    A new benchmark shows that multimodal LLMs, including GPT-4o, consistently fail to recognize objects when their colors are modified, and that larger language models can degrade the vision encoder's performance during fine-tuning.