Recent MLLMs such as Qwen2-VL can match or beat CLIP-style models on several image classification benchmarks, with gains driven mainly by stronger LLMs and more diverse training data.
Why are visually-grounded language models bad at image classi- fication? In NeurIPS, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
contradiction 1
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
contradiction 1polarities
contest 1representative citing papers
citing papers explorer
-
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
Recent MLLMs such as Qwen2-VL can match or beat CLIP-style models on several image classification benchmarks, with gains driven mainly by stronger LLMs and more diverse training data.