REVIEW 13 cited by
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Vision-language models (VLMs) have made significant progress in recent visual-question-answering (VQA) benchmarks that evaluate complex visio-linguistic reasoning. However, are these models truly effective? In this work, we show that VLMs still struggle with natural images and questions that humans can easily answer, which we term natural adversarial samples. We also find it surprisingly easy to generate these VQA samples from natural image-text corpora using off-the-shelf models like CLIP and ChatGPT. We propose a semi-automated approach to collect a new benchmark, NaturalBench, for reliably evaluating VLMs with 10,000 human-verified VQA samples. Crucially, we adopt a $\textbf{vision-centric}$ design by pairing each question with two images that yield different answers, preventing blind solutions from answering without using the images. This makes NaturalBench more challenging than previous benchmarks that can be solved with commonsense priors. We evaluate 53 state-of-the-art VLMs on NaturalBench, showing that models like LLaVA-OneVision, Cambrian-1, Llama3.2-Vision, Molmo, Qwen2-VL, and even GPT-4o lag 50%-70% behind human performance (over 90%). We analyze why NaturalBench is hard from two angles: (1) Compositionality: Solving NaturalBench requires diverse visio-linguistic skills, including understanding attribute bindings, object relationships, and advanced reasoning like logic and counting. To this end, unlike prior work that uses a single tag per sample, we tag each NaturalBench sample with 1 to 8 skill tags for fine-grained evaluation. (2) Biases: NaturalBench exposes severe biases in VLMs, as models often choose the same answer regardless of the image. Lastly, we apply our benchmark curation method to diverse data sources, including long captions (over 100 words) and non-English languages like Chinese and Hindi, highlighting its potential for dynamic evaluations of VLMs.
Forward citations
Cited by 13 Pith papers
-
Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
SpuriVerse, a benchmark of 124 real-world spurious correlation types, shows LVLMs fail on such cases and that diverse synthetic training improves robustness to unseen spurious correlations.
-
Probing Visual Language Priors in VLMs
ViLP shows that vision-language models often answer from text priors instead of image content, and an image-corruption DPO method partially fixes this.
-
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
SAVs extract a sparse set of attention head outputs from a frozen large multimodal model and use them as nearest-centroid features, achieving state-of-the-art few-shot vision-language classification without finetuning.
-
PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models
A new controlled benchmark shows that four open-source vision-language models detect multimodal sarcasm largely from lexical, stylistic, and OCR surface cues, not from pragmatic image-text understanding.
-
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.
-
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
COREVQA introduces a 5,608-pair true/false visual entailment benchmark for crowd images on which the strongest tested vision-language models reach only 77.57% accuracy.
-
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
A new benchmark, DRAMA-X, adds multi-class directional intents, risk labels, and action suggestions for vulnerable road users to frames from the DRAMA dataset, and shows that scene-graph reasoning improves VLM risk sc...
-
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
CAPTURe, a new benchmark for occluded pattern counting, shows that six vision-language models count far worse when objects are hidden, while humans make almost no errors.
-
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Multimodal LLMs show systematic weaknesses in instance-level visual correspondence, and CoLVA, trained with a fine-grained vision expert and object-level contrastive learning, reaches 49.8% accuracy on the new MMVM be...
-
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
FiVL augments vision-language instruction data with GPT-4o-extracted key expressions and segmentation masks, trains LLaVA with a vision-modeling loss that predicts vocabulary tokens for image patches, and measures vis...
-
Turbo3D: Ultra-fast Text-to-3D Generation
A text-to-3D system that generates Gaussian splatting assets in 0.35 seconds through dual-teacher distillation and latent-space reconstruction, with quality on par with slower baselines.
-
Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models
Repeating the question on both sides of the image (question echoing) closes the question-first accuracy gap in five open VLMs and beats standard single-pass orderings on several VQA benchmarks.
-
Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance
A 2B multimodal model trained with HCM-GRPO, a GRPO variant with partial-credit rewards and hard-case oversampling, outperforms larger models on the authors' private physical-plausibility test set.
Discussion (0). Continue with ORCID to comment.