Pith. sign in

REVIEW 13 cited by

NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.14669 v4 pith:G32AW4F6 submitted 2024-10-18 cs.CV cs.CL

classification cs.CVcs.CL
keywords naturalbenchmodelsvlmslikenaturalsamplesimagesadversarial
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Vision-language models (VLMs) have made significant progress in recent visual-question-answering (VQA) benchmarks that evaluate complex visio-linguistic reasoning. However, are these models truly effective? In this work, we show that VLMs still struggle with natural images and questions that humans can easily answer, which we term natural adversarial samples. We also find it surprisingly easy to generate these VQA samples from natural image-text corpora using off-the-shelf models like CLIP and ChatGPT. We propose a semi-automated approach to collect a new benchmark, NaturalBench, for reliably evaluating VLMs with 10,000 human-verified VQA samples. Crucially, we adopt a $\textbf{vision-centric}$ design by pairing each question with two images that yield different answers, preventing blind solutions from answering without using the images. This makes NaturalBench more challenging than previous benchmarks that can be solved with commonsense priors. We evaluate 53 state-of-the-art VLMs on NaturalBench, showing that models like LLaVA-OneVision, Cambrian-1, Llama3.2-Vision, Molmo, Qwen2-VL, and even GPT-4o lag 50%-70% behind human performance (over 90%). We analyze why NaturalBench is hard from two angles: (1) Compositionality: Solving NaturalBench requires diverse visio-linguistic skills, including understanding attribute bindings, object relationships, and advanced reasoning like logic and counting. To this end, unlike prior work that uses a single tag per sample, we tag each NaturalBench sample with 1 to 8 skill tags for fine-grained evaluation. (2) Biases: NaturalBench exposes severe biases in VLMs, as models often choose the same answer regardless of the image. Lastly, we apply our benchmark curation method to diverse data sources, including long captions (over 100 words) and non-English languages like Chinese and Hindi, highlighting its potential for dynamic evaluations of VLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?

    cs.CV 2025-06 conditional novelty 7.0 of 10

    SpuriVerse, a benchmark of 124 real-world spurious correlation types, shows LVLMs fail on such cases and that diverse synthetic training improves robustness to unseen spurious correlations.

  2. Probing Visual Language Priors in VLMs

    cs.CV 2024-12 conditional novelty 7.0 of 10

    ViLP shows that vision-language models often answer from text priors instead of image content, and an image-corruption DPO method partially fixes this.

  3. Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

    cs.CV 2024-11 conditional novelty 7.0 of 10

    SAVs extract a sparse set of attention head outputs from a frozen large multimodal model and use them as nearest-centroid features, achieving state-of-the-art few-shot vision-language classification without finetuning.

  4. PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A new controlled benchmark shows that four open-source vision-language models detect multimodal sarcasm largely from lexical, stylistic, and OCR surface cues, not from pragmatic image-text understanding.

  5. Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.

  6. COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

    cs.CV 2025-07 conditional novelty 6.0 of 10

    COREVQA introduces a 5,608-pair true/false visual entailment benchmark for crowd images on which the strongest tested vision-language models reach only 77.57% accuracy.

  7. DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new benchmark, DRAMA-X, adds multi-class directional intents, risk labels, and action suggestions for vulnerable road users to frames from the DRAMA dataset, and shows that scene-graph reasoning improves VLM risk sc...

  8. CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

    cs.CV 2025-04 conditional novelty 6.0 of 10

    CAPTURe, a new benchmark for occluded pattern counting, shows that six vision-language models count far worse when objects are hidden, while humans make almost no errors.

  9. Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Multimodal LLMs show systematic weaknesses in instance-level visual correspondence, and CoLVA, trained with a fine-grained vision expert and object-level contrastive learning, reaches 49.8% accuracy on the new MMVM be...

  10. FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FiVL augments vision-language instruction data with GPT-4o-extracted key expressions and segmentation masks, trains LLaVA with a vision-modeling loss that predicts vocabulary tokens for image patches, and measures vis...

  11. Turbo3D: Ultra-fast Text-to-3D Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A text-to-3D system that generates Gaussian splatting assets in 0.35 seconds through dual-teacher distillation and latent-space reconstruction, with quality on par with slower baselines.

  12. Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Repeating the question on both sides of the image (question echoing) closes the question-first accuracy gap in five open VLMs and beats standard single-pass orderings on several VQA benchmarks.

  13. Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A 2B multimodal model trained with HCM-GRPO, a GRPO variant with partial-credit rewards and hard-case oversampling, outperforms larger models on the authors' private physical-plausibility test set.

Pith tools