Pith. sign in

REVIEW 2 cited by

ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.19301 v1 pith:3OFEEWFU submitted 2023-10-30 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords reasoningvision-languagebeyondmodelsromecommoncommonsenseknowledge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Humans possess a strong capability for reasoning beyond common sense. For example, given an unconventional image of a goldfish laying on the table next to an empty fishbowl, a human would effortlessly determine that the fish is not inside the fishbowl. The case, however, may be different for a vision-language model, whose reasoning could gravitate towards the common scenario that the fish is inside the bowl, despite the visual input. In this paper, we introduce a novel probing dataset named ROME (reasoning beyond commonsense knowledge) to evaluate whether the state-of-the-art pre-trained vision-language models have the reasoning capability to correctly interpret counter-intuitive content. ROME contains images that defy commonsense knowledge with regards to color, shape, material, size and positional relation. Experiments on the state-of-the-art pre-trained vision-language models reveal that most of these models are still largely incapable of interpreting counter-intuitive scenarios. We hope that ROME will spur further investigations on reasoning beyond commonsense knowledge in vision-language research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation

    cs.CL 2026-01 conditional novelty 6.0 of 10

    LLM judges frequently mark candidates wrong when the gold reference contradicts the model's own knowledge, even if the candidate exactly matches the provided reference.

  2. Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Vision-language models can spot obvious virtual objects in AR photos but frequently miss seamlessly integrated ones, with performance dropping sharply as scene complexity increases.

Pith tools