Pith. sign in

REVIEW 1 cited by

Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.00047 v1 pith:UDGIOVLS submitted 2023-10-31 cs.AI cs.CLcs.CVcs.LG

Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?

classification cs.AI cs.CLcs.CVcs.LG
keywords illusionsvisualhumansmodelsdatasetvlmsworldbetter
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Vision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world. However, known as visual illusions, human's perception of reality isn't always faithful to the physical world. This raises a key question: do VLMs have the similar kind of illusions as humans do, or do they faithfully learn to represent reality? To investigate this question, we build a dataset containing five types of visual illusions and formulate four tasks to examine visual illusions in state-of-the-art VLMs. Our findings have shown that although the overall alignment is low, larger models are closer to human perception and more susceptible to visual illusions. Our dataset and initial findings will promote a better understanding of visual illusions in humans and machines and provide a stepping stone for future computational models that can better align humans and machines in perceiving and communicating about the shared visual world. The code and data are available at https://github.com/vl-illusion/dataset.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions

    cs.CV 2026-03 conditional novelty 6.0

    SMSP, a plug-and-play multi-scale low-pass preprocessing method, lets MLLMs recognize hidden characters in visual illusions by reducing high-frequency background distraction.