Pith. sign in

REVIEW 2 cited by

The Illusion-Illusion: Vision Language Models See Illusions Where There Are None

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18613 v2 pith:CJQEKN6W submitted 2024-12-07 q-bio.NC cs.CLcs.CV

classification q-bio.NCcs.CLcs.CV
keywords illusionsmodelslanguageprocessingsomethingvisionappearscrooked
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something `really is' and how something `appears to be', and this gap helps us understand the mental processing that led to how something appears to be. Illusions are also useful for investigating artificial systems, and much research has examined whether computational models of perception fall prey to the same illusions as people. Here, I invert the standard use of perceptual illusions to examine basic processing errors in current vision language models. I present these models with illusory-illusions, neighbors of common illusions that should not elicit processing errors. These include such things as perfectly reasonable ducks, crooked lines that truly are crooked, circles that seem to have different sizes because they are, in fact, of different sizes, and so on. I show that many current vision language systems mistakenly see these illusion-illusions as illusions. I suggest that such failures are part of broader failures already discussed in the literature.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Adaptive multi-view template retrieval lifts hidden-hate detection on HatefulIllusion to 93.2% balanced accuracy, far above original-view filters and moderators.

  2. Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Large vision-language models that appear to recognize visual illusions often answer fake-illusion questions from prior knowledge, not from actually seeing the images.

Pith tools