REVIEW 4 major objections 4 minor 1 cited by
Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Large vision-language models that appear to recognize visual illusions actually answer from prior knowledge about illusions, not from genuine visual perception, at least for GPT-4o.
desk verdict Useful dataset and honest methodology, but the control condition leaves the prior-knowledge claim more confounded than the abstract suggests. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is the paired genuine-versus-fake illusion design. A genuine illusion image has inducers that create a gap between apparent and actual features; a fake illusion image keeps the same inducers but redraws the content so the apparent feature is also the true feature. Asking both 'which is bigger?' and 'which appears bigger?' with forced-choice options, plus control images stripped of inducers, lets the authors attribute errors to the illusion prior rather than to image difficulty. Human majority votes set the gold-standard answers, including filtering out one Ponzo image that failed to fool human observers.
What would settle it
Show GPT-4o fake-illusion images with questions reworded to avoid 'appears'/'is' wording, for example asking 'according to the measured sizes, which red circle is bigger?'; if the model then answers the actual-feature questions correctly on most items, the prior-knowledge explanation for the original failures is refuted.
Extended reading notes
Core claim
The central claim is that apparent success at illusion recognition is not perceptual competence: when a fake Ebbinghaus image genuinely has a larger left circle, GPT-4o and Claude 3.5 still say the circles are the same size, the answer that would be correct only for the genuine illusion. The same models perform well on genuine-illusion and control questions, so the failure is tied to the conflict between learned illusion priors and the actual image content. The paper proposes the fake-illusion task as a diagnostic that separates 'knows what illusions look like' from 'sees what is in the image.'
Load-bearing premise
The conclusion assumes both that the models understand the wording difference between 'appears' and 'is' and that human majority votes define the correct answers; if either fails, the fake-illusion failures could reflect a language-comprehension or benchmark artifact rather than learned illusion priors.
Editorial extensions
If this is right
- A model that answers both feature questions correctly on genuine illusions has not thereby demonstrated illusion perception; fake-illusion items are needed to expose prior-knowledge guessing.
- Benchmark results that report only genuine-illusion accuracy may overstate LVLM perceptual ability, since the models' apparent advantage over humans disappears on fake-illusion actual-feature questions.
- Prompting variations (zero-shot, one-shot, metacognitive) do not remove the prior-knowledge pattern for GPT-4o and Claude 3.5, so the failure appears robust to surface instruction changes.
- Control images show the errors are not caused solely by illusion inducers; other factors such as a bias toward visual similarity or weak recognition of abstract shapes also contribute, especially for Claude 3.5.
- Determining whether an LVLM's response is driven by stored illusion knowledge or live perception requires probing internal mechanisms, a direction the paper leaves open.
Reading between the lines
- The fake-illusion diagnostic could be applied to other LVLMs and to vision encoders alone; a model that genuinely integrates visual input should answer actual-feature questions on fake illusions correctly even without language priors.
- If the prior-knowledge account is right, one testable prediction is that models will fail more on famous, heavily documented illusions than on obscure or novel ones, since the priors would be stronger for well-known illusions.
- The same dataset could support counterfactual training probes: fine-tuning on fake illusion examples may shift models toward image-grounded answers, which would indicate the priors are corrigible associations rather than architectural constraints.
- For human-AI interaction, the result suggests that a model saying a circle 'appears larger' may be describing its learned idea of the Ebbinghaus illusion rather than what is in front of it, so apparent agreement with human perception should not be read as shared perception.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces FILM, a visual question answering dataset of abstract optical illusions divided into genuine illusions (actual and apparent features differ), fake illusions (features coincide despite inducing configurations), and control images from which inducers were removed. Four LVLMs (GPT-4o, Claude 3.5, LLaVA-NeXT 72b/110b) and human participants were tested with forced-choice questions about actual and apparent features under three prompting schemes. The central claim is that apparent high performance on genuine-illusion VQA is not evidence of visual illusion understanding: on fake-illusion images, GPT-4o and Claude 3.5 often answer apparent-feature questions correctly but actual-feature questions incorrectly, suggesting reliance on prior knowledge of known illusions rather than real-time perception. A control experiment is presented as partial support, mainly for GPT-4o, while the paper concedes that factors other than illusion inducers contribute to Claude's errors.
Significance. If the central claim is borne out, the paper offers a valuable evaluation framework: a released dataset, parallel human data, forced-choice scoring, and explicit separation of actual and apparent features. This is a genuine improvement over earlier illusion-VQA work that used non-abstract images or conflated the two feature types. The use of temperature 0, multiple prompting strategies, and the inclusion of fake illusions as a manipulation of the prior-knowledge hypothesis are all strengths. However, the load-bearing inference about prior knowledge is not as clean as the conclusion suggests: the paper's own control condition shows that a substantial share of the failure pattern persists even when illusion inducers are removed, and the sample sizes are too small for the strength of the claims. The dataset and task design are useful even if the interpretation needs to be narrowed.
major comments (4)
- [Experiment 3 / Table 5] The control condition weakens the attribution of fake-illusion errors to illusion-specific prior knowledge. On control images corresponding to fake illusions, from which the inducers were removed, GPT-4o still gives 'Only Apparent' responses on 38.1% of the considered images and reaches 'Both Correct' on only 52.4%, even though in these images the actual and apparent correct answers coincide and no inducing elements are present. A large share of the failure pattern therefore occurs without any illusion inducer, so it cannot be attributed solely to a learned prior about the specific illusion. The paper's explanation that 'some control images may inadvertently function like fake illusions for the models' is post hoc and untested. A per-image paired analysis showing that the same images are answered correctly after inducer removal is needed before the prior-knowledge interpretation can be regarded as supported for GPT-4o.
- [Task / Figure 3] The inference from fake-illusion failures to prior-knowledge reliance depends on the assumption that the models robustly understand the semantic distinction between 'appears' and 'is'. The paper correctly acknowledges in the Task section that it does not provide independent evidence for this, but this is not a peripheral caveat: if models instead have a systematic bias about how 'appears' questions should be answered, their errors on fake illusions could reflect a language-comprehension gap rather than learned illusion priors. The authors should add a comprehension control, for example by asking both question forms on neutral non-illusion images where the correct answers coincide, and showing that the models answer both forms correctly.
- [Experimental Setup / Tables 2, 3, and 5] The sample sizes are too small to support the strength of the conclusions. Each condition uses at most 28 images, and the fake-illusion analysis is restricted to the subset of images on which the model was 'Both Correct' on the corresponding genuine illusion, leaving only 21 cases for GPT-4o and 17 for Claude 3.5 in the zero-shot condition. Headline differences such as 14.3% vs. 52.4% in Table 5 are differences of a few responses. The paper reports no confidence intervals, exact tests, or other uncertainty quantification. Temperature 0 removes sampling noise from the language model but not stimulus sampling variability, so the authors should report binomial confidence intervals or permutation tests over images.
- [Human Evaluation] The human benchmark is partly constructed by the dataset selection procedure. Fake and control images were retained only when the human majority matched the intended features, and the paper reports 100.0% human consistency on those images; the human rows in Tables 3 and 5 are therefore near-ceiling by design. This does not invalidate the LVLM comparisons, but it does limit the force of the claim that models differ from humans on fake illusions. The authors should report how many candidate images were excluded at the human-evaluation stage for reasons other than the Ponzo images, or otherwise temper the human-performance comparison.
minor comments (4)
- [Throughout] The model name 'LLaV A-NeXT' appears with an inconsistent spacing; it should be 'LLaVA-NeXT'.
- [Table 5] The column headers in Table 5 repeat 'Control' in a confusing way. The first column should be labeled 'Fake Illusion VQA' and the subsequent columns 'Control VQA' for each model, matching Table 4.
- [Abstract / Conclusion] The phrase 'they predict the same answers for both Genuine Illusion and Fake Illusion VQA questions' is imprecise: the correct answers for actual-feature questions differ between the two tasks. What the authors appear to mean is that the models apply the same answer pattern. This should be rephrased for clarity.
- [Human Evaluation] The sentence reporting that 'only 28.9% of participants perceived an apparent difference' in the Ponzo image would be clearer with the number of participants (52) and the exact question format repeated, since this exclusion determines the final dataset composition.
Circularity Check
No circularity found: the central claim is a hedged behavioral inference from independent dataset construction and model evaluations, not a derivation from its own inputs.
full rationale
The paper does not derive its conclusion by fitting parameters and then predicting them, nor does it rest on self-citations. The genuine, fake, and control VQA tasks are constructed from external illusion stimuli and human majority judgments, and the model outputs are evaluated against those judgments. The main inference—that GPT-4o may rely on prior illusion knowledge—is an interpretation of the fake-illusion failure pattern, explicitly hedged as suggestive rather than proven. The paper acknowledges the two main threats to this inference: it provides no independent evidence that LVLMs master the 'appears' versus 'is' distinction, and its own control condition shows GPT-4o still makes Only Apparent errors on 38.1% of inducer-free fake-illusion control images, which the paper admits may occur because 'some control images may inadvertently function like fake illusions.' These limitations weaken the conclusion but do not make it circular, because the conclusion is not equivalent by construction to the dataset definition or to a fitted input. Human judgments are used both to select illusions and as the evaluation gold standard, which is appropriate for a study of human-like perception and does not force the model-side result. Score 0.
Assumptions & free parameters
assumptions (4)
- domain assumption LVLMs correctly understand the semantic distinction between 'appears' and 'is' in questions.
- domain assumption Human majority responses constitute the correct ground truth for actual and apparent features.
- domain assumption The 15 selected illusions and their generated images are representative of the illusion phenomena studied.
- domain assumption Abstract images remove the interpretation ambiguity present in non-abstract illusion images.
Cite this review
Pith. "Pith review of Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?." pith.science (2026). https://pith.science/paper/UIR5A74O
@misc{pith2026250605765,
author = {Pith},
title = {Pith review of: Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?},
year = {2026},
howpublished = {\url{https://pith.science/paper/UIR5A74O}},
note = {Machine review of arXiv:2506.05765}
}
read the original abstract
Humans are susceptible to optical illusions, which serve as valuable tools for investigating sensory and cognitive processes. Inspired by human vision studies, research has begun exploring whether machines, such as large vision language models (LVLMs), exhibit similar susceptibilities to visual illusions. However, studies often have used non-abstract images and have not distinguished actual and apparent features, leading to ambiguous assessments of machine cognition. To address these limitations, we introduce a visual question answering (VQA) dataset, categorized into genuine and fake illusions, along with corresponding control images. Genuine illusions present discrepancies between actual and apparent features, whereas fake illusions have the same actual and apparent features even though they look illusory due to the similar geometric configuration. We evaluate the performance of LVLMs for genuine and fake illusion VQA tasks and investigate whether the models discern actual and apparent features. Our findings indicate that although LVLMs may appear to recognize illusions by correctly answering questions about both feature types, they predict the same answers for both Genuine Illusion and Fake Illusion VQA questions. This suggests that their responses might be based on prior knowledge of illusions rather than genuine visual understanding. The dataset is available at https://github.com/ynklab/FILM
Figures
Forward citations
Cited by 1 Pith paper
-
Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions
Adaptive multi-view template retrieval lifts hidden-hate detection on HatefulIllusion to 93.2% balanced accuracy, far above original-view filters and moderators.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline " cite write " FUNCTION editor.postfix editor num.names #1 > "( )" "( )" if FUNCTION editor.trans.postfix editor num.names #1 > "( )" "( )" if FUNCTION trans.postfix translator num.names #1 > "( )" "( )" if FUNCTION authors.editors.reflist.apa5 'field := 'dot := field num.names 'numnames := numnames 'format.num.names := format.num.names na...
-
[2]
afifi2019else APACrefauthors Afifi, M. \ Brown, M S. APACrefauthors \ 2019 . What else can fool deep learning? A ddressing color constancy errors on deep neural network performance What else can fool deep learning? A ddressing color constancy errors on deep neural network performance . Proceedings of the IEEE/CVF International Conference on Computer Visio...
work page 2019
-
[3]
benjamin2019shared APACrefauthors Benjamin, A. , Qiu, C. , Zhang, L Q. , Kording, K. \ Stocker, A. APACrefauthors \ 2019 . Shared visual illusions between humans and artificial neural networks Shared visual illusions between humans and artificial neural networks . Proceedings of 2019 Conference on Cognitive Computational Neuroscience Proceedings of 2019 c...
work page 2019
-
[4]
day1984nature APACrefauthors Day, R. APACrefauthors \ 1984 . The Nature of Perceptual Illusions The nature of perceptual illusions . Interdisciplinary Science Reviews 9 1 47--58 . APACrefDOI doi:https://doi.org/10.1179/isr.1984.9.1.47 APACrefDOI
-
[5]
gomez2019convolutional APACrefauthors Gomez-Villa, A. , Mart\' i n, A. , Vazquez-Corral, J. \ Bertalm\' i o, M. APACrefauthors \ 2019 . Convolutional neural networks can be deceived by visual illusions Convolutional neural networks can be deceived by visual illusions . Proceedings of the IEEE/CVF conference on computer vision and pattern recognition Proce...
work page 2019
-
[6]
Guan_2024_CVPR APACrefauthors Guan, T. , Liu, F. , Wu, X. , Xian, R. , Li, Z. , Liu, X. Zhou, T. APACrefauthors \ 2024 . Hallusion B ench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models Hallusion B ench: An advanced diagnostic suite for entangled language hallucination and visual illus...
work page 2024
-
[7]
li2024llavanext-72-110 APACrefauthors Li, B. , Zhang, K. , Zhang, H. , Guo, D. , Zhang, R. , Li, F. Li, C. APACrefauthors \ 2024 May . LLaVA-NeXT : Stronger LLMs Supercharge Multimodal Capabilities in the Wild. LLaVA-NeXT : Stronger llms supercharge multimodal capabilities in the wild. APACrefURL https://llava-vl.github.io/blog/2024-05-10-llava-next-stron...
work page 2024
-
[8]
shahgir2024illusionvqa APACrefauthors Shahgir, H S. , Sayeed, K S. , Bhattacharjee, A. , Ahmad, W U. , Dong, Y. \ Shahriyar, R. APACrefauthors \ 2024 . Illusion VQA : A Challenging Optical Illusion Dataset for Vision Language Models (version 3) Illusion VQA : A challenging optical illusion dataset for vision language models (version 3) . Computing Researc...
arXiv 2024
Show all 12 references
-
[9]
\ Dekel, R
sun2021imagenet APACrefauthors Sun, E D. \ Dekel, R. APACrefauthors \ 2021 . ImageNet-trained deep neural networks exhibit illusion-like response to the Scintillating grid Imagenet-trained deep neural networks exhibit illusion-like response to the scintillating grid . Journal ...
2021
-
[10]
APACrefauthors \ 2024
ullman2024illill APACrefauthors Ullman, T. APACrefauthors \ 2024 . The Illusion-Illusion: Vision Language Models See Illusions Where There are None. The illusion-illusion: Vision language models see illusions where there are none. APACrefURL https://arxiv.org/abs/2412.18613 APACrefURL
2024 arXiv
-
[11]
\ Zhao, Y
wang2024metacog APACrefauthors Wang, Y. \ Zhao, Y. APACrefauthors \ 2024 . Metacognitive Prompting Improves Understanding in Large Language Models. Metacognitive prompting improves understanding in large language models. APACrefURL https://arxiv.org/abs/2308.05342 APACrefURL
2024 arXiv
-
[12]
, Pan, J
zhang2023grounding APACrefauthors Zhang, Y. , Pan, J. , Zhou, Y. , Pan, R. \ Chai, J. APACrefauthors \ 2023 . Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans? Grounding visual illusions in language: Do vision-language models per...
2023 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.