REVIEW 4 major objections 7 minor 56 references
Frontier multimodal models still fail basic visual perception: none clear 60% on a benchmark that isolates ten atomic skills.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 04:56 UTC pith:AX6YL2CL
load-bearing objection Solid failure-driven perception benchmark with real diagnostic value; the sub-60% headline is partly baked in by the difficulty filter, so read the absolute claim carefully. the 4 major comments →
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Atomic visual perception remains largely unsolved for frontier multimodal models. On PerceptionBench’s 3,000 verified, perception-isolating questions spanning ten failure-derived capabilities, no evaluated model reaches 60% overall accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent per-capability profiles.
What carries the argument
PerceptionBench: a bottom-up benchmark whose ten atomic perceptual capabilities (localization, attribute, counting, relation, depth & 3D, OCR, comparison, fine-grained recognition, context integration, and perception-related hallucination) are induced from first-error attribution of frontier-model failures on 42 source benchmarks, then populated with short, unambiguous questions verified so that difficulty stems from perception alone.
Load-bearing premise
The method assumes that a stronger analyzer model, plus human decomposition, can correctly pin each failure and each new question to a single perceptual skill as the sole source of difficulty.
What would settle it
If a careful human audit of a large random sample finds that many PerceptionBench items still require non-trivial reasoning or external knowledge to answer, or that first-error labels frequently mis-attribute reasoning failures as perception, the claim that the scores isolate atomic perception would not hold.
If this is right
- Overall leaderboard scores on mixed vision-language tasks will keep overstating visual competence until perception is measured separately.
- Hallucination of non-existent visual content is a distinct, currently weak skill, not just a side effect of weak reasoning.
- Nearly identical aggregate accuracy can hide large gaps in localization, OCR, counting, and other specific skills, so model selection needs capability profiles.
- Stable mean scores can coexist with unstable per-image answers, so reliability checks (pass@k vs always-correct) matter for perception claims.
- As models improve, the same failure-driven pipeline can be re-run to refresh both the taxonomy and the difficulty calibration.
Where Pith is reading between the lines
- Training that rewards aggressive answer commitment may raise scores on answerable items while worsening hallucination probes—suggesting a measurable trade-off worth tracking in future releases.
- Open-source models already sit within a few points of the proprietary leader on this isolation benchmark, so perception gains may be less locked behind closed data than end-to-end task scores imply.
- If first-error labels drift as models get better at early perception, item capability tags will need periodic re-attribution or the benchmark will silently change what it measures.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces PerceptionBench, a 3,000-question benchmark intended to isolate ten atomic visual-perception capabilities in MLLMs. The taxonomy is induced bottom-up from failures on 42 existing benchmarks using first-error attribution; items are then decomposed from failures or newly authored, screened for ensemble saturation and near-duplicate images, and checked by automated and human validators. Sixteen frontier models are evaluated with a unified LLM judge. The main findings are that no model exceeds 60% overall accuracy, hallucination is the weakest category on average, models with similar aggregate scores have different capability profiles, and run-level scores are stable even when per-item behavior is unstable.
Significance. A validated capability-level perception benchmark would be useful for diagnosing a major MLLM bottleneck and for tracking progress more informatively than aggregate task benchmarks. The paper has substantial strengths: public code and data, provenance and license metadata, a transparent multi-stage construction pipeline, a screening ensemble disjoint from the evaluated models, high reported judge–human agreement, released-subset representativeness checks, and repeated-run reliability analysis. The claims would be valuable if the attribution labels and difficulty conditioning are quantitatively validated.
major comments (4)
- [§2.1.2, Abstract, §3.2] The ensemble-saturation filter conditions all reported scores on current frontier-model failures. A sample is removed only when all four ensemble models solve it on every run, but the headline statement that no model reaches 60% and that “atomic perception remains largely unsolved” still invites an absolute interpretation. Please report how many items were discarded overall and by capability, evaluate representative released models on those saturated items, and show sensitivity to a weaker/stronger ensemble or threshold. The abstract and §3.2 should then state the ensemble-conditioned interpretation explicitly.
- [§2.1.1–2.2, Appendix B] The taxonomy and inherited capability labels depend on first-error attribution by an unspecified stronger analyzer model. The first-error-priority rule may classify trajectories with both perception and reasoning errors as perception, and model rationales need not faithfully expose visual processing. Since perceptual isolation is the central claim, please identify the analyzer and prompt and provide a blind human audit stratified by source benchmark and error class, including attribution accuracy/agreement and a confusion analysis showing how residual errors affect capability labels.
- [Table 1; §3.2; Figure 11] The claim that perception-related hallucination is the weakest capability (36.7% average) needs format- and chance-aware analysis. The showcased hallucination items include four-option questions whose reference answer is 0/none, and GPT-5.6-Sol’s 26.9% is close to 25% four-choice chance. Please report the mix of open-ended and multiple-choice hallucination items, chance-adjusted performance with confidence intervals, and response distributions distinguishing unsupported assertions from guessing or abstention-like behavior before drawing conclusions about answer commitment.
- [§3.4, Figure 7] The released-subset validation uses Pearson r=0.84 across only ten capability-level points, and the figure does not clearly state whether accuracies are averaged over models. This does not by itself demonstrate preservation of “relative model performance.” Please report per-model correlations or rank concordance between the released subset and full pool, uncertainty intervals, and ideally variability across multiple possible 3,000-item draws.
minor comments (7)
- [Figure 3] Clarify the normalization of the displayed cell values in Figure 3: the caption calls them shares of failures, but it is not explicit whether each benchmark column sums to one.
- [§2.3, Figure 5] Figure 5 is nearly balanced across categories, while §2.3 says the distribution follows empirical failure frequency with capability balancing. Please distinguish the empirical pre-balancing distribution from the final quotas.
- [§2.1.2 and §2.2] Specify the identity/version of the four screening-ensemble members and the automated capability/grounding verifier, and archive all prompts and random seeds needed to reproduce construction.
- [§3.1, Table 1] The 300-example judge audit should state whether it was stratified across models and capabilities and should discuss the single disagreement; with small per-category differences, confidence intervals would also be useful in Table 1.
- [§3.5, Table 2] The pass@4 versus pass^4 gap demonstrates behavioral inconsistency, but it could arise from decoding or answer-policy stochasticity rather than perception specifically. Please soften the causal wording or provide an analysis separating these factors.
- [§2.1.2] The ISC near-duplicate check is useful, but it does not rule out semantic or benchmark-specific memorization of reused source images. Please discuss this contamination limitation, especially for the 60% decomposed items.
- [Figure 2] Some displayed questions are awkward or ungrammatical, e.g. “Which one has a consistent with the middle image?” in Figure 2. Please proofread released prompts and clarify whether text shown in figures was extracted verbatim.
Circularity Check
No load-bearing circularity: PerceptionBench scores are empirical measurements on filtered items, not predictions forced by fitted inputs or self-definition.
full rationale
This is a benchmark-construction and evaluation paper, not a first-principles derivation. The central claims (no model ≥60% overall; hallucination weakest on average; divergent capability profiles) are direct accuracy measurements on 3,000 verified free-form items judged against short reference answers. The taxonomy is induced empirically from attributed failures on 42 external benchmarks, then used to author/decompose items—an ordinary diagnostic loop, not X defined as Y. Difficulty-aware selection and final screening discard only items solved by all four ensemble models (explicitly disjoint from the sixteen evaluated systems), which raises hardness but does not algebraically force any evaluated model below 60%: a stronger model could still score near ceiling on the retained pool. Authors disclose generation-conditioning and the need to re-induce the taxonomy in Limitations. No self-citation uniqueness theorem, no fitted parameter renamed as prediction, and no ansatz smuggled in via overlapping prior work underwrite the headline numbers. Residual mild conditioning of the item pool on current-frontier failure modes is a methodological caveat, not circular reduction of the reported accuracies.
Axiom & Free-Parameter Ledger
free parameters (4)
- ISC near-duplicate cosine threshold =
0.95
- Ensemble saturation rule (4 frontier MLLMs) =
drop if 4/4 correct; else tier by pass count
- Released subset size and capability quotas =
3000 total; ~255–330 per capability
- Automatic judge model and agreement sample =
GPT-oss-120B; n=300 agreement check
axioms (4)
- domain assumption Visual perception in MLLMs decomposes into discrete atomic capabilities that can be isolated by short questions whose difficulty is perceptual rather than reasoning- or knowledge-based.
- ad hoc to paper First-error-priority attribution: any misread visual fact assigns the failure to the earliest perceptual error type, even if later reasoning also fails.
- domain assumption A stronger frontier analyzer model can recover the earliest erroneous step and an open-vocabulary error label reliably enough to induce a useful taxonomy after clustering.
- standard math Standard practices of multimodal eval: open-ended short answers, license-respecting reuse of public benchmarks, and human annotation sufficiency checks.
invented entities (2)
-
Ten atomic perceptual capabilities (localization, attribute, counting, relation, depth & 3D, OCR, comparison, fine-grained recognition, context integration, perception-related hallucination)
no independent evidence
-
PerceptionBench (3,000-question released suite + >17k in-house pool)
independent evidence
read the original abstract
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.
Figures
Reference graph
Works this paper leans on
-
[1]
Nature Machine Intelligence , volume=
Visual cognition in multimodal large language models , author=. Nature Machine Intelligence , volume=. 2025 , publisher=
2025
-
[2]
Science China Information Sciences , volume=
Ocrbench: on the hidden mystery of ocr in large multimodal models , author=. Science China Information Sciences , volume=
-
[3]
arXiv preprint arXiv:2507.16863 , year=
Pixels, patterns, but no poetry: To see the world like humans , author=. arXiv preprint arXiv:2507.16863 , year=
-
[4]
Advances in Neural Information Processing Systems , volume=
Charxiv: Charting gaps in realistic chart understanding in multimodal llms , author=. Advances in Neural Information Processing Systems , volume=
-
[5]
arXiv preprint arXiv:2606.09669 , year=
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks , author=. arXiv preprint arXiv:2606.09669 , year=
-
[6]
Proceedings of the 33rd ACM International Conference on Multimedia , pages=
Screenspot-pro: Gui grounding for professional high-resolution computer use , author=. Proceedings of the 33rd ACM International Conference on Multimedia , pages=
-
[7]
Proceedings of the IEEE international conference on computer vision , pages=
Vqa: Visual question answering , author=. Proceedings of the IEEE international conference on computer vision , pages=
-
[8]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Gqa: A new dataset for real-world visual reasoning and compositional question answering , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[9]
Proceedings of the IEEE conference on computer vision and pattern recognition , pages=
Visual7w: Grounded question answering in images , author=. Proceedings of the IEEE conference on computer vision and pattern recognition , pages=
-
[10]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[11]
Advances in Neural Information Processing Systems , volume=
Are we on the right way for evaluating large vision-language models? , author=. Advances in Neural Information Processing Systems , volume=
-
[12]
European conference on computer vision , pages=
Mmbench: Is your multi-modal model an all-around player? , author=. European conference on computer vision , pages=. 2024 , organization=
2024
-
[13]
Advances in Neural Information Processing Systems , volume=
Mmlongbench-doc: Benchmarking long-context document understanding with visualizations , author=. Advances in Neural Information Processing Systems , volume=
-
[14]
Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=
Docvqa: A dataset for vqa on document images , author=. Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=
-
[15]
Findings of the association for computational linguistics: ACL 2022 , pages=
Chartqa: A benchmark for question answering about charts with visual and logical reasoning , author=. Findings of the association for computational linguistics: ACL 2022 , pages=
2022
-
[16]
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages=
Infographicvqa , author=. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages=
-
[17]
Advances in Neural Information Processing Systems , volume=
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments , author=. Advances in Neural Information Processing Systems , volume=
-
[18]
European conference on computer vision , pages=
Modeling context in referring expressions , author=. European conference on computer vision , pages=. 2016 , organization=
2016
-
[19]
International journal of computer vision , volume=
Visual genome: Connecting language and vision using crowdsourced dense image annotations , author=. International journal of computer vision , volume=. 2017 , publisher=
2017
-
[20]
The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset , author=. The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=
-
[21]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[22]
International Conference on Learning Representations , volume=
Muirbench: A comprehensive benchmark for robust multi-image understanding , author=. International Conference on Learning Representations , volume=
-
[23]
arXiv preprint arXiv:2505.15929 , year=
PhyX: Does Your Model Have the ``Wits'' for Physical Reasoning? , author=. arXiv preprint arXiv:2505.15929 , year=
-
[24]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[25]
arXiv preprint arXiv:2505.23941 , year=
Vision Language Models are Biased , author=. arXiv preprint arXiv:2505.23941 , year=
-
[26]
arXiv preprint arXiv:2501.05444 , year=
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark , author=. arXiv preprint arXiv:2501.05444 , year=
-
[27]
arXiv preprint arXiv:2407.04973 , year=
Logicvista: Multimodal llm logical reasoning benchmark in visual contexts , author=. arXiv preprint arXiv:2407.04973 , year=
-
[28]
Advances in Neural Information Processing Systems , volume=
Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning , author=. Advances in Neural Information Processing Systems , volume=
-
[29]
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Mathcanvas: Intrinsic visual chain-of-thought for multimodal mathematical reasoning , author=. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[30]
Proceedings of the Asian Conference on Computer Vision , pages=
Vision Language Models Are Blind , author=. Proceedings of the Asian Conference on Computer Vision , pages=
-
[31]
arXiv preprint arXiv:2502.16435 , year=
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs , author=. arXiv preprint arXiv:2502.16435 , year=
-
[32]
arXiv preprint arXiv:2601.06521 , year=
BabyVision: Visual Reasoning Beyond Language , author=. arXiv preprint arXiv:2601.06521 , year=
-
[33]
arXiv preprint arXiv:2505.09990 , year=
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing , author=. arXiv preprint arXiv:2505.09990 , year=
-
[34]
European Conference on Computer Vision , pages=
BLINK: Multimodal Large Language Models Can See but Not Perceive , author=. European Conference on Computer Vision , pages=. 2024 , organization=
2024
-
[35]
arXiv preprint arXiv:2502.09696 , year=
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models , author=. arXiv preprint arXiv:2502.09696 , year=
-
[36]
arXiv preprint arXiv:2510.13804 , year=
Generative Universal Verifier as Multimodal Meta-Reasoner , author=. arXiv preprint arXiv:2510.13804 , year=
-
[37]
arXiv preprint arXiv:2410.23218 , year=
OS-Atlas: A Foundation Action Model for Generalist GUI Agents , author=. arXiv preprint arXiv:2410.23218 , year=
-
[38]
arXiv preprint arXiv:2505.23764 , year=
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence , author=. arXiv preprint arXiv:2505.23764 , year=
-
[39]
arXiv preprint arXiv:2511.03146 , year=
MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity , author=. arXiv preprint arXiv:2511.03146 , year=
-
[40]
The Twelfth International Conference on Learning Representations , year=
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts , author=. The Twelfth International Conference on Learning Representations , year=
-
[41]
European Conference on Computer Vision , pages=
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? , author=. European Conference on Computer Vision , pages=. 2024 , organization=
2024
-
[42]
arXiv preprint arXiv:2505.13227 , year=
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis , author=. arXiv preprint arXiv:2505.13227 , year=
-
[43]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
MMMU -Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025
2025
-
[44]
arXiv preprint arXiv:2507.07999 , year=
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Method , author=. arXiv preprint arXiv:2507.07999 , year=
-
[45]
The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads , author=. The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=
-
[46]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
-
[47]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
Teaching CLIP to Count to Ten , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
-
[48]
Advances in Neural Information Processing Systems , volume=
Depth Anything V2 , author=. Advances in Neural Information Processing Systems , volume=
-
[49]
arXiv preprint arXiv:2506.04308 , year=
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics , author=. arXiv preprint arXiv:2506.04308 , year=
-
[50]
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models , year =
Chengke Zou and Xingang Guo and Rui Yang and Junyu Zhang and Bin Hu and Huan Zhang , booktitle =. DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models , year =
-
[51]
arXiv preprint arXiv:2503.20020 , year=
Gemini Robotics: Bringing AI into the Physical World , author=. arXiv preprint arXiv:2503.20020 , year=
-
[52]
Findings of the Association for Computational Linguistics: ACL 2025
C hart QAP ro: A More Diverse and Challenging Benchmark for Chart Question Answering. Findings of the Association for Computational Linguistics: ACL 2025. 2025
2025
-
[53]
2025 , howpublished=
Visual Physics Comprehension Test , author=. 2025 , howpublished=
2025
-
[54]
2024 , howpublished=
RealWorldQA , author=. 2024 , howpublished=
2024
-
[55]
arXiv preprint arXiv:2504.15279 , year=
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models , author=. arXiv preprint arXiv:2504.15279 , year=
-
[56]
arXiv preprint arXiv:2112.04323 , year=
Contrastive learning with large memory bank and negative embedding subtraction for accurate copy detection , author=. arXiv preprint arXiv:2112.04323 , year=
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.