Pith. sign in

REVIEW 4 major objections 7 minor 56 references

Frontier multimodal models still fail basic visual perception: none clear 60% on a benchmark that isolates ten atomic skills.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 04:56 UTC pith:AX6YL2CL

load-bearing objection Solid failure-driven perception benchmark with real diagnostic value; the sub-60% headline is partly baked in by the difficulty filter, so read the absolute claim carefully. the 4 major comments →

arxiv 2607.24957 v1 pith:AX6YL2CL submitted 2026-07-27 cs.CV

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

classification cs.CV
keywords PerceptionBenchatomic visual perceptionmultimodal large language modelsfailure attributionhallucinationcapability-level evaluationvisual groundingOCR
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that current multimodal models are still weak at the simplest layer of seeing—localizing, counting, reading text, comparing attributes, and avoiding made-up details—because most benchmarks mix perception with reasoning and knowledge. The authors recover ten atomic perceptual capabilities by tracing where frontier models first go wrong on 42 existing benchmarks, then build PerceptionBench: 3,000 short, verified questions each aimed at one capability, with difficulty coming from the image rather than multi-step thought. Across sixteen frontier models, no system reaches 60% overall accuracy; perception-related hallucination is the weakest skill on average; and models with nearly the same total score show very different strength profiles. Aggregate scores are stable across runs, but many individual answers flip, so part of measured accuracy looks like lucky guessing rather than reliable seeing. The result is a capability-level yardstick for diagnosing where visual perception still breaks.

Core claim

Atomic visual perception remains largely unsolved for frontier multimodal models. On PerceptionBench’s 3,000 verified, perception-isolating questions spanning ten failure-derived capabilities, no evaluated model reaches 60% overall accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent per-capability profiles.

What carries the argument

PerceptionBench: a bottom-up benchmark whose ten atomic perceptual capabilities (localization, attribute, counting, relation, depth & 3D, OCR, comparison, fine-grained recognition, context integration, and perception-related hallucination) are induced from first-error attribution of frontier-model failures on 42 source benchmarks, then populated with short, unambiguous questions verified so that difficulty stems from perception alone.

Load-bearing premise

The method assumes that a stronger analyzer model, plus human decomposition, can correctly pin each failure and each new question to a single perceptual skill as the sole source of difficulty.

What would settle it

If a careful human audit of a large random sample finds that many PerceptionBench items still require non-trivial reasoning or external knowledge to answer, or that first-error labels frequently mis-attribute reasoning failures as perception, the claim that the scores isolate atomic perception would not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Overall leaderboard scores on mixed vision-language tasks will keep overstating visual competence until perception is measured separately.
  • Hallucination of non-existent visual content is a distinct, currently weak skill, not just a side effect of weak reasoning.
  • Nearly identical aggregate accuracy can hide large gaps in localization, OCR, counting, and other specific skills, so model selection needs capability profiles.
  • Stable mean scores can coexist with unstable per-image answers, so reliability checks (pass@k vs always-correct) matter for perception claims.
  • As models improve, the same failure-driven pipeline can be re-run to refresh both the taxonomy and the difficulty calibration.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Training that rewards aggressive answer commitment may raise scores on answerable items while worsening hallucination probes—suggesting a measurable trade-off worth tracking in future releases.
  • Open-source models already sit within a few points of the proprietary leader on this isolation benchmark, so perception gains may be less locked behind closed data than end-to-end task scores imply.
  • If first-error labels drift as models get better at early perception, item capability tags will need periodic re-attribution or the benchmark will silently change what it measures.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The manuscript introduces PerceptionBench, a 3,000-question benchmark intended to isolate ten atomic visual-perception capabilities in MLLMs. The taxonomy is induced bottom-up from failures on 42 existing benchmarks using first-error attribution; items are then decomposed from failures or newly authored, screened for ensemble saturation and near-duplicate images, and checked by automated and human validators. Sixteen frontier models are evaluated with a unified LLM judge. The main findings are that no model exceeds 60% overall accuracy, hallucination is the weakest category on average, models with similar aggregate scores have different capability profiles, and run-level scores are stable even when per-item behavior is unstable.

Significance. A validated capability-level perception benchmark would be useful for diagnosing a major MLLM bottleneck and for tracking progress more informatively than aggregate task benchmarks. The paper has substantial strengths: public code and data, provenance and license metadata, a transparent multi-stage construction pipeline, a screening ensemble disjoint from the evaluated models, high reported judge–human agreement, released-subset representativeness checks, and repeated-run reliability analysis. The claims would be valuable if the attribution labels and difficulty conditioning are quantitatively validated.

major comments (4)
  1. [§2.1.2, Abstract, §3.2] The ensemble-saturation filter conditions all reported scores on current frontier-model failures. A sample is removed only when all four ensemble models solve it on every run, but the headline statement that no model reaches 60% and that “atomic perception remains largely unsolved” still invites an absolute interpretation. Please report how many items were discarded overall and by capability, evaluate representative released models on those saturated items, and show sensitivity to a weaker/stronger ensemble or threshold. The abstract and §3.2 should then state the ensemble-conditioned interpretation explicitly.
  2. [§2.1.1–2.2, Appendix B] The taxonomy and inherited capability labels depend on first-error attribution by an unspecified stronger analyzer model. The first-error-priority rule may classify trajectories with both perception and reasoning errors as perception, and model rationales need not faithfully expose visual processing. Since perceptual isolation is the central claim, please identify the analyzer and prompt and provide a blind human audit stratified by source benchmark and error class, including attribution accuracy/agreement and a confusion analysis showing how residual errors affect capability labels.
  3. [Table 1; §3.2; Figure 11] The claim that perception-related hallucination is the weakest capability (36.7% average) needs format- and chance-aware analysis. The showcased hallucination items include four-option questions whose reference answer is 0/none, and GPT-5.6-Sol’s 26.9% is close to 25% four-choice chance. Please report the mix of open-ended and multiple-choice hallucination items, chance-adjusted performance with confidence intervals, and response distributions distinguishing unsupported assertions from guessing or abstention-like behavior before drawing conclusions about answer commitment.
  4. [§3.4, Figure 7] The released-subset validation uses Pearson r=0.84 across only ten capability-level points, and the figure does not clearly state whether accuracies are averaged over models. This does not by itself demonstrate preservation of “relative model performance.” Please report per-model correlations or rank concordance between the released subset and full pool, uncertainty intervals, and ideally variability across multiple possible 3,000-item draws.
minor comments (7)
  1. [Figure 3] Clarify the normalization of the displayed cell values in Figure 3: the caption calls them shares of failures, but it is not explicit whether each benchmark column sums to one.
  2. [§2.3, Figure 5] Figure 5 is nearly balanced across categories, while §2.3 says the distribution follows empirical failure frequency with capability balancing. Please distinguish the empirical pre-balancing distribution from the final quotas.
  3. [§2.1.2 and §2.2] Specify the identity/version of the four screening-ensemble members and the automated capability/grounding verifier, and archive all prompts and random seeds needed to reproduce construction.
  4. [§3.1, Table 1] The 300-example judge audit should state whether it was stratified across models and capabilities and should discuss the single disagreement; with small per-category differences, confidence intervals would also be useful in Table 1.
  5. [§3.5, Table 2] The pass@4 versus pass^4 gap demonstrates behavioral inconsistency, but it could arise from decoding or answer-policy stochasticity rather than perception specifically. Please soften the causal wording or provide an analysis separating these factors.
  6. [§2.1.2] The ISC near-duplicate check is useful, but it does not rule out semantic or benchmark-specific memorization of reused source images. Please discuss this contamination limitation, especially for the 60% decomposed items.
  7. [Figure 2] Some displayed questions are awkward or ungrammatical, e.g. “Which one has a consistent with the middle image?” in Figure 2. Please proofread released prompts and clarify whether text shown in figures was extracted verbatim.

Circularity Check

0 steps flagged

No load-bearing circularity: PerceptionBench scores are empirical measurements on filtered items, not predictions forced by fitted inputs or self-definition.

full rationale

This is a benchmark-construction and evaluation paper, not a first-principles derivation. The central claims (no model ≥60% overall; hallucination weakest on average; divergent capability profiles) are direct accuracy measurements on 3,000 verified free-form items judged against short reference answers. The taxonomy is induced empirically from attributed failures on 42 external benchmarks, then used to author/decompose items—an ordinary diagnostic loop, not X defined as Y. Difficulty-aware selection and final screening discard only items solved by all four ensemble models (explicitly disjoint from the sixteen evaluated systems), which raises hardness but does not algebraically force any evaluated model below 60%: a stronger model could still score near ceiling on the retained pool. Authors disclose generation-conditioning and the need to re-induce the taxonomy in Limitations. No self-citation uniqueness theorem, no fitted parameter renamed as prediction, and no ansatz smuggled in via overlapping prior work underwrite the headline numbers. Residual mild conditioning of the item pool on current-frontier failure modes is a methodological caveat, not circular reduction of the reported accuracies.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

The work is empirical benchmark science, not a formal derivation. Load-bearing commitments are methodological: that early trajectory errors can be labeled into a stable perception taxonomy; that short questions can isolate one skill; that ensemble consensus defines non-saturated difficulty; and that a judge model can score free-form answers against short references. Free parameters are thresholds and sampling choices in construction, not physical constants. The ten capabilities are operational categories induced from clusters, not new physical entities.

free parameters (4)
  • ISC near-duplicate cosine threshold = 0.95
    Images with similarity >0.95 to large pretraining corpora are discarded; the cutoff is a design choice that affects contamination filtering and retained visual diversity.
  • Ensemble saturation rule (4 frontier MLLMs) = drop if 4/4 correct; else tier by pass count
    Items solved by every ensemble member across runs are dropped as trivial; remaining items are stratified by how many models pass. This hand-designed filter sets the difficulty distribution of the benchmark.
  • Released subset size and capability quotas = 3000 total; ~255–330 per capability
    3,000 items subsampled from >17,000 constructed/verified samples with per-capability balancing (e.g., several capabilities fixed at 330). Quotas are designer choices constrained by empirical failure frequencies.
  • Automatic judge model and agreement sample = GPT-oss-120B; n=300 agreement check
    GPT-oss-120B grades free-form answers; human agreement checked on 300 predictions (99.7%). Judge choice and sample size are operational parameters of the reported accuracies.
axioms (4)
  • domain assumption Visual perception in MLLMs decomposes into discrete atomic capabilities that can be isolated by short questions whose difficulty is perceptual rather than reasoning- or knowledge-based.
    Stated throughout Sections 1–2 and enforced via capability-alignment and visual-grounding verification; if isolation fails, scores are not pure perception measures.
  • ad hoc to paper First-error-priority attribution: any misread visual fact assigns the failure to the earliest perceptual error type, even if later reasoning also fails.
    Section 2.1.1 and Appendix B; defines how the perception branch of the taxonomy is populated and how labels transfer to items.
  • domain assumption A stronger frontier analyzer model can recover the earliest erroneous step and an open-vocabulary error label reliably enough to induce a useful taxonomy after clustering.
    Core of failure-driven discovery (Section 2.1.1); Limitations explicitly flags residual attribution error.
  • standard math Standard practices of multimodal eval: open-ended short answers, license-respecting reuse of public benchmarks, and human annotation sufficiency checks.
    Background methodology assumed rather than proved; underpins construction legality and answer uniqueness claims.
invented entities (2)
  • Ten atomic perceptual capabilities (localization, attribute, counting, relation, depth & 3D, OCR, comparison, fine-grained recognition, context integration, perception-related hallucination) no independent evidence
    purpose: Serve as the evaluation axes and label set for PerceptionBench items and leaderboard columns.
    Induced by clustering attributed perception errors rather than taken from a prior standard ontology; operational definitions and disambiguation rules appear in Table 3 / Appendix B.
  • PerceptionBench (3,000-question released suite + >17k in-house pool) independent evidence
    purpose: Provide a capability-balanced, perception-isolating measurement instrument for MLLMs.
    New dataset artifact constructed via decomposition and new authoring guided by the induced taxonomy.

pith-pipeline@v1.2.0-grok45-kimik3 · 24301 in / 3605 out tokens · 77770 ms · 2026-07-31T04:56:18.199432+00:00 · methodology

0 comments
read the original abstract

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.

Figures

Figures reproduced from arXiv: 2607.24957 by Bowen Qu, Chenzhuang Du, Haiming Wang, Haoning Wu, Haotian Yao, Hao Yang, Haoyu Lu, Hongcheng Gao, Jia Chen, Jia Li, Jinguo Zhu, Junwei Yang, Lin Sui, Mengfan Dong, Peizhou Cao, Tongtian Yue, Weihong Li, Xiaoxue Wu, Xinxing Zu, Xinyu Zhou, Yalin Wang, Yangyang Liu, Yao Wang, Y. Charles, Yifeng Xie, Yiping Bao, Yuhao Dong, Zaida Zhou, Zhangyang Qi, Zhiqi Huang, Zichao Lin, Zijia Zhao, Zuhao Yang.

Figure 1
Figure 1. Figure 1: Overall model accuracy on PerceptionBench. While GPT-5.6-Sol (59.7%) and Kimi K3 (58.5%) currently lead the field, no evaluated model surpasses the 60% accuracy threshold, underscoring the significant challenge of atomic visual perception. Accurate visual perception serves as a foundational prerequisite for Multimodal Large Language Models (MLLMs) to perform downstream reasoning and environmental interacti… view at source ↗
Figure 2
Figure 2. Figure 2: PerceptionBench vs. existing benchmarks. PerceptionBench decomposes perception into fine-grained atomic abilities, enabling more detailed and comprehensive evaluation compared to existing benchmarks. Despite this foundational role, current evaluations rarely both isolate visual perception from the capabilities built on top of it and systematically cover the range of perceptual capabilities that models lack… view at source ↗
Figure 3
Figure 3. Figure 3: Failure-type coverage of existing benchmarks. Distribution of attributed failures across error types for each of the 42 aggregated open-source benchmarks (rows: error types, sorted by global frequency; columns: benchmarks, ordered by their dominant error type). Each benchmark’s failures concentrate on one or a few error types, so existing suites probe narrow, complementary slices of the failure space, whil… view at source ↗
Figure 4
Figure 4. Figure 4: Construction pipeline of PerceptionBench. (1) Each model failure on the 42 source benchmarks is attributed to the earliest erroneous step of its reasoning trajectory (left); (2) cluster the attributed error labels into a unified taxonomy whose perception branch defines the ten atomic perceptual capabilities (top right); (3) manual annotation produces new perception questions targeting the ten atomic capabi… view at source ↗
Figure 5
Figure 5. Figure 5: Distribution of PerceptionBench across the ten atomic perceptual capabilities. [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Qualitative results on PerceptionBench. Four source-benchmark items whose original questions require multi-step solutions, decomposed into atomic, perception-only questions (GT: ground truth; colored icons: answers of Kimi K3, GPT-5.6-Sol, Claude-Fable-5, and Gemini-3.1-Pro). Sources: ERQA (Gemini Robotics Team et al. 2025) (top left), MathVision (K. Wang et al. 2024) (top right), MMStar (Lin Chen et al. 2… view at source ↗
Figure 7
Figure 7. Figure 7: Per-capability representativeness of the re￾leased PerceptionBench. Each point represents one per￾ceptual capability. The x- and y-axes show accuracy on the released and full benchmarks, respectively, with point size proportional to the number of samples. Since MLLMs may exhibit stochastic behaviors during inference, we evaluate several frontier models using four independent runs. For each model, we report… view at source ↗
Figure 8
Figure 8. Figure 8: reports the pairwise weighted Jaccard overlap between the error-type distributions of the 42 source benchmarks (benchmarks are ordered as in [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Representative examples from the ten atomic perceptual capabilities in PerceptionBench (1 of 3). [PITH_FULL_IMAGE:figures/full_fig_p017_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Representative examples from the ten atomic perceptual capabilities in PerceptionBench (2 of 3). [PITH_FULL_IMAGE:figures/full_fig_p018_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Representative examples from the ten atomic perceptual capabilities in PerceptionBench (3 of 3). [PITH_FULL_IMAGE:figures/full_fig_p019_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

56 extracted references · 15 linked inside Pith

  1. [1]

    Nature Machine Intelligence , volume=

    Visual cognition in multimodal large language models , author=. Nature Machine Intelligence , volume=. 2025 , publisher=

  2. [2]

    Science China Information Sciences , volume=

    Ocrbench: on the hidden mystery of ocr in large multimodal models , author=. Science China Information Sciences , volume=

  3. [3]

    arXiv preprint arXiv:2507.16863 , year=

    Pixels, patterns, but no poetry: To see the world like humans , author=. arXiv preprint arXiv:2507.16863 , year=

  4. [4]

    Advances in Neural Information Processing Systems , volume=

    Charxiv: Charting gaps in realistic chart understanding in multimodal llms , author=. Advances in Neural Information Processing Systems , volume=

  5. [5]

    arXiv preprint arXiv:2606.09669 , year=

    SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks , author=. arXiv preprint arXiv:2606.09669 , year=

  6. [6]

    Proceedings of the 33rd ACM International Conference on Multimedia , pages=

    Screenspot-pro: Gui grounding for professional high-resolution computer use , author=. Proceedings of the 33rd ACM International Conference on Multimedia , pages=

  7. [7]

    Proceedings of the IEEE international conference on computer vision , pages=

    Vqa: Visual question answering , author=. Proceedings of the IEEE international conference on computer vision , pages=

  8. [8]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

    Gqa: A new dataset for real-world visual reasoning and compositional question answering , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

  9. [9]

    Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

    Visual7w: Grounded question answering in images , author=. Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

  10. [10]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

    Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

  11. [11]

    Advances in Neural Information Processing Systems , volume=

    Are we on the right way for evaluating large vision-language models? , author=. Advances in Neural Information Processing Systems , volume=

  12. [12]

    European conference on computer vision , pages=

    Mmbench: Is your multi-modal model an all-around player? , author=. European conference on computer vision , pages=. 2024 , organization=

  13. [13]

    Advances in Neural Information Processing Systems , volume=

    Mmlongbench-doc: Benchmarking long-context document understanding with visualizations , author=. Advances in Neural Information Processing Systems , volume=

  14. [14]

    Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=

    Docvqa: A dataset for vqa on document images , author=. Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=

  15. [15]

    Findings of the association for computational linguistics: ACL 2022 , pages=

    Chartqa: A benchmark for question answering about charts with visual and logical reasoning , author=. Findings of the association for computational linguistics: ACL 2022 , pages=

  16. [16]

    Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages=

    Infographicvqa , author=. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages=

  17. [17]

    Advances in Neural Information Processing Systems , volume=

    Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments , author=. Advances in Neural Information Processing Systems , volume=

  18. [18]

    European conference on computer vision , pages=

    Modeling context in referring expressions , author=. European conference on computer vision , pages=. 2016 , organization=

  19. [19]

    International journal of computer vision , volume=

    Visual genome: Connecting language and vision using crowdsourced dense image annotations , author=. International journal of computer vision , volume=. 2017 , publisher=

  20. [20]

    The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

    Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset , author=. The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

  21. [21]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  22. [22]

    International Conference on Learning Representations , volume=

    Muirbench: A comprehensive benchmark for robust multi-image understanding , author=. International Conference on Learning Representations , volume=

  23. [23]

    arXiv preprint arXiv:2505.15929 , year=

    PhyX: Does Your Model Have the ``Wits'' for Physical Reasoning? , author=. arXiv preprint arXiv:2505.15929 , year=

  24. [24]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  25. [25]

    arXiv preprint arXiv:2505.23941 , year=

    Vision Language Models are Biased , author=. arXiv preprint arXiv:2505.23941 , year=

  26. [26]

    arXiv preprint arXiv:2501.05444 , year=

    Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark , author=. arXiv preprint arXiv:2501.05444 , year=

  27. [27]

    arXiv preprint arXiv:2407.04973 , year=

    Logicvista: Multimodal llm logical reasoning benchmark in visual contexts , author=. arXiv preprint arXiv:2407.04973 , year=

  28. [28]

    Advances in Neural Information Processing Systems , volume=

    Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning , author=. Advances in Neural Information Processing Systems , volume=

  29. [29]

    Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Mathcanvas: Intrinsic visual chain-of-thought for multimodal mathematical reasoning , author=. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  30. [30]

    Proceedings of the Asian Conference on Computer Vision , pages=

    Vision Language Models Are Blind , author=. Proceedings of the Asian Conference on Computer Vision , pages=

  31. [31]

    arXiv preprint arXiv:2502.16435 , year=

    Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs , author=. arXiv preprint arXiv:2502.16435 , year=

  32. [32]

    arXiv preprint arXiv:2601.06521 , year=

    BabyVision: Visual Reasoning Beyond Language , author=. arXiv preprint arXiv:2601.06521 , year=

  33. [33]

    arXiv preprint arXiv:2505.09990 , year=

    PointArena: Probing Multimodal Grounding Through Language-Guided Pointing , author=. arXiv preprint arXiv:2505.09990 , year=

  34. [34]

    European Conference on Computer Vision , pages=

    BLINK: Multimodal Large Language Models Can See but Not Perceive , author=. European Conference on Computer Vision , pages=. 2024 , organization=

  35. [35]

    arXiv preprint arXiv:2502.09696 , year=

    ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models , author=. arXiv preprint arXiv:2502.09696 , year=

  36. [36]

    arXiv preprint arXiv:2510.13804 , year=

    Generative Universal Verifier as Multimodal Meta-Reasoner , author=. arXiv preprint arXiv:2510.13804 , year=

  37. [37]

    arXiv preprint arXiv:2410.23218 , year=

    OS-Atlas: A Foundation Action Model for Generalist GUI Agents , author=. arXiv preprint arXiv:2410.23218 , year=

  38. [38]

    arXiv preprint arXiv:2505.23764 , year=

    MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence , author=. arXiv preprint arXiv:2505.23764 , year=

  39. [39]

    arXiv preprint arXiv:2511.03146 , year=

    MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity , author=. arXiv preprint arXiv:2511.03146 , year=

  40. [40]

    The Twelfth International Conference on Learning Representations , year=

    MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts , author=. The Twelfth International Conference on Learning Representations , year=

  41. [41]

    European Conference on Computer Vision , pages=

    MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? , author=. European Conference on Computer Vision , pages=. 2024 , organization=

  42. [42]

    arXiv preprint arXiv:2505.13227 , year=

    Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis , author=. arXiv preprint arXiv:2505.13227 , year=

  43. [43]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    MMMU -Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025

  44. [44]

    arXiv preprint arXiv:2507.07999 , year=

    Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Method , author=. arXiv preprint arXiv:2507.07999 , year=

  45. [45]

    The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

    Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads , author=. The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

  46. [46]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

    SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

  47. [47]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

    Teaching CLIP to Count to Ten , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

  48. [48]

    Advances in Neural Information Processing Systems , volume=

    Depth Anything V2 , author=. Advances in Neural Information Processing Systems , volume=

  49. [49]

    arXiv preprint arXiv:2506.04308 , year=

    RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics , author=. arXiv preprint arXiv:2506.04308 , year=

  50. [50]

    DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models , year =

    Chengke Zou and Xingang Guo and Rui Yang and Junyu Zhang and Bin Hu and Huan Zhang , booktitle =. DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models , year =

  51. [51]

    arXiv preprint arXiv:2503.20020 , year=

    Gemini Robotics: Bringing AI into the Physical World , author=. arXiv preprint arXiv:2503.20020 , year=

  52. [52]

    Findings of the Association for Computational Linguistics: ACL 2025

    C hart QAP ro: A More Diverse and Challenging Benchmark for Chart Question Answering. Findings of the Association for Computational Linguistics: ACL 2025. 2025

  53. [53]

    2025 , howpublished=

    Visual Physics Comprehension Test , author=. 2025 , howpublished=

  54. [54]

    2024 , howpublished=

    RealWorldQA , author=. 2024 , howpublished=

  55. [55]

    arXiv preprint arXiv:2504.15279 , year=

    VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models , author=. arXiv preprint arXiv:2504.15279 , year=

  56. [56]

    arXiv preprint arXiv:2112.04323 , year=

    Contrastive learning with large memory bank and negative embedding subtraction for accurate copy detection , author=. arXiv preprint arXiv:2112.04323 , year=