Pith. sign in

REVIEW 3 major objections 5 minor 28 cited by

State-of-the-art multimodal models score 49.7% on basic visual tasks that adults complete at 94.1%, leaving them behind 6-year-olds.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 11:21 UTC pith:MZT5BHUO

load-bearing objection A genuinely useful benchmark whose central age-comparison headline is not supported by the reported measurements; the adult–model gap is real. the 3 major comments →

arxiv 2601.06521 v2 pith:MZT5BHUO submitted 2026-01-10 cs.CV cs.CL

BabyVision: Visual Reasoning Beyond Language

classification cs.CV cs.CL
keywords BabyVisionvisual reasoning benchmarkmultimodal large language modelsearly visionvisual primitivesverbalization bottleneckvisual trackingvisual generation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper's central claim is that the strongest multimodal AI models fail at basic visual tasks that young children solve easily, despite acing knowledge-heavy benchmarks. It establishes this with BabyVision, a 388-question benchmark covering fine-grained visual discrimination, visual tracking, spatial perception, and visual pattern recognition, where the best model scores 49.7% versus 94.1% for adults and lags 6-year-olds by roughly 20 points in a 20-item pilot. The paper argues the failure is not a matter of task difficulty but of a 'verbalization bottleneck': models translate images into language before reasoning, losing visual structure that cannot be verbalized. It also introduces BabyVision-Gen, a generative evaluation that lets models draw or mark answers, and reports that current generation models are still far from reliable. If the claim holds, benchmark scores on knowledge-heavy tasks cannot be taken as evidence of grounded visual understanding.

Core claim

BabyVision is a benchmark of 388 image-only questions built from child-friendly visual tests, organized into four families: fine-grained discrimination, visual tracking, spatial perception, and visual pattern recognition. The paper's measurement claim is that the strongest multimodal model tested reaches 49.7% accuracy, while adult testers average 94.1% and 6-year-olds outperform the model by about 20 points in a smaller pilot. The paper's explanatory claim is that these failures share a single cause—a 'verbalization bottleneck' in which images are compressed into language before reasoning, so fine shape, curve identity, 3D structure, and abstract spatial relations are lost. It further claim

What carries the argument

The load-bearing object is the BabyVision benchmark: 388 items across 22 subtypes in four categories selected from developmental psychology to be solvable by children aged 3–12 without language or cultural knowledge. The mechanism that explains the results is the 'verbalization bottleneck'—the paper's claim that current multimodal models compress images into linguistic tokens before reasoning, losing fine detail, curve identity through intersections, 3D structure, and abstract pattern structure. BabyVision-Gen is the companion machinery: it reformulates 280 tasks into image-annotation prompts and uses a judge model to check generated visual answers against human-annotated ground truth, testi

Load-bearing premise

The age comparison (Section 5.1) assumes the 20-item mini test given to children measures the same skill as the 388-item benchmark on which models were scored; the paper does not report model scores on the mini test or child scores on the full benchmark.

What would settle it

Run the leading models on the exact 20-item BabyVision-Mini test that children took (Section 5.1); if the best model scores at or above the 6-year-old average on identical items, the paper's central developmental-gap claim loses its support.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • High scores on knowledge-heavy multimodal benchmarks should no longer be read as evidence of visual grounding.
  • Visual tracking and 3D spatial perception deserve to be tracked as separate capabilities in model evaluation, since they fail independently of language reasoning.
  • Offering a visual output channel such as drawing, marking, or tracing can reveal reasoning that text answers miss, making generation models a relevant test bed.
  • Reinforcement learning on verifiable rewards improves most BabyVision categories but does not improve visual tracking, consistent with the verbalization-bottleneck explanation.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit: if the verbalization bottleneck is the real cause, natively multimodal models that reason in continuous visual representations should be tested head-to-head on BabyVision; the paper's analysis predicts they will narrow the gap most on tracking and 3D tasks.
  • A direct test the paper does not run: the same subjects and models completing the same items in text-answer versus draw-the-answer formats would separate perception failures from verbalization failures.
  • The four failure modes (fine detail, manifold identity, spatial imagination, pattern induction) could be turned into a reusable diagnostic profile that maps which visual primitives a given model is missing.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces BABYVISION, a 388-item benchmark of 'early vision' tasks in four categories (fine-grained discrimination, visual tracking, spatial perception, visual pattern recognition), together with BABYVISION-GEN, a 280-item generative variant. The authors report that 11 MLLMs achieve at most 49.7% (Gemini3-Pro-Preview) on the full benchmark, whereas 16 adult testers score 94.1%, and claim that frontier models fall below 6-year-old children. They also provide failure-mode analyses, an RLVR fine-tuning study, and an automatic evaluator for the generative setting. The benchmark and code are publicly released.

Significance. If the main model-vs-adult gap is the measured result, this is a valuable diagnostic benchmark: it is carefully taxonomized, publicly released, and the deficit is consistent across all four categories. The generative evaluation and the RLVR experiment broaden the contribution. The paper also reports standard deviations over three runs and validates the generation judge against human judgments on all 280 NanoBanana-Pro outputs, which are concrete reproducibility strengths. However, the strongest advertised claim is the human developmental comparison ('behind 6-year-olds', 'even 3-year-olds'), and that comparison is currently not supported by the data as presented. The adult-model gap remains valid, but the age-based claims require either additional calibration or substantial weakening.

major comments (3)
  1. [Abstract and Section 5.1 ('Human Testers', 'Comparison with Young Humans')] The headline claim that Gemini3-Pro-Preview 'lags behind 6-year-old humans' cross-compares model scores on the full 388-item BABYVISION with child scores on the 20-item BabyVision-Mini. No model score on BabyVision-Mini and no child score on the full benchmark are reported, and the Mini's representativeness is asserted rather than calibrated (e.g., no item-difficulty correlation or adult scores on the Mini). The abstract's 'even 3-year-olds' claim is therefore unsupported. The adult-model gap (94.1 vs. 49.7) is unaffected because both are measured on the full set, but the developmental comparison is load-bearing for the paper's central 'beyond language, child-level' framing. Please report model scores on BabyVision-Mini, child scores on the full set, or item-difficulty statistics demonstrating that the Mini is representative; otherwise restrict the abstract and introduction to the adult
  2. [Section 3.3 and Appendix A] The LLM-as-judge evaluation is reported as having '100% consistency with human evaluators' with no sample size, selection of items, or disagreement analysis. Since every model accuracy and the reported gaps are computed through this judge, a systematic judge bias could change the headline numbers. Please provide the human-judge comparison protocol, sample size, and per-subtype agreement, or report exact-match accuracy as a conservative alternative.
  3. [Section 5.1 ('Human Testers') and Figure 1] The child study is described very briefly: no administration protocol for the 3-year-old group (who reads the questions, how pointing versus verbal responses are recorded), no exclusion criteria, and no item-level results. The footnote about a single school population is a welcome caveat, but the abstract and text still make universal claims. Since the 3-year-old group is used to support the claim that models fail tasks that 'even 3-year-olds can solve effortlessly', this is not a mere reporting detail; please add the testing protocol, inclusion criteria, and per-age-group item-level scores on BabyVision-Mini.
minor comments (5)
  1. [Section 5.1 and Table 1] Grok-4 appears in Table 1 but is not listed in the model enumeration; the text says five proprietary models but Table 1 reports six close-source models. Please align the model list.
  2. [Section 3.1 and Figure 4] Figure 4 says '~50 seed images' while the text says approximately 100 seed examples. Please harmonize these numbers.
  3. [Section 4.2] The automatic evaluator is validated only on NanoBanana-Pro outputs; please state whether the 96.1% agreement is expected to generalize to other generation models, and ideally report a smaller validation on GPT-Image-1.5 and Qwen-Image-Edit.
  4. [Figure 1] The figure would benefit from explicit error bars or confidence intervals for child age groups and a clear label indicating whether the displayed model scores are on the full set or BabyVision-Mini; this directly affects interpretation of the age comparison.
  5. [General] Minor naming inconsistency: 'Nano-Banana-Pro' in the introduction versus 'NanoBanana-Pro' in Tables and Section 5.3; please standardize.

Circularity Check

0 steps flagged

No significant circularity; the paper is an empirical benchmark study with no fitted derivation that reduces to its inputs.

full rationale

BabyVision is a data-collection and evaluation paper, not a derivation. The central claims (MLLMs score 49.7% on the 388-item benchmark; adults score 94.1%) are direct measurements, not outputs of a fitted model, and no equation is used to generate a prediction from a parameter that was fit to the same data. The benchmark taxonomy is motivated by external developmental psychology sources (Johnson 2010; Braddick & Atkinson 2011; Baillarge on et al. 1985), not by the authors' own prior results, so the task design is not a self-citation chain. The automatic judge for BabyVision (Qwen3-Max) is validated at 100% agreement with human evaluators, and the BabyVision-Gen judge (Gemini-3-Flash) is validated on all 280 outputs against PhD-level annotators (96.1% agreement), so the evaluation pipeline has independent support rather than relying on the tested models' self-judgment. The most plausible circularity-adjacent concern is that the child comparison uses a 20-item BabyVision-Mini while model scores come from the full 388-item set; however, this is a measurement-validity and comparability issue, not a circularity issue: the child scores are empirical data, not a renaming or fitting of model outputs. No load-bearing step reduces to its own inputs, so the honest finding is no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The central claim rests on the benchmark's validity as a language-independent measure of early vision, the correctness of its annotations, the comparability of the Mini subset to the full set, and the absence of training-data contamination. No free parameters are fitted; no new physical or theoretical entities are postulated.

axioms (5)
  • domain assumption Tasks solvable by children aged 3–12 are valid measures of foundational, language-independent visual primitives.
    Motivates benchmark design (Sec. 2.2, Sec. 3.1); if false, low model scores may reflect task comprehension or instruction-following rather than visual ability.
  • domain assumption Human adult score (94.1%) and annotator-generated ground truth are correct and unambiguous.
    Assumed from double-blind expert review (Sec. 3.1) and 16 adult testers; no independent item-level validation is reported.
  • domain assumption LLM-as-judge semantic equivalence is valid and 100% consistent with human judgments.
    Stated in Sec. 3.3 without supporting data; if the judge is lenient, model scores could be inflated.
  • domain assumption Models were not contaminated by benchmark images during pretraining.
    Images crawled from the internet (Sec. 3.1) and no contamination analysis is provided; could affect measured scores.
  • domain assumption BabyVision-Mini is representative of the full BabyVision benchmark for age-group comparison.
    20 selected samples; selection criteria not specified (Sec. 5.1), and model scores on the same Mini set are not reported.

pith-pipeline@v1.3.0-alltime-deepseek · 20856 in / 12043 out tokens · 119671 ms · 2026-08-03T11:21:58.300009+00:00 · methodology

0 comments
read the original abstract

While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat-AI/BabyVision for reproduction.

Figures

Figures reproduced from arXiv: 2601.06521 by Baobao Chang, Fangfu Liu, Guopeng Li, Haiyang Shen, Hans Zhao, Haoning Wu, Haoyu Lu, Hongfeng He, Kaiyuan Chen, Kuan Li, Liang Chen, Ming Wu, Shuzheng Si, Tianyu Liu, Weichu Xie, Wendong Xu, Wenhao Chai, Xiaobo Hu, Xuanzhong Chen, Yang Liu, Y. Charles, Yiping Bao, Yixin Ren, Yiyan Liang, Yuan Gong, Yuantao Fan, Zefan Cai, Zhibo Yang, Zhiqi Huang, Ziqi Huang.

Figure 1
Figure 1. Figure 1: Performance on BABYVISION among MLLMs and human of different ages. ∗Equal Core Contributors. Correspondence: Liang Chen <liangchen@unipat.ai>, Kuan Li <kuanli@unipat.ai> 1 arXiv:2601.06521v1 [cs.CV] 10 Jan 2026 [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Fine-grained performance analysis on the full B [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Examples of BABYVISION and BABYVISION-GEN. While BABYVISION evaluates visual un￾derstanding through language output, BABYVISION-GEN evaluates visual reasoning through image generation. To fill this gap, we introduce BABYVISION, a benchmark designed with a scientific and rigorous data curation pipeline to probe the atomic visual skills humans develop early in life. BABYVISION aims to minimize reliance on li… view at source ↗
Figure 4
Figure 4. Figure 4: Overview of the multi-stage data collection and curation pipeline for [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Example questions and the number of examples (#) from [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Four classic vision-centric challenges for MLLMs. All examples highlight failures caused by [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: GRPO training dynamics for Qwen3-VL-8B-Thinking. Both training accuracy and held-out test [PITH_FULL_IMAGE:figures/full_fig_p015_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: BabyVision accuracy before vs. after RLVR fine-tuning. RLVR yields a +4.8 overall accuracy [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Representative examples of visual reasoning results from different image/video generation [PITH_FULL_IMAGE:figures/full_fig_p016_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. An Exam for Active Observers

    cs.CV 2026-07 conditional novelty 7.0

    On a new 17-task benchmark of active visual observation, the best frontier multimodal model solves 10.6% of items and humans solve 96.1%.

  2. Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

    cs.CL 2026-06 unverdicted novelty 7.0

    ZPPO improves distillation to small vision-language models by using binary and negative candidate prompts plus a replay buffer for hard questions, outperforming standard distillation and GRPO on a 31-benchmark suite w...

  3. LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?")

    cs.LG 2026-06 unverdicted novelty 7.0

    VLMs show partial alignment with children's performance on six cognitive tasks, with stronger models matching better at task and item levels but struggling on matrix reasoning and mental rotation.

  4. ATLAS: Agentic Test-time Learning-to-Allocate Scaling

    cs.LG 2026-06 unverdicted novelty 7.0

    ATLAS introduces an LLM-orchestrated agentic framework for dynamic test-time scaling via extensible 'explore' actions, achieving higher accuracy with fewer API calls than fixed-workflow baselines on four benchmarks.

  5. DeepLatent: Think with Images via Parallel Latent Visual Reasoning

    cs.CV 2026-05 unverdicted novelty 7.0

    DeepLatent introduces a parallel latent visual reasoning framework with learnable 2D tokens and continuous RL, trained via distillation then RL, plus a new 180K dataset, claiming SOTA benchmark results.

  6. The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space

    cs.CV 2026-05 unverdicted novelty 7.0

    MLLMs scoring 70-83% on Cartesian visual tasks drop to 31-39% on logically equivalent polar versions, exposing reliance on grid discretization shortcuts instead of topology-invariant reasoning.

  7. The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

    cs.AI 2026-07 conditional novelty 6.5

    A stable middle-layer Visual Relay Window governs grounded VLM reasoning, and TRACE improves grounding-sensitive and reasoning benchmarks by task-adaptively scheduling and anchoring that window at inference time.

  8. Beacon: Knowing When and How to Perform Agentic Visual Reasoning

    cs.CV 2026-07 conditional novelty 6.0

    Beacon improves agentic visual reasoning by teaching models when tools are necessary and how to use them for net gains, via necessity-aware rewards and hint-guided RL.

  9. PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

    cs.CV 2026-07 conditional novelty 6.0

    Frontier MLLMs remain far from mastering atomic visual perception: none reach 60% on a failure-derived, perception-only benchmark of ten capabilities.

  10. The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

    cs.AI 2026-07 conditional novelty 6.0

    Visual evidence in VLMs flows through a depth-wise Visual Relay Window; scheduling that window with the lightweight TRACE controller yields +4.33 points on grounding benchmarks and +3.05 on MathVista across four open-...

  11. RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

    cs.AI 2026-06 conditional novelty 6.0

    A rollout-conditioned contrastive distillation loss plus sparse anchored cross-entropy improves injected-knowledge accuracy in MLLMs while keeping retention close to the base model.

  12. Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    A graph-grounded Combined Road Substrate framework generates traceable QA pairs from road maps to improve small VLMs on compositional road reasoning tasks.

  13. EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

    cs.LG 2026-05 unverdicted novelty 6.0

    Current VLMs depend on tightly aligned curated data and cannot exploit the weakly-aligned egocentric video signals that dominate naturalistic infant input.

  14. Step-wise Rubric Rewards for LLM Reasoning

    cs.LG 2026-05 conditional novelty 6.0

    SRaR attributes rubric items to specific steps via an LLM judge, normalizes per-step scores across rollouts, and combines them with outcome rewards via a decoupled advantage estimator, yielding 3.57-point accuracy gai...

  15. The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space

    cs.CV 2026-05 unverdicted novelty 6.0

    Reformulating 53 visual reasoning tasks in polar coordinates causes frontier MLLMs to drop from 70-83% to 31-39% accuracy while preserving logical equivalence, revealing a Cartesian shortcut in current benchmarks.

  16. Do multimodal models imagine electric sheep?

    cs.CV 2026-05 conditional novelty 6.0

    Fine-tuning VLMs to output action sequences for puzzles causes emergent internal visual representations that improve performance when integrated into reasoning.

  17. Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

    cs.CL 2026-05 unverdicted novelty 6.0

    A plug-and-play RL method adds batch-level distributional supervision via CCC rewards to reduce regression-to-the-mean in MLLMs on imbalanced regression benchmarks.

  18. Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

    cs.CL 2026-05 unverdicted novelty 6.0

    A Group Relative Policy Optimization framework with concordance correlation coefficient rewards improves MLLM regression accuracy on long-tailed distributions, especially in medium- and few-shot regimes, without model...

  19. The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

    cs.CV 2026-04 unverdicted novelty 6.0

    Proposes the Modality Translation Protocol with metrics ToS, CoS, FoS and SSC to quantify visual knowledge bottlenecks in VLMs, plus a Divergence Law hypothesis that scaling language models may increase the penalty.

  20. VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

    cs.CV 2026-04 unverdicted novelty 6.0

    VLMs bypass visual comparison by recovering semantic labels for nameable entities and hallucinate on unnamable ones, as shown by performance gaps and Logit Lens analysis.

  21. Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

    cs.AI 2026-07 conditional novelty 5.0

    A 235-item multimodal stress-test shows frontier closed models outpace open-weight peers by ~10% and leaves shared failures on counting, spatial, and character-level tasks.

  22. V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

    cs.CV 2026-06 unverdicted novelty 5.0

    V-Zero trains MLLMs for visual reasoning without answer labels by gating on-policy distillation trajectories using contrastive evidence from relevant versus negative image crops.

  23. VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following

    cs.CV 2026-05 unverdicted novelty 5.0

    VLMs frequently switch away from a target visual path to nearby similar distractors in controlled tracing tasks, with standard scaling, reasoning, and instruction interventions providing only partial mitigation.

  24. SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

    cs.CV 2026-05 unverdicted novelty 5.0

    SenseNova-U1 presents native unified multimodal models that match top understanding VLMs while delivering strong performance in image generation, infographics, and interleaved tasks via the NEO-unify architecture.

  25. The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

    cs.CV 2026-04 unverdicted novelty 5.0

    Vision-language models exhibit functional blindness by exploiting language priors over visual representations; the Modality Translation Protocol and metrics like Toll, Curse, and Fallacy of Seeing reveal this, support...

  26. Kimi K2.5: Visual Agentic Intelligence

    cs.CL 2026-02 unverdicted novelty 5.0

    Kimi K2.5 combines joint text-vision training with an Agent Swarm parallel orchestration framework to reach claimed state-of-the-art results on coding, vision, reasoning, and agent tasks while cutting latency up to 4.5 times.

  27. Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

    cs.AI 2026-06 unverdicted novelty 2.0

    Seed2.0 model series reports gains in reasoning, visual understanding, search, and reliability on intricate long-horizon tasks via an internal evaluation system.

  28. EXAONE 4.5 Technical Report

    cs.CL 2026-04 unverdicted novelty 2.0

    EXAONE 4.5 is a new open-weight multimodal model that matches general benchmarks and outperforms similar-scale models on document understanding and Korean contextual reasoning.

Reference graph

Works this paper leans on

13 extracted references · 8 linked inside Pith · cited by 24 Pith papers

  1. [2]

    Accessed: 2025-01-09

    URL https: //bagel-ai.org/. Accessed: 2025-01-09. Zefan Cai, Haoyi Qiu, Tianyi Ma, Haozhe Zhao, Gengze Zhou, Kung-Hsiang Huang, Parisa Kordjamshidi, Minjia Zhang, Wen Xiao, Jiuxiang Gu, Nanyun Peng, and Junjie Hu. Mmgr: Multi-modal generative reasoning,

  2. [3]

    Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao

    URLhttps://arxiv.org/abs/2512.14691. Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao. Are we on the right way for evaluating large vision-language models? InAdvances in Neural Information Processing Systems, volume 37,

  3. [5]

    Accessed: 2025-01-09

    URL https://deepmind.googl e/technologies/veo/. Accessed: 2025-01-09. Jinsheng Huang, Liang Chen, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan, Haozhe Zhao, Zhihui Guo, Yichi Zhang, Jingyang Yuan, Wei Ju, Luchen Liu, Tianyu Liu, Baobao Chang, and Ming Zhang. Mmevalpro: Calibrating multimodal benchmarks towards trustworthy and efficient evaluation,

  4. [6]

    Scott P Johnson

    URLhttps://arxiv.org/abs/2407.00468. Scott P Johnson. Development of visual perception.Wiley Interdisciplinary Reviews: Cognitive Science, 1(5): 529–541,

  5. [7]

    Accessed: 2025-01-09

    URL https://openai.com/sora . Accessed: 2025-01-09. ByteDance Seed. Seed1.6 tech introduction,

  6. [8]

    Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu

    URLhttps://seed.bytedance.com/en/seed1_6. Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient rlhf framework.arXiv preprint arXiv: 2409.19256,

  7. [9]

    Core Team, Zihao Yue, Zhenru Lin, Yifan Song, Weikun Wang, Shuhuai Ren, Shuhao Gu, Shicheng Li, et al

    URLhttps://arxiv.org/abs/2507.19427. Core Team, Zihao Yue, Zhenru Lin, Yifan Song, Weikun Wang, Shuhuai Ren, Shuhao Gu, Shicheng Li, et al. Mimo-vl technical report, 2025a. URLhttps://arxiv.org/abs/2506.03569. Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth,...

  8. [11]

    Measuring multimodal mathematical reasoning with math-vision dataset.arXiv preprint arXiv:2402.14804,

    Ke Wang, Junting Pan, Weikang Shi, Zimu Lu, Mingjie Zhan, and Hongsheng Li. Measuring multimodal mathematical reasoning with math-vision dataset.arXiv preprint arXiv:2402.14804,

  9. [12]

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al

    URLhttps://arxiv.org/abs/2508.18265. Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. ...

  10. [13]

    Drvd- bench: Do vision-language models reason like human doctors in medical image diagnosis?arXiv preprint arXiv:2505.24173,

    25 Tianhong Zhou, Yin Xu, Yingtao Zhu, Chuxi Xiao, Haiyang Bian, Lei Wei, and Xuegong Zhang. Drvd- bench: Do vision-language models reason like human doctors in medical image diagnosis?arXiv preprint arXiv:2505.24173,

  11. [2023]

    Glm-4.6v: Open source multimodal models with native tool use, 2025a

    GLM-V Team. Glm-4.6v: Open source multimodal models with native tool use, 2025a. URL https: //z.ai/blog/glm-4.6v. HLE Team. Humanity’s last exam, 2025b. URLhttps://arxiv.org/abs/2501.14249. Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, et al. Kimi-vl technical report, 2025b. URL http...

  12. [2024]

    Mme: A comprehensive evaluation benchmark for multimodal large language models.arXiv preprint arXiv:2306.13394, 2024a

    Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, et al. Mme: A comprehensive evaluation benchmark for multimodal large language models.arXiv preprint arXiv:2306.13394, 2024a. Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranja...

  13. [2025]

    Renée Baillargeon, Elizabeth S Spelke, and Stanley Wasserman

    URLhttps://arxiv.org/abs/2511.21631. Renée Baillargeon, Elizabeth S Spelke, and Stanley Wasserman. Object permanence in five-month-old infants.Cognition, 20(3):191–208,