REVIEW 4 major objections 2 minor 70 references
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
T0 review · 4 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that current multimodal language models lose significant math-solving accuracy when questions come as authentic, phone-captured photos from real K-12 settings.
desk verdict The packet only gives us MathReal's abstract; the full text is an unrelated MLHC paper, so the benchmark's core claims—that the 2,000 images are authentically captured and that the evaluation cleanly separates error types—are unverified, though the idea itself is a real gap worth pursuing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the MathReal benchmark itself: a set of 2,000 real-scene K-12 math problem images, each containing both the question text and visual elements, with a taxonomy of real-world image defects (quality degradation, perspective variation, irrelevant content interference, and 14 finer subcategories). This taxonomy is what lets the paper move from 'models do worse on real photos' to a claim about why they do worse—it connects observed accuracy loss to specific visual distortions that real educational users encounter.
What would settle it
Inspect the dataset's capture provenance and rerun the benchmark with repaired versions of the same questions (e.g., reshot or cleaned images). If the performance drop disappears when the same math content is presented in clean form, and if a curation audit shows images were chosen to match the predefined taxonomy, the central claim that authentic capture conditions challenge MLLMs would be falsified.
Extended reading notes
Core claim
MathReal is a curated benchmark of 2,000 mathematical questions whose images were captured by handheld mobile devices in authentic K-12 scenarios. The paper claims that existing multimodal large language models perform markedly worse on these authentic images than on clean benchmark images. The dataset classifies real-world image issues into three primary categories—image quality degradation, perspective variation, and irrelevant content interference—which are further subdivided into 14 subcategories. The evaluation covers five knowledge and ability categories, three question types, and three difficulty levels, and uses six experimental settings to decompose model failures. The paper's concl
Load-bearing premise
The 2,000 images were genuinely captured by K-12 educational users in authentic scenarios, rather than staged, re-photographed, or selected specifically to fit the paper's degradation taxonomy.
Editorial extensions
If this is right
- If the finding holds, reported mathematical reasoning abilities of MLLMs on clean benchmarks are likely overestimated for real classroom use.
- The three-category taxonomy gives a concrete target for robustness work: models should be tested and trained on degraded, tilted, and cluttered images.
- Error analysis that separates recognition, comprehension, and reasoning failures provides a roadmap for where model improvements are needed most.
- Benchmark designers should include authentic capture conditions, not just synthetic corruptions, to reflect how users actually take photos.
- Performance across difficulty levels, question types, and knowledge categories can identify which kinds of math problems are most sensitive to real-image noise.
Reading between the lines
- The supplied full text is an unrelated manuscript about prognostication of cardiac arrest patients, so the dataset-construction and evaluation claims rest entirely on the abstract; no provenance details, annotation protocol, or raw results were inspectable.
- A natural testable extension is to compare human performance on the same authentic images: if human solvers are barely affected by the degradations, the model drop is a model-specific fragility rather than an intrinsic difficulty of the images.
- The degradation taxonomy may transfer to other visually grounded tasks such as document analysis, chart reading, or visual question answering, where authentic capture conditions are also under-represented in benchmarks.
- The clean-vs-authentic gap could be used as a controlled augmentation protocol: re-formatting the same questions as clean images and as authentic photos would isolate the visual-capture effect from content difficulty.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript as submitted consists of a title and abstract for arXiv:2508.06009, 'MathReal', which proposes a real-scene K-12 math benchmark for multimodal large language models, together with a full text that is a completely different paper: 'Stepwise Fine and Gray: Subject-Specific Variable Selection Shows When Hemodynamic Data Improves Prognostication of Comatose Post-Cardiac Arrest Patients' (arXiv:2508.06023), a competing-risks survival analysis study on 2,278 cardiac arrest patients. The abstract claims the introduction of a 2,000-question dataset with images captured by handheld devices in authentic scenarios, organized into 3 degradation categories and 14 subcategories, and reports that existing MLLMs are 'significantly challenged' in realistic educational contexts. However, none of these content elements—dataset construction, taxonomy definitions, six experimental settings, model evaluation tables, or error analysis—appears anywhere in the supplied full text. The body's sections, equations, tables, and appendices describe a survival analysis method and its clinical application, not a multimodal math reasoning benchmark. Consequently, the paper's central claims are asserted in the abstract but are entirely unsupported by the submitted manuscript and cannot be inspected or verified.
Significance. If the MathReal benchmark exists as described, it addresses a real and valuable gap: most existing MLLM math benchmarks use clean or processed images, and an authentically captured K-12 image benchmark could support robustness research on real-world degradations such as blur, perspective distortion, and irrelevant content. The abstract promises a useful resource with a systematic error taxonomy and a detailed performance analysis. However, significance cannot be assessed from this submission. The submitted full text is an unrelated clinical machine learning paper, so there is no dataset description, no annotation or quality-control protocol, no experimental design, no results, and no reproducible code or data for the MathReal claims. The GitHub link in the abstract is not referenced or described in the body. No credit can be assigned to the benchmark's construction or to any empirical finding, because none of that content is present in the manuscript under review.
major comments (4)
- [Full Text (title, Sections 1–6, Appendices A–E)] The submitted full text is not the MathReal paper. The body's title, abstract, and content (e.g., Section 4 describing a cohort of 2,278 comatose cardiac-arrest patients; Table 1 reporting competing-risks c-indices; Eq. (8) defining an incremental contribution in a survival model) correspond to arXiv:2508.06023, 'Stepwise Fine and Gray'. The claimed MathReal dataset, taxonomy, six experimental settings, MLLM evaluations, and error analyses are absent. This is a load-bearing failure: the paper's central claim exists only in the abstract and has no accompanying manuscript evidence.
- [Abstract (dataset construction and authenticity)] The abstract states that the 2,000 images were 'captured by handheld mobile devices in authentic scenarios' and 'meticulously curated'. The submission contains no description of the data collection process, no inclusion/exclusion criteria, no annotation protocol, no inter-annotator agreement for the 14 subcategories, and no verification that ground-truth answers are correct under the same visual defects the models must overcome. Without an inspectable protocol, the headline claim that MLLMs are 'significantly challenged' by authentic real-scene factors, as opposed to curation or label artifacts, cannot be tested.
- [Abstract (experimental settings and error decomposition)] The abstract promises 'six experimental settings' and a 'thorough analysis of their performance and error patterns, providing insights into their recognition, comprehension, and reasoning capabilities.' None of these settings, the error taxonomy, the model list, or the performance tables is present in the submitted full text. Consequently, the claimed attribution of performance drops to image quality degradation, perspective variation, and irrelevant content interference, and the decomposition into recognition/comprehension/reasoning errors, is not supported by any describable experimental design.
- [Entire submission] Because the body is a different paper, the submission cannot be treated as a coherent manuscript. Replacing the body with the actual MathReal paper is not a revision of the current text; it is a new submission. The current manuscript therefore does not meet the minimal standard of containing its own claimed content, and no amount of local editing can fix the absence of the central empirical study.
minor comments (2)
- [Full Text header] The full text header identifies the paper as 'arXiv:2508.06023v1', while the submission's title and abstract correspond to arXiv:2508.06009. The title/abstract/body mismatch needs to be resolved at the submission level; this is not a stylistic typo but a metadata error that must be corrected.
- [Abstract (data availability)] The abstract includes a GitHub URL for MathReal, but the body (which is a different paper) does not reference this repository or provide a data-availability statement. Should the correct MathReal manuscript be resubmitted, a description of the dataset contents, license, and access instructions is needed.
Circularity Check
No significant circularity; the full-text paper evaluates on held-out test data and its key quantity is an estimated log-subhazard difference, not a prediction defined by its own inputs.
full rationale
The manuscript whose full text is provided is arXiv:2508.06023 (Shen, Elmer, Chen, 'Stepwise Fine and Gray'), not the MathReal benchmark named in the abstract; the MathReal dataset-construction and error-taxonomy sections are therefore not present for inspection. That is an evidentiary gap, not a circularity: under the hard rules, circularity requires quoting a specific reduction, and none can be exhibited for the benchmark claims. For the paper actually in the packet, the load-bearing quantity I_k(h|X_t) in Eq. (8) is an algebraic cancellation of the Phase-1 log-subhazard from the Phase-2 log-subhazard; it is the definition of the model's 'incremental contribution,' not a first-principles prediction that reduces to its own inputs. The learned threshold is selected on a held-out validation set and performance is reported on a separate test split (stratified 64/16/20), so the headline finding ('not always beneficial to use all features') is empirically evaluated rather than forced by construction. Self-citation to Shen et al. (2023) is limited to the claim that Fine and Gray is a strong baseline and is not load-bearing to the new two-phase contribution. The paper's stated limitations—single-center observational data (Section 6) and deviations from i.i.d. in the dynamic setup (Appendix B.3)—are validity risks, not instances of definitional circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Existing math-reasoning benchmarks are predominantly based on clean or processed multimodal inputs and omit real-world K-12 user images.
- domain assumption The taxonomy of three primary categories (image quality degradation, perspective variation, irrelevant content interference) and 14 subcategories is a valid, exhaustive, and non-overlapping description of the real-capture variations in the dataset.
- domain assumption Ground-truth answers and question labels in MathReal are correct despite the visual defects, so that model errors can be attributed to the models rather than to the benchmark.
Cite this review
Pith. "Pith review of MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models." pith.science (2026). https://pith.science/paper/WD4EZAFF
@misc{pith2026250806009,
author = {Pith},
title = {Pith review of: MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/WD4EZAFF}},
note = {Machine review of arXiv:2508.06009}
}
read the original abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in visual mathematical reasoning across various existing benchmarks. However, these benchmarks are predominantly based on clean or processed multimodal inputs, without incorporating the images provided by real-world Kindergarten through 12th grade (K-12) educational users. To address this gap, we introduce MathReal, a meticulously curated dataset comprising 2,000 mathematical questions with images captured by handheld mobile devices in authentic scenarios. Each question is an image, containing the question text and visual element. We systematically classify the real images into three primary categories: image quality degradation, perspective variation, and irrelevant content interference, which are further delineated into 14 subcategories. Additionally, MathReal spans five core knowledge and ability categories, which encompass three question types and are divided into three difficulty levels. To comprehensively evaluate the multimodal mathematical reasoning abilities of state-of-the-art MLLMs in real-world scenarios, we design six experimental settings that enable a systematic analysis of their performance. Through extensive experimentation, we find that the problem-solving abilities of existing MLLMs are significantly challenged in realistic educational contexts. Based on this, we conduct a thorough analysis of their performance and error patterns, providing insights into their recognition, comprehension, and reasoning capabilities, and outlining directions for future improvements. Data and code: https://github.com/junfeng0288/MathReal.
Reference graph
Works this paper leans on
-
[1]
AI, M. 2025 a . Llama 4. https://www.llama.com/
work page 2025
-
[2]
AI, Z. 2025 b . GLM-4.1v-thinking-flashx Model Announcement. https://www.zhipuai.cn/
work page 2025
-
[3]
Amini, A.; Gabriel, S.; Lin, S.; Koncel-Kedziorski, R.; Choi, Y.; and Hajishirzi, H. 2019. MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Paper...
work page 2019
-
[4]
Ankner, Z.; Paul, M.; Cui, B.; Chang, J. D.; and Ammanabrolu, P. 2024. Critique-out-Loud Reward Models. In Pluralistic Alignment Workshop at NeurIPS 2024
work page 2024
-
[5]
Anthropic. 2025. Claude Sonnet 4. https://www.anthropic.com/claude/sonnet
work page 2025
-
[6]
Awais, M.; Ahmed, T.; Aslam, M.; Rehman, A.; Alamri, F. S.; Bahaj, S. A.; and Saba, T. 2024. Mathvision: An accessible intelligent agent for visually impaired people to understand mathematical equations. IEEE Access
work page 2024
-
[7]
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv:2308.12966
arXiv 2023
-
[8]
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. 2025. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923
arXiv 2025
Show all 70 references
-
[9]
Baidu. 2025. ERNIE Technical Report. https://yiyan.baidu.com/blog/publication/ERNIE_Technical_Report.pdf
2025
-
[10]
ByteDance. 2025 a . Doubao-1.5-thinking-vision-pro. https://www.volcengine.com/docs/82379/1554521
2025
-
[11]
ByteDance. 2025 b . Doubao-1.5-vision-pro. https://www.volcengine.com/docs/82379/1553586
2025
-
[12]
ByteDance. 2025 c . Doubao-seed-1.6. https://seed.bytedance.com/en/seed1_6
2025
-
[13]
ByteDance. 2025 d . Doubao-seed-1.6-thinking. https://seed.bytedance.com/en/seed1_6
2025
-
[14]
Chen, H.; Tu, H.; Wang, F.; Liu, H.; Tang, X.; Du, X.; Zhou, Y.; and Xie, C. 2025 a . Sft or rl? an early investigation into training r1-like reasoning large vision-language models. arXiv preprint arXiv:2504.11468
2025 arXiv
-
[15]
Chen, S.; Guo, Y.; Su, Z.; Li, Y.; Wu, Y.; Chen, J.; Chen, J.; Wang, W.; Qu, X.; and Cheng, Y. 2025 b . Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning. arXiv preprint arXiv:2506.04207
2025
-
[16]
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168
2021 arXiv
-
[17]
Comanici, G.; Bieber, E.; Schaekermann, M.; Pasupat, I.; Sachdeva, N.; Dhillon, I.; Blistein, M.; Ram, O.; Zhang, D.; Rosen, E.; et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv ...
2025 arXiv
-
[18]
Deng, Y.; Bansal, H.; Yin, F.; Peng, N.; Wang, W.; and Chang, K.-W. 2025. Openvlthinker: An early exploration to complex vision-language reasoning via iterative self-improvement. arXiv preprint arXiv:2503.17352
2025 arXiv
-
[19]
Fu, D.; Chen, Z.; Xia, R.; Liu, Q.; Feng, Y.; Zhou, H.; Zhang, R.; Feng, S.; Gao, P.; Yan, J.; et al. 2025. Trustgeogen: Scalable and formal-verified data engine for trustworthy multi-modal geometric problem solving. arXiv preprint arXiv:2504.15780
2025
-
[20]
Fu, L.; Kuang, Z.; Song, J.; Huang, M.; Yang, B.; Li, Y.; Zhu, L.; Luo, Q.; Wang, X.; Lu, H.; et al. 2024. Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning. arXiv preprint arXiv:2501.00321
2024 arXiv
-
[21]
Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948
2025 arXiv
-
[22]
L.; Shen, J.; Hu, J.; Han, X.; Huang, Y.; Zhang, Y.; et al
He, C.; Luo, R.; Bai, Y.; Hu, S.; Thai, Z. L.; Shen, J.; Hu, J.; Han, X.; Huang, Y.; Zhang, Y.; et al. 2024. Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems. arXiv preprint arXiv:2402.14008
2024 arXiv
-
[23]
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874
2021 arXiv
-
[24]
Hong, Z.; Wu, H.; Dong, S.; Dong, J.; Xiao, Y.; Zhang, Y.; Wang, Z.; Huang, F.; Li, L.; Yang, H.; et al. 2025. Benchmarking large language models via random variables. arXiv preprint arXiv:2501.11790
2025 arXiv
-
[25]
Huang, M.; Shi, Y.; Peng, D.; Lai, S.; Xie, Z.; and Jin, L. 2025. Ocr-reasoning benchmark: Unveiling the true capabilities of mllms in complex text-rich image reasoning. arXiv preprint arXiv:2505.17163
2025 arXiv
-
[26]
Jaech, A.; Kalai, A.; Lerer, A.; Richardson, A.; El-Kishky, A.; Low, A.; Helyar, A.; Madry, A.; Beutel, A.; Carney, A.; et al. 2024. OpenAI o1 System Card. CoRR
2024
-
[27]
Kamoi, R.; Zhang, Y.; Das, S. S. S.; Zhang, R. H.; and Zhang, R. 2024. Visonlyqa: Large vision language models still struggle with visual perception of geometric information. arXiv preprint arXiv:2412.00947
2024 arXiv
-
[28]
Leng, S.; Wang, J.; Li, J.; Zhang, H.; Hu, Z.; Zhang, B.; Zhang, H.; Jiang, Y.; Li, X.; Zhao, D.; et al. 2025. Mmr1: Advancing the frontiers of multimodal reasoning
2025
-
[29]
Li, B.; Ge, Y.; Chen, Y.; Ge, Y.; Zhang, R.; and Shan, Y. 2024 a . Seed-bench-2-plus: Benchmarking multimodal large language models with text-rich visual comprehension. arXiv preprint arXiv:2404.16790
2024 arXiv
-
[30]
Li, C.; Zhang, T.; Wang, M.; and Huang, H. 2025. VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs. arXiv preprint arXiv:2506.06727
2025
-
[31]
Li, Z.-Z.; Zhang, M.-L.; Yin, F.; Ji, Z.-L.; Bai, J.-F.; Pan, Z.-R.; Zeng, F.-H.; Xu, J.; Zhang, J.-X.; and Liu, C.-L. 2024 b . Cmmath: A chinese multi-modal math skill evaluation benchmark for foundation models. arXiv preprint arXiv:2407.12023
2024 arXiv
-
[32]
Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; et al. 2024 a . Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437
2024 arXiv
-
[33]
Liu, C.; Wei, H.; Chen, J.; Kong, L.; Ge, Z.; Zhu, Z.; Zhao, L.; Sun, J.; Han, C.; and Zhang, X. 2024 b . Focus anywhere for fine-grained multi-page document understanding. arXiv preprint arXiv:2405.14295
2024 arXiv
-
[34]
Liu, H.; Zheng, Z.; Qiao, Y.; Duan, H.; Fei, Z.; Zhou, F.; Zhang, W.; Zhang, S.; Lin, D.; and Chen, K. 2024 c . MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark. In Findings of the Association for Computational Ling...
2024
-
[35]
Lu, P.; Bansal, H.; Xia, T.; Liu, J.; Li, C.; Hajishirzi, H.; Cheng, H.; Chang, K.-W.; Galley, M.; and Gao, J. 2023. MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts. arXiv e-prints, arXiv--2310
2023
-
[36]
X.; Tan, J
Masry, A.; Long, D. X.; Tan, J. Q.; Joty, S.; and Hoque, E. 2022. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. arXiv preprint arXiv:2203.10244
2022 arXiv
-
[37]
Mathew, M.; Karatzas, D.; and Jawahar, C. 2021. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2200--2209
2021
-
[38]
Meng, F.; Du, L.; Liu, Z.; Zhou, Z.; Lu, Q.; Fu, D.; Shi, B.; Wang, W.; He, J.; Zhang, K.; et al. 2025. Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning. CoRR
2025
-
[39]
OpenAI. 2024. GPT-4o. https://platform.openai.com/docs/models/gpt-4o
2024
-
[40]
OpenAI. 2025 a . GPT-4.1. https://openai.com/index/gpt-4-1/
2025
-
[41]
OpenAI. 2025 b . o3-and-o4-mini. https://openai.com/index/introducing-o3-and-o4-mini/
2025
-
[42]
Qiao, R.; Tan, Q.; Dong, G.; Wu, M.; Sun, C.; Song, X.; Gongque, Z.; Lei, S.; Wei, Z.; Zhang, M.; et al. 2024. We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? CoRR
2024
-
[43]
Shen, W.; Pei, J.; Peng, Y.; Song, X.; Liu, Y.; Peng, J.; Sun, H.; Hao, Y.; Wang, P.; and Zhou, Y. 2025. Skywork-R1V3 Technical Report. arXiv preprint arXiv:2507.06167
2025 arXiv
-
[44]
W.; Tay, Y.; Ruder, S.; Zhou, D.; et al
Shi, F.; Suzgun, M.; Freitag, M.; Wang, X.; Srivats, S.; Vosoughi, S.; Chung, H. W.; Tay, Y.; Ruder, S.; Zhou, D.; et al. 2022. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations
2022
-
[45]
Sun, K.; Bai, Y.; Qi, J.; Hou, L.; and Li, J. 2024. MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification. In Findings of the Association for Computational Linguistics: EMNLP 2024, 1358--1375
2024
-
[46]
Sun, Y.; Zhang, S.; Tang, W.; Chen, A.; Koniusz, P.; Zou, K.; Xue, Y.; and van den Hengel, A. 2025. MATHGLANCE: Multimodal Large Language Models Do Not Know Where to Look in Mathematical Diagrams. CoRR
2025
-
[47]
Team, C.; Yue, Z.; Lin, Z.; Song, Y.; Wang, W.; Ren, S.; Gu, S.; Li, S.; Li, P.; Zhao, L.; Li, L.; Bao, K.; Tian, H.; Zhang, H.; Wang, G.; Zhu, D.; Cici; He, C.; Ye, B.; Shen, B.; Zhang, Z.; Jiang, Z.; Zheng, Z.; Song, Z.; Luo, Z.; Yu, Y.; Wang, Y.; Tian, Y.; Tu, Y.; Yan, Y.; ...
2025 arXiv
-
[48]
Team, G.; Kamath, A.; Ferret, J.; Pathak, S.; Vieillard, N.; Merhej, R.; Perrin, S.; Matejovicova, T.; Ram \'e , A.; Rivi \`e re, M.; et al. 2025 b . Gemma 3 technical report. arXiv preprint arXiv:2503.19786
2025 arXiv
-
[49]
Team, K.; Du, A.; Yin, B.; Xing, B.; Qu, B.; Wang, B.; Chen, C.; Zhang, C.; Du, C.; Wei, C.; et al. 2025 c . Kimi-vl technical report. arXiv preprint arXiv:2504.07491
2025 arXiv
-
[50]
K.; Yang, B.; Wen, B.; Liu, C.; Chu, C.; Song, C.; Rao, C.; Yi, C.; Li, D.; Zang, D.; et al
Team, K. K.; Yang, B.; Wen, B.; Liu, C.; Chu, C.; Song, C.; Rao, C.; Yi, C.; Li, D.; Zang, D.; et al. 2025 d . Kwai Keye-VL Technical Report. arXiv preprint arXiv:2507.01949
2025 arXiv
-
[51]
Wang, H.; Qu, C.; Huang, Z.; Chu, W.; Lin, F.; and Chen, W. 2025 a . Vl-rethinker: Incentivizing self-reflection of vision-language models with reinforcement learning. arXiv preprint arXiv:2504.08837
2025 arXiv
-
[52]
Wang, K.; Pan, J.; Shi, W.; Lu, Z.; Ren, H.; Zhou, A.; Zhan, M.; and Li, H. 2024. Measuring multimodal mathematical reasoning with math-vision dataset. Advances in Neural Information Processing Systems, 37: 95095--95169
2024
-
[53]
Wang, P.; Li, Z.-Z.; Yin, F.; Ran, D.; and Liu, C.-L. 2025 b . Mv-math: Evaluating multimodal math reasoning in multi-visual contexts. In Proceedings of the Computer Vision and Pattern Recognition Conference, 19541--19551
2025
-
[54]
Wang, P.; Li, Z.-Z.; Yin, F.; Ran, D.; and Liu, C.-L. 2025 c . Mv-math: Evaluating multimodal math reasoning in multi-visual contexts. In Proceedings of the Computer Vision and Pattern Recognition Conference, 19541--19551
2025
-
[55]
Wang, X.; Yang, Z.; Feng, C.; Lu, H.; Li, L.; Lin, C.-C.; Lin, K.; Huang, F.; and Wang, L. 2025 d . Sota with less: Mcts-guided sample selection for data-efficient visual reasoning self-improvement. arXiv preprint arXiv:2504.07934
2025 arXiv
-
[56]
Wang, Y.; Zhang, P.; Tang, J.; Wei, H.; Yang, B.; Wang, R.; Sun, C.; Sun, F.; Zhang, J.; Wu, J.; et al. 2025 e . Polymath: Evaluating mathematical reasoning in multilingual contexts. arXiv preprint arXiv:2504.18428
2025
-
[57]
Wang, Z.; Sun, J.; Zhang, W.; Hu, Z.; Li, X.; Wang, F.; and Zhao, D. 2025 f . Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency. arXiv preprint arXiv:2504.18589
2025 arXiv
-
[58]
Wei, Y.; Zhao, L.; Sun, J.; Lin, K.; Yin, J.; Hu, J.; Zhang, Y.; Yu, E.; Lv, H.; Weng, Z.; et al. 2025. Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning. arXiv preprint arXiv:2507.05255
2025
-
[59]
xAI. 2025. Grok. https://x.ai/grok
2025
-
[60]
Xiao, Y.; Sun, E.; Liu, T.; and Wang, W. 2024. Logicvista: Multimodal llm logical reasoning benchmark in visual contexts. arXiv preprint arXiv:2407.04973
2024 arXiv
-
[61]
Xu, L.; Xue, H.; Zhu, L.; and Zhao, K. 2024. Superclue-math6: Graded multi-step math reasoning benchmark for llms in chinese. arXiv preprint arXiv:2401.11819
2024 arXiv
-
[62]
Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; et al. 2025 a . Qwen3 technical report. arXiv preprint arXiv:2505.09388
2025 arXiv
-
[63]
Yang, J.; Ma, F.; Wang, Z.; Yin, D.; Rong, K.; Rao, F.; and Zhang, R. 2025 b . WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning. arXiv preprint arXiv:2506.07905
2025 arXiv
-
[64]
Yang, Z.; Tang, J.; Li, Z.; Wang, P.; Wan, J.; Zhong, H.; Liu, X.; Yang, M.; Wang, P.; Bai, S.; et al. 2024. Cc-ocr: A comprehensive and challenging ocr benchmark for evaluating large multimodal models in literacy. arXiv preprint arXiv:2412.02210
2024 arXiv
-
[65]
Yu, T.; Jing, Y.; Zhang, X.; Jiang, W.; Wu, W.; Wang, Y.; Hu, W.; Du, B.; and Tao, D. 2025. Benchmarking reasoning robustness in large language models. arXiv preprint arXiv:2503.04550
2025 arXiv
-
[66]
Zhang, J.; Li, Z.-Z.; Zhang, M.-L.; Yin, F.; Liu, C.-L.; and Moshfeghi, Y. 2024 a . GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving. In Findings of the Association for Computational Linguistics ACL 2024, 1258--1276
2024
-
[67]
Zhang, R.; Jiang, D.; Zhang, Y.; Lin, H.; Guo, Z.; Qiu, P.; Zhou, A.; Lu, P.; Chang, K.-W.; Qiao, Y.; et al. 2024 b . Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, 169--186. Springer
2024
-
[68]
Zheng, M.; Feng, X.; Si, Q.; She, Q.; Lin, Z.; Jiang, W.; and Wang, W. 2024. Multimodal table understanding. arXiv preprint arXiv:2406.08100
2024 arXiv
-
[69]
Zhu, J.; Wang, W.; Chen, Z.; Liu, Z.; Ye, S.; Gu, L.; Tian, H.; Duan, Y.; Su, W.; Shao, J.; et al. 2025. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479
2025 arXiv
-
[70]
Zou, C.; Guo, X.; Yang, R.; Zhang, J.; Hu, B.; and Zhang, H. 2024. Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models. arXiv preprint arXiv:2411.00836
2024 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.