Pith. sign in

REVIEW 2 major objections 5 minor 62 references

Multimodal models often pick useful visual actions, but faithful rendering is the bottleneck—and corrupted visual feedback still changes their answers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 21:36 UTC pith:VVNQO6VN

load-bearing objection Solid evaluation package that cleanly separates visual-state utility from behavioral dependence; residual risk is intervention quality, not the design. the 2 major comments →

arxiv 2607.26769 v1 pith:VVNQO6VN submitted 2026-07-29 cs.CV cs.AI

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

classification cs.CV cs.AI
keywords multimodal reasoningvisual chain-of-thoughtintermediate visual statesaction relevancerender faithfulnessfeedback uptakecorrupted feedbackbenchmark evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether multimodal models that sketch, annotate, crop, or otherwise produce intermediate images during reasoning actually depend on those visual states, or merely decorate a mostly textual chain of thought. It builds a 1,200-problem benchmark of visually dependent tasks across 2D diagrams, 3D scenes, and real-world settings, then runs a closed-loop protocol that records planned visual actions, externally rendered states, and later reasoning under matched conditions—including deliberately corrupted but task-relevant feedback. Across several strong models, no single strategy (text-only, plan-without-render, full closed loop) wins everywhere. Models usually choose relevant operations, yet realizing those operations faithfully is the clearest weak link, and high uptake of returned pixels does not reliably raise accuracy. Still, when the returned state is corrupted in a task-relevant way, accuracy falls—by more than ten points in 3D scenes—showing behavioral dependence even when clean visual feedback was not a free accuracy boost.

Core claim

Genuine intermediate visual-state use is model- and environment-dependent rather than a universal gain: models typically select task-relevant visual operations, faithful rendering is the main bottleneck after planning, high feedback uptake need not improve final accuracy, and task-relevant corrupted feedback still induces measurable behavioral dependence, with accuracy drops over 10 percentage points in 3D scene reasoning under controlled interventions.

What carries the argument

Visual Action-of-Thought (VAoT): an intervenable closed loop that interleaves textual thoughts, structured visual actions, externally rendered image states, and subsequent reasoning, compared across CoT, action planning without rendering, standard VAoT, and task-relevant WrongRender feedback, with process scores for action relevance, render faithfulness, and feedback uptake.

Load-bearing premise

The claim rests on treating an external constrained renderer plus automatic process scores and task-relevant corrupted edits as faithful enough stand-ins for whether a model truly generated, saw, and used an intermediate visual state.

What would settle it

Re-run the paired VAoT versus WrongRender comparison on a large human-verified set where every corruption clearly alters task-critical content while staying visually plausible; if accuracy no longer drops systematically—especially in 3D scenes—behavioral dependence is not established.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Final-answer gains alone cannot certify that a model is thinking with images; action relevance, render faithfulness, and feedback uptake must be measured separately.
  • Improving intermediate visual reasoning should prioritize faithful execution of planned edits over teaching models to request more visual operations.
  • Utility and dependence diverge: a visual workspace can steer decisions under corruption even when clean rendering yields little net accuracy benefit.
  • Benchmarking and tool design should filter text-only shortcuts and use matched corrupted-feedback interventions, not only oracle sketches or end-task scores.
  • The best visual-reasoning regime will remain model- and environment-specific rather than one fixed closed-loop recipe.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Training objectives that reward tool calls or intermediate images without a faithfulness or dependence check may inflate visual-trace metrics while leaving the render bottleneck untouched.
  • If 3D relational tasks show the strongest corruption sensitivity, intermediate-state methods may matter most where structure is not already explicit in the input diagram.
  • Product systems that show users model-drawn highlights should treat those overlays as potentially causal inputs, not just explanations, because answers move when the overlay is wrong.
  • A natural next stress test is learned or model-internal renderers under the same WrongRender pairing, to see whether the bottleneck is the external editor or the model’s use of any returned pixels.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. See2Think asks whether multimodal models genuinely use intermediate visual states rather than merely producing visual traces. The authors introduce See2ThinkBench (1,200 caption-filtered, mostly free-form problems across 12 categories in 2D structured, 3D scene, and real-world settings) and Visual Action-of-Thought (VAoT), which logs thoughts, structured actions, rendered states, and later reasoning under CoT, VAoT-NoRender, VAoT, and task-relevant VAoT-WrongRender. Across GPT-5.5, GPT-o3, Gemini 3.5 Flash, and Qwen3-VL-32B-Instruct they report that no single setting dominates; action relevance is near saturation while render faithfulness is the main bottleneck; high feedback uptake need not raise accuracy; and corrupted feedback still induces behavioral dependence, with large drops especially in 3D scenes. Process scores are human-audited, and WrongRender quality is audited with filtered re-analysis in the appendix.

Significance. The paper addresses a timely gap: existing visual-thinking benchmarks largely score final answers or aggregate process quality without jointly testing action relevance, render faithfulness, and behavioral dependence under matched interventions. The four-way protocol, caption-only shortcut filtering, full 1.2K paired outcome tables, 4,800-trajectory process analysis, outcome-stratified gaps, render benefit/harm transitions, and human audits (process judgments ≥92.9% reasonable-or-partial; WrongRender quality audit with positive filtered drops) are concrete methodological contributions. If the dependence results hold under clearer main-text quality controls, the work gives the community a reusable diagnostic that separates visual-state utility from behavioral dependence—useful for tool-using and “thinking with images” systems beyond any single accuracy leaderboard.

major comments (2)
  1. [§4.3–4.4, Fig. 8, Appendix F.2] §4.3–4.4 and Takeaway 3/6 lean on VAoT–WrongRender accuracy drops (e.g., >10 pp and up to 15.5 pp in 3D at Feedback Uptake=1; Fig. 8, Table 7) as evidence of behavioral dependence. Appendix F.2 reports only 56.7% Strict Pass (68/120) and 78.3% Acceptable on the human WrongRender audit, and the paired drop shrinks from 8.82 pp (strict) to 3.19 pp (relaxed). The direction is preserved, but magnitude and environment-specific claims are sensitive to intervention quality. The main text should report Strict/Acceptable rates and quality-filtered drops alongside Fig. 8, and soften absolute “over 10 percentage points” language where it is not restricted to quality-passed cases.
  2. [§4.4–4.5, Table 3–4, Fig. 6, Appendix F.1] Process diagnosis (Table 3, Fig. 6) rests on a single external judge (GPT-5.4) selecting one key step and scoring Action Relevance / Render Faithfulness / Feedback Uptake on {0, 0.5, 1}. Human audit (Table 4, §4.5) is reassuring at the reasonable-or-partial level (92.9–96.9%) but strict “Reasonable” rates are lower for Render Faithfulness (70.2%) and Action Relevance (74.8%), with large annotator disagreement on the strict boundary (Appendix F.1). Because Takeaway 4 identifies faithful execution—not action selection—as the bottleneck, the paper should either (i) report inter-annotator agreement / score-level confusion on the audited subset for Render Faithfulness, or (ii) show that the correct–incorrect Render gap in 3D (0.097 in Fig. 6b) remains under human-rescored or double-judged subsets. Without that, the stage-localization claim is only partially stress-tested.
minor comments (5)
  1. [§4.2, Table 2] Robot Manipulation accuracy is near floor (0–4% in Table 2) across all settings. A short note in §4.2 on whether this is grounding granularity, action-format mismatch, or benchmark construction would prevent over-reading “no visual benefit” on that category.
  2. [§2.2, Appendix A.2] Caption-only filtering uses a strong MLLM captioner and a ≥3/5 solvability rule (Appendix A.2). State in the main §2.2 how many candidates were removed and whether any sensitivity check (e.g., 2/5 vs 4/5) was run, so readers can gauge residual text-shortcut risk.
  3. [Table 1, §2.1] Table 1’s “Action–Render–Use Diagnosis” checkmark is fair, but a one-sentence clarification that MIRA/ViC/TWI/TwiFF were not re-run under VAoT would avoid implying head-to-head process scores on identical items.
  4. [Abstract, §4.1] Normalize naming of the judge model (GPT-5.4 in §4.1 vs GPT-5.5 as an evaluated model) and fix minor typos in the abstract/intro spacing (“imagesduring”, “itremainsunclear”).
  5. [Figure 5] Figure 5 and group aggregates would be easier to read with error bars or per-model spreads, given the strong model-dependence emphasized in Takeaway 1.

Circularity Check

0 steps flagged

No significant circularity: empirical evaluation with external outcomes and matched interventions, not a self-defining derivation.

full rationale

See2Think is a benchmark-and-protocol paper. Its load-bearing claims are comparative accuracies under CoT, VAoT-NoRender, VAoT, and VAoT-WrongRender, plus process scores (Action Relevance, Render Faithfulness, Feedback Uptake) that are not algebraically identified with final-answer correctness. The paper explicitly treats utility and behavioral dependence as distinct (e.g., high uptake without accuracy gains; WrongRender drops even when VAoT does not beat NoRender). Caption-only filtering removes text-solvable items rather than defining the target quantity. Process judgments are external (GPT-5.4) and human-audited; WrongRender quality is separately audited with quality-filtered paired drops still positive. There is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors to force the result, and no equation that reduces the claimed finding to its inputs by construction. Residual concerns about judge/renderer proxies are validity/measurement issues, not circular derivation.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 3 invented entities

Load-bearing premises are methodological: visual dependence via caption-only solvability, external non-solving renderer, automatic semantic/process judges, and task-relevant corruption as a probe of dependence. No physical constants or fitted scientific parameters; free choices are evaluation thresholds and scoring rubrics.

free parameters (3)
  • Caption-solvability removal threshold (≥3/5 correct) = ≥3 of 5 trials
    Hand-chosen cutoff for discarding caption-only solvable items; changes which samples count as visually dependent.
  • Process score rubric {0, 0.5, 1} on one key step = {0, 0.5, 1}; one key step per trajectory
    Discretization and single-key-step aggregation compress multi-step trajectories into three scalar diagnostics used throughout analysis.
  • WrongRender modify_key intervention policy = modify_key only (modify_non_key unused)
    Corruption is generated by an image editor prompt aiming for natural-looking task-relevant errors; quality is imperfect and magnitude of measured dependence depends on this choice.
axioms (4)
  • domain assumption A sample answered correctly from question+caption in ≥3/5 trials without the image is text-dominant and unsuitable for diagnosing visual-state use.
    Section 2.2 visual dependency filtering; core to claiming the bench is visually dependent.
  • domain assumption The external renderer is a constrained workspace that executes structured edits without the answer or a solving objective, so returned states are inspectable interventions rather than hidden model internals.
    Section 3 renderer principle; required for causal interpretation of VAoT vs WrongRender.
  • domain assumption Semantic-equivalence judging of free-form answers and GPT-5.4 process scoring are sufficiently reliable after human audit for the reported conclusions.
    Sections 4.1, 4.5, Appendix C–F; accuracy and process claims rest on these judges.
  • ad hoc to paper Task-relevant corrupted but natural-looking renders probe behavioral dependence on visual states when the model is not told feedback is corrupted.
    Section 4.1 and Appendix B.5; defines the main causal test distinguishing dependence from utility.
invented entities (3)
  • See2ThinkBench independent evidence
    purpose: 1.2K visually dependent problems across 12 categories and three visual environments for outcome evaluation.
    New benchmark constructed from existing sources with filtering and format normalization; not a physical entity.
  • Visual Action-of-Thought (VAoT) protocol independent evidence
    purpose: Interleaved record of thoughts, structured visual actions, rendered states, and controlled NoRender/WrongRender variants for process diagnosis.
    Inference-time diagnostic protocol invented for this paper; falsifiable via released trajectories and interventions.
  • Action Relevance / Render Faithfulness / Feedback Uptake triad independent evidence
    purpose: Stage-wise scores localizing planning vs execution vs uptake failures.
    Paper-defined diagnostic dimensions with human validation; independent of final accuracy by construction.

pith-pipeline@v1.2.0-daily-grok45 · 28837 in / 3475 out tokens · 77426 ms · 2026-07-30T21:36:16.714309+00:00 · methodology

0 comments
read the original abstract

Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

62 extracted references · 9 linked inside Pith

  1. [1]

    Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux

    Florian Bordes, Quentin Garrido, Justine T. Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux. Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environments, 2025

  2. [2]

    M3CoT: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought

    Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che. M3CoT: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought. InProceedingsofthe 62ndAnnualMeetingofthe Association forComputationalLinguistics(Volume1: LongPapers), pages 8199–8221, Bangkok, Thailand, 2024. Association for ComputationalLinguistics. doi: 10.18653...

  3. [3]

    Do multimodal agents really benefit from tool use? a systematic study of capability gains.arXiv preprint arXiv:2606.02357, 2026

    Garvin Guo, Donglei Yu, Yu Chen, Xiang Wang, Shuai Li, Xinpei Zhao, Huaxing Liu, Qinghao Wang, and Minpeng Liao. Do multimodal agents really benefit from tool use? a systematic study of capability gains.arXiv preprint arXiv:2606.02357, 2026

  4. [4]

    Rbench-v: A primary assessment for visual reasoning models with multi-modal outputs, 2025

    Meng-Hao Guo, Xuanyu Chu, Qianrui Yang, Zhe-Han Mo, Yiqing Shen, Pei-lin Li, Xinjie Lin, Jinnian Zhang, 12 Xin-Sheng Chen, Yi Zhang, Kiyohiro Nakayama, Zhengyang Geng, Houwen Peng, Han Hu, and Shi-Min Hu. Rbench-v: A primary assessment for visual reasoning models with multi-modal outputs, 2025

  5. [5]

    Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark, 2025

    Yunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang, and Yu Cheng. Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark, 2025

  6. [6]

    TVI-CoT: Text-visual interleaved chain-of-thought reasoning for multimodal understanding.arXivpreprint arXiv:2606.08464, 2026

    Lianyu Hu, Xiaoyu Ma, Zeqin Liao, and Yang Liu. TVI-CoT: Text-visual interleaved chain-of-thought reasoning for multimodal understanding.arXivpreprint arXiv:2606.08464, 2026

  7. [7]

    Smith, and Ranjay Krishna

    Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Ranjay Krishna. Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.arXiv preprint arXiv:2406.09403, 2024

  8. [8]

    MME-CoT: Benchmarking chain-of-thought in large multimodal models for reasoning quality, robustness, and efficiency

    Dongzhi Jiang, Renrui Zhang, Ziyu Guo, Yanwei Li, Yu Qi, Xinyan Chen, Liuhui Wang, Jianhan Jin, Claire Guo, Shen Yan, Bo Zhang, Chaoyou Fu, Peng Gao, and Hongsheng Li. MME-CoT: Benchmarking chain-of-thought in large multimodal models for reasoning quality, robustness, and efficiency. InProceedings of the 42nd International Conference on Machine Learning, ...

  9. [9]

    Lawrence Zitnick, and Ross Girshick

    Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. InProceedingsofthe IEEE Conferenceon ComputerVisionand PatternRecognition, pages 1988–1997, 2017

  10. [10]

    DROID: A large-scale in-the-wild robot manipulation dataset, 2024

    Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al. DROID: A large-scale in-the-wild robot manipulation dataset, 2024

  11. [11]

    Reliable thinking with images.arXiv preprint arXiv:2602.12916, 2026

    Haobin Li, Yutong Yang, Yijie Lin, Xiang Dai, Mouxing Yang, and Xi Peng. Reliable thinking with images.arXiv preprint arXiv:2602.12916, 2026

  12. [12]

    S1-VL: Scientific multimodal reasoning model with thinking-with-images.arXivpreprintarXiv:2604.21409, 2026

    Qingxiao Li, Lifeng Xu, QingLi Wang, Yudong Bai, Mingwei Ou, Shu Hu, and Nan Xu. S1-VL: Scientific multimodal reasoning model with thinking-with-images.arXivpreprintarXiv:2604.21409, 2026

  13. [13]

    Super-CLEVR: A virtual benchmark to diagnose domain robustness in visual reasoning

    Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski, Wufei Ma, Benjamin Van Durme, and Alan Yuille. Super-CLEVR: A virtual benchmark to diagnose domain robustness in visual reasoning. InProceedings of the IEEE/CVFConferenceon ComputerVisionandPatternRecognition, 2023

  14. [14]

    TwiFF (think with future frames): A large-scale dataset for dynamic visual reasoning.arXivpreprintarXiv:2602.10675, 2026

    Junhua Liu, Zhangcheng Wang, Zhike Han, Ningli Wang, Guotao Liang, and Kun Kuang. TwiFF (think with future frames): A large-scale dataset for dynamic visual reasoning.arXivpreprintarXiv:2602.10675, 2026

  15. [15]

    Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.arXivpreprint arXiv:2603.21754, 2026

    Xu Liu, Yongheng Zhang, Qiguang Chen, Yao Li, Sheng Wang, and Libo Qin. Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.arXivpreprint arXiv:2603.21754, 2026

  16. [16]

    On the faithfulness of visual thinking: Measurement and enhancement.arXivpreprintarXiv:2510.23482, 2025

    Zujing Liu, Junwen Pan, Qi She, Yuan Gao, and Guisong Xia. On the faithfulness of visual thinking: Measurement and enhancement.arXivpreprintarXiv:2510.23482, 2025

  17. [17]

    MathVista: Evaluating mathematical reasoning of foundation models in visual contexts

    Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. MathVista: Evaluating mathematical reasoning of foundation models in visual contexts. InInternational Conferenceon Learning Representations, 2024

  18. [18]

    Prism-bench: A benchmark of puzzle-based visual tasks with cot error detection, 2025

    Yusu Qian, Cheng Wan, Chao Jia, Yinfei Yang, Qingyu Zhao, and Zhe Gan. Prism-bench: A benchmark of puzzle-based visual tasks with cot error detection, 2025. Withdrawn

  19. [19]

    V-thinker: Interactive thinking with images.arXiv preprint arXiv:2511.04460, 2025

    Runqi Qiao, Qiuna Tan, Minghan Yang, Guanting Dong, Peiqing Yang, Shiqiang Lang, Enhui Wan, Xiaowan Wang, Yida Xu, Lan Yang, Chong Sun, Chen Li, and Honggang Zhang. V-thinker: Interactive thinking with images.arXiv preprint arXiv:2511.04460, 2025

  20. [20]

    Mathcanvas: Intrinsic visual chain-of-thought for multimodal mathematical reasoning.arXiv preprintarXiv:2510.14958, 2025

    Weikang Shi, Aldrich Yu, Rongyao Fang, Houxing Ren, Ke Wang, Aojun Zhou, Changyao Tian, Xinyu Fu, Yuxuan Hu, Zimu Lu, Linjiang Huang, Si Liu, Rui Liu, and Hongsheng Li. Mathcanvas: Intrinsic visual chain-of-thought for multimodal mathematical reasoning.arXiv preprintarXiv:2510.14958, 2025. 13

  21. [21]

    Openthinkimg: Learning to think with images via visual tool reinforcement learning

    Zhaochen Su, Linjie Li, Mingyang Song, Yunzhuo Hao, Zhengyuan Yang, Jun Zhang, Guanjie Chen, Jiawei Gu, Juntao Li, Xiaoye Qu, and Yu Cheng. Openthinkimg: Learning to think with images via visual tool reinforcement learning. arXivpreprintarXiv:2505.08617, 2025

  22. [22]

    Le, and Denny Zhou

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. InAdvancesin NeuralInformation ProcessingSystems, 2022

  23. [23]

    Vic-bench: Benchmarking visual-interleaved chain-of-thought capability in mllms with free-style intermediate state representations.arXivpreprint arXiv:2505.14404, 2025

    Xuecheng Wu, Jiaxing Liu, Danlei Huang, Xiaoyu Li, Yifan Wang, Chen Chen, Liya Ma, Xuezhi Cao, and Junxiao Xue. Vic-bench: Benchmarking visual-interleaved chain-of-thought capability in mllms with free-style intermediate state representations.arXivpreprint arXiv:2505.14404, 2025

  24. [24]

    How and what to imagine? visual thinking in unified multimodal models for cross-view spatial reasoning.arXiv preprint arXiv:2605.27310, 2026

    Qian Yang, Ankur Sikarwar, Huy Le, Le Zhang, Zhuan Shi, Perouz Taslakian, and Aishwarya Agrawal. How and what to imagine? visual thinking in unified multimodal models for cross-view spatial reasoning.arXiv preprint arXiv:2605.27310, 2026

  25. [25]

    Walk the talk: Bridging the reasoning-action gap for thinking with images via multimodal agentic policy optimization.arXivpreprintarXiv:2604.06777, 2026

    Wenhao Yang, Yu Xia, Jinlong Huang, Shiyin Lu, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Yuchen Zhou, Xiaobo Xia, Yuanyu Wan, Lijun Zhang, and Tat-Seng Chua. Walk the talk: Bridging the reasoning-action gap for thinking with images via multimodal agentic policy optimization.arXivpreprintarXiv:2604.06777, 2026

  26. [26]

    MMMU: A massive multi-discipline multimodal understandingandreasoningbenchmarkforexpertAGI

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. MMMU: A massive multi-discipline multimodal understandingandreasoningbenchmarkforexpertAG...

  27. [27]

    VLABench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks, 2024

    Shiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu, Yuan Li, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu-Gang Jiang, and Xipeng Qiu. VLABench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks, 2024

  28. [28]

    Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alexander J. Smola. Multimodal chain-of-thought reasoning in language models.arXivpreprintarXiv:2302.00923, 2023

  29. [29]

    Thinking with images as continuous actions: Numerical visual chain-of-thought.arXivpreprint arXiv:2602.23959, 2026

    Kesen Zhao, Beier Zhu, Junbao Zhou, Xingyu Zhu, Zhongqi Yue, and Hanwang Zhang. Thinking with images as continuous actions: Numerical visual chain-of-thought.arXivpreprint arXiv:2602.23959, 2026

  30. [30]

    thinking with images

    Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. Deepeyes: Incentivizing “thinking with images” via reinforcement learning.arXiv preprintarXiv:2505.14362, 2025

  31. [31]

    Whenvisualizingisthefirststeptoreasoning: Mira, a benchmark for visual chain-of-thought.arXiv preprintarXiv:2511.02779, 2025

    Yiyang Zhou, Haoqin Tu, Zijun Wang, Zeyu Wang, Niklas Muennighoff, Fan Nie, Yejin Choi, James Zou, Chaorui Deng,ShenYan,HaoqiFan,CihangXie,HuaxiuYao,andQinghaoYe. Whenvisualizingisthefirststeptoreasoning: Mira, a benchmark for visual chain-of-thought.arXiv preprintarXiv:2511.02779, 2025

  32. [32]

    What, whether and how? unveiling process reward models for thinking with images reasoning.arXiv preprint arXiv:2602.08346, 2026

    Yujin Zhou, Pengcheng Wen, Jiale Chen, Boqin Yin, Han Zhu, Jiaming Ji, Juntao Dai, Chi-Min Chan, and Sirui Han. What, whether and how? unveiling process reward models for thinking with images reasoning.arXiv preprint arXiv:2602.08346, 2026. 14 Appendix Appendix Contents A Benchmark Construction Details . . . . . . . . . . . . . . . . . . . . . . . . . . ....

  33. [33]

    Never collapse multiple ideas

    **Step-by-Step Logic** Each step must represent **one atomic deduction**. Never collapse multiple ideas

  34. [34]

    stress concentration at joint

    **Visual Necessity (CRITICAL)** -> Generate a visual action **ONLY IF** it *uniquely enables understanding* that text alone cannot provide: - [YES] **Highlighting a specific region** (e.g., ‘"stress concentration at joint"‘) - [YES] **Tracing an irregular path** (e.g., ‘"blood flow through capillary"‘) - [YES] **Zooming on unreadable details** - [YES] **S...

  35. [35]

    Friction

    **Clean & Expressive Communication** - Labels **MUST** be semantic: ‘"Friction"‘, not ‘"Box[200,300]"‘ - When a visual *is* generated, **leverage enhanced styles** if they improve clarity: - ‘callout‘ for key concepts - ‘jitter‘ for organic structures - Domain-appropriate ‘theme‘ (e.g., ‘biology‘ for cells) -> But **never add decoration without purpose**....

  36. [36]

    **Analyze Context** Review ‘Problem‘, ‘Original Image‘, and ‘Previous Steps‘

  37. [37]

    **Identify Next Step** What is the *single next* logical move?

  38. [38]

    The ‘action‘ array MUST contain exactly one item

    **Determine Visual Necessity** * **For Step 1**: Always output **Text Explanation + exactly one visual action object**. The ‘action‘ array MUST contain exactly one item. Do **not** output Final Answer in Step 1. * **For Step 2 and later**: * Default to **Text Explanation only**. * Generate another visual action only when it helps inspect a new region, new...

  39. [39]

    shape":

    **Design Visual Action (If Chosen)** Select the most effective tool -- and **optionally its expressive style**: | Use Case | Recommended Tool + Style | |---------|--------------------------| | Critical concept needing emphasis | ‘annotate‘ + ‘"shape": "callout"‘ + ‘"shadow": true‘ | | Irregular structure (vessel, river) | ‘trace_highlight‘ + ‘"jitter": tr...

  40. [40]

    Generate only the single next step

  41. [41]

    Each step should make one atomic reasoning move

  42. [42]

    Reason only from the problem statement, the original image, and previous text steps

  43. [43]

    You may propose a visual action when it would help, but do not claim it has been executed

  44. [44]

    Do not say that a region is highlighted, zoomed, cropped, traced, or annotated unless that state was already present in the original image

  45. [45]

    Do not use previous proposed actions as evidence

  46. [46]

    step": N,

    For embodied control tasks, predict only the next immediate requested action. **WHEN TO PROPOSE A VISUAL ACTION** Propose a visual action only when it would clearly help locate a key region, trace a path, show a spatial relation, separate overlapping objects, or zoom into dense details. Skip visual actions for pure calculation, logical inference, summary,...

  47. [47]

    If the interference type ismodify_key, please modify key areas in the image so that it produces errors or misleading information in important aspects that affect problem solving

  48. [48]

    question

    If the interference type ismodify_non_key, please modify background or non-key information in the image, keeping the important information correct but creating interference in secondary details. Please analyze the image content, identify the key versus non-key elements for solving the problem, and apply appropriate modifications based on the interference ...

  49. [49]

    Ignore harmless formatting differences, units formatting, LaTeX wrappers, punctuation, and equivalent wording

  50. [50]

    For math, accept algebraically equivalent expressions and numerically equivalent answers when the intended quantity matches

  51. [51]

    For multiple-choice questions, accept either the correct option letter or the correct option text

  52. [52]

    If the model gives extra claims that contradict the reference, mark it incorrect

  53. [53]

    If the model answer is empty, evasive, or says it cannot answer, mark it incorrect

  54. [54]

    correct": true or false,

    Be strict about counts, directions, relations, and named entities. Return only a JSON object with: {{ "correct": true or false, "reason": "brief explanation" }} C.2 Process-Level Evaluation ForeveryVAoTtrajectory,thejudgereceivesthequestion,referenceanswer,modelfinalanswer,andcomplete trajectory. The final response is provided because Feedback Uptake requ...

  55. [55]

    If there is no effective visual operation, use null

    Key Visual Step Selection: Identify the single visual step that is most relevant to the final reasoning or most directly exposes the failure source. If there is no effective visual operation, use null

  56. [56]

    - 1: The visual action directly targets task-relevant evidence and is useful for solving the question

    Action Relevance: Does the model choose visual actions that are useful for solving the task? Score 0 / 0.5 / 1. - 1: The visual action directly targets task-relevant evidence and is useful for solving the question. - 0.5: The visual action is partially relevant, weakly targeted, or contains unnecessary but not harmful operations. - 0: The visual action is...

  57. [57]

    - 1: The rendered visual state faithfully executes the intended action

    Render Faithfulness: Are the rendered visual states faithful to the intended visual actions? Score 0 / 0.5 / 1. - 1: The rendered visual state faithfully executes the intended action. - 0.5: The rendering partially matches the action, but has noticeable ambiguity, imprecision, or minor errors. - 0: The rendering is missing, wrong, misleading, or inconsist...

  58. [58]

    key_step_id

    Feedback Uptake: Does the subsequent reasoning actually use the rendered visual states? Score 0 / 0.5 / 1. - 1: Subsequent reasoning clearly uses the rendered visual state as evidence. - 0.5: Subsequent reasoning weakly or implicitly uses the rendered state, but still relies mostly on text priors. - 0: Subsequent reasoning ignores the rendered state or co...

  59. [59]

    Action Relevance=1with an incorrect VAoT answer

  60. [60]

    Action Relevance=1and Render Faithfulness=0

  61. [61]

    Action Relevance=1and Render Faithfulness=0.5

  62. [62]

    This construction covers high-action/failed-answer cases, rendering failures, partial renders, and weak action selection across all audited models and environments

    Action Relevance in{0,0.5}. This construction covers high-action/failed-answer cases, rendering failures, partial renders, and weak action selection across all audited models and environments. For each trajectory, the annotators inspect the original image and question, the complete VAoT trajectory, the judge-selected key step, the three process scores, an...