{"id":"f4d8e486-5a1d-4138-a7bb-07a4f39f605d","arxiv_id":"2505.16770","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The RBench-V benchmark finds that the best current AI models score 25.8%, far below 82.3% for humans, on visual reasoning problems claimed to require multi-modal outputs.","lead":"RBench-V is a new benchmark of 803 visual reasoning problems that, according to its authors, require drawing or editing images to solve. On it the best AI model, OpenAI o3, scored 25.8% versus 82.3% for human experts, showing a large gap in visual thinking that involves multi-modal outputs.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark's core construct—that solving requires multi-modal output—is unverified because the evaluation never requires image generation; the observed gap could reflect text-only reasoning failure.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern I see: the benchmark never forces multi-modal output, so its accuracy numbers cannot support the paper's central claim about multi-modal generation capability. The paper's own Sec. 4.4 admits that math questions can be solved via text-only algebraic shortcuts, and the authors therefore report results 'w/o math' as a better indicator. This directly concedes that the benchmark does not enforce the target construct for an entire subject area. The same risk applies to the other categories: no experiment shows that drawing is necessary or that models fail because they cannot draw. The LLM-as-a-judge protocol (GPT-4o) adds a further weakness, but it is secondary to construct validity. I therefore agree with the reader's REJECT verdict: the benchmark may be a challenging visual reasoning test, but the specific claim about multi-modal output ability is not supported by the evidence presented. A forced-drawing condition would directly settle whether the measured gap is about output modality or about reasoning more generally.","tokens_in":12722,"tokens_out":3691,"duration_ms":32613,"concrete_test":"Run a forced-drawing evaluation on all 803 questions: after receiving each question, require the model to first output an annotated image (draw auxiliary lines, connect dots, trace paths) or a structured drawing command, then provide the final answer. Score the answer only from the drawing with an image-independent checker, and compare accuracy with the current text-answer-only protocol. If accuracy under forced-drawing is not significantly higher, the paper's claim that models fail because they cannot generate multi-modal outputs is not supported. As a secondary check, have human raters solve a random sample of 100 text-only versions to test whether drawing is truly indispensable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that RBench-V measures 'vision-indispensable reasoning' requiring multi-modal outputs, and that current models fail because they cannot draw during reasoning. For this claim to hold, two things must be true: (1) the questions cannot be solved without generating or modifying images, and (2) the evaluation protocol actually tests that ability. Neither is established. The protocol in Sec. 4 records only a final text answer, judged by GPT-4o; no model is required or incentived to emit an image, and no image output is scored. The paper's own evidence undercuts indispensability: Fig. 4 shows o3 solving the geometry question via coordinate algebra, and Sec. 4.4 explicitly calls this a 'multi-modal reasoning shortcut,' prompting the authors to report results 'w/o math.' If math can be shortcut, the benchmark does not enforce multi-modal output. For the remaining categories, there is no systematic demonstration that drawing is necessary; Table 2's 'win rates' only show that raters judge RBench-V items as more drawing-intensive than MMLU/MMMU on 30 samples, not that drawing is required for correct answers. The reported failures (e.g., o3 describing dots instead of connecting them) are scored wrong, but this is consistent with the model failing at visual understanding or textual reasoning, not specifically with an inability to output drawings. Thus the central claim—that models cannot draw as a thinking tool—is unsupported by the measured accuracy. This is a construct-validity problem: the benchmark could be a hard visual-reasoning test, but the results do not diagnose the modality of the failure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces RBench-V, a benchmark of 803 hand-curated questions across mathematics, physics, counting, and games, intended to assess what the authors call 'vision-indispensable reasoning' that requires multi-modal outputs, such as drawing auxiliary lines, tracing paths, or constructing images. The authors evaluate a broad range of open- and closed-source models, including o3, Gemini 2.5 Pro, GPT-4o, and Qwen2.5-VL, and report that the best model, o3, achieves 25.8% accuracy versus 82.3% for human experts. They also report a 'w/o math' score to control for algebraic shortcuts and argue that current models fail because they cannot generate or manipulate images during reasoning. Data and code are released publicly.","tokens_in":12911,"tokens_out":5805,"duration_ms":51087,"significance":"If the construct were properly validated, RBench-V would fill a genuine gap by evaluating multi-modal chain-of-thought output rather than input understanding alone. The paper's strengths include a broad model suite, a human baseline, public data and code, and an honest acknowledgment that math items can be solved via text-only algebraic reasoning. However, the evaluation protocol never requires or scores an image output, and the judge is GPT-4o, which is itself one of the evaluated models. These issues mean that the headline accuracy numbers are informative measurements of final-answer performance, but they do not yet support the paper's specific central claim about multi-modal output capability or about models' inability to use drawing as a thinking tool.","major_comments":[{"comment":"The central claim that RBench-V measures multi-modal output reasoning is not supported by the evaluation protocol, because models are never required to emit an image and no image output is scored. Section 4 states that a unified LLM-as-a-Judge framework with GPT-4o is used and Top-1 accuracy is reported, which judges only the final text answer. The connect-the-dots example in Section 4.3, where o3 describes the diagram instead of connecting the dots, is scored incorrect, but this outcome is equally consistent with a failure in visual understanding or in text-only reasoning. To support the central claim, the protocol should elicit and score image outputs, or otherwise demonstrate per-item that text-only reasoning cannot produce the correct answer.","section":"Section 4 (evaluation protocol)"},{"comment":"GPT-4o is used as the judge while also being one of the evaluated models, and no human agreement is reported. This creates an uncontrolled potential bias in the model rankings and in the human-versus-model comparison. Please report inter-rater agreement on a sample (e.g., Cohen's kappa) against human graders, and consider using a judge model that is not in the evaluated set, or human grading for open-ended items. In addition, Table 3 reports no confidence intervals or significance tests; several adjacent scores, such as 10.0 and 10.6 for InternVL-3-38B and Qwen2.5VL-72B, are likely within sampling noise, so the ranking claims are not statistically supported.","section":"Section 4 (LLM-as-a-Judge)"},{"comment":"The evidence that drawing is necessary for solving RBench-V items is limited to the design principle and to Table 2's win rates, which are based on only 30 sampled items per benchmark and on subjective ratings by models and experts. High win rates show that RBench-V items are judged as more drawing-intensive than MMLU or MMMU items, not that drawing is required for a correct answer. Moreover, Section 4.4 concedes that math items can be solved by coordinate algebra and therefore reports scores 'w/o math,' but no analogous check is provided for counting, physics, or games. Without category-level controls, such as comparing a text-only reasoning condition against a drawing-allowed condition, the phrase 'vision-indispensable' is not established.","section":"Section 3.1 and Section 4.4 (construct validity)"},{"comment":"The 'w/o math' result is a post-hoc exclusion of 176 of 803 questions, and the headline 25.8% accuracy therefore mixes items with and without the alleged multi-modal requirement. If RBench-V is meant to assess multi-modal output, the math category as currently designed fails for models that can use algebraic shortcuts. The paper should either redesign math items to require geometric construction, or present RBench-V as a mixed benchmark that measures both multi-modal and text-only strategies, with the 'w/o math' score as the primary evidence for the specific multi-modal claim. As written, the 25.8% figure is not a clean measure of multi-modal output ability.","section":"Section 4.3 and Table 3 (math shortcut)"},{"comment":"The human expert score of 82.3% lacks essential reporting details: the number of experts, the number of questions assigned per expert, time limits, whether drawing was allowed or required, and the grading protocol are not specified. Without this information, the headline human-model gap cannot be interpreted as a comparison of multi-modal output capability. The authors should also state whether the same rubric and judge were used for human answers and model answers, since differences in grading criteria could account for part of the gap.","section":"Section 4.2 and Table 3 (human baseline)"}],"minor_comments":[{"comment":"In the caption of Figure 2, 'Rench' should be 'R-Bench'.","section":"Figure 2 caption"},{"comment":"There are several typos: 'vision-indisperential' in the Section 1 bullet list should be 'vision-indispensable', and 'humam' in Section 4 should be 'human'.","section":"Abstract and Section 1"},{"comment":"The deployment details for vLLM and VLMEvalKit are not given (versions, batch sizes, or any non-default generation parameters), which limits reproducibility of the open-source model numbers.","section":"Section 4.2"},{"comment":"The statement that RBench-V includes 40 text-only questions sits awkwardly with the 'vision-indispensable' framing; please clarify whether these items can be solved without visual input and whether they are included in the 'w/o math' analysis.","section":"Section 3.2"},{"comment":"The red lines in Figure 3 illustrate an idealized reasoning trace rather than an actual model output; the caption should make this explicit to avoid implying that any evaluated model produced those drawings.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The central issue is construct validity: the benchmark's claim to measure multi-modal output capability is not supported by the current protocol, which scores only final text answers. This is fixable with additional experiments (an image-output condition, human-judge agreement, and per-category text-only baselines), so I recommend major revision rather than rejection. The authors' self-citation of their companion R-Bench paper appears contextual and acceptable. The released data and code are a strength and would facilitate the requested validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know about RBench-V is that it is a real attempt to measure something nobody else is measuring: reasoning that requires multi-modal outputs, like drawing auxiliary lines or tracing paths. The 803-question set is hand-curated, covers math, physics, counting, and games, and the paper reports a large model evaluation including o3, Gemini 2.5 Pro, and many open models. The headline result—o3 at 25.8% vs. human experts at 82.3%—is striking, and the authors deserve credit for also reporting a 'w/o math' score, since they candidly show that models often solve geometry by coordinate algebra rather than drawing. That honesty is a real strength.\n\nThe main soft spot is construct validity. The paper claims these questions require multi-modal output, but the evaluation protocol never requires a model to emit an image. It only records final text answers, judged by GPT-4o. So the measured gap could be due to failures in visual understanding or text-based reasoning, not specifically an inability to draw. The authors' own example of o3 describing dots instead of connecting them is scored as wrong, but that is consistent with a model that simply couldn't do the visual reasoning at all. The Table 2 'win rates' comparing drawing-necessity against MMMU/MMLU are based on only 30 samples per benchmark and are not a demonstration that drawing is required for correct answers.\n\nA second concern is that GPT-4o serves as both judge and one of the evaluated models, with no reported agreement against human grading. That is a straightforward reliability gap. Confidence intervals and significance tests are also absent, though for a benchmark paper this is a minor issue if the data is released (it is).\n\nStill, I would not reject this out of hand. The benchmark is new, the data is public, and the gap is large enough to matter even if the interpretation is not yet proven. The paper would benefit from a protocol where models can optionally emit images and where success is scored on the image itself, plus a small human-judge validation of the GPT-4o scoring. As is, the contribution is a potentially useful challenging visual reasoning benchmark, not yet a validated measure of multi-modal output capability.\n\nI would send this to peer review. A good referee can push the authors to tighten the protocol and soften the claims. It is the kind of paper that the community will cite regardless, so better to review it seriously than desk-reject.\n\nBest,\n\n[Your name]","headline":"A genuinely new benchmark for multi-modal output reasoning with a striking model-human gap, but the central construct is not yet validated because the evaluation protocol never actually requires or scores image generation.","tokens_in":13629,"tokens_out":1094,"would_cite":true,"duration_ms":10770,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper's benchmark RBench-V asks models to draw while reasoning; the best model scores 25.8 percent to humans' 82.3 percent.","keywords":["RBench-V","multi-modal chain-of-thought","visual reasoning benchmark","omni-model evaluation","image generation for reasoning","LLM-as-a-judge","vision-language models"],"falsifier":"A direct test would be to run o3 and the other top models on RBench-V with all image-generation ability disabled while still allowing the input image to be viewed, then compare accuracy with the reported 25.8 percent. If a text-only-reasoning model scores the same, or if a version of o3 forbidden from drawing shows no drop, then the benchmark is not actually measuring multi-modal output reasoning.","tokens_in":12461,"feed_emoji":"📐","tokens_out":7658,"duration_ms":56946,"temperature":0.7,"pith_summary":"RBench-V is a new benchmark for testing whether multimodal models can reason by producing images, not just by reading them. The paper hand-picks 803 problems in math, physics, counting, and games whose solutions, it argues, require drawing auxiliary lines, tracing paths, or otherwise modifying or creating visual content during the thinking process. Across dozens of open- and closed-source models, the best performer, o3, reaches 25.8 percent accuracy while human experts reach 82.3 percent. The paper takes this gap as evidence that current models cannot yet use drawing as a genuine tool for visual reasoning, and that scaling model size, omni-modal decoding, and long text-only chain-of-thought do not fix the deficiency.","feed_headline":"Best AI scores 25.8% on drawing-based reasoning test vs 82.3% human","feed_subtitle":"New benchmark demands models create or edit images while solving; o3 trails human experts by 56 points.","key_machinery":"The load-bearing object is the benchmark itself: 803 hand-selected question-answer pairs spanning math (176), physics (157), counting (195), and games (275), of which 763 have multimodal inputs and 40 are text-only, split into 356 multiple-choice and 447 open-ended items. The selection criterion is that solving a question should require producing new visual content—drawing a geometric figure, adding auxiliary lines, tracing a trajectory, connecting dots—rather than merely interpreting the input. The evaluation protocol measures reasoning through multi-modal outputs by using a unified LLM-as-a-judge framework with GPT-4o scoring each model's final top-1 text answer.","core_discovery":"The paper's central claim is that RBench-V measures a capability existing benchmarks miss: vision-indispensable reasoning with multi-modal outputs, or multi-modal chain-of-thought. While MMLU and MMMU supply multimodal inputs and demand text answers, RBench-V questions are designed so that the solver must generate novel images, construct auxiliary lines, trace light rays or maze paths, or mark counted objects en route to an answer. The evaluation reports a decisive gap: the best model, o3, scores 25.8 percent overall, far below the human expert score of 82.3 percent, and the best open-source model scores 10.6 percent. The paper also finds that larger models, omni-models with joint text-image decoding, and long text-only reasoning models show little or no improvement on the benchmark, and it documents cases where o3 solves geometry algebraically rather than by drawing, a route it calls a 'multi-modal reasoning shortcut'. It concludes that current foundation models struggle to generate and integrate multi-modal outputs in visual thinking.","pith_inferences":["A stricter protocol that requires the model to actually emit an image (an annotated drawing, a traced path) before the final answer would test whether the low scores come from an inability to draw or from a failure to engage the drawing modality; the current protocol judges only final text.","One testable extension is to give models an external drawing tool and allow them to feed their own generated image back as input; if accuracy rises sharply, the bottleneck is generating the visual step rather than visual perception itself.","Because the paper identifies algebraic shortcuts in math, a focused pure-drawing subset could be repurposed as a training signal: reinforcement learning that rewards correctly drawn auxiliary structures might close the gap faster than scale, which the paper's own scaling results predict.","The same question-design principle could extend beyond still images to video or audio outputs, which the paper names as future work."],"forward_implications":["If RBench-V measures what it claims, omni-models that decode images as well as text do not automatically gain visual reasoning ability: Qwen2.5VL-7B and Qwen2.5-Omni-7B score at similar levels, as do MiniCPM-V-2.6 and MiniCPM-o-2.6.","Scaling model size is not sufficient: raising Qwen2.5VL from 7B to 72B, or scaling InternVL and LLaVA-OneVision similarly, produces no clear accuracy gain on RBench-V.","Long text-only chain-of-thought models barely outperform their non-thinking counterparts, so the missing capability lies in generating multimodal outputs rather than in spending more tokens on text reasoning.","Reporting accuracy after removing math questions, where models can substitute algebra for drawing, widens the human-model gap and is proposed as a cleaner signal of true multimodal reasoning.","o3's large lead over prior models is read as evidence both that RBench-V tracks genuine progress and that the field remains far from human-level visual reasoning."],"supporting_citations":[{"why":"Defines the text-only MMLU benchmark that RBench-V contrasts against, establishing the input-oriented evaluation tradition.","marker":"Hendrycks et al. [2021]"},{"why":"Defines MMMU, the multimodal-input/text-output benchmark that RBench-V positions itself against.","marker":"Yue et al. [2024a]"},{"why":"The o3 model is the best performer evaluated; its 25.8 percent score anchors the paper's central claim about the gap to humans.","marker":"OpenAI [2025c]"},{"why":"The Qwen2.5-VL family is the main open-source scaling case study, used to argue that model size alone does not close the gap.","marker":"Bai et al. [2025]"},{"why":"GPT-4o serves both as the LLM-as-a-judge evaluator and as a representative omni-model baselined on the benchmark.","marker":"OpenAI [2024a]"},{"why":"Supplies the cognitive-science grounding that drawing is a central human reasoning tool, motivating the benchmark's design.","marker":"Fan et al. [2023]"},{"why":"Survey of multimodal chain-of-thought that frames the M-CoT concept RBench-V is designed to measure.","marker":"Wang et al. [2025]"}],"fun_headline_variants":["AI trails humans by 56 points on visual reasoning benchmark","New benchmark requires AI to draw; top model scores 25.8%","o3 scores 25.8% on RBench-V, humans 82.3%: visual reasoning gap","Drawing-based reasoning test: AI 25.8% vs human 82.3%","Humans outscore top AI 82.3% to 25.8% on visual CoT benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's validity depends on the 803 questions truly being solvable only by generating or modifying an image during reasoning, yet the evaluation checks only final text answers and never verifies that any image was produced.","fun_headline_variants_meta":{"raw":{"variants":["AI trails humans by 56 points on visual reasoning benchmark","New benchmark requires AI to draw; top model scores 25.8%","o3 scores 25.8% on RBench-V, humans 82.3%: visual reasoning gap","Drawing-based reasoning test: AI 25.8% vs human 82.3%","Humans outscore top AI 82.3% to 25.8% on visual CoT benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001242,"raw_usage":{"total_tokens":5141,"prompt_tokens":1033,"completion_tokens":4108,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":3993}},"tokens_in":649,"tokens_out":4108,"duration_ms":17934,"temperature":1.0,"reasoning_tokens":3993,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:54:54.886912+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to run o3 and the other top models on RBench-V with all image-generation ability disabled while still allowing the input image to be viewed, then compare accuracy with the reported 25.8 percent. If a text-only-reasoning model scores the same, or if a version of o3 forbidden from drawing shows no drop, then the benchmark is not actually measuring multi-modal output reasoning.","supporting_citations":[],"review_version":1}