{"id":"3f509d70-73e7-4e57-b202-7d16d08d015f","arxiv_id":"2607.04610","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Expert-curated modular Robot-VQA benchmark of 474 questions across 39 robot tasks shows SOTA VLMs have large gaps that correlate with physical robot execution.","lead":"RoboVista is a 474-question expert-annotated VQA benchmark that tests vision-language models on modular robot decisions across agriculture, surgery, industry and more. It reveals large performance gaps in current VLMs and shows that benchmark scores track real robot task success.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Physical correlation rests on a small, non-independent sample of models and tasks that may not establish RoboVista as a reliable proxy.","rationale":"The reader correctly flags representativeness of the expert questions as the weakest assumption and still issues ACCEPT because multi-pass verification and physical validation are present. That validation, however, is the load-bearing support for treating RoboVista accuracy as a proxy; once its statistical thinness and partial circularity are examined, the support is weaker than the HIGH-confidence ACCEPT implies. The rest of the paper (modular RQA construction, multi-domain coverage, thorough zero-shot/CoT/ICL ablations, failure analysis) remains solid. A CONDITIONAL verdict better matches the evidence: accept the benchmark contribution once the physical correlations are shown to be robust under held-out models and independent question sets, or after the authors explicitly qualify the proxy claim as preliminary. This is a refinement of the reader’s concern rather than a disagreement with the overall contribution.","tokens_in":29100,"tokens_out":562,"duration_ms":5495,"concrete_test":"Recompute the Fig. 7b correlations after (i) adding at least three held-out VLMs never used in RoboVista construction or ranking and (ii) reporting leave-one-model-out and bootstrap 95 % CIs on r/ρ. Separately, re-run the dVRK knot-tying sequence with a fresh set of 16 RQA questions that share no images or wording with the RoboVista surgical subset. If either the correlation loses significance or the progress ordering collapses, the proxy claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that RoboVista accuracy is a reliable proxy for real-robot decision quality rests on two physical experiments (Section VI). The bi-manual gripper alignment reports Pearson r = −0.78 / Spearman ρ = −0.93 between RoboVista score and position error, and the surgical knot-tying table shows higher RoboVista-Surgery scores associated with greater progress. Both use only a handful of models (roughly 6–7 points in Fig. 7b; three models in Table IV) that already appear in the main zero-shot ranking; the surgical questions are themselves a subset of the RoboVista surgical domain. With N this small, a single outlier or shared visual prior can dominate the correlation, and the non-independence between the benchmark questions and the closed-loop queries means the association is partly circular. The paper therefore treats a suggestive but under-powered correlation as confirmatory evidence that the 474 expert MCQs are free of linguistic/visual bias and representative of real modular decisions.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces Robot Question Answering (RQA), a modular framework that maps decision points from classical robot pipelines (perception, high-level planning, motion awareness, failure recovery) into robot-centric multiple-choice VQA instances, and instantiates it as RoboVista: 474 expert-annotated questions spanning 39 task types across agriculture, industry, domestic, surgical, driving, and open datasets. Each instance includes robot-visible or onboard imagery, five answer choices, and human rationales. Zero-shot evaluation of open- and closed-source VLMs shows substantial gaps (best overall accuracy 56.5% for Gemini 2.5 Pro). Ablations examine Chain-of-Thought and in-context learning; failure analysis separates perception from reasoning errors. Two physical experiments (bimanual gripper alignment and shared-autonomy dVRK knot-tying) report correlations between RoboVista scores and real-world spatial error / task progress.","tokens_in":29302,"tokens_out":1006,"duration_ms":17219,"significance":"If the reported gaps and the proxy relationship hold, RoboVista supplies a needed diagnostic complement to trajectory-scale embodied VQA benchmarks (Robo2VLM, RoboBrain, ERQA). Its deliberate coverage of data-scarce modular domains (surgery, agriculture, industrial deformable assembly) and expert-verified rationales are genuine strengths. The physical correlation experiments, even if under-powered, are a welcome attempt to link VQA accuracy to closed-loop execution rather than treating the benchmark as self-justifying. The work is therefore of clear interest to the robotics and VLM communities as both an evaluation resource and a design template for modular robot-centric VQA.","major_comments":[{"comment":"Section VI and Fig. 7b: The claimed strong negative correlation between RoboVista score and bimanual position error (Pearson r = −0.78, Spearman ρ = −0.93) rests on roughly 6–7 models that already appear in the main zero-shot ranking. No p-values, confidence intervals, or leave-one-out sensitivity are reported. With such small N a single outlier can dominate; the manuscript should either enlarge the model set, report statistical significance, or explicitly frame the result as exploratory rather than confirmatory of proxy validity.","section":null},{"comment":"Table IV (surgical knot-tying): Only three models are evaluated, and the 16 closed-loop queries are drawn from the same surgical domain (and RQA construction process) that contributes to the RoboVista-Surgery score. This partial non-independence weakens the claim that higher benchmark accuracy predicts greater real-task progress. The paper should either add held-out procedural stages / models or qualify the association more carefully.","section":null},{"comment":"Section IV Quality Control and Appendix A: Inter-annotator agreement (e.g., exact-match or rationale consistency) is not quantified, nor is the procedure for generating and validating the four distractors described beyond “algorithmically grounded.” Because the central claim treats accuracy on these 474 items as a reliable proxy for modular decision quality, a short quantitative reliability analysis (or explicit statement of its absence) is load-bearing and should be added.","section":null}],"minor_comments":[{"comment":"Table I vs. Table II: GPT-5 overall accuracy is listed as 48.1% (zero-shot) in Table I but 55.5% in Table II; clarify whether different checkpoints, decoding, or subset filtering explain the discrepancy.","section":null},{"comment":"Fig. 4 table header says “39 unique task types” while Appendix Table V reports “Unique Tasks 33”; reconcile the counts.","section":null},{"comment":"Section V decoding: main text states temperature 0.7; Appendix B states temperature 0.0 (greedy). Align the reported protocol.","section":null},{"comment":"Figs. 5 and 6 (and Appendix Figs. 8–9) appear to reuse nearly identical failure-analysis diagrams with swapped model names; ensure captions and percentages match the intended model.","section":null},{"comment":"Abstract and Introduction claim “strong correlation”; given the sample-size caveats above, softer language (“suggestive association”) would better match the evidence presented.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The physical-validation section is the weakest link but is presented honestly enough that minor revision (stats + caveats + IAA) should suffice; I would not require a full new physical campaign. The benchmark itself is carefully constructed and fills a real gap. Fit for a robotics or multimodal-evaluation venue is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is that this is a carefully built, multi-domain modular VQA set (474 expert MCQs with reasoning traces) aimed at the decision points that actual robot pipelines use, not just end-to-end trajectories. Best closed model hits only 56.5 % overall, and they show a clear negative correlation between RoboVista score and real gripper alignment error plus better progress on dVRK knot-tying.\n\nWhat is new is the RQA framing: they take the classic modular stack (perception, high-level decision, motion awareness, recovery), turn each decision into a grounded multiple-choice question with human rationale, and deliberately pull from agriculture, surgery, industrial deformable work, and driving where public trajectory data is scarce. That is a genuine complement to Robo2VLM / RoboBrain / ERQA, which mostly mine imitation datasets. The zero-shot / CoT / ICL tables are thorough across model families, the failure breakdown cleanly separates misidentification from spatial/semantic errors, and the physical experiments are concrete (Pearson r = −0.78 on position error, stage-wise knot progress).\n\nSoft spots are real but proportionate. The physical correlations use only a handful of models that already appear in the main ranking, and the surgical questions are a subset of the benchmark itself, so the proxy claim is under-powered and partly non-independent. Construction is expert-heavy (PhD-level annotators, multi-pass review), which keeps quality high but makes scaling and full reproducibility harder; they own this in the limitations. Scale is modest by design, not a bug. No free parameters or circular scoring games. Citations look standard and fair.\n\nThis is for anyone evaluating or fine-tuning VLMs for modular robot systems, especially outside tabletop manipulation. It is solid empirical infrastructure, not a theory paper. I would bring it to reading group, cite the benchmark and the correlation numbers when I need a multi-domain robot VQA reference, and a serious editor should send it to referees rather than desk-reject. Engage with it.","headline":"Useful modular VQA benchmark that fills a real gap for robot VLMs; physical correlations are suggestive but rest on a thin sample.","tokens_in":29962,"tokens_out":513,"would_cite":true,"duration_ms":9514,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Today’s best vision-language models still miss nearly half of the modular decisions real robots must make, and those misses track real physical failure.","keywords":["vision-language models","robot question answering","RoboVista","modular robotics","visual question answering","embodied reasoning","surgical robotics","benchmark"],"falsifier":"Run the same physical bimanual distance-alignment and dVRK knot-tying protocols with a model whose RoboVista score has been deliberately raised (or lowered) by training or prompting, and check whether physical error and task progress move in the predicted direction; if the correlation collapses, the central claim fails.","tokens_in":29992,"feed_emoji":"🤖","tokens_out":658,"duration_ms":9771,"temperature":0.7,"pith_summary":"Robots in factories, farms, homes, and operating rooms do not act as one big black box; they chain many small, modular decisions—what is visible, what to grasp next, whether a motion is safe, how to recover. This paper claims that if vision-language models are to become the common reasoning layer for such systems, we must test those modular decisions directly rather than only scoring end-to-end teleoperated trajectories. The authors therefore introduce Robot Question Answering (RQA), a way to turn real robot pipelines into expert-verified multiple-choice visual questions, and release RoboVista, 474 such questions spanning 39 task types and six application domains. On this benchmark even the strongest models top out near 56 percent accuracy, and the same models’ RoboVista scores strongly predict how well they estimate distances and guide surgical knot-tying on physical hardware. A sympathetic reader cares because the gap is now measurable, domain-by-domain, and because the physical correlation suggests that fixing RoboVista failures would improve actual robots.","feed_headline":"Best VLMs still miss half of real robot decisions","feed_subtitle":"A 474-question modular benchmark tracks physical error and surgical task progress","key_machinery":"Robot Question Answering (RQA): a module-level abstraction that maps each functional block of a robot pipeline—perception, high-level decision making, motion awareness, failure recovery—into a robot-centric visual question with a single verified answer and human rationale, then assembles those questions into the RoboVista benchmark.","core_discovery":"State-of-the-art vision-language models exhibit large, persistent performance gaps on modular robot decision points (best overall accuracy 56.5 percent), and a model’s RoboVista score correlates strongly with real-world spatial estimation error and with progress on closed-loop surgical knot-tying under shared autonomy.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Top VLMs reach just 56.5% on modular robot decisions","RoboVista shows VLMs miss half of real robot choices","SOTA VLMs lag on diverse robot reasoning tasks","Best VLMs correlate poorly with robot task progress","VLMs leave large gaps in agricultural to surgical robots"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The 474 expert-built multiple-choice questions, each tied to one modular decision taken from published robot systems, are representative enough of real decision quality that accuracy on them can stand in for how well a model would help an actual robot.","fun_headline_variants_meta":{"raw":{"variants":["Top VLMs reach just 56.5% on modular robot decisions","RoboVista shows VLMs miss half of real robot choices","SOTA VLMs lag on diverse robot reasoning tasks","Best VLMs correlate poorly with robot task progress","VLMs leave large gaps in agricultural to surgical robots"]},"model":"grok-4.5","effort":"low","cost_usd":0.004552,"raw_usage":{"total_tokens":1268,"prompt_tokens":716,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":45520000,"prompt_tokens_details":{"text_tokens":716,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":488,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":716,"tokens_out":64,"duration_ms":3936,"temperature":1.0,"reasoning_tokens":488,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T16:27:22.563115+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same physical bimanual distance-alignment and dVRK knot-tying protocols with a model whose RoboVista score has been deliberately raised (or lowered) by training or prompting, and check whether physical error and task progress move in the predicted direction; if the correlation collapses, the central claim fails.","supporting_citations":[],"review_version":1}