{"id":"0d0ad556-9ea0-42ea-83ac-5b90d6f6558b","arxiv_id":"2607.17454","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Cross-view depth consistency of predicted robot futures selects better rollouts and a cheap action-future gate decides when to sample, improving success on RoboCasa, LIBERO Long, and RoboTwin 2.0.","lead":"A new method, Gated GeoBoN, uses the imagined-future frames generated by a robot policy to decide when extra compute is worth spending, improving task success on manipulation benchmarks while triggering additional sampling only a quarter of the time. This matters because it shows that test-time scaling for robots can be selective, label-free, and based on geometric self-consistency rather than learned reward models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"VGGT-Ω's reprojection score is never validated on WAM-generated frames; without a negative-control test, the ranking may select VGGT artifacts rather than true cross-view consistency.","rationale":"The reader's weakest assumption is plausible, but the more precise and testable form is that Eq. (5) lacks any validation that VGGT-Ω's depth and pose estimates on WAM generated images behave as a consistency measure. The method's success could be real or could be an artifact of VGGT's own preferences. The proposed negative control directly probes the causal mechanism of the selector: if known-inconsistent pairs do not score higher, then the ranking signal is not cross-view geometry. This is the single most load-bearing step because every downstream result—fixed-budget gains, gating gains, comparison to value heads—depends on it. I agree with the reader's weakest_assumption and do not move the verdict; the condition is exactly that the authors supply this validation or an equivalent.","tokens_in":10070,"tokens_out":10830,"duration_ms":112376,"concrete_test":"On the fixed N=8 candidate dumps used in §4.4, construct negative-control pairs by replacing each candidate's predicted wrist view with the wrist view of another rollout at the same decision point (and, as an unambiguous control, with wrist views from a different decision point or task). These pairs are known to be cross-view inconsistent because the two frames come from different imagined futures. Recompute edepth for the original and mismatched pairs. If the mismatched-pair error is not separated from the original-pair distribution by more than the §4.5 duplicate-pair variation threshold on the large majority of cases, then Eq. (5) is not measuring cross-view consistency on WAM frames, and the central claim fails. If the separation is clear, the VGGT OOD objection is substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. (5) is computed wholly under VGGT-Ω's estimated depth and camera geometry. The paper supplies no evidence that this frozen model is calibrated on WAM-predicted frames, and its own §4.5/Fig. 2 shows the score is noisy: false low-score selections rise from 7% (N=2) to 32% (N=16). Because the closed-loop gains are modest (e.g., +2.1 points on RoboCasa/Cosmos; −4.6 points on Door/Drawer X-WAM), a systematic VGGT bias on synthetic diffusion outputs could easily produce the observed ranking without corresponding to true 3D consistency. The X-WAM rows, which use the backbone's own depth head, cannot validate the VGGT branch. Since the score is used only to compare candidates, correlated VGGT errors across candidates would invalidate the Best-of-N selection claim even if average accuracy on real frames is high.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Gated GeoBoN, a training-free test-time scaling method for World Action Models (WAMs). A cheap action–future consistency gate decides whether to sample additional rollouts; when it triggers, a frozen VGGT-Ω geometry model ranks candidates by cross-view depth reprojection inconsistency (Eq. 5). Fixed-budget GeoBoN is reported to improve N=8 task success in all five benchmark–backbone settings (e.g., RoboCasa/Cosmos 66.3→68.4; RoboCasa/X-WAM 80.8→82.5), and gating recovers on average 74.8% of the always-on gain while triggering on 26.2% of decision points. Additional offline diagnostics compare selectors and gate against alternatives, and identify false low-score selections as a failure mode.","tokens_in":10339,"tokens_out":7917,"duration_ms":80441,"significance":"If the mechanism is correct, this is a useful contribution: it shows that WAM-exposed future frames can support task-label-free, training-free rollout selection and selective compute allocation, with closed-loop evidence across multiple backbones and benchmarks, and it compares favorably to a learned value head. The strengths are the multi-setting closed-loop evaluation, the comparison against future-consensus and value-head baselines, and the attempt to isolate the gate and selector contributions. However, the central geometric evaluator is not validated on the distribution it is applied to (WAM-generated frames), and the false-selection analysis is defined in a way that may be circular; these issues must be resolved before the results can be fully trusted.","major_comments":[{"comment":"The cross-view depth reprojection score is computed entirely under VGGT-Ω's estimated depth and camera pose on predicted future frames. The paper provides no evidence that VGGT-Ω is calibrated on WAM-generated frames; the X-WAM rows use the backbone's native depth head, so they do not validate the VGGT branch. Since the score is used only to compare candidates, correlated VGGT errors across candidates could produce the observed rankings without true geometric consistency. Fig. 2/§4.5 shows false low-score selections increasing from 7% (N=2) to 32% (N=16), confirming that the score is noisy on predicted frames. Given the modest closed-loop gains (RoboCasa/Cosmos +2.1; X-WAM Door/Drawer −4.6), a negative-control experiment is required: e.g., compare VGGT reprojection scores on predicted frames against simulator ground-truth depth/pose, or show that low-score selections correspond to genuin","section":"§3.2, Eq. (5); §4.5/Fig. 2"},{"comment":"The false low-score selection analysis defines 'false' via LPIPS near-duplicates below 0.02 and a duplicate-pair variation threshold. This is circular: a low reprojection score could be a true signal of geometric consistency even if images are perceptually near-identical, and LPIPS < 0.02 is an extremely tight threshold. The analysis also lacks counts of near-duplicate pairs, error bars, and comparison against a ground-truth consistency label. As this analysis is used to explain saturation/degradation at large N, please either define false selections with respect to ground-truth geometry/task outcome, or present it as a heuristic with appropriate caveats and statistical detail.","section":"§4.5/Fig. 2"},{"comment":"All thresholds are fixed, but no sensitivity analysis is reported. The gate threshold τ_gate = −0.2 is particularly surprising: the gate triggers only when the cosine between optical flow and projected action motion is strongly negative, whereas an inconsistency detector would be expected to trigger at low positive or near-zero values. The trigger rates (14.2–34.5%) and the claimed compute savings depend directly on this choice. Please report a threshold sweep on a held-out set, or justify the choice by a principled argument, to show that Gated GeoBoN's benefits are not tied to a single fragile hyperparameter.","section":"§3.1/§4.1"},{"comment":"The selector comparison lacks confidence intervals for online success and error recovery. In RoboTwin/Motus, GeoBoN (89.9%) is actually worse than Consensus (90.5%), and in RoboCasa/Cosmos, VGGT Conf (65.9%) is below baseline (66.3%). The phrase 'most consistent selector' needs statistical support (e.g., paired confidence intervals over seeds/tasks); otherwise, the offline ER advantage may not transfer to closed-loop settings.","section":"Table 3/§4.4"}],"minor_comments":[{"comment":"Clarify how the group average and its confidence interval are computed (e.g., paired by seed/task), and report the number of tasks and seeds per setting.","section":"Table 1"},{"comment":"Add error bars and sample sizes, and define the denominator for the false-selection rate.","section":"Fig. 2"},{"comment":"Typo: 'WA V' should be 'WAV' (World Action Verifier).","section":"Related Work"},{"comment":"Per-task improvements are often within error bars; emphasize the group-level results to avoid overclaiming category-level gains.","section":"§4.2"},{"comment":"Consider releasing evaluation scripts and a reproducibility statement to support the 'training-free' and 'zero-shot' claims.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is promising and likely publishable after major revisions. The main risk is the lack of validation of VGGT-Ω on WAM-generated frames; a negative-control experiment with simulator ground truth would substantially increase confidence. The authors should also clarify the statistical basis for the selector comparison and provide sensitivity analysis for the gate threshold. No code is provided, which hinders reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, honest paper that introduces a sensible two-stage method—a cheap action-future gate decides when to sample more rollouts, and a cross-view depth reprojection consistency score from a frozen VGGT-Ω ranks them. The gains are real but small: about 1-3 points at N=8 across five settings, with group-level improvements outside the confidence intervals. The paper also does a fair job of diagnosing why performance saturates or drops at N=16: false low-score selections from evaluator noise.\n\nWhat's new: using cross-view geometric consistency of predicted futures as a task-label-free selector for WAM rollouts, and combining it with a gate that avoids paying the geometric cost when the initial rollout looks fine. The comparison against a model-specific value head is a nice touch—GeoBoN beats it without training.\n\nThe soft spots are proportionate. The closed-loop gains are modest, and per-task results are often within error bars; the X-WAM Door/Drawer row actually degrades by 4.6 points. The paper is upfront about this. The bigger concern is that Eq. (5) is computed entirely under VGGT-Ω's estimated depth and camera poses, and there's no validation that VGGT-Ω is calibrated on WAM-generated synthetic frames. If its errors are correlated across candidates, the ranking could pick a biased outlier rather than the true best rollout. The X-WAM experiments use the backbone's own depth head, so they don't validate the VGGT branch. The paper's own Fig. 2 shows false selections rising from 7% at N=2 to 32% at N=16, which is consistent with this worry. Still, the method's success across multiple backbones and benchmarks suggests the signal is at least correlated with something useful. A negative-control test—checking whether VGGT-Ω's reprojection error on real vs. predicted frames matches—would settle this.\n\nThe paper does not ship code or data, and the fixed thresholds have no sensitivity analysis. That's limiting but not disqualifying.\n\nBottom line: this is a well-executed, honest contribution that deserves a serious referee. It's not a paradigm shift, but it's a useful step for WAM inference, and the failure-mode analysis is more transparent than most. I'd bring it to a reading group and would cite it if I worked on test-time scaling for robot policies. Recommend peer review, with the calibration question flagged as a required addition.","headline":"Useful training-free test-time scaling for WAMs with modest but consistent gains; the main open question is whether the frozen geometry model's scores are trustworthy on synthetic frames.","tokens_in":10785,"tokens_out":2230,"would_cite":true,"duration_ms":21657,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"World Action Models predict actions plus imagined futures; this paper claims cross-view depth reprojection consistency from a frozen geometry model ranks those futures, enabling label-free robot rollout selection.","keywords":["test-time scaling","world action models","robot manipulation","best-of-N selection","cross-view depth reprojection","geometric consistency","action–future consistency gate","training-free verification"],"falsifier":"Take a fixed set of RoboCasa tasks, sample N=8 rollouts per decision, run every candidate in simulation, and check whether the candidate with the lowest Eq. 5 reprojection error succeeds more often than random or value-head picks; if low reprojection error does not track simulation success on the WAM's predicted frames, the central claim fails. This can be tested offline from the paper's candidate dumps.","tokens_in":9998,"feed_emoji":"🤖","tokens_out":9066,"duration_ms":78078,"temperature":0.7,"pith_summary":"Robot policies called World Action Models (WAMs) generate both an action chunk and a predicted future video at each decision point. This paper tries to establish that the imagined future itself is usable as a test-time selection signal: a cheap action–future consistency gate decides when the initial rollout looks internally inconsistent, and when it does, a frozen geometry model ranks additional rollouts by cross-view depth reprojection inconsistency—how poorly the wrist-camera predicted depth map, projected into the primary camera, matches the primary view's predicted depth. The claim is that this training-free, task-label-free check improves fixed-budget Best-of-N success in every tested setting—for example, RoboCasa group average from 66.3% to 68.4% with Cosmos Policy—and that gating recovers about 74.8% of the always-on gain while sampling extra candidates at only 26.2% of decision points. A sympathetic reader would care because it suggests WAM inference can be scaled selectively using only the model's own multi-view predictions, without learned verifiers or environment feedback.","feed_headline":"Cross-camera agreement picks better robot rollouts","feed_subtitle":"A frozen depth model scores imagined futures; the best-scoring rollout wins, and a cheap gate avoids needless sampling.","key_machinery":"The load-bearing object is the cross-view depth reprojection inconsistency score (Eq. 5): for each sampled rollout, feed the predicted primary and wrist frames into the frozen geometry model VGGT-Ω, then average |log(d_proj/d_vggt)| over pixels where the projected wrist points land in the primary view with positive depth and sufficient model confidence. It converts 'do these two imagined views depict the same 3D scene?' into a single number. The action–future gate (Section 3.1) computes cosine agreement between dense optical flow in a capsule around the projected end-effector trajectory and the motion implied by the generated action chunk; a low agreement triggers additional sampling. Both s","core_discovery":"The central discovery is that cross-view depth reprojection inconsistency is a strong task-label-free selector for WAM rollouts. The paper defines the score as the average absolute log-ratio between the depth predicted directly for the primary view and the wrist-view depth map reprojected into the primary camera, using a frozen geometry model (VGGT-Ω) for depth and pose; lower scores mean the two predicted views correspond to one coherent 3D scene. Around this score the paper builds GeoBoN, a fixed-budget Best-of-N selector that improves N=8 success in all five benchmark–backbone settings, and Gated GeoBoN, which invokes sampling only when an optical-flow action–future consistency test flags","pith_inferences":["A natural extension the paper leaves implicit: the same reprojection score could serve as a self-supervised preference signal for fine-tuning WAMs, since it predicts rollout quality without success labels; one could generate on-policy candidates and optimize against the score.","The gate-plus-evaluator split suggests a compute-adaptive controller that tunes the gate threshold per task or per domain; the paper fixes τgate = −0.2 across all settings, and the diagnostics show gate informativeness varies (e.g., only +4.1 points on RoboTwin/Motus), so adapting the threshold could recover more gain.","Because the geometric evaluator is model-agnostic and operates on multi-view predicted frames, the same recipe could transfer to other multi-view world models or video prediction benchmarks outside manipulation, where no action labels or task rewards are available."],"forward_implications":["Fixed-budget GeoBoN improves N=8 task success in every evaluated setting; on RoboCasa the group average rises from 66.3% to 68.4% with Cosmos Policy and from 80.8% to 82.5% with X-WAM.","Gated GeoBoN recovers on average 74.8% of the always-on success gain while triggering extra sampling at only 26.2% of decision points, cutting latency substantially (e.g., 3.65s to 1.29s on RoboCasa/Cosmos Policy).","Cross-view reprojection is the only selector among the tested training-free signals (confidence-only, future-consensus) that improves closed-loop success in all settings; it also beats the Cosmos Policy value head at identical candidate budgets.","The gate is an informative compute-allocation signal: at equal trigger rates it raises the 'GeoBoN helps' rate above a random trigger in all settings, by up to 24 points on RoboCasa/X-WAM.","Larger candidate pools create a multiple-comparisons failure mode: false low-score selections rise from 7% at N=2 to 32% at N=16, explaining why Best-of-N success saturates or degrades and motivating moderate budgets with selective invocation."],"fun_headline_variants":["Robot rollouts ranked by cross-view depth agreement","GeoBoN: pick rollouts via depth reprojection consistency","Cross-view depth scores select better robot actions","Task-free score: cross-camera depth agreement for WAMs","Gated Best-of-N with geometric consistency beats fixed budget"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The method assumes that a frozen depth-and-pose model, when fed the synthetic frames a WAM imagines, produces errors small and uncorrelated enough across candidates that the lowest reprojection error really marks the most coherent future; if those errors are large or correlated, the selector picks a biased outlier. This premise enters at Eq. 5 (Section 3.2) and is not directly validated.","fun_headline_variants_meta":{"raw":{"variants":["Robot rollouts ranked by cross-view depth agreement","GeoBoN: pick rollouts via depth reprojection consistency","Cross-view depth scores select better robot actions","Task-free score: cross-camera depth agreement for WAMs","Gated Best-of-N with geometric consistency beats fixed budget"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1185,"prompt_tokens":825,"completion_tokens":360,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":281}},"tokens_in":569,"tokens_out":360,"duration_ms":3668,"temperature":1.0,"reasoning_tokens":281,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T17:53:12.472007+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fixed set of RoboCasa tasks, sample N=8 rollouts per decision, run every candidate in simulation, and check whether the candidate with the lowest Eq. 5 reprojection error succeeds more often than random or value-head picks; if low reprojection error does not track simulation success on the WAM's predicted frames, the central claim fails. This can be tested offline from the paper's candidate dumps.","supporting_citations":[],"review_version":1}