{"id":"2ed3c91e-29ec-4db8-9df7-8243ad3a759a","arxiv_id":"2607.26789","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"Action-conditioned world-model verification with conformal first-intervention control and latency-aware suffix repair raises RoboCasa365 success 8.5 points over invocation-matched periodic replanning.","lead":"CheckVLA watches open-loop robot action chunks with a frozen action-conditioned world model and intervenes only when predicted consequences diverge from what is seen. On a large household mobile-manipulation sim benchmark, that timed, latency-aware repair beats matched-budget periodic replanning by 8.5 success points.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Strongest empirical claim is internally consistent; soft spot is repair-coupled “timely” labels plus in-family perturbation eval, not conformal exchangeability alone.","rationale":"The manuscript is a careful systems paper: paired seeds, invocation-matched periodic control, action-shuffled and observation-only monitors, fixed-trigger repair branches, and explicit non-guarantees on conformal scope. No internal contradiction overturns Table 3’s +8.5 or Table 4’s conditioning gap under shared rewriters and empirical FWER matching. The reader correctly flags simulation-only evidence, frozen V-JEPA features, and narrow conformal coverage as reasons for CONDITIONAL rather than ACCEPT. I only re-weight the weakest link for the stated strongest claim: exchangeability is not what carries the comparative recall numbers once empirical FWER is matched; repair-coupled timeliness (App. G) and in-family perturbation evaluation are. Natural-audit attenuation (Table S8) already hints at that. That refinement does not justify REJECT or a harsher label—artifacts and broader validation are still the right bar—so the verdict stays CONDITIONAL / UNCHANGED.","tokens_in":30983,"tokens_out":710,"duration_ms":112627,"concrete_test":"Freeze the Table 4 shadow trajectories and shared randomness, redefine τ_irrev using only a held-out repair rule excluded from validation (wait-for-boundary or RTC inpainting alone), and re-score full vs observation-only vs action-shuffled detectors; also repeat under full-verifier LOFO monitors (Table S10). If the action-cond. timely-recall gap falls below ~15 points or full timely recall falls below ~55%, the headline detection/recovery claim is overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader’s exchangeability worry is real for the formal bound (Eqs. 6–7) but secondary for the strongest empirical claim: Table 4 matches detectors on empirical episode FWER (~5%), so the 77.9% vs 48.6%/37.9% gap does not need the conformal guarantee to be valid as a ranking.\n\nThe more load-bearing soft spot is definitional and distributional. Appendix G defines τ_irrev operationally as the earliest time after which no tested suffix repair restores success, and counts a detection timely only if t*+d_lat < τ_irrev. Absolute timely recall and perturbed success therefore measure detector–repair co-adaptation under the authors’ rewrite set, not recoverability in the abstract. Those metrics are further evaluated mainly on injected impulses/joint offsets/moved-objects that overlap the risk-head supervision taxonomy (App. D, F), with LOFO and the natural-execution audit (Tables S8, S10) as partial outs. On natural held-out configs the execution-mode lift over invocation-matched periodic replanning shrinks to +4.1 points (68.9% vs 64.8%) versus +8.5 on the clean suite build-up (Table 3). The controlled action-conditioning gap under a shared rewriter remains credible; the headline recovery magnitude is softer outside this stack.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"CheckVLA reframes open-loop VLA action chunks as testable predictions of near-future observations and verifies them at execution time with a separately trained, frozen action-conditioned world model. A causal risk head aggregates prediction–observation discrepancies; a split functional conformal threshold bounds the episode-level probability of an unnecessary first intervention on exchangeable nominal successes (Eqs. 6–7). On a trigger, the same VLA rewrites only the latency-feasible suffix under hard prefixing, with exceedance-conditioned retention of the superseded chunk, while an event-driven keyframe bank preserves episodic context. On RoboCasa365 under a common Human300 recipe and matched invocation budget, the system reports 36.1% average success versus 27.6% for periodic replanning (+8.5 points; Table 3). At matched ~5% episode FWER, action conditioning raises timely recall to 77.9% versus 48.6% observation-only and 37.9% action-shuffled controls (Table 4).","tokens_in":31435,"tokens_out":1332,"duration_ms":31645,"significance":"If the controlled results hold, the paper makes a clear systems contribution: action-conditioned consequence checking can restore feedback inside committed chunks without paying full policy-rate inference, and the repair can be made latency-consistent. Strengths include the matched-invocation build-up (Table 3), independently recalibrated detectors at a common episode FWER (Table 4), fixed-trigger repair branches from identical snapshots (Table 6), action-shuffled and observation-only negatives, inference-time binding audits (Table S5), and unusually thorough leakage guards and failure taxonomy in the appendix. The work is a useful complement to policy scaling and adaptive chunking rather than a replacement for either. The main external limit is that all evidence is RoboCasa365 simulation with a frozen V-JEPA encoder and a finite perturbation family; the Discussion already flags recalibration needs under sensing or timing shift.","major_comments":[{"comment":"Appendix G defines τ_irrev operationally as the earliest time after which no tested suffix repair restores success, and counts a detection timely only if t*+d_lat < τ_irrev. Absolute timely recall (77.9% in Table 4) and much of the perturbed-success narrative therefore measure detector–repair co-adaptation under the authors’ rewrite set, not recoverability independent of that set. The shared-rewriter detector ranking remains informative, but the main text (Abstract, Experiments, contributions) should state this coupling explicitly wherever timely recall is headline-quoted, and preferably report a repair-agnostic ranking metric (e.g., event AUPRC / lead before any rewrite) beside it.","section":"Appendix G; Table 4; Abstract"},{"comment":"The risk head is supervised on physics-level impulses, joint offsets, and object displacements (Appendix D), and the locked perturbation study uses the same families (Appendix F). LOFO and the natural-execution audit (Tables S8, S10) partially address this, but on held-out natural configs the lift over invocation-matched periodic replanning shrinks to +4.1 points (68.9% vs 64.8%) versus +8.5 on the clean build-up (Table 3). The central action-conditioning gap is still credible under a shared rewriter; the recovery magnitude should be framed more cautiously in the Abstract and Conclusion as stack- and distribution-dependent, with the natural-execution numbers promoted into the main experimental narrative.","section":"Table 3; Table S8; Appendix D, F"},{"comment":"Eqs. 6–7 give only a first-intervention guarantee on exchangeable nominal successes under a fully fixed pipeline. The paper states this correctly in the Method and Discussion, yet the Abstract’s phrasing (“bounds the episode-level probability of an unnecessary first intervention”) can be read as a deployed safety property. Please keep the bound claim tightly scoped in the Abstract and ensure Table S17’s repeated-intervention / post-repair harm numbers are cited whenever the conformal trigger is presented as the operating point.","section":"Eqs. 6–7; Abstract; Discussion; Table S17"}],"minor_comments":[{"comment":"Table 2’s “state of the art” positioning is labeled descriptive, which is appropriate; consider moving the controlled Table 3 comparison earlier in the Experiments narrative so readers do not overweight cross-system averages.","section":"Table 2; Experiments"},{"comment":"Figure 1 and Figure 2 are helpful; a short pointer in the main text to Appendix Figure S1 (latency hard-prefix schematic) would make d_lat / d_eff easier to parse on first reading.","section":"Figure 1; Appendix A"},{"comment":"Notation for span ℓ(t), active index h_τ, and effective delay d_eff is dense in the Problem Formulation; a one-line symbol table or tighter cross-references to Algorithm S1 would help.","section":"Problem Formulation; Algorithm S1"},{"comment":"Typos / spacing artifacts appear in several places (e.g., “topropagatetheerror”, “CheckVLA, which verifies”, missing spaces after periods in the Abstract and Introduction). A full copy-edit pass is needed.","section":"Abstract; Introduction"},{"comment":"Table 1’s capability marks are useful but explicitly “not re-implementations”; a footnote restating that avoids over-reading the checkmarks as head-to-head results.","section":"Table 1"}],"recommendation":"minor_revision","confidential_remarks":"Solid empirical systems paper with above-average experimental hygiene for this area. I would not block on missing hardware if the authors tighten the timely-recall and natural-execution framing; a major_revision would be warranted only if they refuse to qualify the co-adapted metrics. Fit is appropriate for a robotics / learning systems venue that accepts strong simulation benchmarks with clear limitation statements."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is a careful execution-stack paper, not a new foundation model. Under a shared recipe and matched VLA-call budget on RoboCasa365, CheckVLA gets 36.1% average success versus 27.6% for periodic replanning, and at ~5% episode FWER the action-conditioned verifier hits 77.9% timely recall against 48.6% obs-only and 37.9% action-shuffled. Those gaps are backed by paired seeds, independently recalibrated detectors, fixed-trigger rewrite branches from identical snapshots, and binding audits. That is real internal evidence.\n\nWhat is new is the integration, not any single piece. Action-conditioned world models, conformal first-intervention control, hard-prefix flow rewrite, exceedance-weighted retention, and an event-driven keyframe bank with dual readers already exist in fragments (SAFE, Foresight, RTC, future–reality checks, KEMO-style memory). Table 1 is honest about that. The contribution is turning consequence prediction into a calibrated in-chunk trigger and a latency-feasible repair without observation-rate policy calls, and documenting the stack thoroughly in the appendix.\n\nSoft spots, in proportion. Everything is simulation; no code or weights. Timely recall and perturbed success are defined against an operational irreversibility time under the authors’ own rewrite set, so detector and repair are co-adapted—fair for a systems claim, weaker as abstract recoverability. Injected perturbations overlap the risk-head training taxonomy; LOFO and the natural audit help, and on held-out natural configs the lift over matched periodic shrinks to about +4 points. Conformal exchangeability only bounds unnecessary first interventions on nominal successes; the paper says so. Validation-chosen knobs are numerous but mostly locked before the test suites.\n\nThis is for people building long-horizon VLA mobile manipulators who care about when to spend the next policy call. Math and citation pattern look solid; no load-bearing contradiction in the controlled tables. I would send it to peer review and would bring it to reading group. Cite if you work on chunked execution monitors; skip if you only want foundational world-model theory or hardware results.","headline":"Solid sim systems paper: action-conditioned verification plus latency-aware suffix repair beats matched periodic replanning by a real margin, with unusually careful controls; novelty is integrative and evidence stays inside RoboCasa365.","tokens_in":32109,"tokens_out":548,"would_cite":true,"duration_ms":12458,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"An open-loop action chunk is a testable prediction of future observations, and checking it with an action-conditioned world model restores feedback during long-horizon robot execution.","keywords":["vision-language-action","mobile manipulation","action chunking","action-conditioned world model","execution-time verification","conformal calibration","latency-aware repair","episodic keyframe memory"],"falsifier":"On the same backbone and call budget, if an observation-only or action-shuffled monitor matched the full verifier’s timely recall and perturbed success at the same 5% episode false-alarm rate, or if CheckVLA no longer beat invocation-matched periodic replanning on RoboCasa365, the central claim would fail.","tokens_in":31853,"feed_emoji":"🤖","tokens_out":992,"duration_ms":20767,"temperature":0.7,"pith_summary":"Vision-language-action robots often fire a whole sequence of actions without looking again, which is efficient but dangerous: if a bowl slips or the base drifts mid-chunk, the remaining actions keep executing a plan that is already wrong. CheckVLA treats each committed chunk as a short-horizon forecast of how the scene should evolve, then compares that forecast to what actually arrives using a separately trained, frozen action-conditioned world model. A calibrated risk score decides when the mismatch is serious enough to interrupt, how strongly to keep or discard the old plan, and which suffix can still be rewritten given inference delay, while a keyframe memory keeps completed subgoals from being forgotten. On the RoboCasa365 household benchmark, under the same training recipe and the same number of policy calls, this beats periodic replanning by 8.5 success points, and action conditioning roughly doubles timely failure recall versus observation-only or action-shuffled monitors at a matched 5% false-alarm rate. The practical claim is that verification and latency-aware repair can put feedback back into chunked control without waiting for a stronger open-loop policy alone.","feed_headline":"Robots catch mid-chunk failures by checking what actions predict","feed_subtitle":"Action-conditioned verification beats periodic replanning by 8.5 points at the same call budget","key_machinery":"CheckVLA’s action-conditioned verifier: a short rolling world-model forecast of observation features given remaining committed actions, scored by a causal risk head and compared to a split functional conformal threshold that bounds the episode-level probability of an unnecessary first intervention; threshold exceedance then sets how strongly the rewritten, latency-feasible suffix retains the old chunk.","core_discovery":"A committed action chunk is both a control command and a testable prediction of near-future observations. Verifying that prediction online with a frozen action-conditioned world model, a conformally calibrated first-intervention threshold, risk-adaptive suffix retention, latency-aware hard prefixing, and event-driven episodic memory restores closed-loop feedback during open-loop chunk execution and, at matched policy-call budget on RoboCasa365, raises average success from 27.6% (periodic replanning) to 36.1%.","pith_inferences":["Any robot stack that amortizes compute via open-loop action segments—not only kitchen VLAs—could adopt the same predict-then-verify loop.","Hardware transfer will hinge less on new architecture than on redoing conformal calibration under real sensing noise and timing jitter.","Separating a frozen monitor from the policy path suggests safety auditors could certify the verifier without retraining the controller.","If world-model features stay frozen while tasks shift, the next bottleneck is representation drift, not the conformal math."],"forward_implications":["Chunked VLA systems can regain mid-chunk feedback without querying the full policy at every step.","Action–consequence mismatch detects confidently wrong open-loop failures that commit-time policy uncertainty misses.","Repair strength should scale with calibrated risk exceedance, not only with a binary replan flag.","Inference latency must hard-constrain which suffix is rewritten, or correct alarms still cannot recover the episode.","Persistent keyframe memory is needed so repairs do not erase completed subgoals that left the camera view."],"fun_headline_variants":["CheckVLA verifies action chunks mid-run with frozen world models","Action-conditioned checks lift mobile manip success 8.5 pts vs replan","Committed chunks predict observations; CheckVLA flags deviations online","Conformal risk gates when robots rewrite failing open-loop suffixes","World-model verification restores feedback in chunked VLA execution"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"New successful runs must stay exchangeable with the fixed calibration pipeline so the risk threshold really bounds unnecessary first interruptions; the paper only claims that bound for clean successes, not for recall, post-repair safety, or real-world shift.","fun_headline_variants_meta":{"raw":{"variants":["CheckVLA verifies action chunks mid-run with frozen world models","Action-conditioned checks lift mobile manip success 8.5 pts vs replan","Committed chunks predict observations; CheckVLA flags deviations online","Conformal risk gates when robots rewrite failing open-loop suffixes","World-model verification restores feedback in chunked VLA execution"]},"model":"grok-4.5","effort":"low","cost_usd":0.001873,"raw_usage":{"total_tokens":968,"prompt_tokens":875,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":18728000,"prompt_tokens_details":{"text_tokens":875,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":875,"tokens_out":73,"duration_ms":2458,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T21:02:25.026836+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the same backbone and call budget, if an observation-only or action-shuffled monitor matched the full verifier’s timely recall and perturbed success at the same 5% episode false-alarm rate, or if CheckVLA no longer beat invocation-matched periodic replanning on RoboCasa365, the central claim would fail.","supporting_citations":[],"review_version":1}