Pith. sign in

REVIEW 3 major objections 5 minor 53 references

An open-loop action chunk is a testable prediction of future observations, and checking it with an action-conditioned world model restores feedback during long-horizon robot execution.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 21:02 UTC pith:GBFAZVPW

load-bearing objection Solid sim systems paper: action-conditioned verification plus latency-aware suffix repair beats matched periodic replanning by a real margin, with unusually careful controls; novelty is integrative and evidence stays inside RoboCasa365. the 3 major comments →

arxiv 2607.26789 v1 pith:GBFAZVPW submitted 2026-07-29 cs.RO

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

classification cs.RO
keywords vision-language-actionmobile manipulationaction chunkingaction-conditioned world modelexecution-time verificationconformal calibrationlatency-aware repairepisodic keyframe memory
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Vision-language-action robots often fire a whole sequence of actions without looking again, which is efficient but dangerous: if a bowl slips or the base drifts mid-chunk, the remaining actions keep executing a plan that is already wrong. CheckVLA treats each committed chunk as a short-horizon forecast of how the scene should evolve, then compares that forecast to what actually arrives using a separately trained, frozen action-conditioned world model. A calibrated risk score decides when the mismatch is serious enough to interrupt, how strongly to keep or discard the old plan, and which suffix can still be rewritten given inference delay, while a keyframe memory keeps completed subgoals from being forgotten. On the RoboCasa365 household benchmark, under the same training recipe and the same number of policy calls, this beats periodic replanning by 8.5 success points, and action conditioning roughly doubles timely failure recall versus observation-only or action-shuffled monitors at a matched 5% false-alarm rate. The practical claim is that verification and latency-aware repair can put feedback back into chunked control without waiting for a stronger open-loop policy alone.

Core claim

A committed action chunk is both a control command and a testable prediction of near-future observations. Verifying that prediction online with a frozen action-conditioned world model, a conformally calibrated first-intervention threshold, risk-adaptive suffix retention, latency-aware hard prefixing, and event-driven episodic memory restores closed-loop feedback during open-loop chunk execution and, at matched policy-call budget on RoboCasa365, raises average success from 27.6% (periodic replanning) to 36.1%.

What carries the argument

CheckVLA’s action-conditioned verifier: a short rolling world-model forecast of observation features given remaining committed actions, scored by a causal risk head and compared to a split functional conformal threshold that bounds the episode-level probability of an unnecessary first intervention; threshold exceedance then sets how strongly the rewritten, latency-feasible suffix retains the old chunk.

Load-bearing premise

New successful runs must stay exchangeable with the fixed calibration pipeline so the risk threshold really bounds unnecessary first interruptions; the paper only claims that bound for clean successes, not for recall, post-repair safety, or real-world shift.

What would settle it

On the same backbone and call budget, if an observation-only or action-shuffled monitor matched the full verifier’s timely recall and perturbed success at the same 5% episode false-alarm rate, or if CheckVLA no longer beat invocation-matched periodic replanning on RoboCasa365, the central claim would fail.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Chunked VLA systems can regain mid-chunk feedback without querying the full policy at every step.
  • Action–consequence mismatch detects confidently wrong open-loop failures that commit-time policy uncertainty misses.
  • Repair strength should scale with calibrated risk exceedance, not only with a binary replan flag.
  • Inference latency must hard-constrain which suffix is rewritten, or correct alarms still cannot recover the episode.
  • Persistent keyframe memory is needed so repairs do not erase completed subgoals that left the camera view.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Any robot stack that amortizes compute via open-loop action segments—not only kitchen VLAs—could adopt the same predict-then-verify loop.
  • Hardware transfer will hinge less on new architecture than on redoing conformal calibration under real sensing noise and timing jitter.
  • Separating a frozen monitor from the policy path suggests safety auditors could certify the verifier without retraining the controller.
  • If world-model features stay frozen while tasks shift, the next bottleneck is representation drift, not the conformal math.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. CheckVLA reframes open-loop VLA action chunks as testable predictions of near-future observations and verifies them at execution time with a separately trained, frozen action-conditioned world model. A causal risk head aggregates prediction–observation discrepancies; a split functional conformal threshold bounds the episode-level probability of an unnecessary first intervention on exchangeable nominal successes (Eqs. 6–7). On a trigger, the same VLA rewrites only the latency-feasible suffix under hard prefixing, with exceedance-conditioned retention of the superseded chunk, while an event-driven keyframe bank preserves episodic context. On RoboCasa365 under a common Human300 recipe and matched invocation budget, the system reports 36.1% average success versus 27.6% for periodic replanning (+8.5 points; Table 3). At matched ~5% episode FWER, action conditioning raises timely recall to 77.9% versus 48.6% observation-only and 37.9% action-shuffled controls (Table 4).

Significance. If the controlled results hold, the paper makes a clear systems contribution: action-conditioned consequence checking can restore feedback inside committed chunks without paying full policy-rate inference, and the repair can be made latency-consistent. Strengths include the matched-invocation build-up (Table 3), independently recalibrated detectors at a common episode FWER (Table 4), fixed-trigger repair branches from identical snapshots (Table 6), action-shuffled and observation-only negatives, inference-time binding audits (Table S5), and unusually thorough leakage guards and failure taxonomy in the appendix. The work is a useful complement to policy scaling and adaptive chunking rather than a replacement for either. The main external limit is that all evidence is RoboCasa365 simulation with a frozen V-JEPA encoder and a finite perturbation family; the Discussion already flags recalibration needs under sensing or timing shift.

major comments (3)
  1. [Appendix G; Table 4; Abstract] Appendix G defines τ_irrev operationally as the earliest time after which no tested suffix repair restores success, and counts a detection timely only if t*+d_lat < τ_irrev. Absolute timely recall (77.9% in Table 4) and much of the perturbed-success narrative therefore measure detector–repair co-adaptation under the authors’ rewrite set, not recoverability independent of that set. The shared-rewriter detector ranking remains informative, but the main text (Abstract, Experiments, contributions) should state this coupling explicitly wherever timely recall is headline-quoted, and preferably report a repair-agnostic ranking metric (e.g., event AUPRC / lead before any rewrite) beside it.
  2. [Table 3; Table S8; Appendix D, F] The risk head is supervised on physics-level impulses, joint offsets, and object displacements (Appendix D), and the locked perturbation study uses the same families (Appendix F). LOFO and the natural-execution audit (Tables S8, S10) partially address this, but on held-out natural configs the lift over invocation-matched periodic replanning shrinks to +4.1 points (68.9% vs 64.8%) versus +8.5 on the clean build-up (Table 3). The central action-conditioning gap is still credible under a shared rewriter; the recovery magnitude should be framed more cautiously in the Abstract and Conclusion as stack- and distribution-dependent, with the natural-execution numbers promoted into the main experimental narrative.
  3. [Eqs. 6–7; Abstract; Discussion; Table S17] Eqs. 6–7 give only a first-intervention guarantee on exchangeable nominal successes under a fully fixed pipeline. The paper states this correctly in the Method and Discussion, yet the Abstract’s phrasing (“bounds the episode-level probability of an unnecessary first intervention”) can be read as a deployed safety property. Please keep the bound claim tightly scoped in the Abstract and ensure Table S17’s repeated-intervention / post-repair harm numbers are cited whenever the conformal trigger is presented as the operating point.
minor comments (5)
  1. [Table 2; Experiments] Table 2’s “state of the art” positioning is labeled descriptive, which is appropriate; consider moving the controlled Table 3 comparison earlier in the Experiments narrative so readers do not overweight cross-system averages.
  2. [Figure 1; Appendix A] Figure 1 and Figure 2 are helpful; a short pointer in the main text to Appendix Figure S1 (latency hard-prefix schematic) would make d_lat / d_eff easier to parse on first reading.
  3. [Problem Formulation; Algorithm S1] Notation for span ℓ(t), active index h_τ, and effective delay d_eff is dense in the Problem Formulation; a one-line symbol table or tighter cross-references to Algorithm S1 would help.
  4. [Abstract; Introduction] Typos / spacing artifacts appear in several places (e.g., “topropagatetheerror”, “CheckVLA, which verifies”, missing spaces after periods in the Abstract and Introduction). A full copy-edit pass is needed.
  5. [Table 1] Table 1’s capability marks are useful but explicitly “not re-implementations”; a footnote restating that avoids over-reading the checkmarks as head-to-head results.

Circularity Check

0 steps flagged

No significant circularity: empirical systems results with held-out conformal fits and matched controls, not a derivation that re-labels inputs as predictions.

full rationale

CheckVLA is an empirical robotics systems paper. Its load-bearing claims are comparative success and detector metrics on RoboCasa365 under a common recipe, matched invocation budget, and independently recalibrated detectors—not a first-principles derivation chain. The functional conformal threshold (Eqs. 6–7) is the standard split/finite-sample construction: risk statistics and the quantile q̂_α are fit on nominal-success shadow trajectories disjoint from validation and locked tests, then the bound is stated only for unnecessary first interventions under exchangeability with that fixed pipeline. That is a statistical guarantee by design of conformal calibration, not a fitted quantity renamed as an independent prediction of the same data. Action-conditioning evidence uses negative controls (observation-only and action-shuffled predictors) retrained and recalibrated at the same episode FWER target; rewrite ablations branch from shared first-trigger snapshots so detector quality does not silently define repair wins. Validation-selected maps (exceedance→retention, periodic interval, discrepancy metric) are hyperparameters chosen once before locked tests, not claimed closed-form predictions. Self-citations to related VLA/world-model work are background and non-load-bearing for uniqueness. Operational definition of τ_irrev via the tested repair set couples absolute timely-recall magnitude to the authors’ rewriter, but that is a disclosed metric scope choice evaluated comparatively under a shared rewriter—not a circular reduction of a claimed derivation to its inputs. No step reduces Eq. X to Eq. Y by construction in the sense required here.

Axiom & Free-Parameter Ledger

10 free parameters · 6 axioms · 4 invented entities

Load-bearing content is empirical robotics assumptions plus many validation-chosen scalars, not a short axiom list. The conformal claim needs exchangeability under a frozen pipeline; the success claim needs the sim stack, Human300-only policy data, and the chosen repair/latency model. Invented entities are engineering modules (risk head, keyframe bank, exceedance map), not new physical objects.

free parameters (10)
  • conformal miscoverage α = 0.05
    Target episode-level unnecessary-first-intervention rate; sets operating point of all detector comparisons.
  • prediction span bound k = 8
    Short-horizon WM rollout length; validation-selected; affects discrepancy regimes and suppression.
  • risk window w = 32
    Temporal aggregation length for causal risk head; validation-selected.
  • exceedance sensitivity β and floor w_min = β=1.38, w_min=0.0
    Maps standardized threshold exceedance to reference retention w_0(e); validation-selected, no conformal guarantee.
  • channel decay rates λ_mobility / λ_manip = replan 0.60/0.90; regular 0.10/0.15
    Positional decay of guidance weights on replan vs regular transitions; validation-selected.
  • scheduled latency d_lat = 3 control steps
    Fixed control steps reserved for inference before suffix deploy; nominal 3 from measured VLA latency.
  • chunk length H = 50
    Open-loop action horizon of the backbone policy.
  • keyframe bank budget K and write-filter thresholds = K=8; gap 10; similarity 0.95
    Memory capacity and pause/diversity gates; validation-selected; changes subgoal regression and FWER.
  • span-wise discrepancy stats (μ_d^(ℓ), σ_d^(ℓ)) and risk stats (μ_r,t, σ_r,t) = fit on shadow-mode nominals (200 fit / 400 calib per seed)
    Fitted on held-out nominal successes; enter standardized scores and δ_t.
  • periodic replan interval = 39 control steps (~10.1 calls/ep)
    Validation-chosen to match CheckVLA call budget for the main baseline.
axioms (6)
  • standard math Split functional conformal prediction under exchangeability of nominal-success trajectories with a fully fixed score pipeline yields a trajectory-marginal bound on unnecessary first interventions (Eq. 7).
    Invoked in Calibrated Risk Triggering and Appendix D; standard CP finite-sample quantile construction.
  • domain assumption RoboCasa365 physics, rendering, and task success predicates are an adequate testbed for the claimed execution-time verification benefits.
    All headline metrics are sim-only; Discussion states hardware validation is future work.
  • domain assumption A frozen V-JEPA 2-AC feature space plus short action-conditioned rollouts supply residuals that separate actionable deviations from benign noise when aggregated by R.
    Core of Action-Conditioned Rolling Prediction; supported by shuffled/obs-only ablations but representation is not learned end-to-end for the task.
  • domain assumption Inference latency can be treated as a known scheduled delay d_lat with hard prefix clamp as an empirical deployable projection.
    Suffix Rewriting and Appendix J explicitly note the clamp is not exact conditional sampling.
  • domain assumption Human300-only policy pretraining plus training-side auxiliary rollouts for the verifier introduce no leakage into Composite-Unseen and target-suite scenes.
    Protocol and Appendix F leakage guards; necessary for fair benchmark reading.
  • ad hoc to paper Validation-selected discrepancy metric, guidance map, and suppression rules generalize to locked clean and perturbation tests without retuning.
    Multiple components are chosen once on validation (Table S20); locked-test integrity depends on that discipline holding.
invented entities (4)
  • CheckVLA causal risk head R with episodic summary m(ρ_t) no independent evidence
    purpose: Aggregate standardized multi-step discrepancies into a calibrated intervention score r_t ∈ [0,1].
    Trained module specific to this paper’s monitor; not a prior named law.
  • Exceedance-conditioned retention map W(e) for suffix guidance no independent evidence
    purpose: Tie repair strength to how far risk exceeded the conformal threshold.
    Validation-designed mapping (Eqs. 8–10); ablations support pairing but entity is paper-internal.
  • Event-driven keyframe bank B_t with dual policy/risk readers no independent evidence
    purpose: Preserve completed-subgoal evidence across chunk repairs without flooding the risk head.
    Engineering memory design; related to other keyframe memories but specified here with pause/diversity filter.
  • Latency-aware hard-prefixed flow rewrite of the superseded chunk no independent evidence
    purpose: Replace only deployable suffix actions while clamping irreversible prefix during d_lat/d_eff.
    Systems mechanism combining hard clamp with guided flow; RTC compared as alternative.

pith-pipeline@v1.2.0-daily-grok45 · 34978 in / 4657 out tokens · 88457 ms · 2026-07-30T21:02:25.026836+00:00 · methodology

0 comments
read the original abstract

Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation while the remaining actions continue to propagate the error: commit-time policy confidence cannot react to a deviation that occurs after dispatch, and observation-only anomaly scores lack an action-conditioned reference for separating expected effects from unexplained changes. We propose CheckVLA, which verifies execution with a separately trained, frozen action-conditioned world model. A conformally calibrated risk threshold bounds the episode-level probability of an unnecessary first intervention and determines when to intervene, its exceedance controls how strongly the rewritten suffix retains the superseded chunk, latency-aware hard prefixing restricts replacement to actions that remain deployable, and an event-driven keyframe bank preserves evidence of prior progress across repairs. On RoboCasa365, under a common training recipe and a matched invocation budget, CheckVLA attains a 36.1% average success rate against 27.6% for periodic replanning (+8.5 points). At a matched 5% episode-level false-alarm target, action conditioning raises timely recall to 77.9%, against 48.6% for an observation-only control and 37.9% for an action-shuffled control. These simulation results support action-conditioned verification as a way to restore feedback during chunked execution while keeping the repair consistent with inference latency.

Figures

Figures reproduced from arXiv: 2607.26789 by Chenyu Tang, Fang Chen, Lingfeng Zhang, Peibo Sun, Shoujie Li, Wenbo Ding, Xiao-Ping Zhang, Xintao Chao, Yifan Xie, Yushan Liu, Zhenyang Yang.

Figure 1
Figure 1. Figure 1: Overview of CheckVLA. Open-loop chunks can propagate in-chunk failures such as slippage. CheckVLA compares [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The CheckVLA framework. A separately trained, frozen encoder supplies monitoring features; the world model [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Composite RoboCasa365 episodes (Arrange Bread Basket, top; Gather Tableware, bottom): subtask prompts aligned with third- and first-person keyframes; blue denotes manipulation and orange base motion. Because one unnecessary intervention can spoil an other￾wise successful episode, CheckVLA controls the probability of an unnecessary first intervention over the episode. After fixing the policy, monitor, retri… view at source ↗
Figure 4
Figure 4. Figure 4: Closed-loop diagnostics. (a) Episode-level unnecessary-first-intervention frequency versus target conformal level. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

53 extracted references · 40 linked inside Pith

  1. [1]

    Assran, M.; Bardes, A.; Fan, D.; Garrido, Q.; Howes, R.; Komeili, M.; Muckley, M.; Rizvi, A.; Roberts, C.; Sinha, K.; et al. 2025. V-JEPA 2 : Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv:2506.09985

  2. [2]

    Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. 2024. _0 : A Vision-Language-Action Flow Model for General Robot Control. arXiv:2410.24164

  3. [3]

    Y.; and Levine, S

    Black, K.; Galliker, M. Y.; and Levine, S. 2025. Real-Time Execution of Action Chunking Flow Policies. arXiv:2506.07339

  4. [4]

    Cen, J.; Yu, C.; Yuan, H.; Jiang, Y.; Huang, S.; Guo, J.; Li, X.; Song, Y.; Luo, H.; Wang, F.; Zhao, D.; and Chen, H. 2025. WorldVLA : Towards Autoregressive Action World Model. arXiv:2506.21539

  5. [5]

    Chao, X.; Mu, S.; Liu, Y.; Li, S.; Lyu, C.; Zhang, X.-P.; and Ding, W. 2025. Exo-ViHa : A Cross-Platform Exoskeleton System with Visual and Haptic Feedback for Efficient Dexterous Skill Learning. arXiv:2503.01543

  6. [6]

    Chen, J. 2026. AEGIS : A Backup Reflex for Physical AI . arXiv:2606.06660

  7. [7]

    Chen, R.; Yang, Y.; Tang, Z.; Huo, D.; Lin, T.; Wu, H.; Liu, H.; Chen, Y.; et al. 2026. ABot-M0.5 : Unified Mobility-and-Manipulation World Action Model. arXiv:2607.00678

  8. [8]

    Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; and Song, S. 2023. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv:2303.04137

  9. [9]

    T.; Ichter, B.; Yu, L.; Li-Bell, A.; Pertsch, K.; Ren, A

    Driess, D.; Springenberg, J. T.; Ichter, B.; Yu, L.; Li-Bell, A.; Pertsch, K.; Ren, A. Z.; Walke, H.; Vuong, Q.; Shi, L. X.; and Levine, S. 2025. Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better. arXiv:2505.23705

  10. [10]

    Fan, J.; Liu, Y.; Li, S.; Ren, B.; Li, S.; Zhang, X.-P.; Ding, W.; and Deng, Z. 2026. FUTURE-VLA : Forecasting Unified Trajectories Under Real-time Execution. arXiv:2602.15882

  11. [11]

    Z.; and Finn, C

    Fu, Z.; Zhao, T. Z.; and Finn, C. 2024. Mobile ALOHA : Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. arXiv:2401.02117

  12. [12]

    Gu, Q.; Ju, Y.; Sun, S.; Gilitschenski, I.; Nishimura, H.; Itkina, M.; and Shkurti, F. 2025. SAFE : Multitask Failure Detection for Vision-Language-Action Models. arXiv:2506.09937

  13. [13]

    Guo, Y.; Wang, Y.-J.; Zha, L.; and Chen, J. 2023. DoReMi : Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment. arXiv:2307.00329

  14. [14]

    Hafner, D.; Lillicrap, T.; Ba, J.; and Norouzi, M. 2019. Dream to Control: Learning Behaviors by Latent Imagination. arXiv:1912.01603

  15. [15]

    Hafner, D.; Lillicrap, T.; Fischer, I.; Villegas, R.; Ha, D.; Lee, H.; and Davidson, J. 2018. Learning Latent Dynamics for Planning from Pixels. arXiv:1811.04551

  16. [16]

    F.; Ward, I

    Ho, M.; Ginting, M. F.; Ward, I. R.; Reinke, A.; Kochenderfer, M. J.; Agha-Mohammadi, A.-a.; and Omidshafiei, S. 2026. World Model Failure Classification and Anomaly Detection for Autonomous Inspection. arXiv:2602.16182

  17. [17]

    Jiang, H.; Chen, J.; Bu, Q.; Chen, L.; Shi, M.; Zhang, Y.; Li, D.; Suo, C.; Wang, C.; Peng, Z.; and Li, H. 2025. WholeBodyVLA : Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control. arXiv:2512.11047

  18. [18]

    Jing, D.; Wang, G.; Liu, J.; Tang, W.; Sun, Z.; Yao, Y.; Wei, Z.; Liu, Y.; Lu, Z.; and Ding, M. 2025. Mixture of Horizons in Action Chunking. arXiv:2511.19433

  19. [19]

    Kim, D.; Jang, H.; Koo, M.; Jang, S.; Kim, T.; Kim, B.; Yoon, B.; Jang, C.; Choi, D.; Han, D.; et al. 2026. RLDX-1 Technical Report. arXiv:2605.03269

  20. [20]

    Liu, Y.; Feng, F.; Kong, L.; Lu, W.; Tang, J.; Zhang, K.; Murphy, K.; Finn, C.; and Du, Y. 2026 a . World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry. arXiv:2604.01985

  21. [21]

    I.; Xie, A.; Lee, Y.; Du, M.; and Finn, C

    Liu, Y.; Hamid, J. I.; Xie, A.; Lee, Y.; Du, M.; and Finn, C. 2024. Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling. arXiv:2408.17355

  22. [22]

    Liu, Y.; Lv, T.; Wang, B.; Fan, H.; Zhao, C.; Zheng, H.; Zhong, X.; Xie, Y.; Zhao, C.; Liao, Z.; Luo, L.; Cai, Y.; Zhang, X.-P.; and Ding, W. 2026 b . PerceptDrive : Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving. arXiv:2607.20175

  23. [23]

    Liu, Y.; Mu, S.; Chao, X.; Li, Z.; Mu, Y.; Chen, T.; Li, S.; Lyu, C.; Zhang, X.-P.; and Ding, W. 2025. AVR : Active Vision-Driven Precise Robot Manipulation with Viewpoint and Focal Length Optimization. arXiv:2503.01439

  24. [24]

    Liu, Y.; Sun, P.; Li, S.; Xie, Y.; Zhang, L.; Chao, X.; Dong, S.; Chen, F.; Zhang, X.-P.; and Ding, W. 2026 c . OA-WAM : Object-Addressable World Action Model for Robust Robot Manipulation. arXiv:2605.06481

  25. [25]

    Liu, Z.; Bahety, A.; and Song, S. 2023. REFLECT : Summarizing Robot Experiences for Failure Explanation and Correction. arXiv:2306.15724

  26. [26]

    Nasiriany, S.; Nasiriany, S.; Maddukuri, A.; and Zhu, Y. 2026. RoboCasa365 : A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots. arXiv:2603.04356

  27. [27]

    Bjorck, J.; Blukis, V.; Casta\ neda, F.; Cherniadev, N.; Da, X.; Ding, R.; Fan, L.; Fang, Y.; Fox, D.; Hu, F.; et al. 2025. GR00T N1.5 : An Improved Open Foundation Model for Generalist Humanoid Robots. NVIDIA GEAR, 11 June 2025. https://research.nvidia.com/labs/gear/gr00t-n1_5/

  28. [28]

    GEAR Team ; Azzolini, A.; Bjorck, J.; Blukis, V.; Casta\ neda, F.; Chand, R.; Chang, Y.; Chen, D.; Cherniadev, N.; Da, X.; et al. 2025. GR00T N1.6 : An Improved Open Foundation Model for Generalist Humanoid Robots. NVIDIA GEAR, 15 December 2025. https://research.nvidia.com/labs/gear/gr00t-n1_6/

  29. [29]

    Pan, Y.; Pan, M.; Lu, Q.; Huang, J.; Zhang, M.; Huang, S.; Li, X.; Zhang, J.; Shen, Y.; Zhang, X.; and Zhang, W. 2026. VLA-Corrector : Lightweight Detect-and-Correct Inference for Adaptive Action Horizon. arXiv:2607.01804

  30. [30]

    Physical Intelligence ; Ai, B.; Amin, A.; Aniceto, R.; Balakrishna, A.; Balke, G.; Black, K.; Bokinsky, G.; Cao, S.; et al. 2026. _ 0.7 : a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv:2604.15483

  31. [31]

    Physical Intelligence ; Amin, A.; Aniceto, R.; Balakrishna, A.; Black, K.; Conley, K.; Connors, G.; Darpinian, J.; Dhabalia, K.; et al. 2025 a . ^ * _ 0.6 : a VLA That Learns From Experience. arXiv:2511.14759

  32. [32]

    Physical Intelligence ; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; et al. 2025 b . _ 0.5 : A Vision-Language-Action Model with Open-World Generalization. arXiv:2504.16054

  33. [33]

    Rao, P.; Zhang, W.; Balestriero, R.; LeCun, Y.; and Loianno, G. 2026. SkyJEPA : Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors. arXiv:2606.23444

  34. [34]

    Ross, S.; and Bagnell, D. 2010. Efficient Reductions for Imitation Learning. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (AISTATS), 661--668. PMLR 9

  35. [35]

    J.; and Bagnell, J

    Ross, S.; Gordon, G. J.; and Bagnell, J. A. 2011. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS), 627--635. PMLR 15. arXiv:1011.0686

  36. [36]

    Spencer, J.; Choudhury, S.; Venkatraman, A.; Ziebart, B.; and Bagnell, J. A. 2021. Feedback in Imitation Learning: The Three Regimes of Covariate Shift. arXiv:2102.02872

  37. [37]

    Sun, J.; Zhang, W.; Qi, Z.; Ren, S.; Liu, Z.; Zhu, H.; Sun, G.; Jin, X.; and Chen, Z. 2026 a . VLA-JEPA : Enhancing Vision-Language-Action Model with Latent World Model. arXiv:2602.10098

  38. [38]

    Sun, Z.; Guo, Y.; Sun, H.; Wang, L.; Lu, W.; Ji, J.; Ji, S.; Xiong, J.; and Meng, Z. 2026 b . Pre-VLA : Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts. arXiv:2605.22446

  39. [39]

    Z.; Wang, H.; Tang, J.; Stachowicz, K.; et al

    Torne, M.; Pertsch, K.; Walke, H.; Vedder, K.; Nair, S.; Ichter, B.; Ren, A. Z.; Wang, H.; Tang, J.; Stachowicz, K.; et al. 2026. MEM : Multi-Scale Embodied Memory for Vision Language Action Models. arXiv:2603.03596

  40. [40]

    R.; and Liu, G

    Wang, H.; Zhang, G.; Yan, Y.; Kompella, R. R.; and Liu, G. 2026 a . VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies. arXiv:2602.21445

  41. [41]

    R.; and Liu, G

    Wang, H.; Zhang, G.; Yan, Y.; Shang, Y.; Kompella, R. R.; and Liu, G. 2026 b . Real-Time Robot Execution with Masked Action Chunking. arXiv:2601.20130

  42. [42]

    Wang, R.; Zhang, Y.; Lin, J.; Luo, K.; Wang, J.; Wang, Z.; and Qi, X. 2026 c . When to Trust Imagination: Adaptive Action Execution for World Action Models. arXiv:2605.06222

  43. [43]

    Wang, T.; Hou, H.; Hu, Y.; Liu, Y.; Li, Q.; Jiang, Y.; Wang, Y.; Ma, C.; Wang, R.; and Gao, Y. 2026 d . When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning. arXiv:2607.25593

  44. [44]

    Wang, X.; Zhu, Z.; Huang, G.; Wang, B.; Chen, X.; and Lu, J. 2024. WorldDreamer : Towards General World Models for Video Generation via Predicting Masked Tokens. arXiv:2401.09985

  45. [45]

    Ye, A.; Wang, B.; Ni, C.; Huang, G.; Zhao, G.; Li, H.; Li, H.; Li, J.; Lv, J.; Liu, J.; Cao, M.; Li, P.; Deng, Q.; Mei, W.; Wang, X.; Chen, X.; Zhou, X.; Wang, Y.; Chang, Y.; Li, Y.; Zhou, Y.; Ye, Y.; Liu, Z.; and Zhu, Z. 2026. GigaWorld-Policy : An Efficient Action-Centered World--Action Model. arXiv:2603.17240

  46. [46]

    Yuan, H.; Liang, Z.; Chen, A.; Wang, Y.; Li, H.; Lin, P.; Huang, Y.; Lei, Z.; Zhang, T.; Zhang, J.; Zhang, J.; Fan, J.; Zhou, G.; Peng, Q.; Lv, C.; Chen, X.; Yang, A.; Huang, F.; Lin, J.; Liu, D.; Zhou, J.; Wu, C.; and Chen, X.-H. 2026. Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models. arXiv:2606.17846

  47. [47]

    Zeng, Y.; Ye, M.; Chen, Y.; Shentu, Y.; Wu, P.; Yan, Z.; and Li, Z. 2026. KEMO : Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies. arXiv:2606.23589

  48. [48]

    Zhang, H.; Lu, Y.; Wang, B.; Kang, X.; Kuo, Y.-L.; Cheng, Z.; Wang, M.; and Jenkins, O. C. 2026. Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents. arXiv:2606.23085

  49. [49]

    Zhang, W.; Liu, H.; Qi, Z.; Wang, Y.; Yu, X.; Zhang, J.; Dong, R.; He, J.; Lu, F.; Wang, H.; et al. 2025. DreamVLA : A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge. arXiv:2507.04447

  50. [50]

    Z.; Kumar, V.; Levine, S.; and Finn, C

    Zhao, T. Z.; Kumar, V.; Levine, S.; and Finn, C. 2023. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv:2304.13705

  51. [51]

    Zhong, X.; Zheng, H.; Zhao, C.; Lv, T.; Fan, H.; Wang, B.; Liu, Y.; Gao, L.; Liao, Z.; Luo, L.; Zhao, C.; and Cai, Y. 2026. ForgeDrive : Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving. arXiv:2606.31226

  52. [52]

    Zhou, E.; Su, Q.; Chi, C.; Zhang, Z.; Wang, Z.; Huang, T.; Sheng, L.; and Wang, H. 2024 a . Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection. arXiv:2412.04455

  53. [53]

    Zhou, G.; Pan, H.; LeCun, Y.; and Pinto, L. 2024 b . DINO-WM : World Models on Pre-trained Visual Features enable Zero-shot Planning. arXiv:2411.04983