REVIEW 3 major objections 5 minor 53 references
An open-loop action chunk is a testable prediction of future observations, and checking it with an action-conditioned world model restores feedback during long-horizon robot execution.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 21:02 UTC pith:GBFAZVPW
load-bearing objection Solid sim systems paper: action-conditioned verification plus latency-aware suffix repair beats matched periodic replanning by a real margin, with unusually careful controls; novelty is integrative and evidence stays inside RoboCasa365. the 3 major comments →
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A committed action chunk is both a control command and a testable prediction of near-future observations. Verifying that prediction online with a frozen action-conditioned world model, a conformally calibrated first-intervention threshold, risk-adaptive suffix retention, latency-aware hard prefixing, and event-driven episodic memory restores closed-loop feedback during open-loop chunk execution and, at matched policy-call budget on RoboCasa365, raises average success from 27.6% (periodic replanning) to 36.1%.
What carries the argument
CheckVLA’s action-conditioned verifier: a short rolling world-model forecast of observation features given remaining committed actions, scored by a causal risk head and compared to a split functional conformal threshold that bounds the episode-level probability of an unnecessary first intervention; threshold exceedance then sets how strongly the rewritten, latency-feasible suffix retains the old chunk.
Load-bearing premise
New successful runs must stay exchangeable with the fixed calibration pipeline so the risk threshold really bounds unnecessary first interruptions; the paper only claims that bound for clean successes, not for recall, post-repair safety, or real-world shift.
What would settle it
On the same backbone and call budget, if an observation-only or action-shuffled monitor matched the full verifier’s timely recall and perturbed success at the same 5% episode false-alarm rate, or if CheckVLA no longer beat invocation-matched periodic replanning on RoboCasa365, the central claim would fail.
If this is right
- Chunked VLA systems can regain mid-chunk feedback without querying the full policy at every step.
- Action–consequence mismatch detects confidently wrong open-loop failures that commit-time policy uncertainty misses.
- Repair strength should scale with calibrated risk exceedance, not only with a binary replan flag.
- Inference latency must hard-constrain which suffix is rewritten, or correct alarms still cannot recover the episode.
- Persistent keyframe memory is needed so repairs do not erase completed subgoals that left the camera view.
Where Pith is reading between the lines
- Any robot stack that amortizes compute via open-loop action segments—not only kitchen VLAs—could adopt the same predict-then-verify loop.
- Hardware transfer will hinge less on new architecture than on redoing conformal calibration under real sensing noise and timing jitter.
- Separating a frozen monitor from the policy path suggests safety auditors could certify the verifier without retraining the controller.
- If world-model features stay frozen while tasks shift, the next bottleneck is representation drift, not the conformal math.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. CheckVLA reframes open-loop VLA action chunks as testable predictions of near-future observations and verifies them at execution time with a separately trained, frozen action-conditioned world model. A causal risk head aggregates prediction–observation discrepancies; a split functional conformal threshold bounds the episode-level probability of an unnecessary first intervention on exchangeable nominal successes (Eqs. 6–7). On a trigger, the same VLA rewrites only the latency-feasible suffix under hard prefixing, with exceedance-conditioned retention of the superseded chunk, while an event-driven keyframe bank preserves episodic context. On RoboCasa365 under a common Human300 recipe and matched invocation budget, the system reports 36.1% average success versus 27.6% for periodic replanning (+8.5 points; Table 3). At matched ~5% episode FWER, action conditioning raises timely recall to 77.9% versus 48.6% observation-only and 37.9% action-shuffled controls (Table 4).
Significance. If the controlled results hold, the paper makes a clear systems contribution: action-conditioned consequence checking can restore feedback inside committed chunks without paying full policy-rate inference, and the repair can be made latency-consistent. Strengths include the matched-invocation build-up (Table 3), independently recalibrated detectors at a common episode FWER (Table 4), fixed-trigger repair branches from identical snapshots (Table 6), action-shuffled and observation-only negatives, inference-time binding audits (Table S5), and unusually thorough leakage guards and failure taxonomy in the appendix. The work is a useful complement to policy scaling and adaptive chunking rather than a replacement for either. The main external limit is that all evidence is RoboCasa365 simulation with a frozen V-JEPA encoder and a finite perturbation family; the Discussion already flags recalibration needs under sensing or timing shift.
major comments (3)
- [Appendix G; Table 4; Abstract] Appendix G defines τ_irrev operationally as the earliest time after which no tested suffix repair restores success, and counts a detection timely only if t*+d_lat < τ_irrev. Absolute timely recall (77.9% in Table 4) and much of the perturbed-success narrative therefore measure detector–repair co-adaptation under the authors’ rewrite set, not recoverability independent of that set. The shared-rewriter detector ranking remains informative, but the main text (Abstract, Experiments, contributions) should state this coupling explicitly wherever timely recall is headline-quoted, and preferably report a repair-agnostic ranking metric (e.g., event AUPRC / lead before any rewrite) beside it.
- [Table 3; Table S8; Appendix D, F] The risk head is supervised on physics-level impulses, joint offsets, and object displacements (Appendix D), and the locked perturbation study uses the same families (Appendix F). LOFO and the natural-execution audit (Tables S8, S10) partially address this, but on held-out natural configs the lift over invocation-matched periodic replanning shrinks to +4.1 points (68.9% vs 64.8%) versus +8.5 on the clean build-up (Table 3). The central action-conditioning gap is still credible under a shared rewriter; the recovery magnitude should be framed more cautiously in the Abstract and Conclusion as stack- and distribution-dependent, with the natural-execution numbers promoted into the main experimental narrative.
- [Eqs. 6–7; Abstract; Discussion; Table S17] Eqs. 6–7 give only a first-intervention guarantee on exchangeable nominal successes under a fully fixed pipeline. The paper states this correctly in the Method and Discussion, yet the Abstract’s phrasing (“bounds the episode-level probability of an unnecessary first intervention”) can be read as a deployed safety property. Please keep the bound claim tightly scoped in the Abstract and ensure Table S17’s repeated-intervention / post-repair harm numbers are cited whenever the conformal trigger is presented as the operating point.
minor comments (5)
- [Table 2; Experiments] Table 2’s “state of the art” positioning is labeled descriptive, which is appropriate; consider moving the controlled Table 3 comparison earlier in the Experiments narrative so readers do not overweight cross-system averages.
- [Figure 1; Appendix A] Figure 1 and Figure 2 are helpful; a short pointer in the main text to Appendix Figure S1 (latency hard-prefix schematic) would make d_lat / d_eff easier to parse on first reading.
- [Problem Formulation; Algorithm S1] Notation for span ℓ(t), active index h_τ, and effective delay d_eff is dense in the Problem Formulation; a one-line symbol table or tighter cross-references to Algorithm S1 would help.
- [Abstract; Introduction] Typos / spacing artifacts appear in several places (e.g., “topropagatetheerror”, “CheckVLA, which verifies”, missing spaces after periods in the Abstract and Introduction). A full copy-edit pass is needed.
- [Table 1] Table 1’s capability marks are useful but explicitly “not re-implementations”; a footnote restating that avoids over-reading the checkmarks as head-to-head results.
Circularity Check
No significant circularity: empirical systems results with held-out conformal fits and matched controls, not a derivation that re-labels inputs as predictions.
full rationale
CheckVLA is an empirical robotics systems paper. Its load-bearing claims are comparative success and detector metrics on RoboCasa365 under a common recipe, matched invocation budget, and independently recalibrated detectors—not a first-principles derivation chain. The functional conformal threshold (Eqs. 6–7) is the standard split/finite-sample construction: risk statistics and the quantile q̂_α are fit on nominal-success shadow trajectories disjoint from validation and locked tests, then the bound is stated only for unnecessary first interventions under exchangeability with that fixed pipeline. That is a statistical guarantee by design of conformal calibration, not a fitted quantity renamed as an independent prediction of the same data. Action-conditioning evidence uses negative controls (observation-only and action-shuffled predictors) retrained and recalibrated at the same episode FWER target; rewrite ablations branch from shared first-trigger snapshots so detector quality does not silently define repair wins. Validation-selected maps (exceedance→retention, periodic interval, discrepancy metric) are hyperparameters chosen once before locked tests, not claimed closed-form predictions. Self-citations to related VLA/world-model work are background and non-load-bearing for uniqueness. Operational definition of τ_irrev via the tested repair set couples absolute timely-recall magnitude to the authors’ rewriter, but that is a disclosed metric scope choice evaluated comparatively under a shared rewriter—not a circular reduction of a claimed derivation to its inputs. No step reduces Eq. X to Eq. Y by construction in the sense required here.
Axiom & Free-Parameter Ledger
free parameters (10)
- conformal miscoverage α =
0.05
- prediction span bound k =
8
- risk window w =
32
- exceedance sensitivity β and floor w_min =
β=1.38, w_min=0.0
- channel decay rates λ_mobility / λ_manip =
replan 0.60/0.90; regular 0.10/0.15
- scheduled latency d_lat =
3 control steps
- chunk length H =
50
- keyframe bank budget K and write-filter thresholds =
K=8; gap 10; similarity 0.95
- span-wise discrepancy stats (μ_d^(ℓ), σ_d^(ℓ)) and risk stats (μ_r,t, σ_r,t) =
fit on shadow-mode nominals (200 fit / 400 calib per seed)
- periodic replan interval =
39 control steps (~10.1 calls/ep)
axioms (6)
- standard math Split functional conformal prediction under exchangeability of nominal-success trajectories with a fully fixed score pipeline yields a trajectory-marginal bound on unnecessary first interventions (Eq. 7).
- domain assumption RoboCasa365 physics, rendering, and task success predicates are an adequate testbed for the claimed execution-time verification benefits.
- domain assumption A frozen V-JEPA 2-AC feature space plus short action-conditioned rollouts supply residuals that separate actionable deviations from benign noise when aggregated by R.
- domain assumption Inference latency can be treated as a known scheduled delay d_lat with hard prefix clamp as an empirical deployable projection.
- domain assumption Human300-only policy pretraining plus training-side auxiliary rollouts for the verifier introduce no leakage into Composite-Unseen and target-suite scenes.
- ad hoc to paper Validation-selected discrepancy metric, guidance map, and suppression rules generalize to locked clean and perturbation tests without retuning.
invented entities (4)
-
CheckVLA causal risk head R with episodic summary m(ρ_t)
no independent evidence
-
Exceedance-conditioned retention map W(e) for suffix guidance
no independent evidence
-
Event-driven keyframe bank B_t with dual policy/risk readers
no independent evidence
-
Latency-aware hard-prefixed flow rewrite of the superseded chunk
no independent evidence
read the original abstract
Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation while the remaining actions continue to propagate the error: commit-time policy confidence cannot react to a deviation that occurs after dispatch, and observation-only anomaly scores lack an action-conditioned reference for separating expected effects from unexplained changes. We propose CheckVLA, which verifies execution with a separately trained, frozen action-conditioned world model. A conformally calibrated risk threshold bounds the episode-level probability of an unnecessary first intervention and determines when to intervene, its exceedance controls how strongly the rewritten suffix retains the superseded chunk, latency-aware hard prefixing restricts replacement to actions that remain deployable, and an event-driven keyframe bank preserves evidence of prior progress across repairs. On RoboCasa365, under a common training recipe and a matched invocation budget, CheckVLA attains a 36.1% average success rate against 27.6% for periodic replanning (+8.5 points). At a matched 5% episode-level false-alarm target, action conditioning raises timely recall to 77.9%, against 48.6% for an observation-only control and 37.9% for an action-shuffled control. These simulation results support action-conditioned verification as a way to restore feedback during chunked execution while keeping the repair consistent with inference latency.
Figures
Reference graph
Works this paper leans on
-
[1]
Assran, M.; Bardes, A.; Fan, D.; Garrido, Q.; Howes, R.; Komeili, M.; Muckley, M.; Rizvi, A.; Roberts, C.; Sinha, K.; et al. 2025. V-JEPA 2 : Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv:2506.09985
Pith/arXiv arXiv 2025
-
[2]
Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. 2024. _0 : A Vision-Language-Action Flow Model for General Robot Control. arXiv:2410.24164
Pith/arXiv arXiv 2024
-
[3]
Black, K.; Galliker, M. Y.; and Levine, S. 2025. Real-Time Execution of Action Chunking Flow Policies. arXiv:2506.07339
Pith/arXiv arXiv 2025
-
[4]
Cen, J.; Yu, C.; Yuan, H.; Jiang, Y.; Huang, S.; Guo, J.; Li, X.; Song, Y.; Luo, H.; Wang, F.; Zhao, D.; and Chen, H. 2025. WorldVLA : Towards Autoregressive Action World Model. arXiv:2506.21539
Pith/arXiv arXiv 2025
-
[5]
Chao, X.; Mu, S.; Liu, Y.; Li, S.; Lyu, C.; Zhang, X.-P.; and Ding, W. 2025. Exo-ViHa : A Cross-Platform Exoskeleton System with Visual and Haptic Feedback for Efficient Dexterous Skill Learning. arXiv:2503.01543
Pith/arXiv arXiv 2025
-
[6]
Chen, J. 2026. AEGIS : A Backup Reflex for Physical AI . arXiv:2606.06660
Pith/arXiv arXiv 2026
-
[7]
Chen, R.; Yang, Y.; Tang, Z.; Huo, D.; Lin, T.; Wu, H.; Liu, H.; Chen, Y.; et al. 2026. ABot-M0.5 : Unified Mobility-and-Manipulation World Action Model. arXiv:2607.00678
Pith/arXiv arXiv 2026
-
[8]
Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; and Song, S. 2023. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv:2303.04137
Pith/arXiv arXiv 2023
-
[9]
T.; Ichter, B.; Yu, L.; Li-Bell, A.; Pertsch, K.; Ren, A
Driess, D.; Springenberg, J. T.; Ichter, B.; Yu, L.; Li-Bell, A.; Pertsch, K.; Ren, A. Z.; Walke, H.; Vuong, Q.; Shi, L. X.; and Levine, S. 2025. Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better. arXiv:2505.23705
Pith/arXiv arXiv 2025
-
[10]
Fan, J.; Liu, Y.; Li, S.; Ren, B.; Li, S.; Zhang, X.-P.; Ding, W.; and Deng, Z. 2026. FUTURE-VLA : Forecasting Unified Trajectories Under Real-time Execution. arXiv:2602.15882
arXiv 2026
-
[11]
Fu, Z.; Zhao, T. Z.; and Finn, C. 2024. Mobile ALOHA : Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. arXiv:2401.02117
Pith/arXiv arXiv 2024
-
[12]
Gu, Q.; Ju, Y.; Sun, S.; Gilitschenski, I.; Nishimura, H.; Itkina, M.; and Shkurti, F. 2025. SAFE : Multitask Failure Detection for Vision-Language-Action Models. arXiv:2506.09937
arXiv 2025
-
[13]
Guo, Y.; Wang, Y.-J.; Zha, L.; and Chen, J. 2023. DoReMi : Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment. arXiv:2307.00329
Pith/arXiv arXiv 2023
-
[14]
Hafner, D.; Lillicrap, T.; Ba, J.; and Norouzi, M. 2019. Dream to Control: Learning Behaviors by Latent Imagination. arXiv:1912.01603
Pith/arXiv arXiv 2019
-
[15]
Hafner, D.; Lillicrap, T.; Fischer, I.; Villegas, R.; Ha, D.; Lee, H.; and Davidson, J. 2018. Learning Latent Dynamics for Planning from Pixels. arXiv:1811.04551
Pith/arXiv arXiv 2018
-
[16]
Ho, M.; Ginting, M. F.; Ward, I. R.; Reinke, A.; Kochenderfer, M. J.; Agha-Mohammadi, A.-a.; and Omidshafiei, S. 2026. World Model Failure Classification and Anomaly Detection for Autonomous Inspection. arXiv:2602.16182
arXiv 2026
-
[17]
Jiang, H.; Chen, J.; Bu, Q.; Chen, L.; Shi, M.; Zhang, Y.; Li, D.; Suo, C.; Wang, C.; Peng, Z.; and Li, H. 2025. WholeBodyVLA : Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control. arXiv:2512.11047
arXiv 2025
-
[18]
Jing, D.; Wang, G.; Liu, J.; Tang, W.; Sun, Z.; Yao, Y.; Wei, Z.; Liu, Y.; Lu, Z.; and Ding, M. 2025. Mixture of Horizons in Action Chunking. arXiv:2511.19433
Pith/arXiv arXiv 2025
-
[19]
Kim, D.; Jang, H.; Koo, M.; Jang, S.; Kim, T.; Kim, B.; Yoon, B.; Jang, C.; Choi, D.; Han, D.; et al. 2026. RLDX-1 Technical Report. arXiv:2605.03269
Pith/arXiv arXiv 2026
-
[20]
Liu, Y.; Feng, F.; Kong, L.; Lu, W.; Tang, J.; Zhang, K.; Murphy, K.; Finn, C.; and Du, Y. 2026 a . World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry. arXiv:2604.01985
Pith/arXiv arXiv 2026
-
[21]
I.; Xie, A.; Lee, Y.; Du, M.; and Finn, C
Liu, Y.; Hamid, J. I.; Xie, A.; Lee, Y.; Du, M.; and Finn, C. 2024. Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling. arXiv:2408.17355
Pith/arXiv arXiv 2024
-
[22]
Liu, Y.; Lv, T.; Wang, B.; Fan, H.; Zhao, C.; Zheng, H.; Zhong, X.; Xie, Y.; Zhao, C.; Liao, Z.; Luo, L.; Cai, Y.; Zhang, X.-P.; and Ding, W. 2026 b . PerceptDrive : Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving. arXiv:2607.20175
Pith/arXiv arXiv 2026
-
[23]
Liu, Y.; Mu, S.; Chao, X.; Li, Z.; Mu, Y.; Chen, T.; Li, S.; Lyu, C.; Zhang, X.-P.; and Ding, W. 2025. AVR : Active Vision-Driven Precise Robot Manipulation with Viewpoint and Focal Length Optimization. arXiv:2503.01439
arXiv 2025
-
[24]
Liu, Y.; Sun, P.; Li, S.; Xie, Y.; Zhang, L.; Chao, X.; Dong, S.; Chen, F.; Zhang, X.-P.; and Ding, W. 2026 c . OA-WAM : Object-Addressable World Action Model for Robust Robot Manipulation. arXiv:2605.06481
Pith/arXiv arXiv 2026
-
[25]
Liu, Z.; Bahety, A.; and Song, S. 2023. REFLECT : Summarizing Robot Experiences for Failure Explanation and Correction. arXiv:2306.15724
Pith/arXiv arXiv 2023
-
[26]
Nasiriany, S.; Nasiriany, S.; Maddukuri, A.; and Zhu, Y. 2026. RoboCasa365 : A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots. arXiv:2603.04356
arXiv 2026
-
[27]
Bjorck, J.; Blukis, V.; Casta\ neda, F.; Cherniadev, N.; Da, X.; Ding, R.; Fan, L.; Fang, Y.; Fox, D.; Hu, F.; et al. 2025. GR00T N1.5 : An Improved Open Foundation Model for Generalist Humanoid Robots. NVIDIA GEAR, 11 June 2025. https://research.nvidia.com/labs/gear/gr00t-n1_5/
2025
-
[28]
GEAR Team ; Azzolini, A.; Bjorck, J.; Blukis, V.; Casta\ neda, F.; Chand, R.; Chang, Y.; Chen, D.; Cherniadev, N.; Da, X.; et al. 2025. GR00T N1.6 : An Improved Open Foundation Model for Generalist Humanoid Robots. NVIDIA GEAR, 15 December 2025. https://research.nvidia.com/labs/gear/gr00t-n1_6/
2025
-
[29]
Pan, Y.; Pan, M.; Lu, Q.; Huang, J.; Zhang, M.; Huang, S.; Li, X.; Zhang, J.; Shen, Y.; Zhang, X.; and Zhang, W. 2026. VLA-Corrector : Lightweight Detect-and-Correct Inference for Adaptive Action Horizon. arXiv:2607.01804
Pith/arXiv arXiv 2026
-
[30]
Physical Intelligence ; Ai, B.; Amin, A.; Aniceto, R.; Balakrishna, A.; Balke, G.; Black, K.; Bokinsky, G.; Cao, S.; et al. 2026. _ 0.7 : a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv:2604.15483
Pith/arXiv arXiv 2026
-
[31]
Physical Intelligence ; Amin, A.; Aniceto, R.; Balakrishna, A.; Black, K.; Conley, K.; Connors, G.; Darpinian, J.; Dhabalia, K.; et al. 2025 a . ^ * _ 0.6 : a VLA That Learns From Experience. arXiv:2511.14759
Pith/arXiv arXiv 2025
-
[32]
Physical Intelligence ; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; et al. 2025 b . _ 0.5 : A Vision-Language-Action Model with Open-World Generalization. arXiv:2504.16054
Pith/arXiv arXiv 2025
-
[33]
Rao, P.; Zhang, W.; Balestriero, R.; LeCun, Y.; and Loianno, G. 2026. SkyJEPA : Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors. arXiv:2606.23444
Pith/arXiv arXiv 2026
-
[34]
Ross, S.; and Bagnell, D. 2010. Efficient Reductions for Imitation Learning. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (AISTATS), 661--668. PMLR 9
2010
-
[35]
Ross, S.; Gordon, G. J.; and Bagnell, J. A. 2011. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS), 627--635. PMLR 15. arXiv:1011.0686
Pith/arXiv arXiv 2011
-
[36]
Spencer, J.; Choudhury, S.; Venkatraman, A.; Ziebart, B.; and Bagnell, J. A. 2021. Feedback in Imitation Learning: The Three Regimes of Covariate Shift. arXiv:2102.02872
Pith/arXiv arXiv 2021
-
[37]
Sun, J.; Zhang, W.; Qi, Z.; Ren, S.; Liu, Z.; Zhu, H.; Sun, G.; Jin, X.; and Chen, Z. 2026 a . VLA-JEPA : Enhancing Vision-Language-Action Model with Latent World Model. arXiv:2602.10098
arXiv 2026
-
[38]
Sun, Z.; Guo, Y.; Sun, H.; Wang, L.; Lu, W.; Ji, J.; Ji, S.; Xiong, J.; and Meng, Z. 2026 b . Pre-VLA : Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts. arXiv:2605.22446
Pith/arXiv arXiv 2026
-
[39]
Z.; Wang, H.; Tang, J.; Stachowicz, K.; et al
Torne, M.; Pertsch, K.; Walke, H.; Vedder, K.; Nair, S.; Ichter, B.; Ren, A. Z.; Wang, H.; Tang, J.; Stachowicz, K.; et al. 2026. MEM : Multi-Scale Embodied Memory for Vision Language Action Models. arXiv:2603.03596
arXiv 2026
-
[40]
Wang, H.; Zhang, G.; Yan, Y.; Kompella, R. R.; and Liu, G. 2026 a . VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies. arXiv:2602.21445
Pith/arXiv arXiv 2026
-
[41]
Wang, H.; Zhang, G.; Yan, Y.; Shang, Y.; Kompella, R. R.; and Liu, G. 2026 b . Real-Time Robot Execution with Masked Action Chunking. arXiv:2601.20130
arXiv 2026
-
[42]
Wang, R.; Zhang, Y.; Lin, J.; Luo, K.; Wang, J.; Wang, Z.; and Qi, X. 2026 c . When to Trust Imagination: Adaptive Action Execution for World Action Models. arXiv:2605.06222
Pith/arXiv arXiv 2026
-
[43]
Wang, T.; Hou, H.; Hu, Y.; Liu, Y.; Li, Q.; Jiang, Y.; Wang, Y.; Ma, C.; Wang, R.; and Gao, Y. 2026 d . When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning. arXiv:2607.25593
Pith/arXiv arXiv 2026
-
[44]
Wang, X.; Zhu, Z.; Huang, G.; Wang, B.; Chen, X.; and Lu, J. 2024. WorldDreamer : Towards General World Models for Video Generation via Predicting Masked Tokens. arXiv:2401.09985
Pith/arXiv arXiv 2024
-
[45]
Ye, A.; Wang, B.; Ni, C.; Huang, G.; Zhao, G.; Li, H.; Li, H.; Li, J.; Lv, J.; Liu, J.; Cao, M.; Li, P.; Deng, Q.; Mei, W.; Wang, X.; Chen, X.; Zhou, X.; Wang, Y.; Chang, Y.; Li, Y.; Zhou, Y.; Ye, Y.; Liu, Z.; and Zhu, Z. 2026. GigaWorld-Policy : An Efficient Action-Centered World--Action Model. arXiv:2603.17240
arXiv 2026
-
[46]
Yuan, H.; Liang, Z.; Chen, A.; Wang, Y.; Li, H.; Lin, P.; Huang, Y.; Lei, Z.; Zhang, T.; Zhang, J.; Zhang, J.; Fan, J.; Zhou, G.; Peng, Q.; Lv, C.; Chen, X.; Yang, A.; Huang, F.; Lin, J.; Liu, D.; Zhou, J.; Wu, C.; and Chen, X.-H. 2026. Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models. arXiv:2606.17846
Pith/arXiv arXiv 2026
-
[47]
Zeng, Y.; Ye, M.; Chen, Y.; Shentu, Y.; Wu, P.; Yan, Z.; and Li, Z. 2026. KEMO : Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies. arXiv:2606.23589
Pith/arXiv arXiv 2026
-
[48]
Zhang, H.; Lu, Y.; Wang, B.; Kang, X.; Kuo, Y.-L.; Cheng, Z.; Wang, M.; and Jenkins, O. C. 2026. Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents. arXiv:2606.23085
Pith/arXiv arXiv 2026
-
[49]
Zhang, W.; Liu, H.; Qi, Z.; Wang, Y.; Yu, X.; Zhang, J.; Dong, R.; He, J.; Lu, F.; Wang, H.; et al. 2025. DreamVLA : A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge. arXiv:2507.04447
Pith/arXiv arXiv 2025
-
[50]
Z.; Kumar, V.; Levine, S.; and Finn, C
Zhao, T. Z.; Kumar, V.; Levine, S.; and Finn, C. 2023. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv:2304.13705
Pith/arXiv arXiv 2023
-
[51]
Zhong, X.; Zheng, H.; Zhao, C.; Lv, T.; Fan, H.; Wang, B.; Liu, Y.; Gao, L.; Liao, Z.; Luo, L.; Zhao, C.; and Cai, Y. 2026. ForgeDrive : Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving. arXiv:2606.31226
Pith/arXiv arXiv 2026
-
[52]
Zhou, E.; Su, Q.; Chi, C.; Zhang, Z.; Wang, Z.; Huang, T.; Sheng, L.; and Wang, H. 2024 a . Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection. arXiv:2412.04455
Pith/arXiv arXiv 2024
-
[53]
Zhou, G.; Pan, H.; LeCun, Y.; and Pinto, L. 2024 b . DINO-WM : World Models on Pre-trained Visual Features enable Zero-shot Planning. arXiv:2411.04983
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.