Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Frame-only robot policies can flip short-horizon intents when observations look alike; conditioning action chunks on a compact history-derived intent keeps consecutive chunks consistent and raises success.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 18:58 UTC pith:T3KETLZK

load-bearing objection Useful failure-mode framing plus a real aliasing benchmark; the method is solid short-memory engineering, not a deep new theory of intent. the 3 major comments →

arxiv 2605.14712 v2 pith:T3KETLZK submitted 2026-05-14 cs.RO cs.AIcs.CLcs.CV

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

classification cs.RO cs.AIcs.CLcs.CV
keywords vision-language-actionrobot manipulationshort-horizon intentobservation aliasingimitation learningchunked policieshistory conditioningAliasBench
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Human robot demonstrations are multimodal across episodes but locally committed within an episode: the same image and instruction can legitimately lead to different next action chunks depending on the current phase, path, or recent context. Policies that generate each chunk from only the current frame often re-sample a different intent at every replan, so overlapping chunks conflict and execution becomes unstable. IntentVLA encodes a short window of recent images into a compact short-horizon intent representation, fuses it with the current vision-language context, and conditions the action generator on that fused evidence. The authors also introduce AliasBench, twelve matched training-and-evaluation tasks built to isolate exactly this short-horizon observation aliasing. Across AliasBench and standard simulation suites the history-conditioned policy improves both success rate and inter-chunk consistency over strong frame-only and raw-history baselines.

Core claim

Under short-horizon observation aliasing, conditioning chunked vision-language-action policies on a compact intent representation extracted from recent visual history stabilizes local continuations and outperforms frame-conditioned and raw-history baselines in success and inter-chunk consistency on AliasBench and on SimplerEnv, LIBERO, and RoboCasa.

What carries the argument

The short-horizon intent representation: a frozen geometry encoder turns recent head-camera frames into camera and register tokens; gated cross-attention fuses those tokens into the current vision-language context and an appended compact history-evidence token conditions a flow-matching action head for the next chunk.

Load-bearing premise

A fixed short window of recent head-camera images, encoded only as frozen geometry tokens, is enough at test time to recover the episode’s already-chosen next step even when the robot’s own mistakes make that history look different from the demonstrations.

What would settle it

On AliasBench ambiguity windows, train and evaluate the same short-history intent module against a matched frame-only baseline; if inter-chunk action disagreement (ICC-L2) does not fall and average success stays near the frame-only level, the claim that short-horizon intent conditioning resolves aliasing is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Frame-conditioned chunk VLAs systematically fail when the same visual state recurs with different local goals.
  • Compact history-derived intent conditioning is more effective and memory-efficient than stuffing raw past frames into the language backbone.
  • Inter-chunk consistency on overlapping actions is a practical diagnostic of whether a policy has committed to one short-horizon continuation.
  • Controlled aliasing benchmarks can expose a failure mode that average success on saturated standard suites often hides.
  • Short visual memory alone can improve multi-stage and partially observed manipulation without an explicit long-horizon planner.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Policies that look strong on near-saturated benchmarks may still be fragile in real kitchens or factories where phases and handoffs reappear under partial views.
  • An ambiguity detector that lengthens or refreshes history only when the current frame is mixed could shrink the remaining closed-loop history-shift failures.
  • The same compact intent-token idea may transfer to multi-robot or human-robot handoff settings where visual symmetry creates analogous short-horizon aliasing.
  • Explicit intent labels appear unnecessary if the right compact history features are fused into the action conditioner.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper argues that frame-conditioned chunked VLAs are unstable under short-horizon observation aliasing because demonstrations are multimodal across episodes but locally committed within an episode, while current-frame conditioning can resample different continuations across replans. It introduces IntentVLA, which freezes a VGGT history encoder, retains camera and register tokens from a short visual window, fuses them into a Qwen3-VL context via gated cross-attention, appends a pooled history-evidence token, and conditions a DiT flow-matching action head. It also introduces AliasBench, a 12-task RoboTwin2 benchmark with matched data isolating back-and-forth, crossing-path, bimanual, and multi-goal aliasing, plus an ICC-L2 inter-chunk consistency metric in annotated ambiguity windows. Empirically, IntentVLA raises AliasBench average success from 9.0% (Qwen3VL-GR00T) and 28.1% (best feasible raw-history baseline) to 45.8%, reduces mean ICC-L2 by 17.6%, and improves SimplerEnv (72.9%), LIBERO-Long (97.4%), and RoboCasa (57.0%), with component ablations on SimplerEnv.

Significance. If the result holds, the paper makes two useful contributions for chunked VLA control: (i) a cleanly motivated failure mode—uncommitted multimodality under aliased current observations—and (ii) a practical short-memory design that is more efficient than stuffing raw past frames into the VLM context. AliasBench is a genuine evaluation contribution: matched training/eval, family-level construction around latent factors (phase/source/handoff/target), and a quantitative nearest-neighbor aliasing diagnostic (Fig. 3) make the claim falsifiable rather than purely narrative. The ICC-L2 metric in annotated ambiguity windows is a concrete stability probe beyond success rate. Transfer gains on SimplerEnv, LIBERO-Long, and RoboCasa, plus ablations showing that current-frame VGGT alone does not help while history fusion and the compact evidence token do, strengthen the engineering case. Code is promised at a public GitHub URL. The work is simulation-only and leaves long-horizon memory and closed-loop history shift open, but the short-horizon framing is appropriately scoped.

major comments (3)
  1. §4.1–4.3 and §5.1 / Table 1: The central causal claim is that gains come from recovering the episode’s already-committed short-horizon continuation (phase/source/handoff/active target), not merely from adding short-horizon geometry or motion features. AliasBench is designed so that o_t is aliased while recent history carries the latent factor (Eqs. 1–2; Fig. 3), but the reported ablations do not break that link. Table 5 (SimplerEnv only) shows history fusion and the intent token help and current-frame VGGT does not; it does not test, on AliasBench families, histories that omit or scramble the disambiguating cue (e.g., same-length history from a different phase/source, temporally reversed history, or history without the origin/handoff frames). Without such controls, the jump from 9.0%/28.1% to 45.8% and the ICC-L2 drop can still be read as generic temporal enrichment. A family-level contr
  2. §5.1, Eq. (13) and Fig. 5: ICC-L2 is a good action-level proxy for inter-chunk commitment, but the manuscript does not fully specify the evaluation protocol needed to interpret it as intent consistency. Please state explicitly: (i) whether ICC is computed on all rollouts or only successful ones; (ii) the replan interval r and chunk horizon H used; (iii) how ambiguity windows are annotated and whether they are fixed from demonstration structure or detected online; and (iv) whether lower ICC could arise from more conservative/smoother actions rather than correct latent-factor recovery. Correlating ICC reduction with success within each AliasBench family (or reporting ICC conditional on correct vs incorrect continuation) would make the stability claim tighter.
  3. §5.1 Table 1 and Limitations §7: AliasBench average success remains 45.8%, with bimanual at 17.0% and multi-goal at 31.3%. The paper attributes residual failure partly to the fixed 16-frame window and closed-loop history shift. That is candid, but it is also the weakest assumption of the method (§4.2). The main claim would be stronger if the authors quantified sensitivity to K and to history corruption (e.g., replace recent frames with frames from a wrong-intent neighbor, or inject execution noise so test history diverges from demos) rather than only noting the gap. Even a small controlled study on one back-and-forth and one crossing-path task would show whether the compact intent token remains informative under the failure mode the paper itself highlights.
minor comments (6)
  1. §2.2: Several concurrent intent/memory VLA works are cited; a short table contrasting conditioning signal (current frame vs history), whether intent is supervised, and whether evaluation isolates aliasing would help readers place IntentVLA relative to DIAL, MINT, MemoryVLA, and Mem-0.
  2. Fig. 1 and Fig. 2: The qualitative aliasing examples are clear, but figure captions should state camera viewpoint(s) and whether the shown frames are from training demos or evaluation rollouts.
  3. §4.2: Clarify multi-view handling more precisely—history uses only head-camera frames while the current backbone may use the full current observation. State whether this asymmetry is fixed for all benchmarks and whether wrist/side cameras ever enter the history branch.
  4. Table 2: IntentVLA underperforms the frame-only baseline on Put Spoon on Towel (70.8 vs 83.0) while gaining elsewhere; the discussion in §5.3 is helpful—consider flagging this trade-off earlier when claiming broad robustness.
  5. Appendix A: The mode-switching analysis (Eq. 14) is useful conceptually; note more explicitly that P_switch is not estimated from the model and is only a diagnostic framing for ICC-L2.
  6. Presentation: The manuscript is marked “Work in progress” on the first page; for journal submission, remove that banner and ensure arXiv/version metadata are consistent. Also fix minor redundancy (VFP appears twice in the references).

Circularity Check

0 steps flagged

No significant circularity: standard imitation learning with held-out rollout metrics; mild same-group baseline citations are comparative, not load-bearing.

full rationale

IntentVLA’s chain is architectural and empirical, not a first-principles derivation that collapses into its inputs. The short-horizon intent representation mt = f_φ(ot, ℓ, hK_t) is learned without intent labels (Eqs. 6–12); training is ordinary conditional flow matching on demonstration chunks, while success and ICC-L2 are measured on held-out closed-loop rollouts and annotated ambiguity windows—neither is the training target restated as a prediction. AliasBench is constructed from task design (Eqs. 1–2, families in §3) and a separate embedding nearest-neighbor diagnostic (Fig. 3), not from fitting the policy’s own outputs. Same-group citations (StarVLA, LangForce, PhysBrain, TwinBrainVLA, 3D-Mix) appear as baselines or pipeline references in tables and training protocol; they do not supply a uniqueness theorem, ansatz, or forced result that the central claim reduces to. No fitted parameter is renamed a prediction, and no equation is definitionally equivalent to the reported gains. Score 1 only for non-load-bearing same-group baseline presence; the core claim remains independently evaluated.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

The central claim rests on standard imitation-learning and partial-observability assumptions plus several design choices that are not independently validated outside this paper: finite visual history as a proxy for latent short-horizon intent, frozen VGGT camera/register tokens as the right evidence, and simulation benchmarks as faithful tests of the failure mode. Free parameters are mostly ordinary training/architecture knobs; the invented entities are the compact intent representation and AliasBench itself.

free parameters (5)
  • history window length K (≈16 frames)
    Chosen design hyperparameter that defines how much recent context is available; paper notes remaining failures may come from limited temporal coverage.
  • VGGT token selection (1 camera + 4 register tokens per frame)
    Hand-selected subset of encoder outputs used as history evidence; not learned end-to-end which tokens matter.
  • learned gate scalar α and projection matrices Wh, We
    Trained fusion parameters that control how much history enters the current context; fitted jointly with the policy objective.
  • training schedule (30K steps, lr 1e-5 cosine, batch 256 on 16 H100)
    Compute/optimization settings that affect reported success rates under the 'same budget' comparison.
  • chunk horizon H and replan interval r used for ICC-L2
    Execution/evaluation choices that define the inter-chunk consistency metric and thus part of the stability claim.
axioms (5)
  • domain assumption Demonstrations are multimodal across episodes but locally committed within an episode; frame-only conditioning can break that commitment under partial observability.
    Core motivation in §1 and formalized in §4.1 Eqs. (4)–(5); treated as the problem definition rather than proved.
  • domain assumption A finite recent visual history h^K_t is sufficient evidence to concentrate the short-horizon intent posterior for chunk generation.
    Stated as the modeling choice replacing full interaction history in §4.1; Limitations admit sparse long-range events may need more.
  • ad hoc to paper Frozen VGGT camera and register tokens capture viewpoint change and inter-frame structure useful for active short-horizon intent.
    Implementation claim in §4.2 without independent intent-label validation of those tokens.
  • domain assumption Conditional flow-matching DiT action heads with Euler integration (as in GR00T-style VLAs) are an adequate generation interface once conditioning context is improved.
    Inherited from cited VLA practice (§4.3); not re-derived.
  • domain assumption Simulation environments (RoboTwin2 AliasBench, SimplerEnv, LIBERO, RoboCasa) isolate and measure the intended aliasing failure mode.
    Evaluation premise throughout §3 and §5; real-robot transfer left to future work.
invented entities (3)
  • short-horizon intent representation m_t / condition context C_t no independent evidence
    purpose: Compact history-conditioned embedding used to stabilize chunk generation without explicit intent labels.
    Defined in §4.1–4.3 as the learned fusion of current VLM tokens with history evidence; only validated via policy success and ICC, not external intent probes.
  • AliasBench (12-task ambiguity-aware benchmark) no independent evidence
    purpose: Isolate short-horizon observation aliasing with matched train/eval data across four ambiguity families.
    New evaluation construct in §3 and Appendix B; diagnostic neighbor-mixing supports aliasing but the benchmark itself is paper-introduced.
  • ICC-L2 inter-chunk consistency metric in ambiguity windows no independent evidence
    purpose: Action-space proxy for intent mode switching across adjacent replans.
    Defined in §5.1 Eq. (13) and analyzed in Appendix A; useful but paper-specific operationalization of consistency.

pith-pipeline@v1.1.0-grok45 · 23078 in / 3874 out tokens · 35666 ms · 2026-07-14T18:58:31.586130+00:00 · methodology

0 comments
read the original abstract

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines

Figures

Figures reproduced from arXiv: 2605.14712 by Bin Yu, Changti Wu, Cong Huang, Haishan Liu, Hang Yuan, Kai Chen, Laurence Tianruo Yang, Shijie Lian, Xiaopeng Lin, Yurun Jin, Zhaolong Shen.

Figure 1
Figure 1. Figure 1: An illustrative example of short-horizon intent ambiguity under frame-only conditioning. The task is ordinary: the robot puts a piece of bread into a skillet for cooking and then returns it to the plate. The ambiguity appears because similar bread-in-gripper observations occur before two different continuations: placing the bread into the skillet and returning it to the plate. A frame-conditioned chunk pol… view at source ↗
Figure 2
Figure 2. Figure 2: Representative observation aliasing patterns in AliasBench. The quantitative observation-aliasing diagnostic is shown in [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Quantitative observation-aliasing diagnostic on AliasBench. Back-and-Forth uses intra-episode re￾trieval with a 20-frame temporal gap; all other families use cross-episode retrieval. The diagnostic is not a policy success metric. Instead, it measures whether visually nearby states in the ambiguity window can correspond to different next intents. Left: roughly half of the top-k neighbors (k = 5) come from a… view at source ↗
Figure 4
Figure 4. Figure 4: Overview of IntentVLA. A Qwen3-VL backbone encodes the current image and language instruction, while a frozen VGGT-1B history encoder extracts recent visual evidence. IntentVLA fuses the history tokens with the current visual-language context through gated cross-attention, appends a compact short-horizon intent token, and conditions a DiT-based flow-matching action head for chunk generation. and predicts a… view at source ↗
Figure 5
Figure 5. Figure 5: Inter-chunk consistency in AliasBench ambiguity windows. We compare IntentVLA against the strongest feasible history-as-context baseline in [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

    cs.RO 2026-07 conditional novelty 6.0

    Selecting 50% of robot demonstrations by maximizing exposure to reusable primitive-transition patterns outperforms full-data training while halving training steps.

Reference graph

Works this paper leans on

55 extracted references · 18 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Hongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan, Zhizhong Su, Hang Su, and Jun Zhu. 2025. H-RDT: Human manipulation enhanced bimanual robotic manipulation.arXiv preprint arXiv:2507.23523

  2. [2]

    Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, and 1 others. 2025. GR00T N1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734

  3. [3]

    2024.π 0: A vision- language-action flow model for general robot control.arXiv preprint arXiv:2410.24164

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, and 5 others. 2024.π 0: A vision- language-action flow model for general robot cont...

  4. [4]

    Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xu Huang, Shu Jiang, and 1 others. 2025. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669

  5. [5]

    Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. 2025. Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111

  6. [6]

    Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, and 1 others. 2025. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088

  7. [7]

    Tianxing Chen, Yuran Wang, Mingleyang Li, Yan Qin, Hao Shi, Zixuan Li, Yifan Hu, Yingsheng Zhang, Kaixuan Wang, Yue Chen, and 1 others. 2026. Rmbench: Memory-dependent robotic manipulation benchmark with insights into policy design.arXiv preprint arXiv:2603.01229

  8. [8]

    Yi Chen, Yuying Ge, Hui Zhou, Mingyu Ding, Yixiao Ge, and Xihui Liu. 2026. Dial: Decoupling intent and action via latent world modeling for end-to-end vla.arXiv preprint arXiv:2603.29844

  9. [9]

    StarVLA Community. 2026. Starvla: A lego-like codebase for vision-language-action model developing.arXiv preprint arXiv:2604.05014

  10. [10]

    GEAR-Team, Allison Azzolini, Johan Bjorck, Valts Blukis, Fernando Castañeda, Rahul Chand, and 1 others

  11. [11]

    nvidia.com/labs/gear/gr00t-n1_6/

    Gr00t n1.6: An improved open foundation model for generalist humanoid robots.https://research. nvidia.com/labs/gear/gr00t-n1_6/

  12. [12]

    Renming Huang, Chendong Zeng, Wenjing Tang, Jintian Cai, Cewu Lu, and Panpan Cai. 2026. Mimic intent, not just trajectories.arXiv preprint arXiv:2602.08602

  13. [13]

    Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, and 17 others

    Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y . Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, and 17 others. 2025.π0.5: a vision-language-action model with open-world...

  14. [14]

    Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, and 1 others. 2024. Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945

  15. [15]

    Moo Jin Kim, Chelsea Finn, and Percy Liang. 2025. Fine-tuning vision-language-action models: Optimizing speed and success.arXiv preprint arXiv:2502.19645

  16. [16]

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P Foster, Pannag R Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. 2024. OpenVLA: An open-source vision-language- action model. InConference on Robot Learning (CoRL)

  17. [17]

    Chengmeng Li, Junjie Wen, Yaxin Peng, Yan Peng, and Yichen Zhu. 2026. Pointvla: Injecting the 3d world into vision-language-action models.IEEE Robotics and Automation Letters, 11(3):2506–2513

  18. [18]

    Fuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang, Pengxiang Ding, Donglin Wang, Long Zeng, and Haoang Li. 2025. Spatial forcing: Implicit spatial representation alignment for vision-language-action model.arXiv preprint arXiv:2510.12276. 13

  19. [19]

    Peiyan Li, Yixiang Chen, Hongtao Wu, Xiao Ma, Xiangnan Wu, Yan Huang, Liang Wang, Tao Kong, and Tie- niu Tan. 2025. Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models. InAdvances in neural information processing systems (NeurIPS)

  20. [20]

    Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizhong Zhang, and 1 others. 2024. CogACT: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation.arXiv preprint arXiv:2411.19650

  21. [21]

    Xinghang Li, Peiyan Li, Minghuan Liu, Dong Wang, Jirong Liu, Bingyi Kang, Xiao Ma, Tao Kong, Hanbo Zhang, and Huaping Liu. 2024. Towards generalist robot policies: What matters in building vision-language- action models.arXiv preprint arXiv:2412.14058

  22. [22]

    Xuanlin Li, Kyle Hsu, Jiayuan Gu, Oier Mees, Karl Pertsch, Homer Rich Walke, Chuyuan Fu, Ishikaa Lu- nawat, Isabel Sieh, Sean Kirmani, Sergey Levine, Jiajun Wu, Chelsea Finn, Hao Su, Quan Vuong, and Ted Xiao. 2024. SimplerEnv: Evaluating real-world robot manipulation policies in simulation. InConference on Robot Learning (CoRL)

  23. [23]

    Shijie Lian, Bin Yu, Xiaopeng Lin, Laurence T Yang, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Cong Huang, and Kai Chen. 2026. Langforce: Bayesian decomposition of vision language action models via latent action queries.arXiv e-prints, pages arXiv–2601

  24. [24]

    Tao Lin, Gen Li, Yilei Zhong, Yanwen Zou, Yuxin Du, Jiting Liu, Encheng Gu, and Bo Zhao. 2025. Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416

  25. [25]

    Xiaopeng Lin, Shijie Lian, Bin Yu, Ruoqi Yang, Changti Wu, Yuzhuo Miao, Yurun Jin, Yukun Shi, Cong Huang, Bojun Cheng, and 1 others. 2025. PhysBrain: Human egocentric data as a bridge from vision language models to physical intelligence.arXiv preprint arXiv:2512.16793

  26. [26]

    Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. 2023. LIBERO: Benchmarking knowledge transfer for lifelong robot learning.Advances in neural information processing sys- tems (NeurIPS), 36:44776–44791

  27. [27]

    Songming Liu, Bangguo Li, Kai Ma, Lingxuan Wu, Hengkai Tan, Xiao Ouyang, Hang Su, and Jun Zhu

  28. [28]

    Rdt2: Exploring the scaling limit of umi data towards zero-shot cross-embodiment generalization.arXiv preprint arXiv:2602.03310

  29. [29]

    Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. 2025. RDT-1b: a diffusion foundation model for bimanual manipulation. InInternational Conference on Learning Representations (ICLR)

  30. [30]

    Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. InInternational Conference on Learning Representations (ICLR)

  31. [31]

    Yulin Luo, Hao Chen, Zhuangzhe Wu, Bowen Sui, Jiaming Liu, Chenyang Gu, Zhuoyang Liu, Qiuxuan Feng, Jiale Yu, Shuo Gu, and 1 others. 2026. Look before acting: Enhancing vision foundation representations for vision-language-action models.arXiv preprint arXiv:2603.15618

  32. [32]

    Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Man- dlekar, and Yuke Zhu. 2024. RoboCasa: Large-scale simulation of everyday tasks for generalist robots. In Robotics: Science and Systems

  33. [33]

    Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, and 1 others. 2024. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE

  34. [34]

    Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. 2025. FAST: Efficient action tokenization for vision-language-action models.arXiv preprint arXiv:2501.09747

  35. [35]

    Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Yan Ding, Zhigang Wang, JiaYuan Gu, Bin Zhao, Dong Wang, and 1 others. 2025. Spatialvla: Exploring spatial representations for visual-language-action model.arXiv preprint arXiv:2501.15830

  36. [36]

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimiza- tions enable training deep learning models with over 100 billion parameters. InProceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, pages 3505–3506. 14

  37. [37]

    Yichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang, Yan Lu, Jiaolong Yang, Nanning Zheng, and Baining Guo. 2025. VideoVLA: Video generators can be generalizable robot manipulators. InAdvances in neural information processing systems (NeurIPS)

  38. [38]

    Hao Shi, Bin Xie, Yingfei Liu, Lin Sun, Fengrong Liu, Tiancai Wang, Erjin Zhou, Haoqiang Fan, Xiangyu Zhang, and Gao Huang. 2026. Memoryvla: Perceptual-cognitive memory in vision-language-action models for robotic manipulation. InInternational Conference on Learning Representations (ICLR)

  39. [39]

    Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren, Zezhi Liu, Hanxin Zhu, Guangzhong Sun, Xin Jin, and Zhibo Chen. 2026. Vla-jepa: Enhancing vision-language-action model with latent world model.arXiv preprint arXiv:2602.10098

  40. [40]

    Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, and 1 others. 2024. Octo: An open-source generalist robot policy.arXiv preprint arXiv:2405.12213

  41. [41]

    Homer Rich Walke, Kevin Black, Tony Z Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen-Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, and 1 others. 2023. Bridgedata v2: A dataset for robot learning at scale. InConference on Robot Learning (CoRL), pages 1723–1736. PMLR

  42. [42]

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny

  43. [43]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

    Vggt: Visual geometry grounded transformer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  44. [44]

    Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, and 1 others. 2025. Magma: A foundation model for multimodal ai agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14203– 14214

  45. [45]

    Bin Yu, Shijie Lian, Xiaopeng Lin, Zhaolong Shen, Yuliang Wei, Haishan Liu, Changti Wu, Hang Yuan, Bailing Wang, Cong Huang, and 1 others. 2026. 3d-mix for vla: A plug-and-play module for integrating vggt- based 3d information into vision-language-action models.arXiv preprint arXiv:2603.24393

  46. [46]

    Bin Yu, Shijie Lian, Xiaopeng Lin, Yuliang Wei, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Xinming Wang, Bailing Wang, Cong Huang, and 1 others. 2026. Twinbrainvla: Unleashing the potential of generalist vlms for embodied tasks via asymmetric mixture-of-transformers.arXiv preprint arXiv:2601.14133

  47. [48]

    Xuanran Zhai, Qianyou Zhao, Qiaojun Yu, and Ce Hao. 2025. Vfp: Variational flow-matching policy for multi-modal robot manipulation.arXiv preprint arXiv:2508.01622

  48. [49]

    Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan

  49. [50]

    InInternational conference on machine learning (ICML), pages 61229–61245

    3D-VLA: A 3D vision-language-action generative world model. InInternational conference on machine learning (ICML), pages 61229–61245

  50. [51]

    Jinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu, Xirui Kang, Yuchun Feng, Yinan Zheng, Jiayin Zou, Yilun Chen, Jia Zeng, and 1 others. 2025. X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model.arXiv preprint arXiv:2510.10274

  51. [52]

    Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao, Hal Daumé III, Andrey Kolobov, Furong Huang, and Jianwei Yang. 2025. TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies.arXiv preprint arXiv:2412.10345

  52. [53]

    Linqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong, Maoqing Yao, Si Liu, and Guanghui Ren. 2026. Acot-vla: Action chain-of-thought for vision-language-action models.arXiv preprint arXiv:2601.11404

  53. [54]

    Zheyuan Zhou, Liang Du, Zixun Sun, Xiaoyu Zhou, Ruimin Ye, Qihao Chen, Yinda Chen, and Lemiao Qiu

  54. [55]

    Main-vla: Modeling abstraction of intention and environment for vision-language-action models.arXiv preprint arXiv:2602.02212

  55. [56]

    Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, and 1 others. 2023. Rt-2: Vision-language-action models transfer web knowledge to robotic control. InConference on Robot Learning (CoRL), pages 2165–2183. 15 A Additional Analysis on Intent Consistency and Mode Switching A.1 Mode Swi...