REVIEW 3 major objections 6 minor 1 cited by
Frame-only robot policies can flip short-horizon intents when observations look alike; conditioning action chunks on a compact history-derived intent keeps consecutive chunks consistent and raises success.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 18:58 UTC pith:T3KETLZK
load-bearing objection Useful failure-mode framing plus a real aliasing benchmark; the method is solid short-memory engineering, not a deep new theory of intent. the 3 major comments →
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Under short-horizon observation aliasing, conditioning chunked vision-language-action policies on a compact intent representation extracted from recent visual history stabilizes local continuations and outperforms frame-conditioned and raw-history baselines in success and inter-chunk consistency on AliasBench and on SimplerEnv, LIBERO, and RoboCasa.
What carries the argument
The short-horizon intent representation: a frozen geometry encoder turns recent head-camera frames into camera and register tokens; gated cross-attention fuses those tokens into the current vision-language context and an appended compact history-evidence token conditions a flow-matching action head for the next chunk.
Load-bearing premise
A fixed short window of recent head-camera images, encoded only as frozen geometry tokens, is enough at test time to recover the episode’s already-chosen next step even when the robot’s own mistakes make that history look different from the demonstrations.
What would settle it
On AliasBench ambiguity windows, train and evaluate the same short-history intent module against a matched frame-only baseline; if inter-chunk action disagreement (ICC-L2) does not fall and average success stays near the frame-only level, the claim that short-horizon intent conditioning resolves aliasing is falsified.
If this is right
- Frame-conditioned chunk VLAs systematically fail when the same visual state recurs with different local goals.
- Compact history-derived intent conditioning is more effective and memory-efficient than stuffing raw past frames into the language backbone.
- Inter-chunk consistency on overlapping actions is a practical diagnostic of whether a policy has committed to one short-horizon continuation.
- Controlled aliasing benchmarks can expose a failure mode that average success on saturated standard suites often hides.
- Short visual memory alone can improve multi-stage and partially observed manipulation without an explicit long-horizon planner.
Where Pith is reading between the lines
- Policies that look strong on near-saturated benchmarks may still be fragile in real kitchens or factories where phases and handoffs reappear under partial views.
- An ambiguity detector that lengthens or refreshes history only when the current frame is mixed could shrink the remaining closed-loop history-shift failures.
- The same compact intent-token idea may transfer to multi-robot or human-robot handoff settings where visual symmetry creates analogous short-horizon aliasing.
- Explicit intent labels appear unnecessary if the right compact history features are fused into the action conditioner.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that frame-conditioned chunked VLAs are unstable under short-horizon observation aliasing because demonstrations are multimodal across episodes but locally committed within an episode, while current-frame conditioning can resample different continuations across replans. It introduces IntentVLA, which freezes a VGGT history encoder, retains camera and register tokens from a short visual window, fuses them into a Qwen3-VL context via gated cross-attention, appends a pooled history-evidence token, and conditions a DiT flow-matching action head. It also introduces AliasBench, a 12-task RoboTwin2 benchmark with matched data isolating back-and-forth, crossing-path, bimanual, and multi-goal aliasing, plus an ICC-L2 inter-chunk consistency metric in annotated ambiguity windows. Empirically, IntentVLA raises AliasBench average success from 9.0% (Qwen3VL-GR00T) and 28.1% (best feasible raw-history baseline) to 45.8%, reduces mean ICC-L2 by 17.6%, and improves SimplerEnv (72.9%), LIBERO-Long (97.4%), and RoboCasa (57.0%), with component ablations on SimplerEnv.
Significance. If the result holds, the paper makes two useful contributions for chunked VLA control: (i) a cleanly motivated failure mode—uncommitted multimodality under aliased current observations—and (ii) a practical short-memory design that is more efficient than stuffing raw past frames into the VLM context. AliasBench is a genuine evaluation contribution: matched training/eval, family-level construction around latent factors (phase/source/handoff/target), and a quantitative nearest-neighbor aliasing diagnostic (Fig. 3) make the claim falsifiable rather than purely narrative. The ICC-L2 metric in annotated ambiguity windows is a concrete stability probe beyond success rate. Transfer gains on SimplerEnv, LIBERO-Long, and RoboCasa, plus ablations showing that current-frame VGGT alone does not help while history fusion and the compact evidence token do, strengthen the engineering case. Code is promised at a public GitHub URL. The work is simulation-only and leaves long-horizon memory and closed-loop history shift open, but the short-horizon framing is appropriately scoped.
major comments (3)
- §4.1–4.3 and §5.1 / Table 1: The central causal claim is that gains come from recovering the episode’s already-committed short-horizon continuation (phase/source/handoff/active target), not merely from adding short-horizon geometry or motion features. AliasBench is designed so that o_t is aliased while recent history carries the latent factor (Eqs. 1–2; Fig. 3), but the reported ablations do not break that link. Table 5 (SimplerEnv only) shows history fusion and the intent token help and current-frame VGGT does not; it does not test, on AliasBench families, histories that omit or scramble the disambiguating cue (e.g., same-length history from a different phase/source, temporally reversed history, or history without the origin/handoff frames). Without such controls, the jump from 9.0%/28.1% to 45.8% and the ICC-L2 drop can still be read as generic temporal enrichment. A family-level contr
- §5.1, Eq. (13) and Fig. 5: ICC-L2 is a good action-level proxy for inter-chunk commitment, but the manuscript does not fully specify the evaluation protocol needed to interpret it as intent consistency. Please state explicitly: (i) whether ICC is computed on all rollouts or only successful ones; (ii) the replan interval r and chunk horizon H used; (iii) how ambiguity windows are annotated and whether they are fixed from demonstration structure or detected online; and (iv) whether lower ICC could arise from more conservative/smoother actions rather than correct latent-factor recovery. Correlating ICC reduction with success within each AliasBench family (or reporting ICC conditional on correct vs incorrect continuation) would make the stability claim tighter.
- §5.1 Table 1 and Limitations §7: AliasBench average success remains 45.8%, with bimanual at 17.0% and multi-goal at 31.3%. The paper attributes residual failure partly to the fixed 16-frame window and closed-loop history shift. That is candid, but it is also the weakest assumption of the method (§4.2). The main claim would be stronger if the authors quantified sensitivity to K and to history corruption (e.g., replace recent frames with frames from a wrong-intent neighbor, or inject execution noise so test history diverges from demos) rather than only noting the gap. Even a small controlled study on one back-and-forth and one crossing-path task would show whether the compact intent token remains informative under the failure mode the paper itself highlights.
minor comments (6)
- §2.2: Several concurrent intent/memory VLA works are cited; a short table contrasting conditioning signal (current frame vs history), whether intent is supervised, and whether evaluation isolates aliasing would help readers place IntentVLA relative to DIAL, MINT, MemoryVLA, and Mem-0.
- Fig. 1 and Fig. 2: The qualitative aliasing examples are clear, but figure captions should state camera viewpoint(s) and whether the shown frames are from training demos or evaluation rollouts.
- §4.2: Clarify multi-view handling more precisely—history uses only head-camera frames while the current backbone may use the full current observation. State whether this asymmetry is fixed for all benchmarks and whether wrist/side cameras ever enter the history branch.
- Table 2: IntentVLA underperforms the frame-only baseline on Put Spoon on Towel (70.8 vs 83.0) while gaining elsewhere; the discussion in §5.3 is helpful—consider flagging this trade-off earlier when claiming broad robustness.
- Appendix A: The mode-switching analysis (Eq. 14) is useful conceptually; note more explicitly that P_switch is not estimated from the model and is only a diagnostic framing for ICC-L2.
- Presentation: The manuscript is marked “Work in progress” on the first page; for journal submission, remove that banner and ensure arXiv/version metadata are consistent. Also fix minor redundancy (VFP appears twice in the references).
Circularity Check
No significant circularity: standard imitation learning with held-out rollout metrics; mild same-group baseline citations are comparative, not load-bearing.
full rationale
IntentVLA’s chain is architectural and empirical, not a first-principles derivation that collapses into its inputs. The short-horizon intent representation mt = f_φ(ot, ℓ, hK_t) is learned without intent labels (Eqs. 6–12); training is ordinary conditional flow matching on demonstration chunks, while success and ICC-L2 are measured on held-out closed-loop rollouts and annotated ambiguity windows—neither is the training target restated as a prediction. AliasBench is constructed from task design (Eqs. 1–2, families in §3) and a separate embedding nearest-neighbor diagnostic (Fig. 3), not from fitting the policy’s own outputs. Same-group citations (StarVLA, LangForce, PhysBrain, TwinBrainVLA, 3D-Mix) appear as baselines or pipeline references in tables and training protocol; they do not supply a uniqueness theorem, ansatz, or forced result that the central claim reduces to. No fitted parameter is renamed a prediction, and no equation is definitionally equivalent to the reported gains. Score 1 only for non-load-bearing same-group baseline presence; the core claim remains independently evaluated.
Axiom & Free-Parameter Ledger
free parameters (5)
- history window length K (≈16 frames)
- VGGT token selection (1 camera + 4 register tokens per frame)
- learned gate scalar α and projection matrices Wh, We
- training schedule (30K steps, lr 1e-5 cosine, batch 256 on 16 H100)
- chunk horizon H and replan interval r used for ICC-L2
axioms (5)
- domain assumption Demonstrations are multimodal across episodes but locally committed within an episode; frame-only conditioning can break that commitment under partial observability.
- domain assumption A finite recent visual history h^K_t is sufficient evidence to concentrate the short-horizon intent posterior for chunk generation.
- ad hoc to paper Frozen VGGT camera and register tokens capture viewpoint change and inter-frame structure useful for active short-horizon intent.
- domain assumption Conditional flow-matching DiT action heads with Euler integration (as in GR00T-style VLAs) are an adequate generation interface once conditioning context is improved.
- domain assumption Simulation environments (RoboTwin2 AliasBench, SimplerEnv, LIBERO, RoboCasa) isolate and measure the intended aliasing failure mode.
invented entities (3)
-
short-horizon intent representation m_t / condition context C_t
no independent evidence
-
AliasBench (12-task ambiguity-aware benchmark)
no independent evidence
-
ICC-L2 inter-chunk consistency metric in ambiguity windows
no independent evidence
read the original abstract
Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines
Figures
Forward citations
Cited by 1 Pith paper
-
SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models
Selecting 50% of robot demonstrations by maximizing exposure to reusable primitive-transition patterns outperforms full-data training while halving training steps.
Reference graph
Works this paper leans on
-
[1]
Hongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan, Zhizhong Su, Hang Su, and Jun Zhu. 2025. H-RDT: Human manipulation enhanced bimanual robotic manipulation.arXiv preprint arXiv:2507.23523
Pith/arXiv arXiv 2025
-
[2]
Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, and 1 others. 2025. GR00T N1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734
Pith/arXiv arXiv 2025
-
[3]
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, and 5 others. 2024.π 0: A vision- language-action flow model for general robot cont...
Pith/arXiv arXiv 2024
-
[4]
Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xu Huang, Shu Jiang, and 1 others. 2025. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669
Pith/arXiv arXiv 2025
-
[5]
Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. 2025. Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111
Pith/arXiv arXiv 2025
-
[6]
Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, and 1 others. 2025. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088
Pith/arXiv arXiv 2025
-
[7]
Tianxing Chen, Yuran Wang, Mingleyang Li, Yan Qin, Hao Shi, Zixuan Li, Yifan Hu, Yingsheng Zhang, Kaixuan Wang, Yue Chen, and 1 others. 2026. Rmbench: Memory-dependent robotic manipulation benchmark with insights into policy design.arXiv preprint arXiv:2603.01229
arXiv 2026
-
[8]
Yi Chen, Yuying Ge, Hui Zhou, Mingyu Ding, Yixiao Ge, and Xihui Liu. 2026. Dial: Decoupling intent and action via latent world modeling for end-to-end vla.arXiv preprint arXiv:2603.29844
Pith/arXiv arXiv 2026
-
[9]
StarVLA Community. 2026. Starvla: A lego-like codebase for vision-language-action model developing.arXiv preprint arXiv:2604.05014
Pith/arXiv arXiv 2026
-
[10]
GEAR-Team, Allison Azzolini, Johan Bjorck, Valts Blukis, Fernando Castañeda, Rahul Chand, and 1 others
-
[11]
nvidia.com/labs/gear/gr00t-n1_6/
Gr00t n1.6: An improved open foundation model for generalist humanoid robots.https://research. nvidia.com/labs/gear/gr00t-n1_6/
-
[12]
Renming Huang, Chendong Zeng, Wenjing Tang, Jintian Cai, Cewu Lu, and Panpan Cai. 2026. Mimic intent, not just trajectories.arXiv preprint arXiv:2602.08602
arXiv 2026
-
[13]
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y . Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, and 17 others. 2025.π0.5: a vision-language-action model with open-world...
Pith/arXiv arXiv 2025
-
[14]
Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, and 1 others. 2024. Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945
Pith/arXiv arXiv 2024
-
[15]
Moo Jin Kim, Chelsea Finn, and Percy Liang. 2025. Fine-tuning vision-language-action models: Optimizing speed and success.arXiv preprint arXiv:2502.19645
Pith/arXiv arXiv 2025
-
[16]
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P Foster, Pannag R Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. 2024. OpenVLA: An open-source vision-language- action model. InConference on Robot Learning (CoRL)
2024
-
[17]
Chengmeng Li, Junjie Wen, Yaxin Peng, Yan Peng, and Yichen Zhu. 2026. Pointvla: Injecting the 3d world into vision-language-action models.IEEE Robotics and Automation Letters, 11(3):2506–2513
2026
-
[18]
Fuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang, Pengxiang Ding, Donglin Wang, Long Zeng, and Haoang Li. 2025. Spatial forcing: Implicit spatial representation alignment for vision-language-action model.arXiv preprint arXiv:2510.12276. 13
arXiv 2025
-
[19]
Peiyan Li, Yixiang Chen, Hongtao Wu, Xiao Ma, Xiangnan Wu, Yan Huang, Liang Wang, Tao Kong, and Tie- niu Tan. 2025. Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models. InAdvances in neural information processing systems (NeurIPS)
2025
-
[20]
Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizhong Zhang, and 1 others. 2024. CogACT: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation.arXiv preprint arXiv:2411.19650
Pith/arXiv arXiv 2024
-
[21]
Xinghang Li, Peiyan Li, Minghuan Liu, Dong Wang, Jirong Liu, Bingyi Kang, Xiao Ma, Tao Kong, Hanbo Zhang, and Huaping Liu. 2024. Towards generalist robot policies: What matters in building vision-language- action models.arXiv preprint arXiv:2412.14058
Pith/arXiv arXiv 2024
-
[22]
Xuanlin Li, Kyle Hsu, Jiayuan Gu, Oier Mees, Karl Pertsch, Homer Rich Walke, Chuyuan Fu, Ishikaa Lu- nawat, Isabel Sieh, Sean Kirmani, Sergey Levine, Jiajun Wu, Chelsea Finn, Hao Su, Quan Vuong, and Ted Xiao. 2024. SimplerEnv: Evaluating real-world robot manipulation policies in simulation. InConference on Robot Learning (CoRL)
2024
-
[23]
Shijie Lian, Bin Yu, Xiaopeng Lin, Laurence T Yang, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Cong Huang, and Kai Chen. 2026. Langforce: Bayesian decomposition of vision language action models via latent action queries.arXiv e-prints, pages arXiv–2601
2026
-
[24]
Tao Lin, Gen Li, Yilei Zhong, Yanwen Zou, Yuxin Du, Jiting Liu, Encheng Gu, and Bo Zhao. 2025. Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416
arXiv 2025
-
[25]
Xiaopeng Lin, Shijie Lian, Bin Yu, Ruoqi Yang, Changti Wu, Yuzhuo Miao, Yurun Jin, Yukun Shi, Cong Huang, Bojun Cheng, and 1 others. 2025. PhysBrain: Human egocentric data as a bridge from vision language models to physical intelligence.arXiv preprint arXiv:2512.16793
arXiv 2025
-
[26]
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. 2023. LIBERO: Benchmarking knowledge transfer for lifelong robot learning.Advances in neural information processing sys- tems (NeurIPS), 36:44776–44791
2023
-
[27]
Songming Liu, Bangguo Li, Kai Ma, Lingxuan Wu, Hengkai Tan, Xiao Ouyang, Hang Su, and Jun Zhu
-
[28]
Rdt2: Exploring the scaling limit of umi data towards zero-shot cross-embodiment generalization.arXiv preprint arXiv:2602.03310
-
[29]
Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. 2025. RDT-1b: a diffusion foundation model for bimanual manipulation. InInternational Conference on Learning Representations (ICLR)
2025
-
[30]
Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. InInternational Conference on Learning Representations (ICLR)
2017
-
[31]
Yulin Luo, Hao Chen, Zhuangzhe Wu, Bowen Sui, Jiaming Liu, Chenyang Gu, Zhuoyang Liu, Qiuxuan Feng, Jiale Yu, Shuo Gu, and 1 others. 2026. Look before acting: Enhancing vision foundation representations for vision-language-action models.arXiv preprint arXiv:2603.15618
arXiv 2026
-
[32]
Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Man- dlekar, and Yuke Zhu. 2024. RoboCasa: Large-scale simulation of everyday tasks for generalist robots. In Robotics: Science and Systems
2024
-
[33]
Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, and 1 others. 2024. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE
2024
-
[34]
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. 2025. FAST: Efficient action tokenization for vision-language-action models.arXiv preprint arXiv:2501.09747
Pith/arXiv arXiv 2025
-
[35]
Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Yan Ding, Zhigang Wang, JiaYuan Gu, Bin Zhao, Dong Wang, and 1 others. 2025. Spatialvla: Exploring spatial representations for visual-language-action model.arXiv preprint arXiv:2501.15830
Pith/arXiv arXiv 2025
-
[36]
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimiza- tions enable training deep learning models with over 100 billion parameters. InProceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, pages 3505–3506. 14
2020
-
[37]
Yichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang, Yan Lu, Jiaolong Yang, Nanning Zheng, and Baining Guo. 2025. VideoVLA: Video generators can be generalizable robot manipulators. InAdvances in neural information processing systems (NeurIPS)
2025
-
[38]
Hao Shi, Bin Xie, Yingfei Liu, Lin Sun, Fengrong Liu, Tiancai Wang, Erjin Zhou, Haoqiang Fan, Xiangyu Zhang, and Gao Huang. 2026. Memoryvla: Perceptual-cognitive memory in vision-language-action models for robotic manipulation. InInternational Conference on Learning Representations (ICLR)
2026
-
[39]
Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren, Zezhi Liu, Hanxin Zhu, Guangzhong Sun, Xin Jin, and Zhibo Chen. 2026. Vla-jepa: Enhancing vision-language-action model with latent world model.arXiv preprint arXiv:2602.10098
arXiv 2026
-
[40]
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, and 1 others. 2024. Octo: An open-source generalist robot policy.arXiv preprint arXiv:2405.12213
Pith/arXiv arXiv 2024
-
[41]
Homer Rich Walke, Kevin Black, Tony Z Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen-Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, and 1 others. 2023. Bridgedata v2: A dataset for robot learning at scale. InConference on Robot Learning (CoRL), pages 1723–1736. PMLR
2023
-
[42]
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny
-
[43]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Vggt: Visual geometry grounded transformer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
-
[44]
Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, and 1 others. 2025. Magma: A foundation model for multimodal ai agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14203– 14214
2025
-
[45]
Bin Yu, Shijie Lian, Xiaopeng Lin, Zhaolong Shen, Yuliang Wei, Haishan Liu, Changti Wu, Hang Yuan, Bailing Wang, Cong Huang, and 1 others. 2026. 3d-mix for vla: A plug-and-play module for integrating vggt- based 3d information into vision-language-action models.arXiv preprint arXiv:2603.24393
arXiv 2026
-
[46]
Bin Yu, Shijie Lian, Xiaopeng Lin, Yuliang Wei, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Xinming Wang, Bailing Wang, Cong Huang, and 1 others. 2026. Twinbrainvla: Unleashing the potential of generalist vlms for embodied tasks via asymmetric mixture-of-transformers.arXiv preprint arXiv:2601.14133
arXiv 2026
-
[48]
Xuanran Zhai, Qianyou Zhao, Qiaojun Yu, and Ce Hao. 2025. Vfp: Variational flow-matching policy for multi-modal robot manipulation.arXiv preprint arXiv:2508.01622
arXiv 2025
-
[49]
Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan
-
[50]
InInternational conference on machine learning (ICML), pages 61229–61245
3D-VLA: A 3D vision-language-action generative world model. InInternational conference on machine learning (ICML), pages 61229–61245
-
[51]
Jinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu, Xirui Kang, Yuchun Feng, Yinan Zheng, Jiayin Zou, Yilun Chen, Jia Zeng, and 1 others. 2025. X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model.arXiv preprint arXiv:2510.10274
Pith/arXiv arXiv 2025
-
[52]
Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao, Hal Daumé III, Andrey Kolobov, Furong Huang, and Jianwei Yang. 2025. TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies.arXiv preprint arXiv:2412.10345
Pith/arXiv arXiv 2025
-
[53]
Linqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong, Maoqing Yao, Si Liu, and Guanghui Ren. 2026. Acot-vla: Action chain-of-thought for vision-language-action models.arXiv preprint arXiv:2601.11404
arXiv 2026
-
[54]
Zheyuan Zhou, Liang Du, Zixun Sun, Xiaoyu Zhou, Ruimin Ye, Qihao Chen, Yinda Chen, and Lemiao Qiu
-
[55]
Main-vla: Modeling abstraction of intention and environment for vision-language-action models.arXiv preprint arXiv:2602.02212
-
[56]
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, and 1 others. 2023. Rt-2: Vision-language-action models transfer web knowledge to robotic control. InConference on Robot Learning (CoRL), pages 2165–2183. 15 A Additional Analysis on Intent Consistency and Mode Switching A.1 Mode Swi...
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.