REVIEW 3 major objections 5 minor 55 references
Generalizable robot control is possible with little real-robot data if the model jointly learns 3D world change, visual plans, and actions under mutual constraints.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-07-11 22:50 UTC pith:VYBWY75P
load-bearing objection Solid systems paper with a clean three-expert MoT and strong numbers on little real data; the causal “world-action prior” story is ahead of what the ablations identify. the 3 major comments →
WSA₁: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A robot foundation model that unifies action-conditioned 3D world prediction, 3D-grounded 2D visual thinking, and 3D inverse dynamics inside one latent space can learn generalizable manipulation from far less real-robot data than prevailing vision-language-action or world-action models. Instantiated as WSA1 (3B and 6B), this design yields roughly 93 percent success on RoboTwin2.0 hard and an average gain of about twenty percentage points over strong baselines on seven real tabletop tasks, using a pre-training mix of only six thousand hours of heterogeneous demonstrations.
What carries the argument
3D-centric World-Spatial-Action (WSA) joint modeling: three experts (2D spatial, 3D spatial, 3D action) share a latent space under bidirectional causal attention so that predicted 3D world tokens and action tokens constrain each other, while visual subgoal tokens attend to 3D geometry; trained with MSE losses on visual and 3D latents plus flow-matching on actions.
Load-bearing premise
The claim stands only if the bidirectional attention rules and losses against frozen depth and image tokenizers truly install transferable causal world–action priors rather than just multi-task fitting that helps mainly on the authors’ chosen simulation and tabletop suites.
What would settle it
Train an otherwise identical model that keeps the three prediction heads but replaces bidirectional world–action attention with unidirectional or no cross-attention, using the same six-thousand-hour mix; if real-robot and hard-sim success collapse to the level of ordinary action-only or 2D world-action baselines, the mutual-constraint inductive bias is not doing the work claimed.
If this is right
- Large real-robot teleoperation corpora are no longer a prerequisite for competitive generalist manipulation if 3D world–action co-modeling is used.
- Simulation and egocentric human video become first-class pre-training sources rather than mere supplements, because the 3D joint objective transfers across embodiments.
- Policy architectures can move from reactive 2D understand-then-execute or imagine-then-execute pipelines to closed-loop 3D predict-and-constrain control.
- Scaling robot foundation models can emphasize diverse multi-source data and 3D mutual constraints instead of ever-larger pure real-robot hours.
Where Pith is reading between the lines
- If the mutual-constraint idea is right, the same recipe should transfer to mobile manipulation and multi-room household tasks where 3D layout change is even more critical than on tabletops.
- Freezing off-the-shelf depth and image tokenizers may cap how far the world model can improve; end-to-end learned 3D geometry could be the next leverage point.
- The data-efficiency claim suggests a practical route for labs that cannot afford massive real-robot fleets but can generate simulation and collect modest teleop plus human video.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes WSA1, a robot foundation model based on a 3D-centric World-Spatial-Action (WSA) paradigm that jointly optimizes three objectives in a shared Mixture-of-Transformers latent space: 3D-aware 2D visual thinking (L2D), action-conditioned 3D world prediction (L3D against frozen Depth-Anything targets), and 3D inverse dynamics via flow matching (LACT). Bidirectional attention between 3D scene tokens and action tokens is intended to enforce mutual world–action constraints (Eqs. 5–9, §3.2–3.3). Pre-trained on a 6k-hour heterogeneous mix (only ~1k hours real robot; Table 1), WSA1-B/L report 92.7–93.1% success on RoboTwin2.0 hard, competitive LIBERO averages (~97–98%), and large gains on seven real tabletop tasks (+20% average SR over π0.5/InternVLA-A1; Table 3). Ablations (Figs. 4–5) attribute gains to the joint objectives and pre-training.
Significance. If the data-efficiency claim holds under matched controls, the work would be a practically important contribution: it argues that transferable manipulation priors can be obtained without tens of thousands of hours of real teleoperation by pairing multi-source data with explicit 3D world–action co-modeling. Strengths include multi-benchmark evaluation (RoboTwin2.0 hard, LIBERO, two real embodiments), open-source model scales (3B/6B), a clear MoT architecture with three experts, and ablations that separate visual thinking, 3D prediction, and pre-training. The empirical margins on RoboTwin2.0 and real tasks are large enough to matter for the field if they survive tighter controls.
major comments (3)
- [§3.2–3.3, Eqs. 5–9; Figs. 4–5] Central causal claim is not identified. §3.2 and Eqs. 5–9 assert that bidirectional attention between hg and hact induces mutual world–action constraints rather than multi-task correlation. Figs. 4–5 ablate presence of L2D/L3D and pre-training, but never ablate bidirectional hg–hact attention against unidirectional (predict-then-act or act-then-predict) or action-only attention under matched compute, identical frozen Depth-Anything/VAE targets, and the same data mix. Without that control, the data-efficiency narrative (abstract, §1, §5) remains consistent with ordinary multi-task imitation on a favorable 6k-hour recipe.
- [Table 3; §4.2; abstract] Real-world comparison (Table 3) is under-controlled for the +20% claim. Each of 7 tasks uses ~30 rollouts with no error bars, confidence intervals, or multi-seed variance. Baselines (π0, π0.5, InternVLA-A1) are not shown to have been pre-trained on the same Table-1 mixture or the same post-training budget/embodiment data. The abstract’s “+20% average boosted performance over SOTA RFMs” therefore cannot be attributed cleanly to the WSA inductive bias versus data recipe or fine-tuning protocol.
- [Eq. 7; §3.3; Fig. 6] Frozen 3D encoder as world model is a load-bearing assumption left untested. L3D (Eq. 7) regresses to Depth-Anything latents; if those targets are a weak or non-causal world model for contact-rich dynamics, the claimed “action-caused 3D world prediction” reduces to multi-view depth fitting. A minimal stress test—e.g., replacing Depth-Anything with a weaker depth prior or random 3D tokens, or measuring whether predicted Gt improves action success when actions are held fixed—would be needed to support the causal interpretation.
minor comments (5)
- [Figure 2] Figure 2 panel labels are inconsistent (two panels labeled “(c)”).
- [§3.1; §4.1] Notation for subgoal counts N vs K and action horizon H is introduced in §3.1 but not tabulated with default values used in experiments.
- [Table 1; §3.4] Table 1 sampling weights sum to 1.00 but the text does not state whether they were tuned or fixed a priori; a short sensitivity note would help reproducibility.
- [Table 6; §4.3] LIBERO results (Table 6) are strong but the fine-tuning protocol (epochs, data volume per suite) is thinner than for RoboTwin2.0; a one-sentence protocol match to prior work would clarify fairness.
- [§2; Table 4; Figure 2] Minor typos: “W AMs” spacing, “MotuBrain” vs “Motus” naming consistency in Table 4, and “5.0π” artifact in Figure 2.
Circularity Check
No circularity: empirical multi-task imitation with external rollout metrics; claimed gains do not reduce to training inputs by construction.
full rationale
WSA1 is an architectural/multi-task learning paper, not a first-principles derivation. The three objectives (Eqs. 6–9: MSE to frozen VAE tokens, MSE to Depth-Anything 3D tokens, flow-matching on demonstration actions) are standard supervised targets; success rates on RoboTwin2.0, LIBERO, and real-robot rollouts are external task-completion metrics, not restatements of those losses. Bidirectional attention (Fig. 3 dependency rules) is an inductive bias, not a definition that forces the reported SR. No uniqueness theorem, no fitted scalar renamed as a prediction, and no load-bearing self-citation chain: backbone priors (Qwen3-VL/Wan2.2), Depth-Anything, and flow matching are external. Self-citations (e.g., MiVLA, surveys) are peripheral. Choosing Depth-Anything as the 3D tokenizer is a methodological choice, not circularity of the claimed result. Score 0 is appropriate; any concern about whether mutual constraints truly induce causal priors is a correctness/identification issue, not circular reduction.
Axiom & Free-Parameter Ledger
free parameters (4)
- Data-source sampling weights =
0.47 / 0.07 / 0.17 / 0.19 / 0.10
- Action chunk horizon H and subgoal counts N, K
- Flow-matching noise schedule τ and loss coefficients =
equal sum of three losses
- Model scale / backbone choice (Qwen3-VL-2B vs Wan2.2-5B) =
3B and 6B variants
axioms (5)
- domain assumption Imitation learning on expert demonstrations maximizes the correct policy likelihood for generalist control (Eq. 1).
- domain assumption Frozen Depth-Anything (and VAE) latents are adequate proxies for true 3D world state transitions (Eq. 7, §3.3).
- ad hoc to paper Bidirectional attention between 3D scene tokens and action tokens induces causal world-action mutual constraints rather than mere multi-task correlation (§3.2).
- domain assumption Heterogeneous sim + human + limited real data form a sufficient multi-source pyramid for transferable physical priors (§3.4).
- standard math Flow matching on continuous action chunks is a valid inverse-dynamics objective conditioned on predicted 3D latents (Eq. 8).
invented entities (2)
-
3D-Centric World-Spatial-Action (WSA) joint modeling paradigm
no independent evidence
-
Mixture-of-Transformers with 2D Spatial / 3D Spatial / 3D Action experts under 3D-centric causal attention
no independent evidence
Cite this review
Pith. "Pith review of WSA$_1$: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control." pith.science (2026). https://pith.science/paper/VYBWY75P
@misc{pith2026260703941,
author = {Pith},
title = {Pith review of: WSA$_1$: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control},
year = {2026},
howpublished = {\url{https://pith.science/paper/VYBWY75P}},
note = {Machine review of arXiv:2607.03941}
}
read the original abstract
Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for generalist robotic systems to date. By leveraging imitation learning on extensive robot demonstrations, RFMs have achieved impressive capabilities in mapping visual observations and language instructions to continuous robotic actions. However, current RFMs lack an inherent ability to reason about physical dynamics and the causal effects of robot behaviors on the 3D physical world. This creates a fundamental mismatch between 2D-centric visual perception and 3D-centric embodied interaction, severely limiting the generalization ability of RFMs in real-world tasks.To address this gap, we present WSA$_1$, a novel RFM built upon proposed 3D-Centric World-Spatial-Action modeling paradigm. It not only learns 3D world-aware visual thought for future robot behaviors, but also models mutual constraints between 3D world state transitions and robotic actions to enhance behavior generalization. Notably, WSA$_1$ achieves highly data-efficient pre-training with 6k hours of expert demonstration data (only 1k hours from real robot), while delivering competitive manipulation performance (93% success rate) on RoboTwin2.0 simulation benchmark and achieving +20% average boosted performance over state-of-the-art RFMs on real-world robot control tasks. These results reveal that generalizable RFM can be attained without large-scale real robot data when paired with 3D-centric world-action joint modeling, which offers a practical and affordable pathway to generalist robotic systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, Yunhao Ge, Jinwei Gu, Siddharth Gururani, Ethan He, Jiahui Huang, Jacob Huffman, Pooya Jannaty, Jingyi Jin, S...
Pith/arXiv arXiv 2025
-
[2]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al . 2025. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631(2025)
Pith/arXiv arXiv 2025
-
[3]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...
Pith/arXiv arXiv 2025
-
[4]
Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, Hongyan Zhao, Hanyu Liu, Zhizhong Su, Lei Ma, Hang Su, and Jun Zhu. 2025. Motus: A Unified Latent Action World Model.arXiv preprint arXiv:2512.13030(2025)
Pith/arXiv arXiv 2025
-
[5]
Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi "Jim" Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, Kevin Lin, Guilin Liu, Edith Llontop, Loic Magne, Ajay Mandlekar, Avnish Narayan, Soroush Nasiriany, Scott Reed, You Liang Ta...
Pith/arXiv arXiv 2026
-
[6]
Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Robert Equi, Chelsea Finn, Niccolo Fusai, Manuel Y . Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, brian ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren,...
-
[7]
InCoRL, V ol
π0.5: a Vision-Language-Action Model with Open-World Generalization. InCoRL, V ol. 305. 17–40
-
[8]
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xi- aoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky
-
[9]
π0: A Vision-Language-Action Flow Model for General Robot Control.arXiv preprint arXiv:2410.24164(2024)
Pith/arXiv arXiv 2024
-
[10]
Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xuan Hu, Xu Huang, et al . 2025. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669 (2025)
Pith/arXiv arXiv 2025
-
[11]
Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, et al. 2026. InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation.arXiv preprint arXiv:2601.02456(2026)
arXiv 2026
-
[12]
Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, Weiliang Deng, Yubin Guo, Tian Nian, Xuanbing Xie, Qiangyu Chen, Kailun Su, Tianling Xu, Guodong Liu, Mengkang Hu, Huan ang Gao, Kaixuan Wang, Zhixuan Liang, Yusen Qin, Xiaokang Yang, Ping Luo, and Yao Mu. 2025. RoboTwin 2.0: A Scal...
Pith/arXiv arXiv 2025
-
[13]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InCVPR. 24185–24198
2024
-
[14]
Andy Clark. 2013. Whatever next? Predictive brains, situated agents, and the future of cognitive science.Behavioral and brain sciences36, 3 (2013), 181–204
2013
-
[15]
2011.Frames of Mind: The Theory of Multiple Intelligences
Howard Gardner. 2011.Frames of Mind: The Theory of Multiple Intelligences. Basic Books, New York
2011
-
[16]
Yoon, Mouli Sivapurapu, and Jian Zhang
Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu, and Jian Zhang. 2025. EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video.arXiv preprint arXiv:2505.11709(2025)
Pith/arXiv arXiv 2025
-
[17]
Wenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu, Dieter Fox, Kaichun Mo, and Li Fei-Fei. 2026. PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation. InCVPR
2026
-
[18]
Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, and Jinwei Gu. 2026. Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.arXiv preprint arXiv:2601.16163(2026)
Pith/arXiv arXiv 2026
-
[19]
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P Foster, Pannag R Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. 2025. OpenVLA: An Open-Source Vision-Language-Action Model. InCoRL, V ol. 270. 2679–2713
2025
-
[20]
1996.Image and brain: The resolution of the imagery debate
Stephen M Kosslyn. 1996.Image and brain: The resolution of the imagery debate. MIT press
1996
-
[21]
Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, Yujun Shen, and Yinghao Xu. 2026. Causal World Modeling for Robot Control.arXiv preprint arXiv:2601.21998(2026)
Pith/arXiv arXiv 2026
-
[22]
Peiyan Li, Yixiang Chen, Yuan Xu, Jiabing Yang, Xiangnan Wu, Jun Guo, Nan Sun, Long Qian, Xinghang Li, Xin Xiao, Jing Liu, Nianfeng Liu, Tao Kong, Yan Huang, Liang Wang, and Tieniu Tan. 2026. Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model.arXiv preprint arXiv:2604.03181(2026). 16
Pith/arXiv arXiv 2026
-
[23]
Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. 2025. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647(2025)
Pith/arXiv arXiv 2025
-
[24]
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022. Flow matching for generative modeling. InICLR
2022
-
[25]
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. 2023. LIBERO: benchmarking knowledge transfer for lifelong robot learning. InNeurIPS
2023
-
[26]
Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. 2025. RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation. InICLR, V ol. 2025. 29982–30009
2025
-
[27]
Yunfan Lou, Xiaowei Chi, Xiaojie Zhang, Zezhong Qian, Chengxuan Li, Rongyu Zhang, Yaoxu Lyu, Guoyu Song, Chuyao Fu, Haoxuan Xu, Pengwei Wang, and Shanghang Zhang. 2026. Mask World Model: Predicting What Matters for Robust Robot Policy Learning.arXiv preprint arXiv:2604.19683(2026)
Pith/arXiv arXiv 2026
-
[28]
Hao Luo, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng, Haiweng Xu, Chaoyi Xu, Ziheng Xi, Yuhui Fu, and Zongqing Lu. 2026. Being-H0.7: A Latent World-Action Model from Egocentric Videos.arXiv preprint arXiv:2605.00078(2026)
Pith/arXiv arXiv 2026
-
[29]
Jonas Pai, Liam Achenbach, Victoriano Montesinos, Benedek Forrai, Oier Mees, and Elvis Nava. 2025. mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs.arXiv preprint 2512.15692(2025)
Pith/arXiv arXiv 2025
-
[30]
Jingjing Qian, Boyao Han, Chen Shi, Lei Xiao, Long Yang, Shaoshuai Shi, and Li Jiang. 2025. GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation. InCVPR
2025
-
[31]
Delin Qu, Zeren Gu, Bin Zhao, Dong Wang, and Xuelong Li. 2025. SpatialVLvla:3DVLAA: Exploring Spatial Representations for Visual-Language-Action Models. InRobotics: Science and Systems (RSS)
2025
-
[32]
MotuBrain Team, Chendong Xiang, Fan Bao, Haitian Liu, Hengkai Tan, Hongzhe Bi, James Li, Jiabao Liu, Jingrui Pang, Kiro Jing, Louis Liu, Mengchen Cai, Rongxu Cui, Ruowen Zhao, Runqing Wang, Shuhe Huang, Yao Feng, Yinze Rong, Zeyuan Wang, and Jun Zhu
-
[33]
MotuBrain: An Advanced World Action Model for Robot Control.arXiv preprint arXiv:2604.27792(2026)
Pith/arXiv arXiv 2026
-
[34]
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yun- liang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. 2024. Octo: An Open-Source Generalist Robot Policy.arXiv preprint arXiv:2405.12213 (2024)
Pith/arXiv arXiv 2024
-
[35]
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fang, T...
Pith/arXiv arXiv 2025
-
[36]
Yang Tian, Yuyin Yang, Yiman Xie, Zetao Cai, Xu Shi, Ning Gao, Hangxu Liu, Xuekun Jiang, Zherui Qiu, Feng Yuan, et al. 2025. Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy.arXiv preprint arXiv:2511.16651(2025)
arXiv 2025
-
[37]
Qiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye, Sicheng Xie, Yitao Liu, Junhao Chen, Zhixuan Liang, Jie Zhang, Xintong Hu, Xuhong Huang, Pei Lin, Junyang Lin, Dayiheng Liu, Shuai Bai, Jingren Zhou, Jiazhao Zhang, Haoqi Yuan, Gengze Zhou, Hang Yin, Ye Wang, Yiyang Huang, Zixing Lei, Wujian Peng, Delin Chen, Yingming Zheng, Jingyang Fan, Xianwei Zhuang, Xi...
Pith/arXiv arXiv 2026
-
[38]
Xuanhan Wang, Huimin Deng, Lianli Gao, and Jingkuan Song. 2025. Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models.arXiv preprint arXiv:2503.08201(2025)
Pith/arXiv arXiv 2025
-
[39]
Junjie Wen, Yichen Zhu, Jinming Li, Zhibin Tang, Chaomin Shen, and Feifei Feng. 2025. DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control. In CoRL
2025
-
[40]
Daniel M Wolpert and Zoubin Ghahramani. 2000. Computational principles of movement neuroscience.Nature neuroscience3, 11 (2000), 1212–1217
2000
-
[41]
Adina Yakefu, Bin Xie, Chongyang Xu, Enwen Zhang, Erjin Zhou, Fan Jia, Haitao Yang, Haoqiang Fan, Haowei Zhang, Hongyang Peng, et al . 2025. RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies.arXiv preprint arXiv:2510.17950(2025)
arXiv 2025
-
[42]
Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, Hongxu Yin, Sifei Liu, Song Han, Yao Lu, and Xiaolong Wang. 2025. EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos.arXiv preprint arXiv:2507.12440(2025)
Pith/arXiv arXiv 2025
-
[43]
Yandan Yang, Shuang Zeng, Tong Lin, Xinyuan Chang, Dekang Qi, Junjin Xiao, Haoyun Liu, Ronghan Chen, Yuzhi Chen, Dongjie Huo, Feng Xiong, Xing Wei, Zhiheng Ma, and Mu Xu
-
[44]
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.arXiv preprint arXiv:2602.11236(2026)
Pith/arXiv arXiv 2026
-
[45]
Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, Ayaan Malik, Kyungmin Lee, William Liang, Nadun Ranawaka, Jiasheng Gu, Yinzhen Xu, Guanzhi Wang, Fengyuan Hu, Avnish Narayan, Johan Bjorck, Jing Wang, Gwanghyun Kim, Dantong Niu, Ruijie Zheng, Yuqi Xie, Jimmy Wu, Qi ...
Pith/arXiv arXiv 2026
-
[46]
Zhenhan Yin, Xuanhan Wang, Jiahao Jiang, Kaiyuan Deng, Pengqi Chen, Shuangle Li, Chong Liu, Xing Xu, Jingkuan Song, Lianli Gao, and Heng Tao Shen. 2026. MiVLA: Towards Gener- alizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Find...
2026
-
[47]
Zhaoshu Yu, Bo Wang, Pengpeng Zeng, Haonan Zhang, Ji Zhang, Zheng Wang, Lianli Gao, Jingkuan Song, Nicu Sebe, and Heng Tao Shen. 2025. A survey on efficient vision-language- action models.arXiv preprint arXiv:2510.24795(2025)
arXiv 2025
-
[48]
Tianyuan Yuan, Zibin Dong, Yicheng Liu, and Hang Zhao. 2026. Fast-W AM: Do World Action Models Need Test-time Future Imagination?arXiv preprint arXiv:2603.16666(2026)
Pith/arXiv arXiv 2026
-
[49]
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sigmoid Loss for Language Image Pre-Training. InICCV. 11975–11986
2023
-
[50]
Peng-Fei Zhang, Ying Cheng, Xiaofan Sun, Shijie Wang, Fengling Li, Lei Zhu, and Heng Tao Shen. 2025. A step toward world models: A survey on robotic manipulation.arXiv preprint arXiv:2511.02097(2025)
arXiv 2025
-
[51]
Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang, Xinqiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He, Fan Lu, He Wang, Zhizheng Zhang, Li Yi, Wenjun Zeng, and Xin Jin
-
[52]
InNeurIPS
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge. InNeurIPS. 18
-
[53]
Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, Ankur Handa, Tsung-Yi Lin, Gordon Wetzstein, Ming-Yu Liu, and Donglai Xiang. 2025. CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models. InCVPR. 1702–1713
2025
-
[54]
Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan. 2024. 3D-VLA: A 3D Vision-Language-Action Generative World Model. 61229–61245
2024
-
[55]
Sanketi, Grecia Salazar, Michael S
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, Quan Vuong, Vincent Vanhoucke, Huong Tran, Radu Soricut, Anikait Singh, Jaspiar Singh, Pierre Sermanet, Pannag R. Sanketi, Grecia Salazar, Michael S. Ryoo, Krista Reymann, Kanishka Rao, Karl Pertsch, Igor Mordatch, Henryk Michalewski...
2023
This paper was first reviewed by grok-4.5 on July 11, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.