REVIEW 4 major objections 5 minor 11 cited by
Training Strategies for Efficient Embodied Reasoning
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Robot reasoning pays off in training, not at slow inference.
desk verdict Useful recipes and solid LIBERO ablations; the mechanism claim is under-tested and the Bridge dropout numbers come from a model not trained with the dropout recipe. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a pair of training-time interventions that decouple learning to reason from reasoning at inference: reasoning pre-training, which supervises the VLM on embodied reasoning traces alone and then fine-tunes on actions, and reasoning dropout, which trains the VLA to predict actions while randomly masking or keeping the reasoning tokens in context. Combined with the embodied reasoning annotations themselves (plans, subtask rationales, bounding boxes, gripper positions, and movement rationales), these recipes force the action head to rely on reasoning-shaped representations during training while removing the need to decode reasonings during deployment.
What would settle it
Train the same architecture with the reasoning-dropout protocol but replace the dropped reasoning steps with an equal number of random text tokens at matched sequence length. If the perturbed-split success rates stay near 70-76%, the semantic content of reasoning is not load-bearing and the effect is generic regularization; if performance drops back toward the standard-VLA baseline, the reasoning content itself is doing the work.
Extended reading notes
Core claim
The paper's central claim is that learning to generate reasonings leads to better VLA representations, while attending to the reasonings during training is what lets the policy leverage those features for improved action prediction. Experimentally, reasoning dropout and reasoning pre-training outperform a standard VLA on LIBERO-90 (89.4% and 87.1% versus 82.0% on the standard split) and on real-robot Bridge tasks (69.4% versus 50.5% for pre-training), matching or approaching full embodied chain-of-thought while running at the faster speed of a plain VLA. The paper finds only weak support for the learning-curriculum hypothesis and no support for the expressivity hypothesis, since adding non-semantic thinking tokens actually hurts performance.
Load-bearing premise
The conclusion that reasoning improves policies through semantic representation learning rests on the untested assumption that the gains of reasoning dropout come from the meaning of the reasoning tokens rather than from the generic regularization effect of randomly dropping tokens during training.
Editorial extensions
If this is right
- Policies trained with reasoning pre-training or dropout can be deployed at the control rate of a standard VLA, not the 1-1.2 Hz of full embodied CoT, while matching or beating it on LIBERO-90.
- Reasoning dropout reaches state-of-the-art LIBERO-90 performance (89.4%) without generating any test-time reasoning, indicating that in narrow benchmarks the reasoning traces can be fully internalized.
- Reasoning pre-training lifts real-robot Bridge success from 50.5% to 69.4%, and because it does not need paired reasoning-action data, it can in principle learn from unpaired reasoning data from other embodiments.
- The paper's practical prescription follows from its ablations: use full ECoT to maximize performance, reasoning dropout for narrow task domains, and reasoning pre-training for diverse domains or when unpaired reasoning data are available.
Reading between the lines
- If the mechanism is semantic representation learning, then scaling the diversity of reasoning data should widen performance gaps between the recipes and may make test-time reasoning increasingly valuable, a trend the paper's own LIBERO-versus-Bridge comparison already hints at.
- A direct test of the mechanism would compare reasoning dropout against dropping matched-length random text tokens; if the gains are equivalent, the benefit is generic regularization rather than reasoning content.
- The null result for thinking tokens suggests that extra inference compute without a meaningful training signal will not transfer to manipulation, and a further control could insert semantically meaningful but task-irrelevant tokens to isolate content from mere sequence length.
- Reasoning pre-training may generalize as a recipe for adapting any VLM before action fine-tuning, effectively extending domain-adaptive pretraining to robot control.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates why embodied chain-of-thought (ECoT) reasoning improves vision-language-action (VLA) policies and then proposes two lightweight training recipes, "reasoning pre-training" and "reasoning dropout," that retain most of the benefits of test-time reasoning while maintaining fast inference. The authors hypothesize three mechanisms: better representation learning, an implicit learning curriculum, and increased effective expressivity. They isolate these by testing reasoning pre-training, co-training, dropout, scaffolding, and non-semantic thinking tokens. On LIBERO-90, reasoning dropout reaches 89.4% on the standard split and 76.4% average across variants, versus 82.0%/67.4% for the standard VLA baseline; reasoning pre-training reaches 87.1%/72.8%. On Bridge real-robot tasks, both recipes outperform the standard VLA baseline on aggregate, with reasoning pre-training at 69.4% versus 50.5% for the baseline. The paper concludes that learning to generate reasonings improves VLA representations and that attending to reasonings during training is what matters, while the semantic content of thinking tokens is irrelevant. The two proposed recipes do not generate reasoning at test time and are about 3x faster than full ECoT.
Significance. If the central claim is correct, the paper contributes two practical and well-evaluated training strategies that give a large part of embodied CoT's performance benefit without the inference-time cost, which would be useful for real-time robot control. The empirical scale is a clear strength: 121,500 simulated trials and 444 real-robot trials, with shared evaluation seeds and reported standard errors. The thinking-token control is also a valuable negative result: it suggests that merely adding compute tokens is not enough, consistent with the paper's emphasis on semantically meaningful reasoning. However, the scientific mechanism claim is not yet fully established: the most effective recipe, reasoning dropout, is not compared against a non-semantic token-dropout control, and reasoning pre-training is not compute-matched to the baseline. The paper's own Limitations section (Sec. 9) acknowledges that the learning-dynamics mechanism is not investigated and describes the explanations as "intuitive speculation." The practical recipes may still stand, but the causal interpretation in the abstract requires additional control experiments.
major comments (4)
- [Sec. 5 (Reasoning Dropout), E.3, Table 1] The reasoning-dropout recipe is confounded with generic dropout regularization. The thinking-token condition (E.5.2) inserts filler tokens without applying the same stochastic dropout schedule, and there is no control that drops an equivalent number of non-semantic tokens (e.g., random text tokens or image patches) on the same schedule. Because the LIBERO training set is only 3,917 trajectories, stochastic masking of context tokens could plausibly produce the observed +7.4-point average gain without any reasoning-specific representation change. This missing control leaves the abstract's central claim—that learning to generate reasonings improves VLA representations—underdetermined.
- [Sec. 6.1, E.1, Table 1] Reasoning pre-training is not compute-matched to the standard VLA baseline: it uses 100k steps of reasoning pre-training followed by 200k steps of action tuning (300k total) versus 200k for the standard VLA. The +5.4% average improvement could therefore be due to additional gradient steps rather than to reasoning-specific representation learning. The co-training condition uses 200k steps with a doubled batch size, which changes the number of optimizer updates per sample. A compute-matched baseline (e.g., training a standard VLA for 300k steps) or a learning-curve analysis is needed to attribute the gain to the reasoning content.
- [Sec. 6.2, Table 2] The text says that both ECoT-Lite approaches "improve significantly" over the standard VLA on Bridge, but the reported standard errors do not support this for reasoning dropout: the aggregate is 60.4% ± 4.6% versus 50.5% ± 4.7% for the VLA baseline, a difference of about 1.5 combined standard errors. Additionally, the aggregate hides large per-task reversals (e.g., reasoning pre-training at 37.5% vs. 87.5% for a spatial task). The authors should either report a proper significance test with per-task paired comparisons or soften the significance claim.
- [Sec. 9, A.1, Sec. 6.3] The paper's learning-dynamics explanation for why pre-training outperforms co-training is explicitly labeled as "intuitive speculation" in the Limitations section, and the supporting illustration in Figure 7 is an abstract loss-landscape cartoon. Since the paper's central scientific contribution is the claim that reasoning changes the policy's representations, downstream success alone is only indirect evidence. Adding a direct representation probe (e.g., linear probing for object or subtask features in the intermediate layers after different trainings, or measuring representational similarity) would provide much stronger support for the mechanism and would be within the paper's scope.
minor comments (5)
- [C.1] Typo: "such as Molmo's synthetic reasonings are much ore verbose" should be "much more verbose."
- [Table 2] In the Semantic Gen. row, "Put the screw in the bowel" should be "bowl."
- [A.2] The phrase "Disabling reasonings does not see to affect the spatial generalization tasks" should be "does not seem to affect."
- [Introduction / Abstract] The abstract's claim of outperforming conventional VLAs on BridgeData V2 "by 10-19%" should be qualified with the reported standard errors, since the lower end of that range is within one standard error of the baseline for the dropout variant.
- [E.3] The description of how the reasoning dropout policy can be run with reasoning "turned on" at test time is useful, but the paper does not report the performance of the dropout-trained policy when reasoning is actually enabled at test time on LIBERO; adding this would directly quantify the value of test-time reasoning within the same model.
Circularity Check
No significant circularity: the paper's central claims are empirical ablation results tested against external baselines, not conclusions encoded into the training objectives.
full rationale
The paper's central claim—that learning to generate embodied reasonings improves VLA representations and that attending to reasonings helps leverage these features—is an empirical inference drawn from measured success rates on LIBERO and Bridge, not a quantity that is defined in terms of the outcome it predicts. Each ECoT-Lite recipe is specified by concrete training-objective choices (pre-training vs. co-training, dropout of reasoning tokens, scaffolding without reasoning loss, and non-semantic thinking tokens), and the recipes are compared against standard VLA and full ECoT policies trained on the same demonstration and reasoning data. The thinking-token condition controls for token count and compute, and the reasoning-scaffolding condition controls for in-context presence without generation loss; these are genuine experimental contrasts, not algebraic identities. Citations to prior work, including the authors' own ECoT paper [14], supply the reasoning data, the representative reasoning policy, and a baseline, but the paper's hypothesis tests do not reduce to that citation: the benchmark comparisons to standard VLA and to external state-of-the-art results (Mete et al. [69]) provide independent evidence for the practical recipes. The Limitations section explicitly states that the learning-dynamics mechanism is not investigated and describes the explanations as 'intuitive speculation,' which further shows that the authors are not presenting a derived conclusion as forced. The main weakness—that reasoning dropout is not compared to dropping an equivalent number of non-semantic context tokens—undermines the causal interpretation of the representation-learning mechanism, but underdetermination of a mechanism by experiments is a correctness or validity concern, not circularity. No fitted parameter is renamed as a prediction, no result is assumed by construction, and no self-citation is invoked as a uniqueness proof or as the sole justification for the central claim. Therefore the derivation chain is self-contained with respect to circularity, and the appropriate score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Behavioral cloning on robot demonstrations with next-token prediction yields a valid policy learning approach.
- domain assumption The LIBERO-90 and BridgeData V2 evaluation protocols are representative proxies for robot manipulation generalization.
- domain assumption The embodied reasoning annotations (plans, subtasks, bounding boxes, motions) are sufficiently accurate and informative for the training signal to be useful.
- standard math Pre-trained VLM representations can be adapted to action prediction via fine-tuning.
Cite this review
Pith. "Pith review of Training Strategies for Efficient Embodied Reasoning." pith.science (2026). https://pith.science/paper/WGBIYQP3
@misc{pith2026250508243,
author = {Pith},
title = {Pith review of: Training Strategies for Efficient Embodied Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/WGBIYQP3}},
note = {Machine review of arXiv:2505.08243}
}
read the original abstract
Robot chain-of-thought reasoning (CoT) -- wherein a model predicts helpful intermediate representations before choosing actions -- provides an effective method for improving the generalization and performance of robot policies, especially vision-language-action models (VLAs). While such approaches have been shown to improve performance and generalization, they suffer from core limitations, like needing specialized robot reasoning data and slow inference speeds. To design new robot reasoning approaches that address these issues, a more complete characterization of why reasoning helps policy performance is critical. We hypothesize several mechanisms by which robot reasoning improves policies -- (1) better representation learning, (2) improved learning curricularization, and (3) increased expressivity -- then devise simple variants of robot CoT reasoning to isolate and test each one. We find that learning to generate reasonings does lead to better VLA representations, while attending to the reasonings aids in actually leveraging these features for improved action prediction. Our results provide us with a better understanding of why CoT reasoning helps VLAs, which we use to introduce two simple and lightweight alternative recipes for robot reasoning. Our proposed approaches achieve significant performance gains over non-reasoning policies, state-of-the-art results on the LIBERO-90 benchmark, and a 3x inference speedup compared to standard robot reasoning.
Figures
Figures from the paper (12 more)
Forward citations
Cited by 11 Pith papers
-
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
Physical adversarial patches can steer CoT-reasoning VLAs into attacker-specified manipulation behaviors without changing the user’s instruction.
-
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.
-
Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models
Latent-iterative reasoning makes a vision-language-action robot policy far more fragile to input perturbation, and text-based plan monitors collapse under adaptive attacks.
-
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Pre-training a VLA model on 100k hours of auto-labeled UMI trajectories, then post-training on robot data, yields SOTA simulated manipulation and data-efficient fine-tuning.
-
APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts
A VLM planner that adaptively inserts latent visual thoughts of future states into its reasoning trace beats language-only and prior VLM planners on long-horizon kitchen tasks, especially under tight free space.
-
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
A learned critic of observation grounding and stepwise coherence, used as dense RL reward, improves VLA planner faithfulness and OOD hazard responsiveness while preserving competitive trajectory accuracy.
-
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
Pretraining a vision-language model to output discrete 3D pose tokens on large non-robotic data, before training a robot action head, improves downstream manipulation success and data efficiency.
-
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
A 6K-question embodied-reasoning benchmark plus a flow-matching action tokenizer let one 3B vision-language model reason and manipulate better than continuous- or discrete-action VLA baselines.
-
EVE: A Generator-Verifier System for Generative Policies
Zero-shot VLM verifiers, ensembled and fused via guided diffusion, improve frozen generative robot policies' success rates by 1-2 percentage points on simulated manipulation tasks.
-
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
GraphCoT-VLA couples structured chain-of-thought planning and a real-time 3D object-pose graph to improve robot manipulation under vague instructions.
-
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
A survey that groups VLA efficiency techniques into four categories: model architecture, perception features, action generation, and training/inference strategies.
Reference graph
Works this paper leans on
-
[1]
O’Neill, A
Embodiment Collaboration, A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. Xie, A. Brohan, A. Raffin, A. Sharma, A. Yavary, A. Jain, A. Balakrishna, A. Wahid, B. Bur...
2024
-
[2]
Brohan, N
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manju- nath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsc...
2023
-
[3]
Ghosh, H
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y . Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy. In Proceedings of Robotics: Science and Systems, Delft, Netherlands, 2024
2024
-
[4]
Doshi, H
R. Doshi, H. Walke, O. Mees, S. Dasari, and S. Levine. Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation. In Conference on Robot Learning, 2024
2024
-
[5]
Brohan, N
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y . Lu, H. Michalewski, I. Mordatch, K. Pe...
2023
-
[6]
M. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model. 2024
2024
-
[7]
Ebert, Y
F. Ebert, Y . Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine. Bridge data: Boosting generalization of robotic skills with cross-domain datasets, 2021
2021
-
[8]
Walke, K
H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V . Myers, K. Fang, C. Finn, and S. Levine. Bridgedata v2: A dataset for robot learning at scale. In Conference on Robot Learning (CoRL) , 2023
2023
Show all 86 references
-
[9]
Rosete-Beas, O
E. Rosete-Beas, O. Mees, G. Kalweit, J. Boedecker, and W. Burgard. Latent plans for task agnostic offline reinforcement learning. InProceedings of the 6th Conference on Robot Learning (CoRL), Auckland, New Zealand, 2022
2022
-
[10]
Black, N
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky.π0: A vision-la...
2024
-
[11]
Bjorck, F
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734, 2025
2025 arXiv
-
[12]
Jones, O
J. Jones, O. Mees, C. Sferrazza, K. Stachowicz, P. Abbeel, and S. Levine. Beyond sight: Finetuning generalist robot policies with heterogeneous sensors via language grounding. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) , Atlanta, USA, 2025
2025
-
[13]
H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y . Du, Y . Hong, and C. Gan. 3d-vla: A 3d vision-language-action generative world model. arXiv preprint arXiv:2403.09631, 2024
2024 arXiv
-
[14]
Zawalski, W
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine. Robotic control via embodied chain-of-thought reasoning. In Conference on Robot Learning, 2024
2024
-
[15]
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023
2023
-
[16]
Kojima, S
T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, and Y . Iwasawa. Large language models are zero-shot reasoners, 2023
2023
-
[17]
Zheng, Y
R. Zheng, Y . Liang, S. Huang, J. Gao, H. Daumé III, A. Kolobov, F. Huang, and J. Yang. Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies. arXiv preprint arXiv:2412.10345, 2024
2024 arXiv
-
[18]
R. Xu, J. Zhang, M. Guo, Y . Wen, H. Yang, M. Lin, J. Huang, Z. Li, K. Zhang, L. Wang, et al. A0: An affordance-aware hierarchical model for general robotic manipulation. arXiv preprint arXiv:2504.12636, 2025
2025
-
[19]
Borja-Diaz, O
J. Borja-Diaz, O. Mees, G. Kalweit, L. Hermann, J. Boedecker, and W. Burgard. Affordance learning from play for sample-efficient policy learning. InProceedings of the IEEE International Conference on Robotics and Automation (ICRA) , Philadelphia, USA, 2022
2022
-
[20]
B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning, 2023. URL https://arxiv.org/abs/2306.03310
2023 arXiv
-
[21]
Karamcheti, S
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh. Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024
2024
-
[22]
Beyer, A
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alab- dulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Boš...
2024 arXiv
-
[23]
Steiner, A
A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y . Bitton, A. Gritsenko, M. Min- derer, A. Sherbondy, S. Long, S. Qin, R. Ingle, E. Bugliarello, S. Kazemzadeh, T. Mesnard, I. Alabdulmohsin, L. Beyer, and X. Zhai. Paligemma 2: A family of versatile vlms for transfer,
-
[24]
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
2023
-
[25]
A. Szot, B. Mazoure, O. Attia, A. Timofeev, H. Agrawal, D. Hjelm, Z. Gan, Z. Kira, and A. Toshev. From multimodal llms to generalist embodied agents: Methods and lessons. arXiv preprint arXiv:2412.08442, 2024. 11
2024 arXiv
-
[26]
Huang, F
H. Huang, F. Liu, L. Fu, T. Wu, M. Mukadam, J. Malik, K. Goldberg, and P. Abbeel. Otter: A vision-language-action model with text-aware visual feature extraction. arXiv preprint arXiv:2503.03734, 2025
2025
-
[27]
J. Wen, Y . Zhu, J. Li, M. Zhu, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen, Y . Peng, F. Feng, and J. Tang. Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation. arXiv preprint arXiv:2409.12514, 2024
2024 arXiv
-
[28]
J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu, et al. Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model. arXiv preprint arXiv:2503.10631, 2025
2025 arXiv
-
[29]
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu. Rdt-1b: a diffusion foundation model for bimanual manipulation. arXiv preprint arXiv:2410.07864, 2024
2024 arXiv
-
[30]
Q. Li, Y . Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y . Deng, S. Xu, Y . Zhang, et al. Cogact: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation. arXiv preprint arXiv:2411.19650, 2024
2024 arXiv
-
[31]
Belkhale and D
S. Belkhale and D. Sadigh. Minivla: A better vla with a smaller footprint, 2024. URL https://ai.stanford.edu/blog/minivla/
2024
-
[32]
Pertsch, K
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models, 2025. URL https://arxiv.org/abs/2501.09747
2025 arXiv
-
[33]
Abeyruwan, J
Gemini Robotics Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Arm- strong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, S. Bohez, K. Bousmalis, A. Brohan, T. Buschmann, A. Byravan, S. Cabi, K. Caluwaerts, F. Casarini, O. Chang, J. E. Chen, X. Chen, H.-T....
2025 arXiv
-
[34]
J. Wen, Y . Zhu, J. Li, Z. Tang, C. Shen, and F. Feng. Dexvla: Vision-language model with plug-in diffusion expert for general robot control. arXiv preprint arXiv:2502.05855, 2025
2025 arXiv
-
[35]
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song. Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots. In Proceedings of Robotics: Science and Systems (RSS) , 2024
2024
-
[36]
Etukuru, N
H. Etukuru, N. Naka, Z. Hu, S. Lee, J. Mehu, A. Edsinger, C. Paxton, S. Chintala, L. Pinto, and N. M. M. Shafiullah. Robot utility models: General policies for zero-shot deployment in new environments. arXiv preprint arXiv:2409.05865, 2024
2024 arXiv
-
[37]
N. M. M. Shafiullah, A. Rai, H. Etukuru, Y . Liu, I. Misra, S. Chintala, and L. Pinto. On bringing robots home. arXiv preprint arXiv:2311.16098, 2023
2023 arXiv
-
[38]
F. Lin, Y . Hu, P. Sheng, C. Wen, J. You, and Y . Gao. Data scaling laws in imitation learning for robotic manipulation. arXiv preprint arXiv:2410.18647, 2024. 12
2024 arXiv
-
[39]
Khazatsky, K
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Pa...
2024
-
[40]
Sprague, F
Z. Sprague, F. Yin, J. D. Rodriguez, D. Jiang, M. Wadhwa, P. Singhal, X. Zhao, X. Ye, K. Mahowald, and G. Durrett. To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning. arXiv preprint arXiv:2409.12183, 2024
2024 arXiv
-
[41]
Snell, J
C. Snell, J. Lee, K. Xu, and A. Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024. URL https://arxiv.org/abs/2408.03314
2024 arXiv
-
[42]
G. Feng, B. Zhang, Y . Gu, H. Ye, D. He, and L. Wang. Towards revealing the mystery behind chain of thought: A theoretical perspective, 2023. URL https://arxiv.org/abs/2305. 15408
2023
-
[43]
Merrill and A
W. Merrill and A. Sabharwal. The expressive power of transformers with chain of thought,
-
[44]
Z. Li, H. Liu, D. Zhou, and T. Ma. Chain of thought empowers transformers to solve inherently serial problems, 2024. URL https://arxiv.org/abs/2402.12875
2024 arXiv
-
[45]
URL https://arxiv.org/abs/2310.07923
-
[46]
Q. Zhao, Y . Lu, M. J. Kim, Z. Fu, Z. Zhang, Y . Wu, Z. Li, Q. Ma, S. Han, C. Finn, A. Handa, M.-Y . Liu, D. Xiang, G. Wetzstein, and T.-Y . Lin. Cot-vla: Visual chain-of-thought reasoning for vision-language-action models, 2025. URL https://arxiv.org/abs/2503.22020
2025 arXiv
-
[47]
Hwang, R
J.-J. Hwang, R. Xu, H. Lin, W.-C. Hung, J. Ji, K. Choi, D. Huang, T. He, P. Covington, B. Sapp, Y . Zhou, J. Guo, D. Anguelov, and M. Tan. Emma: End-to-end multimodal model for autonomous driving, 2024. URL https://arxiv.org/abs/2410.23262
2024 arXiv
-
[48]
Sharma, A
P. Sharma, A. Torralba, and J. Andreas. Skill induction and planning with latent language, 2022
2022
-
[49]
Clark, S
J. Clark, S. Mirchandani, D. Sadigh, and S. Belkhale. Action-free reasoning for policy general- ization, 2025. URL https://arxiv.org/abs/2502.03729
2025 arXiv
-
[50]
Liang, W
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng. Code as policies: Language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages 9493–9500. IEEE, 2023
2023
-
[51]
Huang, F
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y . Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter. Inner monologue: Embodied reasoning through planning with language models, 2022
2022
-
[52]
Belkhale, T
S. Belkhale, T. Ding, T. Xiao, P. Sermanet, Q. Vuong, J. Tompson, Y . Chebotar, D. Dwibedi, and D. Sadigh. Rt-h: Action hierarchies using language, 2024. 13
2024
-
[53]
O. Mees, J. Borja-Diaz, and W. Burgard. Grounding language with visual affordances over unstructured data. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK, 2023
2023
-
[54]
D. Niu, Y . Sharma, G. Biamby, J. Quenum, Y . Bai, B. Shi, T. Darrell, and R. Herzig. Llarva: Vision-action instruction tuning enhances robot learning, 2024
2024
-
[55]
J. Gu, S. Kirmani, P. Wohlhart, Y . Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, P. Sundaresan, P. Xu, H. Su, K. Hausman, C. Finn, Q. Vuong, and T. Xiao. Rt- trajectory: Robotic task generalization via hindsight trajectory sketches, 2023. URL https: //arxi...
2023 arXiv
-
[56]
W. Chen, O. Mees, A. Kumar, and S. Levine. Vision-language models provide promptable representations for reinforcement learning. Transactions on Machine Learning Research, 2025
2025
-
[57]
H. Zhou, X. Yao, O. Mees, Y . Meng, T. Xiao, Y . Bisk, J. Oh, E. Johns, M. Shridhar, D. Shah, J. Thomason, K. Huang, J. Chai, Z. Bing, and A. Knoll. Bridging language and action: A survey of language-based robot manipulation. arXiv preprint arXiv:2312.10807, 2024
2024 arXiv
-
[58]
Karamcheti, S
S. Karamcheti, S. Nair, A. S. Chen, T. Kollar, C. Finn, D. Sadigh, and P. Liang. Language-driven representation learning for robotics, 2023
2023
-
[59]
Majumdar, K
A. Majumdar, K. Yadav, S. Arnaud, Y . J. Ma, C. Chen, S. Silwal, A. Jain, V .-P. Berges, P. Abbeel, J. Malik, D. Batra, Y . Lin, O. Maksymets, A. Rajeswaran, and F. Meier. Where are we in the search for an artificial visual cortex for embodied intelligence?, 2024. URL https://...
2024 arXiv
-
[60]
Driess, F
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y . Chebotar, P. Sermanet, D. Duckworth, S. Levine, V . Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence. Palm-e: An emb...
2023
-
[61]
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Driess, P. Florence, D. Sadigh, L. Guibas, and F. Xia. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities, 2024. URL https://arxiv.org/abs/2401.12168
2024 arXiv
-
[63]
Burns, Z
K. Burns, Z. Witzel, J. I. Hamid, T. Yu, C. Finn, and K. Hausman. What makes pre-trained visual representations successful for robust manipulation?, 2023. URL https://arxiv.org/ abs/2312.12444
2023 arXiv
-
[64]
J. Pfau, W. Merrill, and S. R. Bowman. Let’s think dot by dot: Hidden computation in transformer language models, 2024. URL https://arxiv.org/abs/2404.15758
2024 arXiv
-
[65]
K. Kang, A. Setlur, D. Ghosh, J. Steinhardt, C. Tomlin, S. Levine, and A. Kumar. What do learning dynamics reveal about generalization in llm reasoning?, 2024. URL https: //arxiv.org/abs/2411.07681
2024 arXiv
-
[66]
Gururangan, A
S. Gururangan, A. Marasovi ´c, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith. Don’t stop pretraining: Adapt language models to domains and tasks, 2020. URL https://arxiv.org/abs/2004.10964
2020 arXiv
-
[67]
Goyal, Z
S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V . Nagarajan. Think before you speak: Training language models with pause tokens, 2024. URL https://arxiv.org/abs/ 2310.02226
2024 arXiv
-
[68]
Brief, O
M. Brief, O. Ovadia, G. Shenderovitz, N. B. Yoash, R. Lemberg, and E. Sheetrit. Mixing it up: The cocktail effect of multi-task fine-tuning on llm performance – a case study in finance, 2024. URL https://arxiv.org/abs/2410.01109. 14
2024 arXiv
-
[69]
H. Lang, M. Agrawal, Y . Kim, and D. Sontag. Co-training improves prompt-based learning for large language models, 2022. URL https://arxiv.org/abs/2202.00828
2022 arXiv
-
[70]
J. Gao, S. Belkhale, S. Dasari, A. Balakrishna, D. Shah, and D. Sadigh. A taxonomy for evaluating generalist robot policies, 2025. URL https://arxiv.org/abs/2503.01238
2025
-
[71]
A. Mete, H. Xue, A. Wilcox, Y . Chen, and A. Garg. Quest: Self-supervised skill abstractions for learning continuous control, 2024. URL https://arxiv.org/abs/2407.15840
2024 arXiv
-
[72]
Tensorrt-llm
NVIDIA. Tensorrt-llm. https://github.com/NVIDIA/TensorRT-LLM?tab= readme-ov-file, 2024
2024
-
[73]
Y . Ren, S. Guo, W. Bae, and D. J. Sutherland. How to prepare your task head for finetuning,
-
[74]
Kumar, A
A. Kumar, A. Raghunathan, R. Jones, T. Ma, and P. Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution, 2022. URL https://arxiv.org/abs/2202. 10054
2022
-
[75]
M. J. Kim, C. Finn, and P. Liang. Fine-tuning vision-language-action models: Optimizing speed and success, 2025. URL https://arxiv.org/abs/2502.19645
2025 arXiv
-
[76]
W. Chen, M. Zawalski, K. Pertsch, O. Mees, C. Finn, and S. Levine. Tensorrt-openvla, 2025. URL https://github.com/rail-berkeley/tensorrt-openvla
2025
-
[77]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Ho...
2023
-
[78]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. ...
2024
-
[79]
Deitke, C
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y . Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, J. Lu, T. Anderson, E. Bransom, K. Ehsani, H. Ngo, Y . Chen, A. Patel, M. Yatskar, C. Callison-Burch, A. Head, R. Hendrix, F. Bastani, E. VanderBilt, N. Lambert, Y . ...
2024 arXiv
-
[80]
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware, 2023
2023
-
[81]
van den Oord, O
A. van den Oord, O. Vinyals, and K. Kavukcuoglu. Neural discrete representation learning, 2017
2017
-
[82]
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid loss for language image pre-training, 2023
2023
-
[83]
Merrill and A
W. Merrill and A. Sabharwal. The parallelism tradeoff: Limitations of log-precision transform- ers, 2023. URL https://arxiv.org/abs/2207.00729
2023 arXiv
-
[84]
preserve
A. Yao. Circuits and local computation. 16 Table 1: Performance of all proposed approaches and baselines on all Libero-90 benchmark variants (Mean± StdErr). Success rates are computed for 50 trials on all 90 tasks (4500 episodes per variant, or 13500 total episodes per policy)...
-
[85]
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, ...
2024
-
[2023]
URL https://arxiv.org/abs/2302.05779
-
[2024]
URL https://arxiv.org/abs/2412.03555
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.