Pith. sign in

REVIEW 4 major objections 5 minor 11 cited by

Training Strategies for Efficient Embodied Reasoning

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Robot reasoning pays off in training, not at slow inference.

desk verdict Useful recipes and solid LIBERO ablations; the mechanism claim is under-tested and the Bridge dropout numbers come from a model not trained with the dropout recipe. read the letter →

arxiv 2505.08243 v2 pith:WGBIYQP3 submitted 2025-05-13 cs.RO

classification cs.RO
keywords embodiedchain-of-thoughtreasoningvision-language-actionmodelsrepresentationlearningdropoutpre-trainingrobotpolicygeneralizationLIBERO-90real-robotmanipulation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks why chain-of-thought reasoning improves vision-language-action (VLA) robot policies and tests three candidate mechanisms: better representations, a learning curriculum, and extra compute. Its answer is that the main benefit is representational: training a VLA to generate reasoning traces such as plans, bounding boxes, and motion rationales reshapes its internal features, and the policy exploits those features best when actions are trained to attend to them. The paper introduces two lightweight recipes, reasoning pre-training and reasoning dropout, that capture most of this benefit without generating any reasoning at test time. If right, robot policies can gain the generalization advantages of embodied reasoning at the same speed as a standard VLA, a roughly threefold inference speedup over full chain-of-thought control.

What carries the argument

The central machinery is a pair of training-time interventions that decouple learning to reason from reasoning at inference: reasoning pre-training, which supervises the VLM on embodied reasoning traces alone and then fine-tunes on actions, and reasoning dropout, which trains the VLA to predict actions while randomly masking or keeping the reasoning tokens in context. Combined with the embodied reasoning annotations themselves (plans, subtask rationales, bounding boxes, gripper positions, and movement rationales), these recipes force the action head to rely on reasoning-shaped representations during training while removing the need to decode reasonings during deployment.

What would settle it

Train the same architecture with the reasoning-dropout protocol but replace the dropped reasoning steps with an equal number of random text tokens at matched sequence length. If the perturbed-split success rates stay near 70-76%, the semantic content of reasoning is not load-bearing and the effect is generic regularization; if performance drops back toward the standard-VLA baseline, the reasoning content itself is doing the work.

Watch

Extended reading notes

Core claim

The paper's central claim is that learning to generate reasonings leads to better VLA representations, while attending to the reasonings during training is what lets the policy leverage those features for improved action prediction. Experimentally, reasoning dropout and reasoning pre-training outperform a standard VLA on LIBERO-90 (89.4% and 87.1% versus 82.0% on the standard split) and on real-robot Bridge tasks (69.4% versus 50.5% for pre-training), matching or approaching full embodied chain-of-thought while running at the faster speed of a plain VLA. The paper finds only weak support for the learning-curriculum hypothesis and no support for the expressivity hypothesis, since adding non-semantic thinking tokens actually hurts performance.

Load-bearing premise

The conclusion that reasoning improves policies through semantic representation learning rests on the untested assumption that the gains of reasoning dropout come from the meaning of the reasoning tokens rather than from the generic regularization effect of randomly dropping tokens during training.

Editorial extensions

If this is right

  • Policies trained with reasoning pre-training or dropout can be deployed at the control rate of a standard VLA, not the 1-1.2 Hz of full embodied CoT, while matching or beating it on LIBERO-90.
  • Reasoning dropout reaches state-of-the-art LIBERO-90 performance (89.4%) without generating any test-time reasoning, indicating that in narrow benchmarks the reasoning traces can be fully internalized.
  • Reasoning pre-training lifts real-robot Bridge success from 50.5% to 69.4%, and because it does not need paired reasoning-action data, it can in principle learn from unpaired reasoning data from other embodiments.
  • The paper's practical prescription follows from its ablations: use full ECoT to maximize performance, reasoning dropout for narrow task domains, and reasoning pre-training for diverse domains or when unpaired reasoning data are available.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the mechanism is semantic representation learning, then scaling the diversity of reasoning data should widen performance gaps between the recipes and may make test-time reasoning increasingly valuable, a trend the paper's own LIBERO-versus-Bridge comparison already hints at.
  • A direct test of the mechanism would compare reasoning dropout against dropping matched-length random text tokens; if the gains are equivalent, the benefit is generic regularization rather than reasoning content.
  • The null result for thinking tokens suggests that extra inference compute without a meaningful training signal will not transfer to manipulation, and a further control could insert semantically meaningful but task-irrelevant tokens to isolate content from mere sequence length.
  • Reasoning pre-training may generalize as a recipe for adapting any VLM before action fine-tuning, effectively extending domain-adaptive pretraining to robot control.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper investigates why embodied chain-of-thought (ECoT) reasoning improves vision-language-action (VLA) policies and then proposes two lightweight training recipes, "reasoning pre-training" and "reasoning dropout," that retain most of the benefits of test-time reasoning while maintaining fast inference. The authors hypothesize three mechanisms: better representation learning, an implicit learning curriculum, and increased effective expressivity. They isolate these by testing reasoning pre-training, co-training, dropout, scaffolding, and non-semantic thinking tokens. On LIBERO-90, reasoning dropout reaches 89.4% on the standard split and 76.4% average across variants, versus 82.0%/67.4% for the standard VLA baseline; reasoning pre-training reaches 87.1%/72.8%. On Bridge real-robot tasks, both recipes outperform the standard VLA baseline on aggregate, with reasoning pre-training at 69.4% versus 50.5% for the baseline. The paper concludes that learning to generate reasonings improves VLA representations and that attending to reasonings during training is what matters, while the semantic content of thinking tokens is irrelevant. The two proposed recipes do not generate reasoning at test time and are about 3x faster than full ECoT.

Significance. If the central claim is correct, the paper contributes two practical and well-evaluated training strategies that give a large part of embodied CoT's performance benefit without the inference-time cost, which would be useful for real-time robot control. The empirical scale is a clear strength: 121,500 simulated trials and 444 real-robot trials, with shared evaluation seeds and reported standard errors. The thinking-token control is also a valuable negative result: it suggests that merely adding compute tokens is not enough, consistent with the paper's emphasis on semantically meaningful reasoning. However, the scientific mechanism claim is not yet fully established: the most effective recipe, reasoning dropout, is not compared against a non-semantic token-dropout control, and reasoning pre-training is not compute-matched to the baseline. The paper's own Limitations section (Sec. 9) acknowledges that the learning-dynamics mechanism is not investigated and describes the explanations as "intuitive speculation." The practical recipes may still stand, but the causal interpretation in the abstract requires additional control experiments.

major comments (4)
  1. [Sec. 5 (Reasoning Dropout), E.3, Table 1] The reasoning-dropout recipe is confounded with generic dropout regularization. The thinking-token condition (E.5.2) inserts filler tokens without applying the same stochastic dropout schedule, and there is no control that drops an equivalent number of non-semantic tokens (e.g., random text tokens or image patches) on the same schedule. Because the LIBERO training set is only 3,917 trajectories, stochastic masking of context tokens could plausibly produce the observed +7.4-point average gain without any reasoning-specific representation change. This missing control leaves the abstract's central claim—that learning to generate reasonings improves VLA representations—underdetermined.
  2. [Sec. 6.1, E.1, Table 1] Reasoning pre-training is not compute-matched to the standard VLA baseline: it uses 100k steps of reasoning pre-training followed by 200k steps of action tuning (300k total) versus 200k for the standard VLA. The +5.4% average improvement could therefore be due to additional gradient steps rather than to reasoning-specific representation learning. The co-training condition uses 200k steps with a doubled batch size, which changes the number of optimizer updates per sample. A compute-matched baseline (e.g., training a standard VLA for 300k steps) or a learning-curve analysis is needed to attribute the gain to the reasoning content.
  3. [Sec. 6.2, Table 2] The text says that both ECoT-Lite approaches "improve significantly" over the standard VLA on Bridge, but the reported standard errors do not support this for reasoning dropout: the aggregate is 60.4% ± 4.6% versus 50.5% ± 4.7% for the VLA baseline, a difference of about 1.5 combined standard errors. Additionally, the aggregate hides large per-task reversals (e.g., reasoning pre-training at 37.5% vs. 87.5% for a spatial task). The authors should either report a proper significance test with per-task paired comparisons or soften the significance claim.
  4. [Sec. 9, A.1, Sec. 6.3] The paper's learning-dynamics explanation for why pre-training outperforms co-training is explicitly labeled as "intuitive speculation" in the Limitations section, and the supporting illustration in Figure 7 is an abstract loss-landscape cartoon. Since the paper's central scientific contribution is the claim that reasoning changes the policy's representations, downstream success alone is only indirect evidence. Adding a direct representation probe (e.g., linear probing for object or subtask features in the intermediate layers after different trainings, or measuring representational similarity) would provide much stronger support for the mechanism and would be within the paper's scope.
minor comments (5)
  1. [C.1] Typo: "such as Molmo's synthetic reasonings are much ore verbose" should be "much more verbose."
  2. [Table 2] In the Semantic Gen. row, "Put the screw in the bowel" should be "bowl."
  3. [A.2] The phrase "Disabling reasonings does not see to affect the spatial generalization tasks" should be "does not seem to affect."
  4. [Introduction / Abstract] The abstract's claim of outperforming conventional VLAs on BridgeData V2 "by 10-19%" should be qualified with the reported standard errors, since the lower end of that range is within one standard error of the baseline for the dropout variant.
  5. [E.3] The description of how the reasoning dropout policy can be run with reasoning "turned on" at test time is useful, but the paper does not report the performance of the dropout-trained policy when reasoning is actually enabled at test time on LIBERO; adding this would directly quantify the value of test-time reasoning within the same model.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's central claims are empirical ablation results tested against external baselines, not conclusions encoded into the training objectives.

full rationale

The paper's central claim—that learning to generate embodied reasonings improves VLA representations and that attending to reasonings helps leverage these features—is an empirical inference drawn from measured success rates on LIBERO and Bridge, not a quantity that is defined in terms of the outcome it predicts. Each ECoT-Lite recipe is specified by concrete training-objective choices (pre-training vs. co-training, dropout of reasoning tokens, scaffolding without reasoning loss, and non-semantic thinking tokens), and the recipes are compared against standard VLA and full ECoT policies trained on the same demonstration and reasoning data. The thinking-token condition controls for token count and compute, and the reasoning-scaffolding condition controls for in-context presence without generation loss; these are genuine experimental contrasts, not algebraic identities. Citations to prior work, including the authors' own ECoT paper [14], supply the reasoning data, the representative reasoning policy, and a baseline, but the paper's hypothesis tests do not reduce to that citation: the benchmark comparisons to standard VLA and to external state-of-the-art results (Mete et al. [69]) provide independent evidence for the practical recipes. The Limitations section explicitly states that the learning-dynamics mechanism is not investigated and describes the explanations as 'intuitive speculation,' which further shows that the authors are not presenting a derived conclusion as forced. The main weakness—that reasoning dropout is not compared to dropping an equivalent number of non-semantic context tokens—undermines the causal interpretation of the representation-learning mechanism, but underdetermination of a mechanism by experiments is a correctness or validity concern, not circularity. No fitted parameter is renamed as a prediction, no result is assumed by construction, and no self-citation is invoked as a uniqueness proof or as the sole justification for the central claim. Therefore the derivation chain is self-contained with respect to circularity, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper does not fit any data-derived free parameters; it uses fixed hyperparameters and design choices. It relies on standard VLA training assumptions and on the quality of the reasoning data, neither of which is independently validated within the paper.

assumptions (4)
  • domain assumption Behavioral cloning on robot demonstrations with next-token prediction yields a valid policy learning approach.
    This is the standard VLA training framework used throughout, introduced in Section 3, and all conclusions depend on it.
  • domain assumption The LIBERO-90 and BridgeData V2 evaluation protocols are representative proxies for robot manipulation generalization.
    The paper's claims about generalization rely on these benchmarks; Section 6.1 describes the evaluation tasks and distribution shifts.
  • domain assumption The embodied reasoning annotations (plans, subtasks, bounding boxes, motions) are sufficiently accurate and informative for the training signal to be useful.
    The recipes transfer knowledge through these annotations, and their quality is not directly measured; Section C describes the semi-automated generation pipeline.
  • standard math Pre-trained VLM representations can be adapted to action prediction via fine-tuning.
    This is the foundation of all VLA methods cited in Section 2 and is used without re-derivation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Training Strategies for Efficient Embodied Reasoning." pith.science (2026). https://pith.science/paper/WGBIYQP3

@misc{pith2026250508243,
  author       = {Pith},
  title        = {Pith review of: Training Strategies for Efficient Embodied Reasoning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WGBIYQP3}},
  note         = {Machine review of arXiv:2505.08243}
}
read the original abstract

Robot chain-of-thought reasoning (CoT) -- wherein a model predicts helpful intermediate representations before choosing actions -- provides an effective method for improving the generalization and performance of robot policies, especially vision-language-action models (VLAs). While such approaches have been shown to improve performance and generalization, they suffer from core limitations, like needing specialized robot reasoning data and slow inference speeds. To design new robot reasoning approaches that address these issues, a more complete characterization of why reasoning helps policy performance is critical. We hypothesize several mechanisms by which robot reasoning improves policies -- (1) better representation learning, (2) improved learning curricularization, and (3) increased expressivity -- then devise simple variants of robot CoT reasoning to isolate and test each one. We find that learning to generate reasonings does lead to better VLA representations, while attending to the reasonings aids in actually leveraging these features for improved action prediction. Our results provide us with a better understanding of why CoT reasoning helps VLAs, which we use to introduce two simple and lightweight alternative recipes for robot reasoning. Our proposed approaches achieve significant performance gains over non-reasoning policies, state-of-the-art results on the LIBERO-90 benchmark, and a 3x inference speedup compared to standard robot reasoning.

Figures

Figures reproduced from arXiv: 2505.08243 by the authors.

Figure 1
Figure 1. Illustration of our proposed ECoT-Lite approaches. Past robot reasoning policies are performant but slow. By testing numerous hypotheses on why robot reasoning improves policy performance, we find two simple lightweight alternatives for training policies with embodied reasoning data without producing reasonings at test time, boosting performance over non-reasoning VLAs while maintaining fast inference speeds. 1 Intr… view at source ↗
Figure 2
Figure 2. Example intermediate reasoning steps. We use Embodied Chain-of-Thought Reasoning (ECoT [14]) as a representative robot reasoning approach for this work, and thus indicate which steps it does not use with dashed borders (but they are used in other similar works [45, 46]). Robot reasoning. Inspired by the success of CoT for LLMs, there are several works that try to leverage CoT reasoning for robotics [5, 14, 46, 47]. … view at source ↗
Figure 3
Figure 3. ECoT-Lite training recipes. Blue indicates inputs, orange indicates outputs/generations, [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Example ECoT reasonings for LIBERO and Bridge. See [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Top: Performance of all methods on LIBERO-90 benchmarks. The most performant approaches are ECoT and the ECoT-Lite reasoning dropout policy, both of which beat past state-of-the-art on the standard LIBERO-90 evaluations (90.8% and 89.4% vs. 88.6% by Mete et al. [69]). …
Figure 6
Figure 6. Figure 6: More examples of ECoT reasoning steps generated by our policies (LIBERO on top, Bridge on bottom). Red highlights indicate incorrect or hallucinated features (though they still lead to task successes). Note the slight stylistic differences in the LIBERO reasonings, as …
Figure 7
Figure 7. Figure 7: Very abstract illustration of our argument as to why reasoning pre-training seems more effective than co-training, despite using the same data. Blue indicates the loss landscape of the action prediction task, red corresponds to that of reasoning, and darker is lower lo…
Figure 8
Figure 8. Figure 8: Qualitative examples of the importance of test-time robot reasonings. We show policy behaviors on three different Bridge tasks with the reasonings disabled (reasoning dropout) or enabled (full ECoT), as well as the parts of the reasoning that intuitively lead to correc…
Figure 9
Figure 9. Figure 9: First half of the 90 task prompts for LIBERO-90. Note some are repeated in different scenes (and thus they appear multiple times in this list) Disabling reasonings does not see to affect the spatial generalization tasks (which test the policies’ ability to understand s…
Figure 10
Figure 10. Figure 10: Latter half of the 90 task prompts for LIBERO-90. Note some are repeated in different scenes (and thus they appear multiple times in this list) settle) and ends them in failure after 400 more steps have passed without succeeding. The observation and action spaces for …
Figure 11
Figure 11. Figure 11: Full list of task prompts used for our Bridge evaluation, as well as example starting states. We test all combinations of the words in brackets for each scene. these in the scene in regions that do not collide with existing objects. We conduct evaluations under Pertur…
Figure 12
Figure 12. Figure 12: The Llama2 prompt for generating a plan (list of subtasks) from an instruction. I want to analyze each step my robot took to accomplish the task: ’{instruction}.’ To accomplish this task, it takes these subtasks in this order: {numbered list of subtasks} Here is a map…
Figure 13
Figure 13. Figure 13: The Llama2 prompt for aligning the generated subtasks to specific time steps. E.5.2 Practical Implementation Details of Thinking Token Policies To train a thinking token policy, we replace the reasonings in ECoT with a similar number of thinking tokens. Following the …
Figure 14
Figure 14. Figure 14: The Molmo prompt for generating subtask reasonings for each subtask generated via [PITH_FULL_IMAGE:figures/full_fig_p027_14.png]
Figure 15
Figure 15. Figure 15: The Molmo prompt for generating movement reasonings. The VLM also accepts the current observation image. 27 [PITH_FULL_IMAGE:figures/full_fig_p027_15.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches

    cs.CR 2026-03 conditional novelty 7.5 of 10

    Physical adversarial patches can steer CoT-reasoning VLAs into attacker-specified manipulation behaviors without changing the user’s instruction.

  2. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  3. Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Latent-iterative reasoning makes a vision-language-action robot policy far more fragile to input perturbation, and text-based plan monitors collapse under adaptive attacks.

  4. Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Pre-training a VLA model on 100k hours of auto-labeled UMI trajectories, then post-training on robot data, yields SOTA simulated manipulation and data-efficient fine-tuning.

  5. APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A VLM planner that adaptively inserts latent visual thoughts of future states into its reasoning trace beats language-only and prior VLM planners on long-horizon kitchen tasks, especially under tight free space.

  6. Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A learned critic of observation grounding and stepwise coherence, used as dense RL reward, improves VLA planner faithfulness and OOD hazard responsiveness while preserving competitive trajectory accuracy.

  7. PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    Pretraining a vision-language model to output discrete 3D pose tokens on large non-robotic data, before training a robot action head, improves downstream manipulation success and data efficiency.

  8. Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A 6K-question embodied-reasoning benchmark plus a flow-matching action tokenizer let one 3B vision-language model reason and manipulate better than continuous- or discrete-action VLA baselines.

  9. EVE: A Generator-Verifier System for Generative Policies

    cs.RO 2025-12 conditional novelty 6.0 of 10

    Zero-shot VLM verifiers, ensembled and fused via guided diffusion, improve frozen generative robot policies' success rates by 1-2 percentage points on simulated manipulation tasks.

  10. GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions

    cs.RO 2025-08 unverdicted novelty 5.0 of 10

    GraphCoT-VLA couples structured chain-of-thought planning and a real-time 3D object-pose graph to improve robot manipulation under vague instructions.

  11. Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey

    cs.RO 2025-10 conditional novelty 4.0 of 10

    A survey that groups VLA efficiency techniques into four categories: model architecture, perception features, action generation, and training/inference strategies.

Reference graph

Works this paper leans on

86 extracted references · 19 canonical work pages · cited by 11 Pith papers

  1. [1]

    O’Neill, A

    Embodiment Collaboration, A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. Xie, A. Brohan, A. Raffin, A. Sharma, A. Yavary, A. Jain, A. Balakrishna, A. Wahid, B. Bur...

  2. [2]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manju- nath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsc...

  3. [3]

    Ghosh, H

    Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y . Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy. In Proceedings of Robotics: Science and Systems, Delft, Netherlands, 2024

  4. [4]

    Doshi, H

    R. Doshi, H. Walke, O. Mees, S. Dasari, and S. Levine. Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation. In Conference on Robot Learning, 2024

  5. [5]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y . Lu, H. Michalewski, I. Mordatch, K. Pe...

  6. [6]

    M. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model. 2024

  7. [7]

    Ebert, Y

    F. Ebert, Y . Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine. Bridge data: Boosting generalization of robotic skills with cross-domain datasets, 2021

  8. [8]

    Walke, K

    H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V . Myers, K. Fang, C. Finn, and S. Levine. Bridgedata v2: A dataset for robot learning at scale. In Conference on Robot Learning (CoRL) , 2023

Show all 86 references
  1. [9]

    Rosete-Beas, O

    E. Rosete-Beas, O. Mees, G. Kalweit, J. Boedecker, and W. Burgard. Latent plans for task agnostic offline reinforcement learning. InProceedings of the 6th Conference on Robot Learning (CoRL), Auckland, New Zealand, 2022

  2. [10]

    Black, N

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky.π0: A vision-la...

  3. [11]

    Bjorck, F

    J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734, 2025

  4. [12]

    Jones, O

    J. Jones, O. Mees, C. Sferrazza, K. Stachowicz, P. Abbeel, and S. Levine. Beyond sight: Finetuning generalist robot policies with heterogeneous sensors via language grounding. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) , Atlanta, USA, 2025

  5. [13]

    H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y . Du, Y . Hong, and C. Gan. 3d-vla: A 3d vision-language-action generative world model. arXiv preprint arXiv:2403.09631, 2024

  6. [14]

    Zawalski, W

    M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine. Robotic control via embodied chain-of-thought reasoning. In Conference on Robot Learning, 2024

  7. [15]

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023

  8. [16]

    Kojima, S

    T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, and Y . Iwasawa. Large language models are zero-shot reasoners, 2023

  9. [17]

    Zheng, Y

    R. Zheng, Y . Liang, S. Huang, J. Gao, H. Daumé III, A. Kolobov, F. Huang, and J. Yang. Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies. arXiv preprint arXiv:2412.10345, 2024

  10. [18]

    R. Xu, J. Zhang, M. Guo, Y . Wen, H. Yang, M. Lin, J. Huang, Z. Li, K. Zhang, L. Wang, et al. A0: An affordance-aware hierarchical model for general robotic manipulation. arXiv preprint arXiv:2504.12636, 2025

  11. [19]

    Borja-Diaz, O

    J. Borja-Diaz, O. Mees, G. Kalweit, L. Hermann, J. Boedecker, and W. Burgard. Affordance learning from play for sample-efficient policy learning. InProceedings of the IEEE International Conference on Robotics and Automation (ICRA) , Philadelphia, USA, 2022

  12. [20]

    B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning, 2023. URL https://arxiv.org/abs/2306.03310

  13. [21]

    Karamcheti, S

    S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh. Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024

  14. [22]

    Beyer, A

    L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alab- dulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Boš...

  15. [23]

    Steiner, A

    A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y . Bitton, A. Gritsenko, M. Min- derer, A. Sherbondy, S. Long, S. Qin, R. Ingle, E. Bugliarello, S. Kazemzadeh, T. Mesnard, I. Alabdulmohsin, L. Beyer, and X. Zhai. Paligemma 2: A family of versatile vlms for transfer,

  16. [24]

    W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

  17. [25]

    A. Szot, B. Mazoure, O. Attia, A. Timofeev, H. Agrawal, D. Hjelm, Z. Gan, Z. Kira, and A. Toshev. From multimodal llms to generalist embodied agents: Methods and lessons. arXiv preprint arXiv:2412.08442, 2024. 11

  18. [26]

    Huang, F

    H. Huang, F. Liu, L. Fu, T. Wu, M. Mukadam, J. Malik, K. Goldberg, and P. Abbeel. Otter: A vision-language-action model with text-aware visual feature extraction. arXiv preprint arXiv:2503.03734, 2025

  19. [27]

    J. Wen, Y . Zhu, J. Li, M. Zhu, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen, Y . Peng, F. Feng, and J. Tang. Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation. arXiv preprint arXiv:2409.12514, 2024

  20. [28]

    J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu, et al. Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model. arXiv preprint arXiv:2503.10631, 2025

  21. [29]

    S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu. Rdt-1b: a diffusion foundation model for bimanual manipulation. arXiv preprint arXiv:2410.07864, 2024

  22. [30]

    Q. Li, Y . Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y . Deng, S. Xu, Y . Zhang, et al. Cogact: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation. arXiv preprint arXiv:2411.19650, 2024

  23. [31]

    Belkhale and D

    S. Belkhale and D. Sadigh. Minivla: A better vla with a smaller footprint, 2024. URL https://ai.stanford.edu/blog/minivla/

  24. [32]

    Pertsch, K

    K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models, 2025. URL https://arxiv.org/abs/2501.09747

  25. [33]

    Abeyruwan, J

    Gemini Robotics Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Arm- strong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, S. Bohez, K. Bousmalis, A. Brohan, T. Buschmann, A. Byravan, S. Cabi, K. Caluwaerts, F. Casarini, O. Chang, J. E. Chen, X. Chen, H.-T....

  26. [34]

    J. Wen, Y . Zhu, J. Li, Z. Tang, C. Shen, and F. Feng. Dexvla: Vision-language model with plug-in diffusion expert for general robot control. arXiv preprint arXiv:2502.05855, 2025

  27. [35]

    C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song. Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots. In Proceedings of Robotics: Science and Systems (RSS) , 2024

  28. [36]

    Etukuru, N

    H. Etukuru, N. Naka, Z. Hu, S. Lee, J. Mehu, A. Edsinger, C. Paxton, S. Chintala, L. Pinto, and N. M. M. Shafiullah. Robot utility models: General policies for zero-shot deployment in new environments. arXiv preprint arXiv:2409.05865, 2024

  29. [37]

    N. M. M. Shafiullah, A. Rai, H. Etukuru, Y . Liu, I. Misra, S. Chintala, and L. Pinto. On bringing robots home. arXiv preprint arXiv:2311.16098, 2023

  30. [38]

    F. Lin, Y . Hu, P. Sheng, C. Wen, J. You, and Y . Gao. Data scaling laws in imitation learning for robotic manipulation. arXiv preprint arXiv:2410.18647, 2024. 12

  31. [39]

    Khazatsky, K

    A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Pa...

  32. [40]

    Sprague, F

    Z. Sprague, F. Yin, J. D. Rodriguez, D. Jiang, M. Wadhwa, P. Singhal, X. Zhao, X. Ye, K. Mahowald, and G. Durrett. To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning. arXiv preprint arXiv:2409.12183, 2024

  33. [41]

    Snell, J

    C. Snell, J. Lee, K. Xu, and A. Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024. URL https://arxiv.org/abs/2408.03314

  34. [42]

    G. Feng, B. Zhang, Y . Gu, H. Ye, D. He, and L. Wang. Towards revealing the mystery behind chain of thought: A theoretical perspective, 2023. URL https://arxiv.org/abs/2305. 15408

  35. [43]

    Merrill and A

    W. Merrill and A. Sabharwal. The expressive power of transformers with chain of thought,

  36. [44]

    Z. Li, H. Liu, D. Zhou, and T. Ma. Chain of thought empowers transformers to solve inherently serial problems, 2024. URL https://arxiv.org/abs/2402.12875

  37. [45]

    URL https://arxiv.org/abs/2310.07923

  38. [46]

    Q. Zhao, Y . Lu, M. J. Kim, Z. Fu, Z. Zhang, Y . Wu, Z. Li, Q. Ma, S. Han, C. Finn, A. Handa, M.-Y . Liu, D. Xiang, G. Wetzstein, and T.-Y . Lin. Cot-vla: Visual chain-of-thought reasoning for vision-language-action models, 2025. URL https://arxiv.org/abs/2503.22020

  39. [47]

    Hwang, R

    J.-J. Hwang, R. Xu, H. Lin, W.-C. Hung, J. Ji, K. Choi, D. Huang, T. He, P. Covington, B. Sapp, Y . Zhou, J. Guo, D. Anguelov, and M. Tan. Emma: End-to-end multimodal model for autonomous driving, 2024. URL https://arxiv.org/abs/2410.23262

  40. [48]

    Sharma, A

    P. Sharma, A. Torralba, and J. Andreas. Skill induction and planning with latent language, 2022

  41. [49]

    Clark, S

    J. Clark, S. Mirchandani, D. Sadigh, and S. Belkhale. Action-free reasoning for policy general- ization, 2025. URL https://arxiv.org/abs/2502.03729

  42. [50]

    Liang, W

    J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng. Code as policies: Language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages 9493–9500. IEEE, 2023

  43. [51]

    Huang, F

    W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y . Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter. Inner monologue: Embodied reasoning through planning with language models, 2022

  44. [52]

    Belkhale, T

    S. Belkhale, T. Ding, T. Xiao, P. Sermanet, Q. Vuong, J. Tompson, Y . Chebotar, D. Dwibedi, and D. Sadigh. Rt-h: Action hierarchies using language, 2024. 13

  45. [53]

    O. Mees, J. Borja-Diaz, and W. Burgard. Grounding language with visual affordances over unstructured data. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK, 2023

  46. [54]

    D. Niu, Y . Sharma, G. Biamby, J. Quenum, Y . Bai, B. Shi, T. Darrell, and R. Herzig. Llarva: Vision-action instruction tuning enhances robot learning, 2024

  47. [55]

    J. Gu, S. Kirmani, P. Wohlhart, Y . Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, P. Sundaresan, P. Xu, H. Su, K. Hausman, C. Finn, Q. Vuong, and T. Xiao. Rt- trajectory: Robotic task generalization via hindsight trajectory sketches, 2023. URL https: //arxi...

  48. [56]

    W. Chen, O. Mees, A. Kumar, and S. Levine. Vision-language models provide promptable representations for reinforcement learning. Transactions on Machine Learning Research, 2025

  49. [57]

    H. Zhou, X. Yao, O. Mees, Y . Meng, T. Xiao, Y . Bisk, J. Oh, E. Johns, M. Shridhar, D. Shah, J. Thomason, K. Huang, J. Chai, Z. Bing, and A. Knoll. Bridging language and action: A survey of language-based robot manipulation. arXiv preprint arXiv:2312.10807, 2024

  50. [58]

    Karamcheti, S

    S. Karamcheti, S. Nair, A. S. Chen, T. Kollar, C. Finn, D. Sadigh, and P. Liang. Language-driven representation learning for robotics, 2023

  51. [59]

    Majumdar, K

    A. Majumdar, K. Yadav, S. Arnaud, Y . J. Ma, C. Chen, S. Silwal, A. Jain, V .-P. Berges, P. Abbeel, J. Malik, D. Batra, Y . Lin, O. Maksymets, A. Rajeswaran, and F. Meier. Where are we in the search for an artificial visual cortex for embodied intelligence?, 2024. URL https://...

  52. [60]

    Driess, F

    D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y . Chebotar, P. Sermanet, D. Duckworth, S. Levine, V . Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence. Palm-e: An emb...

  53. [61]

    B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Driess, P. Florence, D. Sadigh, L. Guibas, and F. Xia. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities, 2024. URL https://arxiv.org/abs/2401.12168

  54. [63]

    Burns, Z

    K. Burns, Z. Witzel, J. I. Hamid, T. Yu, C. Finn, and K. Hausman. What makes pre-trained visual representations successful for robust manipulation?, 2023. URL https://arxiv.org/ abs/2312.12444

  55. [64]

    J. Pfau, W. Merrill, and S. R. Bowman. Let’s think dot by dot: Hidden computation in transformer language models, 2024. URL https://arxiv.org/abs/2404.15758

  56. [65]

    K. Kang, A. Setlur, D. Ghosh, J. Steinhardt, C. Tomlin, S. Levine, and A. Kumar. What do learning dynamics reveal about generalization in llm reasoning?, 2024. URL https: //arxiv.org/abs/2411.07681

  57. [66]

    Gururangan, A

    S. Gururangan, A. Marasovi ´c, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith. Don’t stop pretraining: Adapt language models to domains and tasks, 2020. URL https://arxiv.org/abs/2004.10964

  58. [67]

    Goyal, Z

    S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V . Nagarajan. Think before you speak: Training language models with pause tokens, 2024. URL https://arxiv.org/abs/ 2310.02226

  59. [68]

    Brief, O

    M. Brief, O. Ovadia, G. Shenderovitz, N. B. Yoash, R. Lemberg, and E. Sheetrit. Mixing it up: The cocktail effect of multi-task fine-tuning on llm performance – a case study in finance, 2024. URL https://arxiv.org/abs/2410.01109. 14

  60. [69]

    H. Lang, M. Agrawal, Y . Kim, and D. Sontag. Co-training improves prompt-based learning for large language models, 2022. URL https://arxiv.org/abs/2202.00828

  61. [70]

    J. Gao, S. Belkhale, S. Dasari, A. Balakrishna, D. Shah, and D. Sadigh. A taxonomy for evaluating generalist robot policies, 2025. URL https://arxiv.org/abs/2503.01238

  62. [71]

    A. Mete, H. Xue, A. Wilcox, Y . Chen, and A. Garg. Quest: Self-supervised skill abstractions for learning continuous control, 2024. URL https://arxiv.org/abs/2407.15840

  63. [72]

    Tensorrt-llm

    NVIDIA. Tensorrt-llm. https://github.com/NVIDIA/TensorRT-LLM?tab= readme-ov-file, 2024

  64. [73]

    Y . Ren, S. Guo, W. Bae, and D. J. Sutherland. How to prepare your task head for finetuning,

  65. [74]

    Kumar, A

    A. Kumar, A. Raghunathan, R. Jones, T. Ma, and P. Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution, 2022. URL https://arxiv.org/abs/2202. 10054

  66. [75]

    M. J. Kim, C. Finn, and P. Liang. Fine-tuning vision-language-action models: Optimizing speed and success, 2025. URL https://arxiv.org/abs/2502.19645

  67. [76]

    W. Chen, M. Zawalski, K. Pertsch, O. Mees, C. Finn, and S. Levine. Tensorrt-openvla, 2025. URL https://github.com/rail-berkeley/tensorrt-openvla

  68. [77]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Ho...

  69. [78]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. ...

  70. [79]

    Deitke, C

    M. Deitke, C. Clark, S. Lee, R. Tripathi, Y . Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, J. Lu, T. Anderson, E. Bransom, K. Ehsani, H. Ngo, Y . Chen, A. Patel, M. Yatskar, C. Callison-Burch, A. Head, R. Hendrix, F. Bastani, E. VanderBilt, N. Lambert, Y . ...

  71. [80]

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware, 2023

  72. [81]

    van den Oord, O

    A. van den Oord, O. Vinyals, and K. Kavukcuoglu. Neural discrete representation learning, 2017

  73. [82]

    X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid loss for language image pre-training, 2023

  74. [83]

    Merrill and A

    W. Merrill and A. Sabharwal. The parallelism tradeoff: Limitations of log-precision transform- ers, 2023. URL https://arxiv.org/abs/2207.00729

  75. [84]

    preserve

    A. Yao. Circuits and local computation. 16 Table 1: Performance of all proposed approaches and baselines on all Libero-90 benchmark variants (Mean± StdErr). Success rates are computed for 50 trials on all 90 tasks (4500 episodes per variant, or 13500 total episodes per policy)...

  76. [85]

    Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, ...

  77. [2023]

    URL https://arxiv.org/abs/2302.05779

  78. [2024]

    URL https://arxiv.org/abs/2412.03555

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.