Pith. sign in

REVIEW 2 major objections 5 minor 175 references

Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review

T0 review · 2 major / 5 minor · reviewed 2026-07-10 · grok-4.5

Pith's one-line read Bimanual VLA recipes for smooth multi-arm control also fit drones, and dual-system designs are the practical path to real deployment.

desk verdict Solid cross-domain survey that actually maps bimanual VLA recipes onto aerial systems; transfer claim is useful analogy, not new experiment. read the letter →

arxiv 2607.06706 v1 pith:CLSO7ZPK submitted 2026-07-07 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords Vision–Language–Actionmodelsbimanualmanipulationunmannedaerialroboticsflowmatchingactionchunkingimitationlearningdual-systemarchitecturesworld
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review of 183 papers argues that Vision–Language–Action models, which turn camera images and plain-language commands into motor commands, have become the main learning framework for robot manipulation, with two-arm coordination as the hardest test. The same structural problem appears in unmanned aerial robotics: a drone must coordinate thrust, attitude, and often gripper commands under tight latency and payload limits. The authors show that the recipes already proven for bimanual arms—continuous chunked action generation (especially flow matching and hybrid designs), diverse cross-embodiment co-training, and reinforcement learning from autonomous practice—transfer to aerial systems. They further conclude that the field is converging on dual-system architectures that pair a slower reasoning module with a faster action module, rather than on a single monolithic end-to-end model. The paper maps fourteen open directions that span both domains, from standardized two-arm benchmarks and safety certification to end-to-end drone VLAs and continuous self-improvement pipelines that close the gap between laboratory scores and industrial reliability.

What carries the argument

Action chunking with continuous generative heads (flow matching and hybrids): instead of emitting one discretized action token at a time, the model predicts a short future trajectory of continuous multi-dimensional actions in a single forward pass, amortizing expensive vision–language inference while preserving inter-actuator correlations needed for two arms or for thrust–attitude–gripper coupling.

What would settle it

A controlled experiment that trains the same continuous-chunk flow-matching VLA on matched bimanual and aerial-manipulation tasks (identical visual backbone, chunk horizon, and co-training mix) and measures whether the aerial policy needs systematically different horizons or denoising steps to match bimanual success rates under wind and latency constraints; large, systematic divergence would undermine the claimed transfer.

Watch

Extended reading notes

Core claim

The coordination strategies, training recipes, and action representations developed for bimanual VLAs transfer to unmanned aerial systems. Continuous, chunked action generation—especially flow-matching and hybrid designs—avoids the quantization and latency bottlenecks of earlier approaches and is the converging solution for tightly coordinated control in both domains; dual-system (slow reasoner + fast actor) architectures are the practical path to real-world deployment.

Load-bearing premise

The paper treats the structural analogy between two multi-joint arms and a single underactuated drone (plus optional gripper) as deep enough that the same chunk lengths, flow schedules, and co-training ratios remain near-optimal, yet the transfer is argued mainly by side-by-side literature rather than controlled cross-embodiment experiments.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This review unifies Vision–Language–Action (VLA) models for bimanual manipulation and unmanned aerial robotics. It surveys 183 works (2017–2026) along a seven-dimension taxonomy (architectures, training recipes, action representations, bimanual coordination, UAV navigation/control, language grounding, and cross-cutting concerns). The central claim is that continuous, chunked action generation—especially flow matching and hybrid designs—avoids autoregressive quantization and multi-step diffusion latency, that training strategy (cross-embodiment diversity, co-training, RECAP-style RL) matters as much as architecture, and that dual-system (slow reasoner + fast actor) designs are the practical path to deployment. The authors argue that bimanual coordination strategies, training recipes, and action representations transfer to aerial systems and list fourteen research directions spanning both domains.

Significance. If the synthesis holds, the paper supplies a timely, cross-embodiment map of a fast-moving field that has lacked a joint treatment of bimanual VLAs and aerial VLAs. Strengths include the breadth of coverage (~183 works), explicit comparison tables (Tables 2–6, 8–15), repeated caveats on self-reported industrial numbers and non-comparable success rates, and a concrete fourteen-direction roadmap. The dual-system and continuous-chunking conclusions are well supported by the cited corpus and align with recent industrial practice (Gemini Robotics, GR00T N1, Helix). As a review, the contribution is organizational and transfer-by-analogy rather than new controlled experiments; that is appropriate to the genre and still useful for both communities.

major comments (2)
  1. The load-bearing transfer claim (Abstract; §8–9; Findings 1, 3, 11–13) rests on structural analogy—joint high-dimensional action chunks for coupled actuators under shared visual–language conditioning—illustrated by Flying Hand’s ACT reuse, dual-arm aerial harvesting as leader–follower, and multi-drone joint spaces. The analogy is coherent and repeatedly flagged, but the manuscript does not report (and does not claim) controlled cross-embodiment ablations that hold flow-matching schedule, chunk horizon H, and co-training ratio fixed across a bimanual arm and a quadrotor. For a review this is a scope limitation rather than an internal error; still, the claim would be stronger if §9 or §12.2 stated more explicitly which hyperparameters are expected to transfer unchanged versus which must be re-tuned for underactuation, ≥100 Hz flight, and outdoor dynamics, and if the corresponding open expe
  2. Table 8 and Figure 10 present approximate success rates and trend lines for bimanual tasks (laundry folding, box assembly, etc.) drawn from original publications under varying protocols. The caption of Table 8 correctly warns that values are not directly comparable, and Figure 10 is labeled “approximate.” Nonetheless, the Discussion (§12.1) and Highlights still treat the rise from ~30% (ACT) to >90% (π*_0) as a primary narrative of progress. A short additional paragraph quantifying protocol differences (object sets, success criteria, number of trials) or restricting the figure to methods evaluated on a common subset would keep the narrative from being read as a strict ranking.
minor comments (5)
  1. Notation: Table 1 and §2.1 introduce o_t, ℓ, q_t, A_t, H, and the flow-matching symbols consistently; a few later sections re-use t for both discrete control steps and continuous flow time without restating the distinction already made in §2.3.
  2. Table 4 and Table 8: bold “best in column” is useful but the captions already note non-comparable conditions; consider adding a footnote that bold is only within the subset of methods that report that column.
  3. Industrial numbers in Table 15 and §12.1 (e.g., DYNA-1 99.4%, Covariant 99%+) are correctly labeled self-reported; a single sentence in the Highlights or Abstract reminding readers of that caveat would match the care already taken in the body.
  4. Figure 1 taxonomy and Figure 2 timeline are clear; ensure that every leaf method cited in the figures appears in the reference list with a consistent year (a few 2025–2026 arXiv entries may shift between submission and publication).
  5. Minor typographical consistency: “UA V” vs “UAV”, “π*_0” vs “π*_0.6”, and occasional missing spaces around em-dashes in the PDF.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: literature review synthesizes external corpus without self-definitional predictions or load-bearing self-citation chains.

full rationale

This is a survey of 183 external contributions (2017–2026) that organizes architectures, training recipes, action representations, bimanual coordination, and aerial systems into a taxonomy and fourteen research directions. The central transfer claim (bimanual strategies apply to UAVs) is advanced by juxtaposition of independently published systems (e.g., Flying Hand reusing ACT, dual-arm aerial harvesting as leader–follower, multi-drone joint action spaces) rather than by fitting parameters to data and then “predicting” related quantities, or by defining X in terms of Y. Formalisms (policy πθ, action chunks At, flow-matching loss LFM, bimanual joint action abi_t) are standard definitions imported from the cited literature, not circular constructions. Tables report success rates and latencies taken from original publications under varying conditions; no new quantitative prediction is forced by construction. Author self-citations, if any, are not load-bearing for the transfer thesis or the fourteen directions. As a review the paper is self-contained against the external corpus it surveys; the structural analogy between bimanual and aerial coordination is an organizational claim, not a derivation that reduces to its inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 1 invented entities

A review paper rests on domain conventions rather than free parameters or invented physical entities. The main load-bearing assumptions are definitional (what counts as a VLA) and analogical (bimanual coordination ≈ multi-actuator aerial control). No numerical constants are fitted; the taxonomy itself is an organizing construct, not a postulated particle or force.

assumptions (3)
  • domain assumption A VLA is defined as a single foundation model that maps camera images + language (+ optional proprioception) to actions via a pre-trained VLM backbone plus an action head.
    Stated in Section 2.1 and used throughout to decide inclusion; classical planners and pure PID controllers are thereby excluded.
  • ad hoc to paper Bimanual joint-action spaces and multi-drone or aerial-manipulation action spaces are sufficiently analogous that the same chunking, flow-matching, and hierarchical recipes remain near-optimal.
    Core transfer claim of Sections 8–9 and Finding 11; not independently measured by a controlled cross-embodiment experiment inside the paper.
  • domain assumption Data diversity across embodiments matters more than raw dataset size for generalization.
    Repeatedly asserted from the surveyed literature (OpenVLA, π0, OXE) and treated as established fact for the training-recipe recommendations.
invented entities (1)
  • Seven-dimension taxonomy (architectures, training, actions, bimanual, aerial, language, cross-cutting)
    purpose: Organizes the 183 papers and structures the transfer argument.
    A useful organizing device introduced by the authors; not claimed to be a physical or mathematical object with independent existence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review." pith.science (2026). https://pith.science/paper/CLSO7ZPK

@misc{pith2026260706706,
  author       = {Pith},
  title        = {Pith review of: Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CLSO7ZPK}},
  note         = {Machine review of arXiv:2607.06706}
}
read the original abstract

Vision Language Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundation model, allowing a robot to follow instructions such as fold the towel or fly to the red building directly from camera images. Because VLAs inherit world knowledge from internet-scale pre-training, they have become the dominant framework for learning-based manipulation, with bimanual coordination serving as the most demanding testbed: two arms with 7 degrees of freedom each must move in concert to fold, assemble, and reorient objects. Unmanned aerial robotics faces a structurally similar challenge: a drone must coordinate thrust, attitude, and increasingly gripper commands from visual observations under strict latency and payload constraints. This review covers 183 contributions spanning 2017-2026 and organized along seven dimensions: VLA architectures, training recipes, action representations, bimanual coordination (2022-2026), unmanned aerial vehicle (UAV) navigation and control (2017-2026), language grounding, and cross-cutting concerns including memory and world models. We show that the coordination strategies, training recipes, and action representations developed for bimanual VLAs transfer to unmanned aerial systems and identify fourteen research directions across both domains.

Figures

Figures reproduced from arXiv: 2607.06706 by the authors.

Figure 1
Figure 1. Taxonomy of VLA models for bimanual manipulation and unmanned aerial robotics. This review is organized along five major dimensions: architectural foundations (autoregressive, flow￾based, diffusion-based, hybrid), training recipes (pre-training, post-training, reinforcement learning), action representations (discrete tokenization, continuous generation), bimanual-specific concerns (coordination strategies, task type… view at source ↗
Figure 2
Figure 2. Timeline of key VLA and bimanual manipulation milestones (2022–2026). Colors indicate the architectural family: autoregressive (blue), flow-based (red), diffusion-based (green), hardware platforms (orange), and hybrid/efficient methods (purple). The field has accelerated rapidly, with the majority of VLA contributions appearing in 2024–2026. (a) Autoregressive (RT-2) Image Tokens Language Tokens VLM Backbone LM Head… view at source ↗
Figure 3
Figure 3. Architectural comparison of the four VLA families. (a) Autoregressive VLAs (RT-2, OpenVLA) discretize actions and generate them as language tokens. (b) Flow-based VLAs (π0) use a flow-matching head that iteratively denoises a noise sample conditioned on VLM features. (c) Diffusion VLAs (RDT-1B) use a Diffusion Transformer to denoise action chunks. (d) Hybrid VLAs (HybridVLA) combine autoregressive and flow heads for… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: The ALOHA bimanual teleoperation platform and representative tasks. A human operator controls two follower arms via leader arms for intuitive demonstration collection. ALOHA and its ACT policy established the standard platform for bimanual VLA research. Reprinted with …
Figure 5
Figure 5. Figure 5: Real-world deployment of π0.5 in homes. A hierarchical VLA decomposes high-level instructions into subgoals, with high success rates on household tasks such as table clearing and laundry folding. Reprinted with permission from Ref. [3]. Copyright 2025, Black et al. htt…
Figure 6
Figure 6. Figure 6: Timeline of unmanned aerial robotics milestones for learning-based drone control (2017– 2026). Colors indicate the research area: RL-based control (blue), vision–language navigation (green), aerial manipulation (red), language-guided planning (orange), and simulation p…
Figure 7
Figure 7. Figure 7: An RL-trained quadrotor recovering from an inverted throw at 5 m/s. The policy maps state to motor commands at 7 µs per step, establishing the viability of learned end-to-end drone control. Reprinted with permission from Ref. [131]. Copyright 2017, Hwangbo et al., IEEE…
Figure 8
Figure 8. Figure 8: Flying Hand: a fully actuated hexarotor with a 4-DOF arm performing writing, peg-in-hole, and pick-and-place via ACT, demonstrating that action chunking transfers from manipulation to aerial systems. The numbers 1–4 along each row index successive video frames of the s…
Figure 9
Figure 9. Figure 9: summarizes the complete VLA training and deployment pipeline that ties together the architectural choices (Section 5), training recipes (Section 6), and deployment considerations discussed above. The interplay among these architectural, training, and deployment conside…
Figure 10
Figure 10. Figure 10: Approximate evolution of VLA performance on bimanual manipulation tasks (2023– 2025). Values are approximate trend values synthesized by the authors from reported results across different evaluation setups and task definitions; they illustrate general trends rather th…
Figure 11
Figure 11. Figure 11: Industrial VLA-powered humanoid robot systems. (a, top left) Boston Dynamics Atlas with TRI Large Behavior Model performing warehouse manipulation. Reprinted with permission from Ref. [178]. Copyright 2025, Boston Dynamics and Toyota Research Institute. (b, top right)…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

175 extracted references · 175 canonical work pages

  1. [1]

    RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

    Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. InProceedings of the 7th Conference on Robot Learning (CoRL); PMLR: Atlanta, GA, USA, 6–9 November 2023; Volume 229, pp. 2165–2183

  2. [2]

    π0: A Vision-Language-Action Flow Model for General Robot Control

    Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. π0: A Vision-Language-Action Flow Model for General Robot Control. InProceedings of Robotics: Science and Systems (RSS), Los Angeles, CA, USA, 21–25 June 2025

  3. [3]

    π0.5: A Vision-Language-Action Model with Open-World Generalization

    Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Galliker, M.Y.; et al. π0.5: A Vision-Language-Action Model with Open-World Generalization. InProceedings of the 9th Conference on Robot Learning (CoRL); PMLR: Seoul, Republic of Korea, 27–30 September 2025; Volume 305, pp. 17–40

  4. [4]

    $\pi^{*}_{0.6}$: a VLA That Learns From Experience

    Amin, A.; Aniceto, R.; Balakrishna, A.; Black, K.; Conley, K.; Connors, G.; Darpinian, J.; Dhabalia, K.; DiCarlo, J.; Driess, D.; et al. π∗ 0.6: A VLA That Learns From Experience.arXiv2025, arXiv:2511.14759

  5. [6]

    Octo: An Open-Source Generalist Robot Policy

    Octo Model Team; Ghosh, D.; Walke, H.; Pertsch, K.; Black, K.; Mees, O.; Dasari, S.; Hejna, J.; Kreiman, T.; Xu, C.; et al. Octo: An Open-Source Generalist Robot Policy. InProceedings of Robotics: Science and Systems (RSS), Delft, The Netherlands, 15–19 July 2024. https://doi.org/10.15607/RSS.2024.XX.090

  6. [7]

    TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.IEEE Robot

    Wen, J.; Zhu, Y.; Li, J.; Zhu, M.; Wu, K.; Xu, Z.; Liu, N.; Cheng, R.; Shen, C.; Peng, Y.; et al. TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.IEEE Robot. Autom. Lett. (RA-L)2025,10, 3988–3995

  7. [8]

    FAST: Efficient Action Tokenization for Vision-Language-Action Models

    Pertsch, K.; Stachowicz, K.; Ichter, B.; Driess, D.; Nair, S.; Vuong, Q.; Mees, O.; Finn, C.; Levine, S. FAST: Efficient Action Tokenization for Vision-Language-Action Models. InProceedings of Robotics: Science and Systems (RSS), Los Angeles, CA, USA, 21–25 June 2025. https://doi.org/10.15607/RSS.2025.XXI.012

  8. [9]

    CognitiveDrone: A VLA Model and Evaluation Benchmark for Real-Time Cognitive Task Solving and Reasoning in UAVs

    Lykov, A.; Serpiva, V .; Khan, M.H.; Sautenkov, O.; Myshlyaev, A.; Tadevosyan, G.; Yaqoot, Y.; Tsetserukou, D. Cognitive- Drone: A VLA Model and Evaluation Benchmark for Real-Time Cognitive Task Solving and Reasoning in UAVs.arXiv2025, arXiv:2503.01378

Show all 175 references
  1. [10]

    DroneVLA: VLA Based Aerial Manipulation

    Mehboob, F.; James, M.; Habel, A.; Sam, J.; Altamirano Cabrera, M.; Tsetserukou, D. DroneVLA: VLA Based Aerial Manipulation. arXiv2026, arXiv:2601.13809

  2. [11]

    AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation.arXiv2026, arXiv:2601.21602

    Sun, J.; Tian, B.; Zhang, Q.; Li, C.; Song, Z.; Cui, Z.; Lv, Y.; Tian, Y. AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation.arXiv2026, arXiv:2601.21602

  3. [12]

    Flying Hand: End-Effector-Centric Framework for Versatile Aerial Manipulation Teleoperation and Policy Learning.arXiv2025, arXiv:2504.10334

    He, G.; Guo, X.; Tang, L.; Zhang, Y.; Mousaei, M.; Xu, J.; Geng, J.; Scherer, S.; Shi, G. Flying Hand: End-Effector-Centric Framework for Versatile Aerial Manipulation Teleoperation and Policy Learning.arXiv2025, arXiv:2504.10334

  4. [13]

    Foundation Models in Robotics: Applications, Challenges, and the Future.Int

    Firoozi, R.; Tucker, J.; Tian, S.; Majumdar, A.; Sun, J.; Liu, W.; Zhu, Y.; Song, S.; Kapoor, A.; Hausman, K.; et al. Foundation Models in Robotics: Applications, Challenges, and the Future.Int. J. Robot. Res.2025,44, 701–739

  5. [14]

    Diffusion Models for Robotic Manipulation: A Survey.Front

    Wolf, R.; Shi, Y.; Liu, S.; Rayyes, R. Diffusion Models for Robotic Manipulation: A Survey.Front. Robot. AI2025,12, 1606247

  6. [15]

    A Systematic Review on Cooperative Dual-Arm Manipulators: Modeling, Planning, Control, and Vision Strategies.Int

    Abbas, M.; Narayan, J.; Dwivedy, S.K. A Systematic Review on Cooperative Dual-Arm Manipulators: Modeling, Planning, Control, and Vision Strategies.Int. J. Intell. Robot. Appl.2023,7, 683–707

  7. [16]

    Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

    Zhao, T.Z.; Kumar, V .; Levine, S.; Finn, C. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. InProceedings of Robotics: Science and Systems (RSS), Daegu, Republic of Korea, 10–14 July 2023. https://doi.org/10.15607/RSS.2023.XIX.016

  8. [17]

    Flow Matching for Generative Modeling

    Lipman, Y.; Chen, R.T.Q.; Ben-Hamu, H.; Nickel, M.; Le, M. Flow Matching for Generative Modeling. InProceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023

  9. [18]

    Multirotor Aerial Vehicles: Modeling, Estimation, and Control of Quadrotor.IEEE Robot

    Mahony, R.; Kumar, V .; Corke, P . Multirotor Aerial Vehicles: Modeling, Estimation, and Control of Quadrotor.IEEE Robot. Autom. Mag.2012,19, 20–32

  10. [19]

    Attention Is All You Need

    Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017

  11. [20]

    An Image is Worth 16 × 16 Words: Transformers for Image Recognition at Scale

    Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. InProceedings of the International Conference on Lea...

  12. [21]

    Learning Transferable Visual Models From Natural Language Supervision

    Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P .; Clark, J.; et al. Learning Transferable Visual Models From Natural Language Supervision. InProceedings of the 38th International Conference on Machine Learning (ICML); ...

  13. [22]

    Language Models Are Few-Shot Learners

    Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P .; Neelakantan, A.; Shyam, P .; Sastry, G.; Askell, A.; et al. Language Models Are Few-Shot Learners. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual, 6–12 December 2020

  14. [23]

    Training Language Models to Follow Instructions with Human Feedback

    Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.L.; Mishkin, P .; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training Language Models to Follow Instructions with Human Feedback. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS),...

  15. [24]

    PaLM-E: An Embodied Multimodal Language Model

    Driess, D.; Xia, F.; Sajjadi, M.S.M.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. PaLM-E: An Embodied Multimodal Language Model. InProceedings of the 40th International Conference on Machine Learning (ICML); PMLR: Honolulu, HI, USA, ...

  16. [25]

    PaliGemma: A Versatile 3B VLM for Transfer.arXiv2024, arXiv:2407.07726

    Beyer, L.; Steiner, A.; Pinto, A.S.; Kolesnikov, A.; Wang, X.; Salz, D.; Neumann, M.; Alabdulmohsin, I.; Tschannen, M.; Bugliarello, E.; et al. PaliGemma: A Versatile 3B VLM for Transfer.arXiv2024, arXiv:2407.07726

  17. [26]

    Gemma: Open Models Based on Gemini Research and Technology.arXiv2024, arXiv:2403.08295

    Gemma Team. Gemma: Open Models Based on Gemini Research and Technology.arXiv2024, arXiv:2403.08295

  18. [27]

    LLaMA: Open and Efficient Foundation Language Models.arXiv2023, arXiv:2302.13971

    Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. LLaMA: Open and Efficient Foundation Language Models.arXiv2023, arXiv:2302.13971

  19. [28]

    Visual Instruction Tuning

    Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual Instruction Tuning. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023

  20. [30]

    ALVINN: An Autonomous Land Vehicle in a Neural Network

    Pomerleau, D.A. ALVINN: An Autonomous Land Vehicle in a Neural Network. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), Denver, CO, USA, 27–30 November 1989

  21. [31]

    A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning

    Ross, S.; Gordon, G.J.; Bagnell, D. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. In Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), Fort Lauderdale, FL, USA, 11–13 April 2011

  22. [32]

    Language-Conditioned Imitation Learning for Robot Manipulation Tasks

    Stepputtis, S.; Campbell, J.; Phielipp, M.; Lee, S.; Baral, C.; Ben Amor, H. Language-Conditioned Imitation Learning for Robot Manipulation Tasks. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual, 6–12 December 2020

  23. [33]

    Auto-Encoding Variational Bayes

    Kingma, D.P .; Welling, M. Auto-Encoding Variational Bayes. InProceedings of the International Conference on Learning Representations (ICLR), Banff, AB, Canada, 14–16 April 2014

  24. [34]

    Generative Adversarial Nets

    Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada, 8–13 December 2014

  25. [35]

    Denoising Diffusion Probabilistic Models

    Ho, J.; Jain, A.; Abbeel, P . Denoising Diffusion Probabilistic Models. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual, 6–12 December 2020

  26. [36]

    Score-Based Generative Modeling through Stochastic Differential Equations

    Song, Y.; Sohl-Dickstein, J.; Kingma, D.P .; Kumar, A.; Ermon, S.; Poole, B. Score-Based Generative Modeling through Stochastic Differential Equations. InProceedings of the International Conference on Learning Representations (ICLR), Virtual, 3–7 May 2021

  27. [37]

    High-Resolution Image Synthesis with Latent Diffusion Models

    Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P .; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 19–24 June 2022

  28. [38]

    Decision Transformer: Reinforcement Learning via Sequence Modeling

    Chen, L.; Lu, K.; Rajeswaran, A.; Lee, K.; Grover, A.; Laskin, M.; Abbeel, P .; Srinivas, A.; Mordatch, I. Decision Transformer: Reinforcement Learning via Sequence Modeling. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2021

  29. [39]

    A Generalist Agent.Trans

    Reed, S.; Zolna, K.; Parisotto, E.; Colmenarejo, S.G.; Novikov, A.; Barth-Maron, G.; Giménez, M.; Sulsky, Y.; Kay, J.; Springenberg, J.T.; et al. A Generalist Agent.Trans. Mach. Learn. Res. (TMLR)2022. Available online: https://openreview.net/forum?id=1ikK0 kHjvj (accessed on ...

  30. [40]

    Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.Int

    Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; Song, S. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.Int. J. Robot. Res. (IJRR)2024,44, 1684–1704. https://doi.org/10.1177/02783649241273668

  31. [41]

    Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

    Liu, X.; Gong, C.; Liu, Q. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. InProceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023

  32. [42]

    Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

    Fu, Z.; Zhao, T.Z.; Finn, C. Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. InProceedings of the 8th Conference on Robot Learning (CoRL); PMLR: Munich, Germany, 6–9 November 2024; Volume 270, pp. 4066–4083

  33. [43]

    Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

    Chi, C.; Xu, Z.; Pan, C.; Cousineau, E.; Burchfiel, B.; Feng, S.; Tedrake, R.; Song, S. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. InProceedings of Robotics: Science and Systems (RSS), Delft, The Netherlands, 15–19 July 2024. https...

  34. [44]

    RDT-1B: A Diffusion Foundation Model for Bimanual Manipulation

    Liu, S.; Wu, L.; Li, B.; Tan, H.; Chen, H.; Wang, Z.; Xu, K.; Su, H.; Zhu, J. RDT-1B: A Diffusion Foundation Model for Bimanual Manipulation. InProceedings of the International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025

  35. [45]

    AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles

    Shah, S.; Dey, D.; Lovett, C.; Kapoor, A. AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles. InField and Service Robotics (FSR), Zurich, Switzerland, 12–15 September 2017; Springer Proceedings in Advanced Robotics, Volume 5, pp. 621–635, published 2018

  36. [46]

    Flightmare: A Flexible Quadrotor Simulator

    Song, Y.; Naji, S.; Kaufmann, E.; Loquercio, A.; Scaramuzza, D. Flightmare: A Flexible Quadrotor Simulator. InProceedings of the 4th Conference on Robot Learning (CoRL); PMLR: Virtual, 16–18 November 2020; Volume 155, pp. 1147–1157

  37. [47]

    LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

    Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; Stone, P . LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023

  38. [48]

    Evaluating Real-World Robot Manipulation Policies in Simulation

    Li, X.; Hsu, K.; Gu, J.; Pertsch, K.; Mees, O.; Walke, H.R.; Fu, C.; Lunawat, I.; Sieh, I.; Kirmani, S.; et al. Evaluating Real-World Robot Manipulation Policies in Simulation. InProceedings of the 8th Conference on Robot Learning (CoRL); PMLR: Munich, Germany, 6–9 November 20...

  39. [49]

    Open X-Embodiment: Robotic Learning Datasets and RT-X Models

    Open X-Embodiment Collaboration. Open X-Embodiment: Robotic Learning Datasets and RT-X Models. InProceedings of the IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; pp. 6892–6903

  40. [50]

    DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset

    Khazatsky, A.; Pertsch, K.; Nair, S.; Balakrishna, A.; Dasari, S.; Karamcheti, S.; Nasiriany, S.; Srirama, M.K.; Chen, L.Y.; Ellis, K.; et al. DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset. InProceedings of Robotics: Science and Systems (RSS), Delft, The Netherla...

  41. [52]

    Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

    Ebert, F.; Yang, Y.; Schmeckpeper, K.; Bucher, B.; Georgakis, G.; Daniilidis, K.; Finn, C.; Levine, S. Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets. InProceedings of Robotics: Science and Systems (RSS), New York, NY, USA, 27 June–1 July 2022

  42. [53]

    GigaBrain-0.5M∗: A VLA That Learns From World Model-Based Reinforcement Learning.arXiv2026, arXiv:2602.12099

    Wang, B.; Li, B.; Ni, C.; Huang, G.; Zhao, G.; Li, H.; Li, J.; Lv, J.; Liu, J.; Feng, L.; et al. GigaBrain-0.5M∗: A VLA That Learns From World Model-Based Reinforcement Learning.arXiv2026, arXiv:2602.12099

  43. [54]

    RLBench: The Robot Learning Benchmark and Learning Environment.IEEE Robot

    James, S.; Ma, Z.; Arrojo, D.R.; Davison, A.J. RLBench: The Robot Learning Benchmark and Learning Environment.IEEE Robot. Autom. Lett. (RA-L)2020,5, 3019–3026

  44. [55]

    Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

    Yu, T.; Quillen, D.; He, Z.; Julian, R.; Hausman, K.; Finn, C.; Levine, S. Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning. In Proceedings of the Conference on Robot Learning (CoRL), 2020

  45. [56]

    robosuite: A Modular Simulation Framework and Benchmark for Robot Learning.arXiv2020, arXiv:2009.12293

    Zhu, Y.; Wong, J.; Mandlekar, A.; Martín-Martín, R.; Joshi, A.; Lin, K.; Maddukuri, A.; Nasiriany, S.; Zhu, Y. robosuite: A Modular Simulation Framework and Benchmark for Robot Learning.arXiv2020, arXiv:2009.12293

  46. [57]

    ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills

    Gu, J.; Xiang, F.; Li, X.; Ling, Z.; Liu, X.; Mu, T.; Tang, Y.; Tao, S.; Wei, X.; Yao, Y.; et al. ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills. InProceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023

  47. [58]

    BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.arXiv 2024, arXiv:2403.09227

    Li, C.; Zhang, R.; Wong, J.; Gokmen, C.; Srivastava, S.; Martín-Martín, R.; Wang, C.; Levine, G.; Ai, W.; Martinez, B.; et al. BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.arXiv 2024, arXiv:2403.09227

  48. [59]

    RT-1: Robotics Transformer for Real-World Control at Scale

    Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Dabis, J.; Finn, C.; Gopalakrishnan, K.; Hausman, K.; Herzog, A.; Hsu, J.; et al. RT-1: Robotics Transformer for Real-World Control at Scale. InProceedings of Robotics: Science and Systems (RSS), Daegu, Republic of Korea, 10–1...

  49. [60]

    Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

    Kim, M.J.; Finn, C.; Liang, P . Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success. InProceedings of Robotics: Science and Systems (RSS), Los Angeles, CA, USA, 21–25 June 2025

  50. [61]

    Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

    Wu, H.; Jing, Y.; Cheang, C.; Chen, G.; Xu, J.; Li, X.; Liu, M.; Li, H.; Kong, T. Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation. InProceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024

  51. [62]

    HAMSTER: Hierarchical Action Models for Open-World Robot Manipulation.arXiv2025, arXiv:2502.05485

    Li, Y.; Deng, Y.; Zhang, J.; Jang, J.; Memmel, M.; Yu, R.; Garrett, C.R.; Ramos, F.; Fox, D.; Li, A.; et al. HAMSTER: Hierarchical Action Models for Open-World Robot Manipulation.arXiv2025, arXiv:2502.05485

  52. [63]

    SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

    Qu, D.; Song, H.; Chen, Q.; Yao, Y.; Ye, X.; Ding, Y.; Wang, Z.; Gu, J.; Zhao, B.; Wang, D.; et al. SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model. InProceedings of Robotics: Science and Systems (RSS), Los Angeles, CA, USA, 21–25 June 2025

  53. [64]

    BAKU: An Efficient Transformer for Multi-Task Policy Learning

    Haldar, S.; Peng, Z.; Pinto, L. BAKU: An Efficient Transformer for Multi-Task Policy Learning. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 9–15 December 2024

  54. [65]

    Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics

    Di Palo, N.; Johns, E. Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics. InProceedings of Robotics: Science and Systems (RSS), Delft, The Netherlands, 15–19 July 2024

  55. [66]

    SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning

    Li, H.; Zuo, Y.; Yu, J.; Zhang, Y.; Yang, Z.; Zhang, K.; Zhu, X.; Zhang, Y.; Chen, T.; Cui, G.; et al. SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning. InProceedings of the International Conference on Learning Representations (ICLR), Available online: https://icl...

  56. [67]

    CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.arXiv2024, arXiv:2411.19650

    Li, Q.; Liang, Y.; Wang, Z.; Luo, L.; Chen, X.; Liao, M.; Wei, F.; Deng, Y.; Xu, S.; Zhang, Y.; et al. CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.arXiv2024, arXiv:2411.19650

  57. [68]

    Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation

    Shridhar, M.; Manuelli, L.; Fox, D. Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation. InProceedings of the Conference on Robot Learning (CoRL), Atlanta, GA, USA, 6–9 November 2023

  58. [69]

    RVT: Robotic View Transformer for 3D Object Manipulation

    Goyal, A.; Xu, J.; Guo, Y.; Blukis, V .; Chao, Y.W.; Fox, D. RVT: Robotic View Transformer for 3D Object Manipulation. In Proceedings of the Conference on Robot Learning (CoRL), Atlanta, GA, USA, 6–9 November 2023

  59. [70]

    3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

    Ze, Y.; Zhang, G.; Zhang, K.; Hu, C.; Wang, M.; Xu, H. 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations. InProceedings of Robotics: Science and Systems (RSS), Delft, The Netherlands, 15–19 July 2024

  60. [71]

    Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

    Zhou, C.; Yu, L.; Babu, A.; Tirumala, K.; Yasunaga, M.; Shamis, L.; Kahn, J.; Ma, X.; Zettlemoyer, L.; Levy, O. Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model. InProceedings of the 13th International Conference on Learning Representations (IC...

  61. [72]

    HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.arXiv2025, arXiv:2503.10631

    Liu, J.; Chen, H.; An, P .; Liu, Z.; Zhang, R.; Gu, C.; Li, X.; Guo, Z.; Chen, S.; Liu, M.; et al. HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.arXiv2025, arXiv:2503.10631

  62. [73]

    MiniVLA: A Better VLA with a Smaller Footprint

    Belkhale, S.; Sadigh, D. MiniVLA: A Better VLA with a Smaller Footprint. 2024. Stanford AI Lab Blog. Available online: https://ai.stanford.edu/blog/minivla/ (accessed on 21 May 2026)

  63. [75]

    MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.arXiv2025, arXiv:2508.19236

    Shi, H.; Xie, B.; Liu, Y.; Sun, L.; Liu, F.; Wang, T.; Zhou, E.; Fan, H.; Zhang, X.; Huang, G. MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.arXiv2025, arXiv:2508.19236

  64. [76]

    ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context.arXiv2025, arXiv:2510.04246

    Jang, H.; Yu, S.; Kwon, H.; Jeon, H.; Seo, Y.; Shin, J. ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context.arXiv2025, arXiv:2510.04246

  65. [77]

    GigaBrain-0: A World Model-Powered Vision-Language-Action Model.arXiv2025, arXiv:2510.19430

    Ye, A.; Wang, B.; Ni, C.; Huang, G.; Zhao, G.; Li, H.; Li, J.; Zhu, J.; Feng, L.; Li, P .; et al. GigaBrain-0: A World Model-Powered Vision-Language-Action Model.arXiv2025, arXiv:2510.19430

  66. [78]

    Causal Video Models Are Data-Efficient Robot Policy Learners

    Rhoda AI. Causal Video Models Are Data-Efficient Robot Policy Learners. 2026 Available online: https://www.rhoda.ai/ research/direct-video-action (accessed on 21 May 2026)

  67. [79]

    WorldVLA: Towards Autoregressive Action World Model.arXiv2025, arXiv:2506.21539

    Cen, J.; Yu, C.; Yuan, H.; Jiang, Y.; Huang, S.; Guo, J.; Li, X.; Song, Y.; Luo, H.; Wang, F.; et al. WorldVLA: Towards Autoregressive Action World Model.arXiv2025, arXiv:2506.21539

  68. [80]

    GigaWorld-0: World Models as Data Engine to Empower Embodied AI.arXiv2025, arXiv:2511.19861

    Ye, A.; Wang, B.; Ni, C.; Huang, G.; Zhao, G.; Li, H.; Zhu, J.; Li, K.; Xu, M.; Deng, Q.; et al. GigaWorld-0: World Models as Data Engine to Empower Embodied AI.arXiv2025, arXiv:2511.19861

  69. [81]

    LoRA: Low-Rank Adaptation of Large Language Models

    Hu, E.J.; Shen, Y.; Wallis, P .; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the International Conference on Learning Representations (ICLR), 2022

  70. [82]

    Direct Preference Optimization: Your Language Model is Secretly a Reward Model

    Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C.D.; Finn, C. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023

  71. [83]

    Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better.arXiv2025, arXiv:2505.23705

    Driess, D.; Springenberg, J.T.; Ichter, B.; Yu, L.; Li-Bell, A.; Pertsch, K.; Ren, A.Z.; Walke, H.; Vuong, Q.; Shi, L.X.; et al. Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better.arXiv2025, arXiv:2505.23705

  72. [84]

    Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance.arXiv2025, arXiv:2509.02055

    Zhang, Y.; Wang, C.; Lu, O.; Zhao, Y.; Ge, Y.; Sun, Z.; Li, X.; Zhang, C.; Bai, C.; Li, X. Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance.arXiv2025, arXiv:2509.02055

  73. [85]

    What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

    Mandlekar, A.; Xu, D.; Wong, J.; Nasiriany, S.; Wang, C.; Kulkarni, R.; Fei-Fei, L.; Savarese, S.; Zhu, Y.; Martín-Martín, R. What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. InProceedings of the 5th Conference on Robot Learning (CoRL); PMLR: ...

  74. [86]

    RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking

    Bharadhwaj, H.; Vakil, J.; Sharma, M.; Gupta, A.; Tulsiani, S.; Kumar, V . RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking. InProceedings of the IEEE International Conference on Robotics and Automation (ICRA), Yokoh...

  75. [87]

    QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation

    Kalashnikov, D.; Irpan, A.; Pastor, P .; Ibarz, J.; Herzog, A.; Jang, E.; Quillen, D.; Holly, E.; Kalakrishnan, M.; Vanhoucke, V .; et al. QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation. InProceedings of the Conference on Robot Learning (CoR...

  76. [88]

    Conservative Q-Learning for Offline Reinforcement Learning

    Kumar, A.; Zhou, A.; Tucker, G.; Levine, S. Conservative Q-Learning for Offline Reinforcement Learning. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual, 6–12 December 2020

  77. [89]

    Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.arXiv2019, arXiv:1910.00177

    Peng, X.B.; Kumar, A.; Zhang, G.; Levine, S. Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.arXiv2019, arXiv:1910.00177

  78. [90]

    Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

    Levine, S.; Kumar, A.; Tucker, G.; Fu, J. Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems. arXiv2020, arXiv:2005.01643

  79. [91]

    Proximal Policy Optimization Algorithms.arXiv2017, arXiv:1707.06347

    Schulman, J.; Wolski, F.; Dhariwal, P .; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms.arXiv2017, arXiv:1707.06347

  80. [92]

    SAIL: Faster-than-Demonstration Execution of Imitation Learning Policies

    Ranawaka Arachchige, N.; Chen, Z.; Jung, W.; Shin, W.C.; Bansal, R.; Barroso, P .; He, Y.H.; Lin, Y.C.; Joffe, B.; Kousik, S.; et al. SAIL: Faster-than-Demonstration Execution of Imitation Learning Policies. InProceedings of the 9th Conference on Robot Learning (CoRL); PMLR: S...

  81. [93]

    VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.arXiv2025, arXiv:2505.18719

    Lu, G.; Guo, W.; Zhang, C.; Zhou, Y.; Jiang, H.; Gao, Z.; Tang, Y.; Wang, Z. VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.arXiv2025, arXiv:2505.18719

  82. [94]

    ConRFT: A Reinforced Fine-Tuning Method for VLA Models via Consistency Policy.arXiv2025, arXiv:2502.05450

    Chen, Y.; Tian, S.; Liu, S.; Zhou, Y.; Li, H.; Zhao, D. ConRFT: A Reinforced Fine-Tuning Method for VLA Models via Consistency Policy.arXiv2025, arXiv:2502.05450

  83. [95]

    Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions

    Chebotar, Y.; Vuong, Q.; Hausman, K.; Xia, F.; Lu, Y.; Irpan, A.; Kumar, A.; Yu, T.; Herzog, A.; Pertsch, K.; et al. Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions. InProceedings of the 7th Conference on Robot Learning (CoRL); PMLR: Atlan...

  84. [96]

    Diffusion Policy Policy Optimization

    Ren, A.Z.; Lidard, J.; Ankile, L.L.; Simeonov, A.; Agrawal, P .; Majumdar, A.; Burchfiel, B.; Dai, H.; Simchowitz, M. Diffusion Policy Policy Optimization. InProceedings of the International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025

  85. [97]

    Self-Improving Embodied Foundation Models.arXiv2025, arXiv:2509.15155

    Ghasemipour, S.K.S.; Wahid, A.; Tompson, J.; Sanketi, P .; Mordatch, I. Self-Improving Embodied Foundation Models.arXiv2025, arXiv:2509.15155

  86. [99]

    Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition

    Ha, H.; Florence, P .; Song, S. Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition. InProceedings of the Conference on Robot Learning (CoRL), Atlanta, GA, USA, 6–9 November 2023

  87. [100]

    Universal Actions for Enhanced Embodied Foundation Models

    Zheng, J.; Li, J.; Liu, D.; Zheng, Y.; Wang, Z.; Ou, Z.; Liu, Y.; Liu, J.; Zhang, Y.Q.; Zhan, X. Universal Actions for Enhanced Embodied Foundation Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025

  88. [101]

    Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning

    Ding, Z.; Jin, C. Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning. InProceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024

  89. [102]

    RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning

    Dai, Y.; Lee, J.; Fazeli, N.; Chai, J. RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning. InProceedings of the IEEE International Conference on Robotics and Automation (ICRA), Atlanta, GA, USA, 19–23 May 2025

  90. [103]

    Real-Time Execution of Action Chunking Flow Policies

    Black, K.; Galliker, M.Y.; Levine, S. Real-Time Execution of Action Chunking Flow Policies. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), San Diego, CA, USA, 2–7 December 2025

  91. [104]

    Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling

    Liu, Y.; Hamid, J.I.; Xie, A.; Lee, Y.; Du, M.; Finn, C. Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling. InProceedings of the International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025

  92. [105]

    Training-Time Action Conditioning for Efficient Real-Time Chunking.arXiv2025, arXiv:2512.05964

    Black, K.; Ren, A.Z.; Equi, M.; Levine, S. Training-Time Action Conditioning for Efficient Real-Time Chunking.arXiv2025, arXiv:2512.05964

  93. [106]

    Stabilize to Act: Learning to Coordinate for Bimanual Manipulation

    Grannen, J.; Wu, Y.; Vu, B.; Sadigh, D. Stabilize to Act: Learning to Coordinate for Bimanual Manipulation. InProceedings of the 7th Conference on Robot Learning (CoRL); PMLR: Atlanta, GA, USA, 6–9 November 2023; Volume 229, pp. 563–576

  94. [107]

    Efficient Bimanual Manipulation Using Learned Task Schemas

    Chitnis, R.; Tulsiani, S.; Gupta, S.; Gupta, A. Efficient Bimanual Manipulation Using Learned Task Schemas. InProceedings of the IEEE International Conference on Robotics and Automation (ICRA), Virtual, 31 May–4 June 2020

  95. [108]

    Code as Policies: Language Model Programs for Embodied Control

    Liang, J.; Huang, W.; Xia, F.; Xu, P .; Hausman, K.; Ichter, B.; Florence, P .; Zeng, A. Code as Policies: Language Model Programs for Embodied Control. InProceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, United Kingdom, 29 May–2 June 2023

  96. [109]

    Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

    Ahn, M.; Brohan, A.; Brown, N.; Chebotar, Y.; Cortes, O.; David, B.; Finn, C.; Fu, C.; Gober, K.; Hausman, K.; et al. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. InProceedings of the 6th Conference on Robot Learning (CoRL); PMLR: Auckland, New Zealand...

  97. [110]

    MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

    Wang, C.; Fan, L.; Sun, J.; Zhang, R.; Fei-Fei, L.; Xu, D.; Zhu, Y.; Anandkumar, A. MimicPlay: Long-Horizon Imitation Learning by Watching Human Play. InProceedings of the Conference on Robot Learning (CoRL), Atlanta, GA, USA, 6–9 November 2023

  98. [111]

    PlayFusion: Skill Acquisition via Diffusion from Language-Annotated Play

    Chen, L.; Bahl, S.; Pathak, D. PlayFusion: Skill Acquisition via Diffusion from Language-Annotated Play. InProceedings of the Conference on Robot Learning (CoRL), Atlanta, GA, USA, 6–9 November 2023

  99. [112]

    Learning Universal Policies via Text-Guided Video Generation

    Du, Y.; Yang, M.; Dai, B.; Dai, H.; Nachum, O.; Tenenbaum, J.B.; Schuurmans, D.; Abbeel, P . Learning Universal Policies via Text-Guided Video Generation. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023

  100. [113]

    Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.arXiv2023, arXiv:2311.17842

    Hu, Y.; Lin, F.; Zhang, T.; Yi, L.; Gao, Y. Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.arXiv2023, arXiv:2311.17842

  101. [114]

    CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling

    Li, H.; Yang, S.; Chen, Y.; Chen, X.; Yang, X.; Tian, Y.; Wang, H.; Wang, T.; Lin, D.; Zhao, F.; et al. CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling. InProceedings of the AAAI Conference on Artificial Intelligence, 2026;f...

  102. [115]

    BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames.arXiv2026, arXiv:2602.15010

    Mark, M.S.; Liang, J.; Attarian, M.; Fu, C.; Dwibedi, D.; Shah, D.; Kumar, A. BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames.arXiv2026, arXiv:2602.15010

  103. [116]

    Learning Long-Context Diffusion Policies via Past-Token Prediction

    Torne, M.; Tang, A.; Liu, Y.; Finn, C. Learning Long-Context Diffusion Policies via Past-Token Prediction. InProceedings of the 9th Conference on Robot Learning (CoRL); PMLR: Seoul, Republic of Korea, 27–30 September 2025; Volume 305

  104. [117]

    SAM2Act: Integrating Visual Foundation Model with a Memory Architecture for Robotic Manipulation

    Fang, H.; Grotz, M.; Pumacay, W.; Wang, Y.R.; Fox, D.; Krishna, R.; Duan, J. SAM2Act: Integrating Visual Foundation Model with a Memory Architecture for Robotic Manipulation. InProceedings of the International Conference on Machine Learning (ICML), Vancouver, BC, Canada, 13–19...

  105. [118]

    MemER: Scaling Up Memory for Robot Control via Experience Retrieval.arXiv2025, arXiv:2510.20328

    Sridhar, A.; Pan, J.; Sharma, S.; Finn, C. MemER: Scaling Up Memory for Robot Control via Experience Retrieval.arXiv2025, arXiv:2510.20328

  106. [119]

    CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding

    Wei, Y.L.; Liao, H.; Lin, Y.; Wang, P .; Liang, Z.; Liu, G.; Zheng, W.S. CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026;forthcoming

  107. [120]

    UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole- body Controllers

    Ha, H.; Gao, Y.; Fu, Z.; Tan, J.; Song, S. UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole- body Controllers. InProceedings of the 8th Conference on Robot Learning (CoRL); PMLR: Munich, Germany, 6–9 November 2024; Volume 270

  108. [122]

    AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.arXiv2025, arXiv:2503.06669

    Bu, Q.; Cai, J.; Chen, L.; Cui, X.; Ding, Y.; Feng, S.; Gao, S.; He, X.; Hu, X.; Huang, X.; et al. AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.arXiv2025, arXiv:2503.06669

  109. [123]

    Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

    Blank, N.; Reuss, M.; Rühle, M.; Ya˘ gmurlu, Ö.E.; Wenzel, F.; Mees, O.; Lioutikov, R. Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models. InProceedings of the 8th Conference on Robot Learning (CoRL); PMLR: Munich, Germany, 6–9 November 2024; Volume 27...

  110. [124]

    AerialVLN: Vision-and-Language Navigation for UAVs

    Liu, S.; Zhang, H.; Qi, Y.; Wang, P .; Zhang, Y.; Wu, Q. AerialVLN: Vision-and-Language Navigation for UAVs. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023

  111. [125]

    Navigation with Large Language Models: Semantic Guesswork as a Heuristic for Planning

    Shah, D.; Equi, M.; Osinski, B.; Xia, F.; Ichter, B.; Levine, S. Navigation with Large Language Models: Semantic Guesswork as a Heuristic for Planning. InProceedings of the Conference on Robot Learning (CoRL), Atlanta, GA, USA, 6–9 November 2023

  112. [126]

    UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation

    Sautenkov, O.; Yaqoot, Y.; Lykov, A.; Mustafa, M.A.; Tadevosyan, G.; Akhmetkazy, A.; Altamirano Cabrera, M.; Martynov, M.; Karaf, S.; Tsetserukou, D. UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation. InProceedings of the 2025 ACM/IEEE Internatio...

  113. [127]

    UAV-VLN: End-to-End Vision Language Guided Navigation for UAVs.arXiv2025, arXiv:2504.21432

    Saxena, P .; Raghuvanshi, N.; Goveas, N. UAV-VLN: End-to-End Vision Language Guided Navigation for UAVs.arXiv2025, arXiv:2504.21432

  114. [128]

    OpenFly: A Comprehensive Platform for Aerial Vision-Language Navigation.arXiv2025, arXiv:2502.18041

    Gao, Y.; Li, C.; You, Z.; Liu, J.; Li, Z.; Chen, P .; Chen, Q.; Tang, Z.; Wang, L.; Yang, P .; et al. OpenFly: A Comprehensive Platform for Aerial Vision-Language Navigation.arXiv2025, arXiv:2502.18041

  115. [129]

    CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory

    Zhang, W.; Gao, C.; Yu, S.; Peng, R.; Zhao, B.; Zhang, Q.; Cui, J.; Chen, X.; Li, Y. CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguis...

  116. [130]

    AgriVLN: Vision-and-Language Navigation for Agricultural Robots.arXiv2025, arXiv:2508.07406

    Zhao, X.; Lyu, X.; Li, X. AgriVLN: Vision-and-Language Navigation for Agricultural Robots.arXiv2025, arXiv:2508.07406

  117. [131]

    Control of a Quadrotor with Reinforcement Learning.IEEE Robot

    Hwangbo, J.; Sa, I.; Siegwart, R.; Hutter, M. Control of a Quadrotor with Reinforcement Learning.IEEE Robot. Autom. Lett.2017, 2, 2096–2103

  118. [132]

    Champion-level drone racing using deep reinforcement learning.Nature2023,620, 982–987

    Kaufmann, E.; Bauersfeld, L.; Loquercio, A.; Müller, M.; Koltun, V .; Scaramuzza, D. Champion-level drone racing using deep reinforcement learning.Nature2023,620, 982–987

  119. [133]

    Neural Lander: Stable Drone Landing Control Using Learned Dynamics

    Shi, G.; Shi, X.; O’Connell, M.; Yu, R.; Azizzadenesheli, K.; Anandkumar, A.; Yue, Y.; Chung, S.J. Neural Lander: Stable Drone Landing Control Using Learned Dynamics. InProceedings of the IEEE International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 20...

  120. [134]

    RaceVLA: VLA-Based Racing Drone Navigation with Human-like Behaviour.arXiv2025, arXiv:2503.02572

    Serpiva, V .; Lykov, A.; Myshlyaev, A.; Khan, M.H.; Abdulkarim, A.A.; Sautenkov, O.; Tsetserukou, D. RaceVLA: VLA-Based Racing Drone Navigation with Human-like Behaviour.arXiv2025, arXiv:2503.02572

  121. [135]

    Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight.arXiv2025, arXiv:2501.14377

    Romero, A.; Shenai, A.; Geles, I.; Aljalbout, E.; Scaramuzza, D. Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight.arXiv2025, arXiv:2501.14377

  122. [136]

    Neural-Fly Enables Rapid Learning for Agile Flight in Strong Winds.Sci

    O’Connell, M.; Shi, G.; Shi, X.; Azizzadenesheli, K.; Anandkumar, A.; Yue, Y.; Chung, S.J. Neural-Fly Enables Rapid Learning for Agile Flight in Strong Winds.Sci. Robot.2022,7, eabm6597. https://doi.org/10.1126/scirobotics.abm6597

  123. [137]

    Vision-assisted Avocado Harvesting with Aerial Bimanual Manipulation.arXiv2024, arXiv:2408.09058

    Liu, Z.; Zhou, J.; Mucchiani, C.; Karydis, K. Vision-assisted Avocado Harvesting with Aerial Bimanual Manipulation.arXiv2024, arXiv:2408.09058

  124. [138]

    Agent as Cerebrum, Controller as Cerebellum: Implementing an Embodied LMM-based Agent on Drones.arXiv2023, arXiv:2311.15033

    Zhao, H.; Pan, F.; Ping, H.; Zhou, Y. Agent as Cerebrum, Controller as Cerebellum: Implementing an Embodied LMM-based Agent on Drones.arXiv2023, arXiv:2311.15033

  125. [139]

    TypeFly: Low-Latency Drone Planning with Large Language Models.IEEE Trans

    Chen, G.; Yu, X.; Ling, N.; Zhong, L. TypeFly: Low-Latency Drone Planning with Large Language Models.IEEE Trans. Mob. Comput.2025,24, 9068–9079. https://doi.org/10.1109/TMC.2025.3561282

  126. [140]

    Interactive Language: Talking to Robots in Real Time.IEEE Robot

    Lynch, C.; Wahid, A.; Tompson, J.; Ding, T.; Betker, J.; Baruch, R.; Armstrong, T.; Florence, P . Interactive Language: Talking to Robots in Real Time.IEEE Robot. Autom. Lett. (RA-L)2023,early access

  127. [141]

    Decentralized Control of Quadrotor Swarms with End-to-end Deep Reinforcement Learning

    Batra, S.; Huang, Z.; Petrenko, A.; Kumar, T.; Molchanov, A.; Sukhatme, G.S. Decentralized Control of Quadrotor Swarms with End-to-end Deep Reinforcement Learning. InProceedings of the Conference on Robot Learning (CoRL), Auckland, New Zealand, 14–18 December 2022

  128. [142]

    Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation

    Doshi, R.; Walke, H.R.; Mees, O.; Dasari, S.; Levine, S. Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation. InProceedings of the 8th Conference on Robot Learning (CoRL); PMLR: Munich, Germany, 6–9 November 2024; Volume 270, pp. 496–512

  129. [143]

    RotorS: A Modular Gazebo MAV Simulator Framework

    Furrer, F.; Burri, M.; Achtelik, M.; Siegwart, R. RotorS: A Modular Gazebo MAV Simulator Framework. InRobot Operating System (ROS): The Complete Reference (Volume 1); Koubaa, A., Ed.; Studies in Computational Intelligence, Volume 625; Springer, 2016; pp. 595–625

  130. [144]

    TartanAir: A Dataset to Push the Limits of Visual SLAM

    Wang, W.; Zhu, D.; Wang, X.; Hu, Y.; Qiu, Y.; Wang, C.; Hu, Y.; Kapoor, A.; Scherer, S. TartanAir: A Dataset to Push the Limits of Visual SLAM. InProceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV , USA, 25–29 October 2020

  131. [146]

    BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

    Jang, E.; Irpan, A.; Khansari, M.; Kappler, D.; Ebert, F.; Lynch, C.; Levine, S.; Finn, C. BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning. InProceedings of the Conference on Robot Learning (CoRL), Auckland, New Zealand, 14–18 December 2022

  132. [147]

    VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

    Huang, W.; Wang, C.; Zhang, R.; Li, Y.; Wu, J.; Fei-Fei, L. VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models. InProceedings of the Conference on Robot Learning (CoRL), Atlanta, GA, USA, 6–9 November 2023

  133. [148]

    Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

    Duan, J.; Yuan, W.; Pumacay, W.; Wang, Y.R.; Ehsani, K.; Fox, D.; Krishna, R. Manipulate-Anything: Automating Real-World Robots using Vision-Language Models. InProceedings of the 8th Conference on Robot Learning (CoRL); PMLR: Munich, Germany, 6–9 November 2024; Volume 270, pp....

  134. [149]

    Chain-of-Thought Predictive Control

    Jia, Z.; Thumuluri, V .; Liu, F.; Chen, L.; Huang, Z.; Su, H. Chain-of-Thought Predictive Control. InProceedings of the International Conference on Machine Learning (ICML), Vienna, Austria, 21–27 July 2024

  135. [150]

    Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution.arXiv2026, arXiv:2602.12684

    Cai, R.; Guo, J.; He, X.; Jin, P .; Li, J.; Lin, B.; Liu, F.; Liu, W.; Ma, F.; Ma, K.; et al. Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution.arXiv2026, arXiv:2602.12684

  136. [151]

    RT-H: Action Hierarchies Using Language

    Belkhale, S.; Ding, T.; Xiao, T.; Sermanet, P .; Vuong, Q.; Tompson, J.; Chebotar, Y.; Dwibedi, D.; Sadigh, D. RT-H: Action Hierarchies Using Language. InProceedings of Robotics: Science and Systems (RSS), Delft, The Netherlands, 15–19 July 2024

  137. [152]

    OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

    Liu, P .; Orru, Y.; Vakil, J.; Paxton, C.; Shafiullah, N.M.M.; Pinto, L. OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics. InProceedings of Robotics: Science and Systems (RSS), Delft, The Netherlands, 15–19 July 2024

  138. [153]

    Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments.arXiv2024, arXiv:2409.05865

    Etukuru, H.; Naka, N.; Hu, Z.; Lee, S.; Mehu, J.; Edsinger, A.; Paxton, C.; Chintala, S.; Pinto, L.; Shafiullah, N.M.M. Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments.arXiv2024, arXiv:2409.05865

  139. [154]

    Transporter Networks: Rearranging the Visual World for Robotic Manipulation

    Zeng, A.; Florence, P .; Tompson, J.; Welker, S.; Chien, J.; Attarian, M.; Armstrong, T.; Krasin, I.; Duong, D.; Sindhwani, V .; et al. Transporter Networks: Rearranging the Visual World for Robotic Manipulation. InProceedings of the Conference on Robot Learning (CoRL), Virtua...

  140. [155]

    R3M: A Universal Visual Representation for Robot Manipulation

    Nair, S.; Rajeswaran, A.; Kumar, V .; Finn, C.; Gupta, A. R3M: A Universal Visual Representation for Robot Manipulation. In Proceedings of the Conference on Robot Learning (CoRL), Auckland, New Zealand, 14–18 December 2022

  141. [156]

    Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation

    Nair, S.; Mitchell, E.; Chen, K.; Ichter, B.; Savarese, S.; Finn, C. Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation. InProceedings of the Conference on Robot Learning (CoRL), Auckland, New Zealand, 14–18 December 2022

  142. [157]

    Majumdar, A.; Yadav, K.; Arnaud, S.; Ma, Y.J.; Chen, C.; Silwal, S.; Jain, A.; Berges, V .P .; Abbeel, P .; Malik, J.; et al. Where Are We in the Search for an Artificial Visual Cortex for Embodied Intelligence? InProceedings of the Advances in Neural Information Processing Sy...

  143. [158]

    SPA: 3D Spatial-Awareness Enables Effective Embodied Representation

    Zhu, H.; Yang, H.; Wang, Y.; Yang, J.; Wang, L.; He, T. SPA: 3D Spatial-Awareness Enables Effective Embodied Representation. In Proceedings of the International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025

  144. [159]

    Gen2Act: Human Video Generation in Novel Scenarios Enables Generalizable Robot Manipulation.arXiv2024, arXiv:2409.16283

    Bharadhwaj, H.; Dwibedi, D.; Gupta, A.; Tulsiani, S.; Doersch, C.; Xiao, T.; Shah, D.; Xia, F.; Sadigh, D.; Kirmani, S. Gen2Act: Human Video Generation in Novel Scenarios Enables Generalizable Robot Manipulation.arXiv2024, arXiv:2409.16283

  145. [160]

    Track2Act: Predicting Point Tracks from Internet Videos Enables General- izable Robot Manipulation

    Bharadhwaj, H.; Mottaghi, R.; Gupta, A.; Tulsiani, S. Track2Act: Predicting Point Tracks from Internet Videos Enables General- izable Robot Manipulation. InProceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024

  146. [161]

    Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

    Hu, Y.; Guo, Y.; Wang, P .; Chen, X.; Wang, Y.J.; Zhang, J.; Sreenath, K.; Lu, C.; Chen, J. Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations. InProceedings of the International Conference on Machine Learning (ICML), Vancouver, BC, Canad...

  147. [162]

    ViPRA: Video Prediction for Robot Actions

    Routray, S.; Pan, H.; Jain, U.; Bahl, S.; Pathak, D. ViPRA: Video Prediction for Robot Actions. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS), San Diego, CA, USA, 2–7 December 2025

  148. [163]

    Mimic-Video: Video-Action Models for Generalizable Robot Control Beyond VLAs.arXiv2025, arXiv:2512.15692

    Pai, J.; Achenbach, L.; Montesinos, V .; Forrai, B.; Mees, O.; Nava, E. Mimic-Video: Video-Action Models for Generalizable Robot Control Beyond VLAs.arXiv2025, arXiv:2512.15692

  149. [164]

    Future Optical Flow Prediction Improves Robot Control & Video Generation.arXiv2026, arXiv:2601.10781

    Ranasinghe, K.; Zhou, H.; Fang, Y.; Yang, L.; Xue, L.; Xu, R.; Xiong, C.; Savarese, S.; Ryoo, M.S.; Niebles, J.C. Future Optical Flow Prediction Improves Robot Control & Video Generation.arXiv2026, arXiv:2601.10781

  150. [165]

    V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.arXiv2025, arXiv:2506.09985

    Assran, M.; Bardes, A.; Fan, D.; Garrido, Q.; Howes, R.; Komeili, M.; Muckley, M.; Rizvi, A.; Roberts, C.; Sinha, K.; et al. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.arXiv2025, arXiv:2506.09985

  151. [166]

    UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

    Zhang, J.; Guo, Y.; Hu, Y.; Chen, X.; Zhu, X.; Chen, J. UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent. InProceedings of the 42nd International Conference on Machine Learning (ICML); PMLR: Vancouver, BC, Canada, 13–19 July 2025

  152. [167]

    Cosmos World Foundation Model Platform for Physical AI.arXiv2025, arXiv:2501.03575

    Agarwal, N.; Ali, A.; Bala, M.; Balaji, Y.; Barker, E.; Cai, T.; Chattopadhyay, P .; Chen, Y.; Cui, Y.; Ding, Y.; et al. Cosmos World Foundation Model Platform for Physical AI.arXiv2025, arXiv:2501.03575

  153. [169]

    TidyBot: Personalized Robot Assistance with Large Language Models.Auton

    Wu, J.; Antonova, R.; Kan, A.; Lepert, M.; Zeng, A.; Song, S.; Bohg, J.; Rusinkiewicz, S.; Funkhouser, T. TidyBot: Personalized Robot Assistance with Large Language Models.Auton. Robot.2023,47, 1087–1102. https://doi.org/10.1007/s10514-023-10139-z (accessed on 21 May 2026)

  154. [170]

    ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data

    Liu, Z.; Chi, C.; Cousineau, E.; Kuppuswamy, N.; Burchfiel, B.; Song, S. ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data. InProceedings of the 8th Conference on Robot Learning (CoRL); PMLR: Munich, Germany, 6–9 November 2024; Volume 270

  155. [171]

    Gemini Robotics: Bringing AI into the Physical World.arXiv2025, arXiv:2503.20020

    Google DeepMind Gemini Robotics Team. Gemini Robotics: Bringing AI into the Physical World.arXiv2025, arXiv:2503.20020

  156. [172]

    GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.arXiv2025, arXiv:2503.14734

    Bjorck, J.; Castañeda, F.; Cherniadev, N.; Da, X.; Ding, R.; Fan, L.; Fang, Y.; Fox, D.; Hu, F.; Huang, S.; et al. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.arXiv2025, arXiv:2503.14734

  157. [173]

    Helix: A Vision-Language-Action Model for Generalist Humanoid Control

    Figure AI. Helix: A Vision-Language-Action Model for Generalist Humanoid Control. 2025. Available online: https://www. figure.ai/news/helix (accessed on 21 May 2026)

  158. [174]

    1X World Model

    1X Technologies. 1X World Model. 2025. Available online: https://www.1x.tech/discover/1x-world-model (accessed on 21 May 2026)

  159. [175]

    DYNA-1: The First Commercial-Ready Robot Foundation Model

    Dyna Robotics. DYNA-1: The First Commercial-Ready Robot Foundation Model. 2025. Available online: https://www.dyna.co/ (accessed on 21 May 2026)

  160. [176]

    ACT-1: A Robot Foundation Model Trained on Zero Robot Data

    Sunday Robotics. ACT-1: A Robot Foundation Model Trained on Zero Robot Data. 2025. Available online: https://www.sunday. ai/journal/no-robot-data (accessed on 21 May 2026)

  161. [177]

    Introducing RFM-1: Giving Robots Human-Like Reasoning Capabilities

    Covariant. Introducing RFM-1: Giving Robots Human-Like Reasoning Capabilities. 2024. Available online: https://en.wikipedia. org/wiki/Covariant_(company) (accessed on 21 May 2026)

  162. [178]

    Large Behavior Models and Atlas Find New Footing

    Boston Dynamics.; Toyota Research Institute. Large Behavior Models and Atlas Find New Footing. 2025. Available online: https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/ (accessed on 21 May 2026)

  163. [179]

    Tesla Optimus: A General-Purpose Humanoid Robot, 2025

    Tesla AI. Tesla Optimus: A General-Purpose Humanoid Robot, 2025. Available online: https://en.wikipedia.org/wiki/Optimus_ (robot) (accessed on 21 May 2026)

  164. [180]

    Tevel Aerobotics Technologies. 2024. Flying Autonomous Robots for Fruit Picking. Available online: https://www.tevel-tech.com (accessed on 8 May 2025)

  165. [181]

    Advances in ground robotic technologies for site-specific weed management in precision agriculture: A review.Comput

    Upadhyay, A.; Zhang, Y.; Koparan, C.; Rai, N.; Howatt, K.; Bajwa, S.; Sun, X. Advances in ground robotic technologies for site-specific weed management in precision agriculture: A review.Comput. Electron. Agric.2024,225, 109363

  166. [182]

    Robotic Harvesters for Fruits and Vegetables

    Anand, S.; Sridharan, B.; Kanchana Devi, V .; Haris, M. Robotic Harvesters for Fruits and Vegetables. InAI-Aided Robotic Applications in Agriculture and Farming; Springer: Berlin/Heidelberg, Germany, 2023

  167. [183]

    HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the Wild.arXiv2026, arXiv:2603.05982

    Zhao, Z.; Wang, S.; Miao, Z.; Xiong, Y. HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the Wild.arXiv2026, arXiv:2603.05982. Disclaimer/Publisher’s Note:The statements, opinions and data contained in all publications are solely those of the ...

Pith tools

Reviewed July 10, 2026 · model on record in the stance chip above.