Pith. sign in

REVIEW 3 major objections 3 minor 5 cited by

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This survey organizes 82 robot imitation-learning methods into one taxonomy.

desk verdict Useful survey structure, but the benchmark tables and reference list have too many verifiable errors to trust as a reference without heavy revision. read the letter →

arxiv 2508.17449 v2 pith:PQRB7N33 submitted 2025-08-24 cs.RO

classification cs.RO
keywords imitationlearningroboticmanipulationvision-language-actionmodelsdiffusionpolicyflowmatchingaffordancepredictionrobotbenchmarksgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to give the field of imitation learning for robotic manipulation a single organizing structure. It identifies 82 representative papers and classifies each by how the robot policy produces actions — diffusion models, flow matching, plain regression, autoregressive generation, classification, or affordance-based planning — then traces how these techniques evolved from 2021 to 2025. It also collects benchmark results across six evaluation suites so that methods can be compared on the same tasks. If the survey's selection and transcriptions are trustworthy, it becomes a reference that newcomers and active researchers can use to locate the main approach families, see what each is good at, and find the open problems that remain.

What carries the argument

The carrying object is the control-strategy taxonomy of Table I: every surveyed method is placed in a grid whose rows distinguish action generation from task planning and whose columns distinguish seven output mechanisms — diffusion, flow matching, Gaussian mixture, naive regression, autoregression, naive classification, and affordance. The survey uses this grid as the spine for its narrative, and pairs it with a chronological timeline (Fig. 2) and six benchmark tables (CALVIN, RLBench, LIBERO, MetaWorld, SimplerEnv, COLOSSEUM) to make the comparison empirical rather than purely taxonomic.

What would settle it

Re-run the survey's literature search for 2021–2025 using an independent citation index, then measure how many of the most-cited imitation-learning manipulation papers appear in the survey's 82-paper selection; if major works are missing or the benchmark numbers in Tables III–VIII disagree with the original papers, the survey's claim to be a comprehensive reference fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that imitation-learning-based robotic manipulation policies can be systematically organized by a two-level taxonomy: first, whether the policy directly generates actions or plans via task-level outputs such as key poses, affordance maps, or heatmaps, and second, which underlying mechanism — diffusion, flow matching, naive regression or classification, autoregression, or affordance prediction — produces the output. On this basis the survey positions each of its 82 representative papers, arranges them on a technology timeline, identifies the dominant trend toward large-scale pretrained vision-language-action models, and compiles benchmark tables from CALVIN, RLBenc

Load-bearing premise

Everything the survey offers depends on its hand-curated selection of 82 representative papers and on the benchmark numbers transcribed from those papers; the paper itself reports 82 analyzed papers in one section and 120 in another, and some table entries cite the same reference for different methods, so the selection and transcriptions are the load-bearing parts.

Editorial extensions

If this is right

  • A newcomer can use the taxonomy to identify the main policy families and the canonical paper in each, shortening the entry path into the field.
  • The benchmark tables provide a current snapshot of which methods lead on long-horizon language-conditioned tasks (CALVIN), keypose prediction (RLBench), lifelong learning (LIBERO), and robustness to perturbation (COLOSSEUM).
  • The timeline documents a shift from single-task diffusion policies toward generalist, large-scale vision-language-action policies, suggesting that scaling data and model size is the field's main current trajectory.
  • The survey's open-challenge list — generalization, embodiment diversity, data efficiency, expert-data dependence, and benchmark standardization — defines where near-term research efforts are most needed.
  • Pretraining on actionless video is highlighted as a way to reduce reliance on expensive action-labeled demonstrations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A reader could test the taxonomy's completeness by applying it to the works the survey explicitly excludes — reinforcement-learning policies and grasp-only papers — and checking whether every new method still falls into one of the existing cells.
  • The survey's numerical inconsistencies (82 analyzed papers in Section I.C versus 120 in Section VII.C, and several different methods sharing the same reference number) suggest that any coverage or citation-count claims should be re-verified from the original sources before being repeated.
  • The juxtaposition of benchmark tables implies an implicit ranking; a natural extension would be to publish per-task numbers and seeds so the tables can be updated as new policies appear.
  • The suggestion that bio-inspired structural priors can compensate for scarce data predicts a testable comparison: whether equivariance-, affordance-, or chain-of-thought-based policies reach the same success as data scaling at a fraction of the data.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper surveys imitation learning (IL) for robotic manipulation (RM). It proposes a hierarchical taxonomy based on control strategy (action generation vs. task planning; diffusion, flow matching, regression, autoregression, classification, affordance), provides structured summaries of representative works, a technological timeline, a discussion of pretraining strategies, benchmark comparisons (CALVIN, RLBench, LIBERO, MetaWorld, SimplerEnv, COLOSSEUM), application categories, and open challenges. The authors claim to identify 82 representative papers and to provide a reproducible resource for locating and comparing IL-based manipulation policies.

Significance. If accurate, this survey would be a valuable entry point and reference: the taxonomy is reasonable, the structured per-paper summaries are useful, and the timeline and benchmark aggregation address a real community need. The paper does not perform circular reasoning; it is a literature survey. However, the resource value depends critically on the correctness of its citations and benchmark tables, and the current version contains several concrete errors in exactly those deliverables. These errors undermine the central claim that readers can trust the tables and references as a reliable comparison resource. The strengths—a first systematic IL-for-RM survey, a sensible taxonomy, and structured paper summaries—are real, but they require a careful fact-checking revision.

major comments (3)
  1. [Section III.A] The paragraph on single-task policies cites CLIPort, ACT, and Diffusion Policy all as [48]; reference [48] is Reactive Diffusion Policy. The correct citations are CLIPort [96], ACT [63], and Diffusion Policy [22]. This breaks the paper's stated goal of letting readers trace methods to their sources. The same problem appears in Section III.C, where RT-1 is cited as [103] (which is RT-2) instead of [14]; RT-1 and RT-2 then both appear as [103] in adjacent sentences. Table III also labels two distinct methods, SIE and DeeR, both as [23], when [23] is SuSIE.
  2. [Table VI] The MetaWorld averages in Table VI are arithmetically incorrect for two of six rows. For Lift3D, the reported Avg. of 84.5 does not match the four difficulty scores (93.1, 82.4, 88.0, 28.0), whose mean is 72.9. For DP3, the reported Avg. of 65.3 does not match (85.7, 49.6, 57.0, 18.0), whose mean is 52.6. Since Table VI is one of the paper's principal quantitative comparisons, these errors mislead readers and must be corrected.
  3. [Section I.C vs. Section VII.C] The paper's scope statement is internally inconsistent: Section I.C says 'we identified 82 of the most representative papers,' while Section VII.C concludes 'We analyzed 120 research papers.' The discrepancy is not explained, and Table I appears to contain roughly 80 entries. Additionally, Table II lists 'GeminiRob [102]' twice, with different strengths/limitations and different average monthly citations (7.0 and 2.32); one row appears to describe OmniManip [101], not GeminiRob. This duplication and citation mismatch further impair the survey's reliability.
minor comments (3)
  1. [Section V.B] In the definition of SPL (Eq. (1)), the symbol M is used but not defined; it should be the number of episodes. Also, in Table II, the CARP row contains the typo 'losed-loop feedback' (should be 'closed-loop').
  2. [Section V.C] Table III is introduced as reporting 'success rate' but the columns and metric are 'Avg. Len' and 'Task completed in a row'; the text should clarify the metric and state the source/version of the 'publicly available leaderboard' used, so readers can verify the numbers.
  3. [Section II.A.b] Minor wording issue: 'Reactive diffusion policy (RDP) [48], a novel slow-fast visual-tactile imitation learning algorithm, allowing robots...' is a sentence fragment; consider restructuring.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey compiles external results and does not fit its own inputs into predictions.

full rationale

This manuscript is a literature survey, not a derivation. Its central deliverables—taxonomy, timeline, paper selection, and benchmark tables—are transcriptions or syntheses of externally published methods and results. There is no equation in the paper that maps a fitted parameter onto a 'predicted' outcome, and no method is defined in terms of the survey's own conclusions. The only self-references (e.g., [109], [110], [112], [115], [152]) are either entries in the surveyed set, dataset/tool citations, or a pointer to a position paper; none of them is load-bearing for the taxonomy or the benchmark comparisons. The reader's concern about arithmetic and citation errors (e.g., Table VI averages that do not match their row entries; Table III labeling SIE and DeeR with the same reference [23]; Section III.A citing CLIPort, ACT, and Diffusion Policy all as [48]) is a legitimate data-quality concern, but it is not circularity: the numbers are copied from external benchmarks and are not regenerated by any model or fit in this paper. The 82-paper count in Section I.C versus the 120-paper count in Section VII.C is likewise an internal inconsistency, not a derivation loop. Because the survey is self-contained as a literature review and its benchmark claims are externally sourced, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey's conclusions rest on the representativeness of the paper selection, the accuracy of transcribed benchmark numbers, and the comparability of results across papers; these are domain assumptions, not argued from first principles.

assumptions (3)
  • domain assumption The hand-curated selection of papers (82 in I.C, 120 in VII.C) is representative of the field.
    Section I.C says papers were selected based on citation counts and community attention, but no reproducible protocol is given and the count is inconsistent.
  • domain assumption Benchmark numbers transcribed from original papers are accurate.
    Tables III-VIII report success rates and other metrics from public leaderboards; the many wrong citation keys raise doubt about whether the numbers were verified.
  • domain assumption Evaluation settings are comparable across the compared papers.
    Section V.C averages results from different papers despite differences in seeds, rollout counts, and environment setups (e.g., Table V notes 20 rollouts, 3 seeds for some methods).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges." pith.science (2026). https://pith.science/paper/PQRB7N33

@misc{pith2026250817449,
  author       = {Pith},
  title        = {Pith review of: Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PQRB7N33}},
  note         = {Machine review of arXiv:2508.17449}
}
read the original abstract

Robotic Manipulation (RM) is central to the advancement of autonomous robots, enabling them to interact with and manipulate objects in real-world environments. This survey focuses on RM methodologies that leverage imitation learning, a powerful technique that allows robots to learn complex manipulation skills by mimicking human demonstrations. We identify and analyze the most influential studies in this domain, selected based on community impact and intrinsic quality. For each paper, we provide a structured summary, covering the research purpose, technical implementation, hierarchical classification, input formats, key priors, strengths and limitations, and citation metrics. Additionally, we trace the chronological development of imitation learning techniques within RM policy (RMP), offering a timeline of key technological advancements. Where available, we report benchmark results and perform quantitative evaluations to compare existing methods. By synthesizing these insights, this review provides a comprehensive resource for researchers and practitioners, highlighting both the state of the art and the challenges that lie ahead in the field of robotic manipulation through imitation learning.

Figures

Figures reproduced from arXiv: 2508.17449 by the authors.

Figure 1
Figure 1. Robotic manipulation classification over four perspectives: their purpose, pretraining strategy, input types, and finally their control strategy (DM, FM, [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Timeline of the models explored in this survey. Each model is classified by its control strategy as presented in Tab. I. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

    cs.RO 2026-07 conditional novelty 6.0 of 10

    S2A2 adds microphone-array spatial audio and spectrograms to imitation-learning policies, substantially improving success on manipulation tasks where vision alone cannot identify the target or destination.

  2. PriGo: Test-Time Primitive Guidance to Diffusion and Flow Policies for Adaptive Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A lightweight primitive classifier and differentiable guidance mechanism improve pretrained diffusion and flow manipulation policies by 3–7 points at test time without retraining.

  3. Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation

    cs.RO 2025-12 conditional novelty 6.0 of 10

    An imitation-learning system that segments demonstrations into VLM-labeled atomic skills, aligns them with contrastive learning, and uses keypose prediction to chain skills, outperforming prior baselines in multi-task...

  4. Multi-Omics Analysis for Cancer Subtype Inference via Unrolling Graph Smoothness Priors

    cs.LG 2025-08 unverdicted novelty 6.0 of 10

    GTMancer unrolls multiplex graph smoothness priors with contrastive learning and dual attention to integrate multi-omics data for cancer subtype classification.

  5. Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    DAMI couples MAML with a 3D diffusion policy and a reference-demonstration conditioning module to improve few-shot adaptation on unseen robotic manipulation tasks.

Reference graph

Works this paper leans on

153 extracted references · 21 canonical work pages · cited by 5 Pith papers

  1. [48]

    Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation,

    H. Xue, J. Ren, W. Chen, G. Zhang, F. Yuan, G. Gu, H. Xu, and C. Lu, “Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation,” in ICRA 2025 Workshop, 2025

  2. [23]

    Zero-shot robotic manipulation with pre-trained image- editing diffusion models,

    K. Black, M. Nakamoto, P. Atreya, H. R. Walke, C. Finn, A. Kumar, and S. Levine, “Zero-shot robotic manipulation with pre-trained image- editing diffusion models,” in The Twelfth ICLR , 2024

  3. [96]

    Cliport: What and where pathways for robotic manipulation,

    M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in CoRL, 2022, pp. 894–906

  4. [63]

    Learning fine-grained bimanual manipulation with low-cost hardware,

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn, “Learning fine-grained bimanual manipulation with low-cost hardware,” in ICML Workshop on New Frontiers in Learning, Control, and Dynamical Systems , 2023

  5. [22]

    Diffusion policy: Visuomotor policy learning via ac- tion diffusion,

    C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via ac- tion diffusion,” The International Journal of Robotics Research , p. 02783649241273668, 2023

  6. [103]

    Rt-2: Vision-language-action models transfer web knowledge to robotic control,

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, X. Chen, K. Choro- manski, T. Ding, D. Driess et al., “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” arXiv:2307.15818, 2023

  7. [14]

    Rt-1: Robotics transformer for real-world control at scale,

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan et al. , “Rt-1: Robotics transformer for real-world control at scale,” arXiv:2212.06817, 2022

  8. [102]

    Gemini robotics: Bringing ai into the physical world,

    G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch et al. , “Gemini robotics: Bringing ai into the physical world,” arXiv:2503.20020, 2025

  9. [101]

    Omnimanip: Towards general robotic manipulation via object-centric interaction primitives as spatial constraints,

    M. Pan, J. Zhang, T. Wu, Y . Zhao, W. Gao, and H. Dong, “Omnimanip: Towards general robotic manipulation via object-centric interaction primitives as spatial constraints,” in CVPR, 2025, pp. 17 359–17 369

Show all 153 references
  1. [1]

    Loeb Classical Library

    Aristotle, Parts of Animals , ser. Loeb Classical Library. Cambridge, MA: Harvard University Press, 1937, vol. 323, book IV , 687a7–9

  2. [2]

    Diels and W

    H. Diels and W. Kranz, Die Fragmente der Vorsokratiker , 6th ed. Berlin: Weidmann, 1951, fragment of Anaxagoras B21a, reporting that man is the most intelligent of animals because he has hands

  3. [3]

    A review of robot learn- ing for manipulation: Challenges, representations, and algorithms,

    O. Kroemer, S. Niekum, and G. Konidaris, “A review of robot learn- ing for manipulation: Challenges, representations, and algorithms,” J. Mach. Learn. Res. , vol. 22, no. 30, pp. 1–82, 2021

  4. [4]

    Deep reinforcement learning for the control of robotic manipulation: a focussed mini-review,

    R. Liu, F. Nageotte, P. Zanne, M. de Mathelin, and B. Dresp-Langley, “Deep reinforcement learning for the control of robotic manipulation: a focussed mini-review,” Robotics, vol. 10, no. 1, p. 22, 2021

  5. [5]

    Rob. auton. syst

    M. Suomalainen, Y . Karayiannidis, and V . Kyrki, “Rob. auton. syst.” Rob. Auton. Syst. , vol. 156, p. 104224, 2022

  6. [6]

    A survey on deep reinforcement learning algorithms for robotic manipulation,

    D. Han, B. Mulyana, V . Stankovic, and S. Cheng, “A survey on deep reinforcement learning algorithms for robotic manipulation,” Sensors, vol. 23, no. 7, p. 3762, 2023

  7. [7]

    Deep learning approaches to grasp synthesis: A review,

    R. Newbury, M. Gu, L. Chumbley, A. Mousavian, C. Eppner, J. Leitner, J. Bohg, A. Morales, T. Asfour, D. Kragic et al. , “Deep learning approaches to grasp synthesis: A review,” IEEE Transactions on Robotics, vol. 39, no. 5, pp. 3994–4015, 2023

  8. [8]

    A survey of embodied learning for object-centric robotic manipulation,

    Y . Zheng, L. Yao, Y . Su, Y . Zhang, Y . Wang, S. Zhao, Y . Zhang, and L.-P. Chau, “A survey of embodied learning for object-centric robotic manipulation,” arXiv:2408.11537, 2024

  9. [9]

    A review of embodied grasping,

    J. Sun, P. Mao, L. Kong, and J. Wang, “A review of embodied grasping,” Sensors (Basel, Switzerland) , vol. 25, no. 3, p. 852, 2025

  10. [10]

    Vision- language-action models: Concepts, progress, applications and chal- lenges,

    R. Sapkota, Y . Cao, K. I. Roumeliotis, and M. Karkee, “Vision- language-action models: Concepts, progress, applications and chal- lenges,” arXiv:2505.04769, 2025

  11. [11]

    A survey on vision- language-action models for embodied ai,

    Y . Ma, Z. Song, Y . Zhuang, J. Hao, and I. King, “A survey on vision- language-action models for embodied ai,” arXiv:2405.14093, 2024

  12. [12]

    A survey on vision-language-action models: An action tokenization perspective,

    Y . Zhong, F. Bai, S. Cai, X. Huang, Z. Chen, X. Zhang, Y . Wang, S. Guo, T. Guan, K. N. Lui et al., “A survey on vision-language-action models: An action tokenization perspective,” arXiv:2507.01925, 2025

  13. [13]

    Vision language action models in robotic manipulation: A systematic review,

    M. U. Din, W. Akram, L. S. Saoud, J. Rosell, and I. Hussain, “Vision language action models in robotic manipulation: A systematic review,” arXiv:2507.10672, 2025

  14. [15]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in NeurIPS, vol. 33, 2020, pp. 6840–6851

  15. [16]

    Denoising diffusion implicit mod- els,

    J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit mod- els,” in ICLR, 2021

  16. [17]

    Score-based generative modeling through stochastic differential equations,

    Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in ICLR, 2021

  17. [18]

    Dpm-ot: A new diffusion probabilistic model based on optimal transport,

    Z. Li, S. Li, Z. Wang, N. Lei, Z. Luo, and X. Gu, “Dpm-ot: A new diffusion probabilistic model based on optimal transport,” in ICCV, 2023, pp. 22 624–22 633

  18. [19]

    Flow matching for generative modeling,

    Y . Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” in ICLR, 2023

  19. [20]

    Vector autoregressive models,

    H. L ¨utkepohl, “Vector autoregressive models,” inHandbook of research methods and applications in empirical macroeconomics. Edward Elgar Publishing, 2013, pp. 139–164

  20. [21]

    Transformation autoregressive networks,

    J. Oliva, A. Dubey, M. Zaheer, B. Poczos, R. Salakhutdinov, E. Xing, and J. Schneider, “Transformation autoregressive networks,” in ICML. PMLR, 2018, pp. 3898–3907

  21. [24]

    Chaineddiffuser: Unifying trajectory diffusion and keypose prediction for robotic manipulation,

    Z. Xian, N. Gkanatsios, T. Gervet, T.-W. Ke, and K. Fragkiadaki, “Chaineddiffuser: Unifying trajectory diffusion and keypose prediction for robotic manipulation,” in CoRL. PMLR, 2023, pp. 2323–2339

  22. [25]

    Learning an actionable discrete diffusion policy via large-scale actionless video pre- training,

    H. He, C. Bai, L. Pan, W. Zhang, B. Zhao, and X. Li, “Learning an actionable discrete diffusion policy via large-scale actionless video pre- training,” NeurIPS, vol. 37, pp. 31 124–31 153, 2025

  23. [26]

    3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations,

    Y . Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu, “3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations,” in ICRA 2024 Workshop, 2024

  24. [27]

    3d diffuser actor: Policy diffusion with 3d scene representations,

    T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki, “3d diffuser actor: Policy diffusion with 3d scene representations,” in 8th Annual CoRL , 2024

  25. [28]

    Equivariant diffusion policy,

    D. Wang, S. Hart, D. Surovik, T. Kelestemur, H. Huang, H. Zhao, M. Yeatman, J. Wang, R. Walters, and R. Platt, “Equivariant diffusion policy,” 8th Annual CoRL , 2024

  26. [29]

    Equibot: Sim (3)-equivariant diffusion policy for generalizable and data efficient learning,

    J. Yang, Z. Cao, C. Deng, R. Antonova, S. Song, and J. Bohg, “Equibot: Sim (3)-equivariant diffusion policy for generalizable and data efficient learning,” in 8th Annual CoRL , 2024

  27. [30]

    Multimodal diffusion transformer: Learning versatile behavior from multimodal goals,

    M. Reuss, ¨O. E. Ya ˘gmurlu, F. Wenzel, and R. Lioutikov, “Multimodal diffusion transformer: Learning versatile behavior from multimodal goals,” arXiv:2407.05996, 2024

  28. [31]

    Sparse diffusion policy: A sparse, reusable, and flexible policy for robot learning,

    Y . Wang, Y . Zhang, M. Huo, T. Tian, X. Zhang, Y . Xie, C. Xu, P. Ji, W. Zhan, M. Ding et al., “Sparse diffusion policy: A sparse, reusable, and flexible policy for robot learning,” in 8th Annual CoRL , 2024

  29. [32]

    Skill expansion and composition in parameter space,

    T. Liu, J. Li, Y . Zheng, H. Niu, Y . Lan, X. Xu, and X. Zhan, “Skill expansion and composition in parameter space,” in ICLR, 2025

  30. [33]

    Adamanip: Adaptive articulated object manipulation environments and policy learning,

    Y . Wang, X. Zhang, R. Wu, Y . Li, Y . Shen, M. Wu, Z. He, Y . Wang, and H. Dong, “Adamanip: Adaptive articulated object manipulation environments and policy learning,” in The Thirteenth ICLR , 2025

  31. [34]

    Afforddp: Generalizable diffusion policy with transferable affordance,

    S. Wu, Y . Zhu, Y . Huang, K. Zhu, J. Gu, J. Yu, Y . Shi, and J. Wang, “Afforddp: Generalizable diffusion policy with transferable affordance,” in CVPR, 2025

  32. [35]

    Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation,

    Q. Lv, H. Li, X. Deng, R. Shao, Y . Li, J. Hao, L. Gao, M. Y . Wang, and L. Nie, “Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation,” in CVPR, 2025

  33. [36]

    Behavior robot suite: Streamlining real-world whole-body manipulation for everyday household activities,

    Y . Jiang, R. Zhang, J. Wong, C. Wang, Y . Ze, H. Yin, C. Gokmen, S. Song, J. Wu, and L. Fei-Fei, “Behavior robot suite: Streamlining real-world whole-body manipulation for everyday household activities,” in RSS 2025 Workshop, 2025

  34. [37]

    Fast flow-based visuomotor policies via conditional optimal transport couplings,

    A. Sochopoulos, N. Malkin, N. Tsagkas, J. Moura, M. Gienger, and S. Vijayakumar, “Fast flow-based visuomotor policies via conditional optimal transport couplings,” arXiv:2505.01179, 2025

  35. [38]

    Octo: An open-source generalist robot policy,

    O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu et al., “Octo: An open-source generalist robot policy,” arXiv:2405.12213, 2024

  36. [39]

    Diffusion-vla: Scaling robot foundation models via unified diffusion and autoregression,

    J. Wen, M. Zhu, Y . Zhu, Z. Tang, J. Li, Z. Zhou, C. Li, X. Liu, Y . Peng, C. Shen et al. , “Diffusion-vla: Scaling robot foundation models via unified diffusion and autoregression,” arXiv:2412.03293, 2024

  37. [40]

    Cogact: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation,

    Q. Li, Y . Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei et al. , “Cogact: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation,” arXiv:2411.19650, 2024

  38. [41]

    Chatvla: Unified multimodal understanding and robot control with vision-language-action model,

    Z. Zhou, Y . Zhu, M. Zhu, J. Wen, N. Liu, Z. Xu, W. Meng, R. Cheng, Y . Peng, C. Shen et al. , “Chatvla: Unified multimodal understanding and robot control with vision-language-action model,” arXiv:2502.14420, 2025. 18 IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIG...

  39. [42]

    Rdt-1b: a diffusion foundation model for bimanual manipulation,

    S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu, “Rdt-1b: a diffusion foundation model for bimanual manipulation,” arXiv:2410.07864, 2024

  40. [43]

    Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems,

    Q. Bu, J. Cai, L. Chen, X. Cui, Y . Ding, S. Feng, S. Gao et al., “Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems,” arXiv:2503.06669, 2025

  41. [44]

    Gr00t n1: An open foundation model for generalist humanoid robots,

    J. Bjorck, F. Casta ˜neda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huang et al., “Gr00t n1: An open foundation model for generalist humanoid robots,” arXiv:2503.14734, 2025

  42. [45]

    Dreamgen: Unlocking generalization in robot learning through neural trajectories,

    J. Jang, S. Ye, Z. Lin, J. Xiang, J. Bjorck, Y . Fang, F. Hu, S. Huang, K. Kundalia, Y .-C. Linet al., “Dreamgen: Unlocking generalization in robot learning through neural trajectories,” arXiv:2505.12705, 2025

  43. [46]

    Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model,

    J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu et al. , “Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model,” arXiv:2503.10631, 2025

  44. [47]

    Affordance-based robot manipulation with flow matching,

    F. Zhang and M. Gienger, “Affordance-based robot manipulation with flow matching,” arXiv:2409.01083, 2024

  45. [49]

    Actionflow: Efficient, accurate, and fast policies with spatially symmetric flow matching,

    N. Funk, J. Urain, J. Carvalho, V . Prasad, G. Chalvatzaki, and J. Pe- ters, “Actionflow: Efficient, accurate, and fast policies with spatially symmetric flow matching,” in R: SS workshop: Structural Priors as Inductive Biases for Learning Robot Dynamics , 2024

  46. [50]

    π0: A vision-language-action flow model for general robot control,

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom et al., “π0: A vision-language-action flow model for general robot control,” arXiv:2410.24164, vol. 2, no. 3, p. 5, 2024

  47. [51]

    Graspvla: a grasping foundation model pre- trained on billion-scale synthetic action data,

    S. Deng, M. Yan, S. Wei, H. Ma, Y . Yang, J. Chen, Z. Zhang, T. Yang, X. Zhang, H. Cui et al., “Graspvla: a grasping foundation model pre- trained on billion-scale synthetic action data,” arXiv:2505.03233, 2025

  48. [52]

    Hi robot: Open-ended instruction following with hierarchical vision-language-action models,

    L. X. Shi, B. Ichter, M. Equi, L. Ke, K. Pertsch, Q. Vuong, J. Tan- ner, A. Walling, H. Wang, N. Fusai et al. , “Hi robot: Open-ended instruction following with hierarchical vision-language-action models,” arXiv:2502.19417, 2025

  49. [53]

    Smolvla: A vision-language-action model for affordable and efficient robotics,

    M. Shukor, D. Aubakirova, F. Capuano, P. Kooijmans, S. Palma, A. Zouitine, M. Aractingi, C. Pascal, M. Russi, A. Marafioti et al. , “Smolvla: A vision-language-action model for affordable and efficient robotics,” arXiv:2506.01844, 2025

  50. [54]

    Masked visual pre- training for motor control,

    T. Xiao, I. Radosavovic, T. Darrell, and J. Malik, “Masked visual pre- training for motor control,” arXiv:2203.06173, 2022

  51. [55]

    Unleashing large-scale video generative pre-training for visual robot manipulation,

    H. Wu, Y . Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong, “Unleashing large-scale video generative pre-training for visual robot manipulation,” in The Twelfth ICLR , 2024

  52. [56]

    Robouniview: Visual-language model with unified view representation for robotic manipulation,

    F. Liu, F. Yan, L. Zheng, C. Feng, Y . Huang, and L. Ma, “Robouniview: Visual-language model with unified view representation for robotic manipulation,” arXiv:2406.18977, 2024

  53. [57]

    Lift3d foundation policy: Lifting 2d large-scale pretrained models for robust 3d robotic manipulation,

    Y . Jia, J. Liu, S. Chen, C. Gu, Z. Wang, L. Luo, L. Lee, P. Wang, Z. Wang et al. , “Lift3d foundation policy: Lifting 2d large-scale pretrained models for robust 3d robotic manipulation,” in CVPR, 2025

  54. [58]

    Sam2act: Integrating visual foundation model with a memory architecture for robotic manipulation,

    H. Fang, M. Grotz, W. Pumacay, Y . R. Wang, D. Fox, R. Krishna, and J. Duan, “Sam2act: Integrating visual foundation model with a memory architecture for robotic manipulation,” arXiv:2501.18564, 2025

  55. [59]

    Fine-tuning vision-language-action models: Optimizing speed and success,

    M. J. Kim, C. Finn, and P. Liang, “Fine-tuning vision-language-action models: Optimizing speed and success,” arXiv:2502.19645, 2025

  56. [60]

    A generalist agent,

    S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y . Sulsky, J. Kay, J. T. Springenberg et al., “A generalist agent,” arXiv:2205.06175, 2022

  57. [61]

    Vima: General robot manipulation with multimodal prompts,

    Y . Jiang, A. Gupta, Z. Zhang, G. Wang, Y . Dou, Y . Chen, L. Fei- Fei, A. Anandkumar, Y . Zhu, and L. Fan, “Vima: General robot manipulation with multimodal prompts,” in NeurIPS 2022 Foundation Models for Decision Making Workshop , 2022

  58. [62]

    Palm-e: An embodied multimodal language model,

    D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu et al., “Palm-e: An embodied multimodal language model,” in ICML. PMLR, 2023, pp. 8469–8488

  59. [64]

    Vision-language foundation models as effective robot imitators,

    X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y . Jing, W. Zhang, H. Liu et al. , “Vision-language foundation models as effective robot imitators,” in The Twelfth ICLR , 2024

  60. [65]

    3d-vla: A 3d vision-language-action generative world model,

    H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y . Du, Y . Hong, and C. Gan, “3d-vla: A 3d vision-language-action generative world model,” in ICML. PMLR, 2024, pp. 61 229–61 245

  61. [66]

    Behavior generation with latent actions,

    S. Lee, Y . Wang, H. Etukuru, H. J. Kim, N. M. M. Shafiullah, and L. Pinto, “Behavior generation with latent actions,” in ICML. PMLR, 2024, pp. 26 991–27 008

  62. [67]

    Openvla: An open-source vision-language-action model,

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong et al., “Openvla: An open-source vision-language-action model,” in 8th Annual CoRL, 2024

  63. [68]

    Quest: Self- supervised skill abstractions for learning continuous control,

    A. Mete, H. Xue, A. Wilcox, Y . Chen, and A. Garg, “Quest: Self- supervised skill abstractions for learning continuous control,” NeurIPS, vol. 37, pp. 4062–4089, 2025

  64. [69]

    Au- toregressive action sequence learning for robotic manipulation,

    X. Zhang, Y . Liu, H. Chang, L. Schramm, and A. Boularias, “Au- toregressive action sequence learning for robotic manipulation,” Rob. Autom. Lett., 2025

  65. [70]

    Carp: Visuomotor policy learning via coarse-to-fine autoregressive prediction,

    Z. Gong, P. Ding, S. Lyu, S. Huang, M. Sun, W. Zhao, Z. Fan, and D. Wang, “Carp: Visuomotor policy learning via coarse-to-fine autoregressive prediction,” arXiv:2412.06782, 2024

  66. [71]

    Tracevla: Visual trace prompting en- hances spatial-temporal awareness for generalist robotic policies,

    R. Zheng, Y . Liang, S. Huang, J. Gao, H. Daum ´e III, A. Kolobov, F. Huang, and J. Yang, “Tracevla: Visual trace prompting en- hances spatial-temporal awareness for generalist robotic policies,” arXiv:2412.10345, 2024

  67. [72]

    Towards generalist robot policies: What matters in building vision-language-action models,

    X. Li, P. Li, M. Liu, D. Wang, J. Liu, B. Kang, X. Ma, T. Kong, H. Zhang, and H. Liu, “Towards generalist robot policies: What matters in building vision-language-action models,” arXiv:2412.14058, 2024

  68. [73]

    Fast: Efficient action tokenization for vision-language-action models,

    K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine, “Fast: Efficient action tokenization for vision-language-action models,” arXiv:2501.09747, 2025

  69. [74]

    Spatialvla: Exploring spatial representations for visual-language-action model,

    D. Qu, H. Song, Q. Chen, Y . Yao, X. Ye, Y . Ding, Z. Wang, J. Gu, B. Zhao, D. Wang et al., “Spatialvla: Exploring spatial representations for visual-language-action model,” arXiv:2501.15830, 2025

  70. [75]

    Vla-cache: Towards efficient vision-language-action model via adaptive token caching in robotic manipulation,

    S. Xu, Y . Wang, C. Xia, D. Zhu, T. Huang, and C. Xu, “Vla-cache: Towards efficient vision-language-action model via adaptive token caching in robotic manipulation,” arXiv:2502.02175, 2025

  71. [76]

    Hamster: Hierarchical action models for open-world robot manipulation,

    Y . Li, Y . Deng, J. Zhang, J. Jang, M. Memmel, R. Yu, C. R. Garrett, F. Ramos, D. Fox, A. Li et al., “Hamster: Hierarchical action models for open-world robot manipulation,” arXiv:2502.05485, 2025

  72. [77]

    Magma: A foundation model for multimodal ai agents,

    J. Yang, R. Tan, Q. Wu, R. Zheng, B. Peng, Y . Liang, Y . Gu, M. Cai, S. Ye, J. Jang et al., “Magma: A foundation model for multimodal ai agents,” in CVPR, 2025, pp. 14 203–14 214

  73. [78]

    Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,

    Q. Zhao, Y . Lu, M. J. Kim, Z. Fu, Z. Zhang, Y . Wu, Z. Li, Q. Ma, S. Han, C. Finn et al., “Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,” in CVPR, 2025, pp. 1702–1713

  74. [79]

    Univla: Learning to act anywhere with task-centric latent actions,

    Q. Bu, Y . Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li, “Univla: Learning to act anywhere with task-centric latent actions,” arXiv:2505.06111, 2025

  75. [80]

    Latent action pretraining from videos,

    S. Ye, J. Jang, B. Jeon, S. J. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y .-W. Chao, B. Y . Lin et al. , “Latent action pretraining from videos,” in CoRL 2024 Workshop, 2024

  76. [81]

    Robotic control via embodied chain-of-thought reasoning,

    M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine, “Robotic control via embodied chain-of-thought reasoning,” in 8th Annual CoRL, 2024

  77. [82]

    Worldvla: Towards autore- gressive action world model,

    J. Cen, C. Yu, H. Yuan, Y . Jiang, S. Huang, J. Guo, X. Li, Y . Song, H. Luo, F. Wang, D. Zhao, and H. Chen, “Worldvla: Towards autore- gressive action world model,” arXiv:2506.21539, 2025

  78. [83]

    Lotus: Continual imitation learning for robot manipulation through unsupervised skill discovery,

    W. Wan, Y . Zhu, R. Shah, and Y . Zhu, “Lotus: Continual imitation learning for robot manipulation through unsupervised skill discovery,” in ICRA, 2024, pp. 537–544

  79. [84]

    What matters in language conditioned robotic imitation learning over unstructured data,

    O. Mees, L. Hermann, and W. Burgard, “What matters in language conditioned robotic imitation learning over unstructured data,” Rob. Autom. Lett., vol. 7, no. 4, pp. 11 205–11 212, 2022

  80. [85]

    Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models,

    P. Li, Y . Chen, H. Wu, X. Ma, X. Wu, Y . Huang, L. Wang, T. Kong, and T. Tan, “Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models,” arXiv:2506.07961, 2025

  81. [86]

    Se (3)-diffusionfields: Learning smooth cost functions for joint grasp and motion optimization through diffusion,

    J. Urain, N. Funk, J. Peters, and G. Chalvatzaki, “Se (3)-diffusionfields: Learning smooth cost functions for joint grasp and motion optimization through diffusion,” in ICRA. IEEE, 2023, pp. 5923–5930

  82. [87]

    A0: An affordance-aware hierarchical model for general robotic manipulation,

    R. Xu, J. Zhang, M. Guo, Y . Wen, H. Yang, M. Lin, J. Huang, Z. Li, K. Zhang, L. Wang et al., “A0: An affordance-aware hierarchical model for general robotic manipulation,” arXiv:2504.12636, 2025

  83. [88]

    Flow matching imitation learning for multi-support manipulation,

    Q. Rouxel, A. Ferrari, S. Ivaldi, and J.-B. Mouret, “Flow matching imitation learning for multi-support manipulation,” in Humanoids. IEEE, 2024, pp. 528–535

  84. [89]

    Polarnet: 3d point clouds for language-guided robotic manipulation,

    S. Chen, R. G. Pinel, C. Schmid, and I. Laptev, “Polarnet: 3d point clouds for language-guided robotic manipulation,” in CoRL. PMLR, 2023, pp. 1761–1781

  85. [90]

    Instruction-driven history-aware policies for robotic ma- nipulations,

    P.-L. Guhur, S. Chen, R. G. Pinel, M. Tapaswi, I. Laptev, and C. Schmid, “Instruction-driven history-aware policies for robotic ma- nipulations,” in CoRL. PMLR, 2023, pp. 175–187

  86. [91]

    Perceiver-actor: A multi-task transformer for robotic manipulation,

    M. Shridhar, L. Manuelli, and D. Fox, “Perceiver-actor: A multi-task transformer for robotic manipulation,” in CoRL, 2023. LI et al.: ROBOTIC MANIPULATION VIA IMITATION LEARNING: TAXONOMY , EVOLUTION, BENCHMARK, AND CHALLENGES 19

  87. [92]

    Rvt: Robotic view transformer for 3d object manipulation,

    A. Goyal, J. Xu, Y . Guo, V . Blukis, Y .-W. Chao, and D. Fox, “Rvt: Robotic view transformer for 3d object manipulation,” in CoRL. PMLR, 2023, pp. 694–710

  88. [93]

    Act3d: 3d feature field transformers for multi-task robotic manipulation,

    T. Gervet, Z. Xian, N. Gkanatsios, and K. Fragkiadaki, “Act3d: 3d feature field transformers for multi-task robotic manipulation,” in CoRL. PMLR, 2023, pp. 3949–3965

  89. [94]

    Sam- e: Leveraging visual foundation model with sequence imitation for embodied manipulation,

    J. Zhang, C. Bai, H. He, Z. Wang, B. Zhao, X. Li, and X. Li, “Sam- e: Leveraging visual foundation model with sequence imitation for embodied manipulation,” in ICML. PMLR, 2024, pp. 58 579–58 598

  90. [95]

    Equact: An se (3)- equivariant multi-task transformer for open-loop robotic manipulation,

    X. Zhu, Y . Qi, Y . Zhu, R. Walters, and R. Platt, “Equact: An se (3)- equivariant multi-task transformer for open-loop robotic manipulation,” arXiv:2505.21351, 2025

  91. [97]

    Moka: Open-vocabulary robotic manipulation through mark-based visual prompting,

    F. Liu, K. Fang, P. Abbeel, and S. Levine, “Moka: Open-vocabulary robotic manipulation through mark-based visual prompting,” in ICRA Workshop, 2024

  92. [98]

    Ram: Retrieval-based affordance transfer for generalizable zero-shot robotic manipulation,

    Y . Kuang, J. Ye, H. Geng, J. Mao, C. Deng, L. Guibas, H. Wang, and Y . Wang, “Ram: Retrieval-based affordance transfer for generalizable zero-shot robotic manipulation,” in 8th Annual CoRL , 2024

  93. [99]

    Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation,

    W. Huang, C. Wang, Y . Li, R. Zhang, and L. Fei-Fei, “Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation,” in 2nd CoRL Workshop, 2024

  94. [100]

    Towards generalizable vision- language robotic manipulation: A benchmark and llm-guided 3d pol- icy,

    R. Garcia, S. Chen, and C. Schmid, “Towards generalizable vision- language robotic manipulation: A benchmark and llm-guided 3d pol- icy,” arXiv:2410.01345, 2024

  95. [104]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Loet al., “Segment anything,” in ICCV, 2023, pp. 4015–4026

  96. [105]

    V oxposer: Composable 3d value maps for robotic manipulation with language models,

    W. Huang, C. Wang, R. Zhang, Y . Li, J. Wu, and L. Fei-Fei, “V oxposer: Composable 3d value maps for robotic manipulation with language models,” arXiv:2307.05973, 2023

  97. [106]

    Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,

    A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee et al., “Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,” in ICRA, 2024, pp. 6892–6903

  98. [107]

    Switchvla: Execution-aware task switching for vision- language-action models,

    M. Li, Z. Zhao, Z. Che, F. Liao, K. Wu, Z. Xu, P. Ren, Z. Jin, N. Liu, and J. Tang, “Switchvla: Execution-aware task switching for vision- language-action models,” arXiv:2506.03574, 2025

  99. [108]

    Rlbench: The robot learning benchmark & learning environment,

    S. James, Z. Ma, D. R. Arrojo, and A. J. Davison, “Rlbench: The robot learning benchmark & learning environment,” Rob. Autom. Lett., vol. 5, no. 2, pp. 3019–3026, 2020

  100. [109]

    Object-centric representations improve policy generalization in robot manipulation,

    A. Chapin, B. Machado, E. Dellandrea, and L. Chen, “Object-centric representations improve policy generalization in robot manipulation,” arXiv:2505.11563, 2025

  101. [110]

    Adpro: a test-time adaptive diffusion policy for robot manipulation via manifold and initial noise constraints,

    Z. Li, R. Yang, R. Chen, Z. Luo, and L. Chen, “Adpro: a test-time adaptive diffusion policy for robot manipulation via manifold and initial noise constraints,” arXiv:2508.06266, 2025

  102. [111]

    Llm-based skill diffusion for zero-shot policy adaptation,

    W. K. Kim, Y . Lee, J. Kim, and H. Woo, “Llm-based skill diffusion for zero-shot policy adaptation,” NeurIPS, vol. 37, pp. 6749–6775, 2024

  103. [112]

    panda-gym: Open-source goal-conditioned environments for robotic learning,

    Q. Gallou ´edec, N. Cazin, E. Dellandr ´ea, and L. Chen, “panda-gym: Open-source goal-conditioned environments for robotic learning,” in 4th NeurIPS Workshop, 2021

  104. [113]

    Roboverse: Towards a unified platform, dataset and benchmark for scalable and generalizable robot learning,

    H. Geng, F. Wang, S. Wei, Y . Li, B. Wang, B. An, C. T. Cheng et al., “Roboverse: Towards a unified platform, dataset and benchmark for scalable and generalizable robot learning,” arXiv:2504.18904, 2025

  105. [114]

    Mujoco playground,

    K. Zakka, B. Tabanpour, Q. Liao, M. Haiderbhai, S. Holt, J. Y . Luo, A. Allshire, E. Frey, K. Sreenath, L. A. Kahrs et al. , “Mujoco playground,” arXiv:2502.08844, 2025

  106. [115]

    Jacquard: A large scale dataset for robotic grasp detection,

    A. Depierre, E. Dellandr ´ea, and L. Chen, “Jacquard: A large scale dataset for robotic grasp detection,” in IROS, 2018, pp. 3511–3516

  107. [116]

    Rh20t: A robotic dataset for learning diverse skills in one-shot,

    H.-S. Fang, H. Fang, Z. Tang, J. Liu, J. Wang, H. Zhu, and C. Lu, “Rh20t: A robotic dataset for learning diverse skills in one-shot,” in RSS 2023 Workshop on Learning for Task and Motion Planning , 2023

  108. [117]

    Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipulation tasks,

    O. Mees, L. Hermann et al. , “Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipulation tasks,” Rob. Autom. Lett. , vol. 7, no. 3, pp. 7327–7334, 2022

  109. [118]

    Pybullet, a python module for physics simulation for games, robotics and machine learning,

    E. Coumans and Y . Bai, “Pybullet, a python module for physics simulation for games, robotics and machine learning,” 2016

  110. [119]

    V-rep: A versatile and scalable robot simulation framework,

    E. Rohmer, S. P. Singh, and M. Freese, “V-rep: A versatile and scalable robot simulation framework,” in IROS. IEEE, 2013, pp. 1321–1326

  111. [120]

    Libero: Benchmarking knowledge transfer for lifelong robot learning,

    B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone, “Libero: Benchmarking knowledge transfer for lifelong robot learning,” NeurIPS, vol. 36, pp. 44 776–44 791, 2023

  112. [121]

    Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,

    T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine, “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in CoRL. PMLR, 2020, pp. 1094–1100

  113. [122]

    robosuite: A modular simulation framework and benchmark for robot learning,

    Y . Zhu, J. Wong, A. Mandlekar, R. Mart ´ın-Mart´ın, A. Joshi, S. Nasiri- any, and Y . Zhu, “robosuite: A modular simulation framework and benchmark for robot learning,” arXiv:2009.12293, 2020

  114. [123]

    What matters in learning from offline human demonstrations for robot manipulation,

    A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei et al. , “What matters in learning from offline human demonstrations for robot manipulation,” in 5th Annual CoRL , 2021

  115. [124]

    Maniskill2: A unified benchmark for generalizable manipulation skills,

    J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, Y . Tang, S. Tao, X. Wei, Y . Yao, X. Yuan, P. Xie et al., “Maniskill2: A unified benchmark for generalizable manipulation skills,” in ICLR, 2023

  116. [125]

    Evaluating real-world robot manipulation policies in simulation,

    X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani et al. , “Evaluating real-world robot manipulation policies in simulation,” arXiv:2405.05941, 2024

  117. [126]

    The colosseum: A benchmark for evaluating generalization for robotic manipulation,

    W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox, “The colosseum: A benchmark for evaluating generalization for robotic manipulation,” in RSS Workshop, 2024

  118. [127]

    D4rl: Datasets for deep data-driven reinforcement learning,

    J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine, “D4rl: Datasets for deep data-driven reinforcement learning,” 2020

  119. [128]

    Language conditioned imitation learning over unstructured data,

    C. Lynch and P. Sermanet, “Language conditioned imitation learning over unstructured data,” arXiv:2005.07648, 2020

  120. [129]

    Language-conditioned imitation learning with base skill priors under unstructured data,

    H. Zhou, Z. Bing, X. Yao, X. Su, C. Yang, K. Huang, and A. Knoll, “Language-conditioned imitation learning with base skill priors under unstructured data,” Rob. Autom. Lett. , 2024

  121. [130]

    Closed-loop visuomotor control with generative expectation for robotic manipulation,

    Y . Yang, “Closed-loop visuomotor control with generative expectation for robotic manipulation,” in The Conference on Neural Information Processing Systems, 2024

  122. [131]

    Diffusion transformer policy,

    Z. Hou, T. Zhang, Y . Xiong, H. Pu, C. Zhao, R. Tong, Y . Qiao, J. Dai, and Y . Chen, “Diffusion transformer policy,” arXiv:2410.15959, 2024

  123. [132]

    Ghil-glue: Hierarchical control with filtered subgoal images,

    K. B. Hatch, A. Balakrishna, O. Mees, S. Nair, S. Park, B. Wulfe, M. Itkina, B. Eysenbach, S. Levine et al. , “Ghil-glue: Hierarchical control with filtered subgoal images,” arXiv:2410.20018, 2024

  124. [133]

    Efficient diffusion transformer policies with mixture of expert denoisers for multitask learning,

    M. Reuss, J. Pari, P. Agrawal, and R. Lioutikov, “Efficient diffusion transformer policies with mixture of expert denoisers for multitask learning,” arXiv:2412.12953, 2024

  125. [134]

    Gr-mg: Leveraging partially-annotated data via multi-modal goal-conditioned policy,

    P. Li, H. Wu, Y . Huang, C. Cheang, L. Wang, and T. Kong, “Gr-mg: Leveraging partially-annotated data via multi-modal goal-conditioned policy,” Rob. Autom. Lett. , 2025

  126. [135]

    Predictive inverse dynamics models are scalable learners for robotic manipulation,

    Y . Tian, S. Yang, J. Zeng, P. Wang, D. Lin, H. Dong, and J. Pang, “Predictive inverse dynamics models are scalable learners for robotic manipulation,” in The Thirteenth ICLR , 2025

  127. [136]

    Rvt- 2: Learning precise manipulation from few demonstrations,

    A. Goyal, V . Blukis, J. Xu, Y . Guo, Y .-W. Chao, and D. Fox, “Rvt- 2: Learning precise manipulation from few demonstrations,” in RSS Workshop, 2024

  128. [137]

    Train a multi-task diffusion policy on rlbench-18 in one day with one gpu,

    Y . Hu, P. S., K. Wen, and R. Detry, “Train a multi-task diffusion policy on rlbench-18 in one day with one gpu,” arXiv:2505.09430, 2025

  129. [138]

    Coarse-to- fine q-attention: Efficient learning for visual robotic manipulation via discretisation,

    S. James, K. Wada, T. Laidlow, and A. J. Davison, “Coarse-to- fine q-attention: Efficient learning for visual robotic manipulation via discretisation,” in CVPR, 2022, pp. 13 739–13 748

  130. [139]

    Mail: Improving imi- tation learning with selective state space models,

    X. Jia, Q. Wang, A. Donat, B. Xing, G. Li, H. Zhou, O. Celik, D. Blessing, R. Lioutikov, and G. Neumann, “Mail: Improving imi- tation learning with selective state space models,” in CoRL, 2024

  131. [140]

    Tinyllava: A framework of small-scale large multimodal models,

    B. Zhou, Y . Hu, X. Weng, J. Jia, J. Luo, X. Liu, J. Wu, and L. Huang, “Tinyllava: A framework of small-scale large multimodal models,” arXiv:2402.14289, 2024

  132. [141]

    Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation,

    Z. Qi, W. Zhang, Y . Ding, R. Dong, X. Yu, J. Li, L. Xu, B. Li, X. He, G. Fan et al. , “Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation,” arXiv:2502.13143, 2025

  133. [142]

    Generative image as action models,

    M. Shridhar, Y . L. Lo, and S. James, “Generative image as action models,” arXiv:2407.07875, 2024

  134. [143]

    R3m: A universal visual representation for robot manipulation,

    S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta, “R3m: A universal visual representation for robot manipulation,” in CoRL. PMLR, 2023, pp. 892–909

  135. [144]

    Manual2skill: Learning to read manuals and acquire robotic skills for furniture assembly using vision-language models,

    C. Tie, S. Sun, J. Zhu, Y . Liu, J. Guo, Y . Hu, H. Chen, J. Chen, R. Wu, and L. Shao, “Manual2skill: Learning to read manuals and acquire robotic skills for furniture assembly using vision-language models,” arXiv:2502.10090, 2025

  136. [145]

    Srsa: Skill retrieval and adaptation for robotic assembly tasks,

    Y . Guo, B. Tang, I. Akinola, D. Fox, A. Gupta, and Y . Narang, “Srsa: Skill retrieval and adaptation for robotic assembly tasks,” in The Fifteenth ICLR , 2025. 20 IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE. PREPRINT VERSION

  137. [146]

    Point-level visual affordance guided retrieval and adaptation for cluttered garments manipulation,

    R. Wu, Z. Zhu, Y . Wang, Y . Chen, J. Wang, and H. Dong, “Point-level visual affordance guided retrieval and adaptation for cluttered garments manipulation,” in CVPR, 2025

  138. [147]

    Do as i can, not as i say: Grounding language in robotic affordances,

    M. Ahn, A. Brohan, N. Brown, Y . Chebotar, O. Cortes, B. David, C. Finn, C. Fu et al., “Do as i can, not as i say: Grounding language in robotic affordances,” arXiv:2204.01691, 2022

  139. [148]

    Shake-vla: Vision-language-action model-based system for bimanual robotic manipulations and liquid mixing,

    M. H. Khan, S. Asfaw, D. Iarchuk, M. A. Cabrera, L. Moreno et al., “Shake-vla: Vision-language-action model-based system for bimanual robotic manipulations and liquid mixing,” arXiv:2501.06919, 2025

  140. [149]

    An improved single short detection method for smart vision-based water garbage cleaning robot,

    A. Haldorai, M. Suriya, M. Balakrishnan et al. , “An improved single short detection method for smart vision-based water garbage cleaning robot,” Cognitive Robotics, vol. 4, pp. 19–29, 2024

  141. [150]

    Learning fusion feature representation for garbage image classification model in human–robot interaction,

    X. Li, T. Li, S. Li, B. Tian, J. Ju et al. , “Learning fusion feature representation for garbage image classification model in human–robot interaction,” Infrared Physics & Technology, vol. 128, p. 104457, 2023

  142. [151]

    Machine learning-based garbage detection and 3d spatial localization for intelligent robotic grasp,

    Z. Lv, T. Chen, Z. Cai, and Z. Chen, “Machine learning-based garbage detection and 3d spatial localization for intelligent robotic grasp,” Applied Sciences, vol. 13, no. 18, p. 10018, 2023

  143. [152]

    Foundational Models for Robotics need to be made Bio-Inspired,

    L. Chen and S. M. Nguyen, “Foundational Models for Robotics need to be made Bio-Inspired,” in ARSO, Jul. 2025

  144. [153]

    Waypoint-based reinforcement learning for robot manipulation tasks,

    S. Mehta, S. Habibian, and D. Losey, “Waypoint-based reinforcement learning for robot manipulation tasks,” in IROS, 2024, pp. 541–548. Zezeng Li received a B.S. degree from Beijing Uni- versity of Technology (BJUT) in 2015 and a Ph.D. degree from Dalian University of Technolog...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.