REVIEW 3 major objections 3 minor 5 cited by
Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This survey organizes 82 robot imitation-learning methods into one taxonomy.
desk verdict Useful survey structure, but the benchmark tables and reference list have too many verifiable errors to trust as a reference without heavy revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the control-strategy taxonomy of Table I: every surveyed method is placed in a grid whose rows distinguish action generation from task planning and whose columns distinguish seven output mechanisms — diffusion, flow matching, Gaussian mixture, naive regression, autoregression, naive classification, and affordance. The survey uses this grid as the spine for its narrative, and pairs it with a chronological timeline (Fig. 2) and six benchmark tables (CALVIN, RLBench, LIBERO, MetaWorld, SimplerEnv, COLOSSEUM) to make the comparison empirical rather than purely taxonomic.
What would settle it
Re-run the survey's literature search for 2021–2025 using an independent citation index, then measure how many of the most-cited imitation-learning manipulation papers appear in the survey's 82-paper selection; if major works are missing or the benchmark numbers in Tables III–VIII disagree with the original papers, the survey's claim to be a comprehensive reference fails.
Extended reading notes
Core claim
The paper's central claim is that imitation-learning-based robotic manipulation policies can be systematically organized by a two-level taxonomy: first, whether the policy directly generates actions or plans via task-level outputs such as key poses, affordance maps, or heatmaps, and second, which underlying mechanism — diffusion, flow matching, naive regression or classification, autoregression, or affordance prediction — produces the output. On this basis the survey positions each of its 82 representative papers, arranges them on a technology timeline, identifies the dominant trend toward large-scale pretrained vision-language-action models, and compiles benchmark tables from CALVIN, RLBenc
Load-bearing premise
Everything the survey offers depends on its hand-curated selection of 82 representative papers and on the benchmark numbers transcribed from those papers; the paper itself reports 82 analyzed papers in one section and 120 in another, and some table entries cite the same reference for different methods, so the selection and transcriptions are the load-bearing parts.
Editorial extensions
If this is right
- A newcomer can use the taxonomy to identify the main policy families and the canonical paper in each, shortening the entry path into the field.
- The benchmark tables provide a current snapshot of which methods lead on long-horizon language-conditioned tasks (CALVIN), keypose prediction (RLBench), lifelong learning (LIBERO), and robustness to perturbation (COLOSSEUM).
- The timeline documents a shift from single-task diffusion policies toward generalist, large-scale vision-language-action policies, suggesting that scaling data and model size is the field's main current trajectory.
- The survey's open-challenge list — generalization, embodiment diversity, data efficiency, expert-data dependence, and benchmark standardization — defines where near-term research efforts are most needed.
- Pretraining on actionless video is highlighted as a way to reduce reliance on expensive action-labeled demonstrations.
Reading between the lines
- A reader could test the taxonomy's completeness by applying it to the works the survey explicitly excludes — reinforcement-learning policies and grasp-only papers — and checking whether every new method still falls into one of the existing cells.
- The survey's numerical inconsistencies (82 analyzed papers in Section I.C versus 120 in Section VII.C, and several different methods sharing the same reference number) suggest that any coverage or citation-count claims should be re-verified from the original sources before being repeated.
- The juxtaposition of benchmark tables implies an implicit ranking; a natural extension would be to publish per-task numbers and seeds so the tables can be updated as new policies appear.
- The suggestion that bio-inspired structural priors can compensate for scarce data predicts a testable comparison: whether equivariance-, affordance-, or chain-of-thought-based policies reach the same success as data scaling at a fraction of the data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys imitation learning (IL) for robotic manipulation (RM). It proposes a hierarchical taxonomy based on control strategy (action generation vs. task planning; diffusion, flow matching, regression, autoregression, classification, affordance), provides structured summaries of representative works, a technological timeline, a discussion of pretraining strategies, benchmark comparisons (CALVIN, RLBench, LIBERO, MetaWorld, SimplerEnv, COLOSSEUM), application categories, and open challenges. The authors claim to identify 82 representative papers and to provide a reproducible resource for locating and comparing IL-based manipulation policies.
Significance. If accurate, this survey would be a valuable entry point and reference: the taxonomy is reasonable, the structured per-paper summaries are useful, and the timeline and benchmark aggregation address a real community need. The paper does not perform circular reasoning; it is a literature survey. However, the resource value depends critically on the correctness of its citations and benchmark tables, and the current version contains several concrete errors in exactly those deliverables. These errors undermine the central claim that readers can trust the tables and references as a reliable comparison resource. The strengths—a first systematic IL-for-RM survey, a sensible taxonomy, and structured paper summaries—are real, but they require a careful fact-checking revision.
major comments (3)
- [Section III.A] The paragraph on single-task policies cites CLIPort, ACT, and Diffusion Policy all as [48]; reference [48] is Reactive Diffusion Policy. The correct citations are CLIPort [96], ACT [63], and Diffusion Policy [22]. This breaks the paper's stated goal of letting readers trace methods to their sources. The same problem appears in Section III.C, where RT-1 is cited as [103] (which is RT-2) instead of [14]; RT-1 and RT-2 then both appear as [103] in adjacent sentences. Table III also labels two distinct methods, SIE and DeeR, both as [23], when [23] is SuSIE.
- [Table VI] The MetaWorld averages in Table VI are arithmetically incorrect for two of six rows. For Lift3D, the reported Avg. of 84.5 does not match the four difficulty scores (93.1, 82.4, 88.0, 28.0), whose mean is 72.9. For DP3, the reported Avg. of 65.3 does not match (85.7, 49.6, 57.0, 18.0), whose mean is 52.6. Since Table VI is one of the paper's principal quantitative comparisons, these errors mislead readers and must be corrected.
- [Section I.C vs. Section VII.C] The paper's scope statement is internally inconsistent: Section I.C says 'we identified 82 of the most representative papers,' while Section VII.C concludes 'We analyzed 120 research papers.' The discrepancy is not explained, and Table I appears to contain roughly 80 entries. Additionally, Table II lists 'GeminiRob [102]' twice, with different strengths/limitations and different average monthly citations (7.0 and 2.32); one row appears to describe OmniManip [101], not GeminiRob. This duplication and citation mismatch further impair the survey's reliability.
minor comments (3)
- [Section V.B] In the definition of SPL (Eq. (1)), the symbol M is used but not defined; it should be the number of episodes. Also, in Table II, the CARP row contains the typo 'losed-loop feedback' (should be 'closed-loop').
- [Section V.C] Table III is introduced as reporting 'success rate' but the columns and metric are 'Avg. Len' and 'Task completed in a row'; the text should clarify the metric and state the source/version of the 'publicly available leaderboard' used, so readers can verify the numbers.
- [Section II.A.b] Minor wording issue: 'Reactive diffusion policy (RDP) [48], a novel slow-fast visual-tactile imitation learning algorithm, allowing robots...' is a sentence fragment; consider restructuring.
Circularity Check
No circularity: the survey compiles external results and does not fit its own inputs into predictions.
full rationale
This manuscript is a literature survey, not a derivation. Its central deliverables—taxonomy, timeline, paper selection, and benchmark tables—are transcriptions or syntheses of externally published methods and results. There is no equation in the paper that maps a fitted parameter onto a 'predicted' outcome, and no method is defined in terms of the survey's own conclusions. The only self-references (e.g., [109], [110], [112], [115], [152]) are either entries in the surveyed set, dataset/tool citations, or a pointer to a position paper; none of them is load-bearing for the taxonomy or the benchmark comparisons. The reader's concern about arithmetic and citation errors (e.g., Table VI averages that do not match their row entries; Table III labeling SIE and DeeR with the same reference [23]; Section III.A citing CLIPort, ACT, and Diffusion Policy all as [48]) is a legitimate data-quality concern, but it is not circularity: the numbers are copied from external benchmarks and are not regenerated by any model or fit in this paper. The 82-paper count in Section I.C versus the 120-paper count in Section VII.C is likewise an internal inconsistency, not a derivation loop. Because the survey is self-contained as a literature review and its benchmark claims are externally sourced, the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The hand-curated selection of papers (82 in I.C, 120 in VII.C) is representative of the field.
- domain assumption Benchmark numbers transcribed from original papers are accurate.
- domain assumption Evaluation settings are comparable across the compared papers.
Cite this review
Pith. "Pith review of Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges." pith.science (2026). https://pith.science/paper/PQRB7N33
@misc{pith2026250817449,
author = {Pith},
title = {Pith review of: Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges},
year = {2026},
howpublished = {\url{https://pith.science/paper/PQRB7N33}},
note = {Machine review of arXiv:2508.17449}
}
read the original abstract
Robotic Manipulation (RM) is central to the advancement of autonomous robots, enabling them to interact with and manipulate objects in real-world environments. This survey focuses on RM methodologies that leverage imitation learning, a powerful technique that allows robots to learn complex manipulation skills by mimicking human demonstrations. We identify and analyze the most influential studies in this domain, selected based on community impact and intrinsic quality. For each paper, we provide a structured summary, covering the research purpose, technical implementation, hierarchical classification, input formats, key priors, strengths and limitations, and citation metrics. Additionally, we trace the chronological development of imitation learning techniques within RM policy (RMP), offering a timeline of key technological advancements. Where available, we report benchmark results and perform quantitative evaluations to compare existing methods. By synthesizing these insights, this review provides a comprehensive resource for researchers and practitioners, highlighting both the state of the art and the challenges that lie ahead in the field of robotic manipulation through imitation learning.
Figures
Forward citations
Cited by 5 Pith papers
-
S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information
S2A2 adds microphone-array spatial audio and spectrograms to imitation-learning policies, substantially improving success on manipulation tasks where vision alone cannot identify the target or destination.
-
PriGo: Test-Time Primitive Guidance to Diffusion and Flow Policies for Adaptive Robotic Manipulation
A lightweight primitive classifier and differentiable guidance mechanism improve pretrained diffusion and flow manipulation policies by 3–7 points at test time without retraining.
-
Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
An imitation-learning system that segments demonstrations into VLM-labeled atomic skills, aligns them with contrastive learning, and uses keypose prediction to chain skills, outperforming prior baselines in multi-task...
-
Multi-Omics Analysis for Cancer Subtype Inference via Unrolling Graph Smoothness Priors
GTMancer unrolls multiplex graph smoothness priors with contrastive learning and dual attention to integrate multi-omics data for cancer subtype classification.
-
Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation
DAMI couples MAML with a 3D diffusion policy and a reference-demonstration conditioning module to improve few-shot adaptation on unseen robotic manipulation tasks.
Reference graph
Works this paper leans on
-
[48]
Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation,
H. Xue, J. Ren, W. Chen, G. Zhang, F. Yuan, G. Gu, H. Xu, and C. Lu, “Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation,” in ICRA 2025 Workshop, 2025
2025
-
[23]
Zero-shot robotic manipulation with pre-trained image- editing diffusion models,
K. Black, M. Nakamoto, P. Atreya, H. R. Walke, C. Finn, A. Kumar, and S. Levine, “Zero-shot robotic manipulation with pre-trained image- editing diffusion models,” in The Twelfth ICLR , 2024
2024
-
[96]
Cliport: What and where pathways for robotic manipulation,
M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in CoRL, 2022, pp. 894–906
2022
-
[63]
Learning fine-grained bimanual manipulation with low-cost hardware,
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn, “Learning fine-grained bimanual manipulation with low-cost hardware,” in ICML Workshop on New Frontiers in Learning, Control, and Dynamical Systems , 2023
2023
-
[22]
Diffusion policy: Visuomotor policy learning via ac- tion diffusion,
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via ac- tion diffusion,” The International Journal of Robotics Research , p. 02783649241273668, 2023
2023
-
[103]
Rt-2: Vision-language-action models transfer web knowledge to robotic control,
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, X. Chen, K. Choro- manski, T. Ding, D. Driess et al., “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” arXiv:2307.15818, 2023
arXiv 2023
-
[14]
Rt-1: Robotics transformer for real-world control at scale,
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan et al. , “Rt-1: Robotics transformer for real-world control at scale,” arXiv:2212.06817, 2022
arXiv 2022
-
[102]
Gemini robotics: Bringing ai into the physical world,
G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch et al. , “Gemini robotics: Bringing ai into the physical world,” arXiv:2503.20020, 2025
arXiv 2025
-
[101]
Omnimanip: Towards general robotic manipulation via object-centric interaction primitives as spatial constraints,
M. Pan, J. Zhang, T. Wu, Y . Zhao, W. Gao, and H. Dong, “Omnimanip: Towards general robotic manipulation via object-centric interaction primitives as spatial constraints,” in CVPR, 2025, pp. 17 359–17 369
2025
Show all 153 references
-
[1]
Loeb Classical Library
Aristotle, Parts of Animals , ser. Loeb Classical Library. Cambridge, MA: Harvard University Press, 1937, vol. 323, book IV , 687a7–9
1937
-
[2]
Diels and W
H. Diels and W. Kranz, Die Fragmente der Vorsokratiker , 6th ed. Berlin: Weidmann, 1951, fragment of Anaxagoras B21a, reporting that man is the most intelligent of animals because he has hands
1951
-
[3]
A review of robot learn- ing for manipulation: Challenges, representations, and algorithms,
O. Kroemer, S. Niekum, and G. Konidaris, “A review of robot learn- ing for manipulation: Challenges, representations, and algorithms,” J. Mach. Learn. Res. , vol. 22, no. 30, pp. 1–82, 2021
2021
-
[4]
Deep reinforcement learning for the control of robotic manipulation: a focussed mini-review,
R. Liu, F. Nageotte, P. Zanne, M. de Mathelin, and B. Dresp-Langley, “Deep reinforcement learning for the control of robotic manipulation: a focussed mini-review,” Robotics, vol. 10, no. 1, p. 22, 2021
2021
-
[5]
Rob. auton. syst
M. Suomalainen, Y . Karayiannidis, and V . Kyrki, “Rob. auton. syst.” Rob. Auton. Syst. , vol. 156, p. 104224, 2022
2022
-
[6]
A survey on deep reinforcement learning algorithms for robotic manipulation,
D. Han, B. Mulyana, V . Stankovic, and S. Cheng, “A survey on deep reinforcement learning algorithms for robotic manipulation,” Sensors, vol. 23, no. 7, p. 3762, 2023
2023
-
[7]
Deep learning approaches to grasp synthesis: A review,
R. Newbury, M. Gu, L. Chumbley, A. Mousavian, C. Eppner, J. Leitner, J. Bohg, A. Morales, T. Asfour, D. Kragic et al. , “Deep learning approaches to grasp synthesis: A review,” IEEE Transactions on Robotics, vol. 39, no. 5, pp. 3994–4015, 2023
2023
-
[8]
A survey of embodied learning for object-centric robotic manipulation,
Y . Zheng, L. Yao, Y . Su, Y . Zhang, Y . Wang, S. Zhao, Y . Zhang, and L.-P. Chau, “A survey of embodied learning for object-centric robotic manipulation,” arXiv:2408.11537, 2024
2024 arXiv
-
[9]
A review of embodied grasping,
J. Sun, P. Mao, L. Kong, and J. Wang, “A review of embodied grasping,” Sensors (Basel, Switzerland) , vol. 25, no. 3, p. 852, 2025
2025
-
[10]
Vision- language-action models: Concepts, progress, applications and chal- lenges,
R. Sapkota, Y . Cao, K. I. Roumeliotis, and M. Karkee, “Vision- language-action models: Concepts, progress, applications and chal- lenges,” arXiv:2505.04769, 2025
2025
-
[11]
A survey on vision- language-action models for embodied ai,
Y . Ma, Z. Song, Y . Zhuang, J. Hao, and I. King, “A survey on vision- language-action models for embodied ai,” arXiv:2405.14093, 2024
2024 arXiv
-
[12]
A survey on vision-language-action models: An action tokenization perspective,
Y . Zhong, F. Bai, S. Cai, X. Huang, Z. Chen, X. Zhang, Y . Wang, S. Guo, T. Guan, K. N. Lui et al., “A survey on vision-language-action models: An action tokenization perspective,” arXiv:2507.01925, 2025
2025 arXiv
-
[13]
Vision language action models in robotic manipulation: A systematic review,
M. U. Din, W. Akram, L. S. Saoud, J. Rosell, and I. Hussain, “Vision language action models in robotic manipulation: A systematic review,” arXiv:2507.10672, 2025
2025
-
[15]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in NeurIPS, vol. 33, 2020, pp. 6840–6851
2020
-
[16]
Denoising diffusion implicit mod- els,
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit mod- els,” in ICLR, 2021
2021
-
[17]
Score-based generative modeling through stochastic differential equations,
Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in ICLR, 2021
2021
-
[18]
Dpm-ot: A new diffusion probabilistic model based on optimal transport,
Z. Li, S. Li, Z. Wang, N. Lei, Z. Luo, and X. Gu, “Dpm-ot: A new diffusion probabilistic model based on optimal transport,” in ICCV, 2023, pp. 22 624–22 633
2023
-
[19]
Flow matching for generative modeling,
Y . Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” in ICLR, 2023
2023
-
[20]
Vector autoregressive models,
H. L ¨utkepohl, “Vector autoregressive models,” inHandbook of research methods and applications in empirical macroeconomics. Edward Elgar Publishing, 2013, pp. 139–164
2013
-
[21]
Transformation autoregressive networks,
J. Oliva, A. Dubey, M. Zaheer, B. Poczos, R. Salakhutdinov, E. Xing, and J. Schneider, “Transformation autoregressive networks,” in ICML. PMLR, 2018, pp. 3898–3907
2018
-
[24]
Chaineddiffuser: Unifying trajectory diffusion and keypose prediction for robotic manipulation,
Z. Xian, N. Gkanatsios, T. Gervet, T.-W. Ke, and K. Fragkiadaki, “Chaineddiffuser: Unifying trajectory diffusion and keypose prediction for robotic manipulation,” in CoRL. PMLR, 2023, pp. 2323–2339
2023
-
[25]
Learning an actionable discrete diffusion policy via large-scale actionless video pre- training,
H. He, C. Bai, L. Pan, W. Zhang, B. Zhao, and X. Li, “Learning an actionable discrete diffusion policy via large-scale actionless video pre- training,” NeurIPS, vol. 37, pp. 31 124–31 153, 2025
2025
-
[26]
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations,
Y . Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu, “3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations,” in ICRA 2024 Workshop, 2024
2024
-
[27]
3d diffuser actor: Policy diffusion with 3d scene representations,
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki, “3d diffuser actor: Policy diffusion with 3d scene representations,” in 8th Annual CoRL , 2024
2024
-
[28]
Equivariant diffusion policy,
D. Wang, S. Hart, D. Surovik, T. Kelestemur, H. Huang, H. Zhao, M. Yeatman, J. Wang, R. Walters, and R. Platt, “Equivariant diffusion policy,” 8th Annual CoRL , 2024
2024
-
[29]
Equibot: Sim (3)-equivariant diffusion policy for generalizable and data efficient learning,
J. Yang, Z. Cao, C. Deng, R. Antonova, S. Song, and J. Bohg, “Equibot: Sim (3)-equivariant diffusion policy for generalizable and data efficient learning,” in 8th Annual CoRL , 2024
2024
-
[30]
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals,
M. Reuss, ¨O. E. Ya ˘gmurlu, F. Wenzel, and R. Lioutikov, “Multimodal diffusion transformer: Learning versatile behavior from multimodal goals,” arXiv:2407.05996, 2024
2024 arXiv
-
[31]
Sparse diffusion policy: A sparse, reusable, and flexible policy for robot learning,
Y . Wang, Y . Zhang, M. Huo, T. Tian, X. Zhang, Y . Xie, C. Xu, P. Ji, W. Zhan, M. Ding et al., “Sparse diffusion policy: A sparse, reusable, and flexible policy for robot learning,” in 8th Annual CoRL , 2024
2024
-
[32]
Skill expansion and composition in parameter space,
T. Liu, J. Li, Y . Zheng, H. Niu, Y . Lan, X. Xu, and X. Zhan, “Skill expansion and composition in parameter space,” in ICLR, 2025
2025
-
[33]
Adamanip: Adaptive articulated object manipulation environments and policy learning,
Y . Wang, X. Zhang, R. Wu, Y . Li, Y . Shen, M. Wu, Z. He, Y . Wang, and H. Dong, “Adamanip: Adaptive articulated object manipulation environments and policy learning,” in The Thirteenth ICLR , 2025
2025
-
[34]
Afforddp: Generalizable diffusion policy with transferable affordance,
S. Wu, Y . Zhu, Y . Huang, K. Zhu, J. Gu, J. Yu, Y . Shi, and J. Wang, “Afforddp: Generalizable diffusion policy with transferable affordance,” in CVPR, 2025
2025
-
[35]
Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation,
Q. Lv, H. Li, X. Deng, R. Shao, Y . Li, J. Hao, L. Gao, M. Y . Wang, and L. Nie, “Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation,” in CVPR, 2025
2025
-
[36]
Behavior robot suite: Streamlining real-world whole-body manipulation for everyday household activities,
Y . Jiang, R. Zhang, J. Wong, C. Wang, Y . Ze, H. Yin, C. Gokmen, S. Song, J. Wu, and L. Fei-Fei, “Behavior robot suite: Streamlining real-world whole-body manipulation for everyday household activities,” in RSS 2025 Workshop, 2025
2025
-
[37]
Fast flow-based visuomotor policies via conditional optimal transport couplings,
A. Sochopoulos, N. Malkin, N. Tsagkas, J. Moura, M. Gienger, and S. Vijayakumar, “Fast flow-based visuomotor policies via conditional optimal transport couplings,” arXiv:2505.01179, 2025
2025 arXiv
-
[38]
Octo: An open-source generalist robot policy,
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu et al., “Octo: An open-source generalist robot policy,” arXiv:2405.12213, 2024
2024 arXiv
-
[39]
Diffusion-vla: Scaling robot foundation models via unified diffusion and autoregression,
J. Wen, M. Zhu, Y . Zhu, Z. Tang, J. Li, Z. Zhou, C. Li, X. Liu, Y . Peng, C. Shen et al. , “Diffusion-vla: Scaling robot foundation models via unified diffusion and autoregression,” arXiv:2412.03293, 2024
2024 arXiv
-
[40]
Cogact: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation,
Q. Li, Y . Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei et al. , “Cogact: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation,” arXiv:2411.19650, 2024
2024 arXiv
-
[41]
Chatvla: Unified multimodal understanding and robot control with vision-language-action model,
Z. Zhou, Y . Zhu, M. Zhu, J. Wen, N. Liu, Z. Xu, W. Meng, R. Cheng, Y . Peng, C. Shen et al. , “Chatvla: Unified multimodal understanding and robot control with vision-language-action model,” arXiv:2502.14420, 2025. 18 IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIG...
2025 arXiv
-
[42]
Rdt-1b: a diffusion foundation model for bimanual manipulation,
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu, “Rdt-1b: a diffusion foundation model for bimanual manipulation,” arXiv:2410.07864, 2024
2024 arXiv
-
[43]
Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems,
Q. Bu, J. Cai, L. Chen, X. Cui, Y . Ding, S. Feng, S. Gao et al., “Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems,” arXiv:2503.06669, 2025
2025 arXiv
-
[44]
Gr00t n1: An open foundation model for generalist humanoid robots,
J. Bjorck, F. Casta ˜neda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huang et al., “Gr00t n1: An open foundation model for generalist humanoid robots,” arXiv:2503.14734, 2025
2025 arXiv
-
[45]
Dreamgen: Unlocking generalization in robot learning through neural trajectories,
J. Jang, S. Ye, Z. Lin, J. Xiang, J. Bjorck, Y . Fang, F. Hu, S. Huang, K. Kundalia, Y .-C. Linet al., “Dreamgen: Unlocking generalization in robot learning through neural trajectories,” arXiv:2505.12705, 2025
2025 arXiv
-
[46]
Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model,
J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu et al. , “Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model,” arXiv:2503.10631, 2025
2025 arXiv
-
[47]
Affordance-based robot manipulation with flow matching,
F. Zhang and M. Gienger, “Affordance-based robot manipulation with flow matching,” arXiv:2409.01083, 2024
2024
-
[49]
Actionflow: Efficient, accurate, and fast policies with spatially symmetric flow matching,
N. Funk, J. Urain, J. Carvalho, V . Prasad, G. Chalvatzaki, and J. Pe- ters, “Actionflow: Efficient, accurate, and fast policies with spatially symmetric flow matching,” in R: SS workshop: Structural Priors as Inductive Biases for Learning Robot Dynamics , 2024
2024
-
[50]
π0: A vision-language-action flow model for general robot control,
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom et al., “π0: A vision-language-action flow model for general robot control,” arXiv:2410.24164, vol. 2, no. 3, p. 5, 2024
2024 arXiv
-
[51]
Graspvla: a grasping foundation model pre- trained on billion-scale synthetic action data,
S. Deng, M. Yan, S. Wei, H. Ma, Y . Yang, J. Chen, Z. Zhang, T. Yang, X. Zhang, H. Cui et al., “Graspvla: a grasping foundation model pre- trained on billion-scale synthetic action data,” arXiv:2505.03233, 2025
2025 arXiv
-
[52]
Hi robot: Open-ended instruction following with hierarchical vision-language-action models,
L. X. Shi, B. Ichter, M. Equi, L. Ke, K. Pertsch, Q. Vuong, J. Tan- ner, A. Walling, H. Wang, N. Fusai et al. , “Hi robot: Open-ended instruction following with hierarchical vision-language-action models,” arXiv:2502.19417, 2025
2025 arXiv
-
[53]
Smolvla: A vision-language-action model for affordable and efficient robotics,
M. Shukor, D. Aubakirova, F. Capuano, P. Kooijmans, S. Palma, A. Zouitine, M. Aractingi, C. Pascal, M. Russi, A. Marafioti et al. , “Smolvla: A vision-language-action model for affordable and efficient robotics,” arXiv:2506.01844, 2025
2025 arXiv
-
[54]
Masked visual pre- training for motor control,
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik, “Masked visual pre- training for motor control,” arXiv:2203.06173, 2022
2022 arXiv
-
[55]
Unleashing large-scale video generative pre-training for visual robot manipulation,
H. Wu, Y . Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong, “Unleashing large-scale video generative pre-training for visual robot manipulation,” in The Twelfth ICLR , 2024
2024
-
[56]
Robouniview: Visual-language model with unified view representation for robotic manipulation,
F. Liu, F. Yan, L. Zheng, C. Feng, Y . Huang, and L. Ma, “Robouniview: Visual-language model with unified view representation for robotic manipulation,” arXiv:2406.18977, 2024
2024 arXiv
-
[57]
Lift3d foundation policy: Lifting 2d large-scale pretrained models for robust 3d robotic manipulation,
Y . Jia, J. Liu, S. Chen, C. Gu, Z. Wang, L. Luo, L. Lee, P. Wang, Z. Wang et al. , “Lift3d foundation policy: Lifting 2d large-scale pretrained models for robust 3d robotic manipulation,” in CVPR, 2025
2025
-
[58]
Sam2act: Integrating visual foundation model with a memory architecture for robotic manipulation,
H. Fang, M. Grotz, W. Pumacay, Y . R. Wang, D. Fox, R. Krishna, and J. Duan, “Sam2act: Integrating visual foundation model with a memory architecture for robotic manipulation,” arXiv:2501.18564, 2025
2025 arXiv
-
[59]
Fine-tuning vision-language-action models: Optimizing speed and success,
M. J. Kim, C. Finn, and P. Liang, “Fine-tuning vision-language-action models: Optimizing speed and success,” arXiv:2502.19645, 2025
2025 arXiv
-
[60]
A generalist agent,
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y . Sulsky, J. Kay, J. T. Springenberg et al., “A generalist agent,” arXiv:2205.06175, 2022
2022 arXiv
-
[61]
Vima: General robot manipulation with multimodal prompts,
Y . Jiang, A. Gupta, Z. Zhang, G. Wang, Y . Dou, Y . Chen, L. Fei- Fei, A. Anandkumar, Y . Zhu, and L. Fan, “Vima: General robot manipulation with multimodal prompts,” in NeurIPS 2022 Foundation Models for Decision Making Workshop , 2022
2022
-
[62]
Palm-e: An embodied multimodal language model,
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu et al., “Palm-e: An embodied multimodal language model,” in ICML. PMLR, 2023, pp. 8469–8488
2023
-
[64]
Vision-language foundation models as effective robot imitators,
X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y . Jing, W. Zhang, H. Liu et al. , “Vision-language foundation models as effective robot imitators,” in The Twelfth ICLR , 2024
2024
-
[65]
3d-vla: A 3d vision-language-action generative world model,
H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y . Du, Y . Hong, and C. Gan, “3d-vla: A 3d vision-language-action generative world model,” in ICML. PMLR, 2024, pp. 61 229–61 245
2024
-
[66]
Behavior generation with latent actions,
S. Lee, Y . Wang, H. Etukuru, H. J. Kim, N. M. M. Shafiullah, and L. Pinto, “Behavior generation with latent actions,” in ICML. PMLR, 2024, pp. 26 991–27 008
2024
-
[67]
Openvla: An open-source vision-language-action model,
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong et al., “Openvla: An open-source vision-language-action model,” in 8th Annual CoRL, 2024
2024
-
[68]
Quest: Self- supervised skill abstractions for learning continuous control,
A. Mete, H. Xue, A. Wilcox, Y . Chen, and A. Garg, “Quest: Self- supervised skill abstractions for learning continuous control,” NeurIPS, vol. 37, pp. 4062–4089, 2025
2025
-
[69]
Au- toregressive action sequence learning for robotic manipulation,
X. Zhang, Y . Liu, H. Chang, L. Schramm, and A. Boularias, “Au- toregressive action sequence learning for robotic manipulation,” Rob. Autom. Lett., 2025
2025
-
[70]
Carp: Visuomotor policy learning via coarse-to-fine autoregressive prediction,
Z. Gong, P. Ding, S. Lyu, S. Huang, M. Sun, W. Zhao, Z. Fan, and D. Wang, “Carp: Visuomotor policy learning via coarse-to-fine autoregressive prediction,” arXiv:2412.06782, 2024
2024 arXiv
-
[71]
Tracevla: Visual trace prompting en- hances spatial-temporal awareness for generalist robotic policies,
R. Zheng, Y . Liang, S. Huang, J. Gao, H. Daum ´e III, A. Kolobov, F. Huang, and J. Yang, “Tracevla: Visual trace prompting en- hances spatial-temporal awareness for generalist robotic policies,” arXiv:2412.10345, 2024
2024 arXiv
-
[72]
Towards generalist robot policies: What matters in building vision-language-action models,
X. Li, P. Li, M. Liu, D. Wang, J. Liu, B. Kang, X. Ma, T. Kong, H. Zhang, and H. Liu, “Towards generalist robot policies: What matters in building vision-language-action models,” arXiv:2412.14058, 2024
2024 arXiv
-
[73]
Fast: Efficient action tokenization for vision-language-action models,
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine, “Fast: Efficient action tokenization for vision-language-action models,” arXiv:2501.09747, 2025
2025 arXiv
-
[74]
Spatialvla: Exploring spatial representations for visual-language-action model,
D. Qu, H. Song, Q. Chen, Y . Yao, X. Ye, Y . Ding, Z. Wang, J. Gu, B. Zhao, D. Wang et al., “Spatialvla: Exploring spatial representations for visual-language-action model,” arXiv:2501.15830, 2025
2025 arXiv
-
[75]
Vla-cache: Towards efficient vision-language-action model via adaptive token caching in robotic manipulation,
S. Xu, Y . Wang, C. Xia, D. Zhu, T. Huang, and C. Xu, “Vla-cache: Towards efficient vision-language-action model via adaptive token caching in robotic manipulation,” arXiv:2502.02175, 2025
2025
-
[76]
Hamster: Hierarchical action models for open-world robot manipulation,
Y . Li, Y . Deng, J. Zhang, J. Jang, M. Memmel, R. Yu, C. R. Garrett, F. Ramos, D. Fox, A. Li et al., “Hamster: Hierarchical action models for open-world robot manipulation,” arXiv:2502.05485, 2025
2025 arXiv
-
[77]
Magma: A foundation model for multimodal ai agents,
J. Yang, R. Tan, Q. Wu, R. Zheng, B. Peng, Y . Liang, Y . Gu, M. Cai, S. Ye, J. Jang et al., “Magma: A foundation model for multimodal ai agents,” in CVPR, 2025, pp. 14 203–14 214
2025
-
[78]
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,
Q. Zhao, Y . Lu, M. J. Kim, Z. Fu, Z. Zhang, Y . Wu, Z. Li, Q. Ma, S. Han, C. Finn et al., “Cot-vla: Visual chain-of-thought reasoning for vision-language-action models,” in CVPR, 2025, pp. 1702–1713
2025
-
[79]
Univla: Learning to act anywhere with task-centric latent actions,
Q. Bu, Y . Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li, “Univla: Learning to act anywhere with task-centric latent actions,” arXiv:2505.06111, 2025
2025 arXiv
-
[80]
Latent action pretraining from videos,
S. Ye, J. Jang, B. Jeon, S. J. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y .-W. Chao, B. Y . Lin et al. , “Latent action pretraining from videos,” in CoRL 2024 Workshop, 2024
2024
-
[81]
Robotic control via embodied chain-of-thought reasoning,
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine, “Robotic control via embodied chain-of-thought reasoning,” in 8th Annual CoRL, 2024
2024
-
[82]
Worldvla: Towards autore- gressive action world model,
J. Cen, C. Yu, H. Yuan, Y . Jiang, S. Huang, J. Guo, X. Li, Y . Song, H. Luo, F. Wang, D. Zhao, and H. Chen, “Worldvla: Towards autore- gressive action world model,” arXiv:2506.21539, 2025
2025 arXiv
-
[83]
Lotus: Continual imitation learning for robot manipulation through unsupervised skill discovery,
W. Wan, Y . Zhu, R. Shah, and Y . Zhu, “Lotus: Continual imitation learning for robot manipulation through unsupervised skill discovery,” in ICRA, 2024, pp. 537–544
2024
-
[84]
What matters in language conditioned robotic imitation learning over unstructured data,
O. Mees, L. Hermann, and W. Burgard, “What matters in language conditioned robotic imitation learning over unstructured data,” Rob. Autom. Lett., vol. 7, no. 4, pp. 11 205–11 212, 2022
2022
-
[85]
Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models,
P. Li, Y . Chen, H. Wu, X. Ma, X. Wu, Y . Huang, L. Wang, T. Kong, and T. Tan, “Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models,” arXiv:2506.07961, 2025
2025
-
[86]
Se (3)-diffusionfields: Learning smooth cost functions for joint grasp and motion optimization through diffusion,
J. Urain, N. Funk, J. Peters, and G. Chalvatzaki, “Se (3)-diffusionfields: Learning smooth cost functions for joint grasp and motion optimization through diffusion,” in ICRA. IEEE, 2023, pp. 5923–5930
2023
-
[87]
A0: An affordance-aware hierarchical model for general robotic manipulation,
R. Xu, J. Zhang, M. Guo, Y . Wen, H. Yang, M. Lin, J. Huang, Z. Li, K. Zhang, L. Wang et al., “A0: An affordance-aware hierarchical model for general robotic manipulation,” arXiv:2504.12636, 2025
2025
-
[88]
Flow matching imitation learning for multi-support manipulation,
Q. Rouxel, A. Ferrari, S. Ivaldi, and J.-B. Mouret, “Flow matching imitation learning for multi-support manipulation,” in Humanoids. IEEE, 2024, pp. 528–535
2024
-
[89]
Polarnet: 3d point clouds for language-guided robotic manipulation,
S. Chen, R. G. Pinel, C. Schmid, and I. Laptev, “Polarnet: 3d point clouds for language-guided robotic manipulation,” in CoRL. PMLR, 2023, pp. 1761–1781
2023
-
[90]
Instruction-driven history-aware policies for robotic ma- nipulations,
P.-L. Guhur, S. Chen, R. G. Pinel, M. Tapaswi, I. Laptev, and C. Schmid, “Instruction-driven history-aware policies for robotic ma- nipulations,” in CoRL. PMLR, 2023, pp. 175–187
2023
-
[91]
Perceiver-actor: A multi-task transformer for robotic manipulation,
M. Shridhar, L. Manuelli, and D. Fox, “Perceiver-actor: A multi-task transformer for robotic manipulation,” in CoRL, 2023. LI et al.: ROBOTIC MANIPULATION VIA IMITATION LEARNING: TAXONOMY , EVOLUTION, BENCHMARK, AND CHALLENGES 19
2023
-
[92]
Rvt: Robotic view transformer for 3d object manipulation,
A. Goyal, J. Xu, Y . Guo, V . Blukis, Y .-W. Chao, and D. Fox, “Rvt: Robotic view transformer for 3d object manipulation,” in CoRL. PMLR, 2023, pp. 694–710
2023
-
[93]
Act3d: 3d feature field transformers for multi-task robotic manipulation,
T. Gervet, Z. Xian, N. Gkanatsios, and K. Fragkiadaki, “Act3d: 3d feature field transformers for multi-task robotic manipulation,” in CoRL. PMLR, 2023, pp. 3949–3965
2023
-
[94]
Sam- e: Leveraging visual foundation model with sequence imitation for embodied manipulation,
J. Zhang, C. Bai, H. He, Z. Wang, B. Zhao, X. Li, and X. Li, “Sam- e: Leveraging visual foundation model with sequence imitation for embodied manipulation,” in ICML. PMLR, 2024, pp. 58 579–58 598
2024
-
[95]
Equact: An se (3)- equivariant multi-task transformer for open-loop robotic manipulation,
X. Zhu, Y . Qi, Y . Zhu, R. Walters, and R. Platt, “Equact: An se (3)- equivariant multi-task transformer for open-loop robotic manipulation,” arXiv:2505.21351, 2025
2025 arXiv
-
[97]
Moka: Open-vocabulary robotic manipulation through mark-based visual prompting,
F. Liu, K. Fang, P. Abbeel, and S. Levine, “Moka: Open-vocabulary robotic manipulation through mark-based visual prompting,” in ICRA Workshop, 2024
2024
-
[98]
Ram: Retrieval-based affordance transfer for generalizable zero-shot robotic manipulation,
Y . Kuang, J. Ye, H. Geng, J. Mao, C. Deng, L. Guibas, H. Wang, and Y . Wang, “Ram: Retrieval-based affordance transfer for generalizable zero-shot robotic manipulation,” in 8th Annual CoRL , 2024
2024
-
[99]
Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation,
W. Huang, C. Wang, Y . Li, R. Zhang, and L. Fei-Fei, “Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation,” in 2nd CoRL Workshop, 2024
2024
-
[100]
Towards generalizable vision- language robotic manipulation: A benchmark and llm-guided 3d pol- icy,
R. Garcia, S. Chen, and C. Schmid, “Towards generalizable vision- language robotic manipulation: A benchmark and llm-guided 3d pol- icy,” arXiv:2410.01345, 2024
2024 arXiv
-
[104]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Loet al., “Segment anything,” in ICCV, 2023, pp. 4015–4026
2023
-
[105]
V oxposer: Composable 3d value maps for robotic manipulation with language models,
W. Huang, C. Wang, R. Zhang, Y . Li, J. Wu, and L. Fei-Fei, “V oxposer: Composable 3d value maps for robotic manipulation with language models,” arXiv:2307.05973, 2023
2023 arXiv
-
[106]
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee et al., “Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,” in ICRA, 2024, pp. 6892–6903
2024
-
[107]
Switchvla: Execution-aware task switching for vision- language-action models,
M. Li, Z. Zhao, Z. Che, F. Liao, K. Wu, Z. Xu, P. Ren, Z. Jin, N. Liu, and J. Tang, “Switchvla: Execution-aware task switching for vision- language-action models,” arXiv:2506.03574, 2025
2025 arXiv
-
[108]
Rlbench: The robot learning benchmark & learning environment,
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison, “Rlbench: The robot learning benchmark & learning environment,” Rob. Autom. Lett., vol. 5, no. 2, pp. 3019–3026, 2020
2020
-
[109]
Object-centric representations improve policy generalization in robot manipulation,
A. Chapin, B. Machado, E. Dellandrea, and L. Chen, “Object-centric representations improve policy generalization in robot manipulation,” arXiv:2505.11563, 2025
2025 arXiv
-
[110]
Adpro: a test-time adaptive diffusion policy for robot manipulation via manifold and initial noise constraints,
Z. Li, R. Yang, R. Chen, Z. Luo, and L. Chen, “Adpro: a test-time adaptive diffusion policy for robot manipulation via manifold and initial noise constraints,” arXiv:2508.06266, 2025
2025
-
[111]
Llm-based skill diffusion for zero-shot policy adaptation,
W. K. Kim, Y . Lee, J. Kim, and H. Woo, “Llm-based skill diffusion for zero-shot policy adaptation,” NeurIPS, vol. 37, pp. 6749–6775, 2024
2024
-
[112]
panda-gym: Open-source goal-conditioned environments for robotic learning,
Q. Gallou ´edec, N. Cazin, E. Dellandr ´ea, and L. Chen, “panda-gym: Open-source goal-conditioned environments for robotic learning,” in 4th NeurIPS Workshop, 2021
2021
-
[113]
Roboverse: Towards a unified platform, dataset and benchmark for scalable and generalizable robot learning,
H. Geng, F. Wang, S. Wei, Y . Li, B. Wang, B. An, C. T. Cheng et al., “Roboverse: Towards a unified platform, dataset and benchmark for scalable and generalizable robot learning,” arXiv:2504.18904, 2025
2025 arXiv
-
[114]
Mujoco playground,
K. Zakka, B. Tabanpour, Q. Liao, M. Haiderbhai, S. Holt, J. Y . Luo, A. Allshire, E. Frey, K. Sreenath, L. A. Kahrs et al. , “Mujoco playground,” arXiv:2502.08844, 2025
2025 arXiv
-
[115]
Jacquard: A large scale dataset for robotic grasp detection,
A. Depierre, E. Dellandr ´ea, and L. Chen, “Jacquard: A large scale dataset for robotic grasp detection,” in IROS, 2018, pp. 3511–3516
2018
-
[116]
Rh20t: A robotic dataset for learning diverse skills in one-shot,
H.-S. Fang, H. Fang, Z. Tang, J. Liu, J. Wang, H. Zhu, and C. Lu, “Rh20t: A robotic dataset for learning diverse skills in one-shot,” in RSS 2023 Workshop on Learning for Task and Motion Planning , 2023
2023
-
[117]
Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipulation tasks,
O. Mees, L. Hermann et al. , “Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipulation tasks,” Rob. Autom. Lett. , vol. 7, no. 3, pp. 7327–7334, 2022
2022
-
[118]
Pybullet, a python module for physics simulation for games, robotics and machine learning,
E. Coumans and Y . Bai, “Pybullet, a python module for physics simulation for games, robotics and machine learning,” 2016
2016
-
[119]
V-rep: A versatile and scalable robot simulation framework,
E. Rohmer, S. P. Singh, and M. Freese, “V-rep: A versatile and scalable robot simulation framework,” in IROS. IEEE, 2013, pp. 1321–1326
2013
-
[120]
Libero: Benchmarking knowledge transfer for lifelong robot learning,
B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone, “Libero: Benchmarking knowledge transfer for lifelong robot learning,” NeurIPS, vol. 36, pp. 44 776–44 791, 2023
2023
-
[121]
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine, “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in CoRL. PMLR, 2020, pp. 1094–1100
2020
-
[122]
robosuite: A modular simulation framework and benchmark for robot learning,
Y . Zhu, J. Wong, A. Mandlekar, R. Mart ´ın-Mart´ın, A. Joshi, S. Nasiri- any, and Y . Zhu, “robosuite: A modular simulation framework and benchmark for robot learning,” arXiv:2009.12293, 2020
2009 arXiv
-
[123]
What matters in learning from offline human demonstrations for robot manipulation,
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei et al. , “What matters in learning from offline human demonstrations for robot manipulation,” in 5th Annual CoRL , 2021
2021
-
[124]
Maniskill2: A unified benchmark for generalizable manipulation skills,
J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, Y . Tang, S. Tao, X. Wei, Y . Yao, X. Yuan, P. Xie et al., “Maniskill2: A unified benchmark for generalizable manipulation skills,” in ICLR, 2023
2023
-
[125]
Evaluating real-world robot manipulation policies in simulation,
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani et al. , “Evaluating real-world robot manipulation policies in simulation,” arXiv:2405.05941, 2024
2024 arXiv
-
[126]
The colosseum: A benchmark for evaluating generalization for robotic manipulation,
W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox, “The colosseum: A benchmark for evaluating generalization for robotic manipulation,” in RSS Workshop, 2024
2024
-
[127]
D4rl: Datasets for deep data-driven reinforcement learning,
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine, “D4rl: Datasets for deep data-driven reinforcement learning,” 2020
2020
-
[128]
Language conditioned imitation learning over unstructured data,
C. Lynch and P. Sermanet, “Language conditioned imitation learning over unstructured data,” arXiv:2005.07648, 2020
2005 arXiv
-
[129]
Language-conditioned imitation learning with base skill priors under unstructured data,
H. Zhou, Z. Bing, X. Yao, X. Su, C. Yang, K. Huang, and A. Knoll, “Language-conditioned imitation learning with base skill priors under unstructured data,” Rob. Autom. Lett. , 2024
2024
-
[130]
Closed-loop visuomotor control with generative expectation for robotic manipulation,
Y . Yang, “Closed-loop visuomotor control with generative expectation for robotic manipulation,” in The Conference on Neural Information Processing Systems, 2024
2024
-
[131]
Diffusion transformer policy,
Z. Hou, T. Zhang, Y . Xiong, H. Pu, C. Zhao, R. Tong, Y . Qiao, J. Dai, and Y . Chen, “Diffusion transformer policy,” arXiv:2410.15959, 2024
2024 arXiv
-
[132]
Ghil-glue: Hierarchical control with filtered subgoal images,
K. B. Hatch, A. Balakrishna, O. Mees, S. Nair, S. Park, B. Wulfe, M. Itkina, B. Eysenbach, S. Levine et al. , “Ghil-glue: Hierarchical control with filtered subgoal images,” arXiv:2410.20018, 2024
2024 arXiv
-
[133]
Efficient diffusion transformer policies with mixture of expert denoisers for multitask learning,
M. Reuss, J. Pari, P. Agrawal, and R. Lioutikov, “Efficient diffusion transformer policies with mixture of expert denoisers for multitask learning,” arXiv:2412.12953, 2024
2024 arXiv
-
[134]
Gr-mg: Leveraging partially-annotated data via multi-modal goal-conditioned policy,
P. Li, H. Wu, Y . Huang, C. Cheang, L. Wang, and T. Kong, “Gr-mg: Leveraging partially-annotated data via multi-modal goal-conditioned policy,” Rob. Autom. Lett. , 2025
2025
-
[135]
Predictive inverse dynamics models are scalable learners for robotic manipulation,
Y . Tian, S. Yang, J. Zeng, P. Wang, D. Lin, H. Dong, and J. Pang, “Predictive inverse dynamics models are scalable learners for robotic manipulation,” in The Thirteenth ICLR , 2025
2025
-
[136]
Rvt- 2: Learning precise manipulation from few demonstrations,
A. Goyal, V . Blukis, J. Xu, Y . Guo, Y .-W. Chao, and D. Fox, “Rvt- 2: Learning precise manipulation from few demonstrations,” in RSS Workshop, 2024
2024
-
[137]
Train a multi-task diffusion policy on rlbench-18 in one day with one gpu,
Y . Hu, P. S., K. Wen, and R. Detry, “Train a multi-task diffusion policy on rlbench-18 in one day with one gpu,” arXiv:2505.09430, 2025
2025 arXiv
-
[138]
Coarse-to- fine q-attention: Efficient learning for visual robotic manipulation via discretisation,
S. James, K. Wada, T. Laidlow, and A. J. Davison, “Coarse-to- fine q-attention: Efficient learning for visual robotic manipulation via discretisation,” in CVPR, 2022, pp. 13 739–13 748
2022
-
[139]
Mail: Improving imi- tation learning with selective state space models,
X. Jia, Q. Wang, A. Donat, B. Xing, G. Li, H. Zhou, O. Celik, D. Blessing, R. Lioutikov, and G. Neumann, “Mail: Improving imi- tation learning with selective state space models,” in CoRL, 2024
2024
-
[140]
Tinyllava: A framework of small-scale large multimodal models,
B. Zhou, Y . Hu, X. Weng, J. Jia, J. Luo, X. Liu, J. Wu, and L. Huang, “Tinyllava: A framework of small-scale large multimodal models,” arXiv:2402.14289, 2024
2024 arXiv
-
[141]
Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation,
Z. Qi, W. Zhang, Y . Ding, R. Dong, X. Yu, J. Li, L. Xu, B. Li, X. He, G. Fan et al. , “Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation,” arXiv:2502.13143, 2025
2025
-
[142]
Generative image as action models,
M. Shridhar, Y . L. Lo, and S. James, “Generative image as action models,” arXiv:2407.07875, 2024
2024 arXiv
-
[143]
R3m: A universal visual representation for robot manipulation,
S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta, “R3m: A universal visual representation for robot manipulation,” in CoRL. PMLR, 2023, pp. 892–909
2023
-
[144]
Manual2skill: Learning to read manuals and acquire robotic skills for furniture assembly using vision-language models,
C. Tie, S. Sun, J. Zhu, Y . Liu, J. Guo, Y . Hu, H. Chen, J. Chen, R. Wu, and L. Shao, “Manual2skill: Learning to read manuals and acquire robotic skills for furniture assembly using vision-language models,” arXiv:2502.10090, 2025
2025
-
[145]
Srsa: Skill retrieval and adaptation for robotic assembly tasks,
Y . Guo, B. Tang, I. Akinola, D. Fox, A. Gupta, and Y . Narang, “Srsa: Skill retrieval and adaptation for robotic assembly tasks,” in The Fifteenth ICLR , 2025. 20 IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE. PREPRINT VERSION
2025
-
[146]
Point-level visual affordance guided retrieval and adaptation for cluttered garments manipulation,
R. Wu, Z. Zhu, Y . Wang, Y . Chen, J. Wang, and H. Dong, “Point-level visual affordance guided retrieval and adaptation for cluttered garments manipulation,” in CVPR, 2025
2025
-
[147]
Do as i can, not as i say: Grounding language in robotic affordances,
M. Ahn, A. Brohan, N. Brown, Y . Chebotar, O. Cortes, B. David, C. Finn, C. Fu et al., “Do as i can, not as i say: Grounding language in robotic affordances,” arXiv:2204.01691, 2022
2022 arXiv
-
[148]
Shake-vla: Vision-language-action model-based system for bimanual robotic manipulations and liquid mixing,
M. H. Khan, S. Asfaw, D. Iarchuk, M. A. Cabrera, L. Moreno et al., “Shake-vla: Vision-language-action model-based system for bimanual robotic manipulations and liquid mixing,” arXiv:2501.06919, 2025
2025 arXiv
-
[149]
An improved single short detection method for smart vision-based water garbage cleaning robot,
A. Haldorai, M. Suriya, M. Balakrishnan et al. , “An improved single short detection method for smart vision-based water garbage cleaning robot,” Cognitive Robotics, vol. 4, pp. 19–29, 2024
2024
-
[150]
Learning fusion feature representation for garbage image classification model in human–robot interaction,
X. Li, T. Li, S. Li, B. Tian, J. Ju et al. , “Learning fusion feature representation for garbage image classification model in human–robot interaction,” Infrared Physics & Technology, vol. 128, p. 104457, 2023
2023
-
[151]
Machine learning-based garbage detection and 3d spatial localization for intelligent robotic grasp,
Z. Lv, T. Chen, Z. Cai, and Z. Chen, “Machine learning-based garbage detection and 3d spatial localization for intelligent robotic grasp,” Applied Sciences, vol. 13, no. 18, p. 10018, 2023
2023
-
[152]
Foundational Models for Robotics need to be made Bio-Inspired,
L. Chen and S. M. Nguyen, “Foundational Models for Robotics need to be made Bio-Inspired,” in ARSO, Jul. 2025
2025
-
[153]
Waypoint-based reinforcement learning for robot manipulation tasks,
S. Mehta, S. Habibian, and D. Losey, “Waypoint-based reinforcement learning for robot manipulation tasks,” in IROS, 2024, pp. 541–548. Zezeng Li received a B.S. degree from Beijing Uni- versity of Technology (BJUT) in 2015 and a Ph.D. degree from Dalian University of Technolog...
2024
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.