Pith. sign in

REVIEW 5 major objections 5 minor 4 cited by

An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper claims that decomposing manipulation tasks into reusable atomic skills, then fine-tuning a vision-language-action model per skill, cuts data costs while matching or exceeding end-to-end success rates.

desk verdict A genuinely new dynamic skill-library idea with real-robot evidence that is weaker than the claims: the missing-skill adaptation scenario is never tested and the "exponential" data-efficiency claim is unmeasured. read the letter →

arxiv 2501.15068 v3 pith:LF23YGAM submitted 2025-01-25 cs.RO

classification cs.RO
keywords atomicskilllibraryembodiedmanipulationvision-language-actionmodelstaskdecompositiondataefficiencyfew-shotfine-tuningrobotic
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that decomposing end-to-end manipulation tasks into atomic skills—small reusable actions like "lift the bottle" or "align and tilt"—cuts the data needed to train a robot manipulator while keeping or improving success rates. The proposed three-wheeled method uses a vision-language-planning agent to split a task into subtasks, a semantic abstraction step to consolidate those subtasks into general skill definitions, and few-shot fine-tuning of a vision-language-action model to realize each skill in an expandable library. The authors report real-world experiments on a dual-arm robot showing that with less data the skill-based method matches end-to-end training, with equal data it outperforms it, and that entirely new task orderings can be executed by recombining existing skills. If the claim holds, the path to general robot manipulation shifts from collecting full task demonstrations toward maintaining a growing library of reusable skills.

What carries the argument

The load-bearing object is the atomic skill library, built by a three-wheeled loop: the VLP agent wheel (GPT-4 prompted with scene descriptions, object bounding boxes from DINO-X, and segmentation masks from SAM-2 to plan subtasks), the VLA wheel (any pretrained vision-language-action model, e.g., RDT-1B or Octo, fine-tuned with few demonstrations per skill), and the atomic skill wheel (a semantic abstraction module, again an LLM, that maps diverse subtasks to a compact set of skill definitions at a granularity set by VLA plasticity and adaptability). The loop makes the library self-updating: when a new task needs a missing skill, only that skill's data is collected and fine-tuned.

What would settle it

A concrete test: assemble a new task entirely from skills already present in the library, but with the objects placed outside every distribution seen during skill training, and measure success; if the success rate is no better than for an end-to-end model that never saw the task, then the claimed cross-task transfer is not real. The paper's block-ordering results, which show 80-100% success on unseen orders, are the positive instance of this test.

Watch

Extended reading notes

Core claim

The central discovery claimed is that a robot policy trained per atomic skill—rather than per end-to-end task—yields a dynamically expandable library whose coverage of new tasks grows with each added skill, so the marginal data cost of a new task is at most the data for the skills it lacks. The paper's Table 1 shows skill-based fine-tuning matching or beating end-to-end fine-tuning on Pour Water, Pick & Place Banana, and Pick & Place Pen, with up to 40 percentage-point gains in out-of-distribution success under equal data budgets, and Table 2 shows that rearranging block-ordering subtasks succeeds at 80-100% without retraining on the new order, where end-to-end models fail entirely. The authors take this as evidence that mapping tasks to atomic skills materially reduces data requirements and enables cross-task generalization.

Load-bearing premise

The method assumes that the VLP agent and the semantic abstraction module can turn each subtask into atomic skill definitions that remain valid and reusable when the task changes, but the paper never measures decomposition or abstraction accuracy, so if skills are inconsistent or task-specific, transfer fails and the data-efficiency claim collapses.

Editorial extensions

If this is right

  • A new task whose subtasks are already covered by the library can be executed with zero new data collection.
  • The data needed to add a task scales with the number of missing skills, not the complexity of the full task.
  • Fixed data budgets can be spent on more diverse object and scene positions per skill, improving out-of-distribution generalization.
  • The framework is agnostic to the choice of VLA backbone, so improvements in pretrained policies directly upgrade the library.
  • As the library expands, the set of addressable tasks grows through skill recombination.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's data-efficiency claim implicitly assumes that semantic abstraction yields skills that transfer across tasks; a direct test would measure whether a skill fine-tuned for one scene or object works in a new scene with zero or very few demonstrations.
  • If VLA plasticity is low, the required skill granularity becomes very fine, and the library may fragment into near-task-specific entries, eroding the data savings; evaluating this trade-off would require systematic variation of VLA capacity.
  • A natural extension is to treat skill reuse as the evaluation metric itself, for example by reporting the percentage of new tasks whose skills are fully covered by an accumulated library, which would quantify the avoidance of data explosion directly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes a three-wheeled framework for building an atomic skill library for embodied manipulation. A VLP agent decomposes a task into subtasks; a semantic abstraction module maps these subtasks to reusable atomic skill definitions; and a VLA model is fine-tuned with small per-skill datasets to construct the library. The authors claim that this decomposition reduces data costs relative to end-to-end training, improves generalization to out-of-distribution object positions, and enables adaptation to new tasks by composing existing skills. Experiments are conducted on a real ALOHA dual-arm robot with two VLA backbones (Octo and RDT-1B) across four tasks, including a block-reordering task used for the new-task evaluation.

Significance. The problem is timely and the general direction is plausible: shifting from end-to-end task data to reusable skill-level data could indeed mitigate the data explosion faced by VLA policies. The paper reports real-robot experiments with two different VLA backbones, which is a strength, and the proposed dynamic skill-library concept goes beyond methods that rely on a fixed predefined skill set. If the central claims were supported by adequate evidence, the work would be a useful contribution to data-efficient manipulation. The current evidence, however, is too thin and confounded to establish the data-efficiency and adaptation claims.

major comments (5)
  1. [§4.1 / Table 1] The evaluation uses 10 trials per condition with no confidence intervals or significance tests. For binary success/failure outcomes, 10 trials cannot resolve the 10-20 percentage point differences that the Q1/Q2 analysis relies on; for example, a change from 40 to 60 is only two trials. The text in §4.2 even says 'All the success rates of our method are no less than the end-to-end method,' but in the Pour Water mug-OOD condition Octo(Ours) grasp is 40 vs Octo(End-to-end) 60. Please report confidence intervals, use more trials or statistical tests, and define a composite task success rate in addition to the per-phase rates.
  2. [§4.2 (Q2)] The 'Ours-plus' comparison is confounded. It is defined as maintaining the same data volume as end-to-end while using a 'larger distribution of data points,' so it changes both the data distribution and the use of the skill decomposition simultaneously. The observed gains (e.g., RDT(Ours-plus) banana both-OOD 80/70 vs RDT(Ours) 60/30) cannot be attributed to the atomic-skill library. A control that trains an end-to-end model on the same position-diverse data with the same total demo count is needed before claiming that skill decomposition, rather than data diversity, drives the improvement.
  3. [§3.1 / §3.4 / Table 2] Table 2 is the only evidence for the new-task adaptation claim, but every condition is a reordering of the same three block-moving skills that were already trained (red, green, blue). The scenario in §3.1 and §3.4 where a genuinely missing skill is detected, newly abstracted, learned from additional data, and then composed is never tested. Please add an experiment with a novel skill absent from the library, report the additional data required, and verify that the VLP agent recognizes the missing skill and that the updated library composes correctly on the new task.
  4. [§3.2 / §3.4] The manuscript never measures the quality of the VLP decomposition or the semantic abstraction of subtasks into atomic skill definitions. Since the data-efficiency and transfer claims depend on skills being reusable across tasks, the absence of any evaluation of decomposition or abstraction accuracy is a load-bearing gap. Please report examples of generated subtasks and skill definitions, a quantitative check of decomposition and abstraction consistency (e.g., human agreement or plan-execution success), and a clear statement of how skill granularity is set in practice.
  5. [§3.3] The data-cost accounting is incomplete. Section 3.3 states that the RDT-1B model was fine-tuned on 6,000 open-source plus 2,000 proprietary trajectories before the few-shot skill experiments, but the experimental demo counts (e.g., 9+9 for Pour Water) exclude this cost, and it is unclear whether the end-to-end baselines include it. The paper should state whether this 8,000-trajectory fine-tuning is a shared prerequisite for both methods and, if so, include it in the total-data-cost comparison.
minor comments (5)
  1. [§1] The phrase 'exponential improvement of data collection efficiency' is not supported by the reported data counts, which show constant-factor reductions; please temper or quantify this claim.
  2. [§4.1] The data count for the Move blocks task is ambiguous: '10 demos of moving red, green, and blue block respectively' could mean 10 per color (30 total) or 10 total; please clarify.
  3. [§4.1 / Table 1] Table 1 reports per-phase success rates (e.g., 'Pick up | Place' and 'Grasp | Pour') but the text frequently refers to 'task success rate'; please define a composite success metric for multi-phase tasks or state explicitly that per-phase rates are reported.
  4. [Table 2] The comparison against end-to-end baselines in Table 2 is informative only as a zero-shot transfer test, since the end-to-end model is trained exclusively on one order and is given no new-task data; please state this more precisely rather than concluding that end-to-end methods 'cannot handle new tasks at all.'
  5. [§3.4] Some implementation details of the semantic abstraction module (e.g., prompts, output schema, and how skill granularity is chosen) are missing; providing these details would improve reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the data-efficiency claims are grounded in matched/held-out comparisons, and the new-task composition test, while limited, is not circular.

full rationale

The paper's central claim is that decomposing end-to-end tasks into atomic skills and fine-tuning a VLA for each skill reduces data cost while maintaining or improving performance. This claim is evaluated in Table 1 by comparing skill-based fine-tuning against end-to-end fine-tuning under matched data distributions and held-out object positions. Those held-out (OOD) evaluations provide an independent check that does not reduce to the training data by construction. The new-task evaluation in Table 2 recombines already-trained block-moving skills into different orders; this tests composition of existing skills, which is exactly the 'covered skills' scenario described in Section 3.1, rather than the missing-skill update scenario. The absence of an experiment that introduces a genuinely missing skill is a limitation in the evidence for the adaptation claim, but it is not circularity: the reported success on permutations does not assume the conclusion that new tasks can be learned with little extra data. The choice of 1,000 fine-tuning steps from a 1k/2k/4k sweep is a hyperparameter selection, not a fitted parameter renamed as a prediction. No load-bearing step is defined in terms of the result it is supposed to establish, and no self-citation chain is used to justify the central mechanism. The VLP and semantic-abstraction modules are not directly evaluated, but that is an evidence gap, not a circular reduction. Overall, the derivation is self-contained with respect to its empirical comparisons, and no specific circular step can be exhibited from the paper's text.

Assumptions & free parameters 5 free parameters · 3 assumptions · 1 invented entities

The method has no formal derivation; it is an empirical system. The free parameters above are choices that affect the reported numbers. The key assumption is that subtask-to-skill abstraction transfers across tasks. No new physical entities are introduced; the 'atomic skill library' is a conceptual construct.

free parameters (5)
  • Few-shot training steps = 1,000 steps
    Selected from a sweep over 1,000/2,000/4,000 steps on 8 L40s GPUs and reported as the best trade-off; used for all skill fine-tuning in the main experiments.
  • Demos per atomic skill = 9/9 for pour water, 9/6 for banana, 9/9 for pen, 10 per block
    Chosen by the authors for the data-efficiency comparison; no procedure is given for how these counts were set.
  • Training position grid = 9 positions on a table
    Used to collect few-shot data for grasp generalization; chosen manually with no stated selection rule.
  • RDT-1B fine-tuning data mix = 6,000 open-source + 2,000 proprietary trajectories
    Global pre-fine-tune dataset for the backbone; composition chosen by the authors without ablations.
  • Skill granularity = Not quantified
    Set by 'plasticity and adaptability' of the VLA in §3.3; no objective criterion is given for how coarse or fine the atomic skill definitions should be.
assumptions (3)
  • domain assumption Subtasks produced by the VLP agent can be abstracted into atomic skill definitions that are reusable across different end-to-end tasks.
    The entire data-efficiency gain rests on this; introduced in §3.4 and used in the 'give guest a cup of water' example.
  • domain assumption The off-the-shelf components (Prismatic, DINO-X, SAM-2, GPT-4) provide sufficiently accurate scene description, detection, segmentation, and planning for task decomposition.
    Used throughout §3.2; decomposition quality is never separately measured.
  • domain assumption A few demonstrations per skill (9-10) are sufficient to teach a robust atomic skill, given a pretrained VLA.
    The experimental protocol in §4.1 assumes this; no scaling study is provided except the 1k/2k/4k step test.
invented entities (1)
  • Atomic skill library
    purpose: Stores fine-tuned VLA policies for each abstracted skill so new tasks can be composed from existing skills.
    The library is the central construct of the method; its only evidence so far is the paper's own success-rate tables, with no external benchmark or released artifact to validate it.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation." pith.science (2026). https://pith.science/paper/LF23YGAM

@misc{pith2026250115068,
  author       = {Pith},
  title        = {Pith review of: An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LF23YGAM}},
  note         = {Machine review of arXiv:2501.15068}
}
read the original abstract

Embodied manipulation is a fundamental ability in the realm of embodied artificial intelligence. Although current embodied manipulation models show certain generalizations in specific settings, they struggle in new environments and tasks due to the complexity and diversity of real-world scenarios. The traditional end-to-end data collection and training manner leads to significant data demands. Decomposing end-to-end tasks into atomic skills helps reduce data requirements and improves the task success rate. However, existing methods are limited by predefined skill sets that cannot be dynamically updated. To address the issue, we introduce a three-wheeled data-driven method to build an atomic skill library. We divide tasks into subtasks using the Vision-Language-Planning (VLP). Then, atomic skill definitions are formed by abstracting the subtasks. Finally, an atomic skill library is constructed via data collection and Vision-Language-Action (VLA) fine-tuning. As the atomic skill library expands dynamically with the three-wheel update strategy, the range of tasks it can cover grows naturally. In this way, our method shifts focus from end-to-end tasks to atomic skills, significantly reducing data costs while maintaining high performance and enabling efficient adaptation to new tasks. Extensive experiments in real-world settings demonstrate the effectiveness and efficiency of our approach.

Figures

Figures reproduced from arXiv: 2501.15068 by the authors.

Figure 1
Figure 1. Three-Wheeled Self-Driven Atomic Skill Library Construction and Inference Pipeline. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The VLP agent reasoning chain framework based on spatial intelligence information. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Task definitions and visualization. Here, we conduct a detailed analysis of the experimental results to address the three questions raised earlier. Q1: From [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition

    cs.RO 2026-07 conditional novelty 6.0 of 10

    VLA skills that score 77-100% in isolation stall from chained states in BEHAVIOR-1K; failures trace to next-skill readiness, target grounding, and control execution.

  2. ROSA: Harnessing Robot States for Vision-Language and Action Alignment

    cs.RO 2025-06 conditional novelty 6.0 of 10

    ROSA trains a VLA model jointly on expert actions and automatically recorded robot states, improving success rates and generalization, particularly with few demonstrations.

  3. SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

    cs.RO 2026-08 conditional novelty 5.0 of 10

    SkillMemo couples MoE-based skill discovery with episodic memory retrieval and reports consistent success-rate gains on diffusion and VLA policies for simulated and real manipulation tasks.

  4. Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents

    cs.RO 2025-05 conditional novelty 4.0 of 10

    An agentic framework using a GPT-4o planner, an OpenVLA executor, and a LoRA-fine-tuned Qwen2.5-VL verifier achieves 79.6% average success on LIBERO by decomposing and verifying subgoals.

Reference graph

Works this paper leans on

23 extracted references · 5 canonical work pages · cited by 4 Pith papers

  1. [1]

    pi0: A vision-language-action flow model for gen- eral robot control

    [Black et al., 2024] Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Nic- colo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. pi0: A vision-language-action flow model for gen- eral robot control. arXiv preprint arXiv:2410.24164,

  2. [5]

    Denoising diffusion probabilistic models

    [Ho et al., 2020] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems , 33:6840–6851,

  3. [6]

    Prismatic vlms: Investigating the design space of visually-conditioned language models

    [Karamcheti et al., 2024] Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and Dorsa Sadigh. Prismatic vlms: Investigating the design space of visually-conditioned language models. arXiv preprint arXiv:2402.07865,

  4. [7]

    Droid: A large-scale in-the-wild robot manipulation dataset

    [Khazatsky et al., 2024] Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. arXiv preprint arXiv:2403.12945,

  5. [8]

    Openvla: An open-source vision-language- action model

    [Kim et al., 2024] Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag San- keti, et al. Openvla: An open-source vision-language- action model. arXiv preprint arXiv:2406.09246,

  6. [9]

    A review of robot learning for manipu- lation: Challenges, representations, and algorithms

    [Kroemer et al., 2021] Oliver Kroemer, Scott Niekum, and George Konidaris. A review of robot learning for manipu- lation: Challenges, representations, and algorithms. Jour- nal of machine learning research, 22(30):1–82,

  7. [11]

    Rdt-1b: a diffusion foundation model for bimanual manipulation

    [Liu et al., 2024] Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. Rdt-1b: a diffusion foundation model for bimanual manipulation. arXiv preprint arXiv:2410.07864,

  8. [12]

    Robo- matrix: A skill-centric hierarchical framework for scalable robot task planning and execution in open-world

    [Mao et al., 2024] Weixin Mao, Weiheng Zhong, Zhou Jiang, Dong Fang, Zhongyue Zhang, Zihan Lan, Fan Jia, Tiancai Wang, Haoqiang Fan, and Osamu Yoshie. Robo- matrix: A skill-centric hierarchical framework for scalable robot task planning and execution in open-world. arXiv preprint arXiv:2412.00171,

Show all 23 references
  1. [13]

    Open x-embodiment: Robotic learning datasets and rt-x models

    [O’Neill et al., 2023] Abby O’Neill, Abdul Rehman, Abhi- nav Gupta, Abhiram Maddukuri, Abhishek Gupta, Ab- hishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864,

  2. [14]

    Scalable diffusion models with transformers

    [Peebles and Xie, 2023] William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 4195–4205,

  3. [15]

    Sam 2: Segment anything in images and videos

    [Ravi et al., 2024] Nikhila Ravi, Valentin Gabeur, Yuan- Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R ¨adle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714,

  4. [16]

    Dino-x: A unified vision model for open-world object detection and understanding

    [Ren et al., 2024] Tianhe Ren, Yihao Chen, Qing Jiang, Zhaoyang Zeng, Yuda Xiong, Wenlong Liu, Zhengyu Ma, Junyi Shen, Yuan Gao, Xiaoke Jiang, et al. Dino-x: A unified vision model for open-world object detection and understanding. arXiv preprint arXiv:2411.14347,

  5. [17]

    High-resolution image synthesis with latent diffusion models

    [Rombach et al., 2022] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 10684–10695,

  6. [19]

    Octo: An open-source generalist robot policy

    [Team et al., 2024] Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, et al. Octo: An open-source generalist robot policy. arXiv preprint arXiv:2405.12213,

  7. [20]

    Hierarchical reinforcement learning with univer- sal policies for multistep robotic manipulation

    [Yang et al., 2021] Xintong Yang, Ze Ji, Jing Wu, Yu- Kun Lai, Changyun Wei, Guoliang Liu, and Rossitza Setchi. Hierarchical reinforcement learning with univer- sal policies for multistep robotic manipulation. IEEE Transactions on Neural Networks and Learning Systems , 33(9):4...

  8. [21]

    Robotic control via embodied chain-of-thought reasoning

    [Zawalski et al., 2024] Michał Zawalski, William Chen, Karl Pertsch, Oier Mees, Chelsea Finn, and Sergey Levine. Robotic control via embodied chain-of-thought reasoning. arXiv preprint arXiv:2407.08693,

  9. [22]

    The system design and robotic manipulation skills of du- alarm autonomous mobile robot for bomb removing

    [Zhao et al., 2022] Dan Zhao, Fuchun Sun, and Linxiang Li. The system design and robotic manipulation skills of du- alarm autonomous mobile robot for bomb removing. In Journal of Physics: Conference Series, volume 2188, page 012006. IOP Publishing,

  10. [23]

    Learning fine-grained biman- ual manipulation with low-cost hardware

    [Zhao et al., 2023] Tony Z Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained biman- ual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023

  11. [2020]

    Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation

    [Fu et al., 2024] Zipeng Fu, Tony Z Zhao, and Chelsea Finn. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation. arXiv preprint arXiv:2401.02117,

  12. [2021]

    Cogact: A founda- tional vision-language-action model for synergizing cog- nition and action in robotic manipulation

    [Li et al., 2024] Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizhong Zhang, et al. Cogact: A founda- tional vision-language-action model for synergizing cog- nition and action in robotic manipulation. arXiv preprint arXi...

  13. [2022]

    Denoising diffusion implicit models

    [Song et al., 2020] Jiaming Song, Chenlin Meng, and Ste- fano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502,

  14. [2023]

    Skillman—a skill-based robotic manipulation framework based on perception and rea- soning

    [Diab et al., 2020] Mohammed Diab, Mihai Pomarlan, Daniel Beßler, Aliakbar Akbari, Jan Rosell, John Bate- man, and Michael Beetz. Skillman—a skill-based robotic manipulation framework based on perception and rea- soning. Robotics and Autonomous Systems , 134:103653,

  15. [2024]

    Diffusion policy: Visuomotor policy learning via action diffusion

    [Chi et al., 2023] Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. The International Journal of Robotics Research, page 02783649241273668,

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.