Pith. sign in

REVIEW 5 cited by

Keypoint Abstraction using Large Models for Object-Relative Imitation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.23254 v1 pith:UAIWRW2S submitted 2024-10-30 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords objectacrosskeypointsmodelsadditionalconsistentenablingenvironments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generalization to novel object configurations and instances across diverse tasks and environments is a critical challenge in robotics. Keypoint-based representations have been proven effective as a succinct representation for capturing essential object features, and for establishing a reference frame in action prediction, enabling data-efficient learning of robot skills. However, their manual design nature and reliance on additional human labels limit their scalability. In this paper, we propose KALM, a framework that leverages large pre-trained vision-language models (LMs) to automatically generate task-relevant and cross-instance consistent keypoints. KALM distills robust and consistent keypoints across views and objects by generating proposals using LMs and verifies them against a small set of robot demonstration data. Based on the generated keypoints, we can train keypoint-conditioned policy models that predict actions in keypoint-centric frames, enabling robots to generalize effectively across varying object poses, camera views, and object instances with similar functional shapes. Our method demonstrates strong performance in the real world, adapting to different tasks and environments from only a handful of demonstrations while requiring no additional labels. Website: https://kalm-il.github.io/

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bilevel Learning for Bilevel Planning

    cs.RO 2025-02 conditional novelty 7.0 of 10

    IVNTR learns neural predicates from demonstrations for bilevel planning and reaches about 77% success on unseen robot tasks, outperforming prior methods that stay below 35%.

  2. Robotic Manipulation Framework Based on Semantic Keypoints for Packing Shoes of Different Sizes, Shapes, and Softness

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A robotic framework using semantic keypoints plus box-edge contact packs shoe pairs from arbitrary initial states into a standard side-by-side configuration.

  3. Knowledge-Driven Imitation Learning: Enabling Generalization Across Diverse Conditions

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A semantic keypoint graph matched to novel objects lets imitation-learned manipulation policies generalize with a quarter of the demonstrations.

  4. AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A three-stage pipeline that turns keypoint tracks into discrete motion tokens, predicts them from action-free video, and decodes them into actions yields large few-shot and zero-shot policy improvements in robot manipulation.

  5. Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A three-layer hierarchical system using a VLM planner and VLM skill monitor with imitation-learned skills and an RL tracking policy achieved 73% success on a real humanoid pick-and-place task.

Pith tools