Pith. sign in

REVIEW 3 cited by

Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11740 v2 pith:BNZUBO6R submitted 2024-06-17 cs.RO cs.AIcs.LG

Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies

classification cs.RO cs.AIcs.LG
keywords policyactionsimaginationtasksactiongenerativeimaginelearning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Humans can imagine goal states during planning and perform actions to match those goals. In this work, we propose Imagination Policy, a novel multi-task key-frame policy network for solving high-precision pick and place tasks. Instead of learning actions directly, Imagination Policy generates point clouds to imagine desired states which are then translated to actions using rigid action estimation. This transforms action inference into a local generative task. We leverage pick and place symmetries underlying the tasks in the generation process and achieve extremely high sample efficiency and generalizability to unseen configurations. Finally, we demonstrate state-of-the-art performance across various tasks on the RLbench benchmark compared with several strong baselines and validate our approach on a real robot.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Equivariant Volumetric Grasping

    cs.RO 2025-07 unverdicted novelty 7.0

    A novel tri-plane equivariant volumetric grasp model adapts GIGA and IGD planners with flow matching and deformable attention to achieve higher real-time performance than non-equivariant baselines.

  2. Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation

    cs.RO 2026-07 conditional novelty 6.0

    Continuous multi-view image-space keypoint trajectories plus per-camera equivariant augmentation beat strong 3D and image baselines on MimicGen and real UR5 tasks.

  3. Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification

    cs.RO 2026-07 conditional novelty 6.0

    Projecting 3D gripper keypoints onto camera pixels and classifying those pixels yields millimeter-precise, multi-modal closed-loop manipulation faster than diffusion policies.