REVIEW 3 cited by
Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies
read the original abstract
Humans can imagine goal states during planning and perform actions to match those goals. In this work, we propose Imagination Policy, a novel multi-task key-frame policy network for solving high-precision pick and place tasks. Instead of learning actions directly, Imagination Policy generates point clouds to imagine desired states which are then translated to actions using rigid action estimation. This transforms action inference into a local generative task. We leverage pick and place symmetries underlying the tasks in the generation process and achieve extremely high sample efficiency and generalizability to unseen configurations. Finally, we demonstrate state-of-the-art performance across various tasks on the RLbench benchmark compared with several strong baselines and validate our approach on a real robot.
Forward citations
Cited by 3 Pith papers
-
Equivariant Volumetric Grasping
A novel tri-plane equivariant volumetric grasp model adapts GIGA and IGD planners with flow matching and deformable attention to achieve higher real-time performance than non-equivariant baselines.
-
Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation
Continuous multi-view image-space keypoint trajectories plus per-camera equivariant augmentation beat strong 3D and image baselines on MimicGen and real UR5 tasks.
-
Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification
Projecting 3D gripper keypoints onto camera pixels and classifying those pixels yields millimeter-precise, multi-modal closed-loop manipulation faster than diffusion policies.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.