REVIEW 6 cited by
CoPa: General Robotic Manipulation through Spatial Constraints of Parts with Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Foundation models pre-trained on web-scale data are shown to encapsulate extensive world knowledge beneficial for robotic manipulation in the form of task planning. However, the actual physical implementation of these plans often relies on task-specific learning methods, which require significant data collection and struggle with generalizability. In this work, we introduce Robotic Manipulation through Spatial Constraints of Parts (CoPa), a novel framework that leverages the common sense knowledge embedded within foundation models to generate a sequence of 6-DoF end-effector poses for open-world robotic manipulation. Specifically, we decompose the manipulation process into two phases: task-oriented grasping and task-aware motion planning. In the task-oriented grasping phase, we employ foundation vision-language models (VLMs) to select the object's grasping part through a novel coarse-to-fine grounding mechanism. During the task-aware motion planning phase, VLMs are utilized again to identify the spatial geometry constraints of task-relevant object parts, which are then used to derive post-grasp poses. We also demonstrate how CoPa can be seamlessly integrated with existing robotic planning algorithms to accomplish complex, long-horizon tasks. Our comprehensive real-world experiments show that CoPa possesses a fine-grained physical understanding of scenes, capable of handling open-set instructions and objects with minimal prompt engineering and without additional training. Project page: https://copa-2024.github.io/
Forward citations
Cited by 6 Pith papers
-
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.
-
RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
A multi-domain affordance benchmark with 273k images and 26k reasoning instructions is introduced, together with a VLM-based grasping pipeline that shows strong zero-shot affordance segmentation and real-robot performance.
-
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
GenManip is a benchmark and simulation platform with LLM-generated scene graphs for testing how robot policies generalize to new instructions, layouts, and objects.
-
CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation
A first benchmark that tests whether robots can read multi-page appliance manuals and then plan and execute manipulation tasks on appliances in simulation.
-
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
UAD distills affordance knowledge from vision-language models and DINOv2 features into a lightweight task-conditioned model that predicts pixel-level manipulation regions and improves few-shot imitation learning gener...
- CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity
Discussion (0). Sign in to comment.