REVIEW 3 cited by
SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce SPOT, an object-centric imitation learning framework. The key idea is to capture each task by an object-centric representation, specifically the SE(3) object pose trajectory relative to the target. This approach decouples embodiment actions from sensory inputs, facilitating learning from various demonstration types, including both action-based and action-less human hand demonstrations, as well as cross-embodiment generalization. Additionally, object pose trajectories inherently capture planning constraints from demonstrations without the need for manually-crafted rules. To guide the robot in executing the task, the object trajectory is used to condition a diffusion policy. We systematically evaluate our method on simulation and real-world tasks. In real-world evaluation, using only eight demonstrations shot on an iPhone, our approach completed all tasks while fully complying with task constraints. Project page: https://nvlabs.github.io/object_centric_diffusion
Forward citations
Cited by 3 Pith papers
-
Knowledge-Driven Imitation Learning: Enabling Generalization Across Diverse Conditions
A semantic keypoint graph matched to novel objects lets imitation-learned manipulation policies generalize with a quarter of the demonstrations.
-
ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models
ControlVLA adapts a DROID-pretrained diffusion VLA policy to new manipulation tasks with 10 to 20 demos by injecting object-centric features through zero-initialized cross-attention layers, achieving 76.7% success acr...
-
Object-centric 3D Motion Field for Robot Learning from Human Videos
A policy trained only on human RGBD videos, with a denoised object-centric 3D motion field as action representation, achieves about 55% average success on five real manipulation tasks where prior flow-based methods st...
Discussion (0). Sign in to comment.