Pith. sign in

REVIEW 1 cited by

Learning Robotic Manipulation through Visual Planning and Acting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.04411 v1 pith:3AKLHX6Y submitted 2019-05-11 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords manipulationmodelplanningvisualdataimaginelearnlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Planning for robotic manipulation requires reasoning about the changes a robot can affect on objects. When such interactions can be modelled analytically, as in domains with rigid objects, efficient planning algorithms exist. However, in both domestic and industrial domains, the objects of interest can be soft, or deformable, and hard to model analytically. For such cases, we posit that a data-driven modelling approach is more suitable. In recent years, progress in deep generative models has produced methods that learn to `imagine' plausible images from data. Building on the recent Causal InfoGAN generative model, in this work we learn to imagine goal-directed object manipulation directly from raw image data of self-supervised interaction of the robot with the object. After learning, given a goal observation of the system, our model can generate an imagined plan -- a sequence of images that transition the object into the desired goal. To execute the plan, we use it as a reference trajectory to track with a visual servoing controller, which we also learn from the data as an inverse dynamics model. In a simulated manipulation task, we show that separating the problem into visual planning and visual tracking control is more sample efficient and more interpretable than alternative data-driven approaches. We further demonstrate our approach on learning to imagine and execute in 3 environments, the final of which is deformable rope manipulation on a PR2 robot.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding Physical Properties of Unseen Deformable Objects by Leveraging Large Language Models and Robot Actions

    cs.RO 2025-06 conditional novelty 5.0 of 10

    Using robot actions and LLM visual reasoning, the system identifies deformability properties of unseen objects with up to 78.57% accuracy, which helps plan bin-packing at over 96% success after replanning.

Pith tools