Pith. sign in

REVIEW 13 cited by

Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18121 v1 pith:S5XLA6MA submitted 2024-09-26 cs.RO cs.CV

classification cs.ROcs.CV
keywords robotobjectpartrsrdmonocularmotionsingletrajectories
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans can learn to manipulate new objects by simply watching others; providing robots with the ability to learn from such demonstrations would enable a natural interface specifying new behaviors. This work develops Robot See Robot Do (RSRD), a method for imitating articulated object manipulation from a single monocular RGB human demonstration given a single static multi-view object scan. We first propose 4D Differentiable Part Models (4D-DPM), a method for recovering 3D part motion from a monocular video with differentiable rendering. This analysis-by-synthesis approach uses part-centric feature fields in an iterative optimization which enables the use of geometric regularizers to recover 3D motions from only a single video. Given this 4D reconstruction, the robot replicates object trajectories by planning bimanual arm motions that induce the demonstrated object part motion. By representing demonstrations as part-centric trajectories, RSRD focuses on replicating the demonstration's intended behavior while considering the robot's own morphological limits, rather than attempting to reproduce the hand's motion. We evaluate 4D-DPM's 3D tracking accuracy on ground truth annotated 3D part trajectories and RSRD's physical execution performance on 9 objects across 10 trials each on a bimanual YuMi robot. Each phase of RSRD achieves an average of 87% success rate, for a total end-to-end success rate of 60% across 90 trials. Notably, this is accomplished using only feature fields distilled from large pretrained vision models -- without any task-specific training, fine-tuning, dataset collection, or annotation. Project page: https://robot-see-robot-do.github.io

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction

    cs.GR 2026-07 conditional novelty 6.0 of 10

    OBB-based part-fitting and contact losses on 3D Gaussians disentangle geometry, appearance, and motion for cleaner articulated reconstruction than photometric-only baselines.

  2. Robots Acquire Manipulation Skills in Seconds from a Single Human Video

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A single human video is enough for a robot to acquire a new manipulation skill at inference time, with no parameter updates, reaching 62% success on 50 novel tasks.

  3. Revisiting Articulated Parts Perception in Robot Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    Proposes GPS representation for articulated parts, uses VR to annotate 41K frames across 234 objects, trains an RGB-D model, and achieves 73% success in heuristic manipulation policies on 9 objects.

  4. NeuROK: Generative 4D Neural Object Kinematics

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    NeuROK learns a data-driven latent kinematic parameterization on a large 4D dataset to generate realistic object deformations by simulating dynamics only in low-dimensional latent space via Lagrangian mechanics.

  5. From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data

    cs.RO 2026-04 accept novelty 6.0 of 10

    Video-to-robot control methods cluster into three interface families, and the field’s main bottleneck is grounding video-derived predictions into dependable closed-loop robot behavior.

  6. PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations

    cs.RO 2026-02 conditional novelty 6.0 of 10

    PokeNet estimates joint types, axes, ranges, and operation order of articulated objects directly from a single-view point cloud video of a human demonstration.

  7. R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation

    cs.RO 2025-10 unverdicted novelty 6.0 of 10

    R2RGen introduces a simulator-free three-stage pipeline that parses, augments, and post-processes real pointcloud observation-action pairs to improve spatial generalization in robotic manipulation policies.

  8. One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.

  9. Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A robot policy trained on one real demonstration plus AI-generated 3D views succeeds from novel initial poses, including opposite-side starts, across six real manipulation tasks.

  10. Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

    cs.RO 2025-07 unverdicted novelty 6.0 of 10

    RIGVid shows that filtered AI-generated videos can serve as effective supervision for complex robotic manipulation tasks without any real demonstrations.

  11. ODeform: Learning Continuous 4D Motion for Shape Deformation with Neural ODEs

    cs.CV 2026-07 conditional novelty 5.0 of 10

    ODeform combines two parallel neural ODEs, one for rigid motion and one for local deformation, to predict arbitrary-time 3D point-cloud deformation from an initial state and physical parameters, outperforming simpler ...

  12. Artiverse: A Diverse and Physically Grounded Dataset for Articulated Objects

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    Artiverse is a new dataset of 5.4K human-authored articulated 3D objects with detailed annotations for parts, multi-DoF joints, interior structures, and physical attributes to enable functional modeling and physics-ba...

  13. From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data

    cs.RO 2026-04 accept novelty 5.0 of 10

    A survey introduces an interface-centric taxonomy for video-to-control methods in robotic manipulation and identifies the robotics integration layer as the central open challenge.

Pith tools