Pith. sign in

REVIEW 17 cited by

FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.04382 v6 pith:RP4KKAPG submitted 2022-05-09 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords objectssystemarticulatedarticulationmotionarticulatefieldmanipulate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We explore a novel method to perceive and manipulate 3D articulated objects that generalizes to enable a robot to articulate unseen classes of objects. We propose a vision-based system that learns to predict the potential motions of the parts of a variety of articulated objects to guide downstream motion planning of the system to articulate the objects. To predict the object motions, we train a neural network to output a dense vector field representing the point-wise motion direction of the points in the point cloud under articulation. We then deploy an analytical motion planner based on this vector field to achieve a policy that yields maximum articulation. We train the vision system entirely in simulation, and we demonstrate the capability of our system to generalize to unseen object instances and novel categories in both simulation and the real world, deploying our policy on a Sawyer robot with no finetuning. Results show that our system achieves state-of-the-art performance in both simulated and real-world experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  2. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  3. 3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Dense 3D point-track prediction from unconstrained human videos plus a track-conditioned closed-loop policy yields large sample-efficiency gains over BC and video-pretraining baselines.

  4. Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A history-aware verifier that scores candidate actions using past interactions cuts failure rates in ambiguous robot manipulation tasks compared to using the generator alone.

  5. GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    GenManip is a benchmark and simulation platform with LLM-generated scene graphs for testing how robot policies generalize to new instructions, layouts, and objects.

  6. RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

    cs.RO 2025-04 conditional novelty 6.0 of 10

    RoboGround uses object and placement-area masks from a grounded vision-language model as intermediate guidance, significantly improving simulated robots' generalization to novel objects, categories, and instructions.

  7. CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World

    cs.RO 2025-02 conditional novelty 6.0 of 10

    CordViP achieves strong real-world dexterous manipulation by feeding a diffusion policy with pose-tracked 3D object models and hand point clouds, pretrained on contact maps and arm-hand coordination.

  8. Locate n' Rotate: Two-stage Openable Part Detection with Foundation Model Priors

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MOPD improves openable part detection and motion parameter prediction by fusing perceptual grouping and geometric priors from foundation models into a two-decoder transformer with a motion-aware optimal transport matc...

  9. Subspace-wise Hybrid RL for Articulated Object Manipulation

    cs.RO 2024-12 conditional novelty 6.0 of 10

    SwRL trains separate RL policies for force magnitude and redundant subspace motion, after decomposing the task space into kinematic, geometric, and redundant subspaces, improving articulated object manipulation performance.

  10. Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Lift3D uses task-aware depth reconstruction and mapped 2D positional embeddings to let pretrained 2D vision transformers act as 3D point-cloud manipulation policies, beating prior methods on average.

  11. GAPartManip: A Large-scale Part-centric Dataset for Material-Agnostic Articulated Object Manipulation

    cs.RO 2024-11 conditional novelty 6.0 of 10

    A new synthetic dataset with material-randomized stereo images and part-level action poses improves depth estimation and articulated object manipulation in simulation and real-world tests.

  12. Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    EgoMono4D estimates depth, camera intrinsics and poses from unlabeled egocentric videos in a single feed-forward pass, reconstructing dense per-frame point clouds better than baseline methods on in-domain and zero-sho...

  13. KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    KAI, a keypoint-and-displacement intermediate with geometric joint priors, matches or beats articulated-manipulation baselines at half the demo data and supports human-video co-training.

  14. Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding

    cs.RO 2025-07 conditional novelty 5.0 of 10

    AdaRPG uses GPT-4o, GroundingDINO, and SAM to locate and segment the movable part, a part-affordance model to choose a grasp, and GPT-4o to write the control loop, outperforming prior methods on new articulated objects.

  15. CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation

    cs.RO 2025-05 conditional novelty 5.0 of 10

    CrayonRobo trains a vision-language-action model to read colored 2D prompt overlays (contact point, end-effector axes, movement direction) and output SE(3) contact poses, enabling step-by-step and long-horizon robotic...

  16. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

  17. CUPS: Improving Human Pose-Shape Estimators with Conformalized Deep Uncertainty

    cs.CV 2024-12 conditional novelty 4.0 of 10

    CUPS learns a deep uncertainty score end-to-end with a video-based SMPL reconstructor and uses it as a conformal score to build calibrated prediction sets despite non-exchangeable video data.

Pith tools