REVIEW 17 cited by
FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We explore a novel method to perceive and manipulate 3D articulated objects that generalizes to enable a robot to articulate unseen classes of objects. We propose a vision-based system that learns to predict the potential motions of the parts of a variety of articulated objects to guide downstream motion planning of the system to articulate the objects. To predict the object motions, we train a neural network to output a dense vector field representing the point-wise motion direction of the points in the point cloud under articulation. We then deploy an analytical motion planner based on this vector field to achieve a policy that yields maximum articulation. We train the vision system entirely in simulation, and we demonstrate the capability of our system to generalize to unseen object instances and novel categories in both simulation and the real world, deploying our policy on a Sawyer robot with no finetuning. Results show that our system achieves state-of-the-art performance in both simulated and real-world experiments.
Forward citations
Cited by 17 Pith papers
-
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.
-
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.
-
3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos
Dense 3D point-track prediction from unconstrained human videos plus a track-conditioned closed-loop policy yields large sample-efficiency gains over BC and video-pretraining baselines.
-
Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online
A history-aware verifier that scores candidate actions using past interactions cuts failure rates in ambiguous robot manipulation tasks compared to using the generator alone.
-
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
GenManip is a benchmark and simulation platform with LLM-generated scene graphs for testing how robot policies generalize to new instructions, layouts, and objects.
-
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
RoboGround uses object and placement-area masks from a grounded vision-language model as intermediate guidance, significantly improving simulated robots' generalization to novel objects, categories, and instructions.
-
CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World
CordViP achieves strong real-world dexterous manipulation by feeding a diffusion policy with pose-tracked 3D object models and hand point clouds, pretrained on contact maps and arm-hand coordination.
-
Locate n' Rotate: Two-stage Openable Part Detection with Foundation Model Priors
MOPD improves openable part detection and motion parameter prediction by fusing perceptual grouping and geometric priors from foundation models into a two-decoder transformer with a motion-aware optimal transport matc...
-
Subspace-wise Hybrid RL for Articulated Object Manipulation
SwRL trains separate RL policies for force magnitude and redundant subspace motion, after decomposing the task space into kinematic, geometric, and redundant subspaces, improving articulated object manipulation performance.
-
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
Lift3D uses task-aware depth reconstruction and mapped 2D positional embeddings to let pretrained 2D vision transformers act as 3D point-cloud manipulation policies, beating prior methods on average.
-
GAPartManip: A Large-scale Part-centric Dataset for Material-Agnostic Articulated Object Manipulation
A new synthetic dataset with material-randomized stereo images and part-level action poses improves depth estimation and articulated object manipulation in simulation and real-world tests.
-
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos
EgoMono4D estimates depth, camera intrinsics and poses from unlabeled egocentric videos in a single feed-forward pass, reconstructing dense per-frame point clouds better than baseline methods on in-domain and zero-sho...
-
KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
KAI, a keypoint-and-displacement intermediate with geometric joint priors, matches or beats articulated-manipulation baselines at half the demo data and supports human-video co-training.
-
Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding
AdaRPG uses GPT-4o, GroundingDINO, and SAM to locate and segment the movable part, a part-affordance model to choose a grasp, and GPT-4o to write the control loop, outperforming prior methods on new articulated objects.
-
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
CrayonRobo trains a vision-language-action model to read colored 2D prompt overlays (contact point, end-effector axes, movement direction) and output SE(3) contact poses, enabling step-by-step and long-horizon robotic...
-
Advances in 4D Representation: Geometry, Motion, and Interaction
A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.
-
CUPS: Improving Human Pose-Shape Estimators with Conformalized Deep Uncertainty
CUPS learns a deep uncertainty score end-to-end with a video-based SMPL reconstructor and uses it as a conformal score to build calibrated prediction sets despite non-exchangeable video data.
Discussion (0). Continue with ORCID to comment.