REVIEW 5 cited by
Learning Robotic Manipulation Policies from Point Clouds with Conditional Flow Matching
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learning from expert demonstrations is a promising approach for training robotic manipulation policies from limited data. However, imitation learning algorithms require a number of design choices ranging from the input modality, training objective, and 6-DoF end-effector pose representation. Diffusion-based methods have gained popularity as they enable predicting long-horizon trajectories and handle multimodal action distributions. Recently, Conditional Flow Matching (CFM) (or Rectified Flow) has been proposed as a more flexible generalization of diffusion models. In this paper, we investigate the application of CFM in the context of robotic policy learning and specifically study the interplay with the other design choices required to build an imitation learning algorithm. We show that CFM gives the best performance when combined with point cloud input observations. Additionally, we study the feasibility of a CFM formulation on the SO(3) manifold and evaluate its suitability with a simplified example. We perform extensive experiments on RLBench which demonstrate that our proposed PointFlowMatch approach achieves a state-of-the-art average success rate of 67.8% over eight tasks, double the performance of the next best method.
Forward citations
Cited by 5 Pith papers
-
High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching
One-step flow-matching visuomotor policy with recursive correction, dual-timestep spectral consistency, and contrastive mode separation matches or exceeds 10-step baselines at 1 NFE.
-
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
Gondola generates multi-view segmentation-mask-grounded next-step plans for robotic manipulation and reports improved generalization on the GemBench benchmark over a prior LLM-based planner.
-
Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation
A force-guided attention module and future-force prediction auxiliary task improve visuo-tactile fusion for dexterous manipulation, reaching 93% average success in real robot trials.
-
Designing for Difference: How Human Characteristics Shape Perceptions of Collaborative Robots
In an online video study, people rated antisocial robot behavior as least acceptable, preferred handover over table placement, and judged collaborations with older adults more sensitively.
-
mimic-one: a Scalable Model Recipe for General Purpose Robot Dexterity
mimic-one reports up to 93.3% out-of-distribution success on three real-world dexterous tasks using a diffusion policy, a custom 16-DoF hand, and a teleoperation data-collection recipe with self-correction trajectories.
Discussion (0). Continue with ORCID to comment.