Pith. sign in

REVIEW 11 cited by

3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.11002 v2 pith:UUAARPAR submitted 2025-08-14 cs.RO

3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation

classification cs.RO
keywords policyactionactordiffusion-basedflowflowmatchlearningmanipulation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We present 3D FlowMatch Actor (3DFA), a 3D policy architecture for robot manipulation that combines flow matching for trajectory prediction with 3D pretrained visual scene representations for learning from demonstration. 3DFA leverages 3D relative attention between action and visual tokens during action denoising, building on prior work in 3D diffusion-based single-arm policy learning. Through a combination of flow matching and targeted system-level and architectural optimizations, 3DFA achieves over 30x faster training and inference than previous 3D diffusion-based policies, without sacrificing performance. On the bimanual PerAct2 benchmark, it establishes a new state of the art, outperforming the next-best method by an absolute margin of 41.4%. In extensive real-world evaluations, it surpasses strong baselines with up to 1000x more parameters and significantly more pretraining. In unimanual settings, it sets a new state of the art on 74 RLBench tasks by directly predicting dense end-effector trajectories, eliminating the need for motion planning. Comprehensive ablation studies underscore the importance of our design choices for both policy effectiveness and efficiency.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Instant-Fold: In-Context Imitation Learning for Deformable Object Manipulation

    cs.RO 2026-06 unverdicted novelty 7.0

    Instant-Fold enables execution of multiple deformable object manipulation modes from a single demonstration via a flow-matching transformer policy that transfers zero-shot from simulation to real robots.

  2. Learning 3D Affordances for Blade Insertion in Cluttered Stowing

    cs.CV 2026-06 conditional novelty 6.5

    VulcanVoxel reconstructs blade occupancy with a 3D masked autoencoder, recovering multi-modal free-space affordances from unimodal warehouse stow data and raising top-5 coverage from 0.71 to 0.89.

  3. High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching

    cs.RO 2026-07 conditional novelty 6.0

    One-step flow-matching visuomotor policy with recursive correction, dual-timestep spectral consistency, and contrastive mode separation matches or exceeds 10-step baselines at 1 NFE.

  4. ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning

    cs.RO 2026-06 conditional novelty 6.0

    Co-training a diffusion visuomotor policy on unified past–current–future sparse 3D object–gripper keypoint flows improves long-horizon and non-Markovian manipulation over action-only and future-only baselines.

  5. ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning

    cs.RO 2026-06 conditional novelty 6.0

    ChronoFlow-Policy improves visuomotor policy learning by co-training action generation with prediction of past-current-future gripper-object keypoint trajectories.

  6. 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

    cs.RO 2026-06 unverdicted novelty 6.0

    3D HAMSTER adds depth encoding and reconstruction to VLMs to produce 3D waypoint sequences that feed directly into pointcloud policies, claiming better generalization than 2D baselines under shifts.

  7. 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

    cs.RO 2026-06 unverdicted novelty 6.0

    3D HAMSTER augments a VLM with depth encoding and reconstruction to predict 3D waypoints that directly guide pointcloud policies, outperforming 2D baselines especially under distribution shifts.

  8. MV-Actor: Aligning Multi-View Semantics and Spatial Awareness for Bimanual Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    MV-Actor proposes a multi-view framework using semantic interaction, semantic-spatial token interaction, and guided depth repair to reach 87.8% success on the PerAct2 bimanual benchmark and outperform RGB/RGB-D baseli...

  9. PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

    cs.RO 2026-07 unverdicted novelty 5.0

    PhysMani couples a physics-principled 3D Gaussian world model with a future-aware policy to achieve higher success rates on dynamic manipulation tasks in simulation and real robots.

  10. R3D: Revisiting 3D Policy Learning

    cs.CV 2026-04 unverdicted novelty 5.0

    A transformer 3D encoder plus diffusion decoder architecture, with 3D-specific augmentations, outperforms prior 3D policy methods on manipulation benchmarks by improving training stability.

  11. ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning

    cs.RO 2026-06 unverdicted novelty 4.0

    ChronoFlow-Policy uses a unified ChronoFlow representation of past-current-future dynamics learned jointly with actions in a diffusion policy, outperforming baselines on 14 simulated and 5 real manipulation tasks.