REVIEW 5 cited by
RISE: 3D Perception Makes Real-World Robot Imitation Simple and Effective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Precise robot manipulations require rich spatial information in imitation learning. Image-based policies model object positions from fixed cameras, which are sensitive to camera view changes. Policies utilizing 3D point clouds usually predict keyframes rather than continuous actions, posing difficulty in dynamic and contact-rich scenarios. To utilize 3D perception efficiently, we present RISE, an end-to-end baseline for real-world imitation learning, which predicts continuous actions directly from single-view point clouds. It compresses the point cloud to tokens with a sparse 3D encoder. After adding sparse positional encoding, the tokens are featurized using a transformer. Finally, the features are decoded into robot actions by a diffusion head. Trained with 50 demonstrations for each real-world task, RISE surpasses currently representative 2D and 3D policies by a large margin, showcasing significant advantages in both accuracy and efficiency. Experiments also demonstrate that RISE is more general and robust to environmental change compared with previous baselines. Project website: rise-policy.github.io.
Forward citations
Cited by 5 Pith papers
-
Asynchronous Multimodal Diffusion Policy Composition via Latency-Aware Guidance Fusion
LAG-Fusion composes asynchronous diffusion policies by rebasing delayed guidance into the current action frame and fusing with latency-aware weights, improving contact-rich manipulation performance.
-
Concurrent Prehensile and Nonprehensile Manipulation: A Practical Approach to Multi-Stage Dexterous Tasks
Object-centric skill decomposition with retrieve-align-execute achieves 66% average success on three dexterous two-object tasks using 3–4 demonstrations per object.
-
SE(3)-Equivariant Diffusion Policy in Spherical Fourier Space
Continuous SE(3) equivariance is embedded in the policy by representing states, actions, and denoising steps in spherical Fourier space, improving generalization to novel 3D arrangements.
-
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
Fusing RGB and point-cloud inputs with training-time modality dropout plus cross-attention makes a diffusion visuomotor policy markedly more robust to visual and spatial shifts than unimodal or naively fused baselines.
-
4D Visual Pre-training for Robot Learning
A next-frame point-cloud diffusion pre-training method (FVP) improves DP3 and RDT-1B manipulation success rates on the paper's own tasks.
Discussion (0). Sign in to comment.