Pith. sign in

REVIEW 5 cited by

3D-MVP: 3D Multiview Pretraining for Robotic Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.18158 v2 pith:2BSMLSOV submitted 2024-06-26 cs.RO cs.CV

classification cs.ROcs.CV
keywords pretrainingd-mvpmanipulationmaskedmulti-viewrobotictransformervisual
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent works have shown that visual pretraining on egocentric datasets using masked autoencoders (MAE) can improve generalization for downstream robotics tasks. However, these approaches pretrain only on 2D images, while many robotics applications require 3D scene understanding. In this work, we propose 3D-MVP, a novel approach for 3D Multi-View Pretraining using masked autoencoders. We leverage Robotic View Transformer (RVT), which uses a multi-view transformer to understand the 3D scene and predict gripper pose actions. We split RVT's multi-view transformer into visual encoder and action decoder, and pretrain its visual encoder using masked autoencoding on large-scale 3D datasets such as Objaverse. We evaluate 3D-MVP on a suite of virtual robot manipulation tasks and demonstrate improved performance over baselines. Our results suggest that 3D-aware pretraining is a promising approach to improve generalization of vision-based robotic manipulation policies. Project site: https://jasonqsy.github.io/3DMVP

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-supervised Learning Of Visual Pose Estimation Without Pose Labels By Classifying LED States

    cs.RO 2025-09 conditional novelty 7.0 of 10

    A self-supervised training scheme where a network estimates a robot's relative pose by predicting the on/off states of its LEDs, needing no pose labels or CAD model.

  2. Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation

    cs.LG 2025-05 conditional novelty 6.0 of 10

    By summing multi-view features and adding single-view features as actor-critic augmentations, MAD produces manipulation policies that learn faster and tolerate missing cameras in simulation.

  3. Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Lift3D uses task-aware depth reconstruction and mapped 2D positional embeddings to let pretrained 2D vision transformers act as 3D point-cloud manipulation policies, beating prior methods on average.

  4. MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A two-stage language-driven grasping system that pools visual features inside a predicted object mask improves grasp accuracy and training efficiency versus CLIP baselines, supported by a new 219M-grasp dataset.

  5. NVSPolicy: Adaptive Novel-View Synthesis for Generalizable Language-Conditioned Policy Learning

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A hierarchical language-conditioned policy that synthesizes an adaptively selected novel view reaches 90.4% single-task success and 2.93 average consecutive-task length on CALVIN, ahead of prior non-foundation-model methods.

Pith tools