Pith. sign in

REVIEW 9 cited by

RVT: Robotic View Transformer for 3D Object Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.14896 v1 pith:75V4SDB2 submitted 2023-06-26 cs.RO cs.CV

classification cs.ROcs.CV
keywords manipulationperactachievingacrosscameraexplicitmodelobject
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

For 3D object manipulation, methods that build an explicit 3D representation perform better than those relying only on camera images. But using explicit 3D representations like voxels comes at large computing cost, adversely affecting scalability. In this work, we propose RVT, a multi-view transformer for 3D manipulation that is both scalable and accurate. Some key features of RVT are an attention mechanism to aggregate information across views and re-rendering of the camera input from virtual views around the robot workspace. In simulations, we find that a single RVT model works well across 18 RLBench tasks with 249 task variations, achieving 26% higher relative success than the existing state-of-the-art method (PerAct). It also trains 36X faster than PerAct for achieving the same performance and achieves 2.3X the inference speed of PerAct. Further, RVT can perform a variety of manipulation tasks in the real world with just a few ($\sim$10) demonstrations per task. Visual results, code, and trained model are provided at https://robotic-view-transformer.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A preprocessing pipeline that reconstructs a 3D point cloud from RGB images and re-renders it from a fixed viewpoint improves viewpoint robustness and data efficiency for vision-based robot policies.

  2. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  3. DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A modular benchmark of 100 dexterous manipulation tasks across 3 arms and 6 hands with 3,180 demonstrations reveals that current policies (Diffusion Policy, DP3, OpenVLA, π0.5) achieve only 34% mean success, exposing ...

  4. Geometry-Aware Motion Latents for Learning Robust Manipulation Policies

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Predicting future 3D pointmaps forces discrete motion latents to encode physical geometric transformations, improving single-view robot manipulation over 2D/static-3D baselines.

  5. ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A hierarchical leader-follower Gaussian world model improves multi-task bimanual manipulation success rates over prior single-arm-based methods.

  6. CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity

    cs.RO 2025-06 conditional novelty 6.0 of 10

    CodeDiffuser uses vision-language-model-generated code to build 3D attention maps that condition a diffusion policy, improving success on ambiguous language manipulation tasks compared with end-to-end baselines.

  7. MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation

    cs.RO 2024-11 conditional novelty 6.0 of 10

    MALMM, a three-agent LLM framework with per-step environment feedback, achieves 81% average success on nine RLBench tasks versus 50% for a single-agent LLM baseline, in zero-shot settings.

  8. QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

    cs.CV 2025-10 conditional novelty 5.0 of 10

    Adding an auxiliary quantized-depth-token prediction task to a VLA policy improves manipulation success rates on LIBERO, Simpler, and real-robot pick-and-place tasks versus the open-pi-zero baseline.

  9. Advances in Transformers for Robotic Applications: A Review

    cs.RO 2024-12 unverdicted

    A survey of Transformer applications in robotic perception, planning, control, human-robot interaction, and reinforcement learning.

Pith tools