Pith. sign in

REVIEW 7 cited by

PolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15596 v1 pith:AX6BYDSZ submitted 2023-09-27 cs.RO cs.CV

classification cs.ROcs.CV
keywords pointcloudmanipulationlanguage-guidedpolarnetapproachesefficientinstructions
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

0 comments
read the original abstract

The ability for robots to comprehend and execute manipulation tasks based on natural language instructions is a long-term goal in robotics. The dominant approaches for language-guided manipulation use 2D image representations, which face difficulties in combining multi-view cameras and inferring precise 3D positions and relationships. To address these limitations, we propose a 3D point cloud based policy called PolarNet for language-guided manipulation. It leverages carefully designed point cloud inputs, efficient point cloud encoders, and multimodal transformers to learn 3D point cloud representations and integrate them with language instructions for action prediction. PolarNet is shown to be effective and data efficient in a variety of experiments conducted on the RLBench benchmark. It outperforms state-of-the-art 2D and 3D approaches in both single-task and multi-task learning. It also achieves promising results on a real robot.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A hierarchical leader-follower Gaussian world model improves multi-task bimanual manipulation success rates over prior single-arm-based methods.

  2. Mini Diffuser: Fast Multi-task Diffusion Policy Training Using Two-level Mini-batches

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Mini Diffuser trains multi-task robotic diffusion policies by reusing each vision-language condition for many noised action samples, reaching 77.6% RLBench success with 4.8% of the training time and 6.6% of the memory...

  3. SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation

    cs.RO 2025-01 conditional novelty 6.0 of 10

    SAM2Act reports 86.8% average success across 18 RLBench tasks, and the memory variant SAM2Act+ reaches 94.3% on the new MemoryBench tasks.

  4. Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Lift3D uses task-aware depth reconstruction and mapped 2D positional embeddings to let pretrained 2D vision transformers act as 3D point-cloud manipulation policies, beating prior methods on average.

  5. DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization

    cs.RO 2025-11 conditional novelty 5.0 of 10

    Fusing RGB and point-cloud inputs with training-time modality dropout plus cross-attention makes a diffusion visuomotor policy markedly more robust to visual and spatial shifts than unimodal or naively fused baselines.

  6. Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation

    cs.RO 2025-01 conditional novelty 5.0 of 10

    A robot framework combining GPT-4V planning with a 3D feature-field skill policy improves long-horizon kitchen manipulation accuracy over LLM baselines, according to small real-robot trials.

  7. Leveraging OS-Level Primitives for Robotic Action Management

    cs.OS 2025-08 conditional novelty 4.0 of 10

    Applying OS-style exception handling, context caching, and replay to robotic action slices raises success rates 7x to 24x and cuts execution steps up to 74% for repetitive manipulation tasks, without retraining the VLA model.

Pith tools