REVIEW 7 cited by
PolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The ability for robots to comprehend and execute manipulation tasks based on natural language instructions is a long-term goal in robotics. The dominant approaches for language-guided manipulation use 2D image representations, which face difficulties in combining multi-view cameras and inferring precise 3D positions and relationships. To address these limitations, we propose a 3D point cloud based policy called PolarNet for language-guided manipulation. It leverages carefully designed point cloud inputs, efficient point cloud encoders, and multimodal transformers to learn 3D point cloud representations and integrate them with language instructions for action prediction. PolarNet is shown to be effective and data efficient in a variety of experiments conducted on the RLBench benchmark. It outperforms state-of-the-art 2D and 3D approaches in both single-task and multi-task learning. It also achieves promising results on a real robot.
Forward citations
Cited by 7 Pith papers
-
ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
A hierarchical leader-follower Gaussian world model improves multi-task bimanual manipulation success rates over prior single-arm-based methods.
-
Mini Diffuser: Fast Multi-task Diffusion Policy Training Using Two-level Mini-batches
Mini Diffuser trains multi-task robotic diffusion policies by reusing each vision-language condition for many noised action samples, reaching 77.6% RLBench success with 4.8% of the training time and 6.6% of the memory...
-
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
SAM2Act reports 86.8% average success across 18 RLBench tasks, and the memory variant SAM2Act+ reaches 94.3% on the new MemoryBench tasks.
-
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
Lift3D uses task-aware depth reconstruction and mapped 2D positional embeddings to let pretrained 2D vision transformers act as 3D point-cloud manipulation policies, beating prior methods on average.
-
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
Fusing RGB and point-cloud inputs with training-time modality dropout plus cross-attention makes a diffusion visuomotor policy markedly more robust to visual and spatial shifts than unimodal or naively fused baselines.
-
Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation
A robot framework combining GPT-4V planning with a 3D feature-field skill policy improves long-horizon kitchen manipulation accuracy over LLM baselines, according to small real-robot trials.
-
Leveraging OS-Level Primitives for Robotic Action Management
Applying OS-style exception handling, context caching, and replay to robotic action slices raises success rates 7x to 24x and cuts execution steps up to 74% for repetitive manipulation tasks, without retraining the VLA model.
Discussion (0). Continue with ORCID to comment.