Pith. sign in

REVIEW 8 cited by

Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.07931 v2 pith:ZJQTJ6DG submitted 2023-07-27 cs.CV cs.AIcs.CLcs.LGcs.RO

Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation

classification cs.CV cs.AIcs.CLcs.LGcs.RO
keywords distilledmanipulationobjectsfeaturefeaturesfew-shotfieldsgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often lacking in 2D image features. This work bridges this 2D-to-3D gap for robotic manipulation by leveraging distilled feature fields to combine accurate 3D geometry with rich semantics from 2D foundation models. We present a few-shot learning method for 6-DOF grasping and placing that harnesses these strong spatial and semantic priors to achieve in-the-wild generalization to unseen objects. Using features distilled from a vision-language model, CLIP, we present a way to designate novel objects for manipulation via free-text natural language, and demonstrate its ability to generalize to unseen expressions and novel categories of objects.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond World-Frame Action Heads: Motion-Centric Action Frames for Vision-Language-Action Models

    cs.AI 2026-05 unverdicted novelty 7.0

    MCF-Proto adds a motion-centric local action frame and prototype parameterization to VLA models, inducing emergent geometric structure and improved robustness from standard demonstrations alone.

  2. Action Images: End-to-End Policy Learning via Multiview Video Generation

    cs.CV 2026-04 unverdicted novelty 7.0

    Action Images turn robot arm motions into interpretable multiview pixel videos, letting video backbones serve as zero-shot policies for end-to-end robot learning.

  3. Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations

    cs.CV 2026-04 unverdicted novelty 6.0

    RecGen achieves state-of-the-art 3D multi-object scene reconstruction from sparse RGB-D views by combining compositional synthetic scene generation with strong 3D shape priors, outperforming SAM3D by 30%+ in shape qua...

  4. C3G: Learning Compact 3D Representations with 2K Gaussians

    cs.CV 2025-12 unverdicted novelty 6.0

    C3G creates compact 3D Gaussian representations with 2K points by guiding placement via learnable tokens that aggregate multi-view features through attention, yielding better efficiency and performance than dense methods.

  5. 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

    cs.RO 2024-03 unverdicted novelty 6.0

    DP3 uses compact 3D representations from sparse point clouds inside diffusion policies to learn generalizable visuomotor skills from few demonstrations, reporting 24% gains in simulation and 85% success on real robots.

  6. Beyond Point-Attached Semantics: Object-Centric Semantic Fields for Generalizable Manipulation

    cs.RO 2026-07 conditional novelty 5.0

    An object-conditioned continuous semantic field queried at explicit 3D locations yields more stable part cues and higher manipulation success than point-attached 2D/3D features.

  7. PhysGraph: A Physics-aware 3D Scene Graph for Perception and Reasoning

    cs.RO 2026-06 unverdicted novelty 5.0

    PhysGraph reconstructs object-centric 3D geometry from RGB-D, decomposes objects into parts, infers materials and articulations via visual reasoning, and reports SOTA results on semantic segmentation, multi-object mas...

  8. Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery

    cs.RO 2026-03 unverdicted novelty 3.0

    A hybrid visual-motor imagery EEG decoder controls a robot for grasping and placement at 40% and 63% accuracy respectively, yielding 21% end-to-end task success in cue-free online use.