Pith. sign in

REVIEW 17 cited by

CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.05663 v3 pith:DQ24J3AP submitted 2022-10-11 cs.RO cs.CV

classification cs.ROcs.CV
keywords clip-fieldssemanticidentificationinstancemappingmemoryonlyscene
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose CLIP-Fields, an implicit scene model that can be used for a variety of tasks, such as segmentation, instance identification, semantic search over space, and view localization. CLIP-Fields learns a mapping from spatial locations to semantic embedding vectors. Importantly, we show that this mapping can be trained with supervision coming only from web-image and web-text trained models such as CLIP, Detic, and Sentence-BERT; and thus uses no direct human supervision. When compared to baselines like Mask-RCNN, our method outperforms on few-shot instance identification or semantic segmentation on the HM3D dataset with only a fraction of the examples. Finally, we show that using CLIP-Fields as a scene memory, robots can perform semantic navigation in real-world environments. Our code and demonstration videos are available here: https://mahis.life/clip-fields

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A 3D Gaussian Splatting SLAM system learns compact 16-dim language features per Gaussian, enabling real-time open-vocabulary mapping, semantic pruning, and language-based loop closure.

  2. Ella: Embodied Social Agents with Lifelong Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Ella, an embodied social agent with a name-centric semantic memory and a spatiotemporal episodic memory, outperformed two re-implemented baselines in social influence and leadership tasks in a 3D simulation.

  3. GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GeoProg3D combines a georeferenced hierarchical 3D language field, geographic vision APIs, and LLM-generated programs to answer natural-language queries about city-scale 3D scenes, and includes a new 952-query benchma...

  4. BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box Fusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    BoxFusion fuses per-frame 3D bounding box proposals from Cubify Anything and CLIP semantics into open-vocabulary 3D detections, reporting state-of-the-art AP among online methods without dense reconstruction.

  5. FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FreeQ-Graph builds a semantically aligned 3D scene graph using CLIP, LLaVA, and an LLM, and reports strong zero-shot results on 3D visual grounding, segmentation, and scene graph prediction.

  6. EquAct: An SE(3)-Equivariant Multi-Task Transformer for Open-Loop Robotic Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    EquAct embeds SE(3) equivariance into a multi-task keyframe manipulation transformer with language conditioning, improving spatial generalization over non-equivariant baselines.

  7. GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GSemSplat predicts open-vocabulary semantic features attached to 3D Gaussians from two uncalibrated images and generalizes across scenes with a single feed-forward pass.

  8. ChatSplat: 3D Conversational Gaussian Splatting

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ChatSplat learns a 3D conversational field in Gaussian Splatting that supports object-, view-, and scene-level chat with an LLM at real-time speeds.

  9. 3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    3D-Mem represents explored and unexplored regions as compact snapshot images that a vision-language model can reason over, improving embodied question answering and lifelong navigation.

  10. FAST-Splat: Fast, Ambiguity-Free Semantics Transfer in Gaussian Splatting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FAST-Splat stores a 3D semantic code on each Gaussian, trains it jointly with the scene, and uses a hash-table of detected objects to give fast, low-memory open-vocabulary 3D segmentation with disambiguated object labels.

  11. LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

    cs.CV 2026-07 conditional novelty 5.0 of 10

    LEEVLA improves VLA robot policies by training-time prioritization of dynamic instruction-relevant patches plus structured latent future-feature prediction with topology constraints.

  12. Adapting Vision-Language Models Without Labels: A Comprehensive Survey

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A survey that organizes unsupervised vision-language model adaptation by unlabeled-data availability into four paradigms: data-free transfer, domain transfer, episodic test-time, and online test-time adaptation.

  13. OpenIN: Open-Vocabulary Instance-Oriented Navigation in Dynamic Domestic Environments

    cs.RO 2025-01 conditional novelty 5.0 of 10

    OpenIN uses a dynamically updated scene graph of carried-by relationships, plus LLM and VLM guidance, to navigate to specific moved objects in homes, reporting higher success than two open-vocabulary baselines.

  14. Semantic Intelligence: Integrating GPT-4 with A Planning in Low-Cost Robotics

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A hybrid system where GPT-4 selects and adjusts A* paths lets a cheap quadruped follow semantic instructions like avoiding a toxic spill or collecting a resource first, with 90-100% success on the authors' tests.

  15. Semantic Mapping in Indoor Embodied AI -- A Survey on Advances, Challenges, and Future Directions

    cs.RO 2025-01 conditional novelty 4.0 of 10

    A review of indoor semantic mapping for embodied AI, categorized by map structure and encoding type, concluding that the field is moving toward open-vocabulary, queryable, task-agnostic maps.

  16. Voxel-Aggregated Feature Synthesis: Efficient Dense Mapping for Simulated 3D Reasoning

    cs.CV 2024-11 conditional novelty 4.0 of 10

    VAFS replaces per-frame embedding in dense 3D mapping with per-object synthetic view embedding, using simulator ground-truth segmentation, achieving faster and higher-IoU semantic maps in simulation.

  17. Foundation Model Driven Robotics: A Comprehensive Review

    cs.RO 2025-07 conditional novelty 2.0 of 10

    A review of foundation-model-driven robotics that synthesizes recent work across perception, planning, control, HRI, simulation, and sim-to-real transfer, and highlights open challenges.

Pith tools