REVIEW 17 cited by
CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose CLIP-Fields, an implicit scene model that can be used for a variety of tasks, such as segmentation, instance identification, semantic search over space, and view localization. CLIP-Fields learns a mapping from spatial locations to semantic embedding vectors. Importantly, we show that this mapping can be trained with supervision coming only from web-image and web-text trained models such as CLIP, Detic, and Sentence-BERT; and thus uses no direct human supervision. When compared to baselines like Mask-RCNN, our method outperforms on few-shot instance identification or semantic segmentation on the HM3D dataset with only a fraction of the examples. Finally, we show that using CLIP-Fields as a scene memory, robots can perform semantic navigation in real-world environments. Our code and demonstration videos are available here: https://mahis.life/clip-fields
Forward citations
Cited by 17 Pith papers
-
LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM
A 3D Gaussian Splatting SLAM system learns compact 16-dim language features per Gaussian, enabling real-time open-vocabulary mapping, semantic pruning, and language-based loop closure.
-
Ella: Embodied Social Agents with Lifelong Memory
Ella, an embodied social agent with a name-centric semantic memory and a spatiotemporal episodic memory, outperformed two re-implemented baselines in social influence and leadership tasks in a 3D simulation.
-
GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
GeoProg3D combines a georeferenced hierarchical 3D language field, geographic vision APIs, and LLM-generated programs to answer natural-language queries about city-scale 3D scenes, and includes a new 952-query benchma...
-
BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box Fusion
BoxFusion fuses per-frame 3D bounding box proposals from Cubify Anything and CLIP semantics into open-vocabulary 3D detections, reporting state-of-the-art AP among online methods without dense reconstruction.
-
FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding
FreeQ-Graph builds a semantically aligned 3D scene graph using CLIP, LLaVA, and an LLM, and reports strong zero-shot results on 3D visual grounding, segmentation, and scene graph prediction.
-
EquAct: An SE(3)-Equivariant Multi-Task Transformer for Open-Loop Robotic Manipulation
EquAct embeds SE(3) equivariance into a multi-task keyframe manipulation transformer with language conditioning, improving spatial generalization over non-equivariant baselines.
-
GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs
GSemSplat predicts open-vocabulary semantic features attached to 3D Gaussians from two uncalibrated images and generalizes across scenes with a single feed-forward pass.
-
ChatSplat: 3D Conversational Gaussian Splatting
ChatSplat learns a 3D conversational field in Gaussian Splatting that supports object-, view-, and scene-level chat with an LLM at real-time speeds.
-
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
3D-Mem represents explored and unexplored regions as compact snapshot images that a vision-language model can reason over, improving embodied question answering and lifelong navigation.
-
FAST-Splat: Fast, Ambiguity-Free Semantics Transfer in Gaussian Splatting
FAST-Splat stores a 3D semantic code on each Gaussian, trains it jointly with the scene, and uses a hash-table of detected objects to give fast, low-memory open-vocabulary 3D segmentation with disambiguated object labels.
-
LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action
LEEVLA improves VLA robot policies by training-time prioritization of dynamic instruction-relevant patches plus structured latent future-feature prediction with topology constraints.
-
Adapting Vision-Language Models Without Labels: A Comprehensive Survey
A survey that organizes unsupervised vision-language model adaptation by unlabeled-data availability into four paradigms: data-free transfer, domain transfer, episodic test-time, and online test-time adaptation.
-
OpenIN: Open-Vocabulary Instance-Oriented Navigation in Dynamic Domestic Environments
OpenIN uses a dynamically updated scene graph of carried-by relationships, plus LLM and VLM guidance, to navigate to specific moved objects in homes, reporting higher success than two open-vocabulary baselines.
-
Semantic Intelligence: Integrating GPT-4 with A Planning in Low-Cost Robotics
A hybrid system where GPT-4 selects and adjusts A* paths lets a cheap quadruped follow semantic instructions like avoiding a toxic spill or collecting a resource first, with 90-100% success on the authors' tests.
-
Semantic Mapping in Indoor Embodied AI -- A Survey on Advances, Challenges, and Future Directions
A review of indoor semantic mapping for embodied AI, categorized by map structure and encoding type, concluding that the field is moving toward open-vocabulary, queryable, task-agnostic maps.
-
Voxel-Aggregated Feature Synthesis: Efficient Dense Mapping for Simulated 3D Reasoning
VAFS replaces per-frame embedding in dense 3D mapping with per-object synthetic view embedding, using simulator ground-truth segmentation, achieving faster and higher-IoU semantic maps in simulation.
-
Foundation Model Driven Robotics: A Comprehensive Review
A review of foundation-model-driven robotics that synthesizes recent work across perception, planning, control, HRI, simulation, and sim-to-real transfer, and highlights open challenges.
Discussion (0). Continue with ORCID to comment.