Pith. sign in

REVIEW 7 cited by

VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03275 v1 pith:USEMGXQ3 submitted 2023-12-06 cs.RO cs.AI

classification cs.ROcs.AI
keywords vlfmnavigationsemanticvision-languageenvironmentsfrontiermapsnavigate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding how humans leverage semantic knowledge to navigate unfamiliar environments and decide where to explore next is pivotal for developing robots capable of human-like search behaviors. We introduce a zero-shot navigation approach, Vision-Language Frontier Maps (VLFM), which is inspired by human reasoning and designed to navigate towards unseen semantic objects in novel environments. VLFM builds occupancy maps from depth observations to identify frontiers, and leverages RGB observations and a pre-trained vision-language model to generate a language-grounded value map. VLFM then uses this map to identify the most promising frontier to explore for finding an instance of a given target object category. We evaluate VLFM in photo-realistic environments from the Gibson, Habitat-Matterport 3D (HM3D), and Matterport 3D (MP3D) datasets within the Habitat simulator. Remarkably, VLFM achieves state-of-the-art results on all three datasets as measured by success weighted by path length (SPL) for the Object Goal Navigation task. Furthermore, we show that VLFM's zero-shot nature enables it to be readily deployed on real-world robots such as the Boston Dynamics Spot mobile manipulation platform. We deploy VLFM on Spot and demonstrate its capability to efficiently navigate to target objects within an office building in the real world, without any prior knowledge of the environment. The accomplishments of VLFM underscore the promising potential of vision-language models in advancing the field of semantic navigation. Videos of real-world deployment can be viewed at naoki.io/vlfm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EAGOR: Embodied Reasoning in Omni-direction

    cs.RO 2026-07 conditional novelty 7.0 of 10

    EAGOR reformulates embodied 360-degree directional reasoning as recursive Bayesian estimation on a spherical manifold using spherical harmonics, achieving training-free, rotation-equivariant target tracking.

  2. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  3. Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring

    cs.RO 2026-07 conditional novelty 6.0 of 10

    An object-centric, training-free pipeline using CLIP-derived room-probability vectors to score frontiers improves zero-shot ObjectNav success by a relative 3% over an image-based baseline on HM3D.

  4. VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical room-and-object memory that persists across independent ObjectNav episodes yields small success-rate gains, but most of the gain comes from within-episode memory rather than the cross-episode component.

  5. CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition

    cs.RO 2025-06 conditional novelty 6.0 of 10

    CARMA combines object detection, person tracking, action detection, and a vision-language model to produce instance-level actor-action-object triplets for human-robot group interactions, achieving up to 72% task succe...

  6. VL-Explore: Zero-shot Vision-Language Exploration and Target Discovery by Mobile Robots

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A monocular, map-free navigation pipeline uses CLIP scores on six image tiles to explore rooms and discover a target in real time.

  7. RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception

    cs.CV 2025-01 conditional novelty 6.0 of 10

    An RL agent generates hard synthetic spatial-reasoning examples to fine-tune VLMs, improving performance on simulated test scenes.

Pith tools