Pith. sign in

REVIEW 6 cited by

VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.11609 v1 pith:MWHNAHMK submitted 2024-11-18 cs.RO

VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation

classification cs.RO
keywords navigationtargetvln-gamelanguageframeworkobjectsearchenvironment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Following human instructions to explore and search for a specified target in an unfamiliar environment is a crucial skill for mobile service robots. Most of the previous works on object goal navigation have typically focused on a single input modality as the target, which may lead to limited consideration of language descriptions containing detailed attributes and spatial relationships. To address this limitation, we propose VLN-Game, a novel zero-shot framework for visual target navigation that can process object names and descriptive language targets effectively. To be more precise, our approach constructs a 3D object-centric spatial map by integrating pre-trained visual-language features with a 3D reconstruction of the physical environment. Then, the framework identifies the most promising areas to explore in search of potential target candidates. A game-theoretic vision language model is employed to determine which target best matches the given language description. Experiments conducted on the Habitat-Matterport 3D (HM3D) dataset demonstrate that the proposed framework achieves state-of-the-art performance in both object goal navigation and language-based navigation tasks. Moreover, we show that VLN-Game can be easily deployed on real-world robots. The success of VLN-Game highlights the promising potential of using game-theoretic methods with compact vision-language models to advance decision-making capabilities in robotic systems. The supplementary video and code can be accessed via the following link: https://sites.google.com/view/vln-game.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Instance-Enriched Semantic Maps for Visual Language Navigation

    cs.RO 2026-07 conditional novelty 6.0

    Instance-enriched 2.5D maps with LLM expert-fusion retrieval improve zero-shot object retrieval and navigation over HOV-SG while cutting map storage by ~96%.

  2. SurveilNav: Collaborative Object Goal Navigation with Robot and Surveillance System

    cs.RO 2026-06 unverdicted novelty 6.0

    SurveilNav integrates robot local perception with multi-view surveillance for improved collaborative object goal navigation and reports SOTA results on HM3D.

  3. Common-agency Games for Multi-Objective Test-Time Alignment

    cs.GT 2026-05 unverdicted novelty 6.0

    CAGE uses common-agency games and an EPEC algorithm to compute equilibrium policies that balance multiple conflicting objectives for test-time LLM alignment.

  4. FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning

    cs.RO 2025-09 unverdicted novelty 6.0

    FiLM-Nav fine-tunes VLMs on a mixture of simulated navigation tasks to reach state-of-the-art SPL and success on HM3D ObjectNav and OVON benchmarks with generalization to unseen categories.

  5. Instance-Enriched Semantic Maps for Visual Language Navigation

    cs.RO 2026-07 unverdicted novelty 5.0

    Instance-enriched 2.5D semantic maps plus LLM expert fusion improve VLN object retrieval by >17% and success by >23% while cutting storage ~96% versus 3D baselines.

  6. Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation

    cs.RO 2026-06 unverdicted novelty 3.0

    HSAN integrates hierarchical semantic graphs, optimal transport-based goal selection, and graph-aware RL to claim SOTA results on VLN-CE tasks.