Pith. sign in

REVIEW 3 cited by

General Scene Adaptation for Vision-and-Language Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.17403 v1 pith:3U3Z447A submitted 2025-01-29 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords instructionsnavigationagentsenvironmentsevaluatedatasetgsa-r2rscene
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-and-Language Navigation (VLN) tasks mainly evaluate agents based on one-time execution of individual instructions across multiple environments, aiming to develop agents capable of functioning in any environment in a zero-shot manner. However, real-world navigation robots often operate in persistent environments with relatively consistent physical layouts, visual observations, and language styles from instructors. Such a gap in the task setting presents an opportunity to improve VLN agents by incorporating continuous adaptation to specific environments. To better reflect these real-world conditions, we introduce GSA-VLN, a novel task requiring agents to execute navigation instructions within a specific scene and simultaneously adapt to it for improved performance over time. To evaluate the proposed task, one has to address two challenges in existing VLN datasets: the lack of OOD data, and the limited number and style diversity of instructions for each scene. Therefore, we propose a new dataset, GSA-R2R, which significantly expands the diversity and quantity of environments and instructions for the R2R dataset to evaluate agent adaptability in both ID and OOD contexts. Furthermore, we design a three-stage instruction orchestration pipeline that leverages LLMs to refine speaker-generated instructions and apply role-playing techniques to rephrase instructions into different speaking styles. This is motivated by the observation that each individual user often has consistent signatures or preferences in their instructions. We conducted extensive experiments on GSA-R2R to thoroughly evaluate our dataset and benchmark various methods. Based on our findings, we propose a novel method, GR-DUET, which incorporates memory-based navigation graphs with an environment-specific training strategy, achieving state-of-the-art results on all GSA-R2R splits.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    LH-AVLN couples multi-goal heterogeneous navigation with alternating goal-linked binaural audio, and existing vision-language and audio-visual agents largely fail full missions while PAG-Nav is only a weak diagnostic ...

  2. UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    UAV-ON is a new benchmark of 14 Unreal Engine environments with 1270 annotated objects that tests whether aerial agents can navigate to goals described by semantic instance-level instructions.

  3. CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking

    cs.AI 2025-07 conditional novelty 5.0 of 10

    CogDDN uses a fast heuristic VLM paired with a slow analytic reflection process and a growing knowledge base to navigate to objects that implicitly satisfy a user's demand, with large reported gains on AI2Thor DDN benchmarks.

Pith tools