Pith. sign in

REVIEW 3 cited by

Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.12667 v3 pith:6A5NSML2 submitted 2022-03-22 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords researchtaskscurrentfuturegoallanguagemethodsnatural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A long-term goal of AI research is to build intelligent agents that can communicate with humans in natural language, perceive the environment, and perform real-world tasks. Vision-and-Language Navigation (VLN) is a fundamental and interdisciplinary research topic towards this goal, and receives increasing attention from natural language processing, computer vision, robotics, and machine learning communities. In this paper, we review contemporary studies in the emerging field of VLN, covering tasks, evaluation metrics, methods, etc. Through structured analysis of current progress and challenges, we highlight the limitations of current VLN and opportunities for future work. This paper serves as a thorough reference for the VLN research community.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

    cs.CL 2026-08 accept novelty 6.0 of 10

    Spatial-memory staleness is a measurable safety failure for VLM agents: stale memory increases deaths, and visual auditing of stale entries is highly model-dependent.

  2. ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ReMeREC introduces a relation-aware multi-entity referring expression comprehension framework and the ReMeX dataset, reporting state-of-the-art grounding and relation prediction, with some evaluation caveats.

  3. Active Test-time Vision-Language Navigation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    ATENA uses episodic success/failure labels and a mixture entropy objective to adapt vision-language navigation policies at test time, improving REVERIE, R2R, and R2R-CE benchmarks.

Pith tools