NavAgent fuses fine-grained landmark detection, a growing scene topology map, and an LLM to improve outdoor vision-and-language navigation, outperforming VELMA on Touchdown and Map2seq.
Beyond the nav-graph: Vision-and-language navigation in continuous environments,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
NavAgent fuses fine-grained landmark detection, a growing scene topology map, and an LLM to improve outdoor vision-and-language navigation, outperforming VELMA on Touchdown and Map2seq.