A frozen Qwen2.5-VL model with prompt engineering and a two-frame egocentric view achieves 5% success on 20 R2R VLN-CE episodes, near the zero-movement baseline.
Vision-and-language navigation: In- terpreting visually-grounded navigation instructions in real environments
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
A Navigation Framework Utilizing Vision-Language Models
A frozen Qwen2.5-VL model with prompt engineering and a two-frame egocentric view achieves 5% success on 20 R2R VLN-CE episodes, near the zero-movement baseline.