Auxiliary cross-modal alignment and future-scene prediction pretraining is claimed to improve VLN agents, but the paper's own ablations show no benefit over no-pretraining baselines.
Vision-and- Language Navigation: Interpreting visually-grounded navigation instructions in real environments
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2019 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Transferable Representation Learning in Vision-and-Language Navigation
Auxiliary cross-modal alignment and future-scene prediction pretraining is claimed to improve VLN agents, but the paper's own ablations show no benefit over no-pretraining baselines.