PGN adapts OpenPangu-7B with Q-Former alignment and LoRA to offline vision-language navigation action prediction, reaching 62.29% normalized action match on 500 held-out expert trajectories.
History aware multimodal transformer for vision-and-language navi- gation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model
PGN adapts OpenPangu-7B with Q-Former alignment and LoRA to offline vision-language navigation action prediction, reaching 62.29% normalized action match on 500 held-out expert trajectories.