A pretrained image-goal navigation model combining early-fusion ViT, auxiliary objectives, and game-video data reports higher success than GNM, ViNT, and NoMaD, though zero-shot generalization is clouded by possible pretraining overlap with test environments.
We define the provided navigation trajectory as: τ = (o0, o1, · · ·, oT ; p0, p1, · · ·, pT ) where T represents the total number of steps in the trajectory
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
A pretrained image-goal navigation model combining early-fusion ViT, auxiliary objectives, and game-video data reports higher success than GNM, ViNT, and NoMaD, though zero-shot generalization is clouded by possible pretraining overlap with test environments.