NavWAM is a diffusion-transformer policy that jointly learns future observation prediction, goal-progress values, and action chunks in a shared latent sequence for goal-conditioned visual navigation.
Navdreamer: Video models as zero-shot 3d navigators
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.RO 6years
2026 6verdicts
UNVERDICTED 6roles
background 3polarities
background 3representative citing papers
ImagineUAV is a 1.3B-parameter cascaded world-action framework that generates instruction-conditioned future observations via latent video diffusion, infers motions, and applies kinodynamic planning to outperform VLN/VLA baselines in aerial navigation.
FlyMirage automates generation of large-scale, diverse, photorealistic aerial VLN datasets with dynamically feasible UAV trajectories by combining LLM scene design, 3D Gaussian Splatting world models, automated exploration, and trajectory planning.
Navigation system transfers image generation models to embodied tasks via BEV-based traversability mask generation from language and cross-view localization for odometry correction, shown on a UAV completing 160m outdoor navigation.
This survey organizes aerial vision-language navigation methods into five architectural categories, critically reviews evaluation infrastructure, and synthesizes seven open problems for LLM/VLM integration.
A survey that clarifies boundaries and organizes World Action Models by generation requirements and predictive substrates, identifying a trend toward generating less of the future.
citing papers explorer
-
NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation
NavWAM is a diffusion-transformer policy that jointly learns future observation prediction, goal-progress values, and action chunks in a shared latent sequence for goal-conditioned visual navigation.
-
ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning
ImagineUAV is a 1.3B-parameter cascaded world-action framework that generates instruction-conditioned future observations via latent video diffusion, infers motions, and applies kinodynamic planning to outperform VLN/VLA baselines in aerial navigation.
-
FlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model
FlyMirage automates generation of large-scale, diverse, photorealistic aerial VLN datasets with dynamically feasible UAV trajectories by combining LLM scene design, 3D Gaussian Splatting world models, automated exploration, and trajectory planning.
-
PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation
Navigation system transfers image generation models to embodied tasks via BEV-based traversability mask generation from language and cross-view localization for odometry correction, shown on a UAV completing 160m outdoor navigation.
-
Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models
This survey organizes aerial vision-language navigation methods into five architectural categories, critically reviews evaluation infrastructure, and synthesizes seven open problems for LLM/VLM integration.
-
World Action Models: A Survey
A survey that clarifies boundaries and organizes World Action Models by generation requirements and predictive substrates, identifying a trend toward generating less of the future.