Pith. sign in

REVIEW 40 cited by

ViNT: A Foundation Model for Visual Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.14846 v2 pith:D6JPYILM submitted 2023-06-26 cs.RO cs.CVcs.LG

ViNT: A Foundation Model for Visual Navigation

classification cs.RO cs.CVcs.LG
keywords navigationvintmodelsdatasetsfoundationmodeltraineddownstream
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

General-purpose pre-trained models ("foundation models") have enabled practitioners to produce generalizable solutions for individual machine learning problems with datasets that are significantly smaller than those required for learning from scratch. Such models are typically trained on large and diverse datasets with weak supervision, consuming much more training data than is available for any individual downstream application. In this paper, we describe the Visual Navigation Transformer (ViNT), a foundation model that aims to bring the success of general-purpose pre-trained models to vision-based robotic navigation. ViNT is trained with a general goal-reaching objective that can be used with any navigation dataset, and employs a flexible Transformer-based architecture to learn navigational affordances and enable efficient adaptation to a variety of downstream navigational tasks. ViNT is trained on a number of existing navigation datasets, comprising hundreds of hours of robotic navigation from a variety of different robotic platforms, and exhibits positive transfer, outperforming specialist models trained on singular datasets. ViNT can be augmented with diffusion-based subgoal proposals to explore novel environments, and can solve kilometer-scale navigation problems when equipped with long-range heuristics. ViNT can also be adapted to novel task specifications with a technique inspired by prompt-tuning, where the goal encoder is replaced by an encoding of another task modality (e.g., GPS waypoints or routing commands) embedded into the same space of goal tokens. This flexibility and ability to accommodate a variety of downstream problem domains establishes ViNT as an effective foundation model for mobile robotics. For videos, code, and model checkpoints, see our project page at https://visualnav-transformer.github.io.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 40 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies

    cs.RO 2026-06 unverdicted novelty 8.0

    TAKO demonstrates real-time adversarial takeover of robotic diffusion policies via reusable universal patches on visual inputs, achieving 100% success in steering attacker-chosen trajectories across multiple tasks, en...

  2. POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation

    cs.RO 2026-05 unverdicted novelty 7.0

    POINav-Bench provides the first high-fidelity real-world benchmark for POI-goal VLN using 3DGS reconstructions of 126k m² with 163 POIs, supported by a Brain-Action framework and 70K real signage-entrance dataset.

  3. Sentinel: Embodied Cooperative Spatial Reasoning and Planning

    cs.CV 2026-05 unverdicted novelty 7.0

    Introduces Sentinel Challenge benchmark and CoSaR framework for cooperative spatial reasoning and planning among 3-5 decentralized embodied agents across 14 city-scale scenes.

  4. World Models as Group Actions

    cs.CV 2026-05 unverdicted novelty 7.0

    Formalizes video world models as group actions on states and uses latent regularization with synthesized supervision to enforce consistency, introducing GAC and GAR metrics that improve structural correctness in SOTA models.

  5. How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace

    cs.AI 2026-04 unverdicted novelty 7.0

    Large multimodal models display emerging but limited spatial action capabilities in goal-oriented urban 3D navigation, remaining far from human-level performance with errors diverging rapidly after critical decision points.

  6. Rectified Schr\"odinger Bridge Matching for Few-Step Visual Navigation

    cs.RO 2026-04 unverdicted novelty 7.0

    RSBM exploits velocity field invariance across regularization levels to achieve over 94% cosine similarity and 92% success in visual navigation using only 3 integration steps.

  7. STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation

    cs.CV 2026-04 conditional novelty 7.0

    STRNet improves goal-conditioned visual navigation by replacing simplistic encoders and pooling with a spatio-temporal fusion module that performs spatial graph reasoning and hybrid temporal modeling.

  8. AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning

    cs.RO 2025-12 conditional novelty 7.0

    AID trains diffusion policies via behavior cloning on existing MAIPP planners followed by RL fine-tuning to achieve faster execution and higher information gain in multi-agent coordination.

  9. Learning Interactive Real-World Simulators

    cs.AI 2023-10 conditional novelty 7.0

    UniSim learns a universal real-world simulator from orchestrated diverse datasets, enabling zero-shot deployment of policies trained purely in simulation.

  10. ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset

    cs.RO 2026-07 conditional novelty 6.0

    ACME is a socially navigated robot and pedestrian trajectory dataset covering 8 sites in 5 countries with 7 robot embodiments, including human-verified BEV tracks and robot speech annotations.

  11. G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation

    cs.RO 2026-07 conditional novelty 6.0

    G2-Nav turns vision-language reasoning about social scenes into a weighted costmap with a safety reflex layer, tested in recorded and live real-world trials.

  12. SeeSE3: Emergence of 3D Space in Vision Features

    cs.CV 2026-07 conditional novelty 6.0

    Self-supervised vision features, especially DINOv2, contain a subspace that a small trained adapter can map to 3D camera motion, enabling pose estimation and latent-space navigation without explicit 3D reconstruction.

  13. Learning to Navigate Efficiently with Only 0.58M Trainable Parameters

    cs.RO 2026-07 conditional novelty 6.0

    Decomposed navigation with analytical geometry interfaces and three small learned modules (0.58M trainable params) approaches SOTA point-goal performance at 50 Hz with lowest collisions.

  14. Learning to Navigate Efficiently with Only 0.58M Trainable Parameters

    cs.RO 2026-07 conditional novelty 6.0

    A 0.58M-trainable-parameter navigation model reaches near-state-of-the-art point-goal success, with 233x fewer trainable parameters, a lower collision rate, and 10+ Hz embedded inference.

  15. NavWM: A Unified Navigation World Model for Foresight-Driven Planning

    cs.RO 2026-06 unverdicted novelty 6.0

    NavWM unifies latent world tokens and anchor-based multimodal trajectory forecasting into a closed-loop planner that improves future state generation and zero-shot navigation.

  16. NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

    cs.RO 2026-06 unverdicted novelty 6.0

    NavWAM is a diffusion-transformer policy that jointly learns future observation prediction, goal-progress values, and action chunks in a shared latent sequence for goal-conditioned visual navigation.

  17. From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation

    cs.RO 2026-06 unverdicted novelty 6.0

    FlowPilot combines anchored flow matching for multimodal action pre-training with human-in-the-loop preference learning to improve long-horizon monocular sidewalk navigation, reporting 42% success in simulation and re...

  18. Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation

    cs.CV 2026-06 unverdicted novelty 6.0

    Goal2Pixel grounds VLN-CE goals to image pixels via VLM prediction plus keyframe memory, reaching 54.1% SR on R2R-CE Val-Unseen with 7.75 calls per episode versus 46.62 for action prediction.

  19. Autonomous Frontier-Based Exploration with VLM Guidance

    cs.RO 2026-05 unverdicted novelty 6.0

    A VLM-based method for selecting exploration frontiers in robotics achieves up to 24% better map coverage than standard geometric heuristics in simulated indoor environments.

  20. Improved Baselines with Representation Autoencoders

    cs.CV 2026-05 conditional novelty 6.0

    RAE v2 reaches gFID 1.06 on ImageNet-256 in 80 epochs by combining multi-layer encoder sums, complementary REPA targets, and free guidance via output reparameterization.

  21. NavOL: Navigation Policy with Online Imitation Learning

    cs.RO 2026-05 unverdicted novelty 6.0

    NavOL collects expert trajectory labels online from a global planner during policy rollouts in simulation to train a diffusion navigation policy, mitigating distribution shift and improving performance on visual navig...

  22. Human Cognition in Machines: A Unified Perspective of World Models

    cs.RO 2026-04 unverdicted novelty 6.0

    The paper introduces a unified framework for world models that fully incorporates all cognitive functions from Cognitive Architecture Theory, highlights under-researched areas in motivation and meta-cognition, and pro...

  23. RAE-NWM: Navigation World Model in Dense Visual Representation Space

    cs.CV 2026-03 conditional novelty 6.0

    Navigation world models trained in dense DINOv2 space with flow-matching CDiT-DH and time-gated action injection improve structural stability and planning over VAE baselines.

  24. Approximate Imitation Learning for Event-based Quadrotor Flight in Cluttered Environments

    cs.RO 2026-03 conditional novelty 6.0

    Approximate imitation learning trains event-to-control quadrotor policies 28× faster by freezing a pretrained event encoder and fine-tuning a shared action decoder via a state-based approximate student, matching onlin...

  25. Learning to Localize Reference Trajectories in Image-Space for Visual Navigation

    cs.RO 2026-02 conditional novelty 6.0

    LoTIS localizes a reference RGB trajectory in the robot's current view, predicting image-space coordinates, visibility, and distance to provide robot-agnostic guidance for navigation.

  26. CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents

    cs.AI 2025-11 reject novelty 6.0

    A navigation benchmark that converts simulated collisions, energy use, and delivery outcomes into dollar costs and revenue, and finds rule-based delivery robot baselines unprofitable per run.

  27. Splatblox: Traversability-Aware Gaussian Splatting for Outdoor Robot Navigation

    cs.RO 2025-11 conditional novelty 6.0

    Splatblox creates a traversability-aware ESDF from RGB-LiDAR fusion via Gaussian Splatting, enabling semantic navigation that outperforms prior methods by over 50% success rate in vegetated field trials on quadruped a...

  28. MATT-Diff: Multimodal Active Target Tracking by Diffusion Policy

    cs.RO 2025-11 unverdicted novelty 6.0

    MATT-Diff uses a diffusion model with vision transformer and attention to generate multimodal actions for active multi-target tracking from expert planner demonstrations.

  29. $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

    cs.LG 2025-04 unverdicted novelty 6.0

    π_{0.5} is a VLA model that achieves long-horizon dexterous manipulation in entirely new homes through co-training on heterogeneous tasks and multi-source data including web and semantic predictions.

  30. Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight

    cs.RO 2025-01 unverdicted novelty 6.0

    DreamerV3 enables pixel-to-control policies for drone racing that reach 9 m/s in both simulation and real hardware-in-the-loop tests.

  31. OpenVLA: An Open-Source Vision-Language-Action Model

    cs.RO 2024-06 unverdicted novelty 6.0

    OpenVLA achieves 16.5% higher task success than the 55B RT-2-X model across 29 tasks with 7x fewer parameters while enabling effective fine-tuning and quantization without performance loss.

  32. Octo: An Open-Source Generalist Robot Policy

    cs.RO 2024-05 unverdicted novelty 6.0

    Octo is an open-source transformer-based generalist robot policy pretrained on 800k trajectories that serves as an effective initialization for finetuning across diverse robotic platforms.

  33. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    cs.RO 2024-03 accept novelty 6.0

    DROID is a new 76k-trajectory in-the-wild robot manipulation dataset spanning 564 scenes and 84 tasks that improves policy performance and generalization when used for training.

  34. Learning Robot Visual Navigation in Crowds via Intention-Aware Scene Representations

    cs.RO 2026-06 unverdicted novelty 5.0

    iCrowdNav encodes egocentric visual observations with occupancy features and human pose intentions to improve DRL policies for crowd navigation, showing better performance than baselines in experiments and real-world tests.

  35. Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation

    cs.RO 2026-06 unverdicted novelty 5.0

    A training-free fusion layer enables stale VLM selections to improve a real-time planner's trajectory scoring for urban sidewalk navigation, yielding 30% ADE reduction in challenging scenarios.

  36. Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

    cs.RO 2026-06 unverdicted novelty 5.0

    SALSA aligns social features and adds future-risk signals in VLA models to cut near-collisions by 86.4% and raise social accuracy from 53% to 93% on SCAND and real robots.

  37. Look Further: Socially-Compliant Navigation System in Residential Buildings

    cs.RO 2026-05 unverdicted novelty 5.0

    Proactive lane-changing at eight meters improves human ratings of robot motion in frontal hallway approaches according to a 42-participant study, with no advantage shown at intersections.

  38. Drift-Resistant Navigation World Model with Anchored Epipolar Guidance

    cs.CV 2026-05 unverdicted novelty 5.0

    A generative navigation world model that uses sparse anchored rollout with epipolar constraints to reduce perceptual and geometric drift.

  39. NavRL++: A System-Level Framework for Improving Sim-to-Real Transfer in Reinforcement Learning-Based Robot Navigation

    cs.RO 2026-05 unverdicted novelty 5.0

    NavRL++ improves sim-to-real transfer for RL navigation via empirical analysis of perturbations, perturbation-aware fine-tuning, and a Transformer temporal policy, with real-world validation showing outperformance ove...

  40. Rectified Schr\"odinger Bridge Matching for Few-Step Visual Navigation

    cs.RO 2026-04 unverdicted novelty 5.0

    Rectified Schrödinger Bridge Matching uses ε-invariant velocity structure and a learned prior so generative navigation policies reach ~94% cosine similarity and 92% success in three integration steps without distillation.