Pith. sign in

REVIEW 18 cited by

NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07896 v1 pith:GRLDICNU submitted 2023-10-11 cs.RO cs.CVcs.LG

NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration

classification cs.RO cs.CVcs.LG
keywords navigationenvironmentsgoalmodelsdiffusionexplorationnovelpolicy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Robotic learning for navigation in unfamiliar environments needs to provide policies for both task-oriented navigation (i.e., reaching a goal that the robot has located), and task-agnostic exploration (i.e., searching for a goal in a novel setting). Typically, these roles are handled by separate models, for example by using subgoal proposals, planning, or separate navigation strategies. In this paper, we describe how we can train a single unified diffusion policy to handle both goal-directed navigation and goal-agnostic exploration, with the latter providing the ability to search novel environments, and the former providing the ability to reach a user-specified goal once it has been located. We show that this unified policy results in better overall performance when navigating to visually indicated goals in novel environments, as compared to approaches that use subgoal proposals from generative models, or prior methods based on latent variable models. We instantiate our method by using a large-scale Transformer-based policy trained on data from multiple ground robots, with a diffusion model decoder to flexibly handle both goal-conditioned and goal-agnostic navigation. Our experiments, conducted on a real-world mobile robot platform, show effective navigation in unseen environments in comparison with five alternative methods, and demonstrate significant improvements in performance and lower collision rates, despite utilizing smaller models than state-of-the-art approaches. For more videos, code, and pre-trained model checkpoints, see https://general-navigation-models.github.io/nomad/

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies

    cs.RO 2026-06 unverdicted novelty 8.0

    TAKO demonstrates real-time adversarial takeover of robotic diffusion policies via reusable universal patches on visual inputs, achieving 100% success in steering attacker-chosen trajectories across multiple tasks, en...

  2. Action Agent: Agentic Video Generation Meets Flow-Constrained Diffusion

    cs.RO 2026-05 unverdicted novelty 7.0

    Action Agent pairs LLM-driven video generation with a flow-constrained diffusion transformer to produce velocity commands, raising video success to 86% and delivering 64.7% real-world navigation on a Unitree G1 humanoid.

  3. FlowPilot: Real-Time World-Action Modeling for Agile UAV Navigation

    cs.RO 2026-08 conditional novelty 6.0

    A dual-stream flow-matching model jointly predicts future depth and Bernstein-polynomial trajectories, enabling real-time onboard quadrotor navigation in clutter.

  4. Distilling Global Traversability Priors for Image-based Affordance Prediction in Off-road Environments

    cs.RO 2026-07 conditional novelty 6.0

    Image-based off-road navigation improves when an affordance model is supervised in heading space with plan-derived labels from satellite traversability maps rather than human demonstrations alone.

  5. Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models

    cs.AI 2026-07 conditional novelty 6.0

    A query-based action interface improves zero-shot sim-to-real navigation by reorganizing inherited vision-language representations before action prediction, cutting instruction OOD outputs and raising closed-loop succ...

  6. Learning to Navigate Efficiently with Only 0.58M Trainable Parameters

    cs.RO 2026-07 conditional novelty 6.0

    A 0.58M-trainable-parameter navigation model reaches near-state-of-the-art point-goal success, with 233x fewer trainable parameters, a lower collision rate, and 10+ Hz embedded inference.

  7. Learning to Navigate Efficiently with Only 0.58M Trainable Parameters

    cs.RO 2026-07 conditional novelty 6.0

    Decomposed navigation with analytical geometry interfaces and three small learned modules (0.58M trainable params) approaches SOTA point-goal performance at 50 Hz with lowest collisions.

  8. NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

    cs.RO 2026-06 unverdicted novelty 6.0

    NavWAM is a diffusion-transformer policy that jointly learns future observation prediction, goal-progress values, and action chunks in a shared latent sequence for goal-conditioned visual navigation.

  9. Adversarial Dual On-Policy Distillation from Expressive Teacher

    cs.LG 2026-05 unverdicted novelty 6.0

    FA-OPD co-trains a flow-matching teacher and MLP student via adversarial dual on-policy distillation, improving robustness over baselines on six robot benchmarks with noisy or limited demonstrations.

  10. Improved Baselines with Representation Autoencoders

    cs.CV 2026-05 conditional novelty 6.0

    RAE v2 reaches gFID 1.06 on ImageNet-256 in 80 epochs by combining multi-layer encoder sums, complementary REPA targets, and free guidance via output reparameterization.

  11. FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning

    cs.LG 2025-10 conditional novelty 6.0

    An online imitation-learning method uses a flow-matching teacher's class-conditional loss as a reward and a regularizer to train a simple MLP policy, beating cloning and adversarial-imitation baselines on five of six tasks.

  12. DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation

    cs.RO 2025-09 conditional novelty 6.0

    DreamNav achieves new zero-shot SOTA on VLN-CE with an egocentric-only pipeline that generates candidate trajectories, imagines their futures, and selects the best by language alignment.

  13. Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight

    cs.RO 2025-01 unverdicted novelty 6.0

    DreamerV3 enables pixel-to-control policies for drone racing that reach 9 m/s in both simulation and real hardware-in-the-loop tests.

  14. Diffusion Policy Policy Optimization

    cs.RO 2024-09 unverdicted novelty 6.0

    DPPO fine-tunes diffusion policies via policy gradients and outperforms prior RL approaches for diffusion policies and PG-tuned alternatives on robot benchmarks while enabling stable training and hardware deployment.

  15. Octo: An Open-Source Generalist Robot Policy

    cs.RO 2024-05 unverdicted novelty 6.0

    Octo is an open-source transformer-based generalist robot policy pretrained on 800k trajectories that serves as an effective initialization for finetuning across diverse robotic platforms.

  16. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    cs.RO 2024-03 accept novelty 6.0

    DROID is a new 76k-trajectory in-the-wild robot manipulation dataset spanning 564 scenes and 84 tasks that improves policy performance and generalization when used for training.

  17. Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative Policies

    cs.LG 2026-05 unverdicted novelty 5.0

    LP-DS improves generative policies for imitation and RL by optimizing latent noise perturbations with a constrained Lagrangian objective, showing up to 25% better returns on manipulation and locomotion tasks.

  18. What Limits Vision-and-Language Navigation ?

    cs.RO 2026-05 unverdicted novelty 5.0

    StereoNav reaches new benchmark highs on R2R-CE and RxR-CE and improves real-robot reliability by supplying persistent target-location priors and stereo-derived geometry that stay stable under lighting changes and blur.