Pith. sign in

REVIEW 7 cited by

GNM: A General Navigation Model to Drive Any Robot

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.03370 v2 pith:3XRM2KEI submitted 2022-10-07 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords navigationrobotsdatatrainedincludingmodelacrossbroad
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning provides a powerful tool for vision-based navigation, but the capabilities of learning-based policies are constrained by limited training data. If we could combine data from all available sources, including multiple kinds of robots, we could train more powerful navigation models. In this paper, we study how a general goal-conditioned model for vision-based navigation can be trained on data obtained from many distinct but structurally similar robots, and enable broad generalization across environments and embodiments. We analyze the necessary design decisions for effective data sharing across robots, including the use of temporal context and standardized action spaces, and demonstrate that an omnipolicy trained from heterogeneous datasets outperforms policies trained on any single dataset. We curate 60 hours of navigation trajectories from 6 distinct robots, and deploy the trained GNM on a range of new robots, including an underactuated quadrotor. We find that training on diverse data leads to robustness against degradation in sensing and actuation. Using a pre-trained navigation model with broad generalization capabilities can bootstrap applications on novel robots going forward, and we hope that the GNM represents a step in that direction. For more information on the datasets, code, and videos, please check out our project page https://sites.google.com/view/drive-any-robot.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  2. Learning to Navigate Efficiently with Only 0.58M Trainable Parameters

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Decomposed navigation with analytical geometry interfaces and three small learned modules (0.58M trainable params) approaches SOTA point-goal performance at 50 Hz with lowest collisions.

  3. Approximate Imitation Learning for Event-based Quadrotor Flight in Cluttered Environments

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Approximate imitation learning trains event-to-control quadrotor policies 28× faster by freezing a pretrained event encoder and fine-tuning a shared action decoder via a state-based approximate student, matching onlin...

  4. CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Counterfactual language-action relabeling raises instruction-following success from about 26% to 53% in real-world navigation tests.

  5. Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Narrate2Nav uses Barlow Twins alignment to distill language-based reasoning from a large teacher into a small RGB-only navigation model, reporting lower trajectory error and higher goal-reaching success than four baselines.

  6. PixelNav: Towards Model-based Vision-Only Navigation with Topological Graphs

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A camera-only navigation system combining visual place recognition, traversability segmentation, and model predictive control over a topological graph, evaluated on a real robot.

  7. Spatially-Enhanced Recurrent Memory for Long-Range Mapless Navigation via End-to-End Reinforcement Learning

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A modified recurrent unit with an input-multiplied gate improves spatial memory and long-range mapless navigation success rates by about 23.5% over standard RNNs in simulation and transfers zero-shot to a real robot.

Pith tools