Pith. sign in

REVIEW 4 cited by

LeLaN: Learning A Language-Conditioned Navigation Policy from In-the-Wild Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03603 v1 pith:7XQOHKZR submitted 2024-10-04 cs.RO

classification cs.RO
keywords datanavigationlanguage-conditionedlelanmodelspolicyvideosaction-free
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The world is filled with a wide variety of objects. For robots to be useful, they need the ability to find arbitrary objects described by people. In this paper, we present LeLaN(Learning Language-conditioned Navigation policy), a novel approach that consumes unlabeled, action-free egocentric data to learn scalable, language-conditioned object navigation. Our framework, LeLaN leverages the semantic knowledge of large vision-language models, as well as robotic foundation models, to label in-the-wild data from a variety of indoor and outdoor environments. We label over 130 hours of data collected in real-world indoor and outdoor environments, including robot observations, YouTube video tours, and human walking data. Extensive experiments with over 1000 real-world trials show that our approach enables training a policy from unlabeled action-free videos that outperforms state-of-the-art robot navigation methods, while being capable of inference at 4 times their speed on edge compute. We open-source our models, datasets and provide supplementary videos on our project page (https://learning-language-navigation.github.io/).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Goal-oriented Navigation Instruction Generation with Tour Video Priors

    cs.CV 2026-08 conditional novelty 6.0 of 10

    VideoNIG tests whether multimodal models can turn tour videos into executable navigation instructions, and a two-stage curriculum with preference optimization improves their spatial reasoning.

  2. VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    VEGA reconstructs local geometry from monocular egocentric video to create supervised trajectories that train a flow-matching VLA policy, yielding lower collision rates on a new benchmark and in real-world tests.

  3. SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    SocialNav-SUB introduces a VQA benchmark for social robot navigation and shows current VLMs underperform rule-based and human-agreement baselines on spatial, spatiotemporal, and social reasoning questions.

  4. CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Navigation policies trained on 2,000+ hours of web videos with visual odometry pseudo-labels achieve higher real-world urban navigation success than fine-tuned prior models.

Pith tools