REVIEW 4 cited by
LeLaN: Learning A Language-Conditioned Navigation Policy from In-the-Wild Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The world is filled with a wide variety of objects. For robots to be useful, they need the ability to find arbitrary objects described by people. In this paper, we present LeLaN(Learning Language-conditioned Navigation policy), a novel approach that consumes unlabeled, action-free egocentric data to learn scalable, language-conditioned object navigation. Our framework, LeLaN leverages the semantic knowledge of large vision-language models, as well as robotic foundation models, to label in-the-wild data from a variety of indoor and outdoor environments. We label over 130 hours of data collected in real-world indoor and outdoor environments, including robot observations, YouTube video tours, and human walking data. Extensive experiments with over 1000 real-world trials show that our approach enables training a policy from unlabeled action-free videos that outperforms state-of-the-art robot navigation methods, while being capable of inference at 4 times their speed on edge compute. We open-source our models, datasets and provide supplementary videos on our project page (https://learning-language-navigation.github.io/).
Forward citations
Cited by 4 Pith papers
-
Goal-oriented Navigation Instruction Generation with Tour Video Priors
VideoNIG tests whether multimodal models can turn tour videos into executable navigation instructions, and a two-stage curriculum with preference optimization improves their spatial reasoning.
-
VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision
VEGA reconstructs local geometry from monocular egocentric video to create supervised trajectories that train a flow-matching VLA policy, yielding lower collision rates on a new benchmark and in real-world tests.
-
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
SocialNav-SUB introduces a VQA benchmark for social robot navigation and shows current VLMs underperform rule-based and human-agreement baselines on spatial, spatiotemporal, and social reasoning questions.
-
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
Navigation policies trained on 2,000+ hours of web videos with visual odometry pseudo-labels achieve higher real-world urban navigation success than fine-tuned prior models.
Discussion (0). Continue with ORCID to comment.