Pith. sign in

REVIEW 1 cited by

Zero-Shot Terrain Generalization for Visual Locomotion Policies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.05513 v1 pith:2Y4ZQ3VN submitted 2020-11-11 cs.RO

classification cs.RO
keywords challengelearninglocomotioncontrollersenvironmentsgeneralizationterrainterrains
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Legged robots have unparalleled mobility on unstructured terrains. However, it remains an open challenge to design locomotion controllers that can operate in a large variety of environments. In this paper, we address this challenge of automatically learning locomotion controllers that can generalize to a diverse collection of terrains often encountered in the real world. We frame this challenge as a multi-task reinforcement learning problem and define each task as a type of terrain that the robot needs to traverse. We propose an end-to-end learning approach that makes direct use of the raw exteroceptive inputs gathered from a simulated 3D LiDAR sensor, thus circumventing the need for ground-truth heightmaps or preprocessing of perception information. As a result, the learned controller demonstrates excellent zero-shot generalization capabilities and can navigate 13 different environments, including stairs, rugged land, cluttered offices, and indoor spaces with humans.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hierarchical Reinforcement Learning and Value Optimization for Challenging Quadruped Locomotion

    cs.RO 2025-06 conditional novelty 4.0 of 10

    A hierarchical quadruped controller uses online optimization over the low-level policy's value function to choose footstep targets, improving normalized reward and reducing collisions over an end-to-end baseline witho...

Pith tools