Pith. sign in

REVIEW 22 cited by

Parkour in the Wild: Learning a General and Extensible Agile Locomotion Policy Using Multi-expert Distillation and RL Fine-tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.11164 v1 pith:D4SWBTLL submitted 2025-05-16 cs.RO

Parkour in the Wild: Learning a General and Extensible Agile Locomotion Policy Using Multi-expert Distillation and RL Fine-tuning

classification cs.RO
keywords policylocomotionfine-tuningleggedrobotsterrainsacrossagile
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Legged robots are well-suited for navigating terrains inaccessible to wheeled robots, making them ideal for applications in search and rescue or space exploration. However, current control methods often struggle to generalize across diverse, unstructured environments. This paper introduces a novel framework for agile locomotion of legged robots by combining multi-expert distillation with reinforcement learning (RL) fine-tuning to achieve robust generalization. Initially, terrain-specific expert policies are trained to develop specialized locomotion skills. These policies are then distilled into a unified foundation policy via the DAgger algorithm. The distilled policy is subsequently fine-tuned using RL on a broader terrain set, including real-world 3D scans. The framework allows further adaptation to new terrains through repeated fine-tuning. The proposed policy leverages depth images as exteroceptive inputs, enabling robust navigation across diverse, unstructured terrains. Experimental results demonstrate significant performance improvements over existing methods in synthesizing multi-terrain skills into a single controller. Deployment on the ANYmal D robot validates the policy's ability to navigate complex environments with agility and robustness, setting a new benchmark for legged robot locomotion.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Light-Loco-Parkour: Versatile Perceptive Whole-Body Locomotion via Multi-Skill Distillation

    cs.RO 2026-08 conditional novelty 7.0

    A single neural-network policy, trained in simulation, makes a humanoid climb, vault, and traverse uneven terrain from onboard depth and a velocity command, with no skill labels or runtime motion graphs.

  2. HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning

    cs.RO 2026-06 unverdicted novelty 7.0

    HumanoidArena is a new benchmark of 7 leg-critical HOI/HSI tasks that evaluates egocentric hierarchical whole-body policies in humanoids and finds performance is strongly conditioned on the low-level GMT used.

  3. roto 2.0: The Robot Tactile Olympiad

    cs.RO 2026-05 unverdicted novelty 7.0

    roto 2.0 provides a standardized benchmark for end-to-end blind tactile RL on 16-24 DOF robots, with open-sourced baselines achieving 13 Baoding ball rotations in 10 seconds.

  4. Extreme-RGMT: Continual Learning of Highly Dynamic Skills for Robust Generalist Humanoid Control

    cs.RO 2026-07 conditional novelty 6.0

    A two-stage continual-learning framework lets a generalist humanoid tracking policy acquire highly dynamic acrobatic skills while preserving its general-purpose motion capabilities.

  5. EgoHTR: Egocentric 4D Demonstrations of Human Terrain Traversal

    cs.RO 2026-07 conditional novelty 6.0

    EgoHTR is a 55-sequence, 150k-frame egocentric 4D human-terrain dataset with a reconstruction pipeline, MoCap-validated benchmark, and perceptive locomotion policies deployed on a Unitree G1.

  6. Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing

    cs.RO 2026-06 unverdicted novelty 6.0

    Foot-mounted proximity sensors provide pre-contact feedback that, when integrated into RL, improves quadruped traversal robustness on discrete terrain with reliable sim-to-real transfer.

  7. StairMaster: Learning to Conquer Risky Hollow Stairs for Agile Quadrupedal Robots

    cs.RO 2026-06 conditional novelty 6.0

    A three-stage RL framework with cross-attention, spatial memory, and depth-noise augmentation lets a Unitree Go2 climb 55° hollow stairs zero-shot from simulation.

  8. StairMaster: Learning to Conquer Risky Hollow Stairs for Agile Quadrupedal Robots

    cs.RO 2026-06 unverdicted novelty 6.0

    StairMaster trains an RL policy that lets a Unitree Go2 quadruped climb hollow stairs up to 55 degrees via zero-shot sim-to-real transfer using cross-attention, SRU memory, and active-perception rewards.

  9. SWAP: Symmetric Equivariant World-Model for Agile Robot Parkour

    cs.RO 2026-06 unverdicted novelty 6.0

    SWAP embeds symmetry equivariance into world models and policies, enabling a quadruped to leap 2.13m gaps and climb 1.63m platforms with robust generalization to mirrored and outdoor terrains.

  10. TAGA: Terrain-aware Active Gaze Learning for Generalizable Agile Humanoid Locomotion

    cs.RO 2026-06 unverdicted novelty 6.0

    TAGA learns terrain-aware active gaze behaviors for humanoid robots via RL alone, enabling generalizable locomotion with 1.2m real-world gap traversal.

  11. Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation

    cs.RO 2026-02 conditional novelty 6.0

    MPAIL2 demonstrates real-world manipulation learning from observation alone, without rewards or action labels, plus positive online transfer.

  12. Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching

    cs.RO 2026-02 unverdicted novelty 6.0

    A modular system uses motion matching to compose long-horizon human skill chains, trains RL experts, and distills them into a depth-based policy that lets a Unitree G1 humanoid autonomously climb, vault, and roll over...

  13. Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning

    cs.RO 2025-11 unverdicted novelty 6.0

    Isaac Lab is a unified GPU-native platform combining high-fidelity physics, photorealistic rendering, multi-frequency sensors, domain randomization, and learning pipelines for scalable multi-modal robot policy training.

  14. Continual-RL for Generalization in Autonomous Racing on the RoboRacer Platform

    cs.RO 2026-07 conditional novelty 5.0

    SAC plus Continual Backpropagation, trained only on real multi-track data, fine-tunes in ~15 minutes on an unseen lower-friction RoboRacer track and outperforms MAP and MPC.

  15. Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing

    cs.RO 2026-06 unverdicted novelty 5.0

    Foot-mounted infrared proximity sensors supply pre-contact signals that, when incorporated into RL, improve quadrupedal traversal of discrete terrain such as gaps and stepping stones.

  16. LadderMan: Learning Humanoid Perceptive Ladder Climbing

    cs.RO 2026-06 unverdicted novelty 5.0

    A hybrid motion-tracking and imitation-reinforcement pipeline produces a depth-based visuomotor policy that lets humanoids climb varied ladders zero-shot on hardware and perform teleoperated manipulation while climbing.

  17. CoRe-MoE: Contrastive Reweighted Mixture of Experts for Multi-Terrain Humanoid Locomotion with Gait Adaptation

    cs.RO 2026-06 unverdicted novelty 5.0

    CoRe-MoE uses a two-stage RL framework with contrastive reweighting in a Mixture-of-Experts architecture to enable gait transitions and multi-terrain adaptation for humanoid locomotion.

  18. SSR: Scaling Surefooted and Symmetric Humanoid Traversal to the Open World

    cs.RO 2026-05 unverdicted novelty 5.0

    SSR is an end-to-end vision-based framework for humanoid traversal that learns imagined foothold guidance, equivariant latent-space symmetry augmentation, and terrain-specific multi-discriminator motion priors to enab...

  19. DPL: Depth-only Perceptive Humanoid Locomotion via Realistic Depth Synthesis and Cross-Attention Terrain Reconstruction

    cs.RO 2025-10 conditional novelty 5.0

    Combining a blind-backbone policy, cross-attention terrain reconstruction from depth plus proprioception, and realistic synthetic depth with noise enables depth-only full-sized humanoid locomotion over stairs, slopes,...

  20. Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient

    cs.RO 2026-05 unverdicted novelty 4.0

    SDPG is a new on-policy visual RL algorithm that estimates gradients via stochastic perturbations of rollouts, achieving faster training and lower memory use than baselines on visual MuJoCo tasks while adding new robo...

  21. Evaluation of an Actuated Spine in Agile Quadruped Locomotion

    cs.RO 2026-05 unverdicted novelty 4.0

    Adding an actuated sagittal spine to a simulated quadruped increases agility and allows it to clear higher obstacles, steeper slopes, and tighter passages than the rigid-spine baseline.

  22. Quadruped Parkour Learning: Sparsely Gated Mixture of Experts with Visual Input

    cs.RO 2026-04 unverdicted novelty 4.0

    Sparsely gated MoE policies double the success rate of a real Unitree Go2 quadruped on large-obstacle parkour versus matched-active-parameter MLP baselines while cutting inference time compared with a scaled-up MLP.