A three-stage RL framework with cross-attention, spatial memory, and depth-noise augmentation lets a Unitree Go2 climb 55° hollow stairs zero-shot from simulation.
hub
Robot parkour learning
16 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
fields
cs.RO 16roles
background 1polarities
background 1representative citing papers
SWAP embeds symmetry equivariance into world models and policies, enabling a quadruped to leap 2.13m gaps and climb 1.63m platforms with robust generalization to mirrored and outdoor terrains.
DreamTIP adds LLM-identified task-invariant properties as auxiliary targets in Dreamer's world model plus a mixed-replay adaptation step, delivering 28.1% average simulated transfer gains and 100% real-world climb success versus 10% for baselines.
A sim-to-real method that distills a privileged camera-based teacher policy into a tactile student policy, improving in-hand rotation and reorientation over proprioception-only policies.
A modular system uses motion matching to compose long-horizon human skill chains, trains RL experts, and distills them into a depth-based policy that lets a Unitree G1 humanoid autonomously climb, vault, and roll over obstacles up to 1.25 m tall.
DreamPolicy integrates an autoregressive diffusion world model with policy learning to produce a single scalable policy that generalizes to unseen composite terrains for humanoid locomotion.
A lightweight RL framework trains terrain-agnostic 3D foothold-tracking policies for humanoids that transfer directly to real-world use as standalone low-level controllers.
A hybrid motion-tracking and imitation-reinforcement pipeline produces a depth-based visuomotor policy that lets humanoids climb varied ladders zero-shot on hardware and perform teleoperated manipulation while climbing.
An end-to-end policy learns robust humanoid locomotion directly from noisy depth images via high-fidelity sensor simulation, vision-aware distillation from privileged maps, and terrain-specific multi-critic reward shaping.
UniCon standardizes states and control logic into modular execution graphs for efficient transfer of learning controllers across heterogeneous robots, with lower latency than ROS.
A four-stage RL system with teacher-student distillation and online constrained adaptation enables humanoid robots to achieve robust ball-kicking accuracy under noisy perception in simulation and on physical hardware.
MSDP pre-trains a transformer encoder with masked multisensory autoencoding, then uses an asymmetric actor-critic bridge (cross-attention for critic, pooling for actor) to accelerate and robustify contact-rich RL across simulation and real robots.
Empirical comparison of blind, critic-perceptive, and fully perceptive variants of morphology-aware RL locomotion controllers shows critic-only perception improves robustness over blind baselines while remaining more stable under perception noise than full perception.
RPG trains a unified humanoid robot policy using motion and temporal randomization to achieve smooth, stable transitions between fighting skills and locomotion.
Sparsely gated MoE policies double the success rate of a real Unitree Go2 quadruped on large-obstacle parkour versus matched-active-parameter MLP baselines while cutting inference time compared with a scaled-up MLP.
citing papers explorer
-
StairMaster: Learning to Conquer Risky Hollow Stairs for Agile Quadrupedal Robots
A three-stage RL framework with cross-attention, spatial memory, and depth-noise augmentation lets a Unitree Go2 climb 55° hollow stairs zero-shot from simulation.
-
SWAP: Symmetric Equivariant World-Model for Agile Robot Parkour
SWAP embeds symmetry equivariance into world models and policies, enabling a quadruped to leap 2.13m gaps and climb 1.63m platforms with robust generalization to mirrored and outdoor terrains.
-
Learning Task-Invariant Properties via Dreamer: Enabling Efficient Policy Transfer for Quadruped Robots
DreamTIP adds LLM-identified task-invariant properties as auxiliary targets in Dreamer's world model plus a mixed-replay adaptation step, delivering 28.1% average simulated transfer gains and 100% real-world climb success versus 10% for baselines.
-
PTLD: Sim-to-real Privileged Tactile Latent Distillation for Dexterous Manipulation
A sim-to-real method that distills a privileged camera-based teacher policy into a tactile student policy, improving in-hand rotation and reorientation over proprioception-only policies.
-
Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching
A modular system uses motion matching to compose long-horizon human skill chains, trains RL experts, and distills them into a depth-based policy that lets a Unitree G1 humanoid autonomously climb, vault, and roll over obstacles up to 1.25 m tall.
-
DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion
DreamPolicy integrates an autoregressive diffusion world model with policy learning to produce a single scalable policy that generalizes to unseen composite terrains for humanoid locomotion.
-
Mind Your Steps: A General Learning Framework for Accurate Humanoid Foothold Tracking
A lightweight RL framework trains terrain-agnostic 3D foothold-tracking policies for humanoids that transfer directly to real-world use as standalone low-level controllers.
-
LadderMan: Learning Humanoid Perceptive Ladder Climbing
A hybrid motion-tracking and imitation-reinforcement pipeline produces a depth-based visuomotor policy that lets humanoids climb varied ladders zero-shot on hardware and perform teleoperated manipulation while climbing.
-
Now You See That: Learning End-to-End Humanoid Locomotion from Raw Pixels
An end-to-end policy learns robust humanoid locomotion directly from noisy depth images via high-fidelity sensor simulation, vision-aware distillation from privileged maps, and terrain-specific multi-critic reward shaping.
-
UniCon: A Unified System for Efficient Robot Learning Transfers
UniCon standardizes states and control logic into modular execution graphs for efficient transfer of learning controllers across heterogeneous robots, with lower latency than ROS.
-
Learning Agile Striker Skills for Humanoid Soccer Robots from Noisy Sensory Input
A four-stage RL system with teacher-student distillation and online constrained adaptation enables humanoid robots to achieve robust ball-kicking accuracy under noisy perception in simulation and on physical hardware.
-
Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning
MSDP pre-trains a transformer encoder with masked multisensory autoencoding, then uses an asymmetric actor-critic bridge (cross-attention for critic, pooling for actor) to accelerate and robustify contact-rich RL across simulation and real robots.
-
Learning Perceptive Platform Adaptive Locomotion Controllers for Quadrupedal Robots
Empirical comparison of blind, critic-perceptive, and fully perceptive variants of morphology-aware RL locomotion controllers shows critic-only perception improves robustness over blind baselines while remaining more stable under perception noise than full perception.
-
RPG: Robust Policy Gating for Smooth Multi-Skill Transitions in Humanoid Fighting
RPG trains a unified humanoid robot policy using motion and temporal randomization to achieve smooth, stable transitions between fighting skills and locomotion.
-
Quadruped Parkour Learning: Sparsely Gated Mixture of Experts with Visual Input
Sparsely gated MoE policies double the success rate of a real Unitree Go2 quadruped on large-obstacle parkour versus matched-active-parameter MLP baselines while cutting inference time compared with a scaled-up MLP.
- GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning