Foot-mounted proximity sensors provide pre-contact feedback that, when integrated into RL, improves quadruped traversal robustness on discrete terrain with reliable sim-to-real transfer.
hub
Learning to walk in minutes using massively parallel deep reinforcement learning
10 Pith papers cite this work. Polarity classification is still indexing.
hub tools
representative citing papers
CLF-guided RL yields exponentially stable optimal controllers, with proofs in continuous and discrete time, numerical checks on double integrator and cart-pole, and implementation on a walking humanoid.
CART learns a vision–proprioception terrain context and uses Temporal Sequence Selection to cut base oscillation by up to 41% in simulation and 22% on Spot outdoors, with a 5% higher sim success rate.
BIFROST learns invariant latent states via cross-domain bisimulation on paired data to enable zero-shot sim2real policy transfer for visual navigation, contact-rich manipulation, and visual servoing.
A thermal residual policy on a pre-trained quadruped locomotion controller prevents motor overheating under payload while preserving performance, lasting over 13 minutes on a Unitree A1 versus ~5 minutes for the nominal policy.
A single reinforcement learning policy jointly trains multiple locomotion skills for wheeled-legged robots with DC-motor constraints and learns a proprioceptive skill selector for adaptive behavior.
Empirical comparison shows a clear sim-to-real gap in reset-free RL for agile driving: TD-MPC2 outperforms the MPPI baseline in the real world while SAC excels in simulation, and residual learning benefits simulation but does not transfer.
A four-stage RL system with teacher-student distillation and online constrained adaptation enables humanoid robots to achieve robust ball-kicking accuracy under noisy perception in simulation and on physical hardware.
FastDSAC adds a truncated Gaussian policy constraint to distributional actor-critic methods to preserve network plasticity and accelerate training for scalable humanoid locomotion in parallel sampling setups.
Sparsely gated MoE policies double the success rate of a real Unitree Go2 quadruped on large-obstacle parkour versus matched-active-parameter MLP baselines while cutting inference time compared with a scaled-up MLP.
citing papers explorer
-
Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing
Foot-mounted proximity sensors provide pre-contact feedback that, when integrated into RL, improves quadruped traversal robustness on discrete terrain with reliable sim-to-real transfer.
-
Stability of Control Lyapunov Function Guided Reinforcement Learning
CLF-guided RL yields exponentially stable optimal controllers, with proofs in continuous and discrete time, numerical checks on double integrator and cart-pole, and implementation on a walking humanoid.
-
CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots
CART learns a vision–proprioception terrain context and uses Temporal Sequence Selection to cut base oscillation by up to 41% in simulation and 22% on Spot outdoors, with a 5% higher sim success rate.
-
BIFROST: Bridging Invariant Feature Representation for Observation-space Sim2Real Transfer
BIFROST learns invariant latent states via cross-domain bisimulation on paired data to enable zero-shot sim2real policy transfer for visual navigation, contact-rich manipulation, and visual servoing.
-
Learning to Balance Motor Thermal Safety and Quadrupedal Locomotion Performance with Residual Policy
A thermal residual policy on a pre-trained quadruped locomotion controller prevents motor overheating under payload while preserving performance, lasting over 13 minutes on a Unitree A1 versus ~5 minutes for the nominal policy.
-
MUJICA: Multi-skill Unified Joint Integration of Control Architecture for Wheeled-Legged Robots
A single reinforcement learning policy jointly trains multiple locomotion skills for wheeled-legged robots with DC-motor constraints and learns a proprioceptive skill selector for adaptive behavior.
-
Reset-Free Reinforcement Learning for Real-World Agile Driving: An Empirical Study
Empirical comparison shows a clear sim-to-real gap in reset-free RL for agile driving: TD-MPC2 outperforms the MPPI baseline in the real world while SAC excels in simulation, and residual learning benefits simulation but does not transfer.
-
Learning Agile Striker Skills for Humanoid Soccer Robots from Noisy Sensory Input
A four-stage RL system with teacher-student distillation and online constrained adaptation enables humanoid robots to achieve robust ball-kicking accuracy under noisy perception in simulation and on physical hardware.
-
FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion
FastDSAC adds a truncated Gaussian policy constraint to distributional actor-critic methods to preserve network plasticity and accelerate training for scalable humanoid locomotion in parallel sampling setups.
-
Quadruped Parkour Learning: Sparsely Gated Mixture of Experts with Visual Input
Sparsely gated MoE policies double the success rate of a real Unitree Go2 quadruped on large-obstacle parkour versus matched-active-parameter MLP baselines while cutting inference time compared with a scaled-up MLP.