REVIEW 4 cited by
Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The ability to recover from a fall is an essential feature for a legged robot to navigate in challenging environments robustly. Until today, there has been very little progress on this topic. Current solutions mostly build upon (heuristically) predefined trajectories, resulting in unnatural behaviors and requiring considerable effort in engineering system-specific components. In this paper, we present an approach based on model-free Deep Reinforcement Learning (RL) to control recovery maneuvers of quadrupedal robots using a hierarchical behavior-based controller. The controller consists of four neural network policies including three behaviors and one behavior selector to coordinate them. Each of them is trained individually in simulation and deployed directly on a real system. We experimentally validate our approach on the quadrupedal robot ANYmal, which is a dog-sized quadrupedal system with 12 degrees of freedom. With our method, ANYmal manifests dynamic and reactive recovery behaviors to recover from an arbitrary fall configuration within less than 5 seconds. We tested the recovery maneuver more than 100 times, and the success rate was higher than 97 %.
Forward citations
Cited by 4 Pith papers
-
MuJoCo Playground
An open-source, MJX-based robot learning framework with integrated batch rendering that provides fast training and demonstrates sim-to-real transfer on six robot platforms.
-
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions
An unbiased PPO gradient that ignores recovery-policy density, plus analytic recovery values and success-gated imitation, cuts training falls by 26–233× on locomotion tasks without sacrificing reward.
-
Learning Humanoid Standing-up Control across Diverse Postures
HoST uses multi-critic reinforcement learning, a force curriculum, and smoothness constraints in simulation so a Unitree G1 humanoid can stand up from diverse postures in the real world without predefined motion trajectories.
-
Sensor-Space Based Robust Kinematic Control of Redundant Soft Manipulator by Learning
A dual learning framework, RL in simulation plus GAIL from human demonstrations, with a pre-calibrated sim-to-real step, achieves kinematic control of a soft manipulator under loads and in confined pipes.
Discussion (0). Continue with ORCID to comment.