Pith. sign in

REVIEW 4 cited by

Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.07517 v1 pith:CZWBDHE2 submitted 2019-01-22 cs.RO cs.AIcs.LGcs.SYeess.SY

classification cs.ROcs.AIcs.LGcs.SYeess.SY
keywords quadrupedalrecoverybehaviorscontrollerrobotanymalapproachdeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability to recover from a fall is an essential feature for a legged robot to navigate in challenging environments robustly. Until today, there has been very little progress on this topic. Current solutions mostly build upon (heuristically) predefined trajectories, resulting in unnatural behaviors and requiring considerable effort in engineering system-specific components. In this paper, we present an approach based on model-free Deep Reinforcement Learning (RL) to control recovery maneuvers of quadrupedal robots using a hierarchical behavior-based controller. The controller consists of four neural network policies including three behaviors and one behavior selector to coordinate them. Each of them is trained individually in simulation and deployed directly on a real system. We experimentally validate our approach on the quadrupedal robot ANYmal, which is a dog-sized quadrupedal system with 12 degrees of freedom. With our method, ANYmal manifests dynamic and reactive recovery behaviors to recover from an arbitrary fall configuration within less than 5 seconds. We tested the recovery maneuver more than 100 times, and the success rate was higher than 97 %.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MuJoCo Playground

    cs.RO 2025-02 conditional novelty 7.0 of 10

    An open-source, MJX-based robot learning framework with integrated batch rendering that provides fast training and demonstrates sim-to-real transfer on six robot platforms.

  2. SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An unbiased PPO gradient that ignores recovery-policy density, plus analytic recovery values and success-gated imitation, cuts training falls by 26–233× on locomotion tasks without sacrificing reward.

  3. Learning Humanoid Standing-up Control across Diverse Postures

    cs.RO 2025-02 conditional novelty 6.0 of 10

    HoST uses multi-critic reinforcement learning, a force curriculum, and smoothness constraints in simulation so a Unitree G1 humanoid can stand up from diverse postures in the real world without predefined motion trajectories.

  4. Sensor-Space Based Robust Kinematic Control of Redundant Soft Manipulator by Learning

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A dual learning framework, RL in simulation plus GAIL from human demonstrations, with a pre-calibrated sim-to-real step, achieves kinematic control of a soft manipulator under loads and in confined pipes.

Pith tools