REVIEW 10 cited by
Reinforcement Learning for Versatile, Dynamic, and Robust Bipedal Locomotion Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper presents a comprehensive study on using deep reinforcement learning (RL) to create dynamic locomotion controllers for bipedal robots. Going beyond focusing on a single locomotion skill, we develop a general control solution that can be used for a range of dynamic bipedal skills, from periodic walking and running to aperiodic jumping and standing. Our RL-based controller incorporates a novel dual-history architecture, utilizing both a long-term and short-term input/output (I/O) history of the robot. This control architecture, when trained through the proposed end-to-end RL approach, consistently outperforms other methods across a diverse range of skills in both simulation and the real world. The study also delves into the adaptivity and robustness introduced by the proposed RL system in developing locomotion controllers. We demonstrate that the proposed architecture can adapt to both time-invariant dynamics shifts and time-variant changes, such as contact events, by effectively using the robot's I/O history. Additionally, we identify task randomization as another key source of robustness, fostering better task generalization and compliance to disturbances. The resulting control policies can be successfully deployed on Cassie, a torque-controlled human-sized bipedal robot. This work pushes the limits of agility for bipedal robots through extensive real-world experiments. We demonstrate a diverse range of locomotion skills, including: robust standing, versatile walking, fast running with a demonstration of a 400-meter dash, and a diverse set of jumping skills, such as standing long jumps and high jumps.
Forward citations
Cited by 10 Pith papers
-
Training and Evaluating Diffusion Policies with Long Context Lengths
Naive long-context Diffusion Policies succeed with UNet+Cross-Attention and sufficient data; variable-history training cuts sample complexity in the low-data regime.
-
DexTrack: Towards Generalizable Neural Tracking Control for Dexterous Manipulation from Human References
A neural controller combining RL and imitation learning on iteratively mined demonstrations tracks human kinematic references for dexterous manipulation, yielding over 10% higher success rates than prior baselines.
-
Learning to Hop for a Single-Legged Robot with Parallel Mechanism
A reinforcement learning policy trained on a simplified serial model, combined with a Jacobian-based torque conversion, enables continuous hopping of a parallel-mechanism single-legged robot in simulation and on hardware.
-
Bridging Adaptivity and Safety: Learning Agile Collision-Free Locomotion Across Varied Physics
A legged-robot controller that estimates payload and friction online and uses those estimates to switch between agile and recovery policies achieves lower collision rates and higher speeds than non-adaptive baselines.
-
ExBody2: Advanced Expressive Humanoid Whole-Body Control
A teacher-student whole-body tracking controller with automated motion-data filtering and specialist fine-tuning outperforms prior methods on a Unitree G1 humanoid.
-
Dynamic Tube MPC: Learning Tube Dynamics with Massively Parallel Simulation for Robust Safety in Practice
A neural network trained with massively parallel simulation predicts tracking-error tubes as a function of planning actions and error history, and Dynamic Tube MPC plans trajectories whose tube stays in free space on ...
-
Neural Internal Model Control: Learning a Robust Control Policy via Predictive Error Feedback
Neural internal model control adds a rigid-body predictive error, the gap between commanded and actual body motion, as a feedback signal to an RL policy, improving disturbance robustness on quadrotors and quadrupeds.
-
Hybrid Data-Driven Predictive Control for Robust and Reactive Exoskeleton Locomotion Synthesis
The paper claims a new HDDPC framework for reactive exoskeleton locomotion, but the supplied full text is a different paper on polymorphic crystals, leaving the claim unverified.
-
EMP: Executable Motion Prior for Humanoid Robot Standing Upper-body Motion Imitation
A state-conditioned executable motion prior network modifies upper-body motion targets so a humanoid can imitate human gestures while maintaining balance.
-
AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control
A hybrid trajectory-optimization and RL framework lets a humanoid robot flex its torso and legs to reach and manipulate objects beyond the range of prior controllers.
Discussion (0). Continue with ORCID to comment.