REVIEW 2 cited by
Learning and Adapting Agile Locomotion Skills by Transferring Experience
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Legged robots have enormous potential in their range of capabilities, from navigating unstructured terrains to high-speed running. However, designing robust controllers for highly agile dynamic motions remains a substantial challenge for roboticists. Reinforcement learning (RL) offers a promising data-driven approach for automatically training such controllers. However, exploration in these high-dimensional, underactuated systems remains a significant hurdle for enabling legged robots to learn performant, naturalistic, and versatile agility skills. We propose a framework for training complex robotic skills by transferring experience from existing controllers to jumpstart learning new tasks. To leverage controllers we can acquire in practice, we design this framework to be flexible in terms of their source -- that is, the controllers may have been optimized for a different objective under different dynamics, or may require different knowledge of the surroundings -- and thus may be highly suboptimal for the target task. We show that our method enables learning complex agile jumping behaviors, navigating to goal locations while walking on hind legs, and adapting to new environments. We also demonstrate that the agile behaviors learned in this way are graceful and safe enough to deploy in the real world.
Forward citations
Cited by 2 Pith papers
-
Unsupervised Skill Discovery as Exploration for Learning Agile Locomotion
A robot training framework, SDAX, uses unsupervised skill discovery as an exploration signal with a learned balancing weight, letting a Unitree A1 learn leap, climb, crawl, and wall-jump behaviors in simulation and on...
-
APEX: Action Priors Enable Efficient Exploration for Robust Motion Tracking on Legged Robots
APEX trains gait-tracking policies with decaying action priors and separate style and task critics, achieving reference-free deployment, faster convergence, and reward-robustness over DeepMimic.
Discussion (0). Continue with ORCID to comment.