Pith. sign in

REVIEW 23 cited by

Learning Agile Robotic Locomotion Skills by Imitating Animals

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.00784 v3 pith:I7JTHJZ2 submitted 2020-04-02 cs.RO cs.LG

classification cs.ROcs.LG
keywords agilebehaviorscontrollerslearninglocomotionableanimalsskills
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Reproducing the diverse and agile locomotion skills of animals has been a longstanding challenge in robotics. While manually-designed controllers have been able to emulate many complex behaviors, building such controllers involves a time-consuming and difficult development process, often requiring substantial expertise of the nuances of each skill. Reinforcement learning provides an appealing alternative for automating the manual effort involved in the development of controllers. However, designing learning objectives that elicit the desired behaviors from an agent can also require a great deal of skill-specific expertise. In this work, we present an imitation learning system that enables legged robots to learn agile locomotion skills by imitating real-world animals. We show that by leveraging reference motion data, a single learning-based approach is able to automatically synthesize controllers for a diverse repertoire behaviors for legged robots. By incorporating sample efficient domain adaptation techniques into the training process, our system is able to learn adaptive policies in simulation that can then be quickly adapted for real-world deployment. To demonstrate the effectiveness of our system, we train an 18-DoF quadruped robot to perform a variety of agile behaviors ranging from different locomotion gaits to dynamic hops and turns.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Sensorimotor Control by Imitating Predictive Models of Human Motion

    cs.RO 2025-08 conditional novelty 7.0 of 10

    A predictive model of human hand motion, trained on human interaction data, can reward a robot policy for tracking predicted future keypoints and enable learning of dexterous manipulation from sparse rewards.

  2. What Matters in Humanoid General Motion Tracking? An Empirical Study

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A controlled ablation of humanoid motion-tracking pipelines shows that explicit reference joint velocities and a short observation history improve tracking, while residual actions and teacher-student training yield on...

  3. StairMaster: Learning to Conquer Risky Hollow Stairs for Agile Quadrupedal Robots

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    StairMaster trains an RL policy that lets a Unitree Go2 quadruped climb hollow stairs up to 55 degrees via zero-shot sim-to-real transfer using cross-attention, SRU memory, and active-perception rewards.

  4. DynaRetarget: Dynamically-Feasible Retargeting using Sampling-Based Trajectory Optimization

    cs.RO 2026-02 unverdicted novelty 6.0 of 10

    DynaRetarget's SBTO, which incrementally extends the optimization horizon while warm-starting from shorter solutions, refines kinematic humanoid demonstrations into dynamically consistent whole-body motions with highe...

  5. Dynamic Policy Learning for Legged Robot with Simplified Model Pretraining and Model-Homotopy-Inspired Transfer

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A model-homotopy curriculum that gradually redistributes mass and inertia from a single-rigid-body model to full-body dynamics lets a quadruped learn flips and wall-assisted maneuvers faster and more stably than direc...

  6. Sampling Strategies for Robust Universal Quadrupedal Locomotion Policies

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A particle-filter-based adaptive sampling of morphologies and wide PD-gain randomization yields a single quadruped locomotion policy that transfers zero-shot to ANYmal hardware.

  7. Musculoskeletal simulation of limb movement biomechanics in Drosophila melanogaster

    q-bio.NC 2025-09 conditional novelty 6.0 of 10

    An anatomically grounded, data-driven muscle model of Drosophila legs is built in OpenSim and MuJoCo and used to replay behaviors, predict synergies, and test passive joint effects.

  8. Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots

    cs.RO 2025-09 conditional novelty 6.0 of 10

    PACE fits a compact set of actuator parameters from brief in-air data and trains energy-aware locomotion policies that transfer zero-shot to real quadrupeds without dynamics randomization.

  9. Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A force-sensing arm acts as teacher for a small humanoid, enabling 20-minute real-world walking speed adaptation and 15-minute swing-up learning from scratch.

  10. CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Causal Diffusion Policy adds historical action conditioning and attention cache sharing to diffusion-based robot policies, improving success rates on most tested manipulation tasks under degraded observations.

  11. CARoL: Context-aware Adaptation for Robot Learning

    cs.RO 2025-06 conditional novelty 6.0 of 10

    CARoL measures task similarity by state-transition prediction errors and uses those similarities to weight prior policies, value functions, or actor-critic knowledge when adapting to a new task.

  12. Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A few-shot LLM planner for ALFRED that relabels suboptimal trajectories with hindsight prompts reaches 25.51 SR on Test Seen, approaching or beating the full-shot HLSM baseline.

  13. Reinforcement Learning from Wild Animal Videos

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A quadruped robot acquires walking, jumping, running-like, and standing skills using only the output of a video classifier trained on wild-animal videos as its reinforcement learning reward.

  14. AffordDP: Generalizable Diffusion Policy with Transferable Affordance

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A diffusion-based manipulation policy conditioned on transferred 3D contact points and post-contact trajectories, with adaptive affordance-guided sampling, generalizes to unseen object instances and categories.

  15. PRISM: Polynomial Representations for Interaction-Structured Motor Control

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Explicit low-degree factorized polynomial proprioceptive features improve robot RL and imitation policies beyond matched-capacity MLPs and induce sensorless compliance-like contact behavior in simulation.

  16. ObjRetarget: An Object-Aware Motion Retargeting Framework with Anthropomorphic Arm Constraints and Polyhedral Hand Modeling

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Decoupled arm–hand retargeting with anthropomorphic arm-plane constraints and polyhedral contact invariants raises real-robot dexterous-task success to 75.8% versus 61.6% and 50.8% for OKAMI and ORION.

  17. Learning to Dock: A Simulation-based Study on Closing the Sim2Real Gap in Autonomous Underwater Docking

    cs.RO 2025-06 conditional novelty 5.0 of 10

    In a simulation study, a naively trained AUV docking policy is competitive with domain-randomized and history-conditioned policies under payload variation, with robustness tricks giving only marginal gains in extreme cases.

  18. Reference Free Platform Adaptive Locomotion for Quadrupedal Robots using a Dynamics Conditioned Policy

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A single dynamics-conditioned RL policy transfers zero-shot across quadrupeds from 12 kg to 50 kg, and diverse reference robots during training clearly improve tracking.

  19. ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills

    cs.RO 2025-02 conditional novelty 5.0 of 10

    ASAP trains a residual action model on real-world rollouts and fine-tunes simulation policies through it, reducing humanoid whole-body motion tracking error in sim-to-real transfer.

  20. FDPP: Fine-tune Diffusion Policy with Human Preference

    cs.RO 2025-01 conditional novelty 5.0 of 10

    FDPP aligns a pre-trained diffusion policy with human preferences by training a reward model from preference labels and fine-tuning the policy with RL plus KL regularization.

  21. Inverse Delayed Reinforcement Learning

    cs.LG 2024-12 reject novelty 5.0 of 10

    The paper proposes an off-policy adversarial IRL method on augmented delayed states and claims, via a Lipschitz bound and MuJoCo experiments, that this outperforms IRL on raw delayed observations.

  22. Hierarchical Reinforcement Learning and Value Optimization for Challenging Quadruped Locomotion

    cs.RO 2025-06 conditional novelty 4.0 of 10

    A hierarchical quadruped controller uses online optimization over the low-level policy's value function to choose footstep targets, improving normalized reward and reducing collisions over an end-to-end baseline witho...

  23. Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

    cs.RO 2025-07 conditional novelty 3.0 of 10

    A diffusion-based visuomotor policy gains 3D and 4D scene awareness from a dynamic Gaussian world model, improving simulated and real robot manipulation success rates.

Pith tools