Pith. sign in

REVIEW 2 cited by

Actuator-Constrained Reinforcement Learning for High-Speed Quadrupedal Locomotion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.17507 v1 pith:MV3GYNKN submitted 2023-12-29 cs.RO

classification cs.RO
keywords motorquadrupedrobotactuatordesignedhigh-speedinfeasiblelearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper presents a method for achieving high-speed running of a quadruped robot by considering the actuator torque-speed operating region in reinforcement learning. The physical properties and constraints of the actuator are included in the training process to reduce state transitions that are infeasible in the real world due to motor torque-speed limitations. The gait reward is designed to distribute motor torque evenly across all legs, contributing to more balanced power usage and mitigating performance bottlenecks due to single-motor saturation. Additionally, we designed a lightweight foot to enhance the robot's agility. We observed that applying the motor operating region as a constraint helps the policy network avoid infeasible areas during sampling. With the trained policy, KAIST Hound, a 45 kg quadruped robot, can run up to 6.5 m/s, which is the fastest speed among electric motor-based quadruped robots.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots

    cs.RO 2025-09 conditional novelty 6.0 of 10

    PACE fits a compact set of actuator parameters from brief in-air data and trains energy-aware locomotion policies that transfer zero-shot to real quadrupeds without dynamics randomization.

  2. High-Performance Reinforcement Learning on Spot: Optimizing Simulation Parameters with Distributional Measures

    cs.LG 2025-04 conditional novelty 6.0 of 10

    An RL policy trained in Isaac Sim and tuned with Wasserstein/MMD distributional gap optimization runs on Spot at over 5.2 m/s, tripling the stock controller's speed.

Pith tools