Pith. sign in

REVIEW 1 cited by

A Real-World Quadrupedal Locomotion Benchmark for Offline Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.16718 v1 pith:I5REBIGJ submitted 2023-09-13 cs.RO cs.LG

classification cs.ROcs.LG
keywords learninglocomotiontasksalgorithmsreinforcementroboticbenchmarkchallenging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Online reinforcement learning (RL) methods are often data-inefficient or unreliable, making them difficult to train on real robotic hardware, especially quadruped robots. Learning robotic tasks from pre-collected data is a promising direction. Meanwhile, agile and stable legged robotic locomotion remains an open question in their general form. Offline reinforcement learning (ORL) has the potential to make breakthroughs in this challenging field, but its current bottleneck lies in the lack of diverse datasets for challenging realistic tasks. To facilitate the development of ORL, we benchmarked 11 ORL algorithms in the realistic quadrupedal locomotion dataset. Such dataset is collected by the classic model predictive control (MPC) method, rather than the model-free online RL method commonly used by previous benchmarks. Extensive experimental results show that the best-performing ORL algorithms can achieve competitive performance compared with the model-free RL, and even surpass it in some tasks. However, there is still a gap between the learning-based methods and MPC, especially in terms of stability and rapid adaptation. Our proposed benchmark will serve as a development platform for testing and evaluating the performance of ORL algorithms in real-world legged locomotion tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Offline Adaptation of Quadruped Locomotion using Diffusion Models

    cs.RO 2024-11 conditional novelty 6.0 of 10

    A diffusion-based quadruped locomotion policy uses classifier-free guidance to adapt to new velocity-tracking rewards after training, enabling multi-skill interpolation and onboard CPU deployment.

Pith tools