Pith. sign in

REVIEW 2 cited by

A Workflow for Offline Model-Free Robotic Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.10813 v2 pith:HMKVI4TM submitted 2021-09-22 cs.LG

classification cs.LG
keywords learningofflineonlinewithoutworkflowpoliciesalgorithmarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Offline reinforcement learning (RL) enables learning control policies by utilizing only prior experience, without any online interaction. This can allow robots to acquire generalizable skills from large and diverse datasets, without any costly or unsafe online data collection. Despite recent algorithmic advances in offline RL, applying these methods to real-world problems has proven challenging. Although offline RL methods can learn from prior data, there is no clear and well-understood process for making various design choices, from model architecture to algorithm hyperparameters, without actually evaluating the learned policies online. In this paper, our aim is to develop a practical workflow for using offline RL analogous to the relatively well-understood workflows for supervised learning problems. To this end, we devise a set of metrics and conditions that can be tracked over the course of offline training, and can inform the practitioner about how the algorithm and model architecture should be adjusted to improve final performance. Our workflow is derived from a conceptual understanding of the behavior of conservative offline RL algorithms and cross-validation in supervised learning. We demonstrate the efficacy of this workflow in producing effective policies without any online tuning, both in several simulated robotic learning scenarios and for three tasks on two distinct real robots, focusing on learning manipulation skills with raw image observations with sparse binary rewards. Explanatory video and additional results can be found at sites.google.com/view/offline-rl-workflow

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

    cs.RO 2025-04 conditional novelty 6.0 of 10

    Pre-training on RoboTwin's generative digital twins and fine-tuning on 20 real demonstrations raises dual-arm task success from about 20% to 62% and single-arm success from about 1% to 72%.

  2. ACL-QL: Adaptive Conservative Level in Q-Learning for Offline Reinforcement Learning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    ACL-QL reports state-of-the-art D4RL results with per-transition adaptive conservatism in Q-learning, but its surrogate losses and CQL anchor are not rigorously justified.

Pith tools