Pith. sign in

REVIEW 1 cited by

Keep Doing What Worked: Behavioral Modelling Priors for Offline Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.08396 v3 pith:64K7YGDH submitted 2020-02-19 cs.LG cs.ROstat.ML

classification cs.LGcs.ROstat.ML
keywords learningalgorithmscontrolbatchbehaviorcontinuousoff-policyreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Off-policy reinforcement learning algorithms promise to be applicable in settings where only a fixed data-set (batch) of environment interactions is available and no new experience can be acquired. This property makes these algorithms appealing for real world problems such as robot control. In practice, however, standard off-policy algorithms fail in the batch setting for continuous control. In this paper, we propose a simple solution to this problem. It admits the use of data generated by arbitrary behavior policies and uses a learned prior -- the advantage-weighted behavior model (ABM) -- to bias the RL policy towards actions that have previously been executed and are likely to be successful on the new task. Our method can be seen as an extension of recent work on batch-RL that enables stable learning from conflicting data-sources. We find improvements on competitive baselines in a variety of RL tasks -- including standard continuous control benchmarks and multi-task learning for simulated and real-world robots.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Offline Diffusion Policy for Multi-User Delay-Constrained Scheduling

    cs.AI 2025-01 conditional novelty 4.0 of 10

    SOCD trains a diffusion-based scheduling policy offline, selects actions via a critic score, and tunes a Lagrange multiplier from the offline dataset to satisfy resource constraints.

Pith tools