Pith. sign in

REVIEW 1 cited by

Offline Reinforcement Learning from Datasets with Structured Non-Stationarity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14114 v2 pith:XD27MZ74 submitted 2024-05-23 cs.LG cs.AI

classification cs.LGcs.AI
keywords offlinemethodpolicydatasetlearningnon-stationarityoftenperforms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current Reinforcement Learning (RL) is often limited by the large amount of data needed to learn a successful policy. Offline RL aims to solve this issue by using transitions collected by a different behavior policy. We address a novel Offline RL problem setting in which, while collecting the dataset, the transition and reward functions gradually change between episodes but stay constant within each episode. We propose a method based on Contrastive Predictive Coding that identifies this non-stationarity in the offline dataset, accounts for it when training a policy, and predicts it during evaluation. We analyze our proposed method and show that it performs well in simple continuous control tasks and challenging, high-dimensional locomotion tasks. We show that our method often achieves the oracle performance and performs better than baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    ADG uses an ambient DDPM to flag corrupted RL transitions, trains a standard DDPM only on the clean subset, then refines the flagged transitions to produce a recovered dataset that improves offline RL policies.

Pith tools