Pith. sign in

REVIEW 1 cited by

DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.08891 v2 pith:UXXJB2FT submitted 2020-10-18 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords offlineapproachcostsdac-mdpdeepderivedenvironmentsinvestigate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study an approach to offline reinforcement learning (RL) based on optimally solving finitely-represented MDPs derived from a static dataset of experience. This approach can be applied on top of any learned representation and has the potential to easily support multiple solution objectives as well as zero-shot adjustment to changing environments and goals. Our main contribution is to introduce the Deep Averagers with Costs MDP (DAC-MDP) and to investigate its solutions for offline RL. DAC-MDPs are a non-parametric model that can leverage deep representations and account for limited data by introducing costs for exploiting under-represented parts of the model. In theory, we show conditions that allow for lower-bounding the performance of DAC-MDP solutions. We also investigate the empirical behavior in a number of environments, including those with image-based observations. Overall, the experiments demonstrate that the framework can work in practice and scale to large complex offline RL problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generating Physically Realistic and Directable Human Motions from Multi-Modal Inputs

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A single reinforcement-learned controller uses masked motion demonstrations to catch up, combine, and complete humanoid motions from sparse multi-modal directives.

Pith tools