Pith. sign in

REVIEW 1 cited by

Robust Batch Policy Learning in Markov Decision Processes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.04185 v4 pith:URMD3BMY submitted 2020-11-09 math.ST cs.LGstat.MLstat.TH

classification math.STcs.LGstat.MLstat.TH
keywords policydecisionrobustdatasetlearningmarkovadaptivityaverage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the offline data-driven sequential decision making problem in the framework of Markov decision process (MDP). In order to enhance the generalizability and adaptivity of the learned policy, we propose to evaluate each policy by a set of the average rewards with respect to distributions centered at the policy induced stationary distribution. Given a pre-collected dataset of multiple trajectories generated by some behavior policy, our goal is to learn a robust policy in a pre-specified policy class that can maximize the smallest value of this set. Leveraging the theory of semi-parametric statistics, we develop a statistically efficient policy learning method for estimating the de ned robust optimal policy. A rate-optimal regret bound up to a logarithmic factor is established in terms of total decision points in the dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online MDP with Transition Prototypes: A Robust Adaptive Approach

    cs.LG 2024-12 reject novelty 4.0 of 10

    An adaptive robust algorithm for online MDPs with finite transition prototypes achieves sublinear regret under a Lipschitz-like structural assumption.

Pith tools