Pith. sign in

REVIEW 1 cited by

Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.08666 v1 pith:HMDN7SCH submitted 2022-09-18 cs.LG stat.ME

classification cs.LGstat.ME
keywords confoundedlearningofflinepolicyvariableschallengescoveragedata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the offline reinforcement learning (RL) in the face of unmeasured confounders. Due to the lack of online interaction with the environment, offline RL is facing the following two significant challenges: (i) the agent may be confounded by the unobserved state variables; (ii) the offline data collected a prior does not provide sufficient coverage for the environment. To tackle the above challenges, we study the policy learning in the confounded MDPs with the aid of instrumental variables. Specifically, we first establish value function (VF)-based and marginalized importance sampling (MIS)-based identification results for the expected total reward in the confounded MDPs. Then by leveraging pessimism and our identification results, we propose various policy learning methods with the finite-sample suboptimality guarantee of finding the optimal in-class policy under minimal data coverage and modeling assumptions. Lastly, our extensive theoretical investigations and one numerical study motivated by the kidney transplantation demonstrate the promising performance of the proposed methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

    cs.LG 2024-12 conditional novelty 7.0 of 10

    A two-way deconfounder algorithm that models unmeasured confounders as per-trajectory and per-timestep latent factors and uses a neural tensor network for off-policy evaluation.

Pith tools