REVIEW 2 cited by
On Covariate Shift of Latent Confounders in Imitation and Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We consider the problem of using expert data with unobserved confounders for imitation and reinforcement learning. We begin by defining the problem of learning from confounded expert data in a contextual MDP setup. We analyze the limitations of learning from such data with and without external reward, and propose an adjustment of standard imitation learning algorithms to fit this setup. We then discuss the problem of distribution shift between the expert data and the online environment when the data is only partially observable. We prove possibility and impossibility results for imitation learning under arbitrary distribution shift of the missing covariates. When additional external reward is provided, we propose a sampling procedure that addresses the unknown shift and prove convergence to an optimal solution. Finally, we validate our claims empirically on challenging assistive healthcare and recommender system simulation tasks.
Forward citations
Cited by 2 Pith papers
-
Confounded Causal Imitation Learning with Instrumental Variables
A causal imitation learning framework that identifies valid instrumental variables from observational data and uses them to learn policies robust to multi-timestep latent confounders.
-
Breaking Habits: On the Role of the Advantage Function in Learning Causal State Representations
Subtracting the state value from action values scales gradient updates to downweight frequent state-action pairs, helping agents learn causal state representations and generalize out-of-trajectory.
Discussion (0). Sign in to comment.