The paper derives the semiparametric efficiency bound for off-policy evaluation in MDPs and introduces a doubly robust, cross-fitted estimator (DRL) that achieves it.
Non-Parametric Inference Adaptive to Intrinsic Dimension
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We consider non-parametric estimation and inference of conditional moment models in high dimensions. We show that even when the dimension $D$ of the conditioning variable is larger than the sample size $n$, estimation and inference is feasible as long as the distribution of the conditioning variable has small intrinsic dimension $d$, as measured by locally low doubling measures. Our estimation is based on a sub-sampled ensemble of the $k$-nearest neighbors ($k$-NN) $Z$-estimator. We show that if the intrinsic dimension of the covariate distribution is equal to $d$, then the finite sample estimation error of our estimator is of order $n^{-1/(d+2)}$ and our estimate is $n^{1/(d+2)}$-asymptotically normal, irrespective of $D$. The sub-sampling size required for achieving these results depends on the unknown intrinsic dimension $d$. We propose an adaptive data-driven approach for choosing this parameter and prove that it achieves the desired rates. We discuss extensions and applications to heterogeneous treatment effect estimation.
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes
The paper derives the semiparametric efficiency bound for off-policy evaluation in MDPs and introduces a doubly robust, cross-fitted estimator (DRL) that achieves it.