The paper derives a bridge-function identification result that supports root-n-rate policy value estimation and policy-gradient policy learning for continuous actions under unmeasured confounding.
(2021), Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders, in International Conference on Artificial Intelligence and Statistics, PMLR, pp
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Reinforcement Learning with Continuous Actions Under Unmeasured Confounding
The paper derives a bridge-function identification result that supports root-n-rate policy value estimation and policy-gradient policy learning for continuous actions under unmeasured confounding.