AIPW estimators with nuisance estimates from no-regret online learning attain near-optimal finite-sample MSE for off-policy evaluation with adaptively collected data.
Counter factual reasoning and learning systems: The example of computational advertising
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Off-policy estimation with adaptively collected data: the power of online learning
AIPW estimators with nuisance estimates from no-regret online learning attain near-optimal finite-sample MSE for off-policy evaluation with adaptively collected data.