AIPW estimators with nuisance estimates from no-regret online learning attain near-optimal finite-sample MSE for off-policy evaluation with adaptively collected data.
Effective evaluation using logged bandit feedback from multiple loggers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Off-policy estimation with adaptively collected data: the power of online learning
AIPW estimators with nuisance estimates from no-regret online learning attain near-optimal finite-sample MSE for off-policy evaluation with adaptively collected data.