A log-sum-exponential (LSE) estimator for off-policy evaluation and learning achieves an O(n^{-epsilon/(1+epsilon)}) regret rate under bounded (1+epsilon)-th moments of weighted reward, at the cost of a tunable pessimistic bias.
For OPL, a separate test set is used to evaluate the estimator’s performance
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
A log-sum-exponential (LSE) estimator for off-policy evaluation and learning achieves an O(n^{-epsilon/(1+epsilon)}) regret rate under bounded (1+epsilon)-th moments of weighted reward, at the cost of a tunable pessimistic bias.