A log-sum-exponential (LSE) estimator for off-policy evaluation and learning achieves an O(n^{-epsilon/(1+epsilon)}) regret rate under bounded (1+epsilon)-th moments of weighted reward, at the cost of a tunable pessimistic bias.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
A log-sum-exponential (LSE) estimator for off-policy evaluation and learning achieves an O(n^{-epsilon/(1+epsilon)}) regret rate under bounded (1+epsilon)-th moments of weighted reward, at the cost of a tunable pessimistic bias.