A new design and a sharper analysis of orthogonalized regression yield optimal regret, logarithmic regret under gaps, and the first PAC and best-arm identification guarantees for semiparametric bandits.
Since T ≳ L(2) T −1X ℓ=1 d2 log(1/δ)4ℓ ≳d2 log(1/δ)4L(2) T , we get the upper bound ofL(2) T as 2L(2) T ≲ s T d2 log(1/δ)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Experimental Design for Semiparametric Bandits
A new design and a sharper analysis of orthogonalized regression yield optimal regret, logarithmic regret under gaps, and the first PAC and best-arm identification guarantees for semiparametric bandits.