A new bandit algorithm, DP-TS-UCB, achieves a tunable privacy-regret trade-off, improving the privacy guarantee of Gaussian Thompson Sampling from O(sqrt(T)) to O(T^0.25) while preserving near-optimal regret.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Connecting Thompson Sampling and UCB: Towards More Efficient Trade-offs Between Privacy and Regret
A new bandit algorithm, DP-TS-UCB, achieves a tunable privacy-regret trade-off, improving the privacy guarantee of Gaussian Thompson Sampling from O(sqrt(T)) to O(T^0.25) while preserving near-optimal regret.