RAVEN-UCB proposes a variance-adaptive UCB algorithm for non-stationary bandits, but the proof of its main regret bound is mathematically invalid.
Recommendation System-based Upper Confidence Bound for Online Advertising
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this paper, the method UCB-RS, which resorts to recommendation system (RS) for enhancing the upper-confidence bound algorithm UCB, is presented. The proposed method is used for dealing with non-stationary and large-state spaces multi-armed bandit problems. The proposed method has been targeted to the problem of the product recommendation in the online advertising. Through extensive testing with RecoGym, an OpenAI Gym-based reinforcement learning environment for the product recommendation in online advertising, the proposed method outperforms the widespread reinforcement learning schemes such as $\epsilon$-Greedy, Upper Confidence (UCB1) and Exponential Weights for Exploration and Exploitation (EXP3).
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
From Theory to Practice with RAVEN-UCB: Addressing Non-Stationarity in Multi-Armed Bandits through Variance Adaptation
RAVEN-UCB proposes a variance-adaptive UCB algorithm for non-stationary bandits, but the proof of its main regret bound is mathematically invalid.