pith. machine review for the scientific record. sign in

arxiv: 1508.07091 · v4 · submitted 2015-08-28 · 💻 cs.LG

Recognition: unknown

Multi-armed Bandit Problem with Known Trend

Authors on Pith no claims yet
classification 💻 cs.LG
keywords banditmulti-armedknownmodelproblemtrendalgorithmdifferent
0
0 comments X
read the original abstract

We consider a variant of the multi-armed bandit model, which we call multi-armed bandit problem with known trend, where the gambler knows the shape of the reward function of each arm but not its distribution. This new problem is motivated by different online problems like active learning, music and interface recommendation applications, where when an arm is sampled by the model the received reward change according to a known trend. By adapting the standard multi-armed bandit algorithm UCB1 to take advantage of this setting, we propose the new algorithm named A-UCB that assumes a stochastic model. We provide upper bounds of the regret which compare favourably with the ones of UCB1. We also confirm that experimentally with different simulations

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.