pith. machine review for the scientific record. sign in

arxiv: 1102.3508 · v1 · submitted 2011-02-17 · 🧮 math.OC · cs.LG

Recognition: unknown

Online Learning of Rested and Restless Bandits

Authors on Pith no claims yet
classification 🧮 math.OC cs.LG
keywords armsrestlessuseraccessbanditslearningmultiarmedonline
0
0 comments X
read the original abstract

In this paper we study the online learning problem involving rested and restless multiarmed bandits with multiple plays. The system consists of a single player/user and a set of K finite-state discrete-time Markov chains (arms) with unknown state spaces and statistics. At each time step the player can play M arms. The objective of the user is to decide for each step which M of the K arms to play over a sequence of trials so as to maximize its long term reward. The restless multiarmed bandit is particularly relevant to the application of opportunistic spectrum access (OSA), where a (secondary) user has access to a set of K channels, each of time-varying condition as a result of random fading and/or certain primary users' activities.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.