A new two-part regret metric and an explore-then-set-cover algorithm for stochastic multi-objective bandits are proposed, with sublinear regret bounds for Pareto-optimal and convex-supported arms.
In: ESANN (2015)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Stochastic Multi-Objective Multi-Armed Bandits: Regret Definition and Algorithm
A new two-part regret metric and an explore-then-set-cover algorithm for stochastic multi-objective bandits are proposed, with sublinear regret bounds for Pareto-optimal and convex-supported arms.