A pessimistic lower-confidence-bound algorithm achieves suboptimality bounds for offline combinatorial multi-armed bandits with probabilistically triggered arms, under coverage conditions requiring observation of each arm of the optimal action.
org/CorpusID:28202810
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Offline Learning for Combinatorial Multi-armed Bandits
A pessimistic lower-confidence-bound algorithm achieves suboptimality bounds for offline combinatorial multi-armed bandits with probabilistically triggered arms, under coverage conditions requiring observation of each arm of the optimal action.