A new continuum-armed bandit algorithm achieves near-optimal regret when feedback is a batched, biased pairwise comparison oracle, and improves known regret bounds for two operations management problems.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Continuum-armed Bandit Optimization with Batch Pairwise Comparison Oracles
A new continuum-armed bandit algorithm achieves near-optimal regret when feedback is a batched, biased pairwise comparison oracle, and improves known regret bounds for two operations management problems.