Introduces a robust max-min benchmark for aggregating calibrated forecasts that is LP-tractable, dominates OIH, and is attained by online algorithms under forecast-only feedback.
A Survey on Contextual Multi-armed Bandits
4 Pith papers cite this work. Polarity classification is still indexing.
abstract
In this survey we cover a few stochastic and adversarial contextual bandit algorithms. We analyze each algorithm's assumption and regret bound.
representative citing papers
Identifiable latent bandits apply nonlinear ICA to observational data to recover representations sufficient for inferring optimal actions in new instances, shortening exploration time.
AutoPilot uses decentralized reinforcement learning to continuously adjust BFT protocol parameters online, achieving 49.8% lower end-to-end latency than static defaults in dynamic environments.
citing papers explorer
-
Robust Aggregation of Calibrated Forecasts
Introduces a robust max-min benchmark for aggregating calibrated forecasts that is LP-tractable, dominates OIH, and is attained by online algorithms under forecast-only feedback.
-
Identifiable Latent Bandits: Leveraging observational data for personalized decision-making
Identifiable latent bandits apply nonlinear ICA to observational data to recover representations sufficient for inferring optimal actions in new instances, shortening exploration time.
-
AutoPilot: Learning to Steer High Speed Robust BFT
AutoPilot uses decentralized reinforcement learning to continuously adjust BFT protocol parameters online, achieving 49.8% lower end-to-end latency than static defaults in dynamic environments.
- Latent Order Bandits