Pith. sign in

A Survey on Contextual Multi-armed Bandits

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it
abstract

In this survey we cover a few stochastic and adversarial contextual bandit algorithms. We analyze each algorithm's assumption and regret bound.

years

2026 3 2024 1

representative citing papers

Robust Aggregation of Calibrated Forecasts

econ.TH · 2026-06-30 · unverdicted · novelty 7.0

Introduces a robust max-min benchmark for aggregating calibrated forecasts that is LP-tractable, dominates OIH, and is attained by online algorithms under forecast-only feedback.

AutoPilot: Learning to Steer High Speed Robust BFT

cs.DC · 2026-06-08 · unverdicted · novelty 5.0

AutoPilot uses decentralized reinforcement learning to continuously adjust BFT protocol parameters online, achieving 49.8% lower end-to-end latency than static defaults in dynamic environments.

citing papers explorer

Showing 4 of 4 citing papers.

  • Robust Aggregation of Calibrated Forecasts econ.TH · 2026-06-30 · unverdicted · none · ref 18 · internal anchor

    Introduces a robust max-min benchmark for aggregating calibrated forecasts that is LP-tractable, dominates OIH, and is attained by online algorithms under forecast-only feedback.

  • Identifiable Latent Bandits: Leveraging observational data for personalized decision-making cs.LG · 2024-07-23 · unverdicted · none · ref 57 · internal anchor

    Identifiable latent bandits apply nonlinear ICA to observational data to recover representations sufficient for inferring optimal actions in new instances, shortening exploration time.

  • AutoPilot: Learning to Steer High Speed Robust BFT cs.DC · 2026-06-08 · unverdicted · none · ref 88 · internal anchor

    AutoPilot uses decentralized reinforcement learning to continuously adjust BFT protocol parameters online, achieving 49.8% lower end-to-end latency than static defaults in dynamic environments.

  • Latent Order Bandits cs.LG · 2026-05-08 · unreviewed · ref 8