REVIEW 9 cited by
A Survey on Contextual Multi-armed Bandits
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this survey we cover a few stochastic and adversarial contextual bandit algorithms. We analyze each algorithm's assumption and regret bound.
Forward citations
Cited by 9 Pith papers
-
Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles
An OCO algorithm with only O(√T) static regret, pluggable as a preconditioner selector, recovers the classical O(1/√T) stationarity rate on smooth stochastic nonconvex problems and the O(T^{-2/7}) rate on nonsmooth ones.
-
Robust Aggregation of Calibrated Forecasts
Introduces a robust max-min benchmark for aggregating calibrated forecasts that is LP-tractable, dominates OIH, and is attained by online algorithms under forecast-only feedback.
-
Contextual Online Decision Making with Infinite-Dimensional Functional Regression
A unified online decision-making framework that learns context-dependent CDFs via infinite-dimensional functional regression, with regret controlled by the eigenvalue decay of a design integral operator.
-
Identifiable Latent Bandits: Leveraging observational data for personalized decision-making
Identifiable latent bandits apply nonlinear ICA to observational data to recover representations sufficient for inferring optimal actions in new instances, shortening exploration time.
-
AutoPilot: Learning to Steer High Speed Robust BFT
AutoPilot uses decentralized reinforcement learning to continuously adjust BFT protocol parameters online, achieving 49.8% lower end-to-end latency than static defaults in dynamic environments.
-
Enhancing Federated Graph Learning via Adaptive Fusion of Structural and Node Characteristics
FedGCF fuses clustered structural models and selected node-feature models with a bandit-tuned ratio, claiming accuracy and communication improvements in federated graph classification, though its test-set-based tuning...
-
BanditWare: A Contextual Bandit-based Framework for Hardware Prediction
BanditWare uses a decaying epsilon-greedy contextual bandit with linear runtime models to recommend hardware for scientific workflows, learning online with far fewer samples than offline ML approaches.
-
In-Domain African Languages Translation Using LLMs and Multi-armed Bandits
Bandit-based model selection matches or slightly improves on the best single NMT system for in-domain English-to-African translation, but the claimed high-confidence statistical support is absent.
-
Selective Reviews of Bandit Problems in AI via a Statistical View
A statistical survey of multi-armed, contextual, and continuum-armed bandits that restates known minimax and regret results, adds an alternative UCB proof, and reports small simulation comparisons.
Discussion (0). Continue with ORCID to comment.