Pith. sign in

REVIEW 1 major objections 5 minor 151 references

Selective Reviews of Bandit Problems in AI via a Statistical View

T0 review · 1 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper argues that stochastic bandit theory is best read as non-asymptotic statistical inference, where concentration inequalities yield the confidence intervals, regret decompositions, and minimax benchmarks that organize multi-armed…

desk verdict A useful but flawed survey: the alternative UCB proof in Appendix B has a load-bearing gap, so the paper's one original contribution is not established as written. read the letter →

arxiv 2412.02251 v3 pith:MTETA5FT submitted 2024-12-03 stat.ML cs.AIcs.LGecon.EMmath.PR

classification stat.MLcs.AIcs.LGecon.EMmath.PR MSC 68W2768T0562E17
keywords stochasticmulti-armedbanditscontinuum-armedsub-Gaussianconcentrationinequalitiesregretboundsminimaxratescontextualexploration-exploitationtrade-offfunctionaldataanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The authors are trying to establish that the theory of stochastic bandits is essentially non-asymptotic statistics: the same toolbox of concentration inequalities, confidence intervals, regret decompositions, and minimax lower bounds explains the guarantees of UCB, MOSS, posterior sampling, and continuum-armed bandits. They make the point by re-deriving the UCB regret bound through an excess-risk-style decomposition rather than the standard proof, and by presenting the sequential decision problem as one of estimating unknown reward distributions under adaptive sampling. A sympathetic reader would take away that bandit guarantees are consequences of tail bounds applied to adaptively collected samples, and that minimax rates separate what is algorithm-specific from what is inherent in the problem. The review itself flags the main caveat: most bounds assume the reward noise is sub-Gaussian with a known variance proxy, usually set to one.

What carries the argument

The object that carries the argument is the regret-decomposition identity for optimistic algorithms: for UCB, $\mathrm{Reg}_T(\pi, v) \le \mathbb{E}\sum_{t=1}^T [\mu^* - \mathrm{UCB}_*(t-1,\delta) + \mathrm{UCB}_{A_t}(t-1,\delta) - \mu_{A_t}]$, where the selected arm maximizes the index so the optimal arm's index term contributes a negative concentration error. This reduces regret to the sum of concentration errors of the chosen and optimal arms; feeding in a sub-Gaussian tail bound plus a union bound yields the $O(K + \sqrt{KT\log T})$ rate. The companion structural objects are the minimax lower bound constructed from Gaussian instances, which fixes the $\sqrt{KT}$ benchmark, and the information gain $\gamma_T$, which quantifies the effective dimensionality a GP-UCB learner must explore over a continuous action space.

What would settle it

Fix a K-armed bandit with near-equal means (small gaps $\Delta_k$) and rewards drawn from a mixture such as $0.4\,N(0,1) + 0.6\,N(0,9)$, so the true sub-Gaussian proxy exceeds 1; run Algorithm 3 with proxy 1 for $T=10^4$ and compare the observed cumulative regret to $3\sum_k \Delta_k + 8\sqrt{TK\log T}$, since a violation at the claimed confidence level would show that the known-proxy assumption is carrying the central theorem.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that the many bandit settings share a single statistical skeleton: every algorithm's exploration is a confidence interval, every regret bound is a concentration calculation, and optimality is judged by minimax rate. The review demonstrates this by reproving the UCB regret theorem—for subG(1) rewards with $\delta = 1/T^2$, $\mathrm{Reg}_T(\pi, v) \le 3\sum_k \Delta_k + 8\sqrt{TK\log T}$, with an alternative proof giving $\mathrm{Reg}_T \le 4|\mu^*|K + 12\sqrt{\log T}\sqrt{KT}$—through a regret decomposition that mirrors excess risk in statistical learning. It then applies the same reading to linear contextual bandits, where LinUCB and linear Thompson sampling attain $\widetilde{O}(d\sqrt{T})$ regret independent of the number of arms, and to continuum-armed bandits, where the GP-UCB regret is controlled by the maximum information gain $\gamma_T$. The review also connects continuum-armed bandits to functional data analysis, treating the unknown reward as a function to be estimated over a continuous domain.

Load-bearing premise

The main theorems assume the reward noise has a known sub-Gaussian variance proxy, normally set to 1, together with a known horizon $T$ and confidence level $\delta$; if that proxy is unknown, misspecified, or the noise is heavy-tailed, the concentration bound in display (19) and Theorem 5 do not hold.

Editorial extensions

If this is right

  • UCB with subG(1) rewards has problem-independent regret $O(\sqrt{TK\log T})$, and the gap to the $\sqrt{KT}$ minimax lower bound is closed by MOSS and the minimax-optimal Thompson-sampling variant MOTS.
  • Contextual linear bandits, both LinUCB and linear Thompson sampling, achieve regret $\widetilde{O}(d\sqrt{T})$ independent of the number of arms, so shared features pay off whenever the feature dimension is manageable.
  • For continuum-armed bandits with a Gaussian-process prior, the cumulative regret is bounded by $\sqrt{C_1 T \beta_T \gamma_T}$, where $\gamma_T$ is the maximum information gain, which gives sublinear regret for common kernels.
  • The causal treatment-effect problem can be modeled as a two-armed bandit, yielding non-asymptotic confidence intervals for the treatment effect when the potential outcomes are sub-Gaussian.
  • The doubling trick converts the horizon-dependent algorithms described in the review into anytime algorithms without changing their regret rates.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Read as a template, the regret-decomposition lemma likely extends to optimistic algorithms beyond UCB: LinUCB and GP-UCB can be viewed as the same concentration-error-plus-estimation-error split with different confidence radii, though the review does not state this unification explicitly.
  • The advertised connection between continuum-armed bandits and functional data analysis suggests transfers the authors leave implicit, such as using functional principal components or basis smoothing inside a continuum bandit when the reward function has low-rank structure, which could reduce the effective action dimension.
  • The simulations in the unknown-variance-proxy section hint at a practical rule the paper does not state: when the variance proxy is unknown, estimating the sub-Gaussian norm online and plugging it into the confidence radius can outperform both asymptotic variance estimators and wrongly applied bounded-reward concentration bounds; this could be tested systematically across reward families.
  • If the statistical reading is correct, then designing a new bandit algorithm reduces to finding the tightest valid concentration inequality for the reward family, meaning any new tail bound would directly yield a new UCB-style algorithm with a corresponding regret bound.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The manuscript is a survey of stochastic bandit problems (finite-armed, contextual, and continuum-armed) viewed through the lens of non-asymptotic statistics. It introduces concentration inequalities for sub-Gaussian and sub-exponential rewards, reviews ETC, UCB, MOSS, Thompson sampling, MOTS, LinUCB, LinTS, GP-UCB, and GP-TS, and discusses connections to functional data analysis and recent topics such as unknown variance proxies. A central feature is an 'alternative proof' of the UCB regret bound based on an excess-risk decomposition from statistical learning theory, presented in Appendix B.

Significance. If correct, the survey would provide a compact bridge between bandit theory and statistical learning theory, with a coherent narrative that many bandit guarantees reduce to concentration inequalities and risk decompositions. The paper is strongest as a compilation: it gathers standard algorithms and bounds, attributes them clearly, and includes simulations that illustrate qualitative differences among algorithms. It also honestly acknowledges the known-variance-proxy limitation in Section 6.4. However, the advertised alternative UCB proof in Appendix B contains a load-bearing gap, and the sub-exponential example in Section 2.2 contains a false inequality. These issues prevent the present version from fully delivering its methodological message, although the standard results quoted are mostly correct.

major comments (1)
  1. [Appendix B, Eqs. (A3)–(A4)] The alternative proof of Theorem 5 drops the term μ* − UCB*(t−1,δ) from Lemma 7 unconditionally. This term is non-positive only on the good event G; on G^c it can be positive and of order sqrt(2 log(1/δ)) even when μ* = 0, and it can persist over many rounds. The displayed bound E{1_{G^c} Σ (UCB_{A_t} − μ_{A_t})} ≤ 2|μ*| T P(G^c) is also invalid, since UCB_{A_t} − μ_{A_t} is not bounded by 2|μ*| on G^c. The martingale-difference term in the preceding display is not handled correctly either: the expectation of an adaptively weighted sum is not bounded by max_t E[Y_{A_t}(τ) − μ_{A_t}]. As a result, the claimed regret bound 4|μ*|K + 12 sqrt(log T) sqrt(KT) does not follow from the written argument. Since this proof is advertised in Section 3.2 as the main pedagogical support for the UCB analysis, it needs to be repaired or replaced; the theorem itself is standard and Appendix A provides a valid proof.
minor comments (5)
  1. [Section 2.2, Example 9] Example 9 asserts e^{-sμ}(1−sμ)^{-1} ≤ e^{s^2 μ^2/2} for |s| ≤ (2μ)^{-1}, but this is false: at s = 0.4/μ the left-hand side is approximately 1.117 and the right-hand side approximately 1.083. The centered exponential is indeed sub-exponential, but the displayed inequality and the implied parameterization (μ, 2μ) should be corrected.
  2. [Appendix B] In the display after Eq. (A3), the quantity S_{k*} should be S_{A_t}: the confidence radius for the selected arm A_t depends on the number of pulls of that arm, not on the number of pulls of the optimal arm.
  3. [Sections 3.2 and 6.4] Theorems 5 and 7 assume a known variance proxy (usually 1), while Section 6.4 discusses the unknown-proxy problem only later; adding a forward reference in Section 3.2 would make the scope of the main bounds clearer to readers.
  4. [Section 2.1, Example 4] The phrase '∑_{i=1}^n X_{ij} i.i.d.∼ N(0, nσ^2)' should read '∑_{i=1}^n X_{ij} ∼ N(0, nσ^2)', since the sum is a single random variable rather than an i.i.d. collection.
  5. [Various] There are minor typos, e.g., 'genalization' in Section 6.1 should be 'generalization', and the caption of Figure 5 has 'Culumative' instead of 'Cumulative'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the UCB derivation reduces to standard sub-Gaussian concentration and regret decomposition, not to its own assumptions.

full rationale

The paper's central derivations are self-contained reductions to textbook material. Theorem 5 is explicitly attributed to Lattimore and Szepesvári's Theorem 7.2, and Appendix A supplies a standard proof from the regret decomposition lemma and sub-Gaussian concentration inequality (14). Lemma 7 decomposes regret into UCB terms, and Appendix B then applies concentration plus integral estimates; this is a genuine derivation chain, not an identity or a fitted-parameter prediction. Self-citations such as [24], [34], and [47] are used as references for background inequalities and for a simulation variant, not as the load-bearing justification for the main regret bounds. The Appendix B proof has a potential gap in bounding the bad-event contribution, but a proof gap is a correctness issue, not circularity. No fitted constants are renamed as predictions, no uniqueness theorem is imported from the authors, and no target result is assumed in defining its own inputs.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The review introduces no new entities. Its mathematical content rests on standard concentration inequalities, quoted theorems, and domain assumptions about sub-Gaussian rewards, known kernels, and known variance proxies. The only ad hoc mathematical assertion is the subexponential MGF bound in Example 9, which is incorrect as stated.

free parameters (3)
  • GP lengthscale l = 0.2575
    Estimated by MLE from 500 synthetic samples in Section 5.3 and used to simulate both GP-UCB and GP-TS; the comparison outcome may change with the kernel lengthscale.
  • GP-UCB exploration parameter beta = 2.0
    Chosen by hand in Section 5.3 with no sensitivity analysis reported.
  • MOTS truncation parameters rho and alpha = rho = 0.8, alpha = 1.5
    The simulation in Section 3.5 uses alpha=1.5 although Theorem 9 requires alpha>=4; the text says this is for a practical confidence bound.
assumptions (6)
  • domain assumption Rewards are sub-Gaussian with known variance proxy, often subG(1), for the MAB and linear bandit bounds.
    Used throughout Sections 3 and 4; see equation (2), Theorem 4, Theorem 5, and Theorem 8.
  • standard math The minimax regret lower bound for K-armed Gaussian bandits is Reg*_T(E) >= (1/27) sqrt((K-1)T).
    Stated as Theorem 6 in Section 3.3.1 and attributed to Theorem 15.1 in [21]; no proof is given.
  • standard math The MOSS, MOTS, LinUCB, LinTS, GP-UCB, and GP-TS regret bounds in Theorems 7, 9, 10, 11, 12, and 13 are correct as quoted.
    These results are imported from [70], [73], [75], [77], and [81] in Sections 3.3.2, 3.5, 4, and 5; the paper provides no independent verification.
  • domain assumption For SCAB, the reward function f is a sample from a Gaussian process with known kernel, or lies in the RKHS of that kernel.
    Assumed in Sections 5.1 and 5.4; the regret bounds depend on the known kernel k and the mutual information gamma_T.
  • domain assumption For the Neyman-Rubin example, potential outcomes are subG(sigma^2) and treatment assignment is independent of potential outcomes.
    Example 8 in Section 2.2; this is the standard causal assumption needed for the confidence interval.
  • ad hoc to paper The centered exponential distribution is claimed to satisfy the subexponential MGF bound with parameters (mu, 2mu) in Example 9.
    The derivation in Section 2.2 uses the false inequality e^{-2t}/(1-2t) <= e^{2t^2}; this is an ad hoc claim of this paper rather than a standard result.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Selective Reviews of Bandit Problems in AI via a Statistical View." pith.science (2026). https://pith.science/paper/MTETA5FT

@misc{pith2026241202251,
  author       = {Pith},
  title        = {Pith review of: Selective Reviews of Bandit Problems in AI via a Statistical View},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MTETA5FT}},
  note         = {Machine review of arXiv:2412.02251}
}
read the original abstract

Reinforcement Learning (RL) is a widely researched area in artificial intelligence that focuses on teaching agents decision-making through interactions with their environment. A key subset includes stochastic multi-armed bandit (MAB) and continuum-armed bandit (SCAB) problems, which model sequential decision-making under uncertainty. This review outlines the foundational models and assumptions of bandit problems, explores non-asymptotic theoretical tools like concentration inequalities and minimax regret bounds, and compares frequentist and Bayesian algorithms for managing exploration-exploitation trade-offs. Additionally, we explore K-armed contextual bandits and SCAB, focusing on their methodologies and regret analyses. We also examine the connections between SCAB problems and functional data analysis. Finally, we highlight recent advances and ongoing challenges in the field.

Figures

Figures reproduced from arXiv: 2412.02251 by the authors.

Figure 1
Figure 1. A player plays at a three-armed bandit machine in a casino. In the field of statistics, the time-uniform confidence sequence problem [26] is often framed as an MAB problem, first introduced by [27] in the context of sequential experi￾mental design. The topic has been extensively studied in the machine learning literature, with significant contributions documented in major journals. For a comprehensive review, see Se… view at source ↗
Figure 2
Figure 2. Cumulative regret comparisons of ETC, UCB, MOSS, TS, and MOTS algorithms [PITH_FULL_IMAGE:figures/full_fig_p025_2.png] view at source ↗
Figure 3
Figure 3. Cumulative regret comparison of LinUCB and LinTS algorithms The advantages and disadvantages of the LinUCB and LinTS algorithms are summa￾rized in [PITH_FULL_IMAGE:figures/full_fig_p030_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Average regret comparisons of GP-UCB and GP-TS algorithms. 5.4. SCAB in Reproducing Kernel Hilbert Space In practical scenarios, the objective function exhibits considerable complexity, necessi￾tating strategies that optimize cumulative rewards or minimize regret over …
Figure 5
Figure 5. Figure 5: Cumulative regret comparisons under mixed Gaussian rewards with an unknown sub￾Gaussian variance proxy. 7. Concluding Remarks and Future Directions Bandit algorithms have gained significant attention and widespread applications across various fields. Accurate uncertain…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

151 extracted references · 68 canonical work pages

  1. [1]

    Reinforcement Learning: An Introduction, 2012.; MIT Press: Cambridge, USA

    Richard, S.; Sutton, S.; Andrew, G. Reinforcement Learning: An Introduction, 2012.; MIT Press: Cambridge, USA

  2. [2]

    Statistical Reinforcement Learning: Modern Machine Learning Approaches; Chapman and Hall/CRC: New York, NY, USA, 2015; pp

    Sugiyama, M. Statistical Reinforcement Learning: Modern Machine Learning Approaches; Chapman and Hall/CRC: New York, NY, USA, 2015; pp. 1-206

  3. [3]

    Autonomous Driving: Technical, Legal and Social Aspects; Springer: Berlin, Germany, 2016; pp

    Maurer, M.; Gerdes, J.C.; Lenz, B.; Winner, H. Autonomous Driving: Technical, Legal and Social Aspects; Springer: Berlin, Germany, 2016; pp. 1-706

  4. [4]

    Large-scale bandit approaches for recommender systems

    Zhou, Q.; Zhang, X.; Xu, J.; Liang, B. Large-scale bandit approaches for recommender systems. In Proceedings of the International Conference on Neural Information Processing (ICONIP 2017), Liu, D.; Xie, S.; Li, Y.; Zhao, D.; El-Alfy, E.S., Eds.; Springer: Cham, 2017; pp. 811–821

  5. [5]

    Unmanned aerial vehicles in agriculture: A survey

    Del Cerro, J.; Cruz Ulloa, C.; Barrientos, A.; de León Rivas, J. Unmanned aerial vehicles in agriculture: A survey. Agronomy 2021, 11, 203

  6. [6]

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P .; Bi, X.; et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv 2025, arXiv:2501.12948

  7. [7]

    Training language models to follow instructions with human feedback

    Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P .; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 2022, 35, 27730–27744

  8. [8]

    Portfolio choices with orthogonal bandit learning

    Shen, W.; Wang, J.; Jiang, Y.G.; Zha, H. Portfolio choices with orthogonal bandit learning. In Proceedings of the 24th International Conference on Artificial Intelligence (IJCAI 2015), Buenos Aires, Argentina, 2015; AAAI Press: 974–980

Show all 151 references
  1. [9]

    Causal Inference: A Statistical Learning Approach

    Wager, S. Causal Inference: A Statistical Learning Approach. 2024. Available online: https://web.stanford.edu/~swager/causal_ inf_book.pdf (accessed on 14th Feb., 2025)

  2. [10]

    Enhancing digital twins through reinforcement learning

    Cronrath, C.; Aderiani, A.R.; Lennartson, B. Enhancing digital twins through reinforcement learning. In Proceedings of the 2019 IEEE 15th International Conference on Automation Science and Engineering (CASE), Vancouver, BC, Canada, 2019; IEEE: 293–298. DOI: https://doi.org/10....

  3. [11]

    A digital twin to train deep reinforcement learning agent for smart manufacturing plants: Environment, interfaces and intelligence

    Xia, K.; Sacco, C.; Kirkpatrick, M.; Saidy, C.; Nguyen, L.; Kircaliali, A.; Harik, R. A digital twin to train deep reinforcement learning agent for smart manufacturing plants: Environment, interfaces and intelligence. J. Manuf. Syst. 2021, 58, 210–230

  4. [12]

    Contextual bandits for adapting treatment in a mouse model of de novo carcinogenesis

    Durand, A.; Achilleos, C.; Iacovides, D.; Strati, K.; Mitsis, G.D.; Pineau, J. Contextual bandits for adapting treatment in a mouse model of de novo carcinogenesis. In Proceedings of the 3rd Machine Learning for Healthcare Conference, 17–18 August 2018, PMLR: 67–82. URL: https...

  5. [13]

    Bandit algorithms for precision medicine

    Lu, Y.; Xu, Z.; Tewari, A. Bandit algorithms for precision medicine. In Handbooks of Modern Statistical Methods; Bühlmann, P ., Drineas, P ., Kane, M., van der Laan, M., Eds.; 2024; Chapter 13

  6. [14]

    A survey on practical applications of multi-armed and contextual bandits

    Bouneffouf, D.; Rish, I. A survey on practical applications of multi-armed and contextual bandits. arXiv 2019, arXiv:1904.10040

  7. [15]

    Introduction to Multi-Armed Bandits

    Slivkins, A. Introduction to Multi-Armed Bandits. Found. Trends® Mach. Learn. 2019, 12, 1–286

  8. [16]

    Survey of multiarmed bandit algorithms applied to recommendation systems

    Elena, G.; Milos, K.; Eugene, I. Survey of multiarmed bandit algorithms applied to recommendation systems. Int. J. Open Inf. Technol. 2021, 9, 12–27

  9. [17]

    A map of bandits for e-commerce

    Liu, Y.; Li, L. A map of bandits for e-commerce. In Proceedings of the KDD 2021 Workshop on Multi-Armed Bandits and Reinforcement Learning (MARBLE), 2021. Version February 20, 2025 submitted to arxiv 49 of 52

  10. [18]

    Bandit algorithms: A comprehensive review and their dynamic selection from a portfolio for multicriteria top-k recommendation

    Letard, A.; Gutowski, N.; Camp, O.; Amghar, T. Bandit algorithms: A comprehensive review and their dynamic selection from a portfolio for multicriteria top-k recommendation. Expert Syst. Appl. 2024, 26, 123151

  11. [19]

    A survey of online experiment design with the stochastic multi-armed bandit

    Burtini, G.; Loeppky, J.; Lawrence, R. A survey of online experiment design with the stochastic multi-armed bandit. arXiv 2015, arXiv:1510.00757

  12. [20]

    A survey on Multi-Armed, Contextual and Causal bandit algorithms for online learning Preprint

    Shah, N. A survey on Multi-Armed, Contextual and Causal bandit algorithms for online learning Preprint. 2020. Available online: https://nihaarshah.github.io/project_pdfs/ML_Theory_Project_Report.pdf (accessed on 14th Feb., 2025)

  13. [21]

    Bandit Algorithms; Cambridge University Press: Cambridge, UK, 2020

    Lattimore, T.; Szepesvári, C. Bandit Algorithms; Cambridge University Press: Cambridge, UK, 2020

  14. [23]

    Zero-Inflated Bandits

    Wei, H.; Wan, R.; Shi, L.; Song, R. Zero-Inflated Bandits. arXiv 2023, arXiv:2312.15595

  15. [24]

    Sharper sub-weibull concentrations

    Zhang, H.; Wei, H. Sharper sub-weibull concentrations. Mathematics 2022, 10, 2252

  16. [25]

    Non-asymptotic guarantees for robust statistical learning under infinite variance assumption

    Xu, L.; Yao, F.; Yao, Q.; Zhang, H. Non-asymptotic guarantees for robust statistical learning under infinite variance assumption. J. Mach. Learn. Res. 2023, 24, 1–46

  17. [26]

    Time-uniform, nonparametric, nonasymptotic confidence sequences

    Howard, S.R.; Ramdas, A.; McAuliffe, J.; Sekhon, J. Time-uniform, nonparametric, nonasymptotic confidence sequences. Ann. Stat. 2021, 49, 1055–1080

  18. [27]

    Some aspects of the sequential design of experiments

    Robbins, H. Some aspects of the sequential design of experiments. Bull. Am. Math. Soc. 1952, 58, 527–535

  19. [28]

    Regret Analysis of Stochastic and Nonstochastic Multi-Armed Bandit Problems

    Bubeck, S.; Cesa-Bianchi, N. Regret Analysis of Stochastic and Nonstochastic Multi-Armed Bandit Problems. In Foundations and Trends® in Machine Learning; Now Publishers, Inc.: 2012; Volume 5, pp. 1–122

  20. [29]

    Asymptotically efficient adaptive allocation rules

    Lai, T.L.; Robbins, H. Asymptotically efficient adaptive allocation rules. Adv. Appl. Math. 1985, 6, 4–22

  21. [30]

    Nearly dimension-independent sparse linear bandit over small action spaces via best subset selection

    Chen, Y.; Wang, Y.; Fang, E.X.; Wang, Z.; Li, R. Nearly dimension-independent sparse linear bandit over small action spaces via best subset selection. J. Am. Stat. Assoc. 2024, 119, 246–258

  22. [31]

    High-dimensional sparse linear bandits

    Hao, B.; Lattimore, T.; Wang, M. High-dimensional sparse linear bandits. Adv. Neural Inf. Process. Syst. 2020, 33, 10753–10763

  23. [32]

    Efficient sparse linear bandits under high dimensional data

    Wang, X.; Wei, M.M.; Yao, T. Efficient sparse linear bandits under high dimensional data. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Long Beach, CA, USA, 2023; pp. 2431–2443

  24. [33]

    Provably Efficient High-Dimensional Bandit Learning with Batched Feedbacks

    Fan, J.; Wang, Z.; Yang, Z.; Ye, C. Provably Efficient High-Dimensional Bandit Learning with Batched Feedbacks. arXiv 2023, arXiv:2311.13180

  25. [34]

    Tight non-asymptotic inference via sub-Gaussian intrinsic moment norm

    Zhang, H.; Wei, H.; Cheng, G. Tight non-asymptotic inference via sub-Gaussian intrinsic moment norm. arXiv 2023, arXiv:2303.07287

  26. [35]

    A perspective on off-policy evaluation in reinforcement learning

    Li, L. A perspective on off-policy evaluation in reinforcement learning. Front. Comput. Sci. 2019, 13, 911–912

  27. [36]

    The continuum-armed bandit problem

    Agrawal, R. The continuum-armed bandit problem. SIAM J. Control Optim. 1995, 33, 1926–1951

  28. [37]

    Optimal Design of Experiments; SIAM, Philadelphia, USA, 2006

    Pukelsheim, F. Optimal Design of Experiments; SIAM, Philadelphia, USA, 2006

  29. [38]

    Bayesian Optimization; Cambridge University Press: Cambridge, UK, 2023

    Garnett, R. Bayesian Optimization; Cambridge University Press: Cambridge, UK, 2023

  30. [39]

    Gaussian processes for regression

    Williams, C.; Rasmussen, C. Gaussian processes for regression. Adv. Neural Inf. Process. Syst. 1995, 8, 514–520

  31. [40]

    On the mathematical foundations of theoretical statistics

    Fisher, R.A. On the mathematical foundations of theoretical statistics. Philos. Trans. R. Soc. Lond. Ser. A Contain. Pap. Math. Phys. Character 1922, 222, 309–368

  32. [41]

    Propriétés locales des fonctions à séries de Fourier aléatoires

    Kahane, J.P . Propriétés locales des fonctions à séries de Fourier aléatoires. Stud. Math. 1960, 19, 1–25

  33. [42]

    Sur un nouveau théoreme-limite de la théorie des probabilités

    Cramér, H. Sur un nouveau théoreme-limite de la théorie des probabilités. Actual. Sci. Ind. 1938, 736, 5–23

  34. [43]

    Probability: A Graduate Course, 2nd ed.; Springer: New York, USA, 2013

    Gut, A. Probability: A Graduate Course, 2nd ed.; Springer: New York, USA, 2013

  35. [44]

    Introduction to High-Dimensional Statistics, 2nd ed.; Chapman and Hall/CRC: Boca Raton, USA, 2021

    Giraud, C. Introduction to High-Dimensional Statistics, 2nd ed.; Chapman and Hall/CRC: Boca Raton, USA, 2021

  36. [45]

    Probability Inequalities for Sums of Bounded Random Variables

    Hoeffding, W. Probability Inequalities for Sums of Bounded Random Variables. J. Am. Stat. Assoc. 1963, 58, 13–30

  37. [46]

    Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator

    Dvoretzky, A.; Kiefer, J.; Wolfowitz, J. Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator. Ann. Math. Stat. 1956, 27, 642–669

  38. [47]

    Concentration inequalities for statistical inference

    Zhang, H.; Chen, S.X. Concentration inequalities for statistical inference. Commun. Math. Res. 2021, 37, 1–85

  39. [48]

    Statistics and Information Theory

    Duchi, J. Statistics and Information Theory. 2024. Available online: https://web.stanford.edu/class/stats311/lecture-notes.pdf (accessed on 14th Feb., 2025)

  40. [49]

    High-Dimensional Statistics: A Non-Asymptotic Viewpoint; Cambridge University Press: Cambridge, UK, 2019

    Wainwright, M.J. High-Dimensional Statistics: A Non-Asymptotic Viewpoint; Cambridge University Press: Cambridge, UK, 2019

  41. [50]

    Petrov, V .V .Limit Theorems of Probability Theory; Sequences of Independent Random Variable ; Oxford University Press: Oxford, UK, 1995

  42. [51]

    Provably optimal algorithms for generalized linear contextual bandits

    Li, L.; Lu, Y.; Zhou, D. Provably optimal algorithms for generalized linear contextual bandits. InProceedings of the 34th International Conference on Machine Learning, 06–11 Aug 2017; Volume 70, pp. 2071–2080

  43. [52]

    Towards practical mean bounds for small samples

    Phan, M.; Thomas, P .; Learned-Miller, E. Towards practical mean bounds for small samples. InProceedings of the 38th International Conference on Machine Learning, 18–24 Jul 2021; Volume 139, pp. 8567–8576

  44. [53]

    Estimating means of bounded random variables by betting

    Waudby-Smith, I.; Ramdas, A. Estimating means of bounded random variables by betting. J. R. Stat. Soc. Ser. B Stat. Methodol. 2024, 86, 1–27

  45. [54]

    Policy optimization using semiparametric models for dynamic pricing

    Fan, J.; Guo, Y.; Yu, M. Policy optimization using semiparametric models for dynamic pricing. J. Am. Stat. Assoc. 2024, 119, 552–564

  46. [55]

    Sums of Independent Random Variables; Springer: Berlin, German, 1975

    Petrov, V .V . Sums of Independent Random Variables; Springer: Berlin, German, 1975

  47. [56]

    Computer Age Statistical Inference, Student Edition: Algorithms, Evidence, and Data Science; Cambridge University Press: Cambridge, UK, 2021

    Efron, B.; Hastie, T. Computer Age Statistical Inference, Student Edition: Algorithms, Evidence, and Data Science; Cambridge University Press: Cambridge, UK, 2021. Version February 20, 2025 submitted to arxiv 50 of 52

  48. [57]

    A Probabilistic Theory of Pattern Recognition; Springer: New York, USA, 1997

    Devroye, L.; Györfi, L.; Lugosi, G. A Probabilistic Theory of Pattern Recognition; Springer: New York, USA, 1997

  49. [58]

    Large deviation methods for approximate probabilistic inference

    Kearns, M.; Saul, L. Large deviation methods for approximate probabilistic inference. In Proceedings of the Fourteenth Conference on Uncertainty in Artificial Intelligence, 1998; pp. 311–319. Morgan Kaufmann Publishers Inc.: San Francisco, CA, USA, 1998

  50. [59]

    Some nonasymptotic results on resampling in high dimension, I: confidence regions

    Arlot, S.; Blanchard, G.; Roquain, E. Some nonasymptotic results on resampling in high dimension, I: confidence regions. Ann. Stat. 2010, 38, 51–82

  51. [60]

    Inference in a class of optimization problems: Confidence regions and finite sample bounds on errors in coverage probabilities

    Horowitz, J.L.; Lee, S. Inference in a class of optimization problems: Confidence regions and finite sample bounds on errors in coverage probabilities. J. Bus. Econ. Stat. 2023, 41, 927–938

  52. [61]

    Mathematical Statistics: A Non-Asymptotic Approach

    Rakhlin, A. Mathematical Statistics: A Non-Asymptotic Approach. Lecture Notes. 2020. Available online: https://web.stanford. edu/class/stats311/lecture-notes.pdf (accessed on 14th Feb., 2025)

  53. [62]

    Finite-time analysis of vector autoregressive models under linear restrictions

    Zheng, Y.; Cheng, G. Finite-time analysis of vector autoregressive models under linear restrictions. Biometrika 2021, 108, 469–489

  54. [63]

    Fast nonasymptotic testing and support recovery for large sparse Toeplitz covariance matrices

    Bettache, N.; Butucea, C.; Sorba, M. Fast nonasymptotic testing and support recovery for large sparse Toeplitz covariance matrices. J. Multivar. Anal. 2021, 190, 104883

  55. [64]

    Valid and Approximately Valid Confidence Intervals for Current Status Data.J

    Kim, S.; Fay, M.P .; Proschan, M.A. Valid and Approximately Valid Confidence Intervals for Current Status Data.J. R. Stat. Soc. Ser. B (Stat. Methodol.) 2021, 83, 438–452

  56. [65]

    Finite sample change point inference and identification for high-dimensional mean vectors

    Yu, M.; Chen, X. Finite sample change point inference and identification for high-dimensional mean vectors. J. R. Stat. Soc. Ser. B (Stat. Methodol.) 2021, 83, 247–270

  57. [66]

    What doubling tricks can and can’t do for multi-armed bandits

    Besson, L.; Kaufmann, E. What doubling tricks can and can’t do for multi-armed bandits. arXiv 2018, arXiv:1803.06971

  58. [67]

    Adaptive treatment allocation and the multi-armed bandit problem

    Lai, T.L. Adaptive treatment allocation and the multi-armed bandit problem. Ann. Stat. 1987, 15, 1091–1114

  59. [68]

    On Lai’s Upper Confidence Bound in Multi-Armed Bandits

    Ren, H.; Zhang, C.H. On Lai’s Upper Confidence Bound in Multi-Armed Bandits. arXiv 2024, arXiv:2410.02279

  60. [69]

    Introduction to Nonparametric Estimation; Springer: New York, USA, 2009

    Alexandre, T. Introduction to Nonparametric Estimation; Springer: New York, USA, 2009

  61. [70]

    Minimax policies for adversarial and stochastic bandits

    Audibert, J.Y.; Bubeck, S. Minimax policies for adversarial and stochastic bandits. In Proceedings of the COLT, 2009; pp. 217–226

  62. [71]

    A tutorial on thompson sampling

    Russo, D.J.; Van Roy, B.; Kazerouni, A.; Osband, I.; Wen, Z. A tutorial on thompson sampling. Found. Trends® Mach. Learn. 2018, 11, 1–96

  63. [72]

    On the likelihood that one unknown probability exceeds another in view of the evidence of two samples

    Thompson, W.R. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika 1933, 25, 285–294

  64. [73]

    MOTS: Minimax Optimal Thompson Sampling

    Jin, T.; Xu, P .; Shi, J.; Xiao, X.; Gu, Q. MOTS: Minimax Optimal Thompson Sampling. In Proceedings of the 38th International Conference on Machine Learning, 2021; pp. 5074–5083. PMLR: 2021

  65. [74]

    A contextual-bandit approach to personalized news article recommendation

    Li, L.; Chu, W.; Langford, J.; Schapire, R.E. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web (WWW ’10), 2010; pp. 661–670

  66. [75]

    Mathematical Analysis of Machine Learning Algorithms; Cambridge University Press: Cambridge, UK, 2023

    Zhang, T. Mathematical Analysis of Machine Learning Algorithms; Cambridge University Press: Cambridge, UK, 2023

  67. [76]

    Stochastic Linear Optimization under Bandit Feedback

    Dani, V .; Hayes, T.P .; Kakade, S.M. Stochastic Linear Optimization under Bandit Feedback. In Proceedings of the COLT, 2008; Volume 2, p. 3

  68. [77]

    Thompson sampling for contextual bandits with linear payoffs

    Agrawal, S.; Goyal, N. Thompson sampling for contextual bandits with linear payoffs. In Proceedings of the 30th International Conference on Machine Learning, 2013; pp. 127–135

  69. [78]

    Learning to optimize via posterior sampling

    Russo, D.; Van Roy, B. Learning to optimize via posterior sampling. Math. Oper. Res. 2014, 39, 1221–1243

  70. [79]

    Neural contextual bandits with UCB-based exploration

    Zhou, D.; Li, L.; Gu, Q. Neural contextual bandits with UCB-based exploration. In Proceedings of the 37th International Conference on Machine Learning, 2020; pp. 11492–11502

  71. [80]

    Neural Thompson Sampling

    Zhang, W.; Zhou, D.; Li, L.; Gu, Q. Neural Thompson Sampling. In Proceedings of the International Conference on Learning Representation (ICLR), 2021

  72. [81]

    Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design

    Srinivas, N.; Krause, A.; Kakade, S.; Seeger, M. Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. In Proceedings of the 27th International Conference on Machine Learning; Omnipress: 2010; pp. 1015–1022

  73. [82]

    A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning

    Brochu, E.; Cora, V .M.; De Freitas, N. A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning. arXiv 2010, arXiv:1012.2599

  74. [83]

    A tutorial on Gaussian process regression: Modelling, exploring, and exploiting functions

    Schulz, E.; Speekenbrink, M.; Krause, A. A tutorial on Gaussian process regression: Modelling, exploring, and exploiting functions. J. Math. Psychol. 2018, 85, 1–16

  75. [84]

    Elements of Information Theory; Wiley-Interscience, Hoboken, USA, 2006

    Thomas, M.; Joy, A.T. Elements of Information Theory; Wiley-Interscience, Hoboken, USA, 2006

  76. [85]

    On kernelized multi-armed bandits

    Chowdhury, S.R.; Gopalan, A. On kernelized multi-armed bandits. In Proceedings of the International Conference on Machine Learning; PMLR: 2017; pp. 844–853

  77. [86]

    Contextual Gaussian Process Bandit Optimization

    Krause, A.; Ong, C. Contextual Gaussian Process Bandit Optimization. Advances in Neural Information Processing Systems 2011, 24, 2447–2455

  78. [87]

    Active learning for level set estimation

    Gotovos, A.; Casati, N.; Hitz, G.; Krause, A. Active learning for level set estimation. In Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence, 2013; pp. 1344–1350

  79. [88]

    Gaussian process classification bandits

    Hayashi, T.; Ito, N.; Tabata, K.; Nakamura, A.; Fujita, K.; Harada, Y.; Komatsuzaki, T. Gaussian process classification bandits. Pattern Recognit. 2024, 149, 110224

  80. [89]

    Gaussian Process Bandits Preprint

    Mathieu, E. Gaussian Process Bandits Preprint. 2016. Available online: https://emilemathieu.fr/files/gpbanditsreport.pdf (accessed on 14th Feb., 2025)

  81. [90]

    Uncertainty Quantification Using Martingales for Misspecified Gaussian Processes

    Neiswanger, W.; Ramdas, A. Uncertainty Quantification Using Martingales for Misspecified Gaussian Processes. Proceedings of the 32nd International Conference on Algorithmic Learning Theory 2021, 132, 963–982. Version February 20, 2025 submitted to arxiv 51 of 52

  82. [91]

    Confidence Bound Minimization for Bayesian Optimization with Student’s-t Processes

    Clare, C.; Hawe, G.; Lin, Z.; McClean, S. Confidence Bound Minimization for Bayesian Optimization with Student’s-t Processes. Proceedings of the 3rd International Conference on Applications of Intelligent Systems 2020, Article 9, 1–5

  83. [92]

    Gaussian Processes for Machine Learning; MIT Press: Cambridge, MA, USA, 2006; Volume 2

    Williams, C.K.; Rasmussen, C.E. Gaussian Processes for Machine Learning; MIT Press: Cambridge, MA, USA, 2006; Volume 2

  84. [93]

    Continuum-Armed Bandits: A Function Space Perspective

    Singh, S. Continuum-Armed Bandits: A Function Space Perspective. Proceedings of The 24th International Conference on Artificial Intelligence and Statistics 2021, 130, 2620–2628

  85. [94]

    Theoretical Foundations of Functional Data Analysis, with an Introduction to Linear Operators; John Wiley & Sons, West Sussex, 2015

    Hsing, T.; Eubank, R. Theoretical Foundations of Functional Data Analysis, with an Introduction to Linear Operators; John Wiley & Sons, West Sussex, 2015

  86. [95]

    Gaussian Process Regression Analysis for Functional Data; CRC Press: 2011

    Shi, J.Q.; Choi, T. Gaussian Process Regression Analysis for Functional Data; CRC Press: 2011

  87. [96]

    Growing-dimensional partially functional linear models: non-asymptotic optimal prediction error

    Zhang, H.; Lei, X. Growing-dimensional partially functional linear models: non-asymptotic optimal prediction error. Phys. Scr. 2023, 98, 095216

  88. [97]

    Functional data analysis for sparse longitudinal data

    Yao, F.; Müller, H.G.; Wang, J.L. Functional data analysis for sparse longitudinal data. J. Am. Stat. Assoc. 2005, 100, 577–590

  89. [98]

    Functional linear regression for discretely observed data: From ideal to reality

    Zhou, H.; Yao, F.; Zhang, H. Functional linear regression for discretely observed data: From ideal to reality. Biometrika 2023, 110, 381–393

  90. [99]

    Stochastic continuum-armed bandits with additive models: Minimax regrets and adaptive algorithm

    Cai, T.T.; Pu, H. Stochastic continuum-armed bandits with additive models: Minimax regrets and adaptive algorithm. Ann. Stat. 2022, 50, 2179–2204

  91. [100]

    Optimal designs for longitudinal and functional data

    Ji, H.; Müller, H.G. Optimal designs for longitudinal and functional data. J. R. Stat. Soc. Ser. B Stat. Methodol. 2017, 79, 859–876

  92. [101]

    One-armed bandit problems with covariates

    Sarkar, J. One-armed bandit problems with covariates. Ann. Stat. 1991, 19, 1978–2002

  93. [102]

    Randomized allocation with nonparametric estimation for a multi-armed bandit problem with covariates

    Yang, Y.; Zhu, D. Randomized allocation with nonparametric estimation for a multi-armed bandit problem with covariates. Ann. Stat. 2002, 30, 100–121

  94. [103]

    The multi-armed bandit problem with covariates

    Perchet, V .; Rigollet, P . The multi-armed bandit problem with covariates. Ann. Stat. 2013, 41, 693–721

  95. [104]

    Sub-sampling for multi-armed bandits

    Baransi, A.; Maillard, O.A.; Mannor, S. Sub-sampling for multi-armed bandits. In Proceedings of the Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2014, Nancy, France, 15–19 September 2014; Proceedings, Part I 14; Springer: 2014; pp. 115–131

  96. [105]

    The Multi-armrd bandit Problem

    Chan, H.P . The Multi-armrd bandit Problem. Ann. Stat. 2020, 48, 346–373

  97. [106]

    A non-parametric solution to the multi-armed bandit problem with covariates

    Ai, M.; Huang, Y.; Yu, J. A non-parametric solution to the multi-armed bandit problem with covariates. J. Stat. Plan. Inference 2021, 211, 402–413

  98. [107]

    Transfer learning for contextual multi-armed bandits

    Cai, C.; Cai, T.T.; Li, H. Transfer learning for contextual multi-armed bandits. Ann. Stat. 2024, 52, 207–232

  99. [108]

    Adaptive algorithm for multi-armed bandit problem with high-dimensional covariates

    Qian, W.; Ing, C.K.; Liu, J. Adaptive algorithm for multi-armed bandit problem with high-dimensional covariates. J. Am. Stat. Assoc. 2024, 119, 970–982

  100. [110]

    Multi-armed bandit for species discovery: A Bayesian nonparametric approach

    Battiston, M.; Favaro, S.; Teh, Y.W. Multi-armed bandit for species discovery: A Bayesian nonparametric approach. J. Am. Stat. Assoc. 2018, 113, 455–466

  101. [111]

    Statistical inference for online decision making: In a contextual bandit setting

    Chen, H.; Lu, W.; Song, R. Statistical inference for online decision making: In a contextual bandit setting. J. Am. Stat. Assoc. 2021, 116, 240–255

  102. [112]

    Stochastic low-rank tensor bandits for multi-dimensional online decision making

    Zhou, J.; Hao, B.; Wen, Z.; Zhang, J.; Sun, W.W. Stochastic low-rank tensor bandits for multi-dimensional online decision making. J. Am. Stat. Assoc. 2024, 1–14. https://doi.org/10.1080/01621459.2024.2311364

  103. [113]

    Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons

    Zhu, B.; Jordan, M.; Jiao, J. Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. Proceedings of the 40th International Conference on Machine Learning 2023, 202, 43037–43067

  104. [114]

    Optimal Design for Reward Modeling in RLHF

    Scheid, A.; Boursier, E.; Durmus, A.; Jordan, M.I.; Ménard, P .; Moulines, E.; Valko, M. Optimal Design for Reward Modeling in RLHF. arXiv 2024, arXiv:2410.17055

  105. [115]

    A review of off-policy evaluation in reinforcement learning

    Uehara, M.; Shi, C.; Kallus, N. A review of off-policy evaluation in reinforcement learning. arXiv 2022, arXiv:2212.06355

  106. [116]

    Bandit problems with infinitely many arms

    Berry, D.A.; Chen, R.W.; Zame, A.; Heath, D.C.; Shepp, L.A. Bandit problems with infinitely many arms. Ann. Stat. 1997, 25, 2103–2116

  107. [117]

    Bayesian nonparametric bandits

    Clayton, M.K.; Berry, D.A. Bayesian nonparametric bandits. Ann. Stat. 1985, 13, 1523–1534

  108. [118]

    Multi-armed bandits with discount factor near one: The Bernoulli case

    Kelly, F. Multi-armed bandits with discount factor near one: The Bernoulli case. Ann. Stat. 1981, 9, 987–1001

  109. [119]

    Bandit processes and dynamic allocation indices

    Gittins, J.C. Bandit processes and dynamic allocation indices. J. R. Stat. Soc. Ser. B Stat. Methodol. 1979, 41, 148–164

  110. [120]

    Multi-armed bandits and the Gittins index

    Whittle, P . Multi-armed bandits and the Gittins index. J. R. Stat. Soc. Ser. B (Methodol.) 1980, 42, 143–149

  111. [121]

    On randomized dynamic allocation indices for the sequential design of experiments

    Glazebrook, K. On randomized dynamic allocation indices for the sequential design of experiments. J. R. Stat. Soc. Ser. B Stat. Methodol. 1980, 42, 342–346

  112. [122]

    Asymptotically efficient strategies for a stochastic scheduling problem with order constraints

    Fuh, C.D.; Hu, I. Asymptotically efficient strategies for a stochastic scheduling problem with order constraints. Ann. Stat. 2000, 28, 1670–1695

  113. [123]

    The learning component of dynamic allocation indices

    Gittins, J.; Wang, Y.G. The learning component of dynamic allocation indices. Ann. Stat. 1992, 20, 1625–1636

  114. [124]

    Strategic two-sample test via the two-armed bandit process

    Chen, Z.; Yan, X.; Zhang, G. Strategic two-sample test via the two-armed bandit process. J. R. Stat. Soc. Ser. B Stat. Methodol. 2023, 85, 1271–1298

  115. [125]

    UCB algorithms for multi-armed bandits: Precise regret and adaptive inference

    Han, Q.; Khamaru, K.; Zhang, C.H. UCB algorithms for multi-armed bandits: Precise regret and adaptive inference. arXiv 2024, arXiv:2412.06126

  116. [126]

    Finite horizon behavior of policies for two-arm bandits

    Fox, B.L. Finite horizon behavior of policies for two-arm bandits. J. Am. Stat. Assoc. 1974, 69, 963–965. Version February 20, 2025 submitted to arxiv 52 of 52

  117. [127]

    Asymptotically efficient allocation rules for two Bernoulli populations

    Li, Z.; Zhang, C.H. Asymptotically efficient allocation rules for two Bernoulli populations. J. R. Stat. Soc. Ser. B Stat. Methodol. 1992, 54, 609–616

  118. [128]

    Kullback-Leibler upper confidence bounds for optimal sequential allocation

    Cappé, O.; Garivier, A.; Maillard, O.A.; Munos, R.; Stoltz, G. Kullback-Leibler upper confidence bounds for optimal sequential allocation. Ann. Stat. 2013, 41, 1516–1541

  119. [129]

    On Bayesian index policies for sequential resource allocation

    Kaufmann, E. On Bayesian index policies for sequential resource allocation. Ann. Stat. 2018, 46, 842–865

  120. [130]

    Response surface bandits

    Ginebra, J.; Clayton, M.K. Response surface bandits. J. R. Stat. Soc. Ser. B Stat. Methodol. 1995, 57, 771–784

  121. [131]

    Pre-trained Gaussian processes for Bayesian optimization

    Wang, Z.; Dahl, G.E.; Swersky, K.; Lee, C.; Nado, Z.; Gilmer, J.; Snoek, J.; Ghahramani, Z. Pre-trained Gaussian processes for Bayesian optimization. J. Mach. Learn. Res. 2024, 25, 1–83

  122. [132]

    Modified two-armed bandit strategies for certain clinical trials

    Berry, D.A. Modified two-armed bandit strategies for certain clinical trials. J. Am. Stat. Assoc. 1978, 73, 339–345

  123. [133]

    Randomized allocation of treatments in sequential experiments

    Bather, J. Randomized allocation of treatments in sequential experiments. J. R. Stat. Soc. Ser. B (Methodol.) 1981, 43, 265–283

  124. [134]

    Optimal adaptive randomized designs for clinical trials

    Cheng, Y.; Berry, D.A. Optimal adaptive randomized designs for clinical trials. Biometrika 2007, 94, 673–689

  125. [135]

    An information theoretic approach for selecting arms in clinical trials

    Mozgunov, P .; Jaki, T. An information theoretic approach for selecting arms in clinical trials. J. R. Stat. Soc. Ser. B Stat. Methodol. 2020, 82, 1223–1247

  126. [136]

    False discovery rate control with e-values

    Wang, R.; Ramdas, A. False discovery rate control with e-values. J. R. Stat. Soc. Ser. B Stat. Methodol. 2022, 84, 822–852

  127. [137]

    Variable selection via Thompson sampling.J

    Liu, Y.; Roˇ cková, V . Variable selection via Thompson sampling.J. Am. Stat. Assoc. 2023, 118, 287–304

  128. [138]

    Dynamic online pricing with incomplete information using multiarmed bandit experiments

    Misra, K.; Schwartz, E.M.; Abernethy, J. Dynamic online pricing with incomplete information using multiarmed bandit experiments. Mark. Sci. 2019, 38, 226–252

  129. [139]

    Effective Adaptive Exploration of Prices and Promotions in Choice-Based Demand Models

    Jain, L.; Li, Z.; Loghmani, E.; Mason, B.; Yoganarasimhan, H. Effective Adaptive Exploration of Prices and Promotions in Choice-Based Demand Models. Mark. Sci. 2024, 43, 925–1151

  130. [140]

    Bandit Interpretability of Deep Models via Confidence Selection

    Duan, X.; Li, H.; Wang, P .; Wang, T.; Liu, B.; Zhang, B. Bandit Interpretability of Deep Models via Confidence Selection. Neurocomputing 2023, 544, 126250

  131. [141]

    Replication or exploration? Sequential design for stochastic simulation experiments

    Binois, M.; Huang, J.; Gramacy, R.B.; Ludkovski, M. Replication or exploration? Sequential design for stochastic simulation experiments. Technometrics 2019, 61, 7–23

  132. [142]

    Risk-averse heteroscedastic bayesian optimization

    Makarova, A.; Usmanova, I.; Bogunovic, I.; Krause, A. Risk-averse heteroscedastic bayesian optimization. Adv. Neural Inf. Process. Syst. 2021, 34, 17235–17245

  133. [143]

    Digital Triplet: A Sequential Methodology for Digital Twin Learning

    Zhang, X.; Lin, D.K.; Wang, L. Digital Triplet: A Sequential Methodology for Digital Twin Learning. Mathematics 2023, 11, 2661

  134. [144]

    Towards a digital twin framework in additive manufacturing: Machine learning and bayesian optimization for time series process optimization

    Karkaria, V .; Goeckner, A.; Zha, R.; Chen, J.; Zhang, J.; Zhu, Q.; Cao, J.; Gao, R.X.; Chen, W. Towards a digital twin framework in additive manufacturing: Machine learning and bayesian optimization for time series process optimization. J. Manuf. Syst. 2024, 75, 322–332

  135. [145]

    A survey on contextual multi-armed bandits

    Zhou, L. A survey on contextual multi-armed bandits. arXiv 2015, arXiv:1508.03326

  136. [146]

    Conservative Bandits

    Wu, Y.; Shariff, R.; Lattimore, T.; Szepesvári, C. Conservative Bandits. Proceedings of the 33rd International Conference on Machine Learning 2016, 48, 1254–1262

  137. [147]

    Residual Bootstrap Exploration for Stochastic Linear Bandit

    Wu, S.; Wang, C.H.; Li, Y.; Cheng, G. Residual Bootstrap Exploration for Stochastic Linear Bandit. Proceedings of the Thirty-Eighth Conference on Uncertainty in Artificial Intelligence 2022, 180, 2117–2127

  138. [148]

    Estimating concentration parameters for bandit algorithms

    Lieber, J. Estimating concentration parameters for bandit algorithms. Job Market Paper 2022. Available online: https://jonaslieber. com/research.html (accessed on 14th Feb., 2025)

  139. [149]

    Multiplier Bootstrap-Based Exploration

    Wan, R.; Wei, H.; Kveton, B.; Song, R. Multiplier Bootstrap-Based Exploration. Proceedings of the 40th International Conference on Machine Learning 2023, 1476, 35444–35490

  140. [150]

    Perturbed-history exploration in stochastic multi-armed bandits

    Kveton, B.; Szepesvári, C.; Ghavamzadeh, M.; Boutilier, C. Perturbed-history exploration in stochastic multi-armed bandits. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, 2019; pp. 2786–2793

  141. [151]

    Perturbed-History Exploration in Stochastic Linear Bandits.Proceedings of the 35th Uncertainty in Artificial Intelligence Conference 2020, 530–540

    Kveton, B.; Szepesvári, C.; Ghavamzadeh, M.; Boutilier, C. Perturbed-History Exploration in Stochastic Linear Bandits.Proceedings of the 35th Uncertainty in Artificial Intelligence Conference 2020, 530–540

  142. [152]

    Universal and data-adaptive algorithms for model selection in linear contextual bandits

    Muthukumar, V .K.; Krishnamurthy, A. Universal and data-adaptive algorithms for model selection in linear contextual bandits. Proceedings of the 39th International Conference on Machine Learning 2022, 16197–16222

  143. [153]

    Model Selection for Contextual Bandits and Reinforcement Learning

    Pacchiano Camacho, A. Model Selection for Contextual Bandits and Reinforcement Learning. Ph.D. Thesis, UC Berkeley, Berkeley, CA, USA, 2021. Under review

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.