REVIEW 4 major objections 1 references
This paper proves that the widely used expected improvement acquisition function in Bayesian optimization has sublinear cumulative regret—no-regret—for two standard incumbent choices, and either sublinear regret or fast simple-regret decay
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Expected improvement with best-posterior-mean or best-sampled-posterior-mean incumbents is proven no-regret for Gaussian process objectives with squared exponential or Matérn kernels.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection Abstract promises a landmark BO theory result, but the full text is a robotics paper — the claims are unverifiable. the 4 major comments →
Bayesian Optimization with Expected Improvement: No Regret and the Choice of Incumbent
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
In the paper's terms: the noisy Gaussian process expected improvement algorithm, with BPMI or BSPMI as the incumbent, is no-regret for squared exponential and Matérn kernels; with BOI, it either achieves sublinear cumulative regret or a fast-converging noisy simple regret bound. This is presented for the first time. The result holds in the Bayesian setting—the objective is a sample from a GP with known kernel hyperparameters—and the bounds are cumulative regret upper bounds, meaning the total deficit relative to the best point found does not grow linearly with the number of evaluations.
What carries the argument
The named object is the acquisition function: expected improvement (EI), defined as the expected amount by which a candidate's posterior predictive value exceeds the incumbent. The paper's load-bearing machinery is the definition of the incumbent, in three flavors: BPMI (max of the posterior mean), BSPMI (max of a sample from the posterior mean), and BOI (best noisy observation). The proof works by bounding the cumulative regret of GP-EI in the Bayesian setting, with the SE and Matérn kernel assumptions supplying the regularity that keeps the posterior from being misled by noise. The incumbent choice determines how aggressively the algorithm treats the current best, and the paper shows that
Load-bearing premise
The proofs assume the objective is a random sample from a Gaussian process with known kernel hyperparameters; for an arbitrary deterministic function the no-regret guarantees are not claimed.
What would settle it
Take a fixed sample path from an SE- or Matérn-kernel GP with known hyperparameters, add Gaussian observation noise, run GP-EI with BPMI/BSPMI for many iterations, and measure cumulative regret. If the regret grows linearly rather than sublinearly, the paper's central claim is wrong.
If this is right
- GP-EI with BPMI or BSPMI can be used in noisy Bayesian optimization without worrying about linear regret: the cumulative regret is guaranteed to grow sublinearly for SE and Matérn kernels.
- The BOI variant—the most natural incumbent from raw observations—is also safe in a weaker sense: it either keeps cumulative regret sublinear or drives the noisy simple regret down quickly.
- The results supply the missing theoretical counterpart to EI's empirical success, putting it on the same footing as acquisition functions that already had regret bounds.
- The choice of incumbent is not merely a numerical detail: the paper's bounds give a formal reason to prefer BPMI/BSPMI for cumulative-regret guarantees while explaining why BOI can still be effective.
- Because the bounds are for the Bayesian setting, they give a baseline against which other acquisition functions and incumbent choices can be compared in future theoretical work.
Where Pith is reading between the lines
- The proof route is likely to extend to other kernels whose posterior concentrates quickly (e.g., other stationary kernels with similar spectral decay), although the paper only states results for SE and Matérn.
- The BOI result may rationalize common practice: practitioners who use BOI and observe good performance may be seeing fast simple-regret convergence even when cumulative regret is not sublinear.
- For finite budgets, the theory says nothing about constants, so the optimal incumbent in practice may be data-dependent even though all three are asymptotically safe.
- If the objective is not actually a GP sample, the Bayesian regret bounds do not directly apply; a testable extension would be to benchmark the same incumbent choices on deterministic benchmark functions and compare empirical regret growth rates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript, as submitted, consists of an abstract claiming new cumulative regret upper bounds for noisy Gaussian-process expected improvement (GP-EI) under the best posterior mean incumbent (BPMI), best sampled posterior mean incumbent (BSPMI), and best observation incumbent (BOI), with no-regret guarantees for squared exponential and Matérn kernels, plus numerical validation. However, the full text supplied is an entirely different paper on dexterous manipulation, 'Exploiting Policy Idling for Dexterous Manipulation.' That text contains no definitions of GP-EI, no theorem statements, no proofs, no analysis of regret, and no experiments on Bayesian optimization. Consequently, the claimed contribution is completely unsupported by the provided manuscript body.
Significance. If the claims in the abstract are correct, they would close a recognized gap in the theoretical understanding of expected improvement, a widely used acquisition function, and would provide concrete guidance on incumbent selection in noisy GP-BO. That would be a valuable contribution to the BO literature. However, the submitted full text provides no evidence for these claims, so the significance cannot be assessed or credited on the basis of this manuscript.
major comments (4)
- [Full text (entire provided manuscript body)] The full text is a robotics paper on 'Exploiting Policy Idling for Dexterous Manipulation.' It contains no mention of Gaussian processes, expected improvement, Bayesian optimization, cumulative regret, BPMI, BSPMI, BOI, or any theorems. Thus the central claim stated in the abstract—first-time cumulative regret upper bounds for GP-EI—has no supporting derivation, proof, or even statement in the manuscript. This is a load-bearing absence: the claimed contribution cannot be checked or accepted.
- [Abstract (no-regret claims for SE and Matérn kernels)] The abstract asserts no-regret guarantees for both squared exponential and Matérn kernels under BPMI and BSPMI, and a sublinear or fast-converging result for BOI. No assumptions, definitions, theorem statements, or proof sketches appear anywhere in the provided text. It is impossible to verify the correctness of these claims, the exact form of the bounds, or whether the proofs rely on circular reasoning. The manuscript must contain the full mathematical development before evaluation is possible.
- [Abstract (numerical experiments)] The abstract states that 'Numerical experiments are conducted to validate our findings,' but the provided full text contains no experiments related to Bayesian optimization or regret. The only experimental content concerns robot manipulation. The claimed empirical validation is therefore entirely missing.
- [References and related work] The manuscript body's references are to robotics literature and do not engage with the existing GP-EI regret literature. Consequently, the claimed novelty ('for the first time') cannot be situated relative to prior work on EI regret bounds, nor can the reader check whether the manuscript builds on, improves, or corrects existing results.
Circularity Check
No circularity identified; the provided full text is an unrelated robotics paper, so the claimed regret bounds have no derivational chain to inspect.
full rationale
The target manuscript (arXiv:2508.15674) is represented only by its abstract, which announces cumulative regret upper bounds for GP-EI with BPMI, BSPMI, and BOI under SE and Matérn kernels. The full text supplied is a different paper, 'Exploiting Policy Idling for Dexterous Manipulation,' containing no theorems, proofs, or equations relevant to Bayesian optimization or expected improvement. There is therefore no derivation chain, no fitted parameter renamed as a prediction, no self-citation used as a load-bearing premise, and no uniqueness theorem imported from the authors' prior work. The abstract itself defines the incumbents (BPMI, BSPMI, BOI) as standard choices and states results about regret bounds; nothing in the presented text shows a conclusion that is equivalent by construction to its inputs. The absence of the proof is a completeness/verifiability problem, not a circularity problem. Per the instruction to reserve circularity findings for cases where the specific reduction can be quoted, no circular step is identified here. The reader's take of 'unverdictable due to missing proof' is appropriate, but this does not raise the circularity score.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption The objective function is a sample from a Gaussian process prior
- domain assumption The kernel is squared exponential or Matérn
Cite this review
Pith. "Pith review of Bayesian Optimization with Expected Improvement: No Regret and the Choice of Incumbent." pith.science (2026). https://pith.science/paper/E3QT75SV
@misc{pith2026250815674,
author = {Pith},
title = {Pith review of: Bayesian Optimization with Expected Improvement: No Regret and the Choice of Incumbent},
year = {2026},
howpublished = {\url{https://pith.science/paper/E3QT75SV}},
note = {Machine review of arXiv:2508.15674}
}
read the original abstract
Expected improvement (EI) is one of the most widely used acquisition functions in Bayesian optimization (BO). Despite its proven empirical success in applications, the cumulative regret upper bound of EI remains an open question. In this paper, we analyze the classic noisy Gaussian process expected improvement (GP-EI) algorithm. We consider the Bayesian setting, where the objective is a sample from a GP. Three commonly used incumbents, namely the best posterior mean incumbent (BPMI), the best sampled posterior mean incumbent (BSPMI), and the best observation incumbent (BOI) are considered as the choices of the current best value in GP-EI. We present for the first time the cumulative regret upper bounds of GP-EI with BPMI and BSPMI. Importantly, we show that in both cases, GP-EI is a no-regret algorithm for both squared exponential (SE) and Mat\'ern kernels. Further, we present for the first time that GP-EI with BOI either achieves a sublinear cumulative regret upper bound or has a fast converging noisy simple regret bound for SE and Mat\'ern kernels. Our results provide theoretical guidance to the choice of incumbent when practitioners apply GP-EI in the noisy setting. Numerical experiments are conducted to validate our findings.
Reference graph
Works this paper leans on
-
[1]
Exploiting Policy Idling for Dexterous Manipulation
Exploiting Policy Idling for Dexterous Manipulation Annie S. Chen 1, Philemon Brakel 2, Antonia Bronars 3, Annie Xie 2, Sandy Huang, Oliver Groth 2, Maria Bauza 2, Markus Wulfmeier 2, Nicolas Heess 2, Dushyant Rao 2 Abstract— Learning-based methods for dexterous manipu- lation have made notable progress in recent years. However, learned policies often sti...
work page internal anchor Pith review Pith/arXiv arXiv 2025
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.