Pith. sign in

REVIEW 5 major objections 7 minor 38 references

An automatic multi-chain HMC warmup controller starts diagonal, promotes to low-rank when evidence supports it, and turns inconclusive or conflicting signals into concrete next actions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 12:26 UTC pith:XWD54HYF

load-bearing objection Real shipped auto-warmup controller with honest theory–practice gap; headline ESS ratios look softer once you read the author’s own ablations. the 5 major comments →

arxiv 2607.23788 v2 pith:XWD54HYF submitted 2026-07-26 stat.CO stat.ME

The Universal Warmup Path: Automatic Preconditioner Selection for HMC

classification stat.CO stat.ME MSC 65C0562F1565C40
keywords Hamiltonian Monte Carlowarmup adaptationmass matrixpreconditioner selectionNUTSlow-rank inverse massmulti-chain diagnosticsESS per gradient
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Euclidean HMC needs a step size and a constant preconditioner chosen from short, nonstationary warmup draws, yet standard practice still forces the user to pick the mass-matrix structure in advance and follow a fixed schedule. This paper offers one multi-chain controller that begins with a diagonal inverse mass matrix, checks evidence only at dimension-derived window endpoints, and either promotes to a low-rank-plus-diagonal matrix (choosing the retained rank under sample-support caps) or stays diagonal. When the evidence is inconclusive it simply gathers the next scheduled window; when chains persistently disagree it keeps a within-region matrix and advises a population or tempering method; when score–position linearity is poor it advises reparameterization. On the headline NUTS benchmarks the controller selected low rank in every run and delivered higher pooled ESS per gradient than both a prespecified Fisher low-rank warmup and a Welford diagonal warmup, with all runs passing the post-warmup quality check. A sympathetic reader cares because the method collapses several common warmup heuristics into one evidence-driven path and converts familiar warning signs into actionable guidance instead of silent failure.

Core claim

A single automatic controller can choose, from multi-chain warmup evidence alone, whether a constant diagonal or low-rank-plus-diagonal inverse mass matrix is supported, which rank to keep, and when to stop adapting the mass matrix—while preserving step-size adaptation and turning persistent within/between disagreement or poor score–position linearity into explicit advice for population methods or reparameterization—and on the evaluated NUTS benchmarks this controller selected low rank in 12/12 runs and improved pooled ESS per gradient relative to both prespecified Fisher low-rank and Welford diagonal baselines.

What carries the argument

The route-indexed attractor: after a confidence-aware latch promotes an estimator (a “route”), finite-window iterates of that estimator’s population update contract toward its fixed target, with an explicit additive budget for starting-law, step-size, whitening, sampling, and regularization error; eligibility is gated by within-chain and between-means structural tests plus a held-out score–position linearity check.

Load-bearing premise

The controller’s fixed diagnostic thresholds are treated as adequate stand-ins for the theory’s confidence sets and five-term finite-window error budget, even though the paper states that connecting the two still requires calibration.

What would settle it

Re-run the preregistered headline NUTS configurations (ill-conditioned Gaussian and German-credit logistic regression, M=8, same gradient budgets and seeds): if the automatic controller fails to select low rank in most runs, or its pooled ESS-per-gradient geometric-mean ratios fall to or below 1 against the Fisher low-rank and Welford diagonal baselines while quality checks still pass, the central efficiency-after-selection claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Users can drop the manual choice of diagonal versus dense/low-rank mass matrix for ordinary Euclidean HMC/NUTS warmup.
  • Persistent within-/between-chain disagreement becomes an explicit handoff signal to population or tempering methods rather than a silent failure of a single constant preconditioner.
  • Poor held-out score–position linearity becomes an explicit reparameterization signal while still leaving a usable within-region matrix.
  • Schedule and rank caps become dimension- and budget-derived defaults instead of per-model tuning knobs.
  • Once thresholds are calibrated to the finite-window error budget, the same controller can be composed with held-out metric ranking or a regional-exploration companion.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same within/between evidence split could gate adaptation in other multi-chain MCMC kernels that currently freeze a single metric by hand.
  • Calibrating the five error terms in the attractor bound would turn the advisory flags into statistically timed stopping rules rather than fixed-threshold heuristics.
  • The controlled GMM sweep suggests a practical diagnostic for when constant Euclidean geometry is simply the wrong global description, independent of any particular sampler.
  • Budget-identical comparisons against single-chain historical policies would clarify how much of the reported gain is multi-chain information versus schedule redesign.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper proposes an automatic multi-chain warmup controller for Euclidean HMC/NUTS that starts with a diagonal inverse mass matrix and, at dimension-derived window endpoints fixed before sampling, decides whether to promote to a low-rank-plus-diagonal matrix, gather another window, retain a within-region matrix with a population/tempering advisory (persistent within/between disagreement), or advise reparameterization (poor held-out score–position linearity). The formal development gives a route-indexed attractor theorem with an explicit five-term finite-window error budget, operator-confidence consequences for rank/subspace decisions, an exact within/between covariance decomposition, and a finite-horizon Le Cam-style local-transcript lower bound. Empirically, on two NUTS benchmarks the controller selected low rank in 12/12 runs and reports geometric-mean pooled ESS/gradient ratios of 2.451–22.572 against prespecified Fisher low-rank and Welford diagonal baselines, with all 36 runs passing quality checks; a controlled GMM sweep holds the marginal spectrum fixed while the within/between evidence and selection change.

Significance. If the empirical claims hold up under uncertainty quantification, this is a useful contribution: an automatic, budget-driven warmup controller that removes a real manual choice (preconditioner structure) from HMC practice, merged into BlackJAX with pinned revisions, frozen manifests, and checksummed outputs. The preregistered baseline roles, float64/explicit-seed discipline, uniform quality gates, and — unusually — the authors' own reported null results in Appendix C are genuine strengths. The theoretical sections are conditional and honestly scoped, and Theorem 6's local-transcript bound is a clean formulation of a real limitation. However, the theory does not currently certify the deployed controller (by the authors' own statement), and the headline efficiency numbers rest on a confounded metric reported without uncertainty, so the practical significance is not yet established at the level the abstract asserts.

major comments (5)
  1. [§7, Table 1] Table 1 reports only bare geometric-mean ESS/gradient ratios: no seed count, no per-seed spread, no interval. Appendix C shows the authors can and do apply a stricter standard: the k=3 matched ablation reports geomean 1.037 with seed-cluster t-interval [0.998, 1.077] and declares a null because the interval covers 1. The headline numbers should meet the same standard: report seeds, per-seed ratios, and intervals for all four Table 1 cells, and state a predeclared go/no-go criterion.
  2. [§7 / §3] The pooled ESS/gradient metric charges all warmup gradients, but the controller chooses its warmup length via the dimension-derived schedule while both baselines run a fixed 312 transitions per chain. A controller that simply stops warmup earlier wins this metric with an identical final preconditioner. The ratio therefore conflates estimator selection with budget allocation. A matched-budget arm (e.g., the controller's frozen metric plugged into the 312-transition schedule, and vice versa, or sampling-only ESS/gradient) is needed to attribute the gain.
  3. [Appendix C vs §7] The matched ablations undermine the mechanism implied by the headline: where low rank is selected under controlled conditions, low-rank/matched-diagonal projection ESS/gradient ratios are 0.869–1.170 (k=2) and 1.037 [0.998, 1.077] (k=3, declared a null). The Discussion concedes 'no general efficiency gain' from the low-rank structure, but the abstract's framing (12/12 low-rank selection alongside 2–22× gains) implies the selection drives the efficiency. These need reconciling: either the headline gains are schedule/allocation effects, or evidence is needed that the low-rank benefit transfers outside the GMM testbed.
  4. [§3 / §7] §3 defines the Fisher-HMC reference prescription as the 30/55/15 schedule of Seyboldt et al. with L=10/L=80 refresh periods, memory wipes, and step-size reinitialization. Yet Table 1's 'Fisher low-rank warmup' baseline used 312 transitions under the same proportional growing-window schedule as the Welford baseline — i.e., the Fisher estimator transplanted into a Stan-like schedule, not the prespecified Fisher-HMC warmup. The comparison should be relabeled accordingly, or an arm running the actual Fisher prescription should be added; as written the baseline name overstates what was compared.
  5. [§7 / §5.4] The advertised 12/12 low-rank selection rate is computed on two targets where low rank is expected a priori (an ill-conditioned Gaussian, where the 22.572 ratio vs diagonal is close to foregone, and a mild logistic posterior). This demonstrates promotion but not discrimination: no headline target where diagonal is the right answer is included. The GMM sweep partially tests selectivity but at d=5 with the marginal spectrum pinned. At least one benchmark where diagonal retention is correct would make the selection claim evidentiary rather than assumed.
minor comments (7)
  1. [§7] 4,996 warmup divergences are reported without comment. Even with zero post-warmup divergences, this count deserves discussion: where in the schedule they occur, whether the 20 gradients/transition allowance or window transitions drive them, and whether they affect the evidence tests.
  2. [Appendix C, Table 3] The SR=10 row shows 1/2 low-rank/diagonal after 0/3 at SR=9 and 9.5 — non-monotone selection at the extreme of the sweep. A sentence on whether this is threshold noise or a real boundary effect would help.
  3. [Title/abstract] 'Universal' in the title is explicitly scoped in §1 to the evaluated HMC-family kernels, which is appropriate, but the abstract does not carry that qualification. Consider scoping the abstract claim as the introduction does.
  4. [Appendix B, Table 2] The one-output estimand's comparator is a historical single-chain 2,500-warmup policy, explicitly 'descriptive and not budget-identical.' Since it is not budget-identical, consider moving it out of Table 2 or visually separating it to avoid readers averaging across the two columns.
  5. [Figure 2] Figure 2 references a 'population/regional companion, Paper 3,' an unpublished companion paper. The manuscript should stand alone; replace with a generic pointer to population/tempering methods.
  6. [§5.2, Eq. (7)] The illustrative Wald radius in Eq. (7) is well-hedged, but the support-recovery assumption ((I−QQᵀ)u = o_p(n^{-1/2})) is strong for a singular Γ_a; a remark on when this is verifiable from the transcript would be useful.
  7. [Throughout] Rendering artifacts: 'bR' appears throughout where R̂ is intended (an ArviZ/Unicode issue); 'LR V' is split. Please fix for readability.

Circularity Check

0 steps flagged

No significant circularity: efficiency claims are external benchmark ratios, and the theory is standard operator/attractor analysis not forced by its inputs.

full rationale

The load-bearing empirical claim is geometric-mean pooled ESS-per-gradient of the automatic controller versus two preregistered external warmups (prespecified Fisher low-rank and Welford diagonal) on fixed targets, with post-warmup quality checks. Those ratios are not quantities fitted from the controller and then re-reported as predictions; the baselines use independent estimator prescriptions and a shared nominal schedule. The formal chain (route-indexed local attractor with explicit finite-window error budget, W/B transcript identity, Fisher geometric-mean bound, operator consequences of a residual ball, Le Cam local-transcript coupling) invokes standard external mathematics and does not define the estimand in terms of the deployed diagnostic thresholds. The paper explicitly separates identification, representativeness, and utility, and states that connecting fixed controller thresholds to Theorem 2 still requires calibration of the five-term error budget—so the theory is not smuggled in as a self-justifying uniqueness result. BlackJAX and the Project Geodesic archive are implementation/host citations, not load-bearing uniqueness theorems. No step reduces a claimed prediction to its own fitted input by construction.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

The central empirical claim rests on engineering schedule constants and fixed diagnostic thresholds more than on deep new physical postulates. Theoretical claims rest on standard MCMC regularity (geometric ergodicity, dissipativity, Stein boundaries) plus the paper's confidence-latch and no-demotion route episode. The largest unpaid bridge is treating fixed thresholds as if they delivered the theorem's calibrated coverage.

free parameters (5)
  • k_cap = min{50, max(⌊d/2⌋, 1)} = min{50, max(floor(d/2), 1)}
    Dimension-derived rank capacity cap chosen by design; directly limits retained low-rank structure and N_min.
  • N_min = 8(k_cap + 1) = 8(k_cap+1)
    Pooled support floor that sets first window length; factor 8 is a hand-chosen sample-support multiplier.
  • nominal 1.5× mass-matrix window growth and final 15% step-size-only phase = 1.5x growth; last 15% step-size only
    Schedule fractions fixed before sampling; not derived from a risk bound in the paper.
  • NUTS gradients-per-transition allowance = 20 = 20
    Conservative budget split ⌊B/(20M)⌋ transitions per chain; engineering constant affecting all window endpoints.
  • fixed diagnostic thresholds for W, T, R² linearity, and disagreement/handoff = implementation-fixed (not numerically tabulated in main text)
    Eligibility and advice gates in the deployed controller; paper admits they are not yet calibrated to theoretical δ_k / e_k budgets.
axioms (5)
  • domain assumption During a route episode the population update field h_a is dissipative and linearly bounded near θ*_a so a small enough step α yields contraction q<1 (Theorem 2).
    Standard SA/MCMC stability ingredient; required for the finite-window attractor bound but not verified per target in the empirical corpus.
  • domain assumption Confidence sets C_k cover the routing estimand η_k with spending ∑δ_k except on a controlled failure event (Lemma 1).
    Soundness of the no-demotion latch; deployed controller substitutes fixed thresholds rather than explicit anytime-valid sets.
  • standard math Frozen-kernel geometric ergodicity and moment conditions for the Markov CLT / long-run covariance of estimator functionals (eq. 6).
    Cited via Jones 2004 / Vats et al. 2018; needed for operator-confidence radii in Theorem 5.
  • standard math Stein boundary conditions and square-integrable score for the full-rank Fisher covariance bound (Proposition 4).
    Classical integration-by-parts setup; paper correctly notes finite-rank/regularized Fisher is not identified from covariance alone.
  • ad hoc to paper Constant Euclidean inverse-mass adequacy is judged from observed warmup transcript (within/between split and held-out score–position linearity), not from global posterior knowledge.
    Operational criterion that licenses promotion, diagonal retention, reparameterization advice, and population handoff; central to the controller's semantics.
invented entities (3)
  • Universal Warmup Path / staged auto metric controller independent evidence
    purpose: Single evidence-driven procedure that schedules windows, selects diagonal vs low-rank-plus-diagonal IMM, and emits advisory handoffs.
    Primary algorithmic object; implemented in BlackJAX meta adaptation. Independent evidence is the frozen benchmark corpus, not an external physical measurement.
  • Route-indexed latch / route episode no independent evidence
    purpose: Formal selected-estimator state with no-demotion rule during a confidence-supported episode.
    Organizes Theorem 2; bookkeeping construct rather than a new physical entity. No external falsifiable handle beyond algorithm behavior.
  • metric_scope / observed_ensemble_evidence / handoff verdict flags no independent evidence
    purpose: User-facing conditioning-scope and population-tempering advice signals from warmup evidence.
    Productized diagnostics tied to the paper's within/between story; validated only inside the paper's GMM and headline runs.

pith-pipeline@v1.2.0-grok45-kimik3 · 22159 in / 4231 out tokens · 80680 ms · 2026-07-30T12:26:49.103434+00:00 · methodology

0 comments
read the original abstract

Euclidean Hamiltonian Monte Carlo (HMC) warmup must choose a step size and constant preconditioner from limited, nonstationary draws. Standard warmup follows a fixed schedule and generally requires the preconditioner structure to be specified in advance. We present a multi-chain controller that starts diagonal and, at dimension-derived window endpoints, selects between diagonal and low-rank-plus-diagonal inverse mass matrices and chooses the retained rank, subject to dimension and sample-support caps. When evidence is inconclusive, the controller gathers another scheduled window. Persistent within-/between-chain disagreement means draws do not support treating one constant preconditioner as an adequate global description; the controller retains its within-region matrix and advises a population or tempering method for regional exploration. Poor held-out score--position linearity advises reparameterization. The controller selected low rank in every evaluated headline benchmark NUTS run ($12/12$). Geometric-mean pooled ESS-per-gradient ratios relative to the prespecified Fisher low-rank warmup and Welford diagonal warmup baselines were respectively $2.451$ and $22.572$ on the synthetic ill-conditioned Gaussian benchmark, and $1.951$ and $6.264$ on the German-credit Bayesian logistic-regression posterior; all compared runs passed the post-warmup quality check. This method unifies common HMC warmup heuristics in one evidence-driven controller, reducing manual choices and turning warning signals into actionable guidance.

Figures

Figures reproduced from arXiv: 2607.23788 by Junpeng Lao.

Figure 1
Figure 1. Figure 1: The evaluated automatic HMC warmup controller. Setup fixes every window boundary [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Practical interpretation of the posterior covariance reference through the observed [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Iid Gaussian spiked-PCA calibration (d = 200). The curves show the asymptotic sample eigenvalue and eigenvector overlap; markers are fixed-seed simulations. The population threshold is 1 + √γ, and the null sample edge is (1 + √γ) 2 . The figure is not a diagnostic threshold for Markov-chain draws. A preconditioner-error claim additionally requires a per-estimator perturbation bound for the moment-to-precon… view at source ↗
Figure 4
Figure 4. Figure 4: Current-corpus controlled GMM boundary, generated from the current [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Schedule evidence and one executed controller trace. The Stan and Fisher-HMC lanes [PITH_FULL_IMAGE:figures/full_fig_p022_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

38 extracted references · 4 linked inside Pith

  1. [1]

    , title =

    Neal, Radford M. , title =. Handbook of Markov Chain Monte Carlo , editor =

  2. [2]

    and Gelman, Andrew , title =

    Hoffman, Matthew D. and Gelman, Andrew , title =. Journal of Machine Learning Research , volume =

  3. [3]

    Duane, Simon and Kennedy, A. D. and Pendleton, Brian J. and Roweth, Duncan , title =. Physics Letters B , volume =

  4. [4]

    and Lee, Daniel and Goodrich, Ben and Betancourt, Michael and Brubaker, Marcus and Guo, Jiqiang and Li, Peter and Riddell, Allen , title =

    Carpenter, Bob and Gelman, Andrew and Hoffman, Matthew D. and Lee, Daniel and Goodrich, Ben and Betancourt, Michael and Brubaker, Marcus and Guo, Jiqiang and Li, Peter and Riddell, Allen , title =. Journal of Statistical Software , volume =. 2017 , doi =

  5. [5]

    2026 , howpublished =

  6. [6]

    Phase Transition of the Largest Eigenvalue for Nonnull Complex Sample Covariance Matrices , journal =

    Baik, Jinho and Ben Arous, G. Phase Transition of the Largest Eigenvalue for Nonnull Complex Sample Covariance Matrices , journal =

  7. [7]

    Rank-Normalization, Folding, and Localization: An Improved

    Vehtari, Aki and Gelman, Andrew and Simpson, Daniel and Carpenter, Bob and B. Rank-Normalization, Folding, and Localization: An Improved. Bayesian Analysis , volume =. 2021 , doi =

  8. [8]

    , title =

    Gelman, Andrew and Rubin, Donald B. , title =. Statistical Science , volume =. 1992 , doi =

  9. [9]

    and Radul, Alexey and Sountsov, Pavel , title =

    Hoffman, Matthew D. and Radul, Alexey and Sountsov, Pavel , title =. Proceedings of the 24th International Conference on Artificial Intelligence and Statistics (AISTATS) , series =

  10. [10]

    and Sountsov, Pavel , title =

    Hoffman, Matthew D. and Sountsov, Pavel , title =. Proceedings of the 25th International Conference on Artificial Intelligence and Statistics (AISTATS) , series =

  11. [11]

    and Carpenter, Bob , title =

    Seyboldt, Adrian and Carlson, Eliot L. and Carpenter, Bob , title =. arXiv preprint arXiv:2603.18845 , year =. 2603.18845 , archivePrefix =

  12. [12]

    arXiv preprint arXiv:2506.18746 , year =

    Bou-Rabee, Nawaf and Carpenter, Bob and Kleppe, Tore Selland and Liu, Sifan , title =. arXiv preprint arXiv:2506.18746 , year =. 2506.18746 , archivePrefix =

  13. [13]

    arXiv preprint arXiv:1905.11916 , year =

    Bales, Ben and Pourzanjani, Arya and Vehtari, Aki and Petzold, Linda , title =. arXiv preprint arXiv:1905.11916 , year =. 1905.11916 , archivePrefix =

  14. [14]

    and Thomas, Joy A

    Dembo, Amir and Cover, Thomas M. and Thomas, Joy A. , title =. IEEE Transactions on Information Theory , volume =. 1991 , doi =

  15. [15]

    arXiv preprint arXiv:2402.10797 , year =

    Cabezas, Alberto and Corenflos, Adrien and Lao, Junpeng and Louf, R. arXiv preprint arXiv:2402.10797 , year =

  16. [16]

    Statistica Sinica , volume =

    Paul, Debashis , title =. Statistica Sinica , volume =

  17. [17]

    , title =

    Jones, Galin L. , title =. Probability Surveys , volume =. 2004 , doi =

  18. [18]

    Stability of Stochastic Approximation under Verifiable Conditions , journal =

    Andrieu, Christophe and Moulines,. Stability of Stochastic Approximation under Verifiable Conditions , journal =. 2005 , doi =

  19. [19]

    Convergence of

    Fort, Gersende and Moulines,. Convergence of. SIAM Journal on Control and Optimization , volume =. 2016 , doi =

  20. [20]

    and Rosenthal, Jeffrey S

    Roberts, Gareth O. and Rosenthal, Jeffrey S. , title =. Journal of Applied Probability , volume =. 2007 , doi =

  21. [21]

    and Jones, Galin L

    Vats, Dootika and Flegal, James M. and Jones, Galin L. , title =. Bernoulli , volume =. 2018 , doi =

  22. [22]

    Information and Inference: A Journal of the IMA , volume =

    Neeman, Joe and Shi, Bobby and Ward, Rachel , title =. Information and Inference: A Journal of the IMA , volume =. 2024 , doi =

  23. [23]

    Bhatia, Rajendra , title =

  24. [24]

    , title =

    Yu, Yi and Wang, Tengyao and Samworth, Richard J. , title =. Biometrika , volume =. 2015 , doi =

  25. [25]

    , title =

    Higham, Nicholas J. , title =. 2008 , doi =

  26. [26]

    , title =

    Baik, Jinho and Silverstein, Jack W. , title =. Journal of Multivariate Analysis , volume =. 2006 , doi =

  27. [27]

    and Hallin, Marc , title =

    Onatski, Alexei and Moreira, Marcelo J. and Hallin, Marc , title =. The Annals of Statistics , volume =. 2013 , doi =

  28. [28]

    Festschrift for Lucien Le Cam , publisher =

    Yu, Bin , title =. Festschrift for Lucien Le Cam , publisher =. 1997 , doi =

  29. [29]

    and Hoffman, Matthew D

    Margossian, Charles C. and Hoffman, Matthew D. and Sountsov, Pavel and Riou-Durand, Lionel and Vehtari, Aki and Gelman, Andrew , title =. Bayesian Analysis , volume =. 2025 , doi =

  30. [30]

    Journal of Machine Learning Research , volume =

    Zhang, Lu and Carpenter, Bob and Gelman, Andrew and Vehtari, Aki , title =. Journal of Machine Learning Research , volume =

  31. [31]

    Communications in Applied Mathematics and Computational Science , volume =

    Goodman, Jonathan and Weare, Jonathan , title =. Communications in Applied Mathematics and Computational Science , volume =. 2010 , doi =

  32. [32]

    and Dellaportas, Petros , title =

    Hirt, Marcel and Titsias, Michalis K. and Dellaportas, Petros , title =. Advances in Neural Information Processing Systems , volume =

  33. [33]

    Journal of Machine Learning Research , volume =

    Hird, Max and Livingstone, Samuel , title =. Journal of Machine Learning Research , volume =

  34. [34]

    Journal of the Royal Statistical Society: Series B , volume =

    Del Moral, Pierre and Doucet, Arnaud and Jasra, Ajay , title =. Journal of the Royal Statistical Society: Series B , volume =. 2006 , doi =

  35. [35]

    and Bartlett, Peter L

    Chatterji, Niladri S. and Bartlett, Peter L. and Long, Philip M. , title =. Bernoulli , volume =. 2022 , doi =

  36. [36]

    and Lu, Chen and Le Gouic, Thibaut and Rigollet, Philippe , title =

    Chewi, Sinho and Gerber, Patrik R. and Lu, Chen and Le Gouic, Thibaut and Rigollet, Philippe , title =. Proceedings of the 35th Conference on Learning Theory , series =. 2022 , publisher =

  37. [37]

    Proceedings of the 38th Conference on Learning Theory , series =

    He, Yuchen and Zhang, Chihao , title =. Proceedings of the 38th Conference on Learning Theory , series =. 2025 , publisher =. 2502.06200 , archivePrefix =

  38. [38]

    , title =

    Neal, Radford M. , title =. Statistics and Computing , volume =. 1996 , doi =