REVIEW 3 major objections 4 minor 1 cited by
On Accelerated Mixing of the No-U-turn Sampler
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The No-U-turn Sampler achieves the diffusive-to-ballistic speed-up of critical Randomized HMC inside an accelerated phase of two-scale Gaussian targets, but outside that phase certain fixed step sizes confine it to short orbits and…
desk verdict First rigorous NUTS mixing bounds in two-scale Gaussians, with a genuinely new concentration argument, but the phase-transition claims outrun the proof in finite dimension. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the concentration of the U-turn diagnostic $f(x,v,t_-,t_+) = \min\{v_+\cdot(x_+-x_-),\, v_-\cdot(x_+-x_-)\}$ around the deterministic uniform term $f_{unif}$, with fluctuations controlled by a Hanson-Wright/Bernstein inequality (Theorem 3). Under the separation condition (19), which keeps the set of possible physical-time orbit lengths $T = \{h(2^k-1)\}$ away from the $\delta$-uncertainty band $\{-\delta \le f_{unif} < \delta\}$ around the zeros of $f_{unif}$, Proposition 4 collapses NUTS's doubling-based orbit selection to a uniform draw from index orbits of a fixed deterministic length $t^*$, independent of position and velocity. This reduction turns NUTS into ordinary HMC with state-independent integration time distribution $\tau^*$, which is then fed through an accept/reject coupling framework (Theorem 11) with Wasserstein contraction and TV-to-Wasserstein regularization (Lemmas 14 and 15) to produce the mixing-time bound.
What would settle it
Simulate NUTS on a two-scale Gaussian with $d_1 = d_2 = 10^4$, $\kappa = 10$, and $d_2/d_1$ outside the accelerated phase $A$ (e.g., $d_2/d_1 = 10$), sweeping $h$ over a fine grid; measure the median selected orbit length $t^*$ and the empirical mixing time. The dichotomy predicts a sharp separation between step sizes with $t^* = \Theta(m_1^{-1/2})$ and mixing about $O(1)$ and step sizes with $t^* = \Theta(m_2^{-1/2})$ and mixing about $O(\kappa)$, with the boundary given by the sets $h = J/(2^k-1)$. If the distribution of orbit lengths does not concentrate near these two values, or if the bad step sizes vanish in measure as $d$ grows more slowly than predicted, the concentration-based reduction of Proposition 4 fails.
Extended reading notes
Core claim
The central discovery is a precise dichotomy for NUTS in two-scale Gaussian targets $\gamma_{2S} = (m_1^{-1/2}\#\gamma_{d_1}) \otimes (m_2^{-1/2}\#\gamma_{d_2})$, with condition number $\kappa = m_2/m_1$. The U-turn diagnostic concentrates around the deterministic uniform term $f^{2S}_{unif}(t) = \sin(m_1^{1/2}t)\, m_1^{-1/2} d_1 + \sin(m_2^{1/2}t)\, m_2^{-1/2} d_2$, so that, under a separation condition on the step size, the adaptive orbit selection reduces to a uniform draw at a fixed physical time $t^*$. If $(\kappa, d_2/d_1)$ lies in the accelerated phase $A = \{g_{\kappa,d_2/d_1}(t) \ge 0 \text{ for all } t \in (0,2\pi)\}$, where $g_{\kappa,d_2/d_1}(t) = \sin(\kappa^{-1/2}t) + \sin(t)\, \kappa^{-1/2} d_2/d_1$, then $t^* = \Omega(m_1^{-1/2})$ for every permitted step size and Theorem 10 gives total-variation mixing time $\widetilde{O}(1)$ transitions, the ballistic speed-up. Outside $A$, there exist two-scale Gaussians and fixed step sizes with $t^* = \Theta(m_2^{-1/2})$ and mixing time $\widetilde{O}(\kappa)$, the diffusive behavior. In the limit $d_1, d_2 \to \infty$ the dichotomy sharpens: inside $A$, long orbits occur for almost all step sizes; outside $A$, short orbits occur for a set of step sizes of positive measure.
Load-bearing premise
The whole reduction assumes the step size $h$ is such that the set of orbit lengths NUTS can check, $T = \{h(2^k-1)\}$, never falls inside the narrow uncertainty band around the zeros of the deterministic U-turn term; the paper proves this only as the dimensions tend to infinity, not for any fixed finite dimension.
Editorial extensions
If this is right
- Inside the accelerated phase $A$, NUTS mixes in $\widetilde{O}(1)$ transitions on two-scale Gaussians, meaning it attains the speed of critically tuned Randomized HMC without knowing $m_1$.
- Outside $A$, fixed step sizes can force NUTS into short orbits with mixing time $\widetilde{O}(\kappa)$, matching the known lower bounds for HMC with short integration times.
- Randomizing the step size (as the paper suggests via its Remark 9) should push all permitted $h$ into the long-orbit regime, a concrete remedy the theory points to.
- For NUTS with leapfrog flow and critical orbit length, the step-size allowance is $\widetilde{\Omega}(m_2^{-1/2} d^{-1/4})$ and the gradient cost is $\widetilde{O}(\kappa^{1/2} d^{1/4})$, consistent with optimal tuning results for HMC in Gaussians.
- The reduction to state-independent HMC provides a template for coupling-based mixing proofs of adaptive HMC methods beyond Gaussian targets.
Reading between the lines
- The separation condition (19) is the paper's weakest point: it guarantees that bad step sizes have measure zero only as $d_1,d_2 \to \infty$. We would predict that, for finite dimension, the set of bad $h$ has small but positive measure shrinking like a negative power of $d$, so the dichotomy still holds for generic $h$ — a testable numerical prediction.
- The phase boundary $g_{\kappa,d_2/d_1}=0$, with the large-$\kappa$ limit at $d_2/d_1 \approx 4.6$, suggests a universal criterion: acceleration survives exactly as long as the fast scale's oscillatory weight in the U-turn diagnostic can overcome the slow scale's negative lobe; this could be checked by measuring empirical NUTS orbit lengths across the $(\kappa, d_2/d_1)$ plane.
- The harmonic-chain example shows the concentration approach degrades when the uniform term is only logarithmic in dimension; we infer that for heavy-tailed or strongly anisotropic targets outside the Gaussian hierarchy, NUTS's self-tuning may not concentrate at all, and the accelerated/non-accelerated distinction is likely not sharp.
- The paper's step-size randomization suggestion implicitly defines a distribution over $\tau^*$; we conjecture that a fully randomized-$h$ NUTS would have a mixing-time bound of $\widetilde{O}(1)$ everywhere in the infinite-dimensional two-scale family, removing the anomalous phase outside $A$.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript analyzes the No-U-turn Sampler (NUTS) on Gaussian targets, combining a concentration-of-measure analysis of the U-turn diagnostic with a coupling-based mixing analysis. It proves a Bernstein-type concentration inequality for the U-turn diagnostic (Theorem 3), derives a deterministic simplification of NUTS's orbit selection under a separation assumption (Proposition 4, Lemmas 5 and 7), identifies an 'accelerated phase' for two-scale Gaussians (Proposition 8), and states total-variation mixing-time bounds for NUTS (Theorem 10). The central claim is that NUTS can achieve the diffusive-to-ballistic speed-up of critical randomized HMC in the accelerated phase, while outside this phase there exist step sizes for which mixing is only diffusive.
Significance. If the main result is correct, this would be the first rigorous evidence that NUTS can attain accelerated mixing, and the paper would make a substantial contribution to the theory of adaptive Hamiltonian Monte Carlo. The proof strategy is natural and the paper has real strengths: Theorem 3 and Lemmas 5, 7, 12--15 are presented with detailed proofs, and the reduction of NUTS to HMC with a state-independent integration-time distribution is an elegant and reusable idea. The analysis is a forward derivation with no fitted constants. However, the finite-dimensional status of the main dichotomy is not established, and the phase diagram as stated contains a demonstrable error. These issues are load-bearing for the paper's central claim, so the current version needs substantial revision.
major comments (3)
- [§3.4, Eq. (48), Figure 6] The definition of the accelerated phase A is incompatible with the caption claim that 'For 1 ≤ κ < 4, (κ, d2/d1) ∈ A for all d2/d1 > 0'. For every 1 < κ < 4 and every q > 0, g_{κ,q}(t) = sin(t/√κ) + q κ^{-1/2} sin(t) is negative at some t ∈ (0, 2π): choose t ∈ (π√κ, 2π), where both sine terms are negative (and for κ = 1 choose t ∈ (π, 2π)). For example, κ = 2 and q = 1 gives g(4.5) ≈ -0.73. Thus A ∩ {1 ≤ κ < 4} is empty, and the accelerated branch of the dichotomy in Proposition 8 and Theorem 10 is vacuous in precisely the regime highlighted by Figure 6. The small-angle approximation sin(κ^{-1/2}t) ≈ κ^{-1/2}t used in the caption is not valid for κ = O(1), and the phase boundary as stated is incorrect.
- [§3.4, Proposition 8 and discussion after Eq. (47)] The dichotomy asserted before Proposition 8 is not proved. The text claims that if f^{2S}_{unif} has no negative values on m_2^{-1/2}(0, 2π) then t* = Ω(m_1^{-1/2}), and otherwise there exist step sizes with t* = Θ(m_2^{-1/2}); neither implication is established in finite dimension. In particular, the first implication requires proving a lower bound on the first negative time of f^{2S}_{unif}, which is a separate statement from non-negativity on one fast period; for κ close to 1 the first negative time can occur inside m_2^{-1/2}(0, 2π) and still be of order m_1^{-1/2}. The existence claim for γ_{2S} ∉ A and h with t* = Θ(m_2^{-1/2}) is justified only by the asymptotic d1, d2 → ∞ discussion after the proposition, so the finite-dimensional dichotomy used by Theorem 10 rests on an unproven and potentially circular step.
- [§3.2–§3.4, Assumption (19)] The simplification to uniform orbit selection requires Assumption (19), T ∩ {-δ ≤ f_unif < δ} = ∅, but the paper gives no finite-dimensional condition on h that ensures this together with max T = Ω(m_1^{-1/2}). Lemma 7 gives δ = O(m_1^{-1/2} d_1^{1/2} + m_2^{-1/2} d_2^{1/2}) via Eq. (39) and (25), so the excluded band around the roots of f_unif has strictly positive width in finite dimension. The only justification that (19) is mild is the asymptotic statement as d1, d2 → ∞ with fixed ratios, where the excluded step sizes have Lebesgue measure zero. Consequently Proposition 4, Proposition 8, and Theorem 10 are conditional in finite dimensions in a way that is not quantified, and the claimed phase transition may fail for every admissible step size in the finite-dimensional targets to which the theorem is applied.
minor comments (4)
- [Definition 2] There are two typos in the definition of f: 'defiinition' and 'the definiition extends' should be 'definition' and 'the definition extends'.
- [§4.1 and Theorem 10] The symbol A is used for the accelerated phase in Eq. (48), for the accept event in Theorem 11 and Lemma 12, and for the 'domain of all model parameters' in Theorem 10. This notational collision is confusing and should be resolved.
- [Theorem 10, Eq. (51)] The step-size restriction h ≤ h̄ is stated with h̄ = eΩ(m_2^{-1/2} d^{-1/4} min(m_1^{1/2} t*, 1)^2), but t* depends on h through T = {h(2^k - 1)}. As written, the bound is self-referential; the statement should clarify that t* is evaluated at the h under consideration or provide a uniform bound independent of h.
- [Figure 6] The vertical axis is labeled d2/d1, while the text and the definition of A use both d2/d1 and the shorthand q; the figure caption should state the relationship explicitly, e.g., q = d2/d1.
Circularity Check
No circularity: the analysis is a forward mathematical derivation from concentration of the U-turn diagnostic to coupling-based mixing bounds, with no fitted quantity or self-cited result substituting for the main conclusion.
full rationale
The paper's derivation chain is self-contained and forward-directed: Theorem 3 establishes concentration of the U-turn diagnostic f via the Hanson-Wright inequality; Lemmas 5 and 7 turn this into explicit, dimension-dependent deviation bounds with stated constants; Proposition 4 derives the simplified uniform orbit selection from the concentration bound and the separation assumption (19); Proposition 8 characterizes the selected orbit length t* through the deterministic uniform term funif; Lemma 12 reduces NUTS to HMC with state-independent integration-time distribution tau* on the relevant event; and Lemmas 14-15, combined in Theorem 10, translate contraction and regularization properties of that HMC process into a mixing-time bound. No parameter is fitted to data, and no predicted quantity is defined in terms of the target mixing-time result: t* is computed from the covariance, step size, and kmax alone, and the final bound is a function of that computed t*. Assumption (19) is a genuine separation condition, and its possible restrictiveness is a correctness or robustness concern, not a circular one. The citations to the author's prior work, notably [12] for Gaussian concentration facts and [11,12] for the accept/reject coupling framework, are used as auxiliary tools with stated assumptions and proof sketches; they do not assert the paper's main dichotomy or mixing bound and therefore are not load-bearing in a circular sense. Thus the paper exhibits no significant circularity.
Assumptions & free parameters
free parameters (2)
- step size h
- max doubling count kmax
assumptions (6)
- standard math Hanson-Wright inequality for Gaussian quadratic forms
- standard math Gaussian spherical concentration (shells D and velocity sets E)
- ad hoc to paper Assumption (19): T avoids the delta-band around roots of funif
- domain assumption Leapfrog stability condition h^2 m2 <= 1
- domain assumption Maximal orbit length max T = Omega(m1^{-1/2})
- domain assumption Background hypocoercivity results (Poincare inequality and critical Randomized HMC relaxation time)
Cite this review
Pith. "Pith review of On Accelerated Mixing of the No-U-turn Sampler." pith.science (2026). https://pith.science/paper/VVTYBIDN
@misc{pith2026250713259,
author = {Pith},
title = {Pith review of: On Accelerated Mixing of the No-U-turn Sampler},
year = {2026},
howpublished = {\url{https://pith.science/paper/VVTYBIDN}},
note = {Machine review of arXiv:2507.13259}
}
read the original abstract
Recent progress on the theory of variational hypocoercivity established that Randomized Hamiltonian Monte Carlo -- at criticality -- can achieve pronounced acceleration in its convergence and hence sampling performance over diffusive dynamics. Manual critical tuning being unfeasible in practice has motivated automated algorithmic solutions, notably the No-U-turn Sampler. Beyond its empirical success, a rigorous study of this method's ability to achieve accelerated convergence has been missing. We initiate this investigation combining a concentration of measure approach to examine the automatic tuning mechanism with a coupling based mixing analysis for Hamiltonian Monte Carlo. In certain Gaussian target distributions, this yields a precise characterization of the sampler's behavior resulting, in particular, in rigorous mixing guarantees describing the algorithm's ability and limitations in achieving accelerated convergence.
Forward citations
Cited by 1 Pith paper
-
A Profile-Separation Framework for Quantitative Convergence of No-U-Turn Samplers
Under a new profile-separation condition, multinomial and biased-progressive NUTS mix in O~(1 + a*^2 kappa^2(1+gamma)^{4/3}) and O~(1 + a*^4 kappa^3(1+gamma)^2) transitions.
Reference graph
Works this paper leans on
-
[1]
D. Albritton, S. Armstrong, J.-C. Mourrat and M. Novack. ‘Variational methods for the kinetic Fokker–Planck equation’. In:Analysis & PDE 17.6 (2024), pp. 1953– 2010
work page 2024
-
[2]
C. Andrieu, N. De Freitas, A. Doucet and M. I. Jordan. ‘An introduction to MCMC for machine learning’. In: Machine learning 50.1-2 (2003), pp. 5–43
work page 2003
- [3]
- [4]
-
[5]
J. M. Bardsley. ‘MCMC-based image reconstruction with uncertainty quantifica- tion’. In: SIAM Journal on Scientific Computing 34.3 (2012), A1316–A1332
work page 2012
- [6]
-
[7]
N. Bou-Rabee, B. Carpenter, T. Kleppe and M. Marsden. Incorporating Local Step- Size Adaptivity into the No-U-Turn Sampler using Gibbs Self Tuning . 2024. arXiv: 2408.08259 [stat.ME]
arXiv 2024
-
[8]
N. Bou-Rabee, B. Carpenter, T. S. Kleppe and S. Liu. The Within-Orbit Adaptive Leapfrog No-U-Turn Sampler. 2025. arXiv: 2506.18746
arXiv 2025
Show all 35 references
-
[9]
Bou-Rabee, B
N. Bou-Rabee, B. Carpenter, S. Liu and S. Oberd¨ orster.The No-Underrun Sampler: A Locally-Adaptive, Gradient-Free MCMC Method. 2025. arXiv: 2501.18548
2025 arXiv
-
[10]
Bou-Rabee, B
N. Bou-Rabee, B. Carpenter and M. Marsden. GIST: Gibbs self-tuning for locally adaptive Hamiltonian Monte Carlo . 2024. arXiv: 2404.15253 [stat.CO]
2024 arXiv
-
[11]
Bou-Rabee and S
N. Bou-Rabee and S. Oberd¨ orster. ‘Mixing of Metropolis-adjusted Markov chains via couplings: The high acceptance regime’. In: Electronic Journal of Probability 29 (2024), pp. 1–27
2024
-
[12]
Bou-Rabee and S
N. Bou-Rabee and S. Oberd¨ orster. Mixing of the No-U-Turn Sampler and the Geometry of Gaussian Concentration . 2024. arXiv: 2410.06978
2024 arXiv
-
[13]
Bou-Rabee and J
N. Bou-Rabee and J. M. Sanz-Serna. ‘Randomized Hamiltonian Monte Carlo’. In: Ann. Appl. Probab. 27.4 (2017), pp. 2159–2194
2017
-
[14]
Y. Cao, J. Lu and L. Wang. ‘On explicit L 2-convergence rate estimate for un- derdamped Langevin dynamics’. In: Archive for Rational Mechanics and Analysis 247.5 (2023), p. 90
2023
-
[15]
Carpenter et al
B. Carpenter et al. ‘Stan: A probabilistic programming language’. In: Journal of Statistical Software 20 (2016), pp. 1–37. 38
2016
-
[16]
Chen and K
Y. Chen and K. Gatmiry. When does Metropolized Hamiltonian Monte Carlo provably outperform Metropolis-adjusted Langevin algorithm? 2023. arXiv: 2304. 04724
2023
-
[17]
de Valpine et al
P. de Valpine et al. ‘Programming with models: writing statistical algorithms for general model structures with NIMBLE’. In: Journal of Computational and Graph- ical Statistics 26 (2017), pp. 403–417
2017
-
[18]
Durmus, S
A. Durmus, S. Gruffaz, M. Kailas, E. Saksman and M. Vihola. ‘On the convergence of dynamic implementations of Hamiltonian Monte Carlo and no U-turn samplers’. In: arXiv preprint 2307.03460 (2023)
2023 arXiv
-
[19]
Eberle, A
A. Eberle, A. Guillin, L. Hahn, F. L¨ orler and M. Michel. Convergence of non- reversible Markov processes via lifting and flow Poincar´ e inequality . 2025. arXiv: 2503.04238
2025 arXiv
-
[20]
Eberle and F
A. Eberle and F. L¨ orler. ‘Non-reversible lifts of reversible diffusion processes and relaxation times’. In: Probab. Theory Relat. Fields (2024)
2024
-
[21]
Eberle and F
A. Eberle and F. L¨ orler. Space-time divergence lemmas and optimal non-reversible lifts of diffusions on Riemannian manifolds with boundary . 2024. arXiv: 2412 . 16710
2024
-
[22]
H. Ge, K. Xu and Z. Ghahramani. ‘Turing: A language for flexible probabilistic inference’. In: International Conference on Artificial Intelligence and Statistics, (AISTATS). 2018, pp. 1682–1690
2018
-
[23]
Gelman et al
A. Gelman et al. Bayesian data analysis . Chapman and Hall/CRC, 2013
2013
-
[24]
Goodman and J
J. Goodman and J. Weare. ‘Ensemble samplers with affine invariance’. In: Com- mun. Appl. Math. Comput. Sci. 5.1 (2010), pp. 65–80
2010
-
[25]
M. D. Hoffman and A. Gelman. ‘The no-U-turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo’. In: Journal of Machine Learning Research 15.1 (2014), pp. 1593–1623
2014
-
[26]
Kaipio and E
J. Kaipio and E. Somersalo. Statistical and Computational Inverse Problems. Vol. 160. Applied Mathematical Sciences. Springer Science & Business Media, 2005
2005
-
[27]
W. Krauth. Hamiltonian Monte Carlo vs. event-chain Monte Carlo: an appraisal of sampling strategies beyond the diffusive regime . 2024. arXiv: 2411.11690
2024 arXiv
-
[28]
Y. T. Lee, R. Shen and K. Tian. Lower Bounds on Metropolized Sampling Methods for Well-Conditioned Distributions . 2021. arXiv: 2106.05480
2021 arXiv
-
[29]
Leli` evre, M
T. Leli` evre, M. Rousset and G. Stoltz. Free Energy Computations: A Mathematical Perspective. 1st. Imperial College Press, 2010
2010
-
[30]
Lu and L
J. Lu and L. Wang. ‘On explicit L 2-convergence rate estimate for piecewise de- terministic Markov processes in MCMC algorithms’. In: Ann. Appl. Probab. 32.2 (2022), pp. 1333–1361
2022
-
[31]
G. A. Pavliotis. Stochastic processes and applications . Vol. 60. Texts in Applied Mathematics. Diffusion processes, the Fokker-Planck and Langevin equations. Springer, New York, 2014, pp. xiv+339. 39
2014
-
[32]
D. Phan, N. Pradhan and M. Jankowiak. ‘Composable effects for flexible and accelerated probabilistic programming in NumPyro’. In:arXiv preprint 1912.11554 (2019)
2019 arXiv
-
[33]
Salvatier, T
J. Salvatier, T. V. Wiecki and C. Fonnesbeck. ‘Probabilistic programming in Py- thon using PyMC3’. In: PeerJ Computer Science 2 (2016), e55
2016
-
[34]
A. M. Stuart. ‘Inverse problems: a Bayesian perspective’. In: Acta Numerica 19 (2010), pp. 451–559
2010
-
[35]
Vershynin
R. Vershynin. High-dimensional probability. Vol. 47. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, Cambridge, 2018. 40
2018
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.