REVIEW 3 major objections 3 minor 42 references
A Diffusion Process Perspective on Posterior Contraction Rates for Parameters
T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper argues that posterior contraction rates for Bayesian parameters are governed by the stationary moments of a Langevin diffusion and, in weakly concave models, by a single fixed-point equation linking the population likelihood's…
desk verdict The diffusion framework is a fresh, promising direction, but Theorem 2's proof has a sign error in Lemma 1 and an unproved Lq convergence step; the promised BvM result also isn't there. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Langevin SDE whose stationary law is the posterior, combined with Itô-calculus moment control for $\mathbb{E}\|\theta_t-\theta^*\|_2^p$. For the weakly concave analysis, the paper defines auxiliary functions $\nu_p(r)=\psi(r^{1/(p-1)})r^{(p-2)/(p-1)}$ and $\tau_p$ by $\tau_p(r^{p-1}\zeta(r))=r^{p-2}\psi(r)$; Assumption W.3 makes these functions convex, so Jensen's inequality turns the moment recursion into the limiting equation (9). The convexity of $r\mapsto\psi(\xi(r))$, where $\xi$ inverts $r\mapsto r\zeta(r)$, is what guarantees uniqueness of the positive solution.
What would settle it
A concrete check would be to exhibit a Bayesian model satisfying Assumptions W.1, W.2, and W.4 whose local curvature and perturbation growth violate W.3—for instance local $\psi(r)=r^2$ with $\zeta(r)=r^2$—and then compute the posterior mass outside the radius predicted by equation (9). A direct numerical or exact calculation for moderate $n$ would show whether inequality (10) still holds; if it fails, the theorem's scope is refuted, and if it holds, W.3 is not necessary.
Extended reading notes
Core claim
The paper's central claim is that posterior contraction is a diffusion-moment phenomenon: because the posterior density is the stationary distribution of the SDE $d\theta_t=\frac12\nabla F_n(\theta_t)dt+\frac{1}{2n}\nabla\log\pi(\theta_t)dt+\frac{1}{\sqrt{n}}dB_t$, bounding $\mathbb{E}\|\theta_t-\theta^*\|_2^p$ along the path and passing to the limit $t\to\infty$ controls posterior mass near $\theta^*$. Under strong concavity of the population log-likelihood $F$, Theorem 1 gives a contraction radius of order $\sqrt{d/(n\mu)}$ plus a perturbation term $\varepsilon_2(n,\delta)/\mu$. Under weak concavity, Theorem 2 shows that the radius is the unique positive solution of the nonlinear equation $\psi(z)=\varepsilon(n,\delta)\zeta(z)z+(B+d\log(1/\delta))/n$, with $\psi$ encoding the weak-concavity geometry and $\zeta$ the stochastic perturbation growth; the result is applied to logistic regression, single index models with $g(r)=r^p$ and $\theta^*=0$, and over-specified location Gaussian mixtures.
Load-bearing premise
The weakly concave result depends on two technical inequalities relating the likelihood's curvature curve $\psi$ and the error-growth curve $\zeta$; they are verified for the paper's three examples but not derived from more basic assumptions. If a model does not satisfy them, the fixed-point radius is not guaranteed.
Editorial extensions
If this is right
- For strongly concave population log-likelihoods, the posterior contracts at the parametric $\sqrt{d/n}$ rate with explicit dependence on the prior concentration constant $B$ and the strong-concavity constant $\mu$, without requiring a bounded parameter space.
- For weakly concave models, the posterior radius is the unique solution of $\psi(z)=\varepsilon(n,\delta)\zeta(z)z+(B+d\log(1/\delta))/n$; for local power forms $\psi(r)=r^\alpha$ and $\zeta(r)=r^\beta$ this yields rates of order $(d/n)^{1/(2(\alpha-\beta-1))}$ or $(d/n)^{1/\alpha}$, whichever is larger.
- Bayesian logistic regression has posterior contraction rate $(d/n)^{1/2}$.
- Bayesian single index models with polynomial link $g(r)=r^p$, $p\ge2$, and true parameter $\theta^*=0$ have posterior contraction rate $(d/n)^{1/(2p)}$.
- Over-specified location Gaussian mixtures have posterior contraction rate of order $(d/n)^{1/4}$, including a novel $d^{1/4}$ dimension dependence and no boundedness assumption on the parameter space.
- Because the SDE moment bounds are non-asymptotic in time, the same technique also addresses approximate posterior distributions produced by Langevin sampling procedures.
Reading between the lines
- Beyond the paper: the same moment-control scheme should give finite-time contraction bounds for Langevin Monte Carlo samplers, since inequality (33) is non-asymptotic in $t$; the paper notes the link to approximate posteriors but does not develop sampler-specific mixing rates.
- Beyond the paper: equation (9) reads as an oracle-style trade-off between statistical bias, $\psi^{-1}((B+d\log(1/\delta))/n)$, and stochastic fluctuation, $\varepsilon\zeta(z)z$; one could try to match $\psi$ and $\zeta$ to a model's Fisher-information degeneracy to conjecture minimax lower bounds of the same shape.
- Beyond the paper: the differential inequalities in Assumption W.3 look like a shape condition on the pair $(\psi,\zeta)$: $\psi$ must be sufficiently more convex than $\zeta$ at every scale. Models with locally flat likelihoods paired with heavy-tailed gradient noise fall outside the theorem, even if the fixed-point equation still has a solution.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a diffusion-process (Langevin SDE) framework for posterior contraction rates of parameters. The posterior is represented as the stationary distribution of the SDE (5), and posterior concentration is bounded through moment control of the SDE combined with Fatou's lemma. Under strong concavity of the population log-likelihood (Assumptions A, B, S.1, S.2), Theorem 1 gives a contraction radius of order sqrt((d + log(1/delta) + B)/(n mu)) + eps_2/mu. Under weak concavity (Assumptions W.1-W.4), Theorem 2 characterizes the radius as the unique positive solution z* of the nonlinear equation psi(z) = eps(n,delta) zeta(z) z + (B + d log(1/delta))/n. Corollaries derive rates (d/n)^{1/2} for Bayesian logistic regression, (d/n)^{1/(2p)} for polynomial single-index models, and (d/n)^{1/4} for over-specified location Gaussian mixtures. Proofs use Ito calculus, Burkholder-Davis-Gundy inequalities, empirical-process concentration, and several auxiliary lemmas deferred to the appendices.
Significance. If the weak-concavity result can be repaired, the framework is a useful complement to density-based posterior contraction techniques: it avoids bounded parameter spaces and gives explicit dimension dependence. The unified fixed-point equation (9) and the new d^{1/4} scaling for over-specified Gaussian mixtures are valuable contributions, and the perturbation bounds in Appendix A are substantial. The paper also provides detailed proofs for the examples. However, because the proof of Theorem 2 currently contains a contradictory lemma and an unproved convergence step, the main weak-concavity result is not established as written.
major comments (3)
- [5.2 / Appendix B.1] Lemma 1 is internally inconsistent as stated. It assumes that phi is non-increasing with phi(c)=0 and phi(t) >= 0 for all t > c; non-increasing together with phi(c)=0 forces phi(t) <= 0 for t > c, so the only function satisfying all three conditions is phi identically zero on [c, infinity). The proof then sets delta = -sup_{s >= c+epsilon} phi(s) < 0, which requires phi(s) < 0 for s > c+epsilon, contradicting the stated sign condition. Moreover, the application in Theorem 2 defines phi(r) = -r + eps tau^{-1}(r) + ((B+(p-1)d)/n) nu^{-1}(r)^{(p-2)/(p-1)}, which is concave with phi(r) > 0 for r < r* and phi(r) < 0 for r > r*; this phi is neither non-increasing nor non-decreasing globally. The intended comparison argument is plausible, but the lemma and its application must be repaired before Theorem 2 is established.
- [5.2] The proof asserts that 'the process (theta_t : t >= 0) converges in Lq norm for arbitrarily large q' and uses this to conclude that lim_{t->infinity} R_p(t) exists. This assertion is unsupported. Proposition 1 guarantees only L2 convergence of the densities of theta_t to the posterior density; it does not imply convergence of the process in Lq norm, nor does it imply convergence of polynomial moments without additional uniform integrability arguments. The step lim_{t->infinity} R_p(t) exists and is bounded by r* is therefore not justified as written. This is load-bearing for Theorem 2.
- [Abstract / Section 6] The abstract promises 'non-asymptotic versions of a Bernstein-von-Mises guarantee for the posterior', but no such theorem appears anywhere in the body. The only related discussion is in Section 6, which states that under weak concavity the posterior cannot be approximated by a Gaussian. Either a non-asymptotic Bernstein-von-Mises result should be stated and proved (presumably for the locally strongly concave case), or the abstract's claim should be removed.
minor comments (3)
- [4.2] The displayed closed-form expression for the single-index population log-likelihood F_I(theta) has the wrong sign and an incorrect constant: since Y = epsilon with theta* = 0, one expects -1/2 - ((2p-1)!!/2) ||theta||^{2p} plus a constant, not '1 + ... / 2'. Only the gradient bound (24a) is used later, so this is a typo, but it should be corrected.
- [5.1] In the proof of Theorem 1 the definition 'alpha = 1/(2 mu) - eps_1(n,delta) > mu/6' is dimensionally and algebraically suspicious; the later combination of J1 and J5 requires alpha <= 2 mu/3 - eps_1(n,delta) rather than the displayed expression. Please check and correct the displayed definition of alpha.
- [4.1] In display (17b) the supremum is written over theta in R^d, but the surrounding text says 'for any r > 0'; the statement should clarify that the bound is uniform over balls B(theta*, r) or over all of R^d, since the two formulations are different.
Circularity Check
No significant circularity: the contraction rates are proved from stated structural assumptions via diffusion moment bounds, and the example rates are derived, not fitted.
full rationale
The paper's central derivation is self-contained in the relevant sense. Theorem 2 characterizes the contraction radius z*(n,δ) as the unique positive solution of equation (9), but that equation is not assumed as an ansatz: it is obtained from the moment bound (33), which is proved from Itô's formula together with Assumptions W.1–W.3 (weak concavity, perturbation control, and convexity of the auxiliary functions). The rates quoted in Corollaries 2–4 are then computed by verifying W.1–W.4 for each model from explicit population log-likelihood inequalities such as (17a), (24a), (29a) and from uniform gradient perturbation bounds such as (17b), (24b), (29b); none of these bounds is tuned to match the final posterior rate. The only citation with overlapping authorship is [9] (Dwivedi et al.), used in Appendix A.3 to obtain the model-specific bounds (29a) and (29b) for the Gaussian mixture example. That cited work is a separate study of EM convergence rates; its bounds are population/empirical-process facts that do not presuppose the posterior contraction rate proved here, so the citation is real evidence rather than a load-bearing circular premise. The manuscript does contain a separate rigor issue in the proof of Theorem 2 — the unproved assertion that (θ_t) converges in Lq for arbitrarily large q and an apparent sign inconsistency in Lemma 1 — but those are validity/correctness concerns, not circular reductions of the result to its own inputs. Accordingly, no circular step is identified and the circularity score is 0.
Assumptions & free parameters
assumptions (6)
- standard math The Langevin SDE defined in (5) has the posterior distribution as its unique stationary distribution and converges in L2.
- standard math Fatou's lemma inequality (6) allows posterior moment bounds to be obtained from limiting diffusion moments.
- domain assumption The empirical gradient perturbation bounds W.2 and S.2 hold uniformly on balls around θ*.
- domain assumption Assumptions A and B on the prior hold (Lipschitz log-density gradient and a finite B).
- domain assumption W.3 differential inequalities hold for the model's ψ and ζ.
- standard math The fixed-point equation (9) has a unique positive solution under W.4.
Cite this review
Pith. "Pith review of A Diffusion Process Perspective on Posterior Contraction Rates for Parameters." pith.science (2026). https://pith.science/paper/OW3O4T5Q
@misc{pith2026190900966,
author = {Pith},
title = {Pith review of: A Diffusion Process Perspective on Posterior Contraction Rates for Parameters},
year = {2026},
howpublished = {\url{https://pith.science/paper/OW3O4T5Q}},
note = {Machine review of arXiv:1909.00966}
}
read the original abstract
We analyze the posterior contraction rates of parameters in Bayesian models via the Langevin diffusion process, in particular by controlling moments of the stochastic process and taking limits. Analogous to the non-asymptotic analysis of statistical M-estimators and stochastic optimization algorithms, our contraction rates depend on the structure of the population log-likelihood function, and stochastic perturbation bounds between the population and sample log-likelihood functions. Convergence rates are determined by a non-linear equation that relates the population-level structure to stochastic perturbation terms, along with a term characterizing the diffusive behavior. Based on this technique, we also prove non-asymptotic versions of a Bernstein-von-Mises guarantee for the posterior. We illustrate this general theory by deriving posterior convergence rates for various concrete examples, as well as approximate posterior distributions computed using Langevin sampling procedures.
Reference graph
Works this paper leans on
- [1]
-
[2]
A. Bhattacharya, D. Pati, , and D. B. Dunson. Anisotropic function estimation using multi-bandwidth Gaussian processes. Annals of Statistics , 42:352–381, 2014. (Cited on page 2.)
work page 2014
-
[3]
D. Blei, A. Ng, and M. Jordan. Latent Dirichlet allocatio n. J. Mach. Learn. Res , 3:993– 1022, 2003. (Cited on pages 2 and 20.)
work page 2003
-
[4]
S. Boucheron, G. Lugosi, and P. Massart. Concentration Inequalities: A Nonasymptotic Theory of Independence . Oxford University Press, 2016. (Cited on pages 27 and 34.)
work page 2016
-
[5]
R. J. Carroll and P. Hall. Optimal rates of convergence fo r deconvolving a density. Journal of American Statistical Association , 83:1184–1186, 1988. (Cited on page 10.)
work page 1988
-
[6]
J. Chen. Optimal rate of convergence for finite mixture mo dels. Annals of Statistics , 23(1):221–233, 1995. (Cited on pages 3 and 14.)
work page 1995
-
[7]
R. de Jonge and J. H. van Zanten. Adaptive nonparametric B ayesian inference using location-scale mixture priors. Annals of Statistics , 38:3300–3320, 2010. (Cited on page 2.)
work page 2010
-
[8]
J. L. Doob. Application of the theory of martingales. Actes du ColloqueInternational Le Calcul des Probabilit´ es et ses applications (Lyon, 28 Juin 3 J uillet, 1948) , pages 23–27,
work page 1948
Show all 42 references
-
[9]
Dwivedi, N
R. Dwivedi, N. Ho, K. Khamaru, M. J. Wainwright, M. I. Jord an, and B. Yu. Singularity, misspecification, and the convergence rate of EM. arXiv preprint arXiv:1810.00828, 2018. (Cited on page 31.)
2018 arXiv
-
[10]
D. A. Freedman. On the asymptotic behavior of Bayes esti mates in the discrete case. Annals of Statistics , 34:13861403, 1963. (Cited on page 1.)
1963
-
[11]
D. A. Freedman. On the asymptotic behavior of Bayes esti mates in the discrete case.II. Annals of Statistics , 36:454456, 1965. (Cited on page 1.) 34
1965
-
[12]
Gao and H
C. Gao and H. H. Zhou. Rate exact Bayesian adaptation wit h modified block priors. Annals of Statistics , 44:318–345, 2016. (Cited on page 2.)
2016
-
[13]
Ghosal, J
S. Ghosal, J. K. Ghosh, and A. van der Vaart. Convergence rates of posterior distributions. Annals of Statistics , 28:500–531, 2000. (Cited on page 2.)
2000
-
[14]
Ghosal and A
S. Ghosal and A. van der Vaart. Entropies and rates of con vergence for maximum likelihood and bayes estimation for mixtures of normal dens ities. Annals of Statistics , 29:1233–1263, 2001. (Cited on pages 2 and 12.)
2001
-
[15]
Ghosal and A
S. Ghosal and A. van der Vaart. Posterior convergence ra tes of Dirichlet mixtures at smooth densities. Annals of Statistics , 35:697–723, 2007. (Cited on page 2.)
2007
-
[16]
Ho and X
N. Ho and X. Nguyen. Convergence rates of parameter esti mation for some weakly identifiable finite mixtures. Annals of Statistics , 44:2726–2755, 2016. (Cited on page 2.)
2016
-
[17]
Ishwaran, L
H. Ishwaran, L. F. James, and J. Sun. Bayesian model sele ction in finite mixtures by marginal density decompositions. Journal of the American Statistical Association , 96:1316–1332, 2001. (Cited on page 14.)
2001
-
[18]
B. J. K. Kleijn and A. W. van der Vaart. Misspecification i n infinite-dimensional Bayesian statistics. Annals of Statistics , 34:837–877, 2006. (Cited on page 2.)
2006
-
[19]
B. Lindsay. Mixture Models: Theory, Geometry and Applications . In NSF-CBMS Re- gional Conference Series in Probability and Statistics. IM S, Hayward, CA., 1995. (Cited on page 12.)
1995
-
[20]
X. Mao. Stochastic Differential Equations and Applications . Elsevier, 2007. (Cited on page 5.)
2007
-
[21]
McCullagh and J
P. McCullagh and J. A. Nelder. Generalized Linear Models . Chapman and Hall/CRC,
-
[22]
X. Nguyen. Convergence of latent mixing measures in fini te and infinite mixture models. Annals of Statistics , 4(1):370–400, 2013. (Cited on pages 2, 3, 13, and 14.)
2013
-
[23]
X. Nguyen. Borrowing strength in hierarchical Bayes: c onvergence of the Dirichlet base measure. Bernoulli, 22:1535–1571, 2016. (Cited on page 2.)
2016
-
[24]
Pritchard, M
J. Pritchard, M. Stephens, and P. Donnelly. Inference o f population structure using multilocus genotype data. Genetics, 155:945959, 2000. (Cited on page 2.)
2000
-
[25]
Revuz and M
D. Revuz and M. Yor. Continuous Martingales and Brownian Motion , volume 293. Springer-Verlag, third edition, 1999. (Cited on pages 5 and 14.)
1999
-
[26]
H. Risken. The Fokker-Planck Equation . Springer, 1996. (Cited on page 5.)
1996
-
[27]
Rousseau
J. Rousseau. Rates of convergence for the posterior dis tributions of mixtures of Betas and adaptive nonparametric estimation of the density. Annals of Statistics , 38:146–180,
-
[28]
Rousseau and K
J. Rousseau and K. Mengersen. Asymptotic behaviour of t he posterior distribution in overfitted mixture models. Journal of the Royal Statistical Society: Series B (Statist ical Methodology), 73:689–710, 2011. (Cited on pages 7 and 12.) 35
2011
-
[29]
Schwartz
L. Schwartz. On Bayes procedures. Zeitschrift f˝ ur Wahrscheinlichkeitstheorie und Ver- wandte Gebiete , 4:10–26, 1965. (Cited on page 1.)
1965
-
[30]
W. Shen, S. R. Tokdar, , and S. Ghosal. Adaptive Bayesian multivariate density estima- tion with Dirichlet mixtures. Biometrika, 100:623640, 2013. (Cited on page 2.)
2013
-
[31]
Shen and L
X. Shen and L. Wasserman. Rates of convergence of poster ior distributions. Annals of Statistics, 29:687–714, 2001. (Cited on page 2.)
2001
-
[32]
A. W. van der Vaart. Asymptotic Statistics . Cambridge University Press, 1998. (Cited on page 20.)
1998
-
[33]
A. W. van der Vaart and J. Wellner. Weak Convergence and Empirical Processes . Springer-Verlag, New York, NY, 1996. (Cited on page 23.)
1996
-
[34]
A. W. van der Vaart and J. A. Wellner. Weak Convergence and Empirical Processes: With Applications to Statistics . Springer-Verlag, New York, NY, 2000. (Cited on page 25.)
2000
-
[35]
M. J. Wainwright. High-Dimensional Statistics: A Non-Asymptotic Viewpoint . Cam- bridge University Press, 2019. (Cited on pages 23, 24, and 27.)
2019
-
[36]
S. Walker. On sufficient conditions for Bayesian consist ency. Annals of Statistics , 90:482– 488, 2003. (Cited on page 1.)
2003
-
[37]
S. Walker. New approaches to Bayesian consistency. Annals of Statistics , 32:2028–2043,
-
[38]
S. G. Walker, A. Lijoi, and I. Prunster. On rates of conve rgence for posterior distributions in infinite-dimensional models. Annals of Statistics , 35:738–746, 2007. (Cited on page 2.)
2007
-
[39]
Y. Wang. Convergence rates of latent topic models under relaxed identifiability conditions. Electronic Journal of Statistics , 13:37–66, 2019. (Cited on page 2.)
2019
-
[40]
Yang and D
Y. Yang and D. B. Dunson. Bayesian manifold regression. Annals of Statistics , 44:876– 905, 2016. (Cited on page 2.)
2016
-
[41]
Yang and S
Y. Yang and S. T. Tokdar. Minimax-optimal nonparametri c regression in high dimensions. Annals of Statistics , 43:652–674, 2015. (Cited on page 2.) 36
2015
-
[1989]
(Cited on pages 7 and 9.)
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.