Pith. sign in

REVIEW 2 major objections 2 minor 1 cited by

Upper and lower bounds for local Lipschitz stability of Bayesian posteriors

T0 review · 2 major / 2 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Bayesian posterior sensitivity is governed by the evidence, and lower bounds show it must grow as the evidence shrinks.

desk verdict The lower-bound inequalities are real and mostly correct, but the advertised 'must increase' monotonicity claim is false as stated, and the paper needs major revision before the technical core is publishable. read the letter →

arxiv 2505.23541 v2 pith:YLHJEPYH submitted 2025-05-29 math.ST stat.TH

classification math.STstat.TH MSC 62F1560B1065J22
keywords BayesianinverseproblemsposteriorperturbationanalysisevidencetotalvariationmetricHellingerKullback-Leiblerdivergence1-WassersteinlocalLipschitzstability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Bayesian inverse problems are rarely stable in a scale-free way: the same perturbation of the misfit or prior can move the posterior a little or a lot. This paper locates the cause in the evidence, the normalizing constant $Z_{\Phi,\mu}=\int e^{-\Phi}d\mu$ in Bayes' formula, and proves lower as well as upper bounds on how far the posterior moves in total variation, Hellinger, Kullback-Leibler, and 1-Wasserstein distances. The lower bounds divide by the evidence, so they force the qualitative conclusion that once a posterior is more concentrated (smaller evidence), its sensitivity to a fixed perturbation must increase, not merely can increase. For the total variation and Hellinger metrics, and for the prior-to-posterior map in 1-Wasserstein, the bounds also show that fixing the evidence level makes the map locally bi-Lipschitz continuous, so the evidence is exactly what carries the instability.

What carries the argument

The load-bearing object is the evidence $Z_{\Phi,\mu}=\int_\Theta e^{-\Phi}\,d\mu$, the normalizing constant in Bayes' rule. The proof machinery is the two-sided local Lipschitz inequality for the exponential and logarithm, $[e^x\wedge e^y]|x-y|\le |e^x-e^y|\le [e^x\vee e^y]|x-y|$ (equivalently $|s-t|/\max(s,t)\le |\log s-\log t|\le |s-t|/\min(s,t)$), applied to the unnormalized likelihoods. This identity converts differences of posterior densities into differences of log-likelihoods multiplied by prefactors in which the evidence sits in the denominator and the bounded-misfit norm sits in the numerator; iterating it through the four metrics produces the lower bounds and, when evidence is conserved, the bi-Lipschitz constants.

What would settle it

Take $\mu$ uniform on $[0,1]$ and $\Phi_n=n\mathbf{1}_{[0,1-1/n]}$. Then $Z_{\Phi_n,\mu}\to0$, but $\exp(-\|\Phi_n\|_\infty)/Z_{\Phi_n,\mu}\to0$, so the prefactor in the paper's total-variation lower bound does not grow. Computing $d_{\mathrm{TV}}(\mu_{\Phi_n}, \mu_{\Phi_n+\mathbf{1}_{[1-1/n,1]}})$ for $n\to\infty$, which saturates at a positive constant rather than growing, would show that the 'must increase' statement fails when the prefactor condition is dropped even though each $\Phi_n$ is itself bounded.

Watch

Extended reading notes

Core claim

Starting from the posterior density $\ell_{\Phi,\mu}=e^{-\Phi}/Z_{\Phi,\mu}$, the paper treats the misfit-to-posterior map $\Phi\mapsto\mu_\Phi$ and the prior-to-posterior map $\mu\mapsto\mu_\Phi$ as objects whose local Lipschitz properties should be quantified with explicit constants. Its central discovery is a family of lower bounds whose prefactors contain inverse powers of the evidence, for example in total variation $d_{\mathrm{TV}}(\mu_{\Phi_1},\mu_{\Phi_2})\ge \frac12(\frac{e^{-\|\Phi_1\|_\infty}}{Z_{\Phi_1,\mu}}\wedge \frac{e^{-\|\Phi_2\|_\infty}}{Z_{\Phi_2,\mu}})\|\Phi_1-\Phi_2+\log(Z_{\Phi_1,\mu}/Z_{\Phi_2,\mu})\|_{L^1_\mu}$. Because $Z_{\Phi,\mu}$ appears in the denominator, the bound grows as the evidence decreases, which the authors read as proving that more concentrated posteriors must be more sensitive. With the evidence held fixed, the matching upper and lower bounds give local bi-Lipschitz continuity on evidence level sets, while without that constraint the evidence map is noninjective and the inversion problem is ill-posed.

Load-bearing premise

The lower bounds require each misfit to be essentially bounded ($\Phi\in L^\infty_\mu$), which excludes quadratic Gaussian misfits; the advertised 'must increase' conclusion also needs the prefactor $\exp(-\|\Phi\|_\infty)/Z$ not to vanish as $Z\to0$, a condition that fails for sequences such as $\Phi_n=n$ on $[0,1-1/n]$ and $0$ elsewhere.

Editorial extensions

If this is right

  • With bounded misfits, smaller evidence implies a larger guaranteed lower bound on posterior movement for the same perturbation in total variation and Hellinger distance.
  • On each evidence level set $\{\Phi: Z_{\Phi,\mu}=c\}$ the misfit-to-posterior map is locally bi-Lipschitz in total variation and Hellinger; the prior-to-posterior map is locally bi-Lipschitz on evidence level sets in total variation, Hellinger, and 1-Wasserstein under the stated support-boundedness conditions.
  • The matching upper and lower bounds show that the earlier local Lipschitz stability analysis is quantitatively sharp up to prefactors when evidence is conserved.
  • Global bi-Lipschitz stability on the full domain is impossible: the misfit-to-evidence and prior-to-evidence maps are noninjective, so distinct misfits or priors with the same evidence are not separated by the posterior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, these bounds suggest a practical rule of thumb for Bayesian computation: when the evidence is small, perturbation sensitivity is an expected feature rather than a bug, and error analyses should report the evidence value together with any stability constant.
  • If this monotone-sensitivity law is correct, then experimental design and model reduction criteria that conserve evidence would automatically preserve worst-case posterior stability; the paper does not test that design implication.
  • A natural extension is to replace the $\exp(-\|\Phi\|_\infty)$ prefactor by a localized quantity such as $\exp(-\operatorname{ess\,inf} \Phi)$ or $\exp(-\Phi(\theta_0))$ for a representative point; if such a bound holds for unbounded quadratic Gaussian misfits, the monotonicity law would extend to the most common Bayesian inverse problems.
  • The noninjectivity results implicitly warn against treating posterior-to-misfit inversion as a single-valued operation: recovering a misfit or prior from a posterior requires a convention for the evidence or a quotient by constant shifts.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper revisits the local Lipschitz stability bounds of Sprungk (2020) for Bayesian posteriors under misfit and prior perturbations, and proves upper and lower bounds with explicit dependence on the evidence in the total variation, Hellinger, Kullback–Leibler, and 1-Wasserstein metrics. The authors claim that their lower bounds show that posterior sensitivity to perturbations not only can but in general must increase as the evidence decreases to zero, and they use the evidence dependence to identify sufficient conditions for local bi-Lipschitz continuity on evidence level sets.

Significance. The individual bounds are derived carefully and several are genuinely new, particularly the lower bounds and the explicit evidence dependence. The sufficient conditions for local bi-Lipschitz continuity on evidence level sets (Sections 3.1, 4.1, and 6.2) are a useful contribution, and the paper provides a detailed, largely self-contained proof apparatus. However, the central advertised claim—that the lower bounds prove a general 'must increase' law of posterior sensitivity as evidence decreases—is false, and the counterexample below lies within the paper's own L∞ assumptions. This invalidates the main interpretive claim of the abstract, Section 1.1, and Section 8, even though the individual inequalities themselves appear mathematically correct.

major comments (2)
  1. [Abstract and §8, with Proposition 3.2(ii)] The asserted monotonicity law is not a consequence of the lower bounds and is false under the paper's assumptions. Every lower bound has a prefactor of the form exp(-||Φ||_∞)/Z times a likelihood distance, e.g. Proposition 3.2(ii). The claim that this grows as Z→0 is valid only if exp(-||Φ||_∞) does not vanish faster than Z, a condition that is neither stated nor satisfied in general. Let μ be uniform on [0,1] and Φ_n = n 1_{[0,1-1/n]}. Then ||Φ_n||_∞ = n and Z_{Φ_n,μ} = (1-1/n)e^{-n} + 1/n → 0, yet exp(-||Φ_n||_∞)/Z_{Φ_n,μ} = e^{-n}/Z_{Φ_n,μ} → 0. For the fixed bounded perturbation Δ(θ)=θ, both posteriors concentrate on [1-1/n,1] and the densities differ by O(1/n), so d_TV(μ_{Φ_n}, μ_{Φ_n+Δ}) → 0 as Z_{Φ_n,μ} → 0. This directly contradicts the abstract's claim that sensitivity 'in general will increase' as evidence decreases, and the same counterexample applies to the conclusion in Section 8.
  2. [Lemma 7.2] The proof of Lemma 7.2 is incomplete. After the display 'exp(-||Φ||_∞) ∧ min_k exp(-||Φ_{n_k}||_∞)/Z^2_{Φ,μ} ||Φ_{n_k} - Φ||^2_{L^2_μ} ≤ ||ℓ_{Φ_{n_k},μ} - ℓ_{Φ,μ}||^2_{L^2_μ}' the proof stops without deriving the claimed convergence of a subsequence of the Φ_n to Φ in L^2_μ. The lemma's second claim is therefore not established as written; the final step relating the displayed inequality to the conclusion is missing.
minor comments (2)
  1. [Proposition 6.4(i) and its proof] In the statement and proof, the displayed inequality contains the typo 'W1(µ2, µ2)'; it should be 'W1(µ1, µ2)'.
  2. [Remark 3.4] Remark 3.4 concedes that L^∞_μ misfits exclude common quadratic Gaussian misfits; this limitation is real, but it does not protect the central claim because the counterexample in my first major comment uses bounded misfits and hence falls squarely within the L∞ assumptions of Proposition 3.2(ii).

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; the bounds are derived from Bayes' rule and elementary inequalities, while the monotonicity claim is an overreach rather than a circular reduction.

full rationale

The derivation chain in this paper is self-contained: posterior densities are defined by Bayes' rule in (1.1), and the upper and lower bounds are obtained by applying the elementary local Lipschitz inequalities for exp and log in (1.5), triangle and reverse-triangle inequalities, Pinsker-type bounds, and Kantorovich-Rubinstein duality. For example, Proposition 3.2(ii) follows by writing the total variation distance as one half the L1 norm of the difference of posterior densities, factoring out the normalized likelihoods, and applying the lower bound in (1.5b); no fitted parameter, no normalization constant chosen to match the conclusion, and no external result carrying the conclusion is used. Citations to Sprungk [15] appear as comparisons or prior context, and the authors' own prior works [4, 11] are mentioned only in related literature, not as load-bearing premises. The advertised claim that sensitivity 'must increase' as evidence decreases is indeed stronger than the theorems warrant, because the prefactor exp(-||Phi||_infty)/Z appearing in the lower bounds can itself tend to zero as Z tends to zero; however, this is an overclaim or unstated-assumption problem, not a circularity in the sense of the target result being assumed or fitted. The incomplete proof of Lemma 7.2 is likewise a proof gap, not a circular step. No equation or argument reduces to its own input by construction, and no self-citation chain forces the conclusions.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper is a pure derivation from standard measure-theoretic probability and optimal transport. No parameters are fitted to data; the only external inputs are the chosen misfit, prior, and metric. The main nonstandard assumptions are the L^∞ boundedness of misfits needed for lower bounds and well-posedness conditions on the evidence.

assumptions (5)
  • standard math Bayes' rule defines the posterior via a normalized likelihood with finite positive evidence Z_{Φ,μ}.
    Used throughout, e.g., Eq. (1.6c) defines the posterior density as exp(-Φ)/Z times the prior.
  • domain assumption For lower bounds, the misfits are μ-essentially bounded (L^∞_μ).
    Required by Propositions 3.2(ii), 3.7(ii), 4.1(ii), 4.3(ii), 5.1(ii), and 6.1(ii); excludes common quadratic Gaussian misfits, as stated in Remark 3.4.
  • domain assumption Posterior measures are well-defined probability measures, meaning Z_{Φ,μ} lies in (0,∞) and the relevant integrability conditions hold.
    Assumed throughout, e.g., before Eq. (1.6c) and in the hypotheses of each proposition.
  • domain assumption For Wasserstein results, measures have finite p-th moments, and either the support is bounded or the relevant functions are Lipschitz.
    Needed for Kantorovich-Rubinstein duality and for attainment of the Wasserstein supremum, as in Corollary F.3 and Propositions 6.1, 6.4, and 6.5.
  • standard math Optimal transport duality theorem from Ambrosio et al.
    Theorem F.1 and Lemma F.2 underpin the 1-Wasserstein upper and lower bounds.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Upper and lower bounds for local Lipschitz stability of Bayesian posteriors." pith.science (2026). https://pith.science/paper/YLHJEPYH

@misc{pith2026250523541,
  author       = {Pith},
  title        = {Pith review of: Upper and lower bounds for local Lipschitz stability of Bayesian posteriors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YLHJEPYH}},
  note         = {Machine review of arXiv:2505.23541}
}
read the original abstract

The work of Sprungk (Inverse Problems, 2020) established the local Lipschitz continuity of the misfit-to-posterior and prior-to-posterior maps with respect to the Kullback--Leibler divergence and the total variation, Hellinger, and 1-Wasserstein metrics, by proving certain upper bounds. The upper bounds were also used to show that if a posterior measure is more concentrated, then it can be more sensitive to perturbations in the misfit or prior. We prove upper bounds and lower bounds that emphasise the importance of the evidence. The lower bounds show that the sensitivity of posteriors to perturbations in the misfit or the prior not only can increase, but in general will increase as the posterior measure becomes more concentrated, i.e. as the evidence decreases to zero. Using the explicit dependence of our bounds on the evidence, we identify sufficient conditions for the misfit-to-posterior and prior-to-posterior maps to be locally bi-Lipschitz continuous.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Convex Approximation Framework for Neural Likelihood-Based Bayesian Inverse Problems

    stat.ML 2026-07 conditional novelty 7.0 of 10

    By folding normalization into a KL-based objective over un-normalized potentials, neural likelihood approximation becomes a strictly convex problem with provable consistency.

Reference graph

Works this paper leans on

19 extracted references · 18 canonical work pages · cited by 1 Pith paper

  1. [1]

    Fabian Altekr¨ uger, Paul Hagemann, and Gabriele Steidl,Conditional generative models are provably robust: Pointwise guarantees for Bayesian inverse problems, Trans. Mach. Learn. Res. (2023), 1–20

  2. [2]

    130, Springer, 2021

    Luigi Ambrosio, Elia Bru´ e, and Daniele Semola,Lectures on optimal transport, Unitext, vol. 130, Springer, 2021

  3. [3]

    6, 1039 – 1051

    Lucien Birg´ e,Model selection for Gaussian regression with random design, Bernoulli10 (2004), no. 6, 1039 – 1051

  4. [4]

    Uncertain

    Nada Cvetkovi´ c, Han Cheng Lie, Harshit Bansal, and Karen Veroy,Choosing observation operators to mitigate model error in Bayesian inverse problems, SIAM/ASA J. Uncertain. Quantif.12(2024), no. 3, 723–758

  5. [5]

    12, 125008

    Duc-Lam Duong, Tapio Helin, and Jose Rodrigo Rojo-Garcia,Stability estimates for the ex- pected utility in Bayesian optimal experimental design, Inverse Problems39(2023), no. 12, 125008. 23

  6. [6]

    Alfredo Garbuno-Inigo, Tapio Helin, Franca Hoffmann, and Bamdad Hosseini,Bayesian posterior perturbation analysis with integral probability metrics, 2023, arXiv:2303.01512

  7. [7]

    44, Cambridge University Press, 2017

    Subhashis Ghosal and Aad van der Vaart,Fundamentals of Nonparametric Bayesian In- ference, vol. 44, Cambridge University Press, 2017

  8. [8]

    Evarist Gin´ e and Richard Nickl,Mathematical foundations of infinite-dimensional statisti- cal models, Camb. Ser. Stat. Probab. Math., vol. 40, Cambridge University Press, 2016

Show all 19 references
  1. [9]

    Michael Habeck, Daniel Rudolf, and Bj¨ orn Sprungk,Stability of doubly-intractable distri- butions, Electron. Commun. Probab.25(2020), 1 – 13

  2. [10]

    3, 831–865

    Jonas Latz,Bayesian inverse problems are usually well-posed, SIAM Rev.65(2023), no. 3, 831–865

  3. [11]

    Sullivan, and Aretha Teckentrup,Random forward models and log- likelihoods in Bayesian inverse problems, SIAM/ASA J

    Han Cheng Lie, T.J. Sullivan, and Aretha Teckentrup,Random forward models and log- likelihoods in Bayesian inverse problems, SIAM/ASA J. Uncertain. Quantif.6(2018), no. 4, 1600–1629

  4. [12]

    Paternain,Statistical guarantees for Bayesian uncertainty quantification in nonlinear inverse problems with Gaussian process priors, Ann

    Fran¸ cois Monard, Richard Nickl, and Gabriel P. Paternain,Statistical guarantees for Bayesian uncertainty quantification in nonlinear inverse problems with Gaussian process priors, Ann. Statist.49(2021), no. 6, 3255 – 3298

  5. [13]

    M¨ ucke, Benjamin Sanderse, Sander M

    Nikolaj T. M¨ ucke, Benjamin Sanderse, Sander M. Boht´ e, and Cornelis W. Oosterlee, Markov chain generative adversarial neural networks for solving Bayesian inverse prob- lems in physics applications, Comput. Math. Appl.147(2023), 278–299

  6. [14]

    Richard Nickl,Bayesian non-linear statistical inverse problems, Zurich Lectures in Ad- vanced Mathematics, European Mathematical Society Press, Berlin, Germany, 2023

  7. [15]

    Bj¨ orn Sprungk,On the local Lipschitz stability of Bayesian inverse problems, Inverse Prob- lems36(2020), 055015

  8. [16]

    Stuart,Inverse problems: A Bayesian perspective, Acta Numerica19(2010), 451–559

    Andrew M. Stuart,Inverse problems: A Bayesian perspective, Acta Numerica19(2010), 451–559

  9. [17]

    Stuart and Aretha L

    Andrew M. Stuart and Aretha L. Teckentrup,Posterior consistency for Gaussian process approximations of Bayesian posterior distributions, Math. Comput.87(2018), no. 310, 721–753

  10. [18]

    Wenpin Tang and Xun Yu Zhou,Regret of exploratory policy improvement andq-learning, 2024, arXiv:2411.01302

  11. [19]

    Old and new, Grundlehren Math

    C´ edric Villani,Optimal transport. Old and new, Grundlehren Math. Wiss., vol. 338, Springer, 2009. A. Auxiliary results Lemma A.1.(i) Letµ, ν∈ P(Θ)withµ≪ν. Then for everyf∈L 0(Θ,R),ess sup µ f≤ ess supν fandess inf µ f≥ess inf ν f. (ii) LetΦ,Ψ∈L 0(Θ,R). Theness sup µ(Φ+Ψ)≤ess...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.