Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

This paper proves that even and elliptical symmetry force any f-divergence variational objective to have a stationary point at the true mean of the target, and under extra conditions a unique minimizer that also recovers the correlation mat

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-04 00:23 UTC pith:5RNTD4T5

load-bearing objection Real partial-symmetry results and a clean stationary-point theorem for f-divergences, but the abstract oversells 'any stationary point' and 'no log-concavity' beyond what the theorems prove. the 3 major comments →

arxiv 2511.01064 v3 pith:5RNTD4T5 submitted 2025-11-02 stat.ML cs.LGstat.CO

Generalized Guarantees for Variational Inference in the Presence of Even and Elliptical Symmetry

classification stat.ML cs.LGstat.CO
keywords variational inferencef-divergenceeven symmetryelliptical symmetrylocation-scale familypartial symmetryhierarchical modelsmean and correlation recovery
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Variational inference replaces a target density p with the best approximation q inside a tractable family, and the usual wisdom is that the choice of divergence changes the answer. This paper establishes a symmetry-matching counterpoint: when p and q share even or elliptical symmetry, different f-divergences all agree on the target's mean, and under additional conditions on the correlation matrix. For any f-divergence—including reverse and forward KL, α-divergences, Hellinger distance, and total variation—a stationary point of the variational objective occurs at the true mean of an even-symmetric p (Theorem 10); uniqueness of that point, and recovery of the normalized covariance matrix, require φ=f∘exp to be convex and strictly decreasing and p to be somewhere-strictly log-concave (Theorem 11), conditions the paper's Table 1 shows are satisfied only by reverse KL. The framework also handles partial symmetry, where only a subset of coordinates is even or elliptically symmetric; this covers funnel-like hierarchical targets and yields guarantees for the recovered mean and correlations along those coordinates. The abstract's claim that any stationary point recovers the mean for all f-divergences is stronger than the theorems, which guarantee that the mean-matching solution is a stationary point and only identify it as the unique minimizer under the extra conditions; Appendix A.5 further concedes that the partial-correlation guarantee fails when hyperparameters enter the prior covariance.

Core claim

On the paper's own terms, the central discovery is that symmetry can be substituted for exactness: the variational family does not need to contain p, or match its tails or independence structure, for VI to recover p's mean and correlations. Theorem 10 shows that if p is even symmetric about μ and q varies over a location family with a symmetric base density, then ν=μ is a stationary point of D_f(p||q) for every differentiable f-divergence. Theorem 11 strengthens this under φ convex and strictly decreasing and p somewhere-strictly log-concave: the objective is strictly convex in location and scale and its unique minimizer sits at ν=μ, S=γ²M, recovering the normalized covariance (and hence cor

What carries the argument

The load-bearing machinery is the log-ratio form of f-divergences, D_φ(p||q)=∫ φ(log(p/q)) q dz with φ=f∘exp, evaluated over a location-scale family q_{ν,S}(z)=q₀(S^{-1/2}(z−ν))|S|^{-1/2} whose base density q₀ is spherically symmetric. Changing variables to ζ=S^{-1/2}(z−ν) makes the integrand depend on p through p(ν+S^{1/2}ζ); even symmetry forces the ζ-gradient of log p to be odd while every other factor is even, so the location-gradient integral vanishes at ν=μ. For elliptical symmetry, the same change of variables plus spherical symmetry of q₀ reduces the scale-gradient condition to a one-dimensional trace equation in γ, whose solution is guaranteed by the sign of g′ and the monotonicity

Load-bearing premise

The guarantees hold only if the target density truly has the stated symmetry about the exact center (possibly after conditioning), and the uniqueness and correlation results additionally require p to be somewhere-strictly log-concave and φ to be convex and strictly decreasing—conditions the paper notes only reverse KL among its tabulated divergences satisfies.

What would settle it

Run minimization of forward KL, Hellinger, and α-divergences from several random starts on an even-symmetric, non-log-concave target (e.g., a symmetric mixture of two well-separated Gaussians) with a Gaussian location family; finding a stationary point with location ν≠μ contradicts the abstract's any-stationary-point claim, while finding that all converged stationary points satisfy ν=μ supports the empirical stronger pattern. Separately, on an elliptically symmetric heavy-tailed target where Theorem 11's conditions fail, check whether scale S remains proportional to the true normalized covaria

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • For any differentiable f-divergence over a location family, an even-symmetric target's mean is guaranteed to be a stationary point of the optimization objective, so divergence choice alone cannot move the mean estimate away from the true value at a solution.
  • When φ is convex and strictly decreasing and p is somewhere-strictly log-concave, reverse-KL location-scale VI uniquely recovers both the mean and the normalized covariance (hence correlations) of an elliptically symmetric target, even if q is misspecified in other ways.
  • Partial-symmetry guarantees apply separately to blocks of coordinates, so in funnel-like hierarchical targets VI provably recovers the mean and correlations of the symmetric block while the other coordinates may remain poorly estimated.
  • For Gaussian q minimizing reverse KL, partial elliptical symmetry with constant center and normalized covariance yields S_σσ proportional to the marginal covariance of the symmetric block; the same reasoning gives conditional recovery when the center varies linearly in the other coordinates.
  • The experiments support a practical workflow: reparameterizing hierarchical models (e.g., non-centered or collapsed forms) lowers the target's measured asymmetry and improves the accuracy of VI's mean estimates along symmetric coordinates.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The fact that all five tabulated divergences empirically recover the mean for a student-t target, where the theorems only guarantee a stationary point, suggests the extra conditions (log-concavity, convex decreasing φ) are sufficient but not necessary; a direct test is whether the mean-matching point is the only stationary point for quasi-concave symmetric p under divergences with nonmonotone φ.
  • The paper's reflected-sample asymmetry score α could be repurposed as a cheap pre-VI diagnostic: estimate α on a small MCMC or prior draw set and use it to decide which coordinates' mean and correlation estimates are protected by symmetry before committing to a variational run.
  • Because the guarantees fail precisely when hyperparameters enter the prior covariance (Appendix A.5), one editorial lesson is that symmetry-based validation and parameterization choice are linked: non-centered or collapsed parameterizations do not just speed up sampling—they restore the conditions under which variational estimates are provably trustworthy.
  • The partial-symmetry proof for reverse KL relies on the chain-rule decomposition of the divergence; objectives that also decompose additively over conditional blocks (e.g., separable divergences or score-based divergences with a similar factorization) may admit analogous blockwise recovery guarantees, though the paper does not pursue this.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper extends recent symmetry-based guarantees for variational inference to f-divergences and to partially symmetric targets. For a target p with even symmetry and a location family Q, Theorem 10 shows that the f-divergence objective has a stationary point at the true mean, and, under additional conditions on φ=f∘exp and p, a unique minimizer. Theorem 11 gives an analogous result for elliptically symmetric targets, asserting a unique minimizer at the mean and a scale multiple of the normalized covariance matrix. Theorems 12 and 13 provide partial-symmetry analogues, with the strongest results restricted to reverse KL. The paper also includes numerical experiments on synthetic and hierarchical models.

Significance. The parity-based stationary-point argument is simple and appears correct, and the paper is admirably explicit in Section 2.1 that the convex/decreasing φ condition needed for uniqueness is verified only by KL(q||p) among the divergences in Table 1. The partial-symmetry results for hierarchical models are potentially useful, and the paper is candid that the strongest correlation guarantees are only a 'nominal generalization' of previous work. However, the abstract and introduction state a substantially stronger theorem than the body proves: the universal claim that 'any stationary point' recovers the mean and correlation matrix, with no log-concavity assumption, is not supported by Theorems 10, 11, and 13. In addition, some proofs in the appendix are sketched rather than fully demonstrated. The core ideas are worth publishing, but the presentation and proof details need revision.

major comments (3)
  1. [Abstract and §§3.1–3.2] The abstract's central claim — 'any stationary point of an f-divergence is guaranteed to recover the mean of p and likewise ... its correlation matrix' — is strictly stronger than what is proved. Theorem 10's first claim is existential: 'a stationary point ... occurs at ν=μ'. The universal statement is obtained only in the second part, under φ convex and strictly decreasing and p somewhere-strictly log concave. Theorem 11 likewise requires these conditions, and Section 2.1/Table 1 notes they are verified only by KL(q||p); Section 3.2 itself says the f-divergence result is a 'nominal generalization'. Please revise the abstract and introduction to state the existential stationary-point result and the exact conditions needed for uniqueness/global claims.
  2. [Theorem 12 and Appendix A.3] The second claim of Theorem 12 ('νσ=μσ at all stationary points') is proved by invoking Lemma 15, but Lemma 15 requires p to be log concave over all of R^d and somewhere-strictly log concave over R^d. Theorem 12 only assumes the conditional densities p(zσ|zbarσ) are somewhere-strictly log concave over R^{|σ|}. The joint density p(z)=p(zσ|zbar)p(zbar) need not be log-concave in zbar, and the strict-concavity sets may vary with zbar. The statement 'applying Lemma 15 ... we have that Dφ(p||qν) is strictly convex in νσ' needs a conditional version of Lemma 15 or additional uniformity assumptions across zbar.
  3. [Appendix A.2, proof of Theorem 11] The strict convexity of Dφ(p||q) in S^{1/2} is asserted in a single sentence: 'By assumption on p, p0 is log somewhere-strictly log concave in S^{1/2}.' This is not immediate: p0 is a density on R^d, and the composition A ↦ p0(Aζ) needs a separate argument, as does the interplay with the strictly concave log|A| term and integration over ζ. Since uniqueness in Theorem 11 depends on this step, please expand the proof or state and prove a precise convexity lemma for the matrix argument.
minor comments (4)
  1. [Appendix A.4, Lemma 16] The last display has missing parentheses and an inconsistent ordering of arguments: 'Cov(z¯σ, zσ)' vs. 'Cov(zσ, z¯σ)'. Please clean up the notation.
  2. [Theorem 11] The statement 'S=γ²M for some γ∈R' should specify γ>0 (or γ≥0), since S is positive definite; the proof later uses γ≥0.
  3. [Theorem 13 proof / Lemma 17] In eqs. (17)–(18) and Lemma 17, the proportionalities involving Cov_p[zσ|zbar] are only correct for the normalized covariance (i.e., correlation) matrix when the scale of the conditional covariance can depend on zbar. Please clarify this to avoid the appearance that the conditional covariance itself is assumed constant.
  4. [Abstract] The abstract says p and q are 'unimodal', but this term is not formally defined in the paper. The theorems use differentiability and, for stronger results, somewhere-strict log-concavity. Consider aligning the abstract's terminology with the body.

Circularity Check

1 steps flagged

No definitional circularity; uniqueness claims import a key inequality from the authors' own Proposition 11, and the abstract's 'any stationary point' exceeds the proved statement.

specific steps
  1. self citation load bearing [Appendix A.1, proof of Lemma 15]
    "We will now derive an inequality on gζ, using Proposition 11 of Margossian and Saul (2025). The latter only applies to functions with a univariate input and so we introduce hζ: R→R, a function whose domain R is the line that goes through ν0 and ν1."

    Lemma 15 is the step that turns the parity-based stationary point at ν=μ into a unique minimizer (Theorems 10 and 11). For condition (i), the strict inequality hζ(νλ) > (1−λ)hζ(ν0)+λhζ(ν1) is not proved here; it is imported from the same authors' prior paper. The uniqueness guarantee therefore rests on a self-citation rather than on a derivation contained in this paper. This is load-bearing but not definitionally circular: the first 'a stationary point at μ' claim is obtained independently by parity.

full rationale

The even-symmetry argument is self-contained: Eq. (11) shows ∇νDφ(p||qν) at ν=μ is an integral of an odd integrand, so μ is a stationary point for every f-divergence; no fitted quantity or prior theorem is needed for that claim. No prediction is a renamed fit, and the experimental sections are benchmarked against MCMC. The only load-bearing self-citation is in Lemma 15, where strict convexity — and hence uniqueness of the minimizer — is obtained via Proposition 11 of Margossian and Saul (2025), a same-author result not proved in this paper. This makes the uniqueness claims of Theorems 10/11 less self-contained, and the paper itself concedes (Section 3.2) that the φ conditions restrict the elliptical result to KL(q||p), providing only a 'nominal generalization' of the authors' prior work. That limitation is stated honestly and is a scope issue, not a circular reduction. Separately, the abstract's 'any stationary point of an f-divergence is guaranteed to recover the mean' is stronger than Theorem 10's proved 'a stationary point ... occurs at ν=μ'; this is an overstatement of the theorem's quantifier, not a circularity. Overall, there is no definitional or fitted-input circularity, so the score is moderate rather than high.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

No data-fitted constants enter the theory; the price is a stack of structural assumptions: spherical q0, log-concavity for uniqueness, constant conditional symmetry for partial results, and a self-cited strict-concavity lemma. The experiments use ad hoc settings (e.g., α=0.5 for the α-divergence, model hyperparameters), but those are illustrative.

axioms (6)
  • standard math Densities are sufficiently differentiable and differentiation can be interchanged with integration (dominated convergence).
    Needed to derive the gradient/stationarity equations in Theorems 10–13; stated in §2.4 but not verified case-by-case.
  • domain assumption The variational base density q0 is spherically symmetric about the origin.
    Definition 8/§2.3; ensures q inherits the needed even/elliptical symmetry about the location parameter.
  • domain assumption For uniqueness/correlation results, p or p(zσ|zbar) is somewhere-strictly log concave on R^d or R^{|σ|}.
    Definition 9; used in Lemma 15 and Theorems 10/11/13. This contradicts the abstract's claim that log-concavity is not required.
  • domain assumption Proposition 11 of Margossian & Saul (2025) is correct and applicable to the univariate strict-concavity argument.
    Invoked in Lemma 15 (Appendix A.1) to prove strict convexity; not reproduced in this paper and authored by the same researchers.
  • domain assumption For Theorem 13, p(zσ|zbar) has a constant point of even symmetry and a constant normalized covariance matrix independent of zbar; Q is Gaussian; φ is linear (reverse KL only).
    Stated in §3.3; fails for hierarchical models where hyperparameters enter the prior covariance (Appendix A.5).
  • domain assumption For Theorem 11, φ(v)=f(exp v) is convex and strictly decreasing, and g is continuously differentiable with |g'(0)|<∞.
    These conditions restrict the f-divergence generalization to KL(q||p), as the paper concedes in §3.2 and §5.

pith-pipeline@v1.3.0-alltime-deepseek · 21724 in / 14382 out tokens · 157529 ms · 2026-08-04T00:23:49.246013+00:00 · methodology

0 comments
read the original abstract

Variational inference (VI) approximates a target density $p$ by the best match $q$ in a family of tractable distributions. The best variational approximation is found by minimizing a divergence between distributions, $D(p||q)$, and several divergences have been proposed as objective functions for VI, with different choices leading to different approximations. We show that even when these divergences have different minimizers, the resulting approximations all abide by certain symmetry-matching principles. Specifically, our results hold for all $f$-divergences, a broad class which includes the reverse and forward Kullback-Leibler divergences and the $\alpha$-divergences. We show that in the presence of even symmetry, any stationary point of an $f$-divergence is guaranteed to recover the mean of $p$ and likewise, in the presence of elliptical symmetry, any stationary point is guaranteed to recover its correlation matrix. To obtain these guarantees we assume that $p$ and $q$ are unimodal, but notably we do not require them to be log-concave, light-tailed, or even everywhere-smooth. These guarantees generalize a previous result obtained for the reverse Kullback-Leibler divergence when $p$ is log-concave. They also extend to cases where the target density $p$ only exhibits symmetry along some but not all of its coordinates. These partial symmetries arise naturally in Bayesian hierarchical models, where the prior induces a challenging geometry but still possesses axes of symmetry.

Figures

Figures reproduced from arXiv: 2511.01064 by Charles C. Margossian, Isaac E. Rankin, Lawrence K. Saul.

Figure 1
Figure 1. Figure 1: VI Gaussian approximation of an elliptical funnel, obtained by minimizing KL(q||p). The funnel is asymmetric along τ but symmetric along θ and so VI provably recovers the mean and correlations of θ. diagonal. The geometry of the funnel is typical of hier￾archical priors and prone to frustrate many inference algorithms. But the funnel also has certain partial symmetries: the conditional distribution p(θ|τ )… view at source ↗
Figure 2
Figure 2. Figure 2: Variational approximation of a multivariate student-t by a Gaussian. Empirically, for each diver￾gence in [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: VI approximations to a skewed normal p with a Laplace distribution. (Left) When p has no skew (κ= 0), its mean is recovered by VI with all the divergences in [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Absolute error in VI’s mean estimate scaled by the target’s standard deviation. The targets are or￾dered, bottom to top, from most to least symmetric. The dotted line is the standard error obtained with 100 independent draws. As a trend, VI returns better es￾timates of the mean for more symmetric targets. The mean is also better estimated in the funnel, crescent, and disease along the coordinates σ whose p… view at source ↗
Figure 5
Figure 5. Figure 5: Error in VI estimates of the correlation. We split the targets into four groups: synthetic targets and implementations of schools, disease, and SKIM. Within each panel, the models are ordered bottom to top from most symmetric to least symmetric according to eq. (21). For the synthetic targets, we obtain better estimates of the correlation for more symmetric targets. There is no clear pattern for other targ… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Corrected Integrated Laplace Approximation for Bayesian Inference in Latent Gaussian Models

    stat.ML 2026-05 unverdicted novelty 5.0

    An importance sampling correction is added to integrated Laplace approximation so that the approximate posterior for latent Gaussian models converges to the true posterior as the number of samples grows.

Reference graph

Works this paper leans on

58 extracted references · 2 linked inside Pith · cited by 1 Pith paper

  1. [1]

    R., and Domke, J

    Agrawal, A., Sheldon, D. R., and Domke, J. (2020). Advances in black-box vi: Normalizing flows, importance weighting, and optimization. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020) , pages 2192--2205

  2. [2]

    H., Trippe, B., and Broderick, T

    Agrawal, R., Huggins, J. H., Trippe, B., and Broderick, T. (2019). The K ernel interaction trick: Fast B ayesian discovery of pairwise interactions in high dimensions. In Proceedings of the 36th International Conference on Machine Learning , pages 141--150

  3. [3]

    and Ridgway, J

    Alquier, P. and Ridgway, J. (2020). Concentration of tempered posteriors and of their variational approximations. The Annals of Statistics , 48(3):1475--1497

  4. [4]

    Betancourt, M. (2018). A conceptual introduction to Hamiltonian Monte Carlo . arXiv:1701.02434v1

  5. [5]

    and Girolami, M

    Betancourt, M. and Girolami, M. (2015). Hamiltonian monte carlo for hierarchical models. In Current Trends in Bayesian Methodology with Applications , page 24. Chapman and Hall/CRC

  6. [6]

    Billingsley, P. (1995). Probability and Measure . John Wiley & Sons, New York, 3 edition

  7. [7]

    and Mackey, L

    Biswas, N. and Mackey, L. (2023). Bounding W asserstein distance with couplings. Journal of the American Statistical Association , 119(548):2947--2958

  8. [8]

    M., Kucukelbir, A., and McAuliffe, J

    Blei, D. M., Kucukelbir, A., and McAuliffe, J. D. (2017). Variational inference: A review for statisticians. Journal of the American Statistical Association , 112:859--877

  9. [9]

    C., Gower, R

    Cai, D., Modi, C., Margossian, C. C., Gower, R. M., Blei, D. M., and Saul, L. K. (2024a). Eigen VI : score-based variational inference with orthogonal function expansions. In Advances in Neural Information Processing Systems 37 , pages 132691--132721

  10. [10]

    C., Gower, R

    Cai, D., Modi, C., Pillaud-Vivien, L., Margossian, C. C., Gower, R. M., Blei, D. M., and Saul, L. K. (2024b). Batch and match: black-box variational inference with a score-based divergence. Proceedings of the 41st International Conference on Machine Learning , page 5258–5297

  11. [11]

    A., Guo, J., Li, P., and Riddell, A

    Carpenter, B., Gelman, A., Hoffman, M., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M. A., Guo, J., Li, P., and Riddell, A. (2017). Stan: A probabilistic programming language. Journal of Statistical Software , 76:1--32

  12. [12]

    and Amari, S.-i

    Cichocki, A. and Amari, S.-i. (2010). Families of alpha- beta- and gamma-divergences: Flexible and robust measures of similarities. Entropy , 12:1532--1568

  13. [13]

    Daudel, K., Benton, J., and Doucet, A. (2023). Alpha-divergence variational inference meets importance weighted auto-encoders: Methodology and asymptotics. Journal of Machine Learning Research , 24(243):1--83

  14. [14]

    K., Catalina, A., Welandawe, M., Andersen, M

    Dhaka, A. K., Catalina, A., Welandawe, M., Andersen, M. R., Huggins, J., and Vehtari, A. (2021). Challenges and opportunities in high dimensional variational inference. In Advances in Neural Information Processing Systems 34 , pages 7787--7798

  15. [15]

    B., Tran, D., Ranganath, R., Paisly, J., and Blei, D

    Dieng, A. B., Tran, D., Ranganath, R., Paisly, J., and Blei, D. M. (2017). Variational inference via upper bound minimization. In Advances in Neural Information Processing Systems 30 , pages 2732--2741

  16. [16]

    B., Stern, H

    Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin, D. B. (2013). Bayesian Data Analysis . Chapman & Hall/CRC Texts in Statistical Science

  17. [17]

    Giordano, R., Broderick, T., and Jordan, M. I. (2018). Covariances, robustness, and variational B ayes. Journal of Machine Learning Research , 19(51):1--49

  18. [18]

    and Rue, H

    Gómez-Rubio, V. and Rue, H. (2018). Markov chain monte carlo with the integrated nested laplace approximation. Statistics and Computing , 28(5):1033--1051

  19. [19]

    Hoffman, M. D. and Gelman, A. (2014). The no- U -turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo . Journal of Machine Learning Research , 15:1593--1623

  20. [20]

    H., Kasprzak, M., Campbell, T., and Broderick, T

    Huggins, J. H., Kasprzak, M., Campbell, T., and Broderick, T. (2020). Validated variational inference via practical posterior error bounds. In Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics , pages 1792--1802

  21. [21]

    I., Ghahramani, Z., Jaakkola, T

    Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., and Saul, L. K. (1999). An introduction to variational methods for graphical models. Machine Learning , 37:183--233

  22. [22]

    and Rigollet, P

    Katsevich, A. and Rigollet, P. (2024). On the approximation accuracy of G aussian variational inference. The Annals of Statistics , 52(4):1384--1409

  23. [23]

    Kingma, D. P. and Welling, M. (2014). Auto-encoding variational B ayes. In Proceedings of the 2nd International Conference on Learning Representations

  24. [24]

    Kucukelbir, A., Tran, D., Ranganath, R., Gelman, A., and Blei, D. (2017). Automatic differentiation variational inference. Journal of Machine Learning Research , 18(14):1--45

  25. [25]

    and Turner, R

    Li, Y. and Turner, R. E. (2016). R\'enyi divergence variational inference. In Advances in Neural Information Processing Systems 29 , pages 1073--1081

  26. [26]

    MacKay, D. J. (2003). Information theory, inference, and learning algorithms . Cambridge University Press

  27. [27]

    C., Pillaud-Vivien, L., and Saul, L

    Margossian, C. C., Pillaud-Vivien, L., and Saul, L. K. (2025). Variational inference for uncertainty quantification: an analysis of trade-offs. Journal of Machine Learning Research . (To appear)

  28. [28]

    Margossian, C. C. and Saul, L. K. (2023). The shrinkage-delinkage trade-off: An analysis of factorized G aussian approximations for variational inference. In Proceedings of the 39th Conference on Uncertainty in Artificial Intelligence , pages 1358--1367

  29. [29]

    Margossian, C. C. and Saul, L. K. (2025). Variational inference in location-scale families: Exact recovery of the mean and correlation matrix. In Proceedings of the 28th Conference on Artificial Intelligence and Statistics , pages 3466--3474

  30. [30]

    C., Vehtari, A., Simpson, D., and Agrawal, R

    Margossian, C. C., Vehtari, A., Simpson, D., and Agrawal, R. (2020). Hamiltonian Monte Carlo using an adjoint-differentiated L aplace approximation: B ayesian inference for latent gaussian models and beyond. In Advances in Neural Information Processing Systems 33 , pages 9086--9097

  31. [31]

    A., Lindsten, F., and Blei, D

    Naesseth, C. A., Lindsten, F., and Blei, D. M. (2020). Markovian score climbing: Variational inference with KL (p||q) . In Advances in Neural Information Processing Systems 34 , pages 15499--15510

  32. [32]

    Neal, R. M. (2001). Annealed importance sampling. Statistics and Computing , 11:125--139

  33. [33]

    O., and Sk \"o ld, M

    Papaspiliopoulos, O., Roberts, G. O., and Sk \"o ld, M. (2007). A general framework for the parametrization of hierarchical models. Statistical Science , 22(1):59--73

  34. [34]

    and Vehtari, A

    Piironen, J. and Vehtari, A. (2017). Sparsity information and regularization in the horseshoe and other shrinkage priors. Electronic Journal of Statistics , 11:5018--5051

  35. [35]

    Ranganath, R., Gerrish, S., and Blei, D. (2014). Black box variational inference. In Proceedings of the Seventeenth International Conference on Artificial Intelligence and Statistics , pages 814--822

  36. [36]

    Rasmussen, C. E. and Williams, C. K. I. (2006). Gaussian Processes for Machine Learning . MIT Press, Cambridge, MA

  37. [37]

    R \'e nyi, A. (1961). On measures of entropy and information. In Le Cam, L. M. and Neyman, J., editors, Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Volume I: Contributions to the theory of statistics , pages 547--561. University of California Press, Berkeley, California

  38. [38]

    Robert, C. P. and Casella, G. (2004). Monte Carlo Statistical Methods . Springer

  39. [39]

    O., Gelman, A., and Gilks, W

    Roberts, G. O., Gelman, A., and Gilks, W. R. (1997). Weak convergence and optimal scaling of random walk metropolis algorithms. The Annals of Applied Probability , 7(1):110--120

  40. [40]

    Roberts, G. O. and Rosenthal, J. S. (2004). General state space markov chains and mcmc algorithms. Probability Surveys , 1:20--71

  41. [41]

    A., Ward, B., Carpenter, B., Syboldt, A., and Axen, S

    Roualdes, E. A., Ward, B., Carpenter, B., Syboldt, A., and Axen, S. D. (2023). Bridgestan: Efficient in-memory access to the methods of a S tan model. Journal of Open Source Software , 8

  42. [42]

    Rubin, D. B. (1981). Estimation in parallelized randomized experiments. Journal of Educational Statistics , 6:377--400

  43. [43]

    Rue, H., Martino, S., and Chopin, N. (2009). Approximate bayesian inference for latent gaussian models by using integrated nested laplace approximations. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , 71(2):319--392

  44. [44]

    Stan modeling language users guide and reference manual

    Stan Development Team (2025). Stan modeling language users guide and reference manual

  45. [45]

    Talts, S., Betancourt, M., Simpson, D., Vehtari, A., and Gelman, A. (2018). Validating B ayesian inference algorithms with simulation‐based calibration. arXiv:1804.06788

  46. [46]

    Tan, L. S. L. and Chen, A. (2024). Variational inference based on a subclass of closed skew normals. Journal of Computational and Graphical Statistics , pages 1--15

  47. [47]

    Turner, R. E. and Sahani, M. (2011). Two problems with variational expectation maximisation for time-series models. In Barber, D., Cemgil, A. T., and Chiappa, S., editors, Bayesian Time series models , chapter 5, pages 109--130. Cambridge University Press

  48. [48]

    van der Vaart, A. (1998). Asymptotic Statistics . Cambridge University Press

  49. [49]

    Vanhatalo, J., Pietil\"ainen, V., and Vehtari, A. (2010). Approximate inference for disease mapping with sparse G aussian processes. Statistics in Medicine , 29:1580--1607

  50. [50]

    P., Schiminovich, and Robert, C

    Vehtari, A., Gelman, A., Sivula, T., Jylanki, P., Tran, D., Sahai, S., Blomstedt, P., Cunningham, J. P., Schiminovich, and Robert, C. P. (2020). Expectation propagation as a way of life: A framework for B ayesian inference on partitioned data. Journal of Machine Learning Research , 21(17):1--53

  51. [51]

    Vehtari, A., Simpson, D., Gelman, A., Yao, Y., and Gabry, J. (2024). Pareto smoothed importance sampling. Journal of Machine Learning Research , 25(72):1--58

  52. [52]

    Wainwright, M. J. and Jordan, M. I. (2008). Graphical models, exponential families, and variational inference. Foundations and Trends in Machine Learning , 1(1--2):1--305

  53. [53]

    and Blei, D

    Wang, Y. and Blei, D. M. (2018). Frequentist consistency of variational B ayes. Journal of the American Statistical Association , 114:1147--1161

  54. [54]

    and Campbell, T

    Xu, Z. and Campbell, T. (2025). Asymptotically exact variational flows via involutive MCMC kernels. In Advances in Neural Information Process Systems 37 . (To appear)

  55. [55]

    Xu, Z., Chen, N., and Campbell, T. (2023). MixFlows: principled variational inference via mixed flows . In Proceedings of the 40th International Conference on Machine Learning , pages 38342--38376

  56. [56]

    I., and Wainwright, M

    Yang, E., Pati, D., Jordan, M. I., and Wainwright, M. J. (2020). -variational inference with statistical guarantees. The Annals of Statistics , 48(2):886--905

  57. [57]

    Yao, Y., Vehtari, A., Simpson, D., and Gelman, A. (2018). Yes, but did it work?: Evaluating variational inference. In Proceedings of the 35th International Conference on Machine Learning , pages 5577--5586

  58. [58]

    and Gao, C

    Zhang, F. and Gao, C. (2020). Convergence rates of variational posterior distributions. The Annals of Statistics , 48(4):2180--2207

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.