Pith. sign in

REVIEW 2 major objections 1 minor 12 references

Positive-definiteness in separable priors: effects on prior interpretability and inference

T0 review · 2 major / 1 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read Truncation to enforce positive definiteness in separable priors can systematically favor sparser matrices unless off-diagonal variances are adjusted with dimension.

desk verdict The truncation bias claim rests on comparing to an untruncated independent-entry distribution that isn't supported on positive definite matrices, so the sparsity effect may be an artifact rather than a real distortion. read the letter →

arxiv 2605.22640 v2 pith:EUXDLDC2 submitted 2026-05-21 stat.ME

classification stat.ME
keywords positive-definitepriorstruncationeffectssparseinferenceBayesiancovarianceestimationpriorinterpretabilityseparableshrinkage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper examines how adding a truncation step to independent-entry priors for positive-definite matrices alters their properties compared to the untruncated version. The truncation ensures the matrix is positive definite but can change the distribution's mass on different sparsity levels. For sparse settings, this leads to the truncated prior and posterior putting more weight on sparser structures than intended. The authors show how to choose prior variances so that these differences diminish as the matrix size increases. A sympathetic reader would care because many Bayesian models for covariance matrices rely on such priors, and unintended shifts in sparsity can affect model selection and inference without the user realizing.

What carries the argument

The truncation operation applied to an independent-entry distribution to restrict support to the set of positive-definite matrices.

What would settle it

Compute or estimate the prior probability of matrices with a given number of zero off-diagonal entries under both the truncated and untruncated versions for fixed variance parameters as dimension increases; persistent deviation from equality for sparser cases would support the claim.

Watch

Extended reading notes

Core claim

The paper claims that for priors on symmetric positive-definite matrices that start with independent entries and then truncate to the positive-definite cone, the resulting distribution differs from the untruncated one in ways that affect interpretability and shrinkage. Specifically, unless the variance of off-diagonal entries is set to decrease appropriately with matrix dimension, the truncated prior assigns higher probability to sparser matrices, and this bias carries over to the posterior.

Load-bearing premise

The untruncated independent-entry distribution is the desired target that truncation should preserve as closely as possible for interpretability and shrinkage characterization.

Editorial extensions

If this is right

  • Setting the variance of off-diagonal entries to scale with dimension mitigates the truncation effect for both dense and sparse matrices.
  • In sparse inference, careful parameter choice prevents the truncated prior and posterior from assigning systematically higher mass to sparser structures.
  • Posterior inference can be affected in unanticipated ways if truncation effects on mass assignment are ignored.
  • The shrinkage properties of the prior become harder to characterise without matching the untruncated margins.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Users of these priors in high-dimensional settings may need explicit scaling rules in software defaults to avoid unintended sparsity bias.
  • Similar truncation adjustments could be required for other constrained matrix distributions such as correlation matrices.
  • Direct Monte Carlo comparison of truncated and untruncated samples in moderate dimensions would quantify the mass shift on sparsity levels.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript examines the effects of adding a truncation to independent-entry priors on symmetric matrices to enforce positive-definiteness. It claims that, unless prior parameters (especially off-diagonal variances) are chosen carefully, the truncated prior and resulting posterior assign systematically higher mass to sparser structures than the untruncated counterpart, for both dense and sparse settings; the paper provides guidance on parameter settings to mitigate this discrepancy as matrix dimension grows.

Significance. If the claimed truncation effects and mitigation rules hold under rigorous derivation, the work would be significant for Bayesian covariance modeling and sparse precision-matrix inference, as it directly addresses interpretability and unintended shrinkage in a widely used prior class.

major comments (2)
  1. [Abstract] Abstract and introduction: the central claim rests on comparing the truncated prior to the untruncated independent-entry distribution as the reference whose sparsity properties should be preserved. However, the untruncated distribution is supported on all symmetric matrices and places positive mass outside the positive-definite cone; any observed difference in mass on sparse structures could therefore be an artifact of the projection onto the cone rather than an intrinsic effect of truncation. This comparison requires explicit justification or re-framing as a diagnostic rather than a normative target.
  2. [Abstract] The mitigation strategy of setting off-diagonal variances to control the effect as dimension grows inherits the same reference-distribution issue; without a clear statement of what properties of the untruncated margins are desirable on the PD cone, it is unclear whether the recommended parameter scaling achieves the intended preservation of interpretability.
minor comments (1)
  1. [Abstract] The abstract states that the analysis covers both dense and sparse matrices, but does not indicate whether the mitigation rules differ between the two regimes or whether the same variance scaling applies.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive comments. We address the major comments point by point below.

read point-by-point responses
  1. Referee: [Abstract] Abstract and introduction: the central claim rests on comparing the truncated prior to the untruncated independent-entry distribution as the reference whose sparsity properties should be preserved. However, the untruncated distribution is supported on all symmetric matrices and places positive mass outside the positive-definite cone; any observed difference in mass on sparse structures could therefore be an artifact of the projection onto the cone rather than an intrinsic effect of truncation. This comparison requires explicit justification or re-framing as a diagnostic rather than a normative target.

    Authors: The untruncated independent-entry prior serves as the natural baseline for separable priors, with truncation applied subsequently to enforce positive-definiteness. Our analysis demonstrates the distortion introduced by this truncation. We have revised the manuscript to explicitly frame the comparison as a diagnostic for assessing truncation effects on interpretability and sparsity, rather than positioning the untruncated distribution as a normative target on the positive-definite cone. This clarification addresses the concern directly. revision: yes

  2. Referee: [Abstract] The mitigation strategy of setting off-diagonal variances to control the effect as dimension grows inherits the same reference-distribution issue; without a clear statement of what properties of the untruncated margins are desirable on the PD cone, it is unclear whether the recommended parameter scaling achieves the intended preservation of interpretability.

    Authors: We agree that the target properties on the PD cone merit explicit statement. The recommended scaling of off-diagonal variances is designed to ensure that the truncated prior's marginal distributions and sparsity characteristics more closely match those of the untruncated prior as dimension increases. We have added clarification in the revised manuscript specifying the desirable properties (matching marginal variances and reduced bias in sparsity) and demonstrating how the scaling achieves this approximation within the positive-definite cone. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; analysis is self-contained

full rationale

The paper directly analyzes the truncation effect on independent-entry priors for positive-definite matrices by comparing truncated and untruncated versions, deriving guidance on parameter settings (e.g., off-diagonal variances) to mitigate sparsity bias as dimension grows. No equations or claims reduce by construction to fitted inputs renamed as predictions, self-definitional loops, or load-bearing self-citations. The untruncated distribution is treated as an explicit reference distribution whose properties are examined post-truncation; this is a modeling choice with independent mathematical content rather than a circular reduction. The derivation chain relies on standard truncation arguments and asymptotic analysis without importing uniqueness theorems or ansatzes from prior self-work.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract supplies no explicit free parameters, axioms, or invented entities beyond the general separable-prior construction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Positive-definiteness in separable priors: effects on prior interpretability and inference." pith.science (2026). https://pith.science/paper/EUXDLDC2

@misc{pith2026260522640,
  author       = {Pith},
  title        = {Pith review of: Positive-definiteness in separable priors: effects on prior interpretability and inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EUXDLDC2}},
  note         = {Machine review of arXiv:2605.22640}
}
read the original abstract

A popular class of priors for symmetric positive-definite matrices assumes independent entries and adds a truncation to ensure positive-definiteness. While conceptually simple and often computationally convenient, unless done carefully this truncation can have unintended effects. If the truncated prior or its margins are significantly different from their untruncated counterpart, then its interpretability may suffer, its shrinkage properties become harder to characterise, and posterior inference may be affected in unanticipated ways. We investigate the effect of the truncation both for dense and sparse matrices, and show how to set prior parameters such as the variance of off-diagonal entries such that said effect is mitigated as the matrix dimension grows. We pay particular attention to sparse inference where, unless prior parameters are set carefully, the truncated prior and hence its corresponding posterior assign systematically higher mass to sparser structures than the untruncated prior.

Figures

Figures reproduced from arXiv: 2605.22640 by the authors.

Figure 1
Figure 1. Monte Carlo estimate of c = P(Θ ≻ 0) for fixed unit diagonal and Gaussian off-diagonals for different δk = µ σ √ k − 2. Left: fixed δk. Right: δk varies with k δk = 0.1 δk = 0.05 δk = 0 δk = − 0.05 δk = − 0.1 0.00 0.25 0.50 0.75 1.00 0 50 100 150 200 k c δk = k−1 2 δk = k−2 3 δk = k−1 δk = k−2 δk = 0 0.92 0.96 1.00 0 50 100 150 200 k c Theorem 2. Let Θ be as in Theorem 1, but with independent stochastic diagonals θi… view at source ↗
Figure 2
Figure 2. Monte Carlo estimate of c = P(Θ ≻ 0) for θii ∼ Exp(1) (left) and θii ∼ Gamma(2, 2) (right) and Gaussian off-diagonals with different standard deviations σ σ = k−2 σ = k−3 2 σ = k−5 4 σ = k−9 8 σ = k−1 0.00 0.25 0.50 0.75 1.00 0 50 100 150 200 k c σ = k−5 4 σ = k−9 8 σ = k−1 σ = k−7 8 σ = k−3 4 0.00 0.25 0.50 0.75 1.00 0 50 100 150 200 k c 3 Sparse matrices An important class of priors are those that induce sparsity … view at source ↗
Figure 3
Figure 3. Monte Carlo estimate of c = P(Θ ≻ 0) in the sparse case for fixed θii = 1 and sparse Gaussian off-diagonals for different sparsity levels ηk = Pp(θij = 0) ηk = 0.5k −1/4 ηk = 0.5k −1/2 δk = 0.1 δk = 0.05 δk = 0 δk = − 0.05 δk = − 0.1 0.00 0.25 0.50 0.75 1.00 0 50 100 150 200 k c δk = 0.1 δk = 0.05 δk = − 0.1 0.00 0.25 0.50 0.75 1.00 0 50 100 150 200 k c 16 [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

12 extracted references · 12 canonical work pages

  1. [1]

    doi: 10.1016/j.jmva.2015.01.015

    ISSN 10957243. doi: 10.1016/j.jmva.2015.01.015. URL http://dx.doi.org/10.1016/j.jmva.2015.01.015. S. Boucheron, G. Lugosi, and P. Massart.Concentration inequalities: A nonasymptotic theory of indepen- dence. Oxford university press, Oxford,

  2. [2]

    doi: 10.1093/biostatistics/kxm045

    ISSN 14654644. doi: 10.1093/biostatistics/kxm045. Lingrui Gan, Naveen N Narisetty, and Feng Liang. Bayesian Regularization for Graphical Models With Unequal Shrinkage.Journal of the American Statistical Association, 114(527):1218–1231,

  3. [3]

    Fengshi Niu, Harsha Nori, Brian Quistorff, Rich Caruana, Donald Ngwe, and Aadharsh Kannan

    ISSN 1537274X. doi: 10.1080/01621459.2018.1482755. Jack Jewson, Li Li, Laura Battaglia, Stephen Hansen, David Rossell, and Piotr Zwiernik. Graphical model inference with external network data.Biometrics, 80(4):ujae151,

  4. [4]

    Locally associated graphical models and mixed convex exponential families.arXiv, 2008.04688:1–34,

    18 REFERENCES REFERENCES Steffen Lauritzen and Piotr Zwiernik. Locally associated graphical models and mixed convex exponential families.arXiv, 2008.04688:1–34,

  5. [5]

    Bayesian computation for high-dimensional gaussian graphical models with spike-and-slab priors.arXiv, 2511.01875:1–139,

    Deborah Sulem, Jack Jewson, and David Rossell. Bayesian computation for high-dimensional gaussian graphical models with spike-and-slab priors.arXiv, 2511.01875:1–139,

  6. [6]

    doi: 10.1214/12-BA729

    ISSN 19360975. doi: 10.1214/12-BA729. Hao Wang. Scaling it up: Stochastic search structure learning in graphical models.Bayesian Analysis, 10 (2):351–377,

  7. [7]

    doi: 10.1214/14-BA916

    ISSN 19316690. doi: 10.1214/14-BA916. 19 A SECTION 1 PROOFS A Section 1 proofs A.1 Proof of Proposition 1 LetSbe the set of symmetric matrices andS + the set of PD matrices. The TV distance is given by TV(p, p+) = sup A⊆S p(A)−p +(A) . The supremum is achieved by taking anyAsuch that{Θ :p(Θ)< p +(Θ)} ⊆A⊆ {Θ :p(Θ)≤p +(Θ)}, provided thatAis measurable. If Θ...

  8. [8]

    off-diagonals with densityπ

    B.3 Proof of Theorem 1 We decompose Θ as Θ =µI+σX k whereX k has zero diagonal and i.i.d. off-diagonals with densityπ. Standard Wigner matrix theory shows that Wk = Xk√ k = Θ−µI σ √ k has eigenvalues converging to the semicircle distribution and, in particular, has minimum eigenvalueλmin(Wk)→ −2 with probability 1 ask→ ∞(Bai and Yin, 1988). Sinceλ min(Θ) ...

Show all 12 references
  1. [9]

    It follows that we still have limk→∞ c= 1 if one sets anyσ= µ (2+δk) √ k such thatk −2/3 =o(δ k)

    This follows from the Wigner matrix theory in Lee and Yin (2014), which shows that deviations of the smallest eigenvalue of eΘ from−2 are of orderk −2/3 in probability. It follows that we still have limk→∞ c= 1 if one sets anyσ= µ (2+δk) √ k such thatk −2/3 =o(δ k). B.4 Proof ...

  2. [10]

    To ease notation, letl z(Θ) andl z′(Θ) be random variables with distribution equal to the conditional distributionλ min(Θ)|Z=zandλ min(Θ)|Z=z ′ respectively

    From this, any otherz ′ such thatz ′ ≥zentry- wise follows by induction. To ease notation, letl z(Θ) andl z′(Θ) be random variables with distribution equal to the conditional distributionλ min(Θ)|Z=zandλ min(Θ)|Z=z ′ respectively. The goal is to show thatE[l z(Θ)]≥E[l z′(Θ)]. ...

  3. [11]

    To apply Theorem 4, we need to find ν=∥E(W 2)∥= X i>j E [zijθij(Eij +E ji)]2 = X i>j zij(Eii +E jj)E(θ2 ij) . where we used thatz ijθij(Eij +E ji) are independent,z 2 ij =z ij, that simple algebra shows that (Eij +E ji)2 = (Eii +E jj), and that for any set independent and zero...

  4. [12]

    (2012), Theorem 2.7 shows that deviations of the maximum eigenvalue of a sparse Wigner matrix from 2 are of the orderk −2/3

    C.8 Proof of Corollaries 5-6 Erd˝ os et al. (2012), Theorem 2.7 shows that deviations of the maximum eigenvalue of a sparse Wigner matrix from 2 are of the orderk −2/3. In particular, the maximum eigenvalue converges to 2 ask→ ∞. Note the condition thatq > N 1/3 whereq= √kηk w...

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.