Pith. sign in

REVIEW 2 major objections 3 minor 15 references

Cumulants of multiinformation density in the case of a multivariate normal distribution

T0 review · 2 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read For Gaussian vectors, the multiinformation density's whole distribution is fixed by regression coefficients.

desk verdict Central theorem is correct and the regression-coefficient parameterization is a useful contribution, but Section 3.1's displayed formulas contain fixable typos that must be corrected before the paper can be cited as a reference. read the letter →

arxiv 1908.06934 v2 pith:PT65PSVN submitted 2019-08-19 math.ST cs.ITmath.ITstat.OTstat.TH

classification math.STcs.ITmath.ITstat.OTstat.TH MSC 62H2062E1560E10
keywords multiinformationdensitycumulant-generatingfunctionmultivariatenormaldistributionregressioncoefficientsmutualinformationcanonicalcorrelations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Multiinformation density is the random variable $id = \ln[f(X)/\prod_n f_n(X_n)]$, whose expectation is multiinformation, a standard scalar measure of dependence. This paper proves that when $X$ is multivariate normal and is split into $N$ subvectors, the entire distribution of $id$ has a closed form: its cumulant-generating function is $t I(X_1;\dots;X_N) - \frac{1}{2}\ln|I_d - t\Gamma|$, where $\Gamma$ is built only from the regression-coefficient blocks $\Gamma_{m|n} = \Sigma_{mn}\Sigma_{nn}^{-1}$. It follows that every cumulant of order $l \ge 2$ equals $\frac{(l-1)!}{2}\operatorname{tr}(\Gamma^l)$, so the moments, variance, and full distribution of multiinformation density depend on the regression structure of the partition and not on marginal variances. If the theorem is right, multiinformation density becomes a tractable, distribution-level probe of Gaussian dependence rather than just a one-number summary.

What carries the argument

The carrying object is the block matrix $\Gamma$ of regression coefficients between subvectors. The proof route is the identity $id = I(X_1;\dots;X_N) + \frac{1}{2}(X-\mu)^T\Phi(X-\mu)$ with $\Phi = \operatorname{diag}(\Sigma_{11},\dots,\Sigma_{NN})^{-1} - \Sigma^{-1}$, which rewrites multiinformation density as a constant plus a centered quadratic form in a normal vector. The moment-generating function of that quadratic form is a Gaussian integral whose value is $|I_d - t\Gamma|^{-1/2}$; expanding the log determinant as $\operatorname{tr}\ln(I_d - t\Gamma)$ and comparing powers of $t$ with the standard cumulant expansion yields the trace formula for $\kappa_l$.

What would settle it

Simulate a $d=4$ normal vector with homogeneous correlation $\rho=0.5$, partition it into four scalar components, compute $id$ on many samples, and compare the empirical second and fourth cumulants with the paper's formulas $\kappa_2 = \rho^2 d(d-1)/2$ and $\kappa_4 = 3\rho^4[d + (d-1)^4 - 1]$; a systematic mismatch beyond sampling error would refute Theorem 1.

Watch

Extended reading notes

Core claim

The paper's central claim, Theorem 1, is an exact cumulant-generating formula: for a multivariate normal $X$ partitioned into $N$ subvectors, $\ln \mathbb{E}(e^{t\,id}) = t I(X_1;\dots;X_N) - \frac{1}{2}\ln|I_d - t\Gamma|$, with $\Gamma = \Sigma\operatorname{diag}(\Sigma_{11},\dots,\Sigma_{NN})^{-1} - I_d$; the off-diagonal blocks of $\Gamma$ are the regression coefficient matrices $\Gamma_{m|n} = \Sigma_{mn}\Sigma_{nn}^{-1}$ for $m \ne n$. The cumulants are $\kappa_1(id) = I(X_1;\dots;X_N)$ and $\kappa_l(id) = \frac{(l-1)!}{2}\operatorname{tr}(\Gamma^l)$ for $l \ge 2$. This determines the entire law of $id$, including its variance, and it specializes cleanly: for two subvectors the odd cumulants vanish and the variance becomes the trace of the product of the two regression blocks, recovering the squared-canonical-correlation quantity; when every subvector is one-dimensional, the variance is the sum of squared correlation coefficients over all pairs.

Load-bearing premise

The load-bearing premise is that $\Sigma^{-1} - t\Phi$ is positive definite for $t$ in a neighborhood of $0$, so the Gaussian integral defining the moment-generating function converges; the proof justifies this by simultaneous diagonalization of the two positive definite matrices.

Editorial extensions

If this is right

  • All cumulants and therefore the whole distribution of multiinformation density are determined by the regression-coefficient block matrix $\Gamma$; marginal variances are irrelevant to the law of $id$.
  • The variance of multiinformation density is $\sum_{1\le m<n\le N} \operatorname{tr}(\Sigma_{mn}\Sigma_{mm}^{-1}\Sigma_{mn}\Sigma_{nn}^{-1})$, reducing in the all-scalar case to the sum of squared pairwise correlation coefficients.
  • For a two-subvector partition, odd cumulants of $id$ are zero and even cumulants are traces of powers of the products $\Gamma_{1|2}\Gamma_{2|1}$ and $\Gamma_{2|1}\Gamma_{1|2}$, linking the distribution to canonical correlations.
  • Multiinformation itself is also a function of $\Gamma$ alone, namely $I(X_1;\dots;X_N) = -\frac{1}{2}\ln|I_d + \Gamma|$.
  • Multiinformation density is identically zero, and has zero variance, exactly when the subvectors are mutually independent, giving distribution-level markers of independence.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the cumulant formula is a function of sample-estimable covariance blocks, a plug-in estimator of $\operatorname{Var}(id)$ follows immediately; an independence test based on that estimator could be compared against permutation tests to see whether it adds power beyond scalar multiinformation.
  • The theorem suggests that $\Gamma$ is the natural dependence parameter for Gaussian partitions generally: any two covariance matrices sharing the same regression blocks produce identical multiinformation-density distributions, even if their marginal variances differ.
  • The appendix's homogeneous-correlation example shows that the standardized multiinformation density is not asymptotically normal as dimension grows, so high-dimensional use of $id$ will need exact cumulant corrections or another limit theory rather than a central limit theorem.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper defines multiinformation density id(X;X_1,...,X_N) = ln(f(X)/∏ f_n(X_n)) for a multivariate normal vector partitioned into N subvectors. Its main result, Theorem 1, states that the cumulant-generating function of id is tI(X_1;...;X_N) - 1/2 ln|I_d - tΓ|, where Γ = Σ diag(Σ_11,...,Σ_NN)^{-1} - I_d has off-diagonal blocks Γ_{m|n} = Σ_mn Σ_nn^{-1}; consequently the first cumulant is the multiinformation and the higher cumulants are (l-1)!/2 tr(Γ^l). The proof uses the decomposition id = I + jd, Gaussian integration, and a log-determinant expansion. The paper then derives consequences for the N=2 case, the independence of the CGF from marginal variances, a graphical loop interpretation of tr(Γ^l), and a check of non-asymptotic normality for homogeneous correlation matrices.

Significance. If correct, Theorem 1 is a compact and useful characterization: the entire law of the multiinformation density for a Gaussian partition is determined by the regression-coefficient blocks, with no fitted parameters. The derivation is self-contained and the main proof steps check out. The graphical loop interpretation and the variance-independence result are valuable additions. However, the manuscript currently contains internal inconsistencies in displayed consequence formulas that must be fixed before final acceptance; these are typos rather than flaws in the central theorem.

major comments (2)
  1. [§3.1] The two displayed cumulant-generating functions in the scalar-X1 case and in the two-scalar case are missing a factor 1/2 in front of the logarithm of the second factor. Theorem 1 gives ln E(e^{t id}) = -t/2 ln(1-R^2) - 1/2 ln(1-t^2R^2) and ln E(e^{t id}) = -t/2 ln(1-ρ^2) - 1/2 ln(1-t^2ρ^2). As printed, the terms '-ln(1-t^2R^2)' and '-ln(1-t^2ρ^2)' would yield even cumulants equal to twice the values stated immediately below and also twice the values implied by the theorem. Please correct these displays and confirm that the stated cumulants follow from the corrected formulas.
  2. [§2.2, Eq. (10)] The general variance formula in Eq. (10) is not a consequence of the theorem. From κ_2(id) = 1/2 tr(Γ^2), the correct expression is Var(id) = ∑_{1≤m<n≤N} tr(Σ_mn Σ_nn^{-1} Σ_nm Σ_mm^{-1}). The printed expression ∑ tr(Σ_mn Σ_mm^{-1} Σ_mn Σ_nn^{-1}) is not equal to this in general and is not even well defined when the subvector dimensions differ. Equation (11) is the correct N=2 case. Please replace Eq. (10) and check the subsequent scalar-case variance statements.
minor comments (3)
  1. [Appendix B] The sentence 'since the two matrices ... are diagonalizable in the same basis' is imprecise. The simultaneous reduction being used is by congruence (F^T A F and F^T B F), not by a common similarity transformation; the displayed equations show the intended statement. Please rephrase so that the reference to Anderson's Theorem A.2.2 is accurate.
  2. [§4] There is a typo in the future-work paragraph: 'couldt contribute' should be 'could contribute'.
  3. [§3.3] The word 'particar' should be 'particular' in the sentence 'This graphical interpretation is in particlar compatible with a partitioning into two subvectors.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Theorem 1 is derived from a Gaussian moment integral and a log-determinant expansion, with no fitted input, predictive claim, or load-bearing self-citation.

full rationale

The paper's central derivation is self-contained. Equation (2) follows from the Gaussian integral E(e^{t j_d}) = |I_d - t Gamma|^{-1/2}, which is obtained by completing the quadratic form in the normal density; no parameter is fitted and no benchmark or data subset is used. The cumulants then follow by the standard Taylor expansion ln|I_d - t Gamma| = -sum tr(Gamma^l) t^l / l and the identification of cumulant-generating function coefficients. The only external inputs are classical results: Gaussian integral identities, the log-determinant expansion, and Anderson's definition of regression coefficients. Appendix B invokes simultaneous diagonalization of two positive definite matrices, a standard theorem; even though the wording 'diagonalizable in the same basis' is technically more appropriate as congruence than similarity, this is a precision issue, not a circular one, and the paper does not assume the conclusion it is proving. The apparent inconsistencies in Section 3.1 (the missing factor 1/2 in the scalar CGF and the block-order error in Eq. (10)) are internal correctness concerns, not instances of the derivation reducing to its own inputs. There are no self-citations used as load-bearing support, and no 'prediction' is constructed from fitted values of the target quantity. The claimed result, that the distribution of multiinformation density for a Gaussian partition is determined by the regression-coefficient block matrix Gamma, is derived directly from the model assumptions rather than assumed or renamed from them. A non-finding is therefore appropriate: no step in the derivation chain is circular by construction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no free parameters and no new entities. It relies only on standard Gaussian integral identities, simultaneous diagonalization of positive definite matrices, and standard cumulant theory. The only quantity that might look like an invented object, the block regression-coefficient matrix Gamma, is directly defined from the covariance matrix and is the paper's contribution rather than a postulated entity.

assumptions (4)
  • domain assumption Multivariate normal distribution with positive definite covariance matrix Sigma.
    The theorem and all derivations are stated for X following N(mu, Sigma) with Sigma positive definite, ensuring all block marginal covariances Sigma_nn are invertible and regression coefficients are well defined.
  • standard math Simultaneous diagonalization of two positive definite matrices by congruence.
    Used in Appendix B to prove Sigma^{-1} - tPhi is positive definite in a neighborhood of t=0. The paper cites Anderson's Theorem A.2.2, though the wording in Appendix B says 'diagonalizable in the same basis'.
  • standard math Cumulant identities: first cumulant is shift-equivariant, higher cumulants are shift-invariant, and the cumulant-generating function is the log of the moment-generating function.
    Used in Section 2.2 to get cumulants of id from those of jd and to identify coefficients from the Taylor expansion.
  • standard math Taylor expansion of ln(I_d - tGamma) for small t and the identity ln|A| = tr(ln A) for positive definite A.
    Used to expand the log determinant and read off cumulants in Section 2.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cumulants of multiinformation density in the case of a multivariate normal distribution." pith.science (2026). https://pith.science/paper/PT65PSVN

@misc{pith2026190806934,
  author       = {Pith},
  title        = {Pith review of: Cumulants of multiinformation density in the case of a multivariate normal distribution},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PT65PSVN}},
  note         = {Machine review of arXiv:1908.06934}
}
abstract

We consider a generalization of information density to a partitioning into $N \geq 2$ subvectors. We calculate its cumulant-generating function and its cumulants, showing that these quantities are only a function of all the regression coefficients associated with the partitioning.

Figures

Figures reproduced from arXiv: 1908.06934 by the authors.

Figure 1
Figure 1. Graphical interpretation of tr(Γ l ). We consider the case N = 4 and l = 3. The directed 3-loop lp = (1 → 2 → 3 → 1) is represented with dark arrows. The value of τ on this loop is equal to τ (l) = tr(Γ1|3Γ3|2Γ2|1 ). tr(Γ 3 ) is obtained by summing τ (p) over all 3-loops. 4 Discussion In the present manuscript, we introduced multiinformation density, a ran￾dom variable that generalizes information density and whose … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

15 extracted references · 15 canonical work pages

  1. [1]

    (Eds.), 1972

    Abramowitz, M., Stegun, I.A. (Eds.), 1972. Handbook of Mathematical Functions. Number 55 in Applied Math., National Bureau of Standards. 12

  2. [2]

    A general class of coefficients of divergence of one distribution from another

    Ali, S., Silvey, S.D., 1966. A general class of coefficients of divergence of one distribution from another. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 28, 131–142

  3. [3]

    An Introduction to Multivariate Statistical Analysis

    Anderson, T.W., 2003. An Introduction to Multivariate Statistical Analysis. Wiley Series in Probability and Mathematical Statistics. 3rd ed., John Wiley and Sons, New York. Csisz´ ar, I., 1963. Eine informationstheoretische Ungleichung und ihre Anwendung auf den Beweis der Ergodizitat von Markoffschen Ketten. A Magyar Tudom´ anyos Akad´ emia Matematikai ´ ...

  4. [4]

    Uncertainty and Structure as Psychological Concepts

    Garner, W.R., 1962. Uncertainty and Structure as Psychological Concepts. John Wiley & Sons, New York

  5. [5]

    Functions of matrices, in: Hogben, L

    Higham, N.J., 2007. Functions of matrices, in: Hogben, L. (Ed.), Handbook of Linear Algebra. Chapman & Hall/CRC Press, Boca Raton. Discrete Mathematics and its Applications. chapter 11

  6. [6]

    Relative entropy measures of multivariate dependence

    Joe, H., 1989. Relative entropy measures of multivariate dependence. Jour- nal of the American Statistical Association 84, 157–164

  7. [7]

    A general correlation coefficient for direc- tional data and related regression problems

    Jupp, P.E., Mardia, K.V., 1980. A general correlation coefficient for direc- tional data and related regression problems. Biometrika 67, 163–173

  8. [8]

    The Advanced Theory of Statistics

    Kendall, M.G., 1945. The Advanced Theory of Statistics. volume 1. 2nd ed., Charles Griffin & Co. Ltd., London

Show all 15 references
  1. [9]

    Multivariate t Distributions and their Appli- cations

    Kotz, S., Nadarajah, S., 2004. Multivariate t Distributions and their Appli- cations. Cambridge University Press, Cambridge, UK

  2. [10]

    Information Theory and Statistics

    Kullback, S., 1968. Information Theory and Statistics. Dover, Mineola, NY

  3. [11]

    The exact moments of a ratio of quadratic forms

    Magnus, J., 1986. The exact moments of a ratio of quadratic forms. Annales d’´ economie et de statistique 4, 95–109

  4. [12]

    Directional Statistics

    Mardia, K.V., Jupp, P.E., 2000. Directional Statistics. Wiley Series in Probability and Statistics, Wiley, Chichester

  5. [13]

    Lecture notes on information theory

    Polyanskiy, Y., Wu, Y., 2017. Lecture notes on information theory. http://www.stat.yale.edu/∼yw562/ln.html. 13 Studen´ y, M., Vejnarov´ a, J., 1998. The multiinformation function as a tool for measuring stochastic dependence, in: Jordan, M.I. (Ed.), Proceedings of the NATO Adv...

  6. [14]

    On thef-divergence and singularity of probability measures

    Vajda, I., 1972. On thef-divergence and singularity of probability measures. Periodica Mathematica Hungarica 2, 223–234

  7. [15]

    Information theoretical analysis of multivariate corre- lation

    Watanabe, S., 1960. Information theoretical analysis of multivariate corre- lation. IBM Journal of Research and Development 4, 66–82. 14

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.