REVIEW 2 major objections 3 minor 15 references
Cumulants of multiinformation density in the case of a multivariate normal distribution
T0 review · 2 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read For Gaussian vectors, the multiinformation density's whole distribution is fixed by regression coefficients.
desk verdict Central theorem is correct and the regression-coefficient parameterization is a useful contribution, but Section 3.1's displayed formulas contain fixable typos that must be corrected before the paper can be cited as a reference. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the block matrix $\Gamma$ of regression coefficients between subvectors. The proof route is the identity $id = I(X_1;\dots;X_N) + \frac{1}{2}(X-\mu)^T\Phi(X-\mu)$ with $\Phi = \operatorname{diag}(\Sigma_{11},\dots,\Sigma_{NN})^{-1} - \Sigma^{-1}$, which rewrites multiinformation density as a constant plus a centered quadratic form in a normal vector. The moment-generating function of that quadratic form is a Gaussian integral whose value is $|I_d - t\Gamma|^{-1/2}$; expanding the log determinant as $\operatorname{tr}\ln(I_d - t\Gamma)$ and comparing powers of $t$ with the standard cumulant expansion yields the trace formula for $\kappa_l$.
What would settle it
Simulate a $d=4$ normal vector with homogeneous correlation $\rho=0.5$, partition it into four scalar components, compute $id$ on many samples, and compare the empirical second and fourth cumulants with the paper's formulas $\kappa_2 = \rho^2 d(d-1)/2$ and $\kappa_4 = 3\rho^4[d + (d-1)^4 - 1]$; a systematic mismatch beyond sampling error would refute Theorem 1.
Extended reading notes
Core claim
The paper's central claim, Theorem 1, is an exact cumulant-generating formula: for a multivariate normal $X$ partitioned into $N$ subvectors, $\ln \mathbb{E}(e^{t\,id}) = t I(X_1;\dots;X_N) - \frac{1}{2}\ln|I_d - t\Gamma|$, with $\Gamma = \Sigma\operatorname{diag}(\Sigma_{11},\dots,\Sigma_{NN})^{-1} - I_d$; the off-diagonal blocks of $\Gamma$ are the regression coefficient matrices $\Gamma_{m|n} = \Sigma_{mn}\Sigma_{nn}^{-1}$ for $m \ne n$. The cumulants are $\kappa_1(id) = I(X_1;\dots;X_N)$ and $\kappa_l(id) = \frac{(l-1)!}{2}\operatorname{tr}(\Gamma^l)$ for $l \ge 2$. This determines the entire law of $id$, including its variance, and it specializes cleanly: for two subvectors the odd cumulants vanish and the variance becomes the trace of the product of the two regression blocks, recovering the squared-canonical-correlation quantity; when every subvector is one-dimensional, the variance is the sum of squared correlation coefficients over all pairs.
Load-bearing premise
The load-bearing premise is that $\Sigma^{-1} - t\Phi$ is positive definite for $t$ in a neighborhood of $0$, so the Gaussian integral defining the moment-generating function converges; the proof justifies this by simultaneous diagonalization of the two positive definite matrices.
Editorial extensions
If this is right
- All cumulants and therefore the whole distribution of multiinformation density are determined by the regression-coefficient block matrix $\Gamma$; marginal variances are irrelevant to the law of $id$.
- The variance of multiinformation density is $\sum_{1\le m<n\le N} \operatorname{tr}(\Sigma_{mn}\Sigma_{mm}^{-1}\Sigma_{mn}\Sigma_{nn}^{-1})$, reducing in the all-scalar case to the sum of squared pairwise correlation coefficients.
- For a two-subvector partition, odd cumulants of $id$ are zero and even cumulants are traces of powers of the products $\Gamma_{1|2}\Gamma_{2|1}$ and $\Gamma_{2|1}\Gamma_{1|2}$, linking the distribution to canonical correlations.
- Multiinformation itself is also a function of $\Gamma$ alone, namely $I(X_1;\dots;X_N) = -\frac{1}{2}\ln|I_d + \Gamma|$.
- Multiinformation density is identically zero, and has zero variance, exactly when the subvectors are mutually independent, giving distribution-level markers of independence.
Reading between the lines
- Because the cumulant formula is a function of sample-estimable covariance blocks, a plug-in estimator of $\operatorname{Var}(id)$ follows immediately; an independence test based on that estimator could be compared against permutation tests to see whether it adds power beyond scalar multiinformation.
- The theorem suggests that $\Gamma$ is the natural dependence parameter for Gaussian partitions generally: any two covariance matrices sharing the same regression blocks produce identical multiinformation-density distributions, even if their marginal variances differ.
- The appendix's homogeneous-correlation example shows that the standardized multiinformation density is not asymptotically normal as dimension grows, so high-dimensional use of $id$ will need exact cumulant corrections or another limit theory rather than a central limit theorem.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper defines multiinformation density id(X;X_1,...,X_N) = ln(f(X)/∏ f_n(X_n)) for a multivariate normal vector partitioned into N subvectors. Its main result, Theorem 1, states that the cumulant-generating function of id is tI(X_1;...;X_N) - 1/2 ln|I_d - tΓ|, where Γ = Σ diag(Σ_11,...,Σ_NN)^{-1} - I_d has off-diagonal blocks Γ_{m|n} = Σ_mn Σ_nn^{-1}; consequently the first cumulant is the multiinformation and the higher cumulants are (l-1)!/2 tr(Γ^l). The proof uses the decomposition id = I + jd, Gaussian integration, and a log-determinant expansion. The paper then derives consequences for the N=2 case, the independence of the CGF from marginal variances, a graphical loop interpretation of tr(Γ^l), and a check of non-asymptotic normality for homogeneous correlation matrices.
Significance. If correct, Theorem 1 is a compact and useful characterization: the entire law of the multiinformation density for a Gaussian partition is determined by the regression-coefficient blocks, with no fitted parameters. The derivation is self-contained and the main proof steps check out. The graphical loop interpretation and the variance-independence result are valuable additions. However, the manuscript currently contains internal inconsistencies in displayed consequence formulas that must be fixed before final acceptance; these are typos rather than flaws in the central theorem.
major comments (2)
- [§3.1] The two displayed cumulant-generating functions in the scalar-X1 case and in the two-scalar case are missing a factor 1/2 in front of the logarithm of the second factor. Theorem 1 gives ln E(e^{t id}) = -t/2 ln(1-R^2) - 1/2 ln(1-t^2R^2) and ln E(e^{t id}) = -t/2 ln(1-ρ^2) - 1/2 ln(1-t^2ρ^2). As printed, the terms '-ln(1-t^2R^2)' and '-ln(1-t^2ρ^2)' would yield even cumulants equal to twice the values stated immediately below and also twice the values implied by the theorem. Please correct these displays and confirm that the stated cumulants follow from the corrected formulas.
- [§2.2, Eq. (10)] The general variance formula in Eq. (10) is not a consequence of the theorem. From κ_2(id) = 1/2 tr(Γ^2), the correct expression is Var(id) = ∑_{1≤m<n≤N} tr(Σ_mn Σ_nn^{-1} Σ_nm Σ_mm^{-1}). The printed expression ∑ tr(Σ_mn Σ_mm^{-1} Σ_mn Σ_nn^{-1}) is not equal to this in general and is not even well defined when the subvector dimensions differ. Equation (11) is the correct N=2 case. Please replace Eq. (10) and check the subsequent scalar-case variance statements.
minor comments (3)
- [Appendix B] The sentence 'since the two matrices ... are diagonalizable in the same basis' is imprecise. The simultaneous reduction being used is by congruence (F^T A F and F^T B F), not by a common similarity transformation; the displayed equations show the intended statement. Please rephrase so that the reference to Anderson's Theorem A.2.2 is accurate.
- [§4] There is a typo in the future-work paragraph: 'couldt contribute' should be 'could contribute'.
- [§3.3] The word 'particar' should be 'particular' in the sentence 'This graphical interpretation is in particlar compatible with a partitioning into two subvectors.'
Circularity Check
No significant circularity: Theorem 1 is derived from a Gaussian moment integral and a log-determinant expansion, with no fitted input, predictive claim, or load-bearing self-citation.
full rationale
The paper's central derivation is self-contained. Equation (2) follows from the Gaussian integral E(e^{t j_d}) = |I_d - t Gamma|^{-1/2}, which is obtained by completing the quadratic form in the normal density; no parameter is fitted and no benchmark or data subset is used. The cumulants then follow by the standard Taylor expansion ln|I_d - t Gamma| = -sum tr(Gamma^l) t^l / l and the identification of cumulant-generating function coefficients. The only external inputs are classical results: Gaussian integral identities, the log-determinant expansion, and Anderson's definition of regression coefficients. Appendix B invokes simultaneous diagonalization of two positive definite matrices, a standard theorem; even though the wording 'diagonalizable in the same basis' is technically more appropriate as congruence than similarity, this is a precision issue, not a circular one, and the paper does not assume the conclusion it is proving. The apparent inconsistencies in Section 3.1 (the missing factor 1/2 in the scalar CGF and the block-order error in Eq. (10)) are internal correctness concerns, not instances of the derivation reducing to its own inputs. There are no self-citations used as load-bearing support, and no 'prediction' is constructed from fitted values of the target quantity. The claimed result, that the distribution of multiinformation density for a Gaussian partition is determined by the regression-coefficient block matrix Gamma, is derived directly from the model assumptions rather than assumed or renamed from them. A non-finding is therefore appropriate: no step in the derivation chain is circular by construction.
Assumptions & free parameters
assumptions (4)
- domain assumption Multivariate normal distribution with positive definite covariance matrix Sigma.
- standard math Simultaneous diagonalization of two positive definite matrices by congruence.
- standard math Cumulant identities: first cumulant is shift-equivariant, higher cumulants are shift-invariant, and the cumulant-generating function is the log of the moment-generating function.
- standard math Taylor expansion of ln(I_d - tGamma) for small t and the identity ln|A| = tr(ln A) for positive definite A.
Cite this review
Pith. "Pith review of Cumulants of multiinformation density in the case of a multivariate normal distribution." pith.science (2026). https://pith.science/paper/PT65PSVN
@misc{pith2026190806934,
author = {Pith},
title = {Pith review of: Cumulants of multiinformation density in the case of a multivariate normal distribution},
year = {2026},
howpublished = {\url{https://pith.science/paper/PT65PSVN}},
note = {Machine review of arXiv:1908.06934}
}
abstract
We consider a generalization of information density to a partitioning into $N \geq 2$ subvectors. We calculate its cumulant-generating function and its cumulants, showing that these quantities are only a function of all the regression coefficients associated with the partitioning.
Figures
Reference graph
Works this paper leans on
-
[1]
Abramowitz, M., Stegun, I.A. (Eds.), 1972. Handbook of Mathematical Functions. Number 55 in Applied Math., National Bureau of Standards. 12
work page 1972
-
[2]
A general class of coefficients of divergence of one distribution from another
Ali, S., Silvey, S.D., 1966. A general class of coefficients of divergence of one distribution from another. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 28, 131–142
work page 1966
-
[3]
An Introduction to Multivariate Statistical Analysis
Anderson, T.W., 2003. An Introduction to Multivariate Statistical Analysis. Wiley Series in Probability and Mathematical Statistics. 3rd ed., John Wiley and Sons, New York. Csisz´ ar, I., 1963. Eine informationstheoretische Ungleichung und ihre Anwendung auf den Beweis der Ergodizitat von Markoffschen Ketten. A Magyar Tudom´ anyos Akad´ emia Matematikai ´ ...
work page 2003
-
[4]
Uncertainty and Structure as Psychological Concepts
Garner, W.R., 1962. Uncertainty and Structure as Psychological Concepts. John Wiley & Sons, New York
work page 1962
-
[5]
Functions of matrices, in: Hogben, L
Higham, N.J., 2007. Functions of matrices, in: Hogben, L. (Ed.), Handbook of Linear Algebra. Chapman & Hall/CRC Press, Boca Raton. Discrete Mathematics and its Applications. chapter 11
work page 2007
-
[6]
Relative entropy measures of multivariate dependence
Joe, H., 1989. Relative entropy measures of multivariate dependence. Jour- nal of the American Statistical Association 84, 157–164
work page 1989
-
[7]
A general correlation coefficient for direc- tional data and related regression problems
Jupp, P.E., Mardia, K.V., 1980. A general correlation coefficient for direc- tional data and related regression problems. Biometrika 67, 163–173
work page 1980
-
[8]
The Advanced Theory of Statistics
Kendall, M.G., 1945. The Advanced Theory of Statistics. volume 1. 2nd ed., Charles Griffin & Co. Ltd., London
work page 1945
Show all 15 references
-
[9]
Multivariate t Distributions and their Appli- cations
Kotz, S., Nadarajah, S., 2004. Multivariate t Distributions and their Appli- cations. Cambridge University Press, Cambridge, UK
2004
-
[10]
Information Theory and Statistics
Kullback, S., 1968. Information Theory and Statistics. Dover, Mineola, NY
1968
-
[11]
The exact moments of a ratio of quadratic forms
Magnus, J., 1986. The exact moments of a ratio of quadratic forms. Annales d’´ economie et de statistique 4, 95–109
1986
-
[12]
Directional Statistics
Mardia, K.V., Jupp, P.E., 2000. Directional Statistics. Wiley Series in Probability and Statistics, Wiley, Chichester
2000
-
[13]
Lecture notes on information theory
Polyanskiy, Y., Wu, Y., 2017. Lecture notes on information theory. http://www.stat.yale.edu/∼yw562/ln.html. 13 Studen´ y, M., Vejnarov´ a, J., 1998. The multiinformation function as a tool for measuring stochastic dependence, in: Jordan, M.I. (Ed.), Proceedings of the NATO Adv...
2017
-
[14]
On thef-divergence and singularity of probability measures
Vajda, I., 1972. On thef-divergence and singularity of probability measures. Periodica Mathematica Hungarica 2, 223–234
1972
-
[15]
Information theoretical analysis of multivariate corre- lation
Watanabe, S., 1960. Information theoretical analysis of multivariate corre- lation. IBM Journal of Research and Development 4, 66–82. 14
1960
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.