{"id":"d7c00c92-5b25-43d6-be8c-bec63b3af1f5","arxiv_id":"1908.06934","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"For a multivariate normal vector partitioned into N blocks, the l-th cumulant of multiinformation density equals (l-1)!/2 times the trace of the l-th power of the block regression-coefficient matrix.","lead":"A new formula gives the full set of statistical cumulants of multiinformation density, a per-observation measure of deviation from independence in a multivariate normal distribution. The result expresses every cumulant in terms of the regression coefficients among the partition blocks, so the whole distribution of this quantity is captured by a single block matrix.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central theorem is sound, but the paper's own Section 3.1 formulas are inconsistent with it: the special-case CGF is missing a factor 1/2 and Eq. (10) has the wrong covariance-block order.","rationale":"The central theorem and its proof are sound. The MGF computation is a standard Gaussian integral, and the Appendix B positive-definiteness argument is valid once interpreted as simultaneous congruence: for two positive definite matrices there exists nonsingular F with F^T A F = I and F^T B F diagonal. The cumulant expansion then follows directly. The genuine issue is internal consistency of the consequence formulas: Section 3.1's scalar CGF is missing a factor 1/2 on the second log term, and Eq. (10)'s variance formula has the covariance blocks in an order that is dimensionally invalid and disagrees with the correct 1/2 tr(Γ^2) expansion. These errors do not undermine the proof of Theorem 1, but they do require correction. The reader's conditional verdict is therefore appropriate; I would neither reject nor upgrade to unconditional acceptance.","tokens_in":7838,"tokens_out":26585,"duration_ms":259085,"concrete_test":"Re-derive the Section 3.1 displays from Theorem 1 and Eq. (10) from kappa_2 = 1/2 tr(Γ^2). Specifically, take two blocks of sizes 1 and 2 with Σ12 = (1,2), Σ11 = 1, Σ22 = I_2, and Σ21 = Σ12^T; compute 1/2 tr(Γ^2) and compare with the printed Eq. (10). The printed term is dimensionally invalid, while 1/2 tr(Γ^2) gives a finite number, which settles that Eq. (10) is a typo. Independently expand the scalar bivariate CGF from Theorem 1 to order t^4 and compare with the displayed Section 3.1 CGF; if they differ, the display is the erroneous formula.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of Theorem 1 is internally correct once Appendix B is read as simultaneous congruence rather than similarity; the positive-definiteness claim follows from the standard simultaneous diagonalization of two positive definite matrices, so that is not the real weakness. The load-bearing problem is that the paper's consequence formulas contradict the theorem. In Section 3.1, for scalar X1 with multiple correlation R^2, Theorem 1 gives ln E(e^{t id}) = -t/2 ln(1-R^2) - 1/2 ln(1-t^2 R^2), hence even cumulants (l-1)!(R^2)^{l/2}; the displayed '- ln(1-t^2 R^2)' would double every even cumulant and is not equivalent. Likewise Eq. (10) writes Var(id) = sum_{m<n} tr(Σ_mn Σ_mm^{-1} Σ_mn Σ_nn^{-1}), which is not even dimensionally defined unless all blocks are square of the same size, and it does not equal 1/2 tr(Γ^2) = sum_{m<n} tr(Σ_mn Σ_nn^{-1} Σ_nm Σ_mm^{-1}); Eq. (11) is the correct N=2 case. These are typos rather than counterexamples to the central theorem, but a reader cannot tell from the manuscript which formula is authoritative, so correction is required before final acceptance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper defines multiinformation density id(X;X_1,...,X_N) = ln(f(X)/∏ f_n(X_n)) for a multivariate normal vector partitioned into N subvectors. Its main result, Theorem 1, states that the cumulant-generating function of id is tI(X_1;...;X_N) - 1/2 ln|I_d - tΓ|, where Γ = Σ diag(Σ_11,...,Σ_NN)^{-1} - I_d has off-diagonal blocks Γ_{m|n} = Σ_mn Σ_nn^{-1}; consequently the first cumulant is the multiinformation and the higher cumulants are (l-1)!/2 tr(Γ^l). The proof uses the decomposition id = I + jd, Gaussian integration, and a log-determinant expansion. The paper then derives consequences for the N=2 case, the independence of the CGF from marginal variances, a graphical loop interpretation of tr(Γ^l), and a check of non-asymptotic normality for homogeneous correlation matrices.","tokens_in":8145,"tokens_out":8712,"duration_ms":79403,"significance":"If correct, Theorem 1 is a compact and useful characterization: the entire law of the multiinformation density for a Gaussian partition is determined by the regression-coefficient blocks, with no fitted parameters. The derivation is self-contained and the main proof steps check out. The graphical loop interpretation and the variance-independence result are valuable additions. However, the manuscript currently contains internal inconsistencies in displayed consequence formulas that must be fixed before final acceptance; these are typos rather than flaws in the central theorem.","major_comments":[{"comment":"The two displayed cumulant-generating functions in the scalar-X1 case and in the two-scalar case are missing a factor 1/2 in front of the logarithm of the second factor. Theorem 1 gives ln E(e^{t id}) = -t/2 ln(1-R^2) - 1/2 ln(1-t^2R^2) and ln E(e^{t id}) = -t/2 ln(1-ρ^2) - 1/2 ln(1-t^2ρ^2). As printed, the terms '-ln(1-t^2R^2)' and '-ln(1-t^2ρ^2)' would yield even cumulants equal to twice the values stated immediately below and also twice the values implied by the theorem. Please correct these displays and confirm that the stated cumulants follow from the corrected formulas.","section":"§3.1"},{"comment":"The general variance formula in Eq. (10) is not a consequence of the theorem. From κ_2(id) = 1/2 tr(Γ^2), the correct expression is Var(id) = ∑_{1≤m<n≤N} tr(Σ_mn Σ_nn^{-1} Σ_nm Σ_mm^{-1}). The printed expression ∑ tr(Σ_mn Σ_mm^{-1} Σ_mn Σ_nn^{-1}) is not equal to this in general and is not even well defined when the subvector dimensions differ. Equation (11) is the correct N=2 case. Please replace Eq. (10) and check the subsequent scalar-case variance statements.","section":"§2.2, Eq. (10)"}],"minor_comments":[{"comment":"The sentence 'since the two matrices ... are diagonalizable in the same basis' is imprecise. The simultaneous reduction being used is by congruence (F^T A F and F^T B F), not by a common similarity transformation; the displayed equations show the intended statement. Please rephrase so that the reference to Anderson's Theorem A.2.2 is accurate.","section":"Appendix B"},{"comment":"There is a typo in the future-work paragraph: 'couldt contribute' should be 'could contribute'.","section":"§4"},{"comment":"The word 'particar' should be 'particular' in the sentence 'This graphical interpretation is in particlar compatible with a partitioning into two subvectors.'","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"The central theorem is correct and the paper is a compact contribution. The errors in Section 2.2 and Section 3.1 are self-contained typos rather than evidence of deeper trouble; I would be comfortable with acceptance after these corrections. The authors should also double-check the exact expansion in the scalar cases and ensure the displayed CGFs match Theorem 1."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know: the central theorem is correct, and the paper's own consequence formulas in Section 3.1 contradict it in a way that looks like typos rather than substantive errors.\n\nThe main result is a closed-form cumulant-generating function for multiinformation density under a multivariate normal partition into N>=2 subvectors. The parameterization by the block matrix Gamma of regression coefficients is a good organizing idea, and the directed-loop interpretation of tr(Gamma^l) is genuinely helpful. The proof is standard and clean: split id = I + jd, evaluate the Gaussian integral, expand the log-determinant. I checked the algebra; Theorem 1 and the cumulants (kappa1=I, kappa_l=(l-1)!/2 tr(Gamma^l) for l>=2) hold. As far as I can tell, the general N>=2 formula is new; the N=2 variance trace is correctly credited to Jupp and Mardia.\n\nThe soft spots are all in the presentation of consequences. Equation (10) as printed, Var(id)=sum_{m<n} tr(Sigma_mn Sigma_mm^{-1} Sigma_mn Sigma_nn^{-1}), has the blocks in the wrong order; it should be tr(Sigma_mn Sigma_nn^{-1} Sigma_nm Sigma_mm^{-1}). As written it is not even conformable when block sizes differ. The N=2 scalar-X1 CGF in Section 3.1 reads -t/2 ln(1-R^2) - ln(1-t^2 R^2), but Theorem 1 gives -1/2 in the second term. The cumulants stated right below match the 1/2 version, so the displayed CGF is simply missing a factor. The same issue appears in the both-scalar case. These are typos, not counterexamples, but they sit in the \"consequences\" section so a reader cannot tell which formula is authoritative until the theorem is re-derived. They must be fixed before acceptance.\n\nAppendix B's simultaneous diagonalization sentence is imprecise-the matrices are congruent, not similar, to their diagonal forms-but the cited result gives the needed positive definiteness, so that is minor.\n\nOverall: a solid theoretical note with a correct main result and a couple of embarrassing typos in the corollaries. It deserves a serious referee, and a conditional accept with minor revision is the right outcome. For a reader working on Gaussian dependence measures, this is worth citing once the typos are cleaned up.","headline":"Central theorem is correct and the regression-coefficient parameterization is a useful contribution, but Section 3.1's displayed formulas contain fixable typos that must be corrected before the paper can be cited as a reference.","tokens_in":8625,"tokens_out":8838,"would_cite":false,"duration_ms":75171,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H20","62E15","60E10"],"pacs":[],"model":"deepseek-v4-flash","headline":"For Gaussian vectors, the multiinformation density's whole distribution is fixed by regression coefficients.","keywords":["multiinformation density","cumulant-generating function","multivariate normal distribution","regression coefficients","mutual information","multiinformation","information density","canonical correlations"],"falsifier":"Simulate a $d=4$ normal vector with homogeneous correlation $\\rho=0.5$, partition it into four scalar components, compute $id$ on many samples, and compare the empirical second and fourth cumulants with the paper's formulas $\\kappa_2 = \\rho^2 d(d-1)/2$ and $\\kappa_4 = 3\\rho^4[d + (d-1)^4 - 1]$; a systematic mismatch beyond sampling error would refute Theorem 1.","tokens_in":7670,"feed_emoji":"📊","tokens_out":10306,"duration_ms":84101,"temperature":0.7,"pith_summary":"Multiinformation density is the random variable $id = \\ln[f(X)/\\prod_n f_n(X_n)]$, whose expectation is multiinformation, a standard scalar measure of dependence. This paper proves that when $X$ is multivariate normal and is split into $N$ subvectors, the entire distribution of $id$ has a closed form: its cumulant-generating function is $t I(X_1;\\dots;X_N) - \\frac{1}{2}\\ln|I_d - t\\Gamma|$, where $\\Gamma$ is built only from the regression-coefficient blocks $\\Gamma_{m|n} = \\Sigma_{mn}\\Sigma_{nn}^{-1}$. It follows that every cumulant of order $l \\ge 2$ equals $\\frac{(l-1)!}{2}\\operatorname{tr}(\\Gamma^l)$, so the moments, variance, and full distribution of multiinformation density depend on the regression structure of the partition and not on marginal variances. If the theorem is right, multiinformation density becomes a tractable, distribution-level probe of Gaussian dependence rather than just a one-number summary.","feed_headline":"One matrix fixes all cumulants of Gaussian multiinformation density","feed_subtitle":"For any normal vector, the full distribution of the log-density ratio is set by regression coefficients alone.","key_machinery":"The carrying object is the block matrix $\\Gamma$ of regression coefficients between subvectors. The proof route is the identity $id = I(X_1;\\dots;X_N) + \\frac{1}{2}(X-\\mu)^T\\Phi(X-\\mu)$ with $\\Phi = \\operatorname{diag}(\\Sigma_{11},\\dots,\\Sigma_{NN})^{-1} - \\Sigma^{-1}$, which rewrites multiinformation density as a constant plus a centered quadratic form in a normal vector. The moment-generating function of that quadratic form is a Gaussian integral whose value is $|I_d - t\\Gamma|^{-1/2}$; expanding the log determinant as $\\operatorname{tr}\\ln(I_d - t\\Gamma)$ and comparing powers of $t$ with the standard cumulant expansion yields the trace formula for $\\kappa_l$.","core_discovery":"The paper's central claim, Theorem 1, is an exact cumulant-generating formula: for a multivariate normal $X$ partitioned into $N$ subvectors, $\\ln \\mathbb{E}(e^{t\\,id}) = t I(X_1;\\dots;X_N) - \\frac{1}{2}\\ln|I_d - t\\Gamma|$, with $\\Gamma = \\Sigma\\operatorname{diag}(\\Sigma_{11},\\dots,\\Sigma_{NN})^{-1} - I_d$; the off-diagonal blocks of $\\Gamma$ are the regression coefficient matrices $\\Gamma_{m|n} = \\Sigma_{mn}\\Sigma_{nn}^{-1}$ for $m \\ne n$. The cumulants are $\\kappa_1(id) = I(X_1;\\dots;X_N)$ and $\\kappa_l(id) = \\frac{(l-1)!}{2}\\operatorname{tr}(\\Gamma^l)$ for $l \\ge 2$. This determines the entire law of $id$, including its variance, and it specializes cleanly: for two subvectors the odd cumulants vanish and the variance becomes the trace of the product of the two regression blocks, recovering the squared-canonical-correlation quantity; when every subvector is one-dimensional, the variance is the sum of squared correlation coefficients over all pairs.","pith_inferences":["Because the cumulant formula is a function of sample-estimable covariance blocks, a plug-in estimator of $\\operatorname{Var}(id)$ follows immediately; an independence test based on that estimator could be compared against permutation tests to see whether it adds power beyond scalar multiinformation.","The theorem suggests that $\\Gamma$ is the natural dependence parameter for Gaussian partitions generally: any two covariance matrices sharing the same regression blocks produce identical multiinformation-density distributions, even if their marginal variances differ.","The appendix's homogeneous-correlation example shows that the standardized multiinformation density is not asymptotically normal as dimension grows, so high-dimensional use of $id$ will need exact cumulant corrections or another limit theory rather than a central limit theorem."],"forward_implications":["All cumulants and therefore the whole distribution of multiinformation density are determined by the regression-coefficient block matrix $\\Gamma$; marginal variances are irrelevant to the law of $id$.","The variance of multiinformation density is $\\sum_{1\\le m<n\\le N} \\operatorname{tr}(\\Sigma_{mn}\\Sigma_{mm}^{-1}\\Sigma_{mn}\\Sigma_{nn}^{-1})$, reducing in the all-scalar case to the sum of squared pairwise correlation coefficients.","For a two-subvector partition, odd cumulants of $id$ are zero and even cumulants are traces of powers of the products $\\Gamma_{1|2}\\Gamma_{2|1}$ and $\\Gamma_{2|1}\\Gamma_{1|2}$, linking the distribution to canonical correlations.","Multiinformation itself is also a function of $\\Gamma$ alone, namely $I(X_1;\\dots;X_N) = -\\frac{1}{2}\\ln|I_d + \\Gamma|$.","Multiinformation density is identically zero, and has zero variance, exactly when the subvectors are mutually independent, giving distribution-level markers of independence."],"supporting_citations":[{"why":"Defines information density for two subvectors, the object this paper generalizes to $N$ subvectors.","marker":"Polyanskiy and Wu, 2017"},{"why":"Supplies the definition of regression coefficients and the simultaneous-diagonalization result used to justify the Gaussian integral's convergence.","marker":"Anderson, 2003"},{"why":"Provides the shift properties of cumulants and the cumulant-expansion form used to read off $\\kappa_l$.","marker":"Kendall, 1945"},{"why":"Gives the cumulants of quadratic forms in normal variables, the alternative derivation the paper cites for the trace formula.","marker":"Magnus, 1986"},{"why":"Introduced the two-subvector variance quantity that the theorem recovers as $\\kappa_2$ in the $N=2$ case.","marker":"Jupp and Mardia, 1980"},{"why":"Supports the log-determinant identity $\\ln|A| = \\operatorname{tr}(\\ln A)$ used to expand the cumulant-generating function.","marker":"Higham, 2007"}],"fun_headline_variants":["All Gaussian info cumulants fixed by one regression matrix","One matrix yields every cumulant of Gaussian info density","Cumulant formula: regression coefficients alone decide all","Gaussian multiinformation: cumulants from regression blocks","Exact law of info density via a single matrix Γ"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that $\\Sigma^{-1} - t\\Phi$ is positive definite for $t$ in a neighborhood of $0$, so the Gaussian integral defining the moment-generating function converges; the proof justifies this by simultaneous diagonalization of the two positive definite matrices.","fun_headline_variants_meta":{"raw":{"variants":["All Gaussian info cumulants fixed by one regression matrix","One matrix yields every cumulant of Gaussian info density","Cumulant formula: regression coefficients alone decide all","Gaussian multiinformation: cumulants from regression blocks","Exact law of info density via a single matrix Γ"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1255,"prompt_tokens":835,"completion_tokens":420,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":451,"completion_tokens_details":{"reasoning_tokens":343}},"tokens_in":451,"tokens_out":420,"duration_ms":5073,"temperature":1.0,"reasoning_tokens":343,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:30:38.993787+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a $d=4$ normal vector with homogeneous correlation $\\rho=0.5$, partition it into four scalar components, compute $id$ on many samples, and compare the empirical second and fourth cumulants with the paper's formulas $\\kappa_2 = \\rho^2 d(d-1)/2$ and $\\kappa_4 = 3\\rho^4[d + (d-1)^4 - 1]$; a systematic mismatch beyond sampling error would refute Theorem 1.","supporting_citations":[{"cited_title":"Lecture notes on information theory","cited_arxiv_id":null,"evidence_quote":"Defines information density for two subvectors, the object this paper generalizes to $N$ subvectors."},{"cited_title":"An Introduction to Multivariate Statistical Analysis","cited_arxiv_id":null,"evidence_quote":"Supplies the definition of regression coefficients and the simultaneous-diagonalization result used to justify the Gaussian integral's convergence."},{"cited_title":"The Advanced Theory of Statistics","cited_arxiv_id":null,"evidence_quote":"Provides the shift properties of cumulants and the cumulant-expansion form used to read off $\\kappa_l$."},{"cited_title":"The exact moments of a ratio of quadratic forms","cited_arxiv_id":null,"evidence_quote":"Gives the cumulants of quadratic forms in normal variables, the alternative derivation the paper cites for the trace formula."},{"cited_title":"A general correlation coeﬃcient for direc- tional data and related regression problems","cited_arxiv_id":null,"evidence_quote":"Introduced the two-subvector variance quantity that the theorem recovers as $\\kappa_2$ in the $N=2$ case."},{"cited_title":"Functions of matrices, in: Hogben, L","cited_arxiv_id":null,"evidence_quote":"Supports the log-determinant identity $\\ln|A| = \\operatorname{tr}(\\ln A)$ used to expand the cumulant-generating function."}],"review_version":1}