Pith. sign in

REVIEW 4 major objections 4 minor 23 references

This paper claims that a second-order stochastic Taylor expansion of mutual information at the independence distribution yields exactly Pearson's chi-square statistic, giving an explicit formula that ties the G² and χ² independence tests.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 15:32 UTC pith:CG3Y6VV2

load-bearing objection The core identity is true but classical; the paper overstates it as exact, and the appendix needs repair before publication. the 4 major comments →

arxiv 2607.18425 v1 pith:CG3Y6VV2 submitted 2026-07-20 math.ST stat.TH

Mutual Information second order expansion is the Pearson's chi-square statistic

classification math.ST stat.TH MSC 62H1762F0562B10
keywords mutual informationPearson chi-square statisticlikelihood-ratio testindependence testingsecond-order expansiondelta methodcontingency tablesG-test
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that the mutual information between two discrete variables, when expanded to second order around the independence distribution through a stochastic Taylor expansion, reproduces exactly Pearson's chi-square statistic as its quadratic term. If that identity holds, the likelihood-ratio test statistic G² and the Pearson chi-square statistic χ² are connected by an explicit formula: G² = χ² plus a third-order remainder that can be computed from third derivatives of mutual information. The author proves by direct algebra that the plug-in version of the quadratic form equals χ², and thereby claims to unveil the exact relation between the two classical tests. A sympathetic reader would care because this bridges information-theoretic dependence measurement and classical hypothesis testing, and offers a way to quantify the finite-sample gap between the two statistics.

Core claim

The paper's central claim is that the second-order term in the δ-method expansion of 2n·MI(p̂_n) is exactly the Pearson chi-square statistic once the Hessian is evaluated at the product-of-marginals distribution. Concretely, with H the Hessian of MI at the independence point, the quadratic form (p̂_n − p̂^0_n)ᵀ H (p̂_n − p̂^0_n) simplifies algebraically to χ²_n = Σ (n_{ij} − n_{i*}n_{*j}/n)² / (n_{i*}n_{*j}). Since 2n·MI(p̂_n) is the likelihood-ratio statistic G²_n, eq. (13) follows: G²_n = χ²_n + 2n R(||p̂_n − p̂^0||²), where the remainder R is composed of third-order derivatives of MI. The paper treats this as the 'exact relation' between G² and χ², with the difference explicitly quantifia

What carries the argument

The load-bearing object is the Hessian matrix H of mutual information at the product-of-marginals distribution, decomposed as H = A + B + C with structured rank-one-like blocks. The proof in the appendix expands the quadratic form cell by cell and shows that all cross-terms cancel, leaving exactly Pearson's sum of (observed−expected)²/expected. The δ-method stochastic Taylor expansion supplies the asymptotic framework, and the plug-in estimate of the marginals is what converts the theoretical quadratic form into the computable chi-square statistic.

Load-bearing premise

The algebraic identity that turns the quadratic form into Pearson's chi-square is carried out at the true independence distribution, but the test statistic uses estimated marginals; the paper only notes that convergence 'should be revisited', so if plug-in yields only asymptotic equality, the claimed exact relation weakens to the classical asymptotic one.

What would settle it

Simulate multinomial samples from an independent product distribution with known marginals; compute G² and χ² from the same table, then compute the remainder term 2n R(||p̂_n − p̂0||²) from third derivatives of MI; if the equation G² = χ² + remainder fails by more than Monte Carlo error for moderate n, the central claim is refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Under independence, G² and χ² are not merely asymptotically equivalent: their difference is a third-order remainder term, which can in principle be estimated from the data.
  • The plug-in quadratic form with the Hessian at the estimated marginals is exactly Pearson's chi-square, not just an approximation, for any contingency table.
  • The remainder term, third derivatives of MI, offers a route to higher-order corrections to the chi-square approximation in finite samples.
  • This gives an information-theoretic interpretation of Pearson's statistic as the leading curvature of MI at independence.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the plug-in substitution of the true marginals by estimates preserves only an asymptotic identity, then eq. (13) may reduce to the classical asymptotic equivalence rather than a finite-sample exact relation; the paper flags this gap itself.
  • The same Hessian-cancellation mechanism likely extends to other f-divergences, suggesting that many divergence-based independence statistics share Pearson's chi-square as their universal second-order form.
  • The spectral assertion that eigenvalues of ΣH collapse to 0 and 1 under the null is not proven here; the distributional claims depend on it and would need a standalone proof.
  • Because the remainder is O_p(n · ||p̂ − p̂0||²), its practical size depends on how fast the joint distribution converges to the product of marginals; in high-dimensional sparse tables the correction may be large.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper studies the relationship between mutual information (MI), the likelihood-ratio statistic G², and Pearson's chi-square statistic for independence in I×J contingency tables. It defines T¹_n = 2n·MI(p̂_n) (= G²_n) and T²_n = n(p̂_n − p_H0)ᵀ H(p_H0)(p̂_n − p_H0), where H is the Hessian of MI. Under H0, Theorem 1 states that 2n·MI(p̂_n) is asymptotically equivalent to a quadratic form with a chi-square distribution. The central claim is Eq. (13): G²_n = χ²_n + 2n·R(||p̂_n − p_H0||²), with R an explicitly quantifiable third-order remainder. The proof in Appendix A aims to show that the plug-in version of T²_n equals χ²_n. The algebraic identity in Appendix A is essentially correct up to a missing factor of n, but the exact equality as stated is not established because the expansion points in Eq. (8) and Appendix A differ, and the remainder R is never actually computed.

Significance. The observation that the Hessian of MI evaluated at the product-of-marginals distribution reproduces Pearson's chi-square statistic is elegant and not widely recognized; I independently verified the 2×2 case and found the underlying quadratic-form identity to be correct. If accompanied by a controlled remainder, Eq. (13) would provide a useful higher-order comparison between two classical independence tests. However, in its present form the manuscript overclaims: the proof establishes at most the classical asymptotic equivalence, the Appendix contains a factor-of-n error in the displayed equality, and the 'explicitly quantified' remainder is never written down. These issues affect the central claim, though they may be repairable within the paper's scope.

major comments (4)
  1. [§3, Eqs. (8), (12)–(13); Appendix A] There is a mismatch of expansion points. Eq. (8) defines T²_n with both the centering and the Hessian evaluated at the unknown true p_H0, while Appendix A proves a quadratic-form identity for the plug-in version in which p_H0 is replaced by p̂0_n (the estimated product of marginals). These are different quadratic forms; since p̂0_n − p_H0 = O_p(n^{-1/2}), after multiplication by n the difference is only o_p(1), which is the classical asymptotic equivalence, not the exact relation claimed in Eq. (13). The manuscript's own §3 lists plug-in substitution as approximation 3 and says 'convergence should be revisited.' To support Eq. (13), the paper must either define T²_n and Eq. (13) consistently with p̂0_n, or state the result as asymptotic with an explicit error bound.
  2. [Appendix A, final displayed equation; Eq. (12)] The final equality in Appendix A states (p̂_n−p̂0_n)ᵀ(A+B+C)(p̂_n−p̂0_n) = Σ_{i,j}(p_{ij}−p0_{ij})²/p0_{ij} = χ²_n. However, by Eq. (12), χ²_n = n·Σ(p̂_{ij}−p̂_i*p̂_*j)²/(p̂_i*p̂_*j), while the computed quadratic form has no factor n. Since p̂_n−p̂0_n = O_p(n^{-1/2}), the quadratic form is O_p(1/n), not O_p(1). The intended identity T²_n = n(p̂_n−p̂0_n)ᵀH(p̂0_n)(p̂_n−p̂0_n) = χ²_n is correct, and the algebra can be repaired by restoring the factor n, but the proof as written is incorrect and the displayed equality is dimensionally inconsistent.
  3. [§3, Eq. (13) and surrounding text] The paper asserts that the difference G²_n − χ²_n is 2n·R(||p̂_n−p_H0||²) with R composed of third-order derivatives, and calls this 'explicitly quantified' in the abstract and §3. But R is never defined, no expression is given, and it is not demonstrated that the Taylor remainder of MI is a function only of ||p̂_n−p_H0||². The sentence 'can be explicitly quantified by computing third order derivatives' is a promise, not a result. As it stands, Eq. (13) is a placeholder for the very quantity the paper claims to derive.
  4. [Theorem 1, §3] Theorem 1 and the subsequent claim that under H0 the eigenvalues of ΣH collapse to 0 and 1 with tr(HΣ) = (I−1)(J−1) are imported from the author's previous work [16,17] and no proof or self-contained statement is given here. This is load-bearing for the distributional conclusion that 2n·MI(p̂_n) is asymptotically χ²_{(I−1)(J−1)}. The manuscript should at least state the spectral result as a lemma with a proof, or give a precise theorem from a publicly available source, rather than relying on a self-citation without verification.
minor comments (4)
  1. [Notation throughout] The symbols p̂0_n, p̂H0, and p_H0 are used interchangeably in different places. Define the notation once and use it consistently, especially in Eqs. (8), (12), and (13).
  2. [Appendix A, around A.1] The text 'p0_IJ / pJ = pI*' appears to be a typo for 'p0_IJ / p*J = pI*'; as written it is confusing because pJ is not a standard marginal symbol.
  3. [Eq. (2)] The notation λ⊤(χ²) is nonstandard. Clarify that λ is a vector of eigenvalues and χ² is a vector of independent chi-square random variables, or use a more conventional quadratic-form notation.
  4. [Introduction, §1] The Introduction contrasts the paper's result with the 'leading-order approximation' known in the literature. Given that the proof as written establishes only an asymptotic equivalence, the novelty claim should be tempered unless the exact remainder is actually supplied.

Circularity Check

2 steps flagged

Self-cited expansion plus a plug-in/true-parameter switch make the exact G^2–χ^2 relation depend on prior work and an unproved substitution; the underlying plug-in quadratic-form identity is independent.

specific steps
  1. self citation load bearing [Theorem 1 (Section 2, Eqs. 2–3) and Section 3 ('Under independence the eigenvalues...')]
    "We write the following theorem adapted from our previous work [16, 17]. ... See [17] for a demonstration and a general second order expansion for non-vanishing MI gradient. ... Under independence the eigenvalues of ΣH collapses to 0,1, and tr(HΣ) = (I−1)(J−1)."

    The asymptotic expansion, the Hessian structure, and the eigenvalue collapse that turn MI into a χ² distribution are not derived here; they are quoted from the same authors' prior work. This is load-bearing for the distributional bridge between MI and Pearson's statistic, though not for the later Appendix-A algebra, which is self-contained. The self-citation does not by itself force the central algebraic identity, but it imports a substantial unproved component into the paper's statistical claims.

  2. self definitional [Section 3, Eq. (8) vs. Appendix A and Eq. (13)]
    "T^2_n = n( ˆpn −p H0 )⊤H|_{pH0}( ˆpn −p H0 ). ... We show in Appendix A that for the plug-in version of T^2_n, its expression reduces to χ^2_n. ... The following calculus stands for any valid probability distribution ˆpn with marginals ˆp0_n."

    Appendix A proves the quadratic-form identity only for the plug-in version, with both the centre and the Hessian evaluated at p̂0, not for the T²_n defined at the unknown p_H0 in Eq. (8). Since p̂0 − p_H0 = O_p(n^{-1/2}), the two forms differ at finite n; after scaling by n the difference is o_p(1), which is classical asymptotic equivalence, not the advertised exact identity. The paper itself concedes 'convergence should be revisited' and lists plug-in substitution as approximation 3. Thus Eq. (13) is obtained by silently switching the expansion point, and the remainder R is never explicitly computed.

full rationale

The paper's genuine algebraic content is Appendix A: for the plug-in Hessian and plug-in product-marginal centre, the quadratic form of MI's Hessian reduces to Pearson's chi-square statistic. That computation is self-contained algebra and is not circular. However, the paper's headline exact relation, Eq. (13), uses the notation p_H0 and claims G²_n = χ²_n + 2nR(||p̂_n−p_H0||²) with an 'explicitly quantified' remainder. The appendix does not prove this for the p_H0-centred statistic defined in Eq. (8); it proves only the plug-in identity, and the paper explicitly flags the plug-in substitution as an approximation ('convergence should be revisited', 'three approximations'). This is a definitional switch more than a circular reduction: the identity is made to appear exact by replacing the stated expansion point with the estimated one. Separately, the distributional and Hessian machinery is imported from the authors' own prior work [16,17] rather than proved here; this is a load-bearing self-citation, but it does not make the Appendix-A algebra circular. There is no fitted parameter renamed as a prediction, no uniqueness theorem used to forbid alternatives, and no ansatz smuggled in by citation that forces the result. Overall, the score is low-to-moderate: the self-citations and unmarked plug-in switch weaken the exactness claim, but the central plug-in quadratic-form identity has independent algebraic content.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No free parameters are fitted — the paper is pure asymptotics. The central identity is supported by the explicit H formula (eq. 3), which I spot-checked numerically against the true MI Hessian; the identity is correct for the plug-in expansion point. The main ledger items are background assumptions imported from the author's prior work and the plug-in substitution whose justification the paper itself defers.

axioms (4)
  • domain assumption Under H0 the eigenvalues of ΣH are 0 or 1 and tr(HΣ) = (I−1)(J−1)
    §3 states this immediately before eqs. (6)–(7) without proof, pointing to the author's own [17]. Classical (likelihood-ratio asymptotics) but load-bearing for the χ²_{(I−1)(J−1)} degrees-of-freedom claim.
  • domain assumption Plug-in substitution of p̂⁰ for p_H0 in T²_n preserves the identity with Pearson's χ²_n
    §3: 'A practical solution is to consider a plug-in estimator... convergence should be revisited.' The appendix's algebra (numerically verified here) proves the identity for H evaluated at p̂⁰, not at the true p_H0 used in the definition of T²_n (eq. 8).
  • domain assumption The Hessian decomposition H = A+B+C (eq. 3) is the MI Hessian on the (IJ−1)-dimensional simplex
    Formula attributed to the author's [16,17]. Verified numerically here for a 2×2 case: δ'ᵀHδ' = Σ(p̂−p̂⁰)²/p̂⁰ for the plug-in expansion point. However, the appendix's intermediate derivations of the B and C pieces are inconsistent with eq. (3) as typeset.
  • domain assumption i.i.d. multinomial sampling from a fixed I×J distribution with all cell probabilities positive
    Implicit throughout §2; required for the multinomial covariance Σ of √n p̂_n and for the δ-method normality step. Standard but unstated.

pith-pipeline@v1.3.0-alltime-deepseek · 7585 in / 57212 out tokens · 404556 ms · 2026-08-01T15:32:01.553057+00:00 · methodology

0 comments
read the original abstract

We show that MI connects subtly and elegantly the two best-known state-of-the-art independence test statistics: the $G^2$ and the Pearson's chi-square statistic $\chi^2$. Furthermore, we show that the MI connects directly those statistics by an elegant formula arising from a stochastic Taylor expansion of MI ($\delta$-method). MI second order term is precisely $\chi^2$ up to a scale factor. As a consequence, by this connection, the difference between $G^2$ and $\chi^2$ can be explicitly quantified.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

23 extracted references · 1 linked inside Pith

  1. [1]

    A mathematical theory of communication,

    C. E. Shannon, “A mathematical theory of communication,”The Bell system technical journal, vol. 27, no. 3, pp. 379–423, 1948

  2. [2]

    T. M. Cover and J. A. Thomas,Elements of information theory. Wiley-Interscience, 2006

  3. [3]

    Multivariate information transmission,

    W. J. McGill, “Multivariate information transmission,”Psychometrika, vol. 19, no. 2, pp. 97–116, 1954

  4. [4]

    Information theoretical analysis of multivariate correlation,

    S. Watanabe, “Information theoretical analysis of multivariate correlation,”IBM Journal of research and develop- ment, vol. 4, no. 1, pp. 66–82, 1960

  5. [5]

    A smoothed bootstrap test for independence based on mutual information,

    E. H. Wu, L. Philip, and W. K. Li, “A smoothed bootstrap test for independence based on mutual information,” Computational statistics & data analysis, vol. 53, no. 7, pp. 2524–2536, 2009

  6. [6]

    Testing unconditional and conditional independence via mutual information,

    C. Ai, L.-H. Sun, Z. Zhang, and L. Zhu, “Testing unconditional and conditional independence via mutual information,”Journal of Econometrics, vol. 240, no. 2, p. 105335, 2024

  7. [7]

    Nonparametric independence testing via mutual information,

    T. B. Berrett and R. J. Samworth, “Nonparametric independence testing via mutual information,”Biometrika, vol. 106, no. 3, pp. 547–566, 2019

  8. [8]

    Measuring and testing dependence by correlation of distances,

    G. J. Székely, M. L. Rizzo, and N. K. Bakirov, “Measuring and testing dependence by correlation of distances,” The Annals of Statistics, vol. 35, no. 6, pp. 2769 – 2794, 2007

  9. [9]

    Brownian distance covariance,

    G. J. Székely and M. L. Rizzo, “Brownian distance covariance,”The Annals of Applied Statistics, vol. 3, no. 4, pp. 1236–1265, 2009

  10. [10]

    Energy statistics: A class of statistics based on distances,

    G. J. Szekely and M. L. Rizzo, “Energy statistics: A class of statistics based on distances,”Journal of statistical planning and inference, vol. 143, no. 8, pp. 1249–1272, 2013

  11. [11]

    Measuring statistical dependence with hilbert-schmidt norms,

    A. Gretton, O. Bousquet, A. Smola, and B. Schölkopf, “Measuring statistical dependence with hilbert-schmidt norms,” inAlgorithmic Learning Theory, S. Jain, H. U. Simon, and E. Tomita, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2005, pp. 63–77

  12. [12]

    The randomized dependence coefficient,

    D. Lopez-Paz, P. Hennig, and B. Schölkopf, “The randomized dependence coefficient,”Advances in neural information processing systems, vol. 26, 2013

  13. [13]

    Agresti,Categorical data analysis, 2nd ed

    A. Agresti,Categorical data analysis, 2nd ed. John Wiley & Sons, 2012

  14. [14]

    A. W. Van der Vaart,Asymptotic statistics. Cambridge university press, 2000, vol. 3

  15. [15]

    A. W. Van Der Vaart and J. A. Wellner,Weak convergence. Springer, 1996

  16. [16]

    On the use of mutual information for testing independence,

    M. Marinescu, “On the use of mutual information for testing independence,”arXiv preprint arXiv:2502.17636, 2025

  17. [17]

    A bias correction for the mutual information sample estimator,

    M. Marinescu and C. Balcau, “A bias correction for the mutual information sample estimator,”Statistics & Probability Letters, p. 110802, 2026

  18. [18]

    The distribution function of a linear combination of chi-squares,

    P. G. Moschopoulos and W. B. Canada, “The distribution function of a linear combination of chi-squares,” Computers & mathematics with applications, vol. 10, no. 4-5, pp. 383–386, 1984

  19. [19]

    New methods to compute the generalized chi-square distribution,

    A. Das, “New methods to compute the generalized chi-square distribution,”Journal of Statistical Computation and Simulation, vol. 95, no. 12, pp. 2608–2642, 2025

  20. [20]

    C. R. Rao,Linear statistical inference and its applications. Wiley New York, 1973, vol. 2

  21. [21]

    Pearson, “X

    K. Pearson, “X. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling,”The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, vol. 50, no. 302, pp. 157–175, 1900

  22. [22]

    Ix. On the problem of the most efficient tests of statistical hypotheses,

    J. Neyman and E. S. Pearson, “Ix. On the problem of the most efficient tests of statistical hypotheses,”Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character, vol. 231, no. 694-706, pp. 289–337, 1933

  23. [23]

    Pearson’s x2 and the loglikelihood ratio statistic g2: a comparative review,

    N. Cressie and T. R. C. Read, “Pearson’s x2 and the loglikelihood ratio statistic g2: a comparative review,” International Statistical Review/Revue Internationale de Statistique, pp. 19–43, 1989. 5 APREPRINT- JULY22, 2026 A Demonstration thatT 2 n is the Pearson’s chi-square statistic Following the notation along this paper, we show that T 2 n = √n( ˆpn −...