Pith. sign in

REVIEW 3 major objections 4 minor 58 references

For data from a multivariate t distribution fitted with a normal model, the expected Kullback–Leibler risk of the MLE density is expanded to second order in n, with explicit dependence on n, k, and ν.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 03:36 UTC pith:CGWNBK3O

load-bearing objection A credible worked example of the author's own expansion; the new 1/n^2 coefficient is asserted, not demonstrated, and the simulation is too coarse to confirm it. the 3 major comments →

arxiv 2607.23092 v1 pith:CGWNBK3O submitted 2026-07-25 stat.ME

Convergence of Estimative Density to Information Projection for Misspecified Normal Distribution Model

classification stat.ME MSC 62B1062E2062F12
keywords Kullback-Leibler divergenceinformation projectionestimative densitymultivariate normal modelmultivariate t distributionasymptotic expansionmaximum likelihoodmisspecified model
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks what happens when a statistician fits a multivariate normal model to data that actually come from a multivariate t-distribution. It first shows that the closest normal density to the true t-density, the information projection, is N(0, (ν/(ν−2)) I_k). It then derives the expected Kullback–Leibler distance between that projection and the normal density obtained by substituting the maximum likelihood estimates. Using a general asymptotic expansion from a previous paper, it obtains the first two terms of this expected distance as explicit functions of sample size n, dimension k, and degrees of freedom ν; the first term is (1/(2n))(p + k(k+2)/(ν−4)) and the second is a displayed polynomial divided by 24 n^2 (ν−6)(ν−4)^2. The closed forms let a practitioner compute, without simulation, the sample size that brings the expected K–L distance below any chosen threshold.

Core claim

The central claim is that the risk R = E[D[g(x;θ*)|g(x;θ̂)]] — where g(x;θ*) = N(0, λ I_k) with λ = ν/(ν−2) is the information projection and θ̂ is the MLE based on n observations from t_k(0, I_k, ν) — satisfies R = f(n,k,ν) + O(n^{-3}), with f(n,k,ν) = (1/(2n))(p + k(k+2)/(ν−4)) + (k/(24 n^2 (ν−6)(ν−4)^2)) times a polynomial in ν with coefficients polynomial in k (equation (8)). The derivation rests on a general second-order expansion of estimative densities in terms of the covariance matrices and higher-order cumulants of the exponential-family sufficient statistics, followed by an explicit enumeration of the third- and fourth-order cumulants of the quadratic statistics x_i x_j under both

What carries the argument

The argument runs through the exponential-family representation of the normal model with sufficient statistics ξ = (x_i, x_i x_j). A general asymptotic expansion (Theorem 2 of the cited previous paper) expresses the expected K–L risk to order n^{-2} in terms of: the covariance matrix G* of ξ under the fitted normal at the projection, the covariance matrix G under the true t-distribution, and third- and fourth-order cumulants κ and κ* of the same statistics. The t-distribution moments are computed via the representation x = √τ z with z ~ N_k(0, I_k) and τ = ν/U, U ~ χ²_ν, which yields λ_1 = ν/(ν−2), λ_2 = ν²/((ν−2)(ν−4)), λ_3 = ν³/((ν−2)(ν−4)(ν−6)). The main technical work is a systematic cla

Load-bearing premise

The paper assumes the general second-order asymptotic expansion for estimative densities applies to the maximum-likelihood normal fit under t-distributed data, but it only checks the moment condition ν>6 and does not verify the expansion's other regularity conditions.

What would settle it

Simulate the risk for a borderline case, e.g., k=1 or k=2, ν=7, n∈{10,20,40,80,160}, with a very large number of replications (10^6); if the difference between the Monte Carlo risk and f(n,k,ν) does not shrink like a constant times n^{-3} (i.e., if n^{3/2} times the difference is not bounded), then either the cumulant tables contain an error or the remainder from the general expansion is not uniformly O(n^{-3}). Also, recompute the index-pattern counts in Tables 2–5 for a fixed small k and check that the sum over classes equals the total number of relevant block choices; any mismatch would ide

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • For any k and ν>6, the expected K–L risk can be evaluated to O(n^{-3}) by a direct formula, eliminating the need for Monte Carlo in this setting.
  • The required sample size n* for a given risk threshold (e.g., C_{0.05}=1/50 corresponding to Bayes error 0.45) can be read off by inverting the formula; Table 1 matches simulations to within about 1%.
  • The first-order term explicitly quantifies the penalty for misspecification: relative to a correct normal model (risk p/(2n)), the t-data add k(k+2)/(ν−4) inside the bracket, which diverges as ν approaches 4 from above.
  • For high dimension k and moderate ν, the second-order term is not negligible; for (k,ν)=(100,8) and n around 100, the 1/n^2 term contributes a nontrivial correction, which the simulations reproduce.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the t-distribution is a scale mixture of normals, the same index-classification machinery should extend to any normal variance mixture (e.g., symmetric stable or Laplace) with finite sixth moments; the formulas would change only through the moments λ_1, λ_2, λ_3.
  • The paper only treats the case where the true covariance is spherical and the projection is a scalar multiple of the identity; for a general covariance t_k(0,Σ,ν), the projection would be N(0, (ν/(ν−2))Σ) and the risk formula would presumably acquire extra terms depending on the eigenstructure of Σ — a direct generalization the paper does not give.
  • A cautious reader should check whether the regularity conditions of the general expansion are satisfied by the MLE in the t model; the paper only verifies moment existence (ν>6), not the theorem's other conditions, so the O(n^{-3}) remainder is an unverified premise.
  • The sample-size numbers in Table 1 suggest that for large k the risk is dominated by the p/(2n) term, so the misspecification correction matters most for small-to-moderate k; this could inform when robust procedures are worth the trouble.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies a misspecified normal model: observations are i.i.d. from a multivariate t-distribution t_k(0,I_k,ν) with ν>6, but the fitted family is the k-variate normal family. The information projection is first identified as N_k(0,(ν/(ν−2))I_k). The main claim is an asymptotic expansion, Eq. (8), for the expected Kullback-Leibler risk R[g(x;θ*)|g(x;θ-hat)] between the information projection and the MLE plug-in normal density. The first-order term is (1/(2n))(p + k(k+2)/(ν−4)); the second-order term is a rational function of k and ν displayed in Eq. (8). The derivation computes the necessary moments, information matrices, and higher-order cumulants under both the t and normal laws, classifying index patterns into tables (Tables 2–5). The final step, however, is compressed: the three second-order sums in Section 5, Eqs. (19)–(21), are stated with 'the details are omitted.'

Significance. If Eq. (8) is correct, the paper provides a closed-form, parameter-free second-order expansion of the expected KL risk for a misspecified multivariate normal model, which can be used directly to compute required sample sizes via the threshold of [2]. The paper carefully derives the information projection, computes the basic moments and cumulant classifications, and includes simulation evidence in four configurations. These are real strengths. The main caveat is that the load-bearing second-order coefficient is obtained from an omitted algebraic step, and the general theorem from [2] is imported without an explicit verification of its hypotheses for the t_k(0,I_k,ν) setting. The contribution is therefore conditional on closing these gaps; it is not yet a fully verified asymptotic result at the claimed O(n^−3) precision.

major comments (3)
  1. [Section 5, Eqs. (19)–(21)] The values assigned to the three sums (19), (20), and (21) are asserted with 'the details are omitted.' These three sums determine the entire 1/n^2 coefficient in Eq. (8). A single missed index pattern or incorrect multiplicity in Tables 2–5 would change the polynomial in Eq. (8). This is a load-bearing gap in the central claim. The author should either include the full summation algebra in the paper or provide a machine-checkable supplementary derivation (e.g., a symbolic computation script).
  2. [Section 1 and Section 5] The paper applies Theorem 2 of [2] to the MLE under t_k(0,I_k,ν) data, but it does not verify the theorem's hypotheses in this setting. The only condition stated is ν>6. Since the O(n^−3) remainder in Eq. (8) is asserted by that theorem, the applicability of the theorem to the present misspecified exponential-family problem must be established explicitly. In particular, conditions involving existence, consistency, moment finiteness, and the rate of the remainder should be checked or cited with precise justification.
  3. [Tables 2–5] Several tables list rows labeled 'Others 0 unknown.' Because Eqs. (19)–(21) sum over all index patterns in R, the classification must either prove that all unlisted patterns contribute zero or enumerate them. If '0' is the value and only the count is unknown, that claim needs a proof; if the value is genuinely unknown for some patterns, the sums are not computable from the information given. This directly affects reproducibility of the central second-order coefficient.
minor comments (4)
  1. [Figure 1 / Table 1] The simulation uses N=1000 with no error bars or standard errors. This is supportive but not decisive for validating the 1/n^2 term, especially in cases where the second-order correction is small. Reporting Monte Carlo error would make the comparison more informative.
  2. [Section 4.4] There is a duplicated sentence: 'In this cumulant-type expression, all disconnected pairings cancel after the products of lower-order moments are subtracted. We now illustrate this cancellation mechanism. products of lower-order moments. We illustrate this cancellation mechanism.' The final two fragments appear to be a typographical repetition.
  3. [Tables 4 and 5] The word 'unkown' should be 'unknown'.
  4. [General] Some notational overload is present in Section 4.4: 'κ*(ij)(lm)(st)(uv)' and the A_ij notation are used without an explicit display of the final cumulant formula. A concise summary equation for this fourth-order cumulant would improve readability.

Circularity Check

0 steps flagged

No circularity: Eq. (8) is an application of a prior general expansion to independently computed moments; the self-citation is a legitimate external input, not a restatement.

full rationale

The derivation chain is not circular. The information projection is obtained directly from the first two moments of the t-distribution (Section 1, eqs. (5)-(6)), not from the target risk. The first-order term is computed in the paper from the covariance matrices G-tilde and G (Section 5, eq. (18)). The second-order coefficient in Eq. (8) is obtained by substituting the cumulant values in Tables 2-5 into the three general sums (19)-(21), which are terms from Theorem 2 of [2]. No parameter is fitted to the simulated risk; the simulation (Figure 1, Table 1) is an external check, and the paper does not adjust Eq. (8) to match it. The main dependency on Theorem 2 of [2] is a self-citation, but it is a general asymptotic expansion for estimative densities rather than a restatement of the special normal-vs-t result (8). The theorem is therefore an input with independent content, and applying it is not circular. There are genuine verification gaps, but they are not circularity: the paper states only 'nu>6' and does not check the regularity conditions of Theorem 2 for the MLE under t_k(0,I_k,nu), and Section 5 reports the second-order sums with 'the details are omitted' and tables include 'unknown' counts. These omissions weaken support for the O(n^-3) remainder and the 1/n^2 coefficient, but they do not make the claimed derivation reduce to its own inputs. The self-citation is load-bearing but legitimate; it does not forbid alternatives or define the target in terms of itself. Hence no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No free parameters are fitted to data and no new entities are introduced. The paper rests on a prior theorem by the same author [2] and on standard moment facts for normal and t distributions; the main unverified input is the applicability and remainder control of that theorem in the present misspecified case.

axioms (4)
  • domain assumption Theorem 2 of Sheena (2023) [2], including its equation (8), gives a valid asymptotic expansion for the expected KL risk of the misspecified MLE and has a negligible remainder in this setting.
    Invoked in Section 1 ('Applying Theorem 2 of [2]') and Section 5, but the theorem's regularity conditions are not stated or verified beyond ν>6.
  • standard math The representation x = √τ z with τ = ν/U, U ∼ χ²_ν independent of z ∼ N_k(0,I_k), gives the moments of t_k(0,I_k,ν), and the moments E[τ], E[τ²], E[τ³] are finite for ν>6.
    Section 2 uses this representation to derive equations (11); it is a standard stochastic representation of the multivariate t distribution.
  • standard math Moments of N_k(0, λ1 I_k) are computed by the pairing formula (9), and all relevant fourth- and sixth-order moments follow from it.
    Equation (9) and its use throughout Sections 3–4 are standard Gaussian moment combinatorics.
  • domain assumption The MLE in the misspecified normal model is consistent for the information projection θ* and its fluctuations are sufficiently regular for the asymptotic expansion.
    Needed for the expansion's validity; the paper does not exhibit or prove these regularity conditions for t_k(0,I_k,ν) data.

pith-pipeline@v1.3.0-alltime-deepseek · 20842 in / 17162 out tokens · 161665 ms · 2026-08-01T03:36:25.246483+00:00 · methodology

0 comments
read the original abstract

This paper investigates the convergence of an estimative multivariate normal density when the true distribution is a misspecified multivariate t-distribution. The statistical model is the family of k-dimensional normal distributions (N_k(\mu,\Sigma)), whereas the observations are assumed to follow (t_k(0,I_k,\nu)), with (\nu>6). The information projection of the true distribution onto the normal model is first identified as the normal distribution with mean zero and covariance matrix (\nu/(\nu-2)I_k). The main objective is to evaluate the expected Kullback-Leibler divergence between this information projection and the normal density obtained by substituting the maximum likelihood estimator into the model. Using a general asymptotic expansion for estimative densities, the paper derives explicit first- and second-order terms of the risk as functions of the sample size (n), the dimension (k), and the degrees of freedom (\nu). To obtain the second-order term, the paper calculates the required moments, information matrices, and higher-order cumulants under both the multivariate normal and multivariate t-distributions. In particular, the complicated third- and fourth-order cumulants involving quadratic sufficient statistics are classified according to their index patterns, and their values and multiplicities are systematically derived.

Figures

Figures reproduced from arXiv: 2607.23092 by Yo Sheena.

Figure 1
Figure 1. Figure 1: Comparison of the risk functions. which are well defined for ν > 6. Using (9) and (10), we can readily derive the moments of x when x ∼ tk(0, Ik, ν). For mutually distinct indices 1 ≤ i, j, l, m, s, t ≤ k, the following equations hold: E[xi ] = 0, E[x 2 i ] = λ1, E[xixj ] = 0, E[x 3 i ] = 0, E[x 2 i xj ] = 0, E[xixjxl ] = 0, E[x 4 i ] = 3λ2, E[x 3 i xj ] = 0, E[x 2 i x 2 j ] = λ2, E[x 2 i xjxl ] = 0, E[xix… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

58 extracted references

  1. [1]

    Arbel and A

    M. Arbel and A. Gretton. Kernel Conditional Exponential Family. Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics

  2. [2]

    S. Amari. Differential Geometry of Curved Exponential Families--Curvature and Information Loss. Annals of Statistics

  3. [3]

    2012 , publisher=

    Differential-geometrical methods in statistics , author=. 2012 , publisher=

  4. [4]

    IEEE Transactions on Information Theory , volume=

    -Divergence Is Unique, Belonging to Both f -Divergence and Bregman Divergence Classes , author=. IEEE Transactions on Information Theory , volume=. 2009 , publisher=

  5. [5]

    S. Amari. Information Geometry and Its Applications

  6. [6]

    Amari and A

    S. Amari and A. Chichocki. Information Geometry of Divergence Function. Bulletin of the Polish Academy of Sciences : Technical Sciences

  7. [7]

    Amari and H

    S. Amari and H. Nagaoka. Methods of Information Geometry

  8. [8]

    O. E. Barndorff-Nielsen. Information and Exponential Families in Statistical Theory

  9. [9]

    A. R. Barron. The convergence in information of probability density estimators. IEEE Int. Symp on Inform. Theory, Kobe, Japan, June 19-24

  10. [10]

    A. R. Barron, L. Gy\'' o orfi and E. C. van der Meulen. Distribution Estimation Consistent in Total Variation and in Two Types of Information Divergence. IEEE Transactions on Information Theory

  11. [11]

    A. R. Barron and C. Sheu. Approximation of Density Functions by Sequences of Exponential Families. Annals of Statistics

  12. [12]

    L. D. Brown. Fundamentals of Statistical Exponential Families

  13. [13]

    Canu and A

    S. Canu and A. Smola. Kernel Methods and the Exponential Family. Neurocomputing

  14. [14]

    N. N. Cencov. Statistical Decision Rules and Optimal Inference

  15. [15]

    2014 , publisher=

    Geometric modeling in probability and statistics , author=. 2014 , publisher=

  16. [16]

    2017 , publisher=

    Differential geometry and statistics , author=. 2017 , publisher=

  17. [17]

    1996 , publisher=

    On the relationships between alpha-connections and the asymptotic properties of predictive distributions , author=. 1996 , publisher=

  18. [18]

    Csisz \' a r

    I. Csisz \' a r. I-divergence geometry of probability distributions and minimization problems. Annals of Probability

  19. [19]

    Csisz \' a r

    I. Csisz \' a r. Why least squares and maximum entropy? An axiomatic approach to inference for linear inverse problems. Annals of Statistics

  20. [20]

    B. Efron. Defining the Curvature of a Statistical Problem (with Applications to Second-order Efficiency). Annals of Statistics

  21. [21]

    Efron and R

    B. Efron and R. Tibshirani. Using Specially Designed Exponential Families for Density Estimation. Annals of Statistics

  22. [22]

    Journal of statistical planning and inference , volume=

    Asymptotical improvement of maximum likelihood estimators on Kullback--Leibler loss , author=. Journal of statistical planning and inference , volume=. 2008 , publisher=

  23. [23]

    H. A. David and H. N. Nagaraja. Ordered Statistics 3rd ed

  24. [24]

    Drezner and D

    Z. Drezner and D. Zerom. A Simple and Effective Discretization of a Continuous Random Variable. Communications in Statistics --Simulation and Computation

  25. [25]

    S. Eguchi. A Differential Geometric Approach to Statistical Inference on the Basis of Contrast Functionals. Hiroshima Mathematical Journal

  26. [26]

    S. Eguchi. Geometry of Minimum Contrast. Hiroshima Mathematical Journal

  27. [27]

    Fukumizu

    K. Fukumizu. Exponential Manifold by Reproducing Kernel Hilbert Spaces. Algebraic and Geometric Methods in Statistics

  28. [28]

    P. Hall. On Kullback--Leibler Loss and Density Estimation. Annals of Statisics

  29. [29]

    J. A. Hartigan. The Maximum Likelihood Prior. Annals of Statisics

  30. [30]

    Konishi and G

    S. Konishi and G. Kitagawa. Generalized Information Criteria in Model Selection. Biometrika

  31. [31]

    Koenker and I

    R. Koenker and I. Mizera. Quasi-concave Density Estimation. Annals of Statistics

  32. [32]

    F. Komaki. On Asymptotic Properties of Predictive Distributions. Biometrika

  33. [33]

    F. Komaki. Asymptotic Properties of Bayesian Predictive Densities When the Distributions of Data and Target Variables are Different. Bayesian Analysis

  34. [34]

    Maji , journal=

    P. Maji , journal=. f -Information Measures for Efficient Selection of Discriminative Genes From Microarray Data , year=

  35. [35]

    Naito and S

    K. Naito and S. Eguchi. Density Estimation with Minimization of U-divergence. Machine Learning

  36. [36]

    S. Portnoy. Asymptotic Behavior of Likelihood Methods for Exponential Families When the Number of Parameters Tends to Infinity. Annals of Statistics

  37. [37]

    Puchkin, S

    N. Puchkin, S. Samsonov, D. Belomestny, E. Moulines and A. Naumov. Rates of convergence for density estimation with generative adversarial networks. Journal of Machine Learning Research

  38. [38]

    IEEE Transactions on Signal Processing , volume=

    A study on invariance of f -divergence and its application to speech recognition , author=. IEEE Transactions on Signal Processing , volume=. 2010 , publisher=

  39. [39]

    C. R. Rao. Statistics And Truth: Putting Chance To Work

  40. [40]

    D. W. Scott. Multivariate Density Estimation

  41. [41]

    Shimazaki and S

    H. Shimazaki and S. Shinomoto. A Method for Selecting the Bin Size of a Time Histogram. Neural Computation

  42. [42]

    Y. Sheena. Asymptotic Expansion of Risk for a Regression Model with respect to -Divergence with an Application to the Sample Size Problem. Far East Journal of Theoretical Statistics

  43. [43]

    Y. Sheena. Asymptotic Expansion of the Risk of Maximum Likelihood Estimator with respect to -divergence as a Measure of the Difficulty of Specifying a Parametric Model. Communications in Statistics -- Theory and Methods

  44. [44]

    Y. Sheena. Asymptotic expansion of the risk of maximum likelihood estimator with respect to -divergence as a measure of the difficulty of specifying a parametric model -- with detailed proof

  45. [45]

    Y. Sheena. Estimation of a Continuous Distribution on the Real Line by Discretization Methods. Metrika

  46. [46]

    Y. Sheena. Convergence of estimative density: criterion for model complexity and sample size. Statistical Papers

  47. [47]

    Y. Sheena. MLE Convergence Speed to Information Projection of Exponential Family: Criterion for Model Dimension and Sample Size -- Complete Proof Version--

  48. [48]

    Sriperumbudur and K

    B. Sriperumbudur and K. Fukumizu and A. Gretton and A. Hyv a rinen and R. Kumar. Density Estimation in Infinite Dimensional Exponential Families. Journal of Machine Learning Research

  49. [49]

    C. J. Stone. Large-sample Inference for Log-spline Models. Annals of Statistics

  50. [50]

    C. J. Stone. An Asymptotically Optimal Histogram Selection Rule

  51. [51]

    Sundberg

    R. Sundberg. Statistical Modeling for Exponential Families

  52. [52]

    Takeuchi

    K. Takeuchi. Distribution of Information Statistics and Criteria for Adequacy of Models (in Japanese). Mathematical Science

  53. [53]

    I. Vajda. Theory of Statistical Inference and Information

  54. [54]

    A. W. Van Der Vaart. Asymptotic Statistics

  55. [55]

    M. J. Wainwright and M. I. Jordan. Graphical Models, Exponential Families, and Variational Inference

  56. [56]

    W. H. Wong and T. A. Severini. On a Maximum Likelihood Estimation in Infinite Dimensional Parameter Spaces. Annals of Statistics

  57. [57]

    F. D. Zhang and Y. M. Shi and H. K. T. Ng and R. B. Wang. Information Geometry of Generalized Bayesian Prediction Using -Divergence as Loss Funtion. IEEE Trans. Inf. Theory

  58. [58]

    1994 , howpublished =

    Nash, Warwick, Sellers, Tracy, Talbot, Simon, Cawthorn, Andrew, and Ford, Wes , title =. 1994 , howpublished =