REVIEW 3 major objections 4 minor 58 references
For data from a multivariate t distribution fitted with a normal model, the expected Kullback–Leibler risk of the MLE density is expanded to second order in n, with explicit dependence on n, k, and ν.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 03:36 UTC pith:CGWNBK3O
load-bearing objection A credible worked example of the author's own expansion; the new 1/n^2 coefficient is asserted, not demonstrated, and the simulation is too coarse to confirm it. the 3 major comments →
Convergence of Estimative Density to Information Projection for Misspecified Normal Distribution Model
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the risk R = E[D[g(x;θ*)|g(x;θ̂)]] — where g(x;θ*) = N(0, λ I_k) with λ = ν/(ν−2) is the information projection and θ̂ is the MLE based on n observations from t_k(0, I_k, ν) — satisfies R = f(n,k,ν) + O(n^{-3}), with f(n,k,ν) = (1/(2n))(p + k(k+2)/(ν−4)) + (k/(24 n^2 (ν−6)(ν−4)^2)) times a polynomial in ν with coefficients polynomial in k (equation (8)). The derivation rests on a general second-order expansion of estimative densities in terms of the covariance matrices and higher-order cumulants of the exponential-family sufficient statistics, followed by an explicit enumeration of the third- and fourth-order cumulants of the quadratic statistics x_i x_j under both
What carries the argument
The argument runs through the exponential-family representation of the normal model with sufficient statistics ξ = (x_i, x_i x_j). A general asymptotic expansion (Theorem 2 of the cited previous paper) expresses the expected K–L risk to order n^{-2} in terms of: the covariance matrix G* of ξ under the fitted normal at the projection, the covariance matrix G under the true t-distribution, and third- and fourth-order cumulants κ and κ* of the same statistics. The t-distribution moments are computed via the representation x = √τ z with z ~ N_k(0, I_k) and τ = ν/U, U ~ χ²_ν, which yields λ_1 = ν/(ν−2), λ_2 = ν²/((ν−2)(ν−4)), λ_3 = ν³/((ν−2)(ν−4)(ν−6)). The main technical work is a systematic cla
Load-bearing premise
The paper assumes the general second-order asymptotic expansion for estimative densities applies to the maximum-likelihood normal fit under t-distributed data, but it only checks the moment condition ν>6 and does not verify the expansion's other regularity conditions.
What would settle it
Simulate the risk for a borderline case, e.g., k=1 or k=2, ν=7, n∈{10,20,40,80,160}, with a very large number of replications (10^6); if the difference between the Monte Carlo risk and f(n,k,ν) does not shrink like a constant times n^{-3} (i.e., if n^{3/2} times the difference is not bounded), then either the cumulant tables contain an error or the remainder from the general expansion is not uniformly O(n^{-3}). Also, recompute the index-pattern counts in Tables 2–5 for a fixed small k and check that the sum over classes equals the total number of relevant block choices; any mismatch would ide
If this is right
- For any k and ν>6, the expected K–L risk can be evaluated to O(n^{-3}) by a direct formula, eliminating the need for Monte Carlo in this setting.
- The required sample size n* for a given risk threshold (e.g., C_{0.05}=1/50 corresponding to Bayes error 0.45) can be read off by inverting the formula; Table 1 matches simulations to within about 1%.
- The first-order term explicitly quantifies the penalty for misspecification: relative to a correct normal model (risk p/(2n)), the t-data add k(k+2)/(ν−4) inside the bracket, which diverges as ν approaches 4 from above.
- For high dimension k and moderate ν, the second-order term is not negligible; for (k,ν)=(100,8) and n around 100, the 1/n^2 term contributes a nontrivial correction, which the simulations reproduce.
Where Pith is reading between the lines
- Because the t-distribution is a scale mixture of normals, the same index-classification machinery should extend to any normal variance mixture (e.g., symmetric stable or Laplace) with finite sixth moments; the formulas would change only through the moments λ_1, λ_2, λ_3.
- The paper only treats the case where the true covariance is spherical and the projection is a scalar multiple of the identity; for a general covariance t_k(0,Σ,ν), the projection would be N(0, (ν/(ν−2))Σ) and the risk formula would presumably acquire extra terms depending on the eigenstructure of Σ — a direct generalization the paper does not give.
- A cautious reader should check whether the regularity conditions of the general expansion are satisfied by the MLE in the t model; the paper only verifies moment existence (ν>6), not the theorem's other conditions, so the O(n^{-3}) remainder is an unverified premise.
- The sample-size numbers in Table 1 suggest that for large k the risk is dominated by the p/(2n) term, so the misspecification correction matters most for small-to-moderate k; this could inform when robust procedures are worth the trouble.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies a misspecified normal model: observations are i.i.d. from a multivariate t-distribution t_k(0,I_k,ν) with ν>6, but the fitted family is the k-variate normal family. The information projection is first identified as N_k(0,(ν/(ν−2))I_k). The main claim is an asymptotic expansion, Eq. (8), for the expected Kullback-Leibler risk R[g(x;θ*)|g(x;θ-hat)] between the information projection and the MLE plug-in normal density. The first-order term is (1/(2n))(p + k(k+2)/(ν−4)); the second-order term is a rational function of k and ν displayed in Eq. (8). The derivation computes the necessary moments, information matrices, and higher-order cumulants under both the t and normal laws, classifying index patterns into tables (Tables 2–5). The final step, however, is compressed: the three second-order sums in Section 5, Eqs. (19)–(21), are stated with 'the details are omitted.'
Significance. If Eq. (8) is correct, the paper provides a closed-form, parameter-free second-order expansion of the expected KL risk for a misspecified multivariate normal model, which can be used directly to compute required sample sizes via the threshold of [2]. The paper carefully derives the information projection, computes the basic moments and cumulant classifications, and includes simulation evidence in four configurations. These are real strengths. The main caveat is that the load-bearing second-order coefficient is obtained from an omitted algebraic step, and the general theorem from [2] is imported without an explicit verification of its hypotheses for the t_k(0,I_k,ν) setting. The contribution is therefore conditional on closing these gaps; it is not yet a fully verified asymptotic result at the claimed O(n^−3) precision.
major comments (3)
- [Section 5, Eqs. (19)–(21)] The values assigned to the three sums (19), (20), and (21) are asserted with 'the details are omitted.' These three sums determine the entire 1/n^2 coefficient in Eq. (8). A single missed index pattern or incorrect multiplicity in Tables 2–5 would change the polynomial in Eq. (8). This is a load-bearing gap in the central claim. The author should either include the full summation algebra in the paper or provide a machine-checkable supplementary derivation (e.g., a symbolic computation script).
- [Section 1 and Section 5] The paper applies Theorem 2 of [2] to the MLE under t_k(0,I_k,ν) data, but it does not verify the theorem's hypotheses in this setting. The only condition stated is ν>6. Since the O(n^−3) remainder in Eq. (8) is asserted by that theorem, the applicability of the theorem to the present misspecified exponential-family problem must be established explicitly. In particular, conditions involving existence, consistency, moment finiteness, and the rate of the remainder should be checked or cited with precise justification.
- [Tables 2–5] Several tables list rows labeled 'Others 0 unknown.' Because Eqs. (19)–(21) sum over all index patterns in R, the classification must either prove that all unlisted patterns contribute zero or enumerate them. If '0' is the value and only the count is unknown, that claim needs a proof; if the value is genuinely unknown for some patterns, the sums are not computable from the information given. This directly affects reproducibility of the central second-order coefficient.
minor comments (4)
- [Figure 1 / Table 1] The simulation uses N=1000 with no error bars or standard errors. This is supportive but not decisive for validating the 1/n^2 term, especially in cases where the second-order correction is small. Reporting Monte Carlo error would make the comparison more informative.
- [Section 4.4] There is a duplicated sentence: 'In this cumulant-type expression, all disconnected pairings cancel after the products of lower-order moments are subtracted. We now illustrate this cancellation mechanism. products of lower-order moments. We illustrate this cancellation mechanism.' The final two fragments appear to be a typographical repetition.
- [Tables 4 and 5] The word 'unkown' should be 'unknown'.
- [General] Some notational overload is present in Section 4.4: 'κ*(ij)(lm)(st)(uv)' and the A_ij notation are used without an explicit display of the final cumulant formula. A concise summary equation for this fourth-order cumulant would improve readability.
Circularity Check
No circularity: Eq. (8) is an application of a prior general expansion to independently computed moments; the self-citation is a legitimate external input, not a restatement.
full rationale
The derivation chain is not circular. The information projection is obtained directly from the first two moments of the t-distribution (Section 1, eqs. (5)-(6)), not from the target risk. The first-order term is computed in the paper from the covariance matrices G-tilde and G (Section 5, eq. (18)). The second-order coefficient in Eq. (8) is obtained by substituting the cumulant values in Tables 2-5 into the three general sums (19)-(21), which are terms from Theorem 2 of [2]. No parameter is fitted to the simulated risk; the simulation (Figure 1, Table 1) is an external check, and the paper does not adjust Eq. (8) to match it. The main dependency on Theorem 2 of [2] is a self-citation, but it is a general asymptotic expansion for estimative densities rather than a restatement of the special normal-vs-t result (8). The theorem is therefore an input with independent content, and applying it is not circular. There are genuine verification gaps, but they are not circularity: the paper states only 'nu>6' and does not check the regularity conditions of Theorem 2 for the MLE under t_k(0,I_k,nu), and Section 5 reports the second-order sums with 'the details are omitted' and tables include 'unknown' counts. These omissions weaken support for the O(n^-3) remainder and the 1/n^2 coefficient, but they do not make the claimed derivation reduce to its own inputs. The self-citation is load-bearing but legitimate; it does not forbid alternatives or define the target in terms of itself. Hence no significant circularity.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Theorem 2 of Sheena (2023) [2], including its equation (8), gives a valid asymptotic expansion for the expected KL risk of the misspecified MLE and has a negligible remainder in this setting.
- standard math The representation x = √τ z with τ = ν/U, U ∼ χ²_ν independent of z ∼ N_k(0,I_k), gives the moments of t_k(0,I_k,ν), and the moments E[τ], E[τ²], E[τ³] are finite for ν>6.
- standard math Moments of N_k(0, λ1 I_k) are computed by the pairing formula (9), and all relevant fourth- and sixth-order moments follow from it.
- domain assumption The MLE in the misspecified normal model is consistent for the information projection θ* and its fluctuations are sufficiently regular for the asymptotic expansion.
read the original abstract
This paper investigates the convergence of an estimative multivariate normal density when the true distribution is a misspecified multivariate t-distribution. The statistical model is the family of k-dimensional normal distributions (N_k(\mu,\Sigma)), whereas the observations are assumed to follow (t_k(0,I_k,\nu)), with (\nu>6). The information projection of the true distribution onto the normal model is first identified as the normal distribution with mean zero and covariance matrix (\nu/(\nu-2)I_k). The main objective is to evaluate the expected Kullback-Leibler divergence between this information projection and the normal density obtained by substituting the maximum likelihood estimator into the model. Using a general asymptotic expansion for estimative densities, the paper derives explicit first- and second-order terms of the risk as functions of the sample size (n), the dimension (k), and the degrees of freedom (\nu). To obtain the second-order term, the paper calculates the required moments, information matrices, and higher-order cumulants under both the multivariate normal and multivariate t-distributions. In particular, the complicated third- and fourth-order cumulants involving quadratic sufficient statistics are classified according to their index patterns, and their values and multiplicities are systematically derived.
Figures
Reference graph
Works this paper leans on
-
[1]
Arbel and A
M. Arbel and A. Gretton. Kernel Conditional Exponential Family. Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics
-
[2]
S. Amari. Differential Geometry of Curved Exponential Families--Curvature and Information Loss. Annals of Statistics
-
[3]
2012 , publisher=
Differential-geometrical methods in statistics , author=. 2012 , publisher=
2012
-
[4]
IEEE Transactions on Information Theory , volume=
-Divergence Is Unique, Belonging to Both f -Divergence and Bregman Divergence Classes , author=. IEEE Transactions on Information Theory , volume=. 2009 , publisher=
2009
-
[5]
S. Amari. Information Geometry and Its Applications
-
[6]
Amari and A
S. Amari and A. Chichocki. Information Geometry of Divergence Function. Bulletin of the Polish Academy of Sciences : Technical Sciences
-
[7]
Amari and H
S. Amari and H. Nagaoka. Methods of Information Geometry
-
[8]
O. E. Barndorff-Nielsen. Information and Exponential Families in Statistical Theory
-
[9]
A. R. Barron. The convergence in information of probability density estimators. IEEE Int. Symp on Inform. Theory, Kobe, Japan, June 19-24
-
[10]
A. R. Barron, L. Gy\'' o orfi and E. C. van der Meulen. Distribution Estimation Consistent in Total Variation and in Two Types of Information Divergence. IEEE Transactions on Information Theory
-
[11]
A. R. Barron and C. Sheu. Approximation of Density Functions by Sequences of Exponential Families. Annals of Statistics
-
[12]
L. D. Brown. Fundamentals of Statistical Exponential Families
-
[13]
Canu and A
S. Canu and A. Smola. Kernel Methods and the Exponential Family. Neurocomputing
-
[14]
N. N. Cencov. Statistical Decision Rules and Optimal Inference
-
[15]
2014 , publisher=
Geometric modeling in probability and statistics , author=. 2014 , publisher=
2014
-
[16]
2017 , publisher=
Differential geometry and statistics , author=. 2017 , publisher=
2017
-
[17]
1996 , publisher=
On the relationships between alpha-connections and the asymptotic properties of predictive distributions , author=. 1996 , publisher=
1996
-
[18]
Csisz \' a r
I. Csisz \' a r. I-divergence geometry of probability distributions and minimization problems. Annals of Probability
-
[19]
Csisz \' a r
I. Csisz \' a r. Why least squares and maximum entropy? An axiomatic approach to inference for linear inverse problems. Annals of Statistics
-
[20]
B. Efron. Defining the Curvature of a Statistical Problem (with Applications to Second-order Efficiency). Annals of Statistics
-
[21]
Efron and R
B. Efron and R. Tibshirani. Using Specially Designed Exponential Families for Density Estimation. Annals of Statistics
-
[22]
Journal of statistical planning and inference , volume=
Asymptotical improvement of maximum likelihood estimators on Kullback--Leibler loss , author=. Journal of statistical planning and inference , volume=. 2008 , publisher=
2008
-
[23]
H. A. David and H. N. Nagaraja. Ordered Statistics 3rd ed
-
[24]
Drezner and D
Z. Drezner and D. Zerom. A Simple and Effective Discretization of a Continuous Random Variable. Communications in Statistics --Simulation and Computation
-
[25]
S. Eguchi. A Differential Geometric Approach to Statistical Inference on the Basis of Contrast Functionals. Hiroshima Mathematical Journal
-
[26]
S. Eguchi. Geometry of Minimum Contrast. Hiroshima Mathematical Journal
-
[27]
Fukumizu
K. Fukumizu. Exponential Manifold by Reproducing Kernel Hilbert Spaces. Algebraic and Geometric Methods in Statistics
-
[28]
P. Hall. On Kullback--Leibler Loss and Density Estimation. Annals of Statisics
-
[29]
J. A. Hartigan. The Maximum Likelihood Prior. Annals of Statisics
-
[30]
Konishi and G
S. Konishi and G. Kitagawa. Generalized Information Criteria in Model Selection. Biometrika
-
[31]
Koenker and I
R. Koenker and I. Mizera. Quasi-concave Density Estimation. Annals of Statistics
-
[32]
F. Komaki. On Asymptotic Properties of Predictive Distributions. Biometrika
-
[33]
F. Komaki. Asymptotic Properties of Bayesian Predictive Densities When the Distributions of Data and Target Variables are Different. Bayesian Analysis
-
[34]
Maji , journal=
P. Maji , journal=. f -Information Measures for Efficient Selection of Discriminative Genes From Microarray Data , year=
-
[35]
Naito and S
K. Naito and S. Eguchi. Density Estimation with Minimization of U-divergence. Machine Learning
-
[36]
S. Portnoy. Asymptotic Behavior of Likelihood Methods for Exponential Families When the Number of Parameters Tends to Infinity. Annals of Statistics
-
[37]
Puchkin, S
N. Puchkin, S. Samsonov, D. Belomestny, E. Moulines and A. Naumov. Rates of convergence for density estimation with generative adversarial networks. Journal of Machine Learning Research
-
[38]
IEEE Transactions on Signal Processing , volume=
A study on invariance of f -divergence and its application to speech recognition , author=. IEEE Transactions on Signal Processing , volume=. 2010 , publisher=
2010
-
[39]
C. R. Rao. Statistics And Truth: Putting Chance To Work
-
[40]
D. W. Scott. Multivariate Density Estimation
-
[41]
Shimazaki and S
H. Shimazaki and S. Shinomoto. A Method for Selecting the Bin Size of a Time Histogram. Neural Computation
-
[42]
Y. Sheena. Asymptotic Expansion of Risk for a Regression Model with respect to -Divergence with an Application to the Sample Size Problem. Far East Journal of Theoretical Statistics
-
[43]
Y. Sheena. Asymptotic Expansion of the Risk of Maximum Likelihood Estimator with respect to -divergence as a Measure of the Difficulty of Specifying a Parametric Model. Communications in Statistics -- Theory and Methods
-
[44]
Y. Sheena. Asymptotic expansion of the risk of maximum likelihood estimator with respect to -divergence as a measure of the difficulty of specifying a parametric model -- with detailed proof
-
[45]
Y. Sheena. Estimation of a Continuous Distribution on the Real Line by Discretization Methods. Metrika
-
[46]
Y. Sheena. Convergence of estimative density: criterion for model complexity and sample size. Statistical Papers
-
[47]
Y. Sheena. MLE Convergence Speed to Information Projection of Exponential Family: Criterion for Model Dimension and Sample Size -- Complete Proof Version--
-
[48]
Sriperumbudur and K
B. Sriperumbudur and K. Fukumizu and A. Gretton and A. Hyv a rinen and R. Kumar. Density Estimation in Infinite Dimensional Exponential Families. Journal of Machine Learning Research
-
[49]
C. J. Stone. Large-sample Inference for Log-spline Models. Annals of Statistics
-
[50]
C. J. Stone. An Asymptotically Optimal Histogram Selection Rule
-
[51]
Sundberg
R. Sundberg. Statistical Modeling for Exponential Families
-
[52]
Takeuchi
K. Takeuchi. Distribution of Information Statistics and Criteria for Adequacy of Models (in Japanese). Mathematical Science
-
[53]
I. Vajda. Theory of Statistical Inference and Information
-
[54]
A. W. Van Der Vaart. Asymptotic Statistics
-
[55]
M. J. Wainwright and M. I. Jordan. Graphical Models, Exponential Families, and Variational Inference
-
[56]
W. H. Wong and T. A. Severini. On a Maximum Likelihood Estimation in Infinite Dimensional Parameter Spaces. Annals of Statistics
-
[57]
F. D. Zhang and Y. M. Shi and H. K. T. Ng and R. B. Wang. Information Geometry of Generalized Bayesian Prediction Using -Divergence as Loss Funtion. IEEE Trans. Inf. Theory
-
[58]
1994 , howpublished =
Nash, Warwick, Sellers, Tracy, Talbot, Simon, Cawthorn, Andrew, and Ford, Wes , title =. 1994 , howpublished =
1994
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.