REVIEW 5 major objections 5 minor 1 cited by
Non-Parametric Goodness-of-Fit Tests Using Tsallis Entropy Measures
T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims a Tsallis-entropy test statistic built from nearest-neighbour estimates converges to 0 under the null model and to a positive constant otherwise, giving a non-parametric goodness-of-fit test.
desk verdict The test statistics are not well-defined, the advertised estimator is missing, and the work adds little over Rényi-based tests; not ready for refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The object that carries the argument is the Tsallis entropy functional $H_q(f) = \frac{1}{1-q}\left(\int_{\mathbb{R}^m} f^q(x)\,dx - 1\right)$, paired with the $k$-nearest-neighbour estimator $\hat{S}_{k,N,q} = \frac{1}{N}\sum_{i=1}^N (\zeta_{i,k,N})^{1-q}$, where $\zeta_{i,k,N}$ is built from the Euclidean distance to the $k$th nearest neighbour. The test statistic is the gap between that estimate and the claimed maximum-entropy value of the hypothesized distribution, so the mechanism that converts entropy estimation into hypothesis testing is the identification of $\tfrac{1}{2}\log|\hat\Sigma_N| + T(x; a, q, \sigma)$ as that maximum. The convergence argument then depends on the moment conditions in Theorem 3 that guarantee $\hat{S}_{k,N,q}$ converges.
What would settle it
Simulate data from a multivariate q-Gaussian with known q and covariance, compute the closed-form Tsallis entropy given in Section 3 of the paper, form the test statistic with that value in place of the unexplained $T(x; a, q, \sigma)$, and check whether the Monte Carlo average tends to 0 as $N$ grows; if it does not, the claimed consistency under the null is false.
Extended reading notes
Core claim
The paper claims to establish a class of test statistics of the form $Q_{N,k}^{\mathrm{Tsallis}} = H_q^{\mathrm{upper}} - \hat{S}_{k,N,q}$, where $\hat{S}_{k,N,q}$ is the $k$-nearest-neighbour estimator of Tsallis entropy and $H_q^{\mathrm{upper}} = \tfrac{1}{2}\log|\hat\Sigma_N| + T(x; a, q, \sigma)$ is meant to be the maximum Tsallis entropy of the assumed model. The argument is that the covariance estimate converges to the true covariance, the nearest-neighbour entropy estimator converges to the true Tsallis entropy under the moment conditions of Theorem 3, and so by a standard convergence argument $Q$ converges in probability to 0 under the null and to a positive constant $c>0$ otherwise. This 0-versus-positive separation is what turns an entropy estimate into a goodness-of-fit test. The simulation study reports convergence of the statistic, stable critical values, and an empirical approach to normality as $q\to 1$.
Load-bearing premise
The whole test depends on the assumption that the paper has correctly identified a formula for the maximum Tsallis entropy of the null model, but the formula it uses is never defined and, as written, does not even have matching units, so if that identification is wrong the claimed convergence to 0 under the null collapses.
Editorial extensions
If this is right
- If the consistency claim is correct, the tests offer a non-parametric check of whether multivariate data follow a q-Gaussian or generalized Gaussian law, with no need to estimate the density before testing.
- The tabulated Monte Carlo critical values give immediate 5% thresholds for dimensions 2 and 3 and sample sizes up to 1000, so the method is usable without deriving an analytic null distribution.
- Because the statistic converges to a positive constant under alternatives, the tests can in principle detect departures in tail weight or shape parameter q, not just location-scale shifts.
- The empirical approach to normality as q approaches 1 suggests that near-Gaussian nulls can be calibrated with standard normal quantiles rather than a full Monte Carlo table.
- The estimator's moment conditions tie practical applicability to tail behavior: convergence requires finite q-weighted moments, so extremely heavy-tailed alternatives may need larger samples.
Reading between the lines
- Editorial inference: A direct check the reader could run is to replace the unexplained $T(x; a, q, \sigma)$ term in $H_q^{\mathrm{upper}}$ with the closed-form Tsallis entropy of the q-Gaussian derived in Section 3, and see whether the simulated statistic still converges to 0 under the null; the paper's own maximum-entropy discussion suggests this substitution is what the term is meant to be.
- Editorial inference: If the maximum-entropy identification fails, the statistic no longer has a clean 0-versus-positive interpretation; it would reduce to comparing a nearest-neighbour entropy estimate with a constant, which is a weaker model-checking heuristic rather than a calibrated goodness-of-fit test.
- Editorial inference: The same construction could be lifted to other maximum-entropy families by replacing the closed-form entropy term with any correctly derived maximum-entropy expression, but the consistency proof would need to be re-run because it depends on the specific moment and support conditions.
- Editorial inference: A testable extension is an adaptive choice of the neighbourhood size k, since the paper notes sensitivity to k and a data-driven rule (for instance, minimising a variance-bias proxy of $\hat{S}_{k,N,q}$) could make the test fully automatic.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes goodness-of-fit tests for multivariate generalized Gaussian and q-Gaussian distributions based on Tsallis entropy. The test statistics QTsallis_N,k and QTsallis*_N,k are defined in Eqs. (11) and (12) as the difference between a claimed maximum Tsallis entropy H_upper_q of the null model and a k-nearest-neighbor estimator of Tsallis entropy. The paper claims, in §5.2, that these statistics converge in probability to 0 under the null and to a positive constant under alternatives, and it presents Monte Carlo critical values and convergence plots for various parameter settings.
Significance. If the proposed statistics were well-defined and the asymptotic claims correct, the paper would offer a genuinely useful entropy-based approach to multivariate goodness-of-fit testing for heavy-tailed and compactly supported distributions, complementing likelihood-based methods. The paper also provides a large set of Monte Carlo tables and graphs, which is a useful resource if the underlying statistics are valid. However, the central construction is not mathematically well-defined: the quantities T1 and T2 are never defined, the expression for H_upper_q is dimensionally inconsistent, and the claimed limit Q → 0 under H0 is asserted without a derivation and contradicts the paper's own §3.3 entropy formula. These defects undermine the core contribution, so the potential significance is not realized in the current version.
major comments (5)
- [§5.1, Eqs. (11) and (12)] The test statistics are not well-defined. H_upper_q is given as (1/2) log|Σhat_N| + T1(x; a, q, σ) (resp. T2), but T1 and T2 are never defined anywhere in the manuscript; §5 only says they are distributions in class K. If T1(x; a, q, σ) is a density evaluated at a point x, then H_upper_q depends on a single observation x, whereas the k-NN estimator Rhat_TS_k,N,q is a scalar computed from the full sample; the difference is therefore not a meaningful statistic. If T1 is instead meant to be a constant, its value and its derivation are absent. The null hypotheses 'X ∼ T1(x; a, q, σ)' and 'X ∼ T2(x; a, q, σ)' are thus not identifiable from the text.
- [§5.2 and §3.3] The claimed convergence Q → 0 under H0 is unsupported. H_upper_q is asserted to be the maximum Tsallis entropy of the assumed model, but this identification is never proved and is inconsistent with the paper's own computation in §3.3, where the Tsallis entropy of the multivariate q-Gaussian is given as a power-law expression in |Σ|: H_q = 1/(q−1)(1 − C_q^q |Σ|^{(1−q)/2} 2^{m/2}Γ(m/2)/((1−q)^{m/2}Γ(m/2 + 1/(q−1)))). This is not (1/2) log|Σ| plus a constant/density term. Consequently, the equality H_upper_q = S_q under the null, on which the limit in §5.2 rests, does not follow from the manuscript. The Monte Carlo critical values are therefore calibrated for a statistic whose null mean is an uncomputed, generally nonzero constant.
- [§5 and Theorem 3] The asymptotic justification mixes incompatible parameter regimes. Theorem 3 is stated only for q ∈ (0,1), yet the first test statistic (11) is proposed for q ∈ (1,3). Remark 4 cites a different consistency result for q ∈ (1,(k+1)/2), but this range excludes many configurations used in the simulations (e.g., k=1 with q=2.5 or 3.0), and Remark 4 is not integrated into the main consistency argument. Thus the statement that 'By Theorem 3, the distributions T1 and T2 are included in this class' does not cover the parameter settings actually used.
- [Definition 1] Definition 1 misstates the r-th moment. It defines K_r(f) = E(||X||^r) = 1/(q−1) ∫ ||x||^r f^q(x) dx, but the left-hand side E(||X||^r) equals ∫ ||x||^r f(x) dx, which is not equal to the right-hand side. The subsequent critical moment r_c(f) and the moment conditions (9) and (10) rely on this quantity, so the conditions as stated are ambiguous or incorrect.
- [§6, Table 1 and Table 2] The numerical results do not support the claimed convergence to zero. Table 1 reports 5% critical values of Q_T_N,k(m,q) that remain around 0.03 across all sample sizes N from 100 to 1000 and all parameter settings, rather than decreasing to 0. Table 2 reports slopes β in a log-log regression of |E[Q]| on N that are mostly positive or near zero, which is inconsistent with the statement that E[Q] → 0 as N → ∞. The sentence 'Furthermore numerical applications are proof of the theoretical idea that E[Q_T_N,k(m,q)] → 0' is also methodologically circular: simulations cannot prove an asymptotic limit, and the tabulated values do not exhibit the claimed behavior.
minor comments (5)
- [Abstract and §1] The term 'non-parametric' is used loosely: the test statistics require estimation of the covariance matrix Σ and refer to specific parametric null families (q-Gaussian and generalized Gaussian). A more precise wording would avoid the implication that the tests are fully distribution-free.
- [§1 (last paragraph)] The organization paragraph omits Section 4 ('Tsallis Entropy: Statistical Estimation Method') and says 'Section 5 introduces entropy-based goodness-of-fit test statistics', which is correct, but the list skips from Section 3 to Section 5; this should be corrected.
- [Eq. (5)] The normalization constant in the multivariate exponential power density appears to lack a factor of 2^{m/s} in the denominator relative to standard forms, and the use of Γ(m/s + 1) versus Γ(m/s) should be checked against the cited literature.
- [Figure 2 caption] The caption states that 'decreasing q sharpens the peak and widens the tails', but for q in the heavy-tailed regime (q < 1), the displayed behavior may be non-monotonic in q; a precise statement of which q range is being described would improve clarity.
- [References] Reference [5] and the first part of the bibliography entry [4] appear to refer to the same article (Berrett, Samworth, and Yuan); duplicate entries should be merged or clearly distinguished.
Circularity Check
The claimed consistency Q→0 under H0 is built into the definitional label of H_upper_q as the model's Tsallis entropy and is never derived from the paper's own entropy formula.
-
self definitional
[Section 5.1, Eqs. (11)–(12); Section 5.2]
"Q^{Tsallis}_{N,k}(m,q) = H^{upper}_q − \hat S_{k,N,q}, where H^{upper}_q = 1/2 log |\hatΣ_N| + T1(x; a, q, σ) denotes the maximum Tsallis entropy under the assumed model. ... According to Theorem 3 ... the test statistics converge in probability as: lim_{N→∞} Q^{Tsallis}_{N,k}(m,q) → 0 if X ∼ T1(x; a, q, σ), c>0 otherwise."
Under H0, Theorem 3 gives \hat S_{k,N,q} → S_q(T1), so the claimed limit Q→0 is equivalent to H^{upper}_q = S_q(T1). The paper does not compute S_q(T1) for T1; it simply labels the expression (1/2)log|\hatΣ_N| + T1(x;...) as 'the maximum Tsallis entropy under the assumed model.' Thus the consistency result is not derived from a calculated entropy but is a corollary of the definitional label attached to H^{upper}_q. The label is also unsupported: T1 is never defined, and §3.3 derives for the q-Gaussian a power-law entropy in |Σ|, not (1/2)log|Σ| plus a scalar. Hence the central asymptotic claim reduces by construction to an unproved identification.
full rationale
The only substantial circular move is in Section 5.1–5.2: the test statistic is defined as H^{upper}_q − \hat S_{k,N,q}, and H^{upper}_q is asserted, without derivation, to be the maximum Tsallis entropy of the assumed model. Once that assertion is accepted, the claimed limit Q→0 under H0 is immediate from the cited kNN estimator consistency and contains no independent predictive content. The paper's own §3.3 expression for the q-Gaussian Tsallis entropy is a power law in |Σ|, not the log-determinant-plus-scalar form used in (11)–(12), and T1/T2 are never defined, so the identification is not merely unproved but inconsistent with the paper's own equations. The kNN estimator consistency itself is cited to the author's earlier work [28,8], but the estimator and related results also trace to the external literature [20,24], so this self-citation is not independently load-bearing. The Monte Carlo critical values are standard null calibration rather than a fitted parameter renamed as a prediction. Because the central asymptotic claim reduces to a definitional identification, a score of 6 is appropriate; the failure is partial circularity compounded by an unverified and dimensionally suspect formula.
Assumptions & free parameters
free parameters (2)
- q (Tsallis entropic index)
- k (number of nearest neighbors)
assumptions (4)
- domain assumption The density f is Lebesgue-continuous and satisfies the moment conditions (9) or (10) so that the k-NN entropy estimator is consistent.
- ad hoc to paper The maximum Tsallis entropy under the null q-Gaussian model is given by H_upper_q = 1/2 log|Sigma_hat_N| + T1(x; a, q, sigma), and analogously for T2.
- ad hoc to paper T1(x; a, q, sigma) and T2(x; a, q, sigma) are well-defined distributions in the class K for which the estimator is consistent.
- standard math Standard measure-theoretic and Gamma/Beta integral identities used to evaluate the q-Gaussian entropy in Section 3.3.
Cite this review
Pith. "Pith review of Non-Parametric Goodness-of-Fit Tests Using Tsallis Entropy Measures." pith.science (2026). https://pith.science/paper/2YQHHHYZ
@misc{pith2026250614242,
author = {Pith},
title = {Pith review of: Non-Parametric Goodness-of-Fit Tests Using Tsallis Entropy Measures},
year = {2026},
howpublished = {\url{https://pith.science/paper/2YQHHHYZ}},
note = {Machine review of arXiv:2506.14242}
}
abstract
In this paper, we investigate new procedures for statistical testing based on Tsallis entropy, a parametric generalization of Shannon entropy. Focusing on multivariate generalized Gaussian and $q$-Gaussian distributions, we develop entropy-based goodness-of-fit tests based on maximum entropy formulations and nearest neighbour entropy estimators. Furthermore, we propose a novel iterative approach for estimating the shape parameters of the distributions, which is crucial for practical inference. This method extends entropy estimation techniques beyond traditional approaches, improving precision in heavy-tailed and non-Gaussian contexts. The numerical experiments are demonstrative of the statistical properties and convergence behaviour of the proposed tests. These findings are important for disciplines that require robust distributional tests, such as machine learning, signal processing, and information theory.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
A recursive fine-tuning study claims quality filtering reverses model collapse, but the paper's own data show the filtered model only matched its starting score.
Reference graph
Works this paper leans on
-
[1]
Sumiyoshi Abe. Heat and entropy in nonextensive thermodynamics: transmutation from tsallis theory to r´ enyi-entropy-based theory. Physica A: Statistical Mechanics and its Applications , 300(3-4):417–423, 2001
work page 2001
-
[2]
Further properties of tsallis entropy and its application
Ghadah Alomani and Mohamed Kayid. Further properties of tsallis entropy and its application. Entropy, 25(2):199, 2023
work page 2023
-
[3]
T. B. Berrett and R. J. Samworth. Nonparametric independence testing via mutual information. Biometrika, 106(3):547–566, 2019
work page 2019
-
[5]
T. B. Berrett, R. J. Samworth, and M. Yuan. Efficient multivariate entropy estima- tion via k-nearest neighbour distances. Annals of Statistics , 47(1):288–318, 2019
work page 2019
-
[6]
A. Bulinski and D. Dimitrov. Statistical estimation of the Shannon entropy. Acta Mathematica Sinica, English Series , 35(1):17–46, 2019
work page 2019
-
[7]
A. Bulinski and A. Kozhevin. Statistical estimation of conditional Shannon entropy. ESAIM Probability and Statistics , 23:350–386, 2019
work page 2019
-
[8]
Entropy-based goodness-of-fit tests for multivariate distributions
Mehmet Cadirci. Entropy-based goodness-of-fit tests for multivariate distributions . PhD thesis, Cardiff University, 2021
work page 2021
-
[9]
Entropy-based test for generalised gaussian distributions
Mehmet Siddik Cadirci, Dafydd Evans, Nikolai Leonenko, and Vitalii Makogin. Entropy-based test for generalised gaussian distributions. Computational Statistics & Data Analysis , 173:107502, 2022
work page 2022
Show all 30 references
-
[10]
Su una estensione dello schema delle curve normali di ordine r alle variabili doppie
S De Simoni. Su una estensione dello schema delle curve normali di ordine r alle variabili doppie. Statistica, 37(447-474):63, 1968
1968
-
[11]
Delattre and N
S. Delattre and N. Fournier. On the Kozachenko–Leonenko entropy estimator. Jour- nal of Statistical Planning and Inference , 185:69–93, 2017
2017
-
[12]
The strong uniform consistency of nearest neighbor density estimates
Luc P Devroye and Terry J Wagner. The strong uniform consistency of nearest neighbor density estimates. The Annals of Statistics , pages 536–540, 1977
1977
-
[13]
Generalization of shannon’s theorem for tsallis entropy
Roberto JV dos Santos. Generalization of shannon’s theorem for tsallis entropy. Journal of Mathematical Physics , 38(8):4104–4107, 1997
1997
-
[14]
K. T. Fang and S. Kotz. Symmetric multivariate and related distributions. Mono- graphs on Statistics and Applied Probability , 36, 1990
1990
-
[15]
On the maximum entropy principle and the minimization of the fisher information in tsallis statistics
Shigeru Furuichi. On the maximum entropy principle and the minimization of the fisher information in tsallis statistics. Journal of Mathematical Physics , 50(1), 2009
2009
-
[16]
A multivariate gen- eralization of the power exponential family of distributions
Eusebio G´ omez, MA Gomez-Viilegas, and J Miguel Mar ´ ın. A multivariate gen- eralization of the power exponential family of distributions. Communications in Statistics-Theory and Methods , 27(3):589–600, 1998. 19
1998
-
[17]
Consistency property of elliptic probability density functions
Yutaka Kano. Consistency property of elliptic probability density functions. Journal of Multivariate Analysis , 51(1):139–147, 1994
1994
-
[18]
Some results on tsallis entropy measure and k-record values
Vikas Kumar. Some results on tsallis entropy measure and k-record values. Physica A: Statistical Mechanics and its Applications , 462:667–673, 2016
2016
-
[19]
Characterization results based on dynamic tsallis cumulative residual entropy
Vikas Kumar. Characterization results based on dynamic tsallis cumulative residual entropy. Communications in Statistics-Theory and Methods, 46(17):8343–8354, 2017
2017
-
[20]
N. N. Leonenko, L. Pronzato, and V. Savani. A class of R´ enyi information estimators for multidimensional densities. Annals of Statistics , 36(5):2153–2182, 2008
2008
-
[21]
A nonparametric estimate of a multivariate density function
Don O Loftsgaarden and Charles P Quesenberry. A nonparametric estimate of a multivariate density function. The Annals of Mathematical Statistics , 36(3):1049– 1051, 1965
1965
-
[22]
Tsallis’ entropy maximization procedure revisited
S Martınez, F Nicol´ as, Flavia Pennini, and A Plastino. Tsallis’ entropy maximization procedure revisited. Physica A: Statistical Mechanics and its Applications , 286(3- 4):489–502, 2000
2000
-
[23]
On r´ enyi and tsallis entropies and divergences for exponential families
Frank Nielsen and Richard Nock. On r´ enyi and tsallis entropies and divergences for exponential families. arXiv preprint, arXiv:1105.3259, 2011
2011 arXiv
-
[24]
M. D. Penrose and J. E. Yukich. Laws of large numbers and nearest neighbor dis- tances. In Advances in Directional and Linear Statistics , pages 189–199. Springer, 2011
2011
-
[25]
Some characterization results on dynamic cumulative residual tsallis entropy
Madan Mohan Sati, Nitin Gupta, et al. Some characterization results on dynamic cumulative residual tsallis entropy. Journal of probability and statistics , 2015, 2015
2015
-
[26]
A mathematical theory of communication
Claude Elwood Shannon. A mathematical theory of communication. The Bell system technical journal, 27(3):379–423, 1948
1948
-
[27]
S. S. Shapiro and M. B. Wilk. An analysis of variance test for normality (complete samples). Biometrika, 52:591–611, 1965
1965
-
[28]
Sta- tistical tests based on r´ enyi entropy estimation
Mehmet Siddik Cadirci, Dafydd Evans, Nikolai Leonenko, and Oleg Seleznjev. Sta- tistical tests based on r´ enyi entropy estimation. arXiv e-prints , pages arXiv–2106, 2021
2021
-
[29]
Random variate generation from multivariate exponential power distribution
Nadia Solaro et al. Random variate generation from multivariate exponential power distribution. Statistica & Applicazioni , 2(2):25–44, 2004
2004
-
[30]
The unique non self-referential q-canonical distribution and the phys- ical temperature derived from the maximum entropy principle in tsallis statistics
Hiroki Suyari. The unique non self-referential q-canonical distribution and the phys- ical temperature derived from the maximum entropy principle in tsallis statistics. Progress of Theoretical Physics Supplement , 162:79–86, 2006
2006
-
[31]
Possible generalization of boltzmann-gibbs statistics
Constantino Tsallis. Possible generalization of boltzmann-gibbs statistics. Journal of statistical physics , 52:479–487, 1988. 20
1988
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.