{"id":"3b12ca5f-e05d-4d73-996f-e1545ed61248","arxiv_id":"2506.14242","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper proposes Tsallis-entropy goodness-of-fit tests for q-Gaussian data, but the test statistics are under-defined and the advertised iterative shape-parameter estimator is absent.","lead":"A statistics preprint proposes goodness-of-fit tests for multivariate q-Gaussian distributions using Tsallis entropy estimated from nearest-neighbor distances. The tests are meant to check distributional assumptions for heavy-tailed or compactly supported data, but key formulas are incomplete as written.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The statistic H_upper_q in eqs. (11)–(12) is never derived as the maximum Tsallis entropy, is dimensionally inconsistent, and does not match the paper's own entropy formula in §3.3; the claimed Q → 0 under H0 is therefore unsupported.","rationale":"The paper’s central claim is that QTsallis_N,k and QTsallis*_N,k are consistent goodness-of-fit statistics. The only argument for consistency under H0 is Slutsky’s theorem combined with the assertion that H_upper_q is the maximum Tsallis entropy of the null distribution. That assertion is unsupported and internally inconsistent: T1 and T2 are never defined; H_upper_q adds a density function to a scalar, making eqs. (11)–(12) not well-defined statistics as written; and the expression (1/2) log|Σ| + T1 does not match the paper’s own exact computation of Tsallis entropy for the q-Gaussian in §3.3. Without a correct H_upper_q, the claimed limit of 0 under H0 does not follow, and the Monte Carlo critical values in Table 1 lack a valid null interpretation. This is more fundamental than the absence of size/power simulations or the limited novelty relative to Rényi entropy: the proposed test statistic itself must be redefined before any empirical evaluation can be meaningful. The reader’s weakest_assumption identified this same issue, and I agree. Since the reader’s verdict is REJECT and this concern supports rejection, no adjustment to the verdict is needed.","tokens_in":10112,"tokens_out":4260,"duration_ms":40120,"concrete_test":"Independently re-derive H_upper_q for the null model. Take the multivariate q-Gaussian density from §3.3, plug it into the definition of Tsallis entropy, eq. (4), and evaluate the integral to obtain the exact S_q(f). Then compare this expression with (1/2) log|ΣTN| + T1(x; a, q, σ) as written in eq. (11). If the two disagree — in particular, if the |Σ|-dependence is a power law rather than log|Σ| — then eq. (11) does not converge to 0 under H0 and the claimed consistency fails. A quick computational check: for m = 2, q = 1.5, simulate N = 5000 from the null q-Gaussian, compute QTsallis_N,k using the paper’s formula, and compare the empirical mean to S_q(f) − TS_k,N,q; if the mean is not near 0, the statistic is mis-centered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equations (11) and (12) define QTsallis_N,k = H_upper_q − TS_k,N,q with H_upper_q = (1/2) log|ΣTN| + T1(x; a, q, σ) (resp. T2), described as “the maximum Tsallis entropy under the assumed model.” This identification is load-bearing for the central consistency claim in §5.2, but it is never proved and appears incorrect. Three concrete defects: (i) neither T1 nor T2 is defined anywhere in the manuscript; §5 merely says they are distributions in class K, so the null hypothesis is not identifiable. (ii) If T1(x; ·) is a density evaluated at x, then H_upper_q depends on one observation x while TS_k,N,q is a scalar, so QTsallis_N,k is not a well-defined statistic; if T1 is meant to be a constant, it is never specified and no derivation is given. (iii) The formula does not match the paper’s own Tsallis entropy computation in §3.3, which gives H_q(G_{m,q}) = (1/(q−1))(1 − C_q^q |Σ|^{(1−q)/2} [2^{m/2}Γ(m/2)/((1−q)^{m/2}Γ(m/2 + 1/(q−1)))]) for the multivariate q-Gaussian. This is a power law in |Σ|, not (1/2) log|Σ| plus a scalar. Hence H_upper_q is not the Tsallis entropy of the null density, and the asserted limit Q → 0 under H0 does not follow. The Monte Carlo tables estimate critical values for a statistic whose null mean is an uncomputed, generally nonzero constant; a decision based on those critical values would not test the stated null hypothesis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes goodness-of-fit tests for multivariate generalized Gaussian and q-Gaussian distributions based on Tsallis entropy. The test statistics QTsallis_N,k and QTsallis*_N,k are defined in Eqs. (11) and (12) as the difference between a claimed maximum Tsallis entropy H_upper_q of the null model and a k-nearest-neighbor estimator of Tsallis entropy. The paper claims, in §5.2, that these statistics converge in probability to 0 under the null and to a positive constant under alternatives, and it presents Monte Carlo critical values and convergence plots for various parameter settings.","tokens_in":10532,"tokens_out":4654,"duration_ms":43699,"significance":"If the proposed statistics were well-defined and the asymptotic claims correct, the paper would offer a genuinely useful entropy-based approach to multivariate goodness-of-fit testing for heavy-tailed and compactly supported distributions, complementing likelihood-based methods. The paper also provides a large set of Monte Carlo tables and graphs, which is a useful resource if the underlying statistics are valid. However, the central construction is not mathematically well-defined: the quantities T1 and T2 are never defined, the expression for H_upper_q is dimensionally inconsistent, and the claimed limit Q → 0 under H0 is asserted without a derivation and contradicts the paper's own §3.3 entropy formula. These defects undermine the core contribution, so the potential significance is not realized in the current version.","major_comments":[{"comment":"The test statistics are not well-defined. H_upper_q is given as (1/2) log|Σhat_N| + T1(x; a, q, σ) (resp. T2), but T1 and T2 are never defined anywhere in the manuscript; §5 only says they are distributions in class K. If T1(x; a, q, σ) is a density evaluated at a point x, then H_upper_q depends on a single observation x, whereas the k-NN estimator Rhat_TS_k,N,q is a scalar computed from the full sample; the difference is therefore not a meaningful statistic. If T1 is instead meant to be a constant, its value and its derivation are absent. The null hypotheses 'X ∼ T1(x; a, q, σ)' and 'X ∼ T2(x; a, q, σ)' are thus not identifiable from the text.","section":"§5.1, Eqs. (11) and (12)"},{"comment":"The claimed convergence Q → 0 under H0 is unsupported. H_upper_q is asserted to be the maximum Tsallis entropy of the assumed model, but this identification is never proved and is inconsistent with the paper's own computation in §3.3, where the Tsallis entropy of the multivariate q-Gaussian is given as a power-law expression in |Σ|: H_q = 1/(q−1)(1 − C_q^q |Σ|^{(1−q)/2} 2^{m/2}Γ(m/2)/((1−q)^{m/2}Γ(m/2 + 1/(q−1)))). This is not (1/2) log|Σ| plus a constant/density term. Consequently, the equality H_upper_q = S_q under the null, on which the limit in §5.2 rests, does not follow from the manuscript. The Monte Carlo critical values are therefore calibrated for a statistic whose null mean is an uncomputed, generally nonzero constant.","section":"§5.2 and §3.3"},{"comment":"The asymptotic justification mixes incompatible parameter regimes. Theorem 3 is stated only for q ∈ (0,1), yet the first test statistic (11) is proposed for q ∈ (1,3). Remark 4 cites a different consistency result for q ∈ (1,(k+1)/2), but this range excludes many configurations used in the simulations (e.g., k=1 with q=2.5 or 3.0), and Remark 4 is not integrated into the main consistency argument. Thus the statement that 'By Theorem 3, the distributions T1 and T2 are included in this class' does not cover the parameter settings actually used.","section":"§5 and Theorem 3"},{"comment":"Definition 1 misstates the r-th moment. It defines K_r(f) = E(||X||^r) = 1/(q−1) ∫ ||x||^r f^q(x) dx, but the left-hand side E(||X||^r) equals ∫ ||x||^r f(x) dx, which is not equal to the right-hand side. The subsequent critical moment r_c(f) and the moment conditions (9) and (10) rely on this quantity, so the conditions as stated are ambiguous or incorrect.","section":"Definition 1"},{"comment":"The numerical results do not support the claimed convergence to zero. Table 1 reports 5% critical values of Q_T_N,k(m,q) that remain around 0.03 across all sample sizes N from 100 to 1000 and all parameter settings, rather than decreasing to 0. Table 2 reports slopes β in a log-log regression of |E[Q]| on N that are mostly positive or near zero, which is inconsistent with the statement that E[Q] → 0 as N → ∞. The sentence 'Furthermore numerical applications are proof of the theoretical idea that E[Q_T_N,k(m,q)] → 0' is also methodologically circular: simulations cannot prove an asymptotic limit, and the tabulated values do not exhibit the claimed behavior.","section":"§6, Table 1 and Table 2"}],"minor_comments":[{"comment":"The term 'non-parametric' is used loosely: the test statistics require estimation of the covariance matrix Σ and refer to specific parametric null families (q-Gaussian and generalized Gaussian). A more precise wording would avoid the implication that the tests are fully distribution-free.","section":"Abstract and §1"},{"comment":"The organization paragraph omits Section 4 ('Tsallis Entropy: Statistical Estimation Method') and says 'Section 5 introduces entropy-based goodness-of-fit test statistics', which is correct, but the list skips from Section 3 to Section 5; this should be corrected.","section":"§1 (last paragraph)"},{"comment":"The normalization constant in the multivariate exponential power density appears to lack a factor of 2^{m/s} in the denominator relative to standard forms, and the use of Γ(m/s + 1) versus Γ(m/s) should be checked against the cited literature.","section":"Eq. (5)"},{"comment":"The caption states that 'decreasing q sharpens the peak and widens the tails', but for q in the heavy-tailed regime (q < 1), the displayed behavior may be non-monotonic in q; a precise statement of which q range is being described would improve clarity.","section":"Figure 2 caption"},{"comment":"Reference [5] and the first part of the bibliography entry [4] appear to refer to the same article (Berrett, Samworth, and Yuan); duplicate entries should be merged or clearly distinguished.","section":"References"}],"recommendation":"reject","confidential_remarks":"The central test statistics in Eqs. (11) and (12) are undefined, the claimed maximum entropy identification is not derived and contradicts the paper's own formula in §3.3, and the simulations do not exhibit the asserted convergence to zero. These are load-bearing defects that cannot be fixed by simple editing; the paper would need a fundamental re-derivation of the entropy quantities and a redefinition of the test statistics. In addition, the use of Theorem 3 for q ∈ (1,3) is outside the stated scope, and Definition 1 contains a false equality. The manuscript is not publishable in its current form, and the path to a publishable version is not a local revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This one isn't ready for refereeing. The central test statistic is not well-defined, and the paper advertises an iterative shape-parameter estimator that never appears in the text. I agree with the skeptical report, but let me give credit first: the paper correctly cites the k-NN estimator literature, includes standard Tsallis entropy formulas for the q-Gaussian, and the Monte Carlo plots at least show the statistic converges as N grows. That part is fine. The problems are load-bearing. Equations (11) and (12) define H_upper^q as (1/2) log|Sigma_hat| + T1(x; a, q, sigma), and call it the maximum Tsallis entropy under the null. T1 and T2 are never defined, so the null hypothesis is not identifiable. If T1 is a density evaluated at x, then H_upper^q depends on one observation while the k-NN estimator is a sample average, so the statistic is not a statistic. If T1 is a constant, no derivation is given. The stress-test note is right that this doesn't match the paper's own entropy formula in Section 3.3, which gives a power law in |Sigma|, not (1/2) log|Sigma| plus a scalar. So the claimed Q to 0 under H0 does not follow. This is a hard defect, not a stylistic quibble. There is also a novelty problem. Tsallis entropy is a monotone transform of Rényi entropy for any fixed q, so tests based on the k-NN estimator of one are equivalent to tests based on the other in terms of ordering. The paper doesn't acknowledge this, and it means the proposed statistics carry no new testing information over existing Rényi-based tests. The consistency argument also has a range mismatch: Theorem 3 is stated for q in (0,1), but the first test uses q in (1,3), with only a remark citing Leonenko-Pronzato-Savani for that range. The numerical section never evaluates size or power. The critical value tables are calibrated under the null, but without alternatives you cannot tell whether the test detects anything. The abstract promises an iterative shape-parameter estimator; it is absent from the body. That alone would justify major revision. Who is this for? A reader interested in entropy-based GoF might use the literature review, but the paper as it stands does not support its claims. I would desk-reject, not send to referees, because the statistic is undefined and the central claim is unsupported. If the author repairs the definitions and runs real power studies, it could become a modest incremental paper, but that would be a new submission.","headline":"The test statistics are not well-defined, the advertised estimator is missing, and the work adds little over Rényi-based tests; not ready for refereeing.","tokens_in":754,"tokens_out":960,"would_cite":false,"duration_ms":24988,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62B10","62H15","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a Tsallis-entropy test statistic built from nearest-neighbour estimates converges to 0 under the null model and to a positive constant otherwise, giving a non-parametric goodness-of-fit test.","keywords":["Tsallis entropy","goodness-of-fit tests","k-nearest neighbour estimator","q-Gaussian distribution","generalized Gaussian distribution","maximum entropy principle","Monte Carlo critical values","multivariate heavy tails"],"falsifier":"Simulate data from a multivariate q-Gaussian with known q and covariance, compute the closed-form Tsallis entropy given in Section 3 of the paper, form the test statistic with that value in place of the unexplained $T(x; a, q, \\sigma)$, and check whether the Monte Carlo average tends to 0 as $N$ grows; if it does not, the claimed consistency under the null is false.","tokens_in":9844,"feed_emoji":"📊","tokens_out":15338,"duration_ms":124559,"temperature":0.7,"pith_summary":"This paper proposes goodness-of-fit tests for multivariate generalized Gaussian and q-Gaussian distributions based on Tsallis entropy, a one-parameter generalization of classical differential entropy. The test statistic compares a k-nearest-neighbour estimate of the sample's Tsallis entropy with the maximum Tsallis entropy of the hypothesized model. The central claim is that this difference converges in probability to 0 when the data follow the null distribution and to a positive constant otherwise, so that Monte Carlo critical values yield a practical significance test. Simulations are reported showing convergence of the statistic, stability across neighbourhood sizes and dimensions, and an approach to normality as the q parameter nears 1. If the claim holds, the tests give a density-estimation-free way to check heavy-tailed or compactly supported multivariate fits.","feed_headline":"Entropy-based tests converge to zero under the right model","feed_subtitle":"Nearest-neighbour Tsallis entropy estimates plus Monte Carlo thresholds test heavy-tailed multivariate fits.","key_machinery":"The object that carries the argument is the Tsallis entropy functional $H_q(f) = \\frac{1}{1-q}\\left(\\int_{\\mathbb{R}^m} f^q(x)\\,dx - 1\\right)$, paired with the $k$-nearest-neighbour estimator $\\hat{S}_{k,N,q} = \\frac{1}{N}\\sum_{i=1}^N (\\zeta_{i,k,N})^{1-q}$, where $\\zeta_{i,k,N}$ is built from the Euclidean distance to the $k$th nearest neighbour. The test statistic is the gap between that estimate and the claimed maximum-entropy value of the hypothesized distribution, so the mechanism that converts entropy estimation into hypothesis testing is the identification of $\\tfrac{1}{2}\\log|\\hat\\Sigma_N| + T(x; a, q, \\sigma)$ as that maximum. The convergence argument then depends on the moment conditions in Theorem 3 that guarantee $\\hat{S}_{k,N,q}$ converges.","core_discovery":"The paper claims to establish a class of test statistics of the form $Q_{N,k}^{\\mathrm{Tsallis}} = H_q^{\\mathrm{upper}} - \\hat{S}_{k,N,q}$, where $\\hat{S}_{k,N,q}$ is the $k$-nearest-neighbour estimator of Tsallis entropy and $H_q^{\\mathrm{upper}} = \\tfrac{1}{2}\\log|\\hat\\Sigma_N| + T(x; a, q, \\sigma)$ is meant to be the maximum Tsallis entropy of the assumed model. The argument is that the covariance estimate converges to the true covariance, the nearest-neighbour entropy estimator converges to the true Tsallis entropy under the moment conditions of Theorem 3, and so by a standard convergence argument $Q$ converges in probability to 0 under the null and to a positive constant $c>0$ otherwise. This 0-versus-positive separation is what turns an entropy estimate into a goodness-of-fit test. The simulation study reports convergence of the statistic, stable critical values, and an empirical approach to normality as $q\\to 1$.","pith_inferences":["Editorial inference: A direct check the reader could run is to replace the unexplained $T(x; a, q, \\sigma)$ term in $H_q^{\\mathrm{upper}}$ with the closed-form Tsallis entropy of the q-Gaussian derived in Section 3, and see whether the simulated statistic still converges to 0 under the null; the paper's own maximum-entropy discussion suggests this substitution is what the term is meant to be.","Editorial inference: If the maximum-entropy identification fails, the statistic no longer has a clean 0-versus-positive interpretation; it would reduce to comparing a nearest-neighbour entropy estimate with a constant, which is a weaker model-checking heuristic rather than a calibrated goodness-of-fit test.","Editorial inference: The same construction could be lifted to other maximum-entropy families by replacing the closed-form entropy term with any correctly derived maximum-entropy expression, but the consistency proof would need to be re-run because it depends on the specific moment and support conditions.","Editorial inference: A testable extension is an adaptive choice of the neighbourhood size k, since the paper notes sensitivity to k and a data-driven rule (for instance, minimising a variance-bias proxy of $\\hat{S}_{k,N,q}$) could make the test fully automatic."],"forward_implications":["If the consistency claim is correct, the tests offer a non-parametric check of whether multivariate data follow a q-Gaussian or generalized Gaussian law, with no need to estimate the density before testing.","The tabulated Monte Carlo critical values give immediate 5% thresholds for dimensions 2 and 3 and sample sizes up to 1000, so the method is usable without deriving an analytic null distribution.","Because the statistic converges to a positive constant under alternatives, the tests can in principle detect departures in tail weight or shape parameter q, not just location-scale shifts.","The empirical approach to normality as q approaches 1 suggests that near-Gaussian nulls can be calibrated with standard normal quantiles rather than a full Monte Carlo table.","The estimator's moment conditions tie practical applicability to tail behavior: convergence requires finite q-weighted moments, so extremely heavy-tailed alternatives may need larger samples."],"supporting_citations":[{"why":"Supplies the k-nearest-neighbour entropy estimator that the proposed test statistic is built from.","marker":"[20]"},{"why":"States the consistency theorem with moment conditions that the paper relies on for convergence of the entropy estimator.","marker":"[28, 8]"},{"why":"Gives the maximum-entropy result for Tsallis statistics used to justify the form of the maximum-entropy term.","marker":"[15]"},{"why":"Defines the generalized Gaussian target family and earlier entropy-based testing for it, which the proposed tests extend.","marker":"[9]"},{"why":"Supplies the law of large numbers for nearest-neighbour distances used in the asymptotic convergence step.","marker":"[24]"},{"why":"Provides the normality test used in the simulations to assess the Gaussian approximation of the statistic.","marker":"[27]"}],"fun_headline_variants":["Tsallis entropy tests zero in on correct model","Nearest-neighbour Tsallis entropy spots wrong fits","New entropy test separates null from alternative","Goodness-of-fit from Tsallis entropy gap","Heavy-tailed fit test via Tsallis entropy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole test depends on the assumption that the paper has correctly identified a formula for the maximum Tsallis entropy of the null model, but the formula it uses is never defined and, as written, does not even have matching units, so if that identification is wrong the claimed convergence to 0 under the null collapses.","fun_headline_variants_meta":{"raw":{"variants":["Tsallis entropy tests zero in on correct model","Nearest-neighbour Tsallis entropy spots wrong fits","New entropy test separates null from alternative","Goodness-of-fit from Tsallis entropy gap","Heavy-tailed fit test via Tsallis entropy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1389,"prompt_tokens":878,"completion_tokens":511,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":441}},"tokens_in":494,"tokens_out":511,"duration_ms":5404,"temperature":1.0,"reasoning_tokens":441,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:17:57.199898+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data from a multivariate q-Gaussian with known q and covariance, compute the closed-form Tsallis entropy given in Section 3 of the paper, form the test statistic with that value in place of the unexplained $T(x; a, q, \\sigma)$, and check whether the Monte Carlo average tends to 0 as $N$ grows; if it does not, the claimed consistency under the null is false.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the k-nearest-neighbour entropy estimator that the proposed test statistic is built from."},{"cited_title":"On the maximum entropy principle and the minimization of the fisher information in tsallis statistics","cited_arxiv_id":null,"evidence_quote":"Gives the maximum-entropy result for Tsallis statistics used to justify the form of the maximum-entropy term."},{"cited_title":"Entropy-based test for generalised gaussian distributions","cited_arxiv_id":null,"evidence_quote":"Defines the generalized Gaussian target family and earlier entropy-based testing for it, which the proposed tests extend."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the law of large numbers for nearest-neighbour distances used in the asymptotic convergence step."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the normality test used in the simulations to assess the Gaussian approximation of the statistic."}],"review_version":1}