REVIEW 4 major objections 3 minor 26 references
Ball-codifference, a moment-free dependency measure, is proposed as a sure-screening statistic for heavy-tailed ultra-high-dimensional predictors.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 09:15 UTC pith:53NYB3RB
load-bearing objection New statistic worth a look, but the sure-screening proof is conditional and the data example is misreported in the paper's own tables. the 4 major comments →
Ball-Codifference Screening for Heavy-Tailed Predictors
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that dependence for screening can be measured by the Ball-codifference: the squared difference between the joint probability of a random product ball and the product of its marginal probabilities, multiplied by a local codifference weight formed from conditional cosine expectations. This quantity is a bounded V-statistic, giving asymptotic normality through empirical-process and functional-delta arguments, and it induces a marginal utility that detects association without requiring finite moments. The paper's main theorem states that, under a population separation condition and a uniform empirical-error condition, the BCodifCor-SIS procedure retains every active predicto
What carries the argument
The load-bearing object is the Ball-codifference functional, BCodif_s^2(X,Y) = E[D_{12}^2 κ_{12}], where D_{12} compares joint and marginal probabilities of the random ball A_{12} and κ_{12} is the local codifference inside that ball. The codifference weight is a ratio of characteristic-function-type expectations; like a trigonometric knot, it remains finite under infinite variance. The empirical version rewrites every piece as sums of indicator and cosine terms, making the statistic a bounded V-statistic whose first Hoeffding projection controls its normal limit. This combination is what gives the method its moment-free, model-free character.
Load-bearing premise
The theorem's guarantee depends on every estimated Ball-codifference score being uniformly close to its population value (error o(c_n) with probability tending to one) while active and inactive population scores are separated by at least 2 c_n; the paper asserts this uniform-error condition can be obtained from exponential inequalities but does not prove it for the new statistic.
What would settle it
Simulate p = 1000, n = 150 with independent symmetric α-stable predictors (α = 0.8) and no response association, then track max_r |ω̂_r − ω_r| across replications. If this maximum does not converge to zero at the theorem's required rate — for instance, if heavy-tailed cosine contrasts keep the error large — the uniform-error premise fails and the sure-screening guarantee would not hold.
If this is right
- If the claims hold, practitioners can screen ultra-high-dimensional predictors without first verifying finite moments, extending valid inference to α-stable and other heavy-tailed data.
- The sure-screening theorem means the top-d selection rule will, with probability tending to one, include every active predictor once the uniform-error and separation conditions are met.
- The bounded V-statistic representation yields a normal limit, so standard errors and approximate inference for the screening scores are available under the stated assumptions.
- In the paper's simulations, the codifference-weighted score has higher average retention of strongly associated predictors than Ball-covariance screening across Gaussian and heavy-tailed designs.
- On the riboflavin benchmark, pre-screening with Ball-codifference reduces linear-regression prediction error relative to Ball-covariance pre-screening for most pre-selected model sizes, with the largest reported improvements above 30%.
Where Pith is reading between the lines
- Editorial extension: the bounded-contrast construction is portable; the same 'squared ball discrepancy times local codifference' weighting could be grafted onto other geometric dependence measures, potentially yielding moment-free versions of distance-correlation screening.
- Editorial extension: the sure-screening theorem is conditional on a uniform-error bound that is asserted rather than proved; a central open step is a concentration inequality for max_r |ω̂_r − ω_r| that would make the screening guarantee unconditional.
- Editorial extension: since the statistic only uses metrics and cosines, it should extend to functional or non-Euclidean predictors; a testable next step is comparing Ball-codifference screening on curve or spatial data against competitor screening utilities.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Ball-codifference, a marginal screening utility that combines random-ball geometry with a characteristic-function-based codifference weight, aiming to screen predictors under heavy-tailed distributions where covariance or correlation may not exist. The main theoretical results are an asymptotic normality claim for the empirical Ball-codifference statistic (Theorem 4.2) and a sure-screening property for a proposed BCodifCor-SIS procedure (Theorem 4.3). The numerical section compares BCodifCor-SIS with BCor-SIS on simulated Gaussian and sub-Gaussian stable designs, and a real-data example (Riboflavin) is used to claim improved prediction accuracy. The paper also provides R code for reproducibility.
Significance. If established, a moment-free screening statistic that is robust to heavy tails and still detects dependence would be a useful addition to the high-dimensional screening toolbox. The paper has a plausible heuristic and the supplied code is a positive feature. However, the central theoretical guarantees are conditional on unverified assumptions, and the main empirical claim in the data example is contradicted by the paper's own table. As it stands, the manuscript does not make a convincing case for either the sure-screening property or the practical superiority of the proposed method.
major comments (4)
- [Section 4, Theorem 4.3 and Appendix C] The sure-screening conclusion is derived only from the assumed uniform error condition max_{1≤r≤p}|ω̂_r − ω_r| ≤ ε_n with probability tending to one and the population separation condition. Neither condition is proved for the Ball-codifference utility. The text after the theorem states that the uniform error condition 'can be obtained from exponential inequalities for bounded empirical processes', but no such derivation is supplied. The statistic (13) contains inverse powers of the empirical joint-ball count Δ_{ij}^{XY}, so it is not a bounded kernel in the usual sense; the cited exponential inequalities are therefore not directly applicable. Appendix D only controls the inverse denominators on the event min_{i,j} Δ_{ij}^{XY} ≥ π_0/2, but no uniform lower bound on the local ball probabilities π_ij (defined in (26)) appears in Assumption 4.1 or in Theorem 4.3. Thus the load-bearing premis
- [Appendix D / Theorem 4.2] The claim that the statistic in (13) is a bounded V-statistic is not correct as stated. As shown in equations (24)–(25), the statistic has terms of the form (Δ_{ij}^{XY})^{-1} and (Δ_{ij}^{XY})^{-2} unless one conditions on all local denominators being bounded away from zero. Appendix D invokes a lower bound π_ij ≥ π_0 > 0, but Assumption 4.1(iii) only allows a trimming sequence a_n ↓ 0 with n a_n → ∞. The bias introduced by trimming Δ_{ij}^{XY} to Δ_{ij}^{XY} ∨ a_n is not analyzed; when π_ij is of order smaller than a_n, the trimming error need not be o_p(n^{-1/2}). Hence the asymptotic normality result is conditional on unstated additional assumptions, and the proof does not establish the theorem as stated.
- [Section 6, Table 3] The column 'Improvement (%)' is computed as (RMSE_BCodif − RMSE_BCorr)/RMSE_BCorr × 100; for example, ncov=2 gives 0.0192/0.7322 ≈ 2.6%. Thus a positive 'Improvement' means BCodif has a higher RMSE, i.e., worse prediction accuracy. The table has positive values in 26 of 28 rows, meaning BCodifCor-SIS is worse than BCor-SIS in almost all configurations. The text, however, claims 'the BCodifCor-SIS screening method dominates BCor-SIS significantly, except for the two instances (ncov=6,7) of small better performance for BCor-SIS.' In fact, for ncov=6 and 7 the Diffmean is negative, indicating BCodif is better. The abstract's claim that 'our variable screening method significantly improves prediction accuracy' is therefore directly contradicted by the paper's own reported numbers.
- [Section 5, Tables 1 and 2] The statement that 'Ball-codifference dominates Ball-Covariance' is an overstatement based only on the pa (average retention) columns. The exact-rank recovery probabilities pm often favor BCor-SIS; examples include Table 1 (α=0.9, ρ=0.95): pm(X15) is 0.22 for BCodifCor-SIS vs 0.40 for BCor-SIS; Table 2 (ρ=0.95): pm(X10) 0.68 vs 0.80, pm(X5) 0.90 vs 0.94. The text acknowledges that exact-rank recovery is 'more variable', but the broad claim of dominance is not supported. The selective emphasis on pa while de-emphasizing pm without a principled basis weakens the simulation evidence.
minor comments (3)
- [Throughout] There are several typographical errors: 'Ball-codifierence' (Section 1), 'MEPE' instead of 'MSPE' (Section 6), and 'Section 4' in Section 3.1 where the numerical experiments actually appear in Section 5. The phrase 'The second form is used in the numerical experiments in Section 4' should refer to the correct section.
- [Section 5, first paragraph] The simulation setup states 'We generated (Y, X_1, ..., X_{p-1}) with p = 1000' and then says 'The response is the first component, Y = Z_1'. This is a notational inconsistency; the number of predictors and the index of the response should be clarified.
- [Section 4, after Theorem 4.3] The sentence 'The second condition can be obtained from exponential inequalities for bounded empirical processes when log p = o(n c_n^2)' is vague. Even if a proof were supplied, the phrase 'the exact rate depends on the metric entropy' leaves the condition unactionable. At minimum, a reference or a formal statement of the required entropy condition should be given.
Circularity Check
No significant circularity: the main claims are conditional theorems and the key same-author result is reproved in-appendix.
full rationale
The derivation chain is not circular. The new Ball-codifference statistic is explicitly constructed (Eqs. 11-13) from ball indicators and bounded cosine contrasts, so the moment-free property follows directly from the definition rather than being imported from a fit or from a prior result. The only same-author citation, Maroufy et al. (2025), supplies the extended codifference, but the load-bearing independence characterization used in the paper is proved in Appendix A (Lemma 2.1) via the spectral-measure representation, so it does not rest on an unexamined self-citation. Theorem 4.3 is a conditional sure-screening statement: it assumes a uniform empirical-error event and a population separation, and the proof in Appendix C derives P(A ⊆ Â_n) → 1 from those assumptions. It does not claim to prove the uniform-error condition for the Ball-codifference statistic; the sentence that this condition 'can be obtained from exponential inequalities for bounded empirical processes' is an unsubstantiated assertion, and Appendix D openly notes that the statistic is a V-statistic only 'apart from the inverse powers of Δ^XY_ij'. These are genuine technical gaps and correctness risks, as is the sign-reversed reading of Table 3's 'Improvement' column. But a missing proof of a condition, or a mislabeled numerical comparison, is not a circular reduction of the result to its inputs. The central theoretical statements are logical implications from stated, albeit partially unverified, assumptions. Hence the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (1)
- Trimming sequence a_n =
unspecified
axioms (5)
- domain assumption The class of product balls B_X(x1,x2) × B_Y(y1,y2) is Glivenko-Cantelli and Donsker under F (Assumption 4.1(ii)).
- domain assumption Local ball probabilities are bounded away from zero or are handled by a trimming sequence a_n (Assumption 4.1(iii)).
- domain assumption Population utilities are separated, min_{r∈A} ω_r − max_{r∈M} ω_r ≥ 2c_n, and the uniform empirical error is bounded by ε_n = o(c_n) (Theorem 4.3).
- domain assumption The variables are symmetric stable (or symmetric) so that the cosine-based estimators (8)-(9) are valid (Section 2).
- domain assumption The first Hoeffding projection of the limiting kernel has positive finite variance (Assumption 4.1(iv)).
read the original abstract
High-dimensional screening is commonly built on covariance, correlation, or least-squares measures. These summary measures can be unstable or even undefined, when predictors are sparse or have heavy-tailed distributions. Building on our recent work on extended codifference and the idea of Ball-covariance, we develop Ball-codifference for marginal screening in statistical modeling with heavy-tailed predictors and responses. The proposed statistic combines the rank-type geometry of random balls with the codifference as a dependency measure constructed based on the characteristic function, so it can be computed without requiring well-defined finite first or second moments. We define Ball-codifference and its normalized screening utility, and formulate a sure independence screening procedure. Large-sample normality follows from a bounded V-statistic and functional-delta-method argument under standard nondegeneracy and regularity conditions. Simulation studies under Gaussian and sub-Gaussian stable designs show that codifference-weighted Ball screening gives competitive or improved recovery of highly associated predictors, especially when tail heaviness is pronounced. Also, our data example illustrates that our variable screening method significantly improves prediction accuracy in linear regression.
Reference graph
Works this paper leans on
-
[1]
and Nolan, J
Alparslan, U. and Nolan, J. P. (2016). Measure of dependence for stable distributions.Extremes19, 303–323
2016
-
[2]
and Afhami, B., (2025)
Maroufy, V., Rezapour, M., Zarepour, M. and Afhami, B., (2025). On Measure of Association for Heavy-tailed Random Variables.Statistica Sinica, 38(4)
2025
-
[3]
and Z¨ ahle, H
Beutner, E. and Z¨ ahle, H. (2010). A modified functional delta method and its application to the estimation of risk functionals.Journal of Multivariate Analysis101, 2452–2463
2010
-
[4]
and Z¨ ahle, H
Beutner, E. and Z¨ ahle, H. (2012). Deriving the asymptotic distribution of U- and V-statistics of dependent data using weighted empirical processes.Bernoulli18, 803–822
2012
-
[5]
and Z¨ ahle, H
Beutner, E. and Z¨ ahle, H. (2014). Continuous mapping approach to the asymptotics of U- and V-statistics.Bernoulli20, 846–877
2014
-
[6]
and Paulauskas, V
Damarackas, J. and Paulauskas, V. (2014). Properties of spectral covariance for linear processes with infinite variance.Lithuanian Mathematical Journal54, 252–276
2014
-
[7]
and Paulauskas, V
Damarackas, J. and Paulauskas, V. (2017). Spectral covariance and limit theorems for random fields with infinite variance.Journal of Multivariate Analysis153, 156–175
2017
-
[8]
and Li, R
Fan, J. and Li, R. (2001). Variable selection via nonconcave penalized likelihood and its oracle properties.Journal of the American Statistical Association96, 1348–1360
2001
-
[9]
and Lv, J
Fan, J. and Lv, J. (2008). Sure independence screening for ultra-high dimensional feature space. Journal of the Royal Statistical Society: Series B70, 849–911
2008
-
[10]
and Kodia, B
Garel, B. and Kodia, B. (2009). Signed symmetric covariation coefficient for alpha-stable dependence modeling.Comptes Rendus Mathematique347, 347–352
2009
-
[11]
and Garel, B
Kodia, B. and Garel, B. (2014). Estimation and comparison of signed symmetric covariation coefficient and generalized association parameter for alpha-stable dependence modeling.Communications in Statistics–Theory and Methods43, 5156–5174. 16
2014
-
[12]
Kokoszka, P. S. and Taqqu, M. S. (1993). Asymptotic dependence of moving average type self-similar stable random fields.Nagoya Mathematical Journal130, 85–100
1993
-
[13]
and Lv, J
Kong, Y., Li, D., Fan, Y. and Lv, J. (2017). Interaction pursuit in high-dimensional multi-response regression via distance correlation.The Annals of Statistics45, 897–922
2017
-
[14]
and Gajda, J
Kruczek, P., Wy loma´ nska, A., Teuerle, M. and Gajda, J. (2017). The modified Yule-Walker method forα-stable time series models.Physica A469, 588–603
2017
-
[15]
Levy, J. B. and Taqqu, M. S. (2012). On the codifference of linear fractional stable motion.Seminaire et Congres28, 89–113
2012
-
[16]
Levy, J. B. and Taqqu, M. S. (2014). The asymptotic codifference and covariation of log-fractional stable noise.Journal of Econometrics181, 34–43
2014
-
[17]
and Zhu, L
Li, R., Zhong, W. and Zhu, L. (2012). Feature screening via distance correlation learning.Journal of the American Statistical Association107, 1129–1139
2012
-
[18]
and Afhami, B
Maroufy, V., Rezapour, M., Zarepour, M. and Afhami, B. (2025). On measure of association for heavy-tailed random variables.Statistica Sinicapreprint SS-2025-0370
2025
-
[19]
Nolan, J. P. (2016).Stable Distributions: Models for Heavy Tailed Data. Springer, New York
2016
-
[20]
and Zhu, H
Pan, W., Wang, X., Xiao, W. and Zhu, H. (2019). A generic sure independence screening procedure. Journal of the American Statistical Association114, 928–937
2019
-
[21]
Paulauskas, V. (1976). Some remarks on multivariate stable distributions.Journal of Multivariate Analysis6, 356–368
1976
-
[22]
Paulauskas, V. (2013). On α-covariance, long, short, and negative memories for sequences of random variables with infinite variance. arXiv preprint
2013
-
[23]
Press, S. J. (1972). Multivariate stable distributions.Journal of Multivariate Analysis2, 444–462
1972
-
[24]
Rosadi, D. (2006). Order identification for Gaussian moving averages using the codifference function. Journal of Statistical Computation and Simulation76, 553–559
2006
-
[25]
and Taqqu, M
Samorodnitsky, G. and Taqqu, M. S. (1994).Stable Non-Gaussian Random Processes: Stochastic Models with Infinite Variance. Chapman & Hall/CRC, Boca Raton
1994
-
[26]
Tibshirani, R. (1996). Regression shrinkage and selection via the lasso.Journal of the Royal Statistical Society: Series B58, 267–288. Wy loma´ nska, A., Chechkin, A., Gajda, J. and Sokolov, I. M. (2015). Codifference as a practical tool to measure interdependence.Physica A421, 412–429. 17
1996
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.