REVIEW 3 major objections 4 minor 40 references
A Modified Dependence Measure Related to Chatterjee's Rank Correlation: Theoretical Properties and Asymptotic Analysis
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper proposes a nonnegative modification of the DSS dependence measure and derives an asymptotic chi-square-mixture null limit for its estimator.
desk verdict A well-intentioned reformulation of Chatterjee's correlation whose central null distribution theorem is false. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the ratio $\xi_+(X,Y) = \int \operatorname{Var}(E(1\{Y>t\}|X))\,dF_Y(t) \,/\, \int \operatorname{Var}(1\{Y>t\})\,dF_Y(t)$, written through right-continuous conditional CDFs. Its squared-distance representation, $\xi_+ = \int E[(F_{Y|X}(t)-F_Y(t))^2]\,dF_Y(t) \,/\, \int [F_Y(t)(1-F_Y(t))]\,dF_Y(t)$, connects the measure to the mean-variance index and lets the empirical version inherit U-statistic asymptotics: the categorical plug-in estimator's identity with six times that index is what transfers the chi-square-mixture null limit, and the same representation provides the influence function behind the non-null normal limit.
What would settle it
Compute equation (3.5) when $X$ has a single category ($R=1$) and $Y$ is continuous: the proposed estimator equals $3/n + 1/n^2$, while the mean-variance index is identically $0$; $n$ times the estimator therefore tends to $3$, not to the claimed null limit $0$. This one calculation is enough to test whether the premise of Theorem 3.2 is valid.
Extended reading notes
Core claim
The central discovery, stated on the paper's own terms, is that replacing the left-continuous CDF in the DSS measure by the right-continuous CDF yields a measure $\xi_+$ that coincides with the original at the extremes—zero exactly under independence and one exactly when $Y$ is a measurable function of $X$—while satisfying cleaner analytic conventions. For a categorical $X$ with $R$ classes, the paper identifies $\xi_+$ with six times the mean-variance index, and uses the known null distribution of that index to claim that under $H_0$, $n\xi_{+,n}$ converges in distribution to $6\sum_{j=1}^\infty \chi^2_j(R-1)/(\pi^2 j^2)$. It further claims strong consistency of the plug-in estimator and asymptotic normality with variance $36\sigma_\tau^2$ under fixed alternatives, and presents a parametrized family $\xi_{k,l,r,s}$ together with a conditional-dependence analogue.
Load-bearing premise
The null-distribution theorem depends on the proposed estimator being exactly six times an earlier mean-variance index; in the simplest one-category case the two are not equal, so the theorem's premise is the equality itself.
Editorial extensions
If this is right
- Under $H_0$, for fixed $R$, the test statistic $n\xi_{+,n}$ has a non-normal limit that is an infinite mixture of chi-square variables, so critical values can be computed without bootstrapping.
- For alternatives, $\xi_{+,n}$ is consistent and asymptotically normal at $\sqrt{n}$ rate, so confidence intervals for dependence strength follow from a plug-in variance estimator.
- Because $\xi_+$ is nonnegative and vanishes only under independence, the proposed test repairs the finite-sample sign problem of Chatterjee's coefficient, which is negative about half the time under the null.
- The extension family $\xi_{k,l,r,s}$ preserves the 0/1 characterization and can be tuned through $k,l,r,s$ for different distributional sensitivity.
- The conditional version $\xi_+(Y,Z|X)$ avoids left limits and is estimated by plugging out-of-sample conditional CDFs, giving a cleaner asymptotic theory for conditional independence.
Reading between the lines
- If the stated equality with the mean-variance index fails for degenerate category counts, a centered or debiased version of the estimator would be needed to preserve the chi-square-mixture calibration.
- The null limit's dependence on $R$ suggests that permutation or bootstrap critical values could be compared against the mixture quantiles, giving a robustness check that does not rely on the exact equality.
- Because $\xi_+$ is an integrated squared distance between conditional and marginal CDFs, the same construction should extend to multivariate $Y$ or vector-valued $X$ with kernel estimators, where the rank-based baseline does not directly apply.
- A direct simulation with $R=2$ and continuous $Y$, comparing empirical quantiles of $n\xi_{+,n}$ under $H_0$ to the claimed mixture, would reveal the theorem's practical scope at moderate sample sizes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a modified DSS-type dependence measure ξ+(X,Y), gives equivalent representations, constructs sample estimators for categorical and continuous X, and claims strong consistency, an asymptotic null distribution as a scaled infinite chi-square mixture for categorical X, asymptotic normality under fixed alternatives, and an extension to conditional dependence. The central theoretical result is Theorem 3.2, which states that under H0: X⊥Y, nξ+,n converges to 6Σ_j χ²_j(R-1)/(π²j²) for the estimator in (3.5). The paper also reports simulations suggesting improved Type I error control relative to Chatterjee's coefficient.
Significance. If Theorem 3.2 were correct, the paper would offer a nonnegative estimator of a dependence measure with a known null distribution for categorical X, thereby addressing a known drawback of Chatterjee's coefficient. The paper properly credits earlier work (Gamboa, Klein and Lagnoux; Cui and Zhong) rather than creating a citation loop. However, the main asymptotic claim is false: the identity on which the proof rests is algebraically incorrect, so the proposed test is not asymptotically calibrated. The claimed theoretical contribution is therefore not established, and the favorable Type I error statements in the abstract and conclusions are not supported.
major comments (3)
- [Section 3.2, Theorem 3.1(ii)] The proof of strong consistency for continuous X is not a valid proof. It asserts 'by the strong law of large numbers' that R_i/n → F(Y_i) and R_{i,j}/n_j → F(Y_i|X_j), but these are not sums of i.i.d. random variables for fixed indices i,j; the observations Y_i and X_j depend on n, the kernel weights involve the bandwidth h_n, and the denominator n_j can be zero or small. The regularity conditions (C1)–(C4) are not used anywhere in the proof. Consequently, the almost-sure limit for the kernel estimator is unproven, which also undermines Theorem 3.3 and the simulation-based claims for the kernel version of the estimator.
- [Section 3.5, Asymptotic Power] The local alternative claim is unsupported. The paper simply states that under H1n: ξ+ = δ/√n, √nξ+,n → N(δ, σ²), without proving contiguity, deriving the limiting variance under the local sequence, or verifying that the fixed-alternative variance σ² applies. Moreover, the proposed test statistic uses the standardization √nξ+,n/σ̂_n and normal quantiles, whereas under H0 the correct scaling is n (as claimed in Theorem 3.2). Because Theorem 3.2 is false, the description of the level-α test and the power function β(δ) is not justified.
- [Section 6, Table 2 and the Abstract] The simulation study uses the kernel estimator (3.2) with continuous X, but the only null distribution provided in the paper, Theorem 3.2, applies to the categorical estimator (3.5). No null limit is derived for the kernel estimator, so the claimed 'better Type I error control' (Abstract and Section 7) is not supported by any theoretical result. The r=0 rows of Table 2 report means of ξ+,n, yet these are not compared with the claimed mixture distribution or with the correct null limit, so the simulation does not validate the asymptotic calibration of the test.
minor comments (4)
- [Section 2, Proposition 2.2 proof] In the proof of necessity of (iii), the displayed expression '∫ E[(F^2_{Y|X}(t) − F_Y(t)]dF_Y(t) = 0' is a typo; it should be '∫ (E[F_{Y|X}(t)^2] − F_Y(t)^2) dF_Y(t) = 0' (or the equivalent form derived from (2.5)). The subsequent reasoning is understandable but the formula as written is not.
- [Throughout] There are numerous typos and grammatical errors, including 'remak' (Introduction), 'As showed by Chatterjee' (Introduction), 'as follow' (Introduction), and 'the sample estimator of ξ+(X,Y) defined in (2.2)' (Section 3.1), where (2.2) is a population expression. A careful proofreading pass is needed.
- [Sections 3.1–3.5] The same symbol ξ+,n is used for the categorical estimator (3.1)/(3.5), the kernel estimator (3.2), and the conditional-dependence estimator (5.1), without explicit distinction. This creates ambiguity in Theorems 3.2–3.4 and in the simulation section, where the reader must infer which estimator is being discussed.
- [Section 3.4, Theorem 3.4] The statement 'Since ξ+,n = 6τ̂n + o_p(n^{-1/2}) by construction' is not a consequence of the construction; the exact relation is ξ+,n = 6τ̂n + 3/n + 1/n², which is indeed o_p(n^{-1/2}) under H1. The proof should be corrected to state and use the exact relation, rather than an unsupported identity.
Circularity Check
No significant circularity: the Theorem 3.2 issue is a non-circular algebraic mismatch with an external theorem, not a derivation that reduces to its own inputs.
full rationale
The paper's derivation chain does not exhibit circularity. The proposed measure ξ+ is defined directly in Definition 2.1, and the population identity ξ+ = 6τ in equation (3.4) is an algebraic expansion of the mean-variance index of Cui and Zhong, an external result cited from [13]. The estimator in (3.5) is a plug-in estimator, and the proof of Theorem 3.2 attempts to apply Cui-Zhong's Theorem 1 to it. Whether the plug-in estimator actually equals 6τ̂ is a factual algebraic question; as the reader notes, the equality fails because ∫F̂²dF̂ = 1/3 + O(1/n) rather than exactly 1/3, so nξ_{+,n} need not converge to the stated limit for the estimator in (3.5). This is a non-circular mathematical error, or a mismatch between estimators, not a case of a fitted parameter renamed as a prediction, a self-citation chain, or a definition that smuggles in the conclusion. The paper's key external inputs—Cui and Zhong's limit theorem, Chatterjee's consistency result, Gamboa et al.'s representation—are independent of the present paper's own claims, and no load-bearing step reduces to the paper's own output. The paper also contains no relevant self-citations. Therefore the appropriate circularity score is 0; the correctness defect in Theorem 3.2 is a soundness issue, not a circularity issue.
Assumptions & free parameters
free parameters (1)
- Kernel bandwidth h =
1.06 sigma_hat_X n^{-1/5}
assumptions (3)
- standard math Variance decomposition Var(A) = Var(E(A|X)) + E(Var(A|X)) for indicator variables.
- domain assumption Cui and Zhong's Theorems 1 and 2 give the null and alternative asymptotic distributions of the mean-variance index tau_hat.
- domain assumption Under conditions C1-C4, kernel conditional CDF estimators converge uniformly, justifying the SLLN step for the double sum in Theorem 3.1(ii).
Cite this review
Pith. "Pith review of A Modified Dependence Measure Related to Chatterjee's Rank Correlation: Theoretical Properties and Asymptotic Analysis." pith.science (2026). https://pith.science/paper/U6BBU7RN
@misc{pith2026260807844,
author = {Pith},
title = {Pith review of: A Modified Dependence Measure Related to Chatterjee's Rank Correlation: Theoretical Properties and Asymptotic Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/U6BBU7RN}},
note = {Machine review of arXiv:2608.07844}
}
abstract
In his recent breakthrough work [JASA, 2021], Chatterjee proposed a rank-based correlation coefficient $\xi(X,Y)$ to measure the dependence of a random variable $Y$ on $X$. Unlike classical measures such as Pearson, Spearman, or Kendall, $\xi$ satisfies $\xi=0$ if and only if $X$ and $Y$ are independent, and $\xi=1$ if and only if $Y$ is a measurable function of $X$, without requiring monotonicity or linearity. This paper proposes a refined measure of $\xi(X,Y)$ and investigates its theoretical properties. We derive several equivalent representations, construct an estimator, and establish its strong consistency as well as its asymptotic distribution. A natural extension of the proposed framework is also presented. The Monte Carlo simulations show that the proposed estimator outperforms Chatterjee's rank correlation in terms of finite-sample performance, particularly in controlling Type I error under the null.
Reference graph
Works this paper leans on
-
[1]
Auddy, A., Deb, N., & Nandy, S. (2024). Exact detection thresholds and minimax optimality of Chatterjee’s correlation. Bernoulli 30(2), 1056-1079
work page 2024
-
[2]
Ansari, J., Langthaler, Patrick B., Fuchs, S., & Trutschnig, W. (2026). Quantifying and estimating dependence via sensitivity of conditional distributions. Bernoulli 32 (1): 179-204
work page 2026
-
[3]
Ansari, J., & Rockel, M. (2026). The exact region and an inequality between Chatter- jee’s and Spearman’s rank correlations. Journal of Multivariate Analysis 214, 105630
work page 2026
-
[4]
Azadkia, M., & Chatterjee, S. (2021). A simple measure of conditional dependence. Annals of Statistics 49(6), 3070-3102
work page 2021
-
[5]
Bahari, F. (2026). Improvement of Chatterjee’s correlation with missing at random data in theYvariable. Statistical Papers (2026) 67: 56
work page 2026
-
[6]
Bickel, P. J. (2022). Measures of independence and functional dependence. Available at arXiv:2206.13663
arXiv 2022
-
[7]
Cao, S., & Bickel, P. J. (2020). Correlations with tailored extremal properties. Avail- able at arXiv:2008.10177v2
arXiv 2020
-
[8]
Chao, C.C., Bai, Z. & Liang, W.Q. (1993). Asymptotic normality for oscillation of permutation. Probability in the Engineering and Informational Sciences 7, 227-235
work page 1993
Show all 40 references
-
[9]
Chatterjee, S. (2021). A new coefficient of correlation. Journal of the American Sta- tistical Association 116(536), 2009-2022
2021
-
[10]
Chatterjee, S. (2022). A survey of some recent developments in measures of associa- tion. Available at arXiv:2211.04702
2022 arXiv
-
[11]
Chierichetti, F., Giacchini, M., & Kumar, R. (2026). On the metricity of the Chat- terjee correlation coefficient. The American Statistician 80(2), 310-317. 27
2026
-
[12]
Cui, H.J., Li, R.Z., & Zhong, W. (2015). Model-free feature screening for ultrahigh dimensional discriminant analysis. Journal of the American Statistical Association 110, 630-641
2015
-
[13]
Cui, H.J., & Zhong, W. 2019. A distribution-free test of independence based on mean variance index. Computational Statistics and Data Analysis 139, 117-133
2019
-
[14]
Dalitz, C., Arning, G., & Goebbels, S. (2024). A simple bias reduction for Chatterjee’s correlation. Journal of Statistical Theory and Practice 18, 51. https://doi.org/10.1007/s42519-024-00399-y
2024 doi
-
[15]
Deb, N., Ghosal, P., & Sen, B. (2020). Measuring association on topological spaces using kernels and geometric graphs. Available at arXiv:2010.01768v2
2020 arXiv
-
[16]
Dette, H., & Kroll, M. (2025). A simple bootstrap for Chatterjee’s rank correlation. Biometrika 112(1), 2025, asae045, https://doi.org/10.1093/biomet/asae045
2025 doi
-
[17]
Dette, H., Siburg, K. F. & Stoimenov, P. A. (2013). A copula-based non-parametric measure of regression dependence. Scandinavian Journal of Statistics 40(1), 21-41
2013
-
[18]
Gamboa, F., Gremaud, P., Klein, T., & Lagnoux, A. (2022). Global sensitivity analy- sis: A novel generation of mighty estimators based on rank statistics. Bernoulli 28(4), 2345-2374
2022
-
[19]
Gamboa, F., Klein, T., & Lagnoux, A. (2018). Sensitivity analysis based on Cram´ er- von Mises distance. SIAM/ASA J. Uncertain. Quantif., 6(2), 522-548
2018
-
[20]
R., & Trutschnig, W
Griessenberger, F., Junker, R. R., & Trutschnig, W. (2022). On a multivariate copula- based dependence measure and its estimation. Electron. J. Stat., 16(1), 2206-2251
2022
-
[21]
Han, F. (2021). On extensions of rank correlation coefficients to multivariate spaces. Bernoulli 28, 7-11
2021
-
[22]
& Huang, Z
Han, F. & Huang, Z. (2024). Azadkia-Chatterjee’s correlation coefficient adapts to manifold data. The Annals of Applied Probability 34(6), 5172-5210
2024
-
[23]
He, S., Ma, S., & Xu, W. (2019). A modified mean-variance feature-screening proce- dure for ultrahigh-dimensional discriminant analysis. Computational Statistics and Data Analysis 137, 155-169
2019
-
[24]
H¨ ormann, S., & Strenger, D. (2026). Azadkia-Chatterjee’s dependence coefficient for infinite dimensional data. Bernoulli 32(1), 467-492. 28
2026
-
[25]
Huang, Z., Deb, N., & Sen, B. (2022). Kernel partial correlation coefficient-a measure of conditional dependence. Journal of Machine Learning Research 23, 1-58
2022
-
[26]
S., & Borovskich, Y
Korolyuk, V. S., & Borovskich, Y. V. (1994). Theory ofU-Statistics (Mathematics and Its Applications). Dordrecht: Kluwer Academic Publishers
1994
-
[27]
Kroll, M. (2025). Debiased Chatterjee’s correlation: Asymptotic normality without continuity. arXiv:2408.11547
2025 arXiv
-
[29]
Lin, Z., and Han, F. (2023). On Boosting the Power of Chatterjee’s Rank Correlation, Biometrika 110(2), 283-299
2023
-
[30]
Lin Z, & Han F. (2025). Limit theorems of Chatterjee’s rank correlation. arXiv:2204.08031v4
2025 arXiv
-
[31]
Mai, Q., & Zou, H. (2015). The fused kolmogorov filter: A nonparametric model-free screening method. Ann. Statist. 43, 1471-1497
2015
-
[32]
Shi, H., Drton, M., & Han, F. (2022). On the power of Chatterjee’s rank correlation. Biometrika 109(2), 317-333
2022
-
[33]
Shi, H., Drton, M., & Han, F. (2023). On the asymptotic normality of the conditional Chatterjee correlation. Bernoulli(to appear)
2023
-
[34]
Strothmann, C., Dette, H., & iburg, K. F. (2024). Rearranged dependence measures. Bernoulli 30(2), 1055-1078
2024
-
[35]
Wang, L., Li, X.,Wang, X., & Lai, P. (2022). Unified mean-variance feature screening for ultrahigh-dimensional regression. Computational Statistics, 37, 1887-1918
2022
-
[36]
Xia, L, Cao, R, Du, J, & Chen, X. (2025). The improved correlation coefficient of Chatterjee. Journal of Nonparametric Statistics 37(2), 265-281
2025
-
[37]
Yan, X., Tang, N., Xie, J., Ding, X., & Wang, Z. (2018). Fused mean-variance filter for feature screening. Computational Statistics and Data Analysis 122, 18-32
2018
-
[38]
Yang, S. S. (1977). General distribution theory of the concomitants of order statistics. The Annals of Statistics 5 (5), 996-1002. 29
1977
-
[39]
Zhang, Q. (2023). On the asymptotic distribution of the symmetrized Chatterjee’s correlation coefficient. Stat. Probabil. Lett., 194, 109759
2023
-
[40]
Zhang, Q. (2025). On the properties of distance covariance for categorical data: Robustness, sure screening, and approximate null distributions. Scandinavian Journal of Statistics 52(7), 777-804
2025
-
[41]
Zhang, Q. (2026). On the extensions of the Chatterjee-Spearman test. Journal of Nonparametric Statistics 38(2), 347-376. 30
2026
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.