{"id":"1141028b-da0d-4c71-8889-53bd0a0618c5","arxiv_id":"2608.07844","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"The proposed right-continuous variant of the DSS measure is not new for continuous variables, and the theorem giving its null distribution is contradicted by a simple calculation.","lead":"This paper proposes a slightly modified version of Chatterjee's rank-based dependence measure, using right-continuous instead of left-continuous distribution functions, and claims a better finite-sample estimator. The main asymptotic theorem is misstated for the estimator it names, and the simulation evidence for better Type I error control is not supported by the reported numbers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3.2's proof relies on the false identity ξ_{+,n}=6τ̂; for R=1 the estimator in (3.5) gives nξ_{+,n}→3, not the claimed 0, so the null limit theorem is false.","rationale":"I examined the central claim: Theorem 3.2 gives a null limit distribution for the plug-in estimator (3.5). The proof hinges on the equality ξ_{+,n}=6τ̂. I re-derived both expressions. For any empirical CDF, 6τ̂ = 6[Σ p̂_r ∫F̂_r² dF̂ − ∫F̂² dF̂], while (3.5) = 6Σ p̂_r∫F̂_r² dF̂ − 2. Equality requires ∫F̂² dF̂ = 1/3, which fails by an O(1/n) term. For R=1, all conditional empirical CDFs equal the marginal, so (3.5) = 6∫F̂²dF̂−2 = 3/n+1/n², giving nξ→3. The claimed limit for R=1 is 0 because χ²_0 is 0. Hence the theorem is false as stated. This breaks the paper's headline contribution: the asymptotic null distribution used for Type I error claims is not established and, in the R=1 case, is demonstrably wrong. The simulation section uses estimator (3.2), not (3.5), so it does not rescue the theorem. I therefore agree with the reader's assessment, and no adjustment to the verdict is needed.","tokens_in":17425,"tokens_out":9890,"duration_ms":92128,"concrete_test":"Perform the explicit two-point computation: with R=1, n=2, Y=(0,1), compute (3.5) directly: \\hat F(Y_1)=1/2, \\hat F(Y_2)=1, so ξ_{+,n}=6·(1/2)(1/4+1)−2 = 7/4, while 6τ̂=0. Thus the identity used in the proof fails. Alternatively, for general n with continuous Y and R=1, derive nξ_{+,n}=3+1/n and compare with the claimed limit 0; a Monte Carlo with n=1000 will show mean ≈3, confirming that the null distribution theorem is incorrect for the estimator (3.5).","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3's Theorem 3.2 is the central asymptotic result: under H0, n times the plug-in estimator (3.5) is claimed to converge to 6Σ χ²_j(R−1)/(π²j²). The proof applies the continuous mapping theorem to ξ_{+,n}=6τ̂, where τ̂ is the Cui–Zhong MV-index estimator. This equality is algebraically false. With \\hat F_r(t)=\\hat F_{Y|X}(t|x_r) and \\hat F(t)=\\hat F_Y(t)=Σ_r \\hat p_r \\hat F_r(t), one obtains 6τ̂ = 6Σ_r \\hat p_r ∫(\\hat F_r−\\hat F)² d\\hat F = 6[Σ_r \\hat p_r ∫\\hat F_r² d\\hat F − ∫\\hat F² d\\hat F], whereas (3.5) is 6Σ_r \\hat p_r ∫\\hat F_r² d\\hat F − 2. Equality would require ∫\\hat F² d\\hat F = 1/3 exactly; for continuous Y with no ties, ∫\\hat F² d\\hat F = (n+1)(2n+1)/(6n²) = 1/3 + 1/(2n) + 1/(6n²), off by O(1/n) terms that survive multiplication by n. In the simplest case R=1 (which satisfies the theorem's assumptions), \\hat F_1=\\hat F, so (3.5)=6∫\\hat F² d\\hat F−2 = 3/n+1/n² and nξ_{+,n}→3, while the claimed limit (R−1=0) is 0. Thus the theorem is false as stated, and the claimed asymptotic calibration of the test does not hold for estimator (3.5).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a modified DSS-type dependence measure ξ+(X,Y), gives equivalent representations, constructs sample estimators for categorical and continuous X, and claims strong consistency, an asymptotic null distribution as a scaled infinite chi-square mixture for categorical X, asymptotic normality under fixed alternatives, and an extension to conditional dependence. The central theoretical result is Theorem 3.2, which states that under H0: X⊥Y, nξ+,n converges to 6Σ_j χ²_j(R-1)/(π²j²) for the estimator in (3.5). The paper also reports simulations suggesting improved Type I error control relative to Chatterjee's coefficient.","tokens_in":17821,"tokens_out":7668,"duration_ms":74832,"significance":"If Theorem 3.2 were correct, the paper would offer a nonnegative estimator of a dependence measure with a known null distribution for categorical X, thereby addressing a known drawback of Chatterjee's coefficient. The paper properly credits earlier work (Gamboa, Klein and Lagnoux; Cui and Zhong) rather than creating a citation loop. However, the main asymptotic claim is false: the identity on which the proof rests is algebraically incorrect, so the proposed test is not asymptotically calibrated. The claimed theoretical contribution is therefore not established, and the favorable Type I error statements in the abstract and conclusions are not supported.","major_comments":[{"comment":"The proof of strong consistency for continuous X is not a valid proof. It asserts 'by the strong law of large numbers' that R_i/n → F(Y_i) and R_{i,j}/n_j → F(Y_i|X_j), but these are not sums of i.i.d. random variables for fixed indices i,j; the observations Y_i and X_j depend on n, the kernel weights involve the bandwidth h_n, and the denominator n_j can be zero or small. The regularity conditions (C1)–(C4) are not used anywhere in the proof. Consequently, the almost-sure limit for the kernel estimator is unproven, which also undermines Theorem 3.3 and the simulation-based claims for the kernel version of the estimator.","section":"Section 3.2, Theorem 3.1(ii)"},{"comment":"The local alternative claim is unsupported. The paper simply states that under H1n: ξ+ = δ/√n, √nξ+,n → N(δ, σ²), without proving contiguity, deriving the limiting variance under the local sequence, or verifying that the fixed-alternative variance σ² applies. Moreover, the proposed test statistic uses the standardization √nξ+,n/σ̂_n and normal quantiles, whereas under H0 the correct scaling is n (as claimed in Theorem 3.2). Because Theorem 3.2 is false, the description of the level-α test and the power function β(δ) is not justified.","section":"Section 3.5, Asymptotic Power"},{"comment":"The simulation study uses the kernel estimator (3.2) with continuous X, but the only null distribution provided in the paper, Theorem 3.2, applies to the categorical estimator (3.5). No null limit is derived for the kernel estimator, so the claimed 'better Type I error control' (Abstract and Section 7) is not supported by any theoretical result. The r=0 rows of Table 2 report means of ξ+,n, yet these are not compared with the claimed mixture distribution or with the correct null limit, so the simulation does not validate the asymptotic calibration of the test.","section":"Section 6, Table 2 and the Abstract"}],"minor_comments":[{"comment":"In the proof of necessity of (iii), the displayed expression '∫ E[(F^2_{Y|X}(t) − F_Y(t)]dF_Y(t) = 0' is a typo; it should be '∫ (E[F_{Y|X}(t)^2] − F_Y(t)^2) dF_Y(t) = 0' (or the equivalent form derived from (2.5)). The subsequent reasoning is understandable but the formula as written is not.","section":"Section 2, Proposition 2.2 proof"},{"comment":"There are numerous typos and grammatical errors, including 'remak' (Introduction), 'As showed by Chatterjee' (Introduction), 'as follow' (Introduction), and 'the sample estimator of ξ+(X,Y) defined in (2.2)' (Section 3.1), where (2.2) is a population expression. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The same symbol ξ+,n is used for the categorical estimator (3.1)/(3.5), the kernel estimator (3.2), and the conditional-dependence estimator (5.1), without explicit distinction. This creates ambiguity in Theorems 3.2–3.4 and in the simulation section, where the reader must infer which estimator is being discussed.","section":"Sections 3.1–3.5"},{"comment":"The statement 'Since ξ+,n = 6τ̂n + o_p(n^{-1/2}) by construction' is not a consequence of the construction; the exact relation is ξ+,n = 6τ̂n + 3/n + 1/n², which is indeed o_p(n^{-1/2}) under H1. The proof should be corrected to state and use the exact relation, rather than an unsupported identity.","section":"Section 3.4, Theorem 3.4"}],"recommendation":"reject","confidential_remarks":"The central asymptotic result is false in a way that cannot be repaired by local editing: the null distribution of the proposed estimator is not what is claimed, and the test as described is not calibrated. The authors would need to derive the correct null limit for their estimator (or for an appropriately modified estimator) and rework the inference sections before the paper could be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Chuancun Yin's paper is a serious but flawed attempt to fix a small defect in Chatterjee's correlation. The central null distribution theorem is false, and the error is not a gap in the proof but an algebraic identity that doesn't hold. I'd send it back for major revision.\n\nWhat's genuinely good: the paper gives proper credit. It explicitly attributes the population measure to Dette, Siburg, and Stoimenov and notes the equivalent representation from Gamboa, Klein, and Lagnoux. The equivalent representations are clean, and the extension to conditional dependence via plug-in CDF estimates is a reasonable direction. The consistency proof for the kernel estimator (Theorem 3.1) looks plausible under the stated conditions. The simulations are honestly reported with bias and standard deviation tables.\n\nThe problem is Theorem 3.2. The estimator (3.5) is not equal to six times the Cui–Zhong tau-hat. The stress-test note gives a concrete counterexample: for R=1, which the theorem allows, (3.5) equals 3/n + 1/n^2, so n times it converges to 3, while the claimed mixture limit with R-1=0 degrees of freedom is 0. The proof's identity would require ∫F_hat^2 dF_hat = 1/3 exactly; for continuous Y this is 1/3 + O(1/n), and the O(1/n) terms survive multiplication by n. So the null distribution is false as stated. That also invalidates the Type I error claims. The simulation section never computes an actual Type I error rate; it only compares variance under the null, which is not the same as calibration.\n\nThere are smaller issues. Remark 2.1's claim that xi+ = xi = rho^2 for binary Y is not true in general; it holds only when the conditional mean is linear in X (e.g., binary X). Section 3.5's local power claim uses a normal limit under local alternatives that the paper itself says cannot hold under H0; the argument is hand-wavy.\n\nNet: the paper is a sincere contribution with careful attributions and a plausible extension, but the main asymptotic result is wrong. A referee should be asked to verify the identity, because if I'm right, the test is not calibrated and the paper needs major revision. This is worth refereeing, not desk-rejecting, because the error is subtle and the rest of the paper is competent.","headline":"A well-intentioned reformulation of Chatterjee's correlation whose central null distribution theorem is false.","tokens_in":18344,"tokens_out":5860,"would_cite":false,"duration_ms":57440,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G20","62H20","62G10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a nonnegative modification of the DSS dependence measure and derives an asymptotic chi-square-mixture null limit for its estimator.","keywords":["dependence measure","DSS measure","Chatterjee rank correlation","mean-variance index","rank statistics","asymptotic distribution","strong consistency","independence test"],"falsifier":"Compute equation (3.5) when $X$ has a single category ($R=1$) and $Y$ is continuous: the proposed estimator equals $3/n + 1/n^2$, while the mean-variance index is identically $0$; $n$ times the estimator therefore tends to $3$, not to the claimed null limit $0$. This one calculation is enough to test whether the premise of Theorem 3.2 is valid.","tokens_in":17200,"feed_emoji":"📊","tokens_out":9932,"duration_ms":97259,"temperature":0.7,"pith_summary":"The paper proposes a modification of the Dette–Siburg–Stoimenov dependence measure, replacing the left-continuous CDF with a right-continuous one, so the population quantity is well-defined for discrete and mixed distributions. The new coefficient still characterizes independence (value 0) and functional dependence (value 1) exactly, and the paper's plug-in estimator is designed to stay nonnegative, fixing the sign problem of Chatterjee's rank correlation, whose null distribution puts mass on negative values. The paper claims strong consistency for the estimator, an n-rate null limit that is an infinite chi-square mixture, and sqrt-n asymptotic normality under fixed alternatives, plus a family of generalizations. If those claims hold, the estimator gives an independence test with better finite-sample Type I error control than Chatterjee's coefficient.","feed_headline":"Nonnegative rank correlation with a clean chi-square null limit","feed_subtitle":"New estimator stays nonnegative and gives a calibrated limit for testing independence.","key_machinery":"The load-bearing object is the ratio $\\xi_+(X,Y) = \\int \\operatorname{Var}(E(1\\{Y>t\\}|X))\\,dF_Y(t) \\,/\\, \\int \\operatorname{Var}(1\\{Y>t\\})\\,dF_Y(t)$, written through right-continuous conditional CDFs. Its squared-distance representation, $\\xi_+ = \\int E[(F_{Y|X}(t)-F_Y(t))^2]\\,dF_Y(t) \\,/\\, \\int [F_Y(t)(1-F_Y(t))]\\,dF_Y(t)$, connects the measure to the mean-variance index and lets the empirical version inherit U-statistic asymptotics: the categorical plug-in estimator's identity with six times that index is what transfers the chi-square-mixture null limit, and the same representation provides the influence function behind the non-null normal limit.","core_discovery":"The central discovery, stated on the paper's own terms, is that replacing the left-continuous CDF in the DSS measure by the right-continuous CDF yields a measure $\\xi_+$ that coincides with the original at the extremes—zero exactly under independence and one exactly when $Y$ is a measurable function of $X$—while satisfying cleaner analytic conventions. For a categorical $X$ with $R$ classes, the paper identifies $\\xi_+$ with six times the mean-variance index, and uses the known null distribution of that index to claim that under $H_0$, $n\\xi_{+,n}$ converges in distribution to $6\\sum_{j=1}^\\infty \\chi^2_j(R-1)/(\\pi^2 j^2)$. It further claims strong consistency of the plug-in estimator and asymptotic normality with variance $36\\sigma_\\tau^2$ under fixed alternatives, and presents a parametrized family $\\xi_{k,l,r,s}$ together with a conditional-dependence analogue.","pith_inferences":["If the stated equality with the mean-variance index fails for degenerate category counts, a centered or debiased version of the estimator would be needed to preserve the chi-square-mixture calibration.","The null limit's dependence on $R$ suggests that permutation or bootstrap critical values could be compared against the mixture quantiles, giving a robustness check that does not rely on the exact equality.","Because $\\xi_+$ is an integrated squared distance between conditional and marginal CDFs, the same construction should extend to multivariate $Y$ or vector-valued $X$ with kernel estimators, where the rank-based baseline does not directly apply.","A direct simulation with $R=2$ and continuous $Y$, comparing empirical quantiles of $n\\xi_{+,n}$ under $H_0$ to the claimed mixture, would reveal the theorem's practical scope at moderate sample sizes."],"forward_implications":["Under $H_0$, for fixed $R$, the test statistic $n\\xi_{+,n}$ has a non-normal limit that is an infinite mixture of chi-square variables, so critical values can be computed without bootstrapping.","For alternatives, $\\xi_{+,n}$ is consistent and asymptotically normal at $\\sqrt{n}$ rate, so confidence intervals for dependence strength follow from a plug-in variance estimator.","Because $\\xi_+$ is nonnegative and vanishes only under independence, the proposed test repairs the finite-sample sign problem of Chatterjee's coefficient, which is negative about half the time under the null.","The extension family $\\xi_{k,l,r,s}$ preserves the 0/1 characterization and can be tuned through $k,l,r,s$ for different distributional sensitivity.","The conditional version $\\xi_+(Y,Z|X)$ avoids left limits and is estimated by plugging out-of-sample conditional CDFs, giving a cleaner asymptotic theory for conditional independence."],"supporting_citations":[{"why":"Introduces the rank-based coefficient and the population measure this paper modifies, and supplies the baseline to be improved.","marker":"[9]"},{"why":"Defines the earlier regression-dependence measure whose right-continuous version is proposed here.","marker":"[17]"},{"why":"Supplies the mean-variance index and its chi-square-mixture null limit, which the paper transfers to its estimator.","marker":"[13]"},{"why":"Records the squared-distance representation used in the paper's equivalent formulations.","marker":"[19]"},{"why":"Provides the U-statistics theory behind the degenerate null and non-degenerate alternative asymptotics.","marker":"[26]"},{"why":"Gives the recent asymptotic-normality framework for the baseline rank correlation that motivates the inference setting.","marker":"[30]"},{"why":"Defines the conditional-dependence measure whose right-continuous version and plug-in estimator are developed in Section 5.","marker":"[4]"}],"fun_headline_variants":["Modified Chatterjee measure: exact zero independence, better null","Refined rank dependence: cleaner chi-square null, lower Type I error","New rank correlation: exact zero independence, calibrated tests","Improved Chatterjee-type measure: beats original in Type I error"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The null-distribution theorem depends on the proposed estimator being exactly six times an earlier mean-variance index; in the simplest one-category case the two are not equal, so the theorem's premise is the equality itself.","fun_headline_variants_meta":{"raw":{"variants":["Modified Chatterjee measure: exact zero independence, better null","Refined rank dependence: cleaner chi-square null, lower Type I error","New rank correlation: exact zero independence, calibrated tests","Improved Chatterjee-type measure: beats original in Type I error"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000491,"raw_usage":{"total_tokens":2395,"prompt_tokens":910,"completion_tokens":1485,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":1415}},"tokens_in":526,"tokens_out":1485,"duration_ms":14420,"temperature":1.0,"reasoning_tokens":1415,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:49:21.712776+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute equation (3.5) when $X$ has a single category ($R=1$) and $Y$ is continuous: the proposed estimator equals $3/n + 1/n^2$, while the mean-variance index is identically $0$; $n$ times the estimator therefore tends to $3$, not to the claimed null limit $0$. This one calculation is enough to test whether the premise of Theorem 3.2 is valid.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the rank-based coefficient and the population measure this paper modifies, and supplies the baseline to be improved."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the earlier regression-dependence measure whose right-continuous version is proposed here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the mean-variance index and its chi-square-mixture null limit, which the paper transfers to its estimator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Records the squared-distance representation used in the paper's equivalent formulations."},{"cited_title":"S., & Borovskich, Y","cited_arxiv_id":null,"evidence_quote":"Provides the U-statistics theory behind the degenerate null and non-degenerate alternative asymptotics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the conditional-dependence measure whose right-continuous version and plug-in estimator are developed in Section 5."}],"review_version":1}