{"id":"d175e74f-a779-4d0a-9b66-146b5aebab65","arxiv_id":"2506.13671","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"For rare-event independence tests, power and convergence rate depend on the number of cases n1, so a subsampled control set of size s times n1 achieves nearly the same power as the full sample.","lead":"This paper shows that in independence tests for rare events, statistical power is set by the number of rare cases, not the total sample size, and proposes a rescaled test and a subsampled 'boosted' test that keep nearly the same power while using far fewer non-events. The finding matters for large imbalanced datasets, where keeping all controls can be wasteful once the case count is fixed.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"High-dimensional normality of RIT/BIT relies on unverified condition (2.6); the distance/projection covariance examples never check it, leaving the central high-dimensional claim unsupported.","rationale":"The paper's headline phenomenon is conditional on extreme imbalance n1/n0 -> 0; the reader is right that outside this regime 'more controls' does carry information. That is a scope restriction, explicitly acknowledged by the paper, and it does not threaten the internal logic of Theorem 1. The more serious gap is the unverified condition (2.6) in the high-dimensional theorems. Section 2.3 presents T_dcov and T_IPcov as special cases, but the asymptotic normality in the high-dimensional case depends on a martingale CLT that the authors reduce to condition (2.6). The manuscript even states that for fixed p the condition fails, so the p -> infinity regime is not a formality. A concrete verification for the two kernels is required. Since these statistics are the ones used in the simulations and real-data sections, the missing check is load-bearing. The reader's CONDITIONAL verdict remains appropriate, but for a different reason than the one emphasized.","tokens_in":42192,"tokens_out":4232,"duration_ms":42891,"concrete_test":"Take X with independent N(0,1) coordinates, with dimension p = p(n1). For the distance-covariance kernel h_{0,2}(x,y) = D(x)+D(y)-||x-y|| - gamma*, compute or asymptotically bound E[G^2], E[h_{0,2}^4], and (E[h_{0,2}^2])^2 as n1 and p grow along at least two paths: n1 = o(p) and p = o(n1). Check whether the ratio in condition (2.6) tends to 0. Repeat the same calculation for the improved projection covariance kernel A(.,.). If the ratio does not vanish, Theorems 2 and 4 cannot be applied to these kernels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central high-dimensional claim (Theorems 2 and 4, and the derived limits for T_dcov and T_IPcov in Section 2.3 items 4(2) and 5(2)) is conditioned on condition (2.6), which requires E[G^2(X1,X2)] + n1^{-1} E[h_{0,2}^4] to be o((E[h_{0,2}^2])^2). The paper nowhere verifies condition (2.6) for the distance covariance or improved projection covariance kernels; Appendix C.2 only states that the martingale CLT verification assumes (2.6) holds. Since the paper itself notes that the condition fails when p is fixed, its behavior for p going to infinity is delicate. Without an explicit verification, the asymptotic N(0,2) result for T_dcov, which is used in the simulations and real-data analyses, is not established. The reader's extreme-imbalance concern is a scope boundary; this is a missing support inside the stated theorem, and it directly affects the flagship high-dimensional examples.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies independence testing when one binary class is rare, i.e., when n1/n0 goes to 0. It introduces a rescaled independence test (RIT) based on generalized U-statistics and proves that, under this extreme-imbalance regime, the convergence rate of the statistic is determined by the number of rare events n1 rather than the total sample size n. It then proposes a boosted independence test (BIT) that subsamples controls with a user-chosen ratio s and shows that the same convergence rate is preserved. The theoretical results cover first-order and second-order kernels in both fixed and high-dimensional settings, include a local power analysis under mixture alternatives, and extend to multiple rare-event classes. Simulations and two real-data applications illustrate the effectiveness and computational savings of the proposed procedures.","tokens_in":42330,"tokens_out":7029,"duration_ms":69945,"significance":"If the main theorems hold, the paper delivers a practically valuable and counterintuitive message: under extreme class imbalance, adding control observations does not improve the power of independence tests, and a small subsample of controls can nearly match the full-sample performance. The first-order asymptotic proofs are standard and carefully executed, and the subsampling construction uses a user-chosen ratio rather than data-fitted constants, which is a methodological strength. The multi-class extension broadens the framework. However, the high-dimensional normal limits for the distance and projection covariance examples rest on an unverified condition, and the headline claim in the abstract is stated without the extreme-imbalance qualifier that the theorems require. These issues are fixable but currently leave the flagship examples only partially supported.","major_comments":[{"comment":"Condition (2.6) is not verified for the distance covariance or improved projection covariance kernels. The paper states after Theorem 2 that the condition fails when p is fixed, and Appendix C.2 only uses (2.6) as an assumption in the martingale central limit theorem verification; it does not show that the kernels h_0,2 for T_dcov or T_IPcov satisfy the required moment ratios when p diverges. Since the asymptotic N(0,2) limits in items 4(2) and 5(2) are used in the simulation studies (p = 50) and the real-data analyses (p = 5408), this missing verification directly affects the evidence for the paper's flagship high-dimensional claims. The authors should either provide a proof of (2.6) for these kernels under explicit dimension and moment assumptions, or restrict the statements to kernels for which the condition is verified.","section":"Section 2.3, items 4(2) and 5(2); Theorems 2 and 4; Appendix C.2"},{"comment":"The proof that Var{n1(TS - V)} tends to 0 uses the claim Var(n1TS) converges to m1^2(m1-1)^2 xi_{0,2}/2 'from formula (C.10)', but formula (C.10) also contains the terms m1^2 m0^2 xi_{1,1}/(s n1^2) and m0^2(m0-1)^2 xi_{2,0}/(2 s^2 n1^2). After multiplying by n1^2, these terms vanish only if xi_{1,1} = o(s) and xi_{2,0} = o(s^2). Since xi_{1,1} and xi_{2,0} may depend on n1 and p, the condition s tending to infinity alone does not ensure these negligibility conditions. The theorem needs an additional explicit condition, or a proof that (2.6) implies these ratios are negligible.","section":"Appendix C.4, proof of Theorem 4"},{"comment":"The central statement that 'the power of the test is determined by the number of rare events rather than the total sample size' is presented without the qualifier n1/n0 -> 0. Theorem 1 is proved only under n1/n -> 0, and for the second-order case with xi_{1,0} not zero it additionally requires n1^2/n0 -> 0. Outside this regime the asymptotic variance of the rescaled statistic includes the control contribution m0^2 xi_{1,0}/n0, so increasing n0 does carry information. The abstract and conclusion should state the extreme-imbalance regime explicitly so that the headline claim is not overgeneralized.","section":"Abstract and Section 2.2"}],"minor_comments":[{"comment":"The first variance term in the statement appears to be a typo: it should be m1^2(m1-1)^2 zeta_{2,1}/2 rather than m_k^2(m1-1)^2 zeta_{2,1}/2.","section":"Corollary 3(2)"},{"comment":"In the count of xi_{0,2} terms in the proof of Theorem 1(2), the combinations involving n1 are written with m0 in the lower entries; they should be (n1-2 choose m1-2) and (n1-m1 choose m1-2).","section":"Appendix C.1, Step 1"},{"comment":"Tables 4 and 5 report p-values as <0.001, but the text does not state how the null distribution was obtained (asymptotic normal approximation via Theorem 2, permutation, or another method). Please specify the procedure.","section":"Section 3.3"},{"comment":"The phrase 'Key W ords' contains a typo and should be 'Key Words'.","section":"Abstract"},{"comment":"The text refers to Figure 1, but the figure itself is not shown in the manuscript; please ensure it is included in the published version.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The central phenomenon is real and clearly demonstrated: in the n1/n0 -> 0 regime, first-order RIT statistics converge at the n1 rate, and the BIT subsampling scheme preserves that rate with far less computation. The toy example and simulations make the point convincingly, and the generalized U-statistic framework is a nice way to organize the rescaled Pearson, Kendall, distance, and projection covariance statistics. That part is solid and useful.\n\nThe soft spots are where the theory reaches for high dimensions. Condition (2.6) is load-bearing for Theorems 2 and 4, and the paper never verifies it for the distance covariance or improved projection covariance kernels. The appendix just says the martingale CLT verification assumes (2.6) holds. Since the paper itself notes the condition fails for fixed p, and the high-dimensional behavior is delicate, the N(0,2) results for T_dcov and T_IPcov are not actually established. That is a missing support inside the stated theorem, not just a scope boundary. I would want that fixed before publication: either verify (2.6) for those kernels under explicit growth conditions on p, or weaken the claims to 'conditional on (2.6)'.\n\nAlso, Theorem 3(2) only asserts a nondegenerate limit with a given variance, not a named distribution. That is acceptable for a variance calculation but weaker than Theorem 1, and the paper could say so more clearly. The real-data analyses lack code and permutation details; minor but easy to address.\n\nThe scope caveat the reader raised is real but not fatal: if n1/n0 does not vanish, control observations do carry information, and the paper's headline statement stops being true. The authors do state the regime clearly, so this is a defined scope rather than an overclaim. Still, the abstract's 'surprisingly, determined by the number of rare events' should come with a qualifier like 'in the extreme imbalance regime' to avoid misleading casual readers.\n\nBottom line: the first-order theory and the subsampling idea are worth publishing. The high-dimensional claims need the missing verification or a conditional statement. I would send this to a serious referee; with that fix it would be a solid contribution.","headline":"The rare-event independence-test framework is a genuinely useful contribution, but the high-dimensional normality results rest on an unverified condition (2.6) that must be addressed before the paper is fully trustworthy.","tokens_in":42925,"tokens_out":2259,"would_cite":false,"duration_ms":22856,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62G20","62H20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Rare-event independence tests get their power from the number of cases, not the total sample size.","keywords":["independence test","rare events","imbalanced data","generalized U-statistics","subsampling","local power","high-dimensional testing","distance correlation"],"falsifier":"Run the rescaled Kendall tau test with $n_1=50$ fixed and a fixed shift alternative, and compare power at $n_0=500$, $5{,}000$, and $500{,}000$: the paper predicts the power stays essentially constant once $n_1/n_0$ is small. If the power climbs systematically with $n_0$ in that comparison, the rate result would fail; conversely, running the same comparison at $n_1/n_0=0.5$ should show power increasing with $n_0$, confirming that the no-information conclusion is specific to the vanishing-proportion regime.","tokens_in":41930,"feed_emoji":"📊","tokens_out":9201,"duration_ms":85298,"temperature":0.7,"pith_summary":"Independence tests are normally expected to improve as more data arrive. This paper shows that when one outcome class is rare, the power of an independence test is governed by the number of rare cases $n_1$, not the total sample size $n$: if $n_1$ is fixed while $n_1/n$ shrinks, correlations shrink to zero and test power plateaus below one. To recover the lost information, the paper rescales common statistics — Pearson correlation, Kendall's tau, distance covariance, and projection covariance — as two-sample U-statistics, and proves that the rescaled statistic converges at a rate set only by $n_1$. It then proposes the Boosted Independence Test (BIT), which keeps every case and only about $s n_1$ subsampled controls; with fixed $s\\geq 2$ it has the same convergence rate and nearly the same local power as using all controls, at a fraction of the computation. The theory extends to multi-class rare events.","feed_headline":"Rare events: test power comes from cases, not sample size","feed_subtitle":"Subsampling controls while keeping all rare cases preserves nearly the full power of independence tests.","key_machinery":"The carrying object is the rescaled two-sample U-statistic $T$ of (2.3), with kernel $h$ acting on $m_0$ controls and $m_1$ cases, and its projections $h_{a,b}(X^{(0)}_1,\\dots,X^{(0)}_a; X^{(1)}_1,\\dots,X^{(1)}_b)$ with variances $\\xi_{a,b}=\\mathrm{Var}(h_{a,b})$. The proof's leverage is the variance expansion $\\mathrm{Var}(T)=m_1^2\\xi_{0,1}/n_1 + m_0^2\\xi_{1,0}/n_0 + O(n_0^{-1}n_1^{-1})$; under $n_1/n_0\\to 0$, the $n_0$ term drops out, so the convergence rate of the whole statistic is $n_1$. The BIT version (2.7) multiplies each control in the kernel by an independent Bernoulli indicator $\\delta_i$ with $\\mathrm{P}(\\delta_i=1)=s n_1/n_0$, so only about $s n_1$ controls are used; this adds the variance term $m_0^2\\xi_{1,0}/s$ but leaves the $n_1$-rate intact, reducing computational complexity from $O(p n^2)$ to $O(p s^2 n_1^2)$ for the second-order statistics.","core_discovery":"On the paper's own terms, the central discovery is a rate collapse: in the extreme-imbalance regime $n_1/n_0\\to 0$, the asymptotic law of any kernel-based independence statistic of the form (2.3) is fixed by the case count alone. For a first-order kernel, $n_1^{1/2}T \\xrightarrow{d} N(0, m_1^2 \\xi_{0,1})$; for a second-order (degenerate) kernel with $\\xi_{0,1}=0$, $n_1 T$ converges to a weighted chi-square distribution unless a high-dimensional scaling (2.6) turns it normal. The same rate controls the subsampled statistic $T_S$ built from all $n_1$ cases and $s n_1$ controls: $n_1^{1/2}T_S \\xrightarrow{d} N(0, m_1^2\\xi_{0,1}+m_0^2\\xi_{1,0}/s)$. This is why the classical statistics in the toy example plateau: they are effectively using a rate tied to $n_1$, so adding controls alone cannot push their power to one. The paper verifies the phenomenon and the rescaled remedy for Pearson, Kendall's tau, distance covariance, improved projection covariance, and multi-class extensions.","pith_inferences":["If the conclusion holds, then in study design for rare outcomes the relevant currency is the number of observed cases: a database with few cases and millions of controls is informationally equivalent for independence testing to its control subsample, so retrospective collections should prioritize case accretion.","The paper's simulations suggest a practical rule of thumb not stated as a theorem: a sampling ratio around $s=5$ recovers essentially full-sample power for both first- and second-order statistics, so the computational saving can be attained without an explicit optimal-$s$ procedure.","The variance formula makes a testable prediction for mildly imbalanced data: at $n_1/n_0$ small but not vanishing, power should improve only by the amount predicted by the $m_0^2\\xi_{1,0}/n_0$ term, which is negligible compared with $m_1^2\\xi_{0,1}/n_1$; this gives a quantitative boundary for when more controls help.","The same projection argument should apply to other kernel-based dependence measures, so the rescaled-statistic construction could be transplanted to, for example, kernel correlation or rank-based measures without new rate calculations."],"forward_implications":["Classical tests such as Pearson correlation, Kendall's tau, distance correlation, and projection correlation lose power in rare-event settings, and the loss cannot be repaired by adding controls.","Rescaled statistics detect dependence at local alternatives of size $n_1^{-1/2}$ (first-order) or $n_1^{-1}$ (high-dimensional second-order); the classical $n^{-1/2}$ scale is unattainable.","BIT with $s n_1$ controls has the same convergence rate as the full-sample RIT, and as $s\\to\\infty$ its distribution matches the full statistic.","The multi-class extension shows that when several rare classes have comparable sizes, the convergence rate is still set by $n_1$; when one class is rarer than all others, the rate is set by that class alone.","In high-dimensional settings, the second-order BIT is asymptotically normal under condition (2.6), giving an explicit null distribution that avoids permutation."],"supporting_citations":[{"why":"Establishes asymptotic normality of non-degenerate U-statistics; supplies the baseline limit theory that Lemma 1 contrasts with the rare-event result.","marker":"Hoeffding, 1948a"},{"why":"Gives the degenerate U-statistic weighted-chi-square limit used for the second-order RIT null distribution.","marker":"Gregory, 1977"},{"why":"Introduces generalized U-statistics whose two-sample form grounds the rescaled statistic T in (2.3).","marker":"Sen, 1974"},{"why":"Provides the projection decomposition and Lemma 3 for degenerate kernels that carry the variance and limit derivations.","marker":"Serfling, 2009"},{"why":"Martingale central limit technique adapted in Appendix C.2 to prove high-dimensional asymptotic normality under condition (2.6).","marker":"Zheng, 1996"},{"why":"Defines the mixture alternative family used for local power analysis in Corollary 1.","marker":"Farlie, 1961"},{"why":"Earlier demonstration that under-sampling maintains convergence rates in imbalanced data; the paper extends this to independence testing.","marker":"Wang, 2020"},{"why":"Random-permutation calibration recommended for the intractable second-order null distribution.","marker":"Berrett et al., 2020"}],"fun_headline_variants":["Rare events: power follows cases, not sample size","Rare event tests: power comes from case count, not sample size","Subsample controls, keep rare cases, retain test power","Case count, not sample size, drives rare-event test power"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the case fraction vanishes: $n_1/n_0\\to 0$ while $n_1\\to\\infty$ (plus, for second-order kernels, either $\\xi_{1,0}=0$ or $n_1^2/n_0\\to 0$); if the class proportion does not vanish, the $m_0^2\\xi_{1,0}/n_0$ term survives in the limiting variance and additional controls genuinely add information, so the paper's central answer to its title question no longer holds.","fun_headline_variants_meta":{"raw":{"variants":["Rare events: power follows cases, not sample size","Rare event tests: power comes from case count, not sample size","Subsample controls, keep rare cases, retain test power","Case count, not sample size, drives rare-event test power"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001036,"raw_usage":{"total_tokens":4388,"prompt_tokens":1002,"completion_tokens":3386,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":3315}},"tokens_in":618,"tokens_out":3386,"duration_ms":25803,"temperature":1.0,"reasoning_tokens":3315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:28:43.886468+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the rescaled Kendall tau test with $n_1=50$ fixed and a fixed shift alternative, and compare power at $n_0=500$, $5{,}000$, and $500{,}000$: the paper predicts the power stays essentially constant once $n_1/n_0$ is small. If the power climbs systematically with $n_0$ in that comparison, the rate result would fail; conversely, running the same comparison at $n_1/n_0=0.5$ should show power increasing with $n_0$, confirming that the no-information conclusion is specific to the vanishing-proportion regime.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the degenerate U-statistic weighted-chi-square limit used for the second-order RIT null distribution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces generalized U-statistics whose two-sample form grounds the rescaled statistic T in (2.3)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the projection decomposition and Lemma 3 for degenerate kernels that carry the variance and limit derivations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Martingale central limit technique adapted in Appendix C.2 to prove high-dimensional asymptotic normality under condition (2.6)."},{"cited_title":"(1961), The asymptotic efficiency of Daniels’s generalized correlation coefficients, Journal of the Royal Statistical Society Series B: Statistical Methodology, 23, 128--142","cited_arxiv_id":null,"evidence_quote":"Defines the mixture alternative family used for local power analysis in Corollary 1."},{"cited_title":"(2020), Logistic Regression for Massive Data with Rare Events, Journal of Machine Learning Research, 21, 1--34","cited_arxiv_id":null,"evidence_quote":"Earlier demonstration that under-sampling maintains convergence rates in imbalanced data; the paper extends this to independence testing."},{"cited_title":"B., Wang, Y., Barber, R","cited_arxiv_id":null,"evidence_quote":"Random-permutation calibration recommended for the intractable second-order null distribution."}],"review_version":1}