{"id":"e9c86d86-c3a3-4fb8-ba98-9a5ca2c878ba","arxiv_id":"2509.10817","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A model-X conditional independence test using a Gaussian-kernel energy distance between observed data and conditionally independent variants, calibrated by random coordinate swaps.","lead":"This paper builds a conditional independence test that compares real data with artificially generated 'conditionally independent variants' using a kernel distance. It proves exact type I error control, consistency, and nontrivial local power, but only when the conditional distribution of X given Z is known or well approximated.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Exact model-X sampling is the sole guarantee of type I error control; the abstract promises a treatment of estimated X|Z that the full text never provides, and Remark 1's GenAI suggestion (sampling X' given Y,Z) violates the exchangeability condition needed by Proposition 2.1.","rationale":"The reader's weakest-assumption diagnosis is correct: the finite-sample validity of Proposition 2.1 requires exact sampling from the true conditional distribution of X given Z, with X'_i independent of Y_i given Z_i. The abstract explicitly promises an investigation of the estimated-X|Z case, but the full text contains no such result; Remark 1 even proposes a procedure (sampling X' given Y,Z) that violates the independence condition and would break the proof of Proposition 2.1. This is a genuine gap between the paper's claims and its content, and it is the dominant reason the verdict should remain CONDITIONAL rather than ACCEPT. I do not find a flaw that would overturn Proposition 2.1 itself under exact model-X sampling: the exchangeability argument is standard and the fixed-alternative consistency follows from Proposition 2.2 and the positive-definiteness of the Gaussian kernel. Secondary technical issues include Theorem 2.5 stating N(0,1) while its proof gives variance 4*sigma_1^2, and Lemma A.2/Theorem 3.2 implicitly requiring G and F to share the same conditional X|Z distribution for the local-power analysis; these are real but less central than the omitted estimation analysis. The CONDITIONAL verdict is therefore unchanged.","tokens_in":29987,"tokens_out":15317,"duration_ms":136320,"concrete_test":"Simulate H0 data with non-Gaussian X|Z (e.g., X = Z + epsilon with epsilon ~ t_3), estimate the conditional distribution by a misspecified Gaussian fit, generate X'_i from the fitted model, and compute the rejection rate of the proposed test at alpha = 0.05 over 1000 replications. If the empirical rejection rate substantially exceeds 5%, the missing analysis of estimated X|Z is essential; Proposition 2.1's exact-exchangeability guarantee does not transfer to the estimated setting the abstract promises to cover. Separately, check whether generating X' from X|(Y,Z) as suggested in Remark 1 inflates type I error under H0.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing condition is Definition 1 plus the sampling step in Section 2.3: X'_i must be drawn from the true conditional distribution of X_i given Z_i, independently of Y_i. Proposition 2.1's coordinate-swap exchangeability of (X_i, X'_i) given (Y_i, Z_i) holds only under this independence. The abstract promises that 'The effect of estimating the conditional distribution used to generate the exchangeable pairs is also investigated, and condition under which validity and power properties are preserved is established.' No such theorem, lemma, or even a formal condition appears in the full text; Section 5 only gestures at generative-model heuristics with no theoretical properties. Worse, Remark 1 explicitly suggests generating X' from (Y,Z) via GenAI, which would make X' dependent on Y and destroy the exchangeability on which Proposition 2.1 rests. Thus for the realistic setting where X|Z is unknown and estimated, the paper provides no type I error control; the central finite-sample guarantee is conditional on an assumption the paper itself does not deliver on, and the abstract overclaims the scope of the results.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a conditional independence test based on an exchangeable-pairs construction in the model-X framework. For each observation (X_i,Y_i,Z_i), an exchangeable counterpart X'_i is drawn from the true conditional distribution of X given Z, independently of Y given Z. The population discrepancy ζσ(P) is the Gaussian-kernel MMD between (X,Y,Z) and its conditionally independent variant (X',Y,Z), and it is estimated by a bounded U-statistic on the augmented sample. A coordinate-swap resampling scheme calibrates the test, yielding a finite-sample type I error guarantee under exact model-X sampling (Proposition 2.1) and consistency against fixed alternatives (Proposition 2.3). The paper further claims asymptotic null and alternative distributions, local asymptotic power and \"Pitman efficiency\" against contiguous alternatives, high-dimensional consistency, and competitive empirical performance.","tokens_in":30226,"tokens_out":17274,"duration_ms":153331,"significance":"The core idea is attractive and, under the oracle model-X assumption, the central randomization argument is sound: the coordinate-swap p-value gives exact finite-sample type I error control, the estimator is a bounded U-statistic with exponential concentration, and the fixed-alternative consistency argument is simple and convincing. This provides a useful complement to the CRT that requires only one augmented sample and reuses pairwise distances. However, several advertised contributions are not actually established: the abstract promises an analysis of estimated X|Z that never appears, Theorem 2.5 is internally inconsistent, the label \"Pitman efficiency\" is not supported by any optimality comparison, and the local-power derivation has a substantial gap. With careful correction and re-scoping, the exchangeable-pairs methodology could be a solid contribution, but the current version overclaims.","major_comments":[{"comment":"The abstract promises that the effect of estimating the conditional distribution used to generate the exchangeable pairs is investigated and that a condition preserving validity and power is established. No such theorem, lemma, or formal condition appears in the full text; Section 5 only says it would be interesting to study generative-model-based sampling. More seriously, Remark 1 suggests using GenAI to generate X given observations on (Y,Z), which makes X' dependent on Y given Z and destroys the coordinate-swap exchangeability on which Proposition 2.1 rests. This is both a missing promised result and an internal inconsistency: the finite-sample type I error guarantee is only established for oracle sampling from the true P_{X|Z}.","section":"Abstract; Section 2.3; Remark 1"},{"comment":"Theorem 2.5 states that under H1, sqrt(n)(ζ̂_{n,σ} − ζσ(P)) converges in distribution to N(0,1). The proof in Appendix A concludes asymptotic normality with variance 4σ1², where σ1² = Var(g1), and no argument shows that 4σ1² = 1. Since σ1² depends on P and σ, the limiting variance is not generally 1. The theorem is false as stated; it should either state N(0,4σ1²) or the statistic must be studentized by a consistent estimator of its asymptotic variance.","section":"Theorem 2.5 and Appendix A"},{"comment":"The label \"Pitman efficient\" is not justified. Theorem 3.2 derives a limiting power function L(β), but it does not compare the test with an optimal benchmark or compute an asymptotic relative efficiency, which is what Pitman efficiency standardly means. In addition, the abstract claims that the test attains the minimax separation rate for the proposed discrepancy measure; no minimax separation result is stated or proved anywhere in the text. These claims should be substantiated or removed.","section":"Section 3.1; Theorem 3.2"},{"comment":"The local alternative F_{1−βn/√n} = (1−βn/√n)G + (βn/√n)F changes the conditional distribution of X|Z, but the procedure requires drawing X'_i from the true P_{X|Z}; under the mixture alternative, X'_i must be generated from the mixture conditional rather than from G_{X|Z}. The contiguity calculation in the proof of Theorem 3.1 computes the log-likelihood ratio only for V_i and omits the likelihood ratio contribution of the generated X'_i given Z_i, so the shift term βE_F{ψ_i} is not derived for the actual augmented-data experiment. Moreover, the displayed limiting covariance matrix in the same proof has a negative lower-right entry (−β²/2 E[...]) instead of a nonnegative variance. The local-power analysis needs to be reformulated (for example, by restricting to alternatives with fixed X|Z or by deriving the full likelihood ratio for the augmented data) before Theorems 3.1 and 3.2 can be considered established.","section":"Theorems 3.1 and 3.2; proof of Theorem 3.1"}],"minor_comments":[{"comment":"There are numerous typographical errors, including \"Keywards\" in the abstract, \"hypothesises\" in Section 1, \"afficacy\" in Section 5, and \"Piman efficiency\" in the text following Theorem 3.2.","section":"Abstract and throughout"},{"comment":"The sentence \"A test of spherical symmetry is also proposed\" is incorrect; the paper proposes a test of conditional independence, not a test of spherical symmetry.","section":"Section 5"},{"comment":"The displayed power formula in Theorem 3.2(b) has the form (Z_i + βE_F{ψ_i})², but the proof in Appendix A writes \"λ_i(Z_i + β²E[ψ_i(V_1,V'_1)] − 1)\" and omits the square; the notation should be harmonized.","section":"Theorem 3.2 statement and proof"},{"comment":"The legend uses \"GCM.fix test\" while the text defines \"wGCM.fix test\"; these should be made consistent.","section":"Figure 3"},{"comment":"Lemma A.1 refers to \"spherically symmetric variants\" although the construction is coordinate swapping, not spherical symmetry. In Remark 5, the justification \"V1 =D V'2\" is incorrect; the cancellation follows from V1 =D V'1 when ζσ(G) = 0.","section":"Lemma A.1 and Remark 5"}],"recommendation":"major_revision","confidential_remarks":"The central finite-sample randomization argument (Proposition 2.1) and the fixed-alternative consistency result are sound and provide a publishable core if the paper is re-scoped. The current version, however, contains a false theorem (Theorem 2.5), an unfulfilled promise about estimated X|Z, unsupported claims of Pitman efficiency and minimax optimality, and a gap in the local-power derivation. I would encourage the editor to invite a major revision rather than reject, but the revision must either fix or explicitly remove these claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First: the paper has a real, useful core. Proposition 2.1 is the right kind of result—exact finite-sample type I error under the model-X assumption—and the coordinate-swap calibration combined with the U-statistic discrepancy is a genuinely new construction. Fixed-alternative consistency follows cleanly from c_hat = O_P(n^{-1}), and Theorem 3.3's high-dimensional condition (n ζ_σ(P) → ∞) is a reasonable sufficient condition. The basic idea of comparing observed data to conditionally independent variants is worth taking seriously; Proposition 2.2's bound is sound, even if the proof is rougher than it should be.\n\nThe soft spots are real and need work. Theorem 2.5 states that √n(ζ̂ - ζ) ⇒ N(0,1), but the proof reports asymptotic variance 4σ₁² and never shows that this equals 1. That is a technical misstatement, not a deep flaw, but it has to be fixed. The \"Pitman efficiency\" language is also inflated: the paper shows the power converges to a nontrivial limit L(β), but Pitman efficiency normally means a comparison to a benchmark test, and no such comparison is made. The abstract that came with the submission promises both minimax separation rates and an analysis of estimating X|Z. Neither appears in the full text. Section 5 only gestures at generative models. Worse, Remark 1 suggests generating X' from (Y,Z), which makes X' dependent on Y given Z and destroys the exchangeability on which Proposition 2.1 rests.\n\nNone of this sinks the exact-model-X version of the test. The randomization argument is sound. What sinks the current submission is the gap between what is claimed and what is proved. The author should either deliver the estimated-X|Z theory and the minimax statement, or cut them from the abstract and discussion. Remark 1 needs to be rewritten to say X'|Z, or removed. There is also a leftover \"test of spherical symmetry\" in Section 5 and assorted typos; the manuscript is not ready as is.\n\nThe intended audience is people working on model-X conditional independence testing, especially those interested in fast resampling calibration. It deserves a serious referee, but the referee should focus on the exact-sampling case and require that the overclaims be brought in line. I would not cite it in its current form, but I would watch for a revised version.","headline":"A clever exchangeable-pairs CI test with a sound exact-model-X randomization core, but the abstract promises an estimated-X|Z analysis the full text never delivers, and one remark explicitly breaks the exchangeability assumption.","tokens_in":30718,"tokens_out":3546,"would_cite":false,"duration_ms":34189,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Exchangeable pairs make conditional independence tests exact at every n","keywords":["conditional independence","exchangeable pairs","model-X framework","maximum mean discrepancy","U-statistic","resampling test","Pitman efficiency","high-dimensional inference"],"falsifier":"Simulate i.i.d. data under $H_0$ from a known distribution with accessible $P_{X\\mid Z}$, sample the CI variants exactly, and run the coordinate-flip test over a grid of $n$, $d$, and $\\alpha$ with many replications: the claim is that the rejection frequency never exceeds $\\alpha$, so any stable exceedance would falsify Proposition 2.1. The same design with $X'_i$ drawn from a deliberately misspecified conditional distribution would show how the guarantee erodes when the model-X assumption is relaxed.","tokens_in":29782,"feed_emoji":"🔀","tokens_out":7081,"duration_ms":48847,"temperature":0.7,"pith_summary":"This paper proposes a test of conditional independence $X \\perp\\!\\!\\!\\perp Y \\mid Z$ built from exchangeable pairs. The idea is to sample a conditionally independent variant $X'_i$ from the true distribution of $X \\mid Z_i$ for each observation, and then compare the joint law of $(X_i,Y_i,Z_i)$ with that of $(X'_i,Y_i,Z_i)$; under the null the two laws coincide, while under the alternative they differ. The comparison is made by a Gaussian-kernel maximum mean discrepancy $\\zeta_\\sigma(P)$, estimated by a U-statistic whose concentration is exponential and free of the dimension $d$. The test is calibrated by flipping each $X_i$ with $X'_i$ in all $2^n$ ways; under the null this coordinate swap is distributionally exact, so the conditional p-value satisfies $P\\{p_n < \\alpha\\}\\le\\alpha$ for every $n$ and $d$. Because the discrepancy $\\zeta_\\sigma(P)$ is positive under any fixed alternative and the critical value is $O_P(n^{-1})$, the same test is consistent, has nontrivial Pitman power against local contiguous alternatives, and is consistent when $d\\to\\infty$ with $n$ provided $n\\zeta_\\sigma(P)\\to\\infty$.","feed_headline":"Exchangeable pairs make conditional independence tests exact at every n","feed_subtitle":"Sampling a matched X′ from X|Z turns conditional independence into a two-sample test with error control at any n.","key_machinery":"The load-bearing objects are the CI variant $(X',Y,Z)$ and the coordinate-swap resampling scheme. Given each $X'_i\\sim P_{X\\mid Z_i}$ independent of $Y_i$ given $Z_i$, the paper forms ordered pairs $(X_i,X'_i,Y_i,Z_i)$ and measures dependence by the Gaussian-kernel MMD $\\zeta_\\sigma(P)=\\mathbb{E}K(V_1,V_2)+\\mathbb{E}K(V'_1,V'_2)-2\\mathbb{E}K(V_1,V'_2)$ with $K(a,b)=\\exp(-\\sigma^2\\|a-b\\|^2/2)$. The estimator is a degree-two U-statistic with a bounded core, which yields dimension-free exponential concentration; the null degeneracy of the core gives the $\\sum_i\\lambda_i(U_i^2-1)$ limit. Calibration uses the exact exchangeability of $X_i$ and $X'_i$ under $H_0$: flipping the two coordinates in every subset $\\pi\\in\\{0,1\\}^n$ produces a resampling distribution whose $(1-\\alpha)$-quantile is bounded by $2(\\alpha(n-1))^{-1}$ almost surely, so the p-value is finite-sample valid.","core_discovery":"The central claim is that conditional independence testing can be reformulated as a two-sample problem: compare $(X,Y,Z)$ with its conditionally independent variant $(X',Y,Z)$, where $X'\\mid Z$ has the same law as $X\\mid Z$ and $X'\\perp\\!\\!\\!\\perp Y\\mid Z$. The paper proves that the Gaussian-kernel discrepancy $\\zeta_\\sigma(P)$ between these two laws is zero exactly under $H_0$ and positive under $H_1$, that its U-statistic estimator $\\hat\\zeta_{n,\\sigma}$ concentrates exponentially fast with constants free of dimension, and that the coordinate-flip resampling p-value $p_n$ obeys $P\\{p_n<\\alpha\\}\\le\\alpha$ for all $n$ and $d$ under $H_0$. It then shows the test is consistent against fixed alternatives, attains a nontrivial limiting power $L(\\beta)$ against local contiguous alternatives at rate $n^{-1/2}$, and remains consistent when $d$ grows with $n$ as long as $n\\zeta_\\sigma(P)$ diverges.","pith_inferences":["The paper's abstract announces an investigation of estimating $P_{X\\mid Z}$, but the full text proves no theorem for estimated or misspecified conditional distributions; if $X'_i$ is generated from an estimate or a generative model, the exchangeability behind Proposition 2.1 no longer follows, and a direct simulation under $H_0$ with misspecified sampling would show how rejection rates depart from","The coordinate-flip calibration is a general exchangeability principle: any statistic that is invariant under swapping each $X_i$ with $X'_i$ inherits the same finite-sample type I error argument, suggesting the exchangeable-pairs construction could be combined with dependence measures other than the Gaussian-kernel MMD.","Because $\\zeta_\\sigma(P)$ is a kernel MMD between a vector and its CI variant, a finite battery of bounded kernels (Laplace, inverse quadratic, or others) could in principle be used to detect different dependence geometries; the paper notes the extension to other bounded kernels but does not study the power of such a battery.","The model-X assumption is likely replaceable by a doubly robust variant in which only the regression of $X$ on $Z$ is estimated well, as is done in recent work on conditional randomization tests; testing that variant would be a natural next step beyond the paper's exact-exchangeability setting."],"forward_implications":["Under exact model-X sampling, the test controls the type I error at level $\\alpha$ for every finite $n$ and every dimension $d$, without smoothness or moment assumptions beyond the existence of $X\\mid Z$.","The test is consistent against any fixed alternative because $\\zeta_\\sigma(P)>0$ under $H_1$ and the resampling critical value $\\hat c_{1-\\alpha}$ is $O_P(n^{-1})$.","Against local contiguous alternatives at distance $n^{-1/2}$, the power converges to a nontrivial limit $L(\\beta)$ that increases from $\\alpha$ to one as $\\beta$ grows.","When the dimension $d$ grows with the sample size, power converges to one provided $n\\zeta_\\sigma(P)\\to\\infty$; dimension enters only through the size of the population discrepancy.","The randomized p-value $p_{n,B}$ approximates the exact flip p-value $p_n$ with error of order $O_P(B^{-1/2})$ plus $(B+1)^{-1}$, so moderate $B$ suffices in practice."],"supporting_citations":[{"why":"Supplies the model-X framework and the conditional randomization test; the CI-variant construction borrows the assumption that $X\\mid Z$ is known and sampleable.","marker":"Candes et al. (2018)"},{"why":"Establishes the kernel maximum mean discrepancy formulation that the measure $\\zeta_\\sigma(P)$ directly instantiates.","marker":"Gretton et al. (2012)"},{"why":"Provides the U-statistic limit theory used for the null distribution, asymptotic normality, and variance bounds.","marker":"Lee (1990)"},{"why":"Introduces the conditional permutation test and studies robustness under approximate knowledge of the conditional distribution, the benchmark for model-X resampling.","marker":"Berrett et al. (2019)"},{"why":"Documents the no-free-lunch hardness of nonparametric conditional independence testing, motivating the model-X assumption as the setting where finite-sample control is feasible.","marker":"Shah and Peters (2020)"},{"why":"Supplies the Dvoretzky-Kiefer-Wolfowitz bound used to control the randomized p-value approximation.","marker":"Massart (1990)"},{"why":"Provides the bounded-difference inequality used to prove the dimension-free exponential concentration of the estimator.","marker":"Wainwright (2019)"},{"why":"Supplies Le Cam's third lemma and contiguity tools used to derive the local limiting distribution under contiguous alternatives.","marker":"Van der Vaart (2000)"}],"fun_headline_variants":["CI testing via exchangeable pairs: exact p-values for any n and d","Turn conditional independence into a two-sample test with exact error control","Exchangeable pairs give finite-sample-valid conditional independence tests","New test for conditional independence: valid for all n and d","Conditional independence test with finite-sample guarantees via exchangeability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the analyst can sample $X'_i$ from the true conditional distribution of $X$ given $Z_i$, independently of $Y_i$ given $Z_i$ for every observation, so that without exact model-X sampling the coordinate-swap exchangeability behind the finite-sample bound fails.","fun_headline_variants_meta":{"raw":{"variants":["CI testing via exchangeable pairs: exact p-values for any n and d","Turn conditional independence into a two-sample test with exact error control","Exchangeable pairs give finite-sample-valid conditional independence tests","New test for conditional independence: valid for all n and d","Conditional independence test with finite-sample guarantees via exchangeability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000726,"raw_usage":{"total_tokens":3282,"prompt_tokens":1001,"completion_tokens":2281,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":2194}},"tokens_in":617,"tokens_out":2281,"duration_ms":360615,"temperature":1.0,"reasoning_tokens":2194,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:53:00.933643+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate i.i.d. data under $H_0$ from a known distribution with accessible $P_{X\\mid Z}$, sample the CI variants exactly, and run the coordinate-flip test over a grid of $n$, $d$, and $\\alpha$ with many replications: the claim is that the rejection frequency never exceeds $\\alpha$, so any stable exceedance would falsify Proposition 2.1. The same design with $X'_i$ drawn from a deliberately misspecified conditional distribution would show how the guarantee erodes when the model-X assumption is relaxed.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the model-X framework and the conditional randomization test; the CI-variant construction borrows the assumption that $X\\mid Z$ is known and sampleable."},{"cited_title":"M., Rasch, M","cited_arxiv_id":null,"evidence_quote":"Establishes the kernel maximum mean discrepancy formulation that the measure $\\zeta_\\sigma(P)$ directly instantiates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the U-statistic limit theory used for the null distribution, asymptotic normality, and variance bounds."},{"cited_title":"B., Wang, Y., Barber, R","cited_arxiv_id":null,"evidence_quote":"Introduces the conditional permutation test and studies robustness under approximate knowledge of the conditional distribution, the benchmark for model-X resampling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the no-free-lunch hardness of nonparametric conditional independence testing, motivating the model-X assumption as the setting where finite-sample control is feasible."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Dvoretzky-Kiefer-Wolfowitz bound used to control the randomized p-value approximation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the bounded-difference inequality used to prove the dimension-free exponential concentration of the estimator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Le Cam's third lemma and contiguity tools used to derive the local limiting distribution under contiguous alternatives."}],"review_version":2}