Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Conditional Independence Testing Using Exchangeable Pairs

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Exchangeable pairs make conditional independence tests exact at every n

desk verdict A clever exchangeable-pairs CI test with a sound exact-model-X randomization core, but the abstract promises an estimated-X|Z analysis the full text never delivers, and one remark explicitly breaks the exchangeability assumption. read the letter →

arxiv 2509.10817 v2 pith:PBZBCN5I submitted 2025-09-13 math.ST stat.MEstat.TH

classification math.STstat.MEstat.TH MSC 62G1062G20
keywords conditionalindependenceexchangeablepairsmodel-XframeworkmaximummeandiscrepancyU-statisticresamplingtestPitmanefficiencyhigh-dimensionalinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a test of conditional independence $X \perp\!\!\!\perp Y \mid Z$ built from exchangeable pairs. The idea is to sample a conditionally independent variant $X'_i$ from the true distribution of $X \mid Z_i$ for each observation, and then compare the joint law of $(X_i,Y_i,Z_i)$ with that of $(X'_i,Y_i,Z_i)$; under the null the two laws coincide, while under the alternative they differ. The comparison is made by a Gaussian-kernel maximum mean discrepancy $\zeta_\sigma(P)$, estimated by a U-statistic whose concentration is exponential and free of the dimension $d$. The test is calibrated by flipping each $X_i$ with $X'_i$ in all $2^n$ ways; under the null this coordinate swap is distributionally exact, so the conditional p-value satisfies $P\{p_n < \alpha\}\le\alpha$ for every $n$ and $d$. Because the discrepancy $\zeta_\sigma(P)$ is positive under any fixed alternative and the critical value is $O_P(n^{-1})$, the same test is consistent, has nontrivial Pitman power against local contiguous alternatives, and is consistent when $d\to\infty$ with $n$ provided $n\zeta_\sigma(P)\to\infty$.

What carries the argument

The load-bearing objects are the CI variant $(X',Y,Z)$ and the coordinate-swap resampling scheme. Given each $X'_i\sim P_{X\mid Z_i}$ independent of $Y_i$ given $Z_i$, the paper forms ordered pairs $(X_i,X'_i,Y_i,Z_i)$ and measures dependence by the Gaussian-kernel MMD $\zeta_\sigma(P)=\mathbb{E}K(V_1,V_2)+\mathbb{E}K(V'_1,V'_2)-2\mathbb{E}K(V_1,V'_2)$ with $K(a,b)=\exp(-\sigma^2\|a-b\|^2/2)$. The estimator is a degree-two U-statistic with a bounded core, which yields dimension-free exponential concentration; the null degeneracy of the core gives the $\sum_i\lambda_i(U_i^2-1)$ limit. Calibration uses the exact exchangeability of $X_i$ and $X'_i$ under $H_0$: flipping the two coordinates in every subset $\pi\in\{0,1\}^n$ produces a resampling distribution whose $(1-\alpha)$-quantile is bounded by $2(\alpha(n-1))^{-1}$ almost surely, so the p-value is finite-sample valid.

What would settle it

Simulate i.i.d. data under $H_0$ from a known distribution with accessible $P_{X\mid Z}$, sample the CI variants exactly, and run the coordinate-flip test over a grid of $n$, $d$, and $\alpha$ with many replications: the claim is that the rejection frequency never exceeds $\alpha$, so any stable exceedance would falsify Proposition 2.1. The same design with $X'_i$ drawn from a deliberately misspecified conditional distribution would show how the guarantee erodes when the model-X assumption is relaxed.

Watch

Extended reading notes

Core claim

The central claim is that conditional independence testing can be reformulated as a two-sample problem: compare $(X,Y,Z)$ with its conditionally independent variant $(X',Y,Z)$, where $X'\mid Z$ has the same law as $X\mid Z$ and $X'\perp\!\!\!\perp Y\mid Z$. The paper proves that the Gaussian-kernel discrepancy $\zeta_\sigma(P)$ between these two laws is zero exactly under $H_0$ and positive under $H_1$, that its U-statistic estimator $\hat\zeta_{n,\sigma}$ concentrates exponentially fast with constants free of dimension, and that the coordinate-flip resampling p-value $p_n$ obeys $P\{p_n<\alpha\}\le\alpha$ for all $n$ and $d$ under $H_0$. It then shows the test is consistent against fixed alternatives, attains a nontrivial limiting power $L(\beta)$ against local contiguous alternatives at rate $n^{-1/2}$, and remains consistent when $d$ grows with $n$ as long as $n\zeta_\sigma(P)$ diverges.

Load-bearing premise

The load-bearing premise is that the analyst can sample $X'_i$ from the true conditional distribution of $X$ given $Z_i$, independently of $Y_i$ given $Z_i$ for every observation, so that without exact model-X sampling the coordinate-swap exchangeability behind the finite-sample bound fails.

Editorial extensions

If this is right

  • Under exact model-X sampling, the test controls the type I error at level $\alpha$ for every finite $n$ and every dimension $d$, without smoothness or moment assumptions beyond the existence of $X\mid Z$.
  • The test is consistent against any fixed alternative because $\zeta_\sigma(P)>0$ under $H_1$ and the resampling critical value $\hat c_{1-\alpha}$ is $O_P(n^{-1})$.
  • Against local contiguous alternatives at distance $n^{-1/2}$, the power converges to a nontrivial limit $L(\beta)$ that increases from $\alpha$ to one as $\beta$ grows.
  • When the dimension $d$ grows with the sample size, power converges to one provided $n\zeta_\sigma(P)\to\infty$; dimension enters only through the size of the population discrepancy.
  • The randomized p-value $p_{n,B}$ approximates the exact flip p-value $p_n$ with error of order $O_P(B^{-1/2})$ plus $(B+1)^{-1}$, so moderate $B$ suffices in practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's abstract announces an investigation of estimating $P_{X\mid Z}$, but the full text proves no theorem for estimated or misspecified conditional distributions; if $X'_i$ is generated from an estimate or a generative model, the exchangeability behind Proposition 2.1 no longer follows, and a direct simulation under $H_0$ with misspecified sampling would show how rejection rates depart from
  • The coordinate-flip calibration is a general exchangeability principle: any statistic that is invariant under swapping each $X_i$ with $X'_i$ inherits the same finite-sample type I error argument, suggesting the exchangeable-pairs construction could be combined with dependence measures other than the Gaussian-kernel MMD.
  • Because $\zeta_\sigma(P)$ is a kernel MMD between a vector and its CI variant, a finite battery of bounded kernels (Laplace, inverse quadratic, or others) could in principle be used to detect different dependence geometries; the paper notes the extension to other bounded kernels but does not study the power of such a battery.
  • The model-X assumption is likely replaceable by a doubly robust variant in which only the regression of $X$ on $Z$ is estimated well, as is done in recent work on conditional randomization tests; testing that variant would be a natural next step beyond the paper's exact-exchangeability setting.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a conditional independence test based on an exchangeable-pairs construction in the model-X framework. For each observation (X_i,Y_i,Z_i), an exchangeable counterpart X'_i is drawn from the true conditional distribution of X given Z, independently of Y given Z. The population discrepancy ζσ(P) is the Gaussian-kernel MMD between (X,Y,Z) and its conditionally independent variant (X',Y,Z), and it is estimated by a bounded U-statistic on the augmented sample. A coordinate-swap resampling scheme calibrates the test, yielding a finite-sample type I error guarantee under exact model-X sampling (Proposition 2.1) and consistency against fixed alternatives (Proposition 2.3). The paper further claims asymptotic null and alternative distributions, local asymptotic power and "Pitman efficiency" against contiguous alternatives, high-dimensional consistency, and competitive empirical performance.

Significance. The core idea is attractive and, under the oracle model-X assumption, the central randomization argument is sound: the coordinate-swap p-value gives exact finite-sample type I error control, the estimator is a bounded U-statistic with exponential concentration, and the fixed-alternative consistency argument is simple and convincing. This provides a useful complement to the CRT that requires only one augmented sample and reuses pairwise distances. However, several advertised contributions are not actually established: the abstract promises an analysis of estimated X|Z that never appears, Theorem 2.5 is internally inconsistent, the label "Pitman efficiency" is not supported by any optimality comparison, and the local-power derivation has a substantial gap. With careful correction and re-scoping, the exchangeable-pairs methodology could be a solid contribution, but the current version overclaims.

major comments (4)
  1. [Abstract; Section 2.3; Remark 1] The abstract promises that the effect of estimating the conditional distribution used to generate the exchangeable pairs is investigated and that a condition preserving validity and power is established. No such theorem, lemma, or formal condition appears in the full text; Section 5 only says it would be interesting to study generative-model-based sampling. More seriously, Remark 1 suggests using GenAI to generate X given observations on (Y,Z), which makes X' dependent on Y given Z and destroys the coordinate-swap exchangeability on which Proposition 2.1 rests. This is both a missing promised result and an internal inconsistency: the finite-sample type I error guarantee is only established for oracle sampling from the true P_{X|Z}.
  2. [Theorem 2.5 and Appendix A] Theorem 2.5 states that under H1, sqrt(n)(ζ̂_{n,σ} − ζσ(P)) converges in distribution to N(0,1). The proof in Appendix A concludes asymptotic normality with variance 4σ1², where σ1² = Var(g1), and no argument shows that 4σ1² = 1. Since σ1² depends on P and σ, the limiting variance is not generally 1. The theorem is false as stated; it should either state N(0,4σ1²) or the statistic must be studentized by a consistent estimator of its asymptotic variance.
  3. [Section 3.1; Theorem 3.2] The label "Pitman efficient" is not justified. Theorem 3.2 derives a limiting power function L(β), but it does not compare the test with an optimal benchmark or compute an asymptotic relative efficiency, which is what Pitman efficiency standardly means. In addition, the abstract claims that the test attains the minimax separation rate for the proposed discrepancy measure; no minimax separation result is stated or proved anywhere in the text. These claims should be substantiated or removed.
  4. [Theorems 3.1 and 3.2; proof of Theorem 3.1] The local alternative F_{1−βn/√n} = (1−βn/√n)G + (βn/√n)F changes the conditional distribution of X|Z, but the procedure requires drawing X'_i from the true P_{X|Z}; under the mixture alternative, X'_i must be generated from the mixture conditional rather than from G_{X|Z}. The contiguity calculation in the proof of Theorem 3.1 computes the log-likelihood ratio only for V_i and omits the likelihood ratio contribution of the generated X'_i given Z_i, so the shift term βE_F{ψ_i} is not derived for the actual augmented-data experiment. Moreover, the displayed limiting covariance matrix in the same proof has a negative lower-right entry (−β²/2 E[...]) instead of a nonnegative variance. The local-power analysis needs to be reformulated (for example, by restricting to alternatives with fixed X|Z or by deriving the full likelihood ratio for the augmented data) before Theorems 3.1 and 3.2 can be considered established.
minor comments (5)
  1. [Abstract and throughout] There are numerous typographical errors, including "Keywards" in the abstract, "hypothesises" in Section 1, "afficacy" in Section 5, and "Piman efficiency" in the text following Theorem 3.2.
  2. [Section 5] The sentence "A test of spherical symmetry is also proposed" is incorrect; the paper proposes a test of conditional independence, not a test of spherical symmetry.
  3. [Theorem 3.2 statement and proof] The displayed power formula in Theorem 3.2(b) has the form (Z_i + βE_F{ψ_i})², but the proof in Appendix A writes "λ_i(Z_i + β²E[ψ_i(V_1,V'_1)] − 1)" and omits the square; the notation should be harmonized.
  4. [Figure 3] The legend uses "GCM.fix test" while the text defines "wGCM.fix test"; these should be made consistent.
  5. [Lemma A.1 and Remark 5] Lemma A.1 refers to "spherically symmetric variants" although the construction is coordinate swapping, not spherical symmetry. In Remark 5, the justification "V1 =D V'2" is incorrect; the cancellation follows from V1 =D V'1 when ζσ(G) = 0.

Circularity Check

0 steps flagged · score 0.0 of 10

No material circularity: the finite-sample validity and consistency follow from the model-X exchangeability construction and standard resampling arguments; the only self-citation is a minor technical lemma, not load-bearing.

full rationale

The derivation chain is self-contained given the model-X assumption. Definition 1 defines the CI variant by X'|Z =D X|Z and X' ⊥⊥ Y|Z; Theorem 2.1 proves ζσ(P)=0 iff X⊥⊥Y|Z via characteristic functions, not by assuming the conclusion. Eq. (4) estimates ζσ(P), and Theorems 2.3-2.5 establish concentration and limit laws by bounded-difference and U-statistic arguments. Proposition 2.1's P{p_n<α}≤α is the standard exact randomization-test property: under H0, coordinate-swapping X_i and X'_i preserves the joint law of (X_i,X'_i,Y_i,Z_i), so the resampling p-value in (6) is stochastically uniform; this is a theorem about the constructed statistic, not a fitted input renamed as a prediction. Proposition 2.2 bounds the resampling critical value by 2(α(n-1))^{-1}, and Proposition 2.3 combines this with consistency of ζ̂ to get power → 1. No uniqueness theorem, ansatz-via-citation, or fitted parameter carries the argument. The one self-citation ('as in Lemma 1 from Banerjee (2024)') appears only in a technical approximation step of the proof of Theorem 3.2 and is accompanied by standard references (Lee 1990; Van der Vaart 2000); it is not load-bearing. A separate completeness issue: the top abstract promises a study of estimating the conditional distribution used to generate X', but the full text contains no such theorem, and Remark 1's suggestion to generate X' from (Y,Z) is not covered by Proposition 2.1. This is an overclaim/limitation, not circularity.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The paper's central claim depends on external knowledge of X|Z and a user-chosen kernel bandwidth σ. No new physical or mathematical entities are introduced.

free parameters (1)
  • Gaussian kernel bandwidth σ
    Appears in ζ_σ(P) and the estimator; must be chosen by the user. Simulations do not specify the choice, and the test's power depends on it, especially in high dimension.
assumptions (3)
  • domain assumption Exact knowledge or accurate generative model for X|Z (model-X assumption)
    Required to sample CI variants X' that satisfy X'|Z = X|Z and X' ⊥ Y|Z. Invoked in Definition 1 and the augmentation step in Section 2.3; validity of Proposition 2.1 depends on it.
  • standard math Gaussian characteristic function identity used to close the MMD expression (Theorem 2.2)
    Standard Fourier identity; no issue.
  • domain assumption Contiguity and second-moment conditions on the likelihood ratio f/g in the local alternative analysis
    Assumed in Proposition 3.1 and used to apply Le Cam's third lemma in Theorem 3.1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Conditional Independence Testing Using Exchangeable Pairs." pith.science (2026). https://pith.science/paper/PBZBCN5I

@misc{pith2026250910817,
  author       = {Pith},
  title        = {Pith review of: Conditional Independence Testing Using Exchangeable Pairs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PBZBCN5I}},
  note         = {Machine review of arXiv:2509.10817}
}
abstract

This article considers the problem of testing conditional independence between two random vectors \(bm X\) and \(\bm Y\) given a confounding random vector \(\bm Z\). An exchangeable-pairs framework is introduced through which the conditional independence testing problem is reformulated as a two-sample testing problem. The framework is motivated by ideas from the model-X literature and is based on a fundamental exchangeability property that holds under the null hypothesis of conditional independence. An energy-distance/maximum mean discrepancy type measure is employed on the resulting exchangeable pairs to quantify departures from conditional independence. A consistent estimator of the proposed discrepancy measure is constructed and its theoretical properties are established under general assumptions. A conditional independence test is then developed using this estimator as a test statistic and is calibrated through a suitable resampling procedure. It is shown that the proposed test is consistent against fixed alternatives, possesses nontrivial asymptotic power against local contiguous alternatives, attains the minimax separation rate for detecting alternatives characterized by the proposed discrepancy measure, and remains consistent when the data dimension diverges with the sample size. The effect of estimating the conditional distribution used to generate the exchangeable pairs is also investigated, and condition under which validity and power properties are preserved is established. Extensive simulation studies demonstrate that the proposed procedure performs competitively with some state-of-the-art methods.

Figures

Figures reproduced from arXiv: 2509.10817 by the authors.

Figure 1
Figure 1. The contour of the probability density function of ( [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Power of the proposed test for different values of [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Powers of the AUG test ( ), the AUG.CRT test (♦), the cDC.CRT test (■), the AC test (♦), the GCM test (×), the GCM.fix test (▲) and the wGCM.est test (⋆) in Examples 1 (a)-(c). Example 2. Here, some bivariate distributions are considered with geometric dependence structures. Specifically, suppose Z ∼ Unif(0, 1), X = Z + η1, and Y = Z + η2, where (η1, η2) are generated from an equal mixture of two bivariate normal di… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Powers of the AUG test ( ), the AUG.CRT test (♦), the cDC.CRT test (■), the AC test (♦), the GCM test (×), the GCM.fix test (▲) and the wGCM.est test (⋆) in Examples 2 (a)-(b). Now consider a high-dimensional alternative of Example 2. Here, X, Y are uni-dimensional, bu…
Figure 5
Figure 5. Figure 5: Powers of the AUG test ( ), the AUG.CRT test (♦) and the AC test (♦) in Examples 3 (a) and (b). Example 3. Let X = f(Z) + η1 and Y = f(Z) + η2 where η1, η2 are i.i.d. observations as in Example 2, and Z is sampled from a d-dimensional standard normal distribution indep…
Figure 6
Figure 6. Figure 6: Powers of the AUG test ( ), the AUG.CRT test (♦) and the AC test (♦) in Examples 4 (a) and (b). Example 4. Consider the same random variables as in Example 3 but now generate n = d 2 + 20 many observations from the data where d is the dimension of the random vector Z. …

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CRT*: Conditional Randomization Testing with Heterogeneous External and Unlabeled Data

    stat.ME 2026-07 conditional novelty 6.0 of 10

    CRT* adaptively fuses internal, external, and unlabeled data via transfer learning and smooth residual bootstrap to give valid and more powerful conditional randomization tests under distributional heterogeneity.

Reference graph

Works this paper leans on

64 extracted references · 53 canonical work pages · cited by 1 Pith paper

  1. [1]

    and Chatterjee, S

    Azadkia, M. and Chatterjee, S. (2021). A simple measure of conditional dependence. Ann. Statist. , 49(6):3070--3102

  2. [2]

    Azadkia, M., Chen, L., and Han, F. (2025). Bias correction for chatterjee's graph-based correlation coefficient. arXiv preprint arXiv:2508.09040

  3. [3]

    Banerjee, B. (2024). Testing distributional equality for functional random variables. J. Multivariate Anal. , 203:105318

  4. [4]

    A Ball Divergence Based Measure For Conditional Independence Testing

    Banerjee, B., Bhattacharya, B. B., and Ghosh, A. K. (2024). A ball divergence based measure for conditional independence testing. arXiv preprint arXiv:2407.21456

  5. [5]

    Barber, R. F. and Cand\`es, E. J. (2019). A knockoff filter for high-dimensional selective inference. Ann. Statist. , 47(5):2504--2537

  6. [6]

    Bergsma, W. (2004). Testing conditional independence for continuous random variables . Report Eurandom. Eurandom

  7. [7]

    B., Wang, Y., Barber, R

    Berrett, T. B., Wang, Y., Barber, R. F., and Samworth, R. J. (2019). The conditional permutation test for independence while controlling for confounders. J. R. Stat. Soc., B: Stat. Methodol. , 82(1):175--197

  8. [8]

    Cai, Z., Li, R., and Zhang, Y. (2022). A distribution free conditional independence test with applications to causal discovery. J. Mach. Learn. Res. , 23:Paper No. [85], 41

Show all 64 references
  1. [9]

    Candes, E., Fan, Y., Janson, L., and Lv, J. (2018). Panning for gold: ‘model- X ’ knockoffs for high dimensional controlled variable selection. J. R. Stat. Soc., B: Stat. Methodol. , 80(3):551--577

  2. [10]

    Cook, R. D. and Li, B. (2002). Dimension reduction for conditional mean in regression. Ann. Statist. , 30(2):455--474

  3. [11]

    Dawid, A. P. (1980). Conditional independence for statistical operations. Ann. Statist. , 8(3):598--617

  4. [12]

    Deb, N., Ghosal, P., and Sen, B. (2020). Measuring association on topological spaces using kernels and geometric graphs. arXiv preprint arXiv:2010.01768

  5. [13]

    R., Yao, G., and West, M

    Dobra, A., Hans, C., Jones, B., Nevins, J. R., Yao, G., and West, M. (2004). Sparse graphical models for exploring gene expression data. Journal of Multivariate Analysis , 90(1):196--212

  6. [14]

    Doran, G., Muandet, K., Zhang, K., and Sch \"o lkopf, B. (2014). A permutation-based kernel conditional independence test. In UAI , pages 132--141

  7. [15]

    Fan, J., Feng, Y., and Xia, L. (2020). A projection-based conditional dependence measure with applications to high-dimensional undirected graphical models. J. Econometrics , 218(1):119--139

  8. [16]

    Fisher, R. A. (1924). The distribution of the partial correlation coefficient. Metron , 3:329--332

  9. [17]

    Fukumizu, K., Gretton, A., Sun, X., and Sch \"o lkopf, B. (2007). Kernel measures of conditional dependence. Advances in neural information processing systems , 20

  10. [18]

    M., Rasch, M

    Gretton, A., Borgwardt, K. M., Rasch, M. J., Sch \"o lkopf, B., and Smola, A. (2012). A kernel two-sample test. J. Mach. Learn. Res. , 13(1):723--773

  11. [19]

    Gy \"o rfi, L., Kohler, M., Krzy \.z ak, A., and Walk, H. (2002). A distribution-free theory of nonparametric regression . Springer

  12. [20]

    Huang, T.-M. (2010). Testing conditional independence using maximal nonlinear conditional correlation. Ann. Statist. , 38(4):2047--2091

  13. [21]

    Huang, Z., Deb, N., and Sen, B. (2022). Kernel partial correlation coefficient---a measure of conditional dependence. J. Mach. Learn. Res. , 23:Paper No. [216], 58

  14. [22]

    and Ramdas, A

    Katsevich, E. and Ramdas, A. (2022). On the power of conditional independence testing under model- X . Electron. J. Stat. , 16(2):6348--6394

  15. [23]

    and Sabatti, C

    Katsevich, E. and Sabatti, C. (2019). Multilayer knockoff filter: controlled variable selection at multiple resolutions. Ann. Appl. Stat. , 13(1):1--33

  16. [24]

    Kim, I., Neykov, M., Balakrishnan, S., and Wasserman, L. (2022). Local permutation tests for conditional independence. Ann. Statist. , 50(6):3388--3414

  17. [25]

    Lauritzen, S. L. (1996). Graphical models , volume 17. Clarendon Press

  18. [26]

    Lee, A. J. (1990). U - S tatistics: T heory and P ractice , volume 110 of Statistics: Textbooks and Monographs . Marcel Dekker, Inc., New York

  19. [27]

    Li, B. (2018). Sufficient dimension reduction: Methods and applications with R . Chapman and Hall/CRC

  20. [28]

    and Fan, X

    Li, C. and Fan, X. (2020). On nonparametric conditional independence tests for continuous variables. Wiley Interdiscip. Rev. Comput. Stat. , 12(3):e1489, 11

  21. [29]

    Li, S., Zhang, Y., Zhu, H., Wang, C., Shu, H., Chen, Z., Sun, Z., and Yang, Y. (2024). K-nearest-neighbor local sampling based conditional independence testing. Advances in Neural Information Processing Systems , 36

  22. [30]

    and Gozalo, P

    Linton, O. and Gozalo, P. (1996). Conditional independence restrictions: Testing and estimation. Technical report, Cowles Foundation for Research in Economics, Yale University

  23. [31]

    Liu, M., Katsevich, E., Janson, L., and Ramdas, A. (2021). Fast and powerful conditional randomization testing via distillation . 109(2):277--293

  24. [32]

    Maathuis, M., Drton, M., Lauritzen, S., and Wainwright, M. (2018). Handbook of graphical models . CRC Press

  25. [33]

    and Spang, R

    Markowetz, F. and Spang, R. (2007). Inferring cellular networks--a review. BMC bioinformatics , 8:1--17

  26. [34]

    Massart, P. (1990). The tight constant in the D voretzky- K iefer- W olfowitz inequality. Ann. Probab. , 18(3):1269--1283

  27. [35]

    Neykov, M., Balakrishnan, S., and Wasserman, L. (2021). Minimax optimal conditional independence testing. Ann. Statist. , 49(4):2151--2177

  28. [36]

    Niu, Z., Chakraborty, A., Dukes, O., and Katsevich, E. (2022). Reconciling model-x and doubly robust approaches to conditional independence testing. arXiv preprint arXiv:2211.14698

  29. [37]

    K., Sen, B., and Sz \'e kely, G

    Patra, R. K., Sen, B., and Sz \'e kely, G. J. (2016). On a nonparametric notion of residual and its applications. Statistics & Probability Letters , 109:208--213

  30. [38]

    Pearl, J. (1988). Probabilistic reasoning in intelligent systems: networks of plausible inference . Morgan Kaufmann, San Mateo, CA

  31. [39]

    Peters, J., Janzing, D., and Sch \"o lkopf, B. (2017). Elements of causal inference: foundations and learning algorithms . The MIT Press

  32. [40]

    and Hansen, N

    Petersen, L. and Hansen, N. R. (2021). Testing conditional independence via quantile regression based partial copulas. J. Mach. Learn. Res. , 22:Paper No. 70, 47

  33. [41]

    J., and Gretton, A

    Pogodin, R., Schrab, A., Li, Y., Sutherland, D. J., and Gretton, A. (2024). Practical kernel tests of conditional independence. arXiv preprint arXiv:2402.13196

  34. [42]

    Runge, J. (2018). Conditional independence testing based on a nearest-neighbor estimator of conditional mutual information. In International Conference on Artificial Intelligence and Statistics , pages 938--947. PMLR

  35. [43]

    and Ghosh, A

    Sarkar, S. and Ghosh, A. K. (2018). Some multivariate tests of independence based on ranks of nearest neighbors. Technometrics , 60(1):101--111

  36. [44]

    Scetbon, M., Meunier, L., and Romano, Y. (2022). An asymptotic test for conditional independence using analytic kernel embeddings. In International Conference on Machine Learning , pages 19328--19346. PMLR

  37. [45]

    o rrmann, J., and B\

    Scheidegger, C., H\" o rrmann, J., and B\" u hlmann, P. (2022). The weighted generalised covariance measure. J. Mach. Learn. Res. , 23:Paper No. [273], 68

  38. [46]

    Sesia, M., Sabatti, C., and Cand\`es, E. J. (2019). Gene hunting with hidden M arkov model knockoffs. Biometrika , 106(1):1--18

  39. [47]

    Shah, R. D. and Peters, J. (2020). The hardness of conditional independence testing and the generalised covariance measure. Ann. Statist. , 48(3):1514--1538

  40. [48]

    and Sriperumbudur, B

    Sheng, T. and Sriperumbudur, B. K. (2023). On distance and kernel measures of conditional dependence. Journal of Machine Learning Research , 24(7):1--16

  41. [49]

    Shi, H., Drton, M., and Han, F. (2022). On the power of chatterjee's rank correlation. Biometrika , 109(2):317--333

  42. [50]

    Shi, H., Drton, M., and Han, F. (2024). On A zadkia-- C hatterjee's conditional dependence coefficient. Bernoulli , 30(2):851--877

  43. [51]

    Song, K. (2009). Testing conditional independence via R osenblatt transforms. Ann. Statist. , 37(6B):4011--4045

  44. [52]

    V., Zhang, K., and Visweswaran, S

    Strobl, E. V., Zhang, K., and Visweswaran, S. (2019). Approximate kernel-based conditional independence tests for fast non-parametric causal discovery. Journal of Causal Inference , 7(1):20180017

  45. [53]

    and White, H

    Su, L. and White, H. (2007). A consistent characteristic function-based test for conditional independence. J. Econometrics , 141(2):807--834

  46. [54]

    and White, H

    Su, L. and White, H. (2008). A nonparametric H ellinger metric test for conditional independence. Econometric Theory , 24(4):829--864

  47. [55]

    and White, H

    Su, L. and White, H. (2014). Testing conditional independence via empirical likelihood. J. Econometrics , 182(1):27--44

  48. [56]

    Sz\' e kely, G. J. and Rizzo, M. L. (2014). Partial distance correlation with methods for dissimilarities. Ann. Statist. , 42(6):2382--2412

  49. [57]

    Sz \'e kely, G. J. and Rizzo, M. L. (2023). The Energy of Data and Distance Correlation . Chapman and Hall/CRC

  50. [58]

    and Han, F

    Tran, L. and Han, F. (2024). On a rank-based azadkia-chatterjee correlation coefficient. arXiv preprint arXiv:2412.02668

  51. [59]

    Van der Vaart, A. W. (2000). Asymptotic S tatistics , volume 3. Cambridge university press

  52. [60]

    Veraverbeke, N., Omelka, M., and Gijbels, I. (2011). Estimation of a conditional copula and association measures. Scandinavian Journal of Statistics , 38(4):766--780

  53. [61]

    Wainwright, M. J. (2019). High-dimensional statistics: A non-asymptotic viewpoint , volume 48 of Cambridge Series in Statistical and Probabilistic Mathematics . Cambridge University Press, Cambridge

  54. [62]

    Wang, X., Pan, W., Hu, W., Tian, Y., and Zhang, H. (2015). Conditional distance correlation. J. Amer. Statist. Assoc. , 110(512):1726--1734

  55. [63]

    Zhang, K., Peters, J., Janzing, D., and Sch \"o lkopf, B. (2012). Kernel-based conditional independence test and application in causal discovery. arXiv preprint arXiv:1202.3775

  56. [64]

    Zhang, Y., Huang, L., Yang, Y., and Shao, X. (2025). Doubly robust conditional independence testing with generative neural networks. J. R. Stat. Soc. Ser. B. Stat. Methodol. , page qkaf047

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.