{"id":"dc2dbce6-ad31-44b1-9b07-5a92b89cf6fe","arxiv_id":"2608.11156","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that organizes conditional independence tests for constraint-based causal discovery into six families and connects test-level properties to graph-level errors.","lead":"This paper surveys conditional independence tests, the statistical checks that let causal discovery algorithms like PC and FCI decide whether two variables are directly linked or only connected through other variables. It organizes these tests into six families, compares their software support, and gives practical guidance for high-dimensional and mixed-type biomedical data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Decision diagrams (Figs 6–8) are not validated at graph level; the survey's own §10.5 concedes the required systematic comparison is missing, so the practitioner guidance rests on an untested transfer assumption.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the practitioner-facing decision frameworks assume that isolated per-test properties transfer to integrated constraint-based discovery. I agree, and the manuscript itself confirms the gap in Section 10.5. This concern is not fatal because the survey is explicitly a synthesis and its open-problems section is honest, but it does justify a conditional acceptance: the central practical deliverable should either be accompanied by systematic empirical validation or be explicitly framed as provisional. I considered the Benjamini-Hochberg statement in Section 6.2.5 as an alternative concern; it is factually imprecise (BH controls FDR under positive regression dependency, not only under independence), but the surrounding recommendation to prefer Benjamini-Yekutieli for the dependent CI-test sequence in PC remains defensible, so I do not treat it as load-bearing. I found no internal inconsistency in the survey's review of CI test families, and the paper's consolidation of software availability and test-family properties is a genuine contribution. The appropriate verdict remains CONDITIONAL, unchanged from the reader's assessment.","tokens_in":27879,"tokens_out":4658,"duration_ms":47421,"concrete_test":"Replicate the Raghu et al. (2018) mixed-data benchmark setup with the full set of CI test families (Fisher's Z, G2/chi-square, regression/GCM, KNN-CMI, RCoT/RCIT, and KCI) plugged into the same PC implementation, and vary n in {100, 500, 2000}, maximum conditioning-set depth in {2, 4, 6}, and graph density in {sparse, dense}. Compare per-test calibration and power against graph-level SHD, F1, and orientation accuracy within each regime specified by Figures 6–8; if the recommended test family does not consistently yield the best graph-level metrics in its regime, the decision diagrams are unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central practical contribution is the heuristic decision framework in Figures 6–8 (contribution 4 in §2), which translates per-test properties into recommended CI tests for continuous, categorical, and mixed data. For these recommendations to be correct, per-test calibration and power must transfer to graph-level performance inside iterative algorithms such as PC and FCI. This transfer is not automatic: §6.2.2 and §6.2.5 explain that type I and type II errors propagate asymmetrically and that repeated adaptive testing changes the operating characteristics of a test. A test that is well calibrated in isolation can behave differently when hypotheses are selected from the same data, when conditioning sets are grown adaptively, and when hundreds of tests are performed; §6.2.5 itself notes that graph-level error is not controlled by per-test type I rate. The survey does not supply evidence that the regimes in Figures 6–8—keyed to data type, compute budget, and conditioning-set size—actually minimize graph-level error. Section 10.5 explicitly concedes that 'a systematic empirical comparison ... is still lacking' and that the proposed heuristics 'need confirmation or refinement.' Since contribution 4 is a stated goal and the figures are its deliverable, this unvalidated transfer is the load-bearing weak point of the survey's central claim that graph quality is governed by CI test choice. This is a conditional concern rather than a rejection: the theoretical properties reviewed are relevant, but the practitioner-facing guidance outruns the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey treats conditional independence testing as the core statistical engine of constraint-based causal discovery, arguing that the validity of CI decisions controls skeleton recovery and edge orientation in algorithms such as PC and FCI. It organizes the literature into six families—partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based—and discusses their assumptions, calibration, computational cost, and robustness layers. It also links test-level properties such as power decay with conditioning-set size and asymmetric type I/II error costs to graph-level failure modes, compares software support across R and Python libraries, and closes with practitioner-oriented decision diagrams and open problems. The paper is a synthesis rather than a new methodological contribution.","tokens_in":28091,"tokens_out":6837,"duration_ms":65751,"significance":"If the issues below are addressed, this survey would be a genuinely useful reference for entering researchers and practitioners. Its strengths are a clear family taxonomy, compact and mostly accurate descriptions of each CI test with appropriate citations, detailed software comparison tables, and a candid treatment of open problems—including, notably, the admission in Section 10.5 that the survey's own decision heuristics are not yet validated. The organizing theme that per-test CI behavior propagates to graph-level errors is valuable and gives the survey a coherent message. The paper does not introduce fitted parameters or new entities, and its claims are summaries of published results, so circularity is not a concern.","major_comments":[{"comment":"The statement that 'Benjamini–Hochberg, meanwhile, controls the FDR only for independent test statistics' is incorrect. The Benjamini–Hochberg procedure controls the FDR under positive regression dependence on the subset of true null hypotheses, and under arbitrary dependence it still controls the FDR at level (m0/m) times the nominal level. The true issue for PC is whether the collection of CI test statistics satisfies the required dependence conditions, not that BH requires independence. This paragraph should be corrected and the motivation for recommending Benjamini–Yekutieli rephrased accordingly.","section":"Section 6.2.5"},{"comment":"The practitioner decision frameworks in Figures 6–8 are presented as contribution 4, but the manuscript itself concedes in Section 10.5 that a systematic empirical comparison across data types, sample sizes, conditioning-set depths, and graph densities is still lacking, and that the proposed heuristics 'need confirmation or refinement.' This matters because Sections 6.2.2 and 6.2.5 explain that type I/II errors propagate asymmetrically through iterative discovery and that repeated adaptive testing changes the operating characteristics of a test; therefore, per-test calibration and power do not automatically transfer to graph-level performance in PC or FCI. I recommend either clearly reframing Figures 6–8 as provisional hypotheses with the Section 10.5 caveat repeated in Section 6.3 and the figure captions, or adding a focused validation study. As written, the practical guidance is stronger than the evidence presented.","section":"Section 2 / Figures 6–8 / Section 10.5"}],"minor_comments":[{"comment":"The degrees-of-freedom expression for G2 and chi-square is ambiguous as typeset: it should explicitly sum over the strata of the conditioning variable Z and clarify whether the category counts r_x and r_y are stratum-dependent.","section":"Section 5.2"},{"comment":"The claim that 'empirical evidence suggests that KNN-CMI is better calibrated at small |Z|, while RCoT/RCIT achieve higher power at large |Z|' needs a precise pointer to the specific experiments in Runge (2018); as written, the claim is not checkable.","section":"Figure 6 caption"},{"comment":"Saying that small-sample CI tests can exhibit 'inflated false discovery rates' is imprecise; the per-test quantity that inflates is the false positive rate, whereas FDR is a property of a multiple-testing procedure.","section":"Section 10.2"},{"comment":"The decision diagrams use qualitative labels such as 'Small', 'Med.', and 'Large' without any quantitative anchors; adding rough sample-size or conditioning-depth thresholds would make the guidance more actionable.","section":"Section 6.3"}],"recommendation":"major_revision","confidential_remarks":"The survey is within scope and has no circularity or novelty concerns. The main revision burden is to align the strength of the practitioner guidance with the admitted lack of validation and to correct the false claim about Benjamini–Hochberg. With those changes, the paper could be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, useful survey of conditional independence testing for constraint-based causal discovery, and it deserves real peer review. The organizing choice—treating CI tests as the load-bearing statistical component of PC/FCI—pays off. The six-family taxonomy (partial correlation, contingency table, regression, KNN, kernel, ML-based) is accurate and well illustrated, and the linking of per-test power and error asymmetries to skeleton and orientation failures is one of the clearest treatments I've seen in one place. The software availability tables and the practitioner-facing decision diagrams are genuinely new artifacts, even if they are compilations rather than new results.\n\nThe main soft spot is the multiple-testing discussion in Section 6.2.5. The paper says Benjamini-Hochberg controls the FDR only for independent test statistics. That is not right; BH controls FDR under positive regression dependence, not merely independence. The broader point—that PC's CI tests are not independent and may not satisfy PRDS, making BH unsafe—is still valid, but the stated condition is wrong and should be corrected. Minor fix, but it matters because the section is otherwise careful.\n\nThe other concern, which the authors themselves flag in Section 10.5, is that the decision diagrams in Figures 6–8 are heuristic. They rest on the assumption that per-test calibration and power transfer to graph-level performance inside iterative discovery algorithms. That transfer is not automatic, and the authors concede that a systematic empirical comparison is still missing. For a survey this is acceptable, but the labeling could be more prominent: I'd like to see a one-line disclaimer directly under each figure saying the recommendations are provisional and unvalidated at the graph level. As written, the paper does own this limitation, so it is a moderate caveat rather than a fatal flaw.\n\nThe survey is honest about what it does not cover, the references are broad and appropriate (including Shah & Peters, Runge, Colombo & Maathuis), and the authors' self-citations appear only as application examples, not as support for the methodological core. No circularity issue.\n\nWho should read this: PhD students entering causal discovery, and practitioners who need a map of CI test options and their trade-offs. It consolidates scattered package documentation and gives a workable starting point for test selection. It is not a research breakthrough, but it is a credible reference survey.\n\nRecommendation: send to serious peer review. With the BH correction and clearer hedging on the figures, I would accept with minor revisions.","headline":"A genuinely useful CI-test survey with an accurate core synthesis, a real but local technical error about Benjamini-Hochberg, and heuristic guidance that is honestly labeled as unvalidated.","tokens_in":28632,"tokens_out":2420,"would_cite":true,"duration_ms":24381,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H20","62H15","62G10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Conditional-independence tests decide every edge in constraint-based causal discovery, and six families of tests come with sharply different failure modes.","keywords":["conditional independence testing","constraint-based causal discovery","PC algorithm","FCI algorithm","skeleton recovery","v-structure orientation","mixed-type data","statistical power"],"falsifier":"A controlled benchmark holding the discovery algorithm fixed and varying only the CI-test family, sample size, conditioning-set depth, and graph density, on standard benchmark networks and semi-synthetic mixed-type data, would settle the central link: if per-test calibration and power rankings stop predicting graph-level skeleton and orientation accuracy—for example, a better-calibrated test producing worse graphs—the survey's core claim about CI tests as the decisive engine would fail. Section 10.5 notes that exactly this experiment is still missing.","tokens_in":27636,"feed_emoji":"🧬","tokens_out":9047,"duration_ms":77293,"temperature":0.7,"pith_summary":"Conditional-independence tests are the statistical engine of constraint-based causal discovery: algorithms such as PC (Peter-Clark), FCI (Fast Causal Inference), and Grow-Shrink remove edges and orient v-structures purely on the outcome of tests of the form $X \\perp Y \\mid Z$. This survey argues that graph quality is only as good as the validity of those CI decisions, and organizes the field into six test families—partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based—to make the choice tractable. It connects test-level behavior (power decaying as conditioning sets grow, asymmetric consequences of false positives and false negatives) to graph-level failures (spurious or missing edges, wrong v-structure orientations). It then translates those links into heuristic decision diagrams for continuous, categorical, and mixed-type data, and compares which families are actually available in R and Python software. A reader should care because the same CI test choice that governs every edge also shapes interpretability, scalability, and reliability in high-stakes biomedical applications.","feed_headline":"Every edge in a causal graph rests on a conditional-independence test","feed_subtitle":"Wrong conditional-independence tests corrupt skeletons and orientations; the survey maps the trade-offs.","key_machinery":"The load-bearing object is the conditional-independence test itself, treated as the engine of skeleton pruning and orientation. The paper's organizing machinery is a six-family taxonomy—partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based—augmented by robustness layers (shrinkage, permutation, test aggregation) that wrap base tests without becoming separate families. The mechanism that carries the argument from tests to graphs is the decision chain: a test of $X \\perp Y \\mid Z$ removes or keeps an edge, the conditioning set that yields independence becomes the separating set for that edge, separating sets determine whether an unshielded triple is oriented as a v-structure, and orientations propagate through Meek's rules into a completed partially directed acyclic graph or partial ancestral graph. Within that chain, power decay of every family as $|Z|$ grows, and the asymmetric impact of type I versus type II errors, are the quantities that explain graph-level failure modes.","core_discovery":"On its own terms, the paper claims that conditional-independence testing, not scoring or global optimization, is the decisive statistical component of constraint-based causal discovery. In PC, FCI, and related algorithms, a CI test outcome either keeps or deletes an edge, stores the separating set, and later sets the orientation of unshielded triples; those orientations then propagate through orientation rules into the final completed partially directed acyclic graph or partial ancestral graph. The survey's central assertion is that the validity of the learned graph is only as good as the validity of these CI decisions, and that test selection is the main lever a practitioner controls. It reviews six families of tests, states for each when its decisions reflect the data-generating distribution and when they fail, and shows that errors propagate asymmetrically: false positives add spurious edges, while false negatives delete true edges and can destroy separating sets needed for correct v-structure orientation, with cascading effects through subsequent orientation rules. It closes with heuristic selection frameworks and an explicit list of open problems, mixed-type testing without discretization and small-sample error control among them.","pith_inferences":["The paper's Section 10.5 concedes that the decision diagrams in Figures 6–8 have not been validated by a systematic empirical comparison; the natural next step is a controlled benchmark that plugs each family into the same PC/FCI implementation while varying sample size, conditioning depth, and graph density.","The asymmetric-error argument suggests that safety-critical applications should tune $\\alpha$ toward avoiding false negatives rather than toward global false-positive control, because type II errors cascade through orientation—a consequence the survey notes but does not quantify.","If power decay with $|Z|$ is as steep as the survey describes, algorithmic restructuring that reduces conditioning depth (local Markov-blanket search, divide-and-conquer skeleton learning) could improve graph accuracy even without improving per-test power.","The gap between methodological development and tooling implies that near-term practical gains are more likely to come from robust wrappers around existing tests (permutation, wild bootstrap, aggregation) than from entirely new CI tests."],"forward_implications":["Any change in CI-test calibration propagates directly into the graph: inflated false positives add spurious edges, and inflated false negatives remove true edges and can erase the separating sets that later v-structure orientation depends on.","Because power decays with conditioning-set size for every family, capping $|Z|$ or switching to Markov-blanket and local discovery methods is the practical safeguard when the sample size is small relative to graph degree.","For continuous Gaussian-like data Fisher's $Z$ is the fastest and best-powered default; for categorical data with small-to-moderate conditioning sets $G^2/\\chi^2$ tests are effective; for nonlinear continuous data the choice hinges on budget and conditioning depth (KNN-CMI for small $|Z|$, RCIT/RCoT for large $|Z|$, full KCI only with ample compute).","No major library surveyed ships a native machine-learning-based CI test, so DML-style and generative tests are currently research tools rather than drop-in options for practitioners.","Reporting the CI test, its significance level, the maximum conditioning-set size, and key separating sets should become standard practice, because the audit trail of a constraint-based graph is exactly the sequence of CI decisions."],"supporting_citations":[{"why":"Supplies the foundational assumption framework (Causal Markov, Faithfulness, Causal Sufficiency) and the PC/FCI pipeline that CI tests drive.","marker":"(Spirtes et al., 2000)"},{"why":"Provides the high-dimensional consistency analysis of PC and the shrinking significance level $\\alpha_n$ that frames the power and multiple-testing discussion.","marker":"(Kalisch & Bühlmann, 2007)"},{"why":"Supplies Stable-PC order-independent skeleton discovery and the caveat that per-test error control does not guarantee graph-level accuracy.","marker":"(Colombo & Maathuis, 2014)"},{"why":"Establishes that unrestricted CI testing is impossible and introduces the Generalised Covariance Measure used as the ML-based family's core.","marker":"(Shah & Peters, 2020)"},{"why":"Introduces the KNN conditional-mutual-information estimator and permutation test, the basis for the nearest-neighbor family and for the Figure 6 split between KNN-CMI and RCoT/RCIT.","marker":"(Runge, 2018)"},{"why":"Supplies regression-based CI tests for mixed data, the main route the survey recommends for mixed-type discovery without discretization.","marker":"(Tsagris et al., 2018)"},{"why":"Provides RCIT/RCoT randomized kernel tests with approximately linear scaling, the practical option for nonlinear continuous data at large conditioning depth.","marker":"(Strobl et al., 2019b)"},{"why":"Formalizes when zero partial correlation is equivalent to conditional independence, delimiting the validity of the most widely used test family.","marker":"(Baba et al., 2004)"}],"fun_headline_variants":["Causal discovery lives or dies on CI test choice","CI tests decide every edge and orientation in causal graphs","Survey maps six CI test families and their trade-offs","Your causal graph is only as good as your CI test","Wrong CI tests corrupt causal graphs; right ones build them"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The guidance assumes that the strengths and weaknesses of each test family, measured in isolation, still hold when the same test is embedded in an iterative PC- or FCI-style search across real mixed-type datasets—an assumption the paper itself states has not been systematically validated.","fun_headline_variants_meta":{"raw":{"variants":["Causal discovery lives or dies on CI test choice","CI tests decide every edge and orientation in causal graphs","Survey maps six CI test families and their trade-offs","Your causal graph is only as good as your CI test","Wrong CI tests corrupt causal graphs; right ones build them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001106,"raw_usage":{"total_tokens":4613,"prompt_tokens":950,"completion_tokens":3663,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":3585}},"tokens_in":566,"tokens_out":3663,"duration_ms":24979,"temperature":1.0,"reasoning_tokens":3585,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:05:20.898045+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled benchmark holding the discovery algorithm fixed and varying only the CI-test family, sample size, conditioning-set depth, and graph density, on standard benchmark networks and semi-synthetic mixed-type data, would settle the central link: if per-test calibration and power rankings stop predicting graph-level skeleton and orientation accuracy—for example, a better-calibrated test producing worse graphs—the survey's core claim about CI tests as the decisive engine would fail. Section 10.5 notes that exactly this experiment is still missing.","supporting_citations":[],"review_version":1}