Pith. sign in

REVIEW 3 cited by

Comprehensive Review and Empirical Evaluation of Causal Discovery Algorithms for Numerical Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13054 v2 pith:L4UF7JSY submitted 2024-07-17 cs.AI

Comprehensive Review and Empirical Evaluation of Causal Discovery Algorithms for Numerical Data

classification cs.AI
keywords causaldiscoveryalgorithmscomprehensivedatamethodsalgorithmdatasets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Causal analysis has become an essential component in understanding the underlying causes of phenomena across various fields. Despite its significance, existing literature on causal discovery algorithms is fragmented, with inconsistent methodologies, i.e., there is no universal classification standard for existing methods, and a lack of comprehensive evaluations, i.e., data characteristics are often ignored to be jointly analyzed when benchmarking algorithms. This study addresses these gaps by conducting an exhaustive review and empirical evaluation for causal discovery methods on numerical data, aiming to provide a clearer and more structured understanding of the field. Our research begins with a comprehensive literature review spanning over two decades, analyzing over 200 academic articles and identifying more than 40 representative algorithms. This extensive analysis leads to the development of a structured taxonomy tailored to the complexities of causal discovery, categorizing methods into six main types. To address the lack of comprehensive evaluations, our study conducts an extensive empirical assessment of 29 causal discovery algorithms on multiple synthetic and real-world datasets. We categorize synthetic datasets based on size, linearity, and noise distribution, employing five evaluation metrics, and summarize the top-3 algorithm recommendations, providing guidelines for users in various data scenarios. Our results highlight a significant impact of dataset characteristics on algorithm performance. Moreover, a metadata extraction strategy with an accuracy exceeding 80% is developed to assist users in algorithm selection on unknown datasets. Based on these insights, we offer professional and practical guidelines to help users choose the most suitable causal discovery methods for their specific dataset.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations

    cs.LG 2026-05 unverdicted novelty 7.0

    TCD-Arena is a new customizable testing framework that runs millions of experiments to map how 33 different assumption violations affect time series causal discovery methods and shows ensembles can boost overall robustness.

  2. DKCD: Domain Knowledge-Enhanced Causal Discovery from Unstructured Data

    cs.CL 2026-07 conditional novelty 6.0

    DKCD uses retrieved domain-KG subgraphs to let LLMs discover latent factors and generate causal clues, yielding higher node F1 and lower extended SHD than COAT-style baselines on two synthetic medical datasets.

  3. AutoCause: A Python framework that automates expert decisions in environmental time-series causal discovery

    cs.LG 2026-07 conditional novelty 5.0

    An auditable consensus workflow around four causal-discovery methods gives higher-precision links on synthetic benchmarks, but the precision gain disappears on a real river network where the reference graph is incomplete.