Pith. sign in

REVIEW 3 cited by

Integrating Large Language Models in Causal Discovery: A Statistical Causal Approach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.01454 v5 pith:6ERZ64TP submitted 2024-02-02 cs.LG cs.AIstat.MEstat.ML

classification cs.LGcs.AIstat.MEstat.ML
keywords causalknowledgedatasetapproachchallengesdomaininferencellms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In practical statistical causal discovery (SCD), embedding domain expert knowledge as constraints into the algorithm is important for reasonable causal models reflecting the broad knowledge of domain experts, despite the challenges in the systematic acquisition of background knowledge. To overcome these challenges, this paper proposes a novel method for causal inference, in which SCD and knowledge-based causal inference (KBCI) with a large language model (LLM) are synthesized through ``statistical causal prompting (SCP)'' for LLMs and prior knowledge augmentation for SCD. The experiments in this work have revealed that the results of LLM-KBCI and SCD augmented with LLM-KBCI approach the ground truths, more than the SCD result without prior knowledge. These experiments have also revealed that the SCD result can be further improved if the LLM undergoes SCP. Furthermore, with an unpublished real-world dataset, we have demonstrated that the background knowledge provided by the LLM can improve the SCD on this dataset, even if this dataset has never been included in the training data of the LLM. For future practical application of this proposed method across important domains such as healthcare, we also thoroughly discuss the limitations, risks of critical errors, expected improvement of techniques around LLMs, and realistic integration of expert checks of the results into this automatic process, with SCP simulations under various conditions both in successful and failure scenarios. The careful and appropriate application of the proposed approach in this work, with improvement and customization for each domain, can thus address challenges such as dataset biases and limitations, illustrating the potential of LLMs to improve data-driven causal inference across diverse scientific domains. The code used in this work is publicly available at: www.github.com/mas-takayama/LLM-and-SCD

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM Cannot Discover Causality, and Should Be Restricted to Non-Decisional Support in Causal Discovery

    cs.LG 2025-06 conditional novelty 6.0 of 10

    LLMs are unreliable causal reasoners, so they should be limited to non-decisional search support in causal discovery algorithms.

  2. Causal-Invariant Cross-Domain Out-of-Distribution Recommendation

    cs.IR 2025-05 conditional novelty 6.0 of 10

    CICDOR learns two causal DAGs for shared and domain-specific user preferences, uses an LLM guided by the FCI algorithm to extract confounders from reviews, and reports consistent accuracy gains over twelve baselines o...

  3. Paths to Causality: Finding Informative Subgraphs Within Knowledge Graphs for Knowledge-Based Causal Discovery

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A learning-to-rank model chooses informative knowledge-graph paths between entity pairs, and adding the top path to zero-shot prompts improves LLM causal classification by up to 44.4 F1 points.

Pith tools