Pith. sign in

REVIEW 5 cited by

Understanding Causality with Large Language Models: Feasibility and Opportunities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.05524 v1 pith:DJ535NPP submitted 2023-04-11 cs.LG cs.CL

classification cs.LGcs.CL
keywords causalllmsanswerquestionsenableknowledgelanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We assess the ability of large language models (LLMs) to answer causal questions by analyzing their strengths and weaknesses against three types of causal question. We believe that current LLMs can answer causal questions with existing causal knowledge as combined domain experts. However, they are not yet able to provide satisfactory answers for discovering new knowledge or for high-stakes decision-making tasks with high precision. We discuss possible future directions and opportunities, such as enabling explicit and implicit causal modules as well as deep causal-aware LLMs. These will not only enable LLMs to answer many different types of causal questions for greater impact but also enable LLMs to be more trustworthy and efficient in general.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Linear-LLM-SCM: Benchmarking LLMs for Coefficient Elicitation in Linear-Gaussian Causal Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    LLMs asked to fill in linear-Gaussian causal equations give inaccurate, unstable, and perturbation-sensitive coefficients; the open-source Linear-LLM-SCM benchmark measures this, with Gemini 2.5 Flash leading on scale...

  2. Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLMs perform much worse on causal questions built from post-cutoff news articles, suggesting their apparent causal skill is mostly memorization, and a general-knowledge prompt method only partly closes the gap.

  3. LLM Cannot Discover Causality, and Should Be Restricted to Non-Decisional Support in Causal Discovery

    cs.LG 2025-06 conditional novelty 6.0 of 10

    LLMs are unreliable causal reasoners, so they should be limited to non-decisional search support in causal discovery algorithms.

  4. Paths to Causality: Finding Informative Subgraphs Within Knowledge Graphs for Knowledge-Based Causal Discovery

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A learning-to-rank model chooses informative knowledge-graph paths between entity pairs, and adding the top path to zero-shot prompts improves LLM causal classification by up to 44.4 F1 points.

  5. Causal Distillation: Transferring Structured Explanations from Large to Compact Language Models

    cs.CL 2025-05 reject novelty 3.0 of 10

    Small language models fine-tuned on GPT-4 causal explanations score high on a new teacher-similarity metric, but the paper provides no independent evidence that causal reasoning was transferred.

Pith tools