Pith. sign in

REVIEW 4 cited by

LLMs Are Prone to Fallacies in Causal Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12158 v1 pith:OZPPC66O submitted 2024-06-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords causalrelationsllmsdatafactstemporalbeforecauses
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent work shows that causal facts can be effectively extracted from LLMs through prompting, facilitating the creation of causal graphs for causal inference tasks. However, it is unclear if this success is limited to explicitly-mentioned causal facts in the pretraining data which the model can memorize. Thus, this work investigates: Can LLMs infer causal relations from other relational data in text? To disentangle the role of memorized causal facts vs inferred causal relations, we finetune LLMs on synthetic data containing temporal, spatial and counterfactual relations, and measure whether the LLM can then infer causal relations. We find that: (a) LLMs are susceptible to inferring causal relations from the order of two entity mentions in text (e.g. X mentioned before Y implies X causes Y); (b) if the order is randomized, LLMs still suffer from the post hoc fallacy, i.e. X occurs before Y (temporal relation) implies X causes Y. We also find that while LLMs can correctly deduce the absence of causal relations from temporal and spatial relations, they have difficulty inferring causal relations from counterfactuals, questioning their understanding of causality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

    stat.ML 2026-07 conditional novelty 7.0 of 10

    CausalForge is a Lean-grounded, self-improving agentic framework that proposes, proves, and statement-audits causal inference theorems; its runs produced nine accepted results including a new ATE minimax upper bound.

  2. Do Large Language Models Show Biases in Causal Learning?

    cs.AI 2024-12 conditional novelty 6.0 of 10

    LLMs show a causal illusion bias: they often report causal relationships when evidence is purely correlational, null-contingency, or temporally impossible, particularly in 0-100 scaled judgments.

  3. Tagged for Direction: Pinning Down Causal Edge Directions with Precision

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Multi-tag variable annotations, generated by LLMs, allow causal discovery algorithms to orient undirected edges by transferring direction statistics from already directed edges with the same tag pairs.

  4. Reasoning Capabilities and Invariability of Large Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A new four-variant benchmark of simple geometric logic questions shows most LLMs score near chance, with performance largely stable across small language variations.

Pith tools