Pith. sign in

REVIEW 7 cited by

Causal Evaluation of Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.00622 v1 pith:4FCW7KVR submitted 2024-05-01 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords causallanguagemodelscalmevaluationresultsacrossadaptations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Causal reasoning is viewed as crucial for achieving human-level machine intelligence. Recent advances in language models have expanded the horizons of artificial intelligence across various domains, sparking inquiries into their potential for causal reasoning. In this work, we introduce Causal evaluation of Language Models (CaLM), which, to the best of our knowledge, is the first comprehensive benchmark for evaluating the causal reasoning capabilities of language models. First, we propose the CaLM framework, which establishes a foundational taxonomy consisting of four modules: causal target (i.e., what to evaluate), adaptation (i.e., how to obtain the results), metric (i.e., how to measure the results), and error (i.e., how to analyze the bad results). This taxonomy defines a broad evaluation design space while systematically selecting criteria and priorities. Second, we compose the CaLM dataset, comprising 126,334 data samples, to provide curated sets of causal targets, adaptations, metrics, and errors, offering extensive coverage for diverse research pursuits. Third, we conduct an extensive evaluation of 28 leading language models on a core set of 92 causal targets, 9 adaptations, 7 metrics, and 12 error types. Fourth, we perform detailed analyses of the evaluation results across various dimensions (e.g., adaptation, scale). Fifth, we present 50 high-level empirical findings across 9 dimensions (e.g., model), providing valuable guidance for future language model development. Finally, we develop a multifaceted platform, including a website, leaderboards, datasets, and toolkits, to support scalable and adaptable assessments. We envision CaLM as an ever-evolving benchmark for the community, systematically updated with new causal targets, adaptations, models, metrics, and error types to reflect ongoing research advancements. Project website is at https://opencausalab.github.io/CaLM.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

    stat.ML 2026-07 conditional novelty 7.0 of 10

    CausalForge is a Lean-grounded, self-improving agentic framework that proposes, proves, and statement-audits causal inference theorems; its runs produced nine accepted results including a new ATE minimax upper bound.

  2. IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data

    cs.CL 2025-06 conditional novelty 6.0 of 10

    IP-Dialog is a synthetic dialogue benchmark, with 1,000 test and 10,790 training items, for evaluating implicit personalization: inferring hidden user attributes from history and tailoring responses.

  3. LLM Cannot Discover Causality, and Should Be Restricted to Non-Decisional Support in Causal Discovery

    cs.LG 2025-06 conditional novelty 6.0 of 10

    LLMs are unreliable causal reasoners, so they should be limited to non-decisional search support in causal discovery algorithms.

  4. EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models

    cs.CL 2025-10 reject novelty 5.0 of 10

    A submitted paper whose abstract promises a context-aware economic causal-reasoning benchmark, but whose body implements a different four-task benchmark with no context annotations.

  5. Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models

    cs.CL 2025-06 reject novelty 5.0 of 10

    ProCI uses LLMs to iteratively generate and impute hidden confounders, then validates them with a conditional independence test to improve treatment effect estimation.

  6. Causal MAS: A Survey of Large Language Model Architectures for Discovery and Effect Estimation

    cs.AI 2025-08 conditional novelty 3.0 of 10

    A structured survey that defines and catalogs multi-agent LLM systems for causal reasoning, discovery, and effect estimation, including their architectures, benchmarks, and applications.

  7. Causal Distillation: Transferring Structured Explanations from Large to Compact Language Models

    cs.CL 2025-05 reject novelty 3.0 of 10

    Small language models fine-tuned on GPT-4 causal explanations score high on a new teacher-similarity metric, but the paper provides no independent evidence that causal reasoning was transferred.

Pith tools