Pith. sign in

REVIEW 3 cited by

CRAB: Assessing the Strength of Causal Relationships Between Real-world Events

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.04284 v1 pith:PVK35LDR submitted 2023-11-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords causaleventscrabreasoningmodelsnarrativesreal-worldrelationships
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding narratives requires reasoning about the cause-and-effect relationships between events mentioned in the text. While existing foundation models yield impressive results in many NLP tasks requiring reasoning, it is unclear whether they understand the complexity of the underlying network of causal relationships of events in narratives. In this work, we present CRAB, a new Causal Reasoning Assessment Benchmark designed to evaluate causal understanding of events in real-world narratives. CRAB contains fine-grained, contextual causality annotations for ~2.7K pairs of real-world events that describe various newsworthy event timelines (e.g., the acquisition of Twitter by Elon Musk). Using CRAB, we measure the performance of several large language models, demonstrating that most systems achieve poor performance on the task. Motivated by classical causal principles, we also analyze the causal structures of groups of events in CRAB, and find that models perform worse on causal reasoning when events are derived from complex causal structures compared to simple linear causal chains. We make our dataset and code available to the research community.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLMs perform much worse on causal questions built from post-cutoff news articles, suggesting their apparent causal skill is mostly memorization, and a general-knowledge prompt method only partly closes the gap.

  2. WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts

    cs.CL 2025-06 conditional novelty 6.0 of 10

    WikiMixQA is a new 1,000-question benchmark for cross-modal table-and-chart reasoning, on which proprietary models drop from ~70% to ~55% accuracy when full Wikipedia pages are provided.

  3. Causal Graph based Event Reasoning using Semantic Relation Experts

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A multi-agent LLM debate with four semantic-relation experts builds causal event graphs that improve explainable event likelihood prediction and match fine-tuned models on forecasting and next-event prediction.

Pith tools