Pith. sign in

REVIEW 2 cited by

CRAFT: A Benchmark for Causal Reasoning About Forces and inTeractions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.04293 v3 pith:BZC5SDMF submitted 2020-12-08 cs.AI cs.CLcs.CV

classification cs.AIcs.CLcs.CV
keywords causalcraftquestioninteractionsbenchmarkforceshumansintroduce
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans are able to perceive, understand and reason about causal events. Developing models with similar physical and causal understanding capabilities is a long-standing goal of artificial intelligence. As a step towards this direction, we introduce CRAFT, a new video question answering dataset that requires causal reasoning about physical forces and object interactions. It contains 58K video and question pairs that are generated from 10K videos from 20 different virtual environments, containing various objects in motion that interact with each other and the scene. Two question categories in CRAFT include previously studied descriptive and counterfactual questions. Additionally, inspired by the Force Dynamics Theory in cognitive linguistics, we introduce a new causal question category that involves understanding the causal interactions between objects through notions like cause, enable, and prevent. Our results show that even though the questions in CRAFT are easy for humans, the tested baseline models, including existing state-of-the-art methods, do not yet deal with the challenges posed in our benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment

    cs.AI 2026-07 conditional novelty 6.0 of 10

    VAORA aligns VLM chain-of-thought reasoning with visual scene observations and post-action outcomes via structured symbolic rewards, achieving cross-task and cross-environment generalization on physical reasoning benchmarks.

  2. Think before You Simulate: Symbolic Reasoning to Orchestrate Neural Computation for Counterfactual Question Answering

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A neuro-symbolic system constructs a causal graph from perceived collisions and uses it to selectively start physical simulation, improving counterfactual question answering on CLEVRER and CRAFT.

Pith tools