Pith. sign in

REVIEW 4 cited by

CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.06177 v2 pith:MGCOCDUT submitted 2019-08-16 cs.LG cs.CLcs.LOstat.ML

classification cs.LGcs.CLcs.LOstat.ML
keywords clutrrmodelrobustnessallowsbenchmarkdiagnosticgeneralizationinductive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent success of natural language understanding (NLU) systems has been troubled by results highlighting the failure of these models to generalize in a systematic and robust way. In this work, we introduce a diagnostic benchmark suite, named CLUTRR, to clarify some key issues related to the robustness and systematicity of NLU systems. Motivated by classic work on inductive logic programming, CLUTRR requires that an NLU system infer kinship relations between characters in short stories. Successful performance on this task requires both extracting relationships between entities, as well as inferring the logical rules governing these relationships. CLUTRR allows us to precisely measure a model's ability for systematic generalization by evaluating on held-out combinations of logical rules, and it allows us to evaluate a model's robustness by adding curated noise facts. Our empirical results highlight a substantial performance gap between state-of-the-art NLU models (e.g., BERT and MAC) and a graph neural network model that works directly with symbolic inputs---with the graph-based model exhibiting both stronger generalization and greater robustness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

    cs.AI 2025-11 unverdicted novelty 7.0 of 10

    DecompSR is a large, symbolically verified benchmark dataset and generation framework that independently varies productivity, substitutivity, overgeneralisation, and systematicity to probe compositional multihop spati...

  2. The Knowledge-Reasoning Dissociation: Fundamental Limitations of LLMs in Clinical Natural Language Inference

    cs.AI 2025-08 reject novelty 6.0 of 10

    Across four clinical inference tasks, six LLMs answer paired knowledge probes at 92% accuracy but the main reasoning tasks at 25%, indicating a systematic knowledge-reasoning gap.

  3. A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning

    cs.AI 2025-08 unverdicted novelty 4.0 of 10

    A comparison of two neurosymbolic designs concludes that the hybrid design, pairing an LLM with a separate symbolic solver, is the more promising path to general logical reasoning.

  4. Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning

    cs.LG 2025-06 reject novelty 4.0 of 10

    GoCE swaps CoM's chain structure for a differentiable causal graph and reports accuracy gains on CLUTRR, CLadder, EX-FEVER, and CausalQA, but the evidence is sandbox-generated and unauditable.

Pith tools