Pith. sign in

Paper Citation Record · LEDGER

CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:1908.06177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.06177 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:37:02.501468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T20:57:23.984225Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 67de0736-678d-41c7-b20d-abf295c31a6c · inbound

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity cites this paper.

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:58:48.132125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T19:58:48.066562Z digest=sha256:c48f258733ac05203a2e4d0e7ef03e054f00c81bc45730333eacbb420e5c69fd

Observation 29a4fef6-7651-49b6-a773-479980ce5c6b · inbound

Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning cites this paper.

Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:02.501468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:02.501468Z digest=sha256:51a5f5c5f3c119ddd702071171c90401b6944264b88fd9042c860b59305218f9

Observation 3f9015cc-7a64-4831-b3b2-83f33f1acdc6 · inbound

A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning cites this paper.

A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:33:36.559851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:33:36.559851Z digest=sha256:95885589930ac36747c6459e86ccd895a9605e24260a09ed8b6512a66f1fd055

Observation 166b57ea-7d5a-4df8-9b45-0f307d17155a · inbound

The Knowledge-Reasoning Dissociation: Fundamental Limitations of LLMs in Clinical Natural Language Inference cites this paper.

The Knowledge-Reasoning Dissociation: Fundamental Limitations of LLMs in Clinical Natural Language Inference CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:19:26.281676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:19:26.281676Z digest=sha256:064a21f1b115641248878ce9f900af77e767d3d215a480c4c2e259872719f22c

Observation cf2c98ca-9ec7-44e2-9cfb-802b4e33a740 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:20:34.428264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:c7672a6804c1cd405b7d44ffec872d59fa4f6ccf800ef7e56fc348160ddfeb34

Observation f1262741-742d-46bd-9a07-e678af826b99 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:52.045761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:52.045761Z digest=sha256:6e8f798b8e8cead4de3ffd03f8697bc07cd122e67846762276d29e05783e2a68

Observation 265b363a-d799-4ec7-a651-a82870d5be20 · inbound

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task cites this paper.

Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:07.545195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:34:02.796850Z digest=sha256:6b137223048baf5d549930e074f8269b421f357ff7f8afcdf598d41fd938f14b

Observation 3da4fc31-ca78-4f37-9e84-4a7f609f8bcd · inbound

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination cites this paper.

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.985716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T19:58:32.016341Z digest=sha256:de069f71af7fa7df9510ad311ae9b4d303dfbe36251481685326146a03f9c557

Observation 6aa2899d-1f28-4e4a-9ac1-4207623cdaf4 · inbound

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling cites this paper.

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:22.080918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:37:38.371341Z digest=sha256:71277958ce15cec8d626d2b8431fd4577d4d614305d699a4b31166e8928926d1