Pith. sign in

Paper Citation Record · LEDGER

Large Language Models Are Not Strong Abstract Reasoners

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2305.19555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.19555 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:20:43.028897Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.072171Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b5d5559f-a62b-42a0-9ce0-880d5263b7bd · inbound

Efficient Causal Graph Discovery Using Large Language Models cites this paper.

Efficient Causal Graph Discovery Using Large Language Models Large Language Models Are Not Strong Abstract Reasoners

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:25:28.683305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T04:10:17.251713Z digest=sha256:9eba8e9cb8f594096013b16612aaee78ae16570dafb5f27ac31c58e537b75c62

Observation a3fb7fdb-ef50-4083-b3b4-67b221cc8d9f · inbound

EXP-Bench: Can AI Conduct AI Research Experiments? cites this paper.

EXP-Bench: Can AI Conduct AI Research Experiments? Large Language Models Are Not Strong Abstract Reasoners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:20:43.028897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:20:43.028897Z digest=sha256:b31ec5ddd9e141dfd166104c4b8d257380295717bd6a7c984e4915fd12ebaadd

Observation a6ace9b7-614f-4344-a2af-50d24872b524 · inbound

Adaptive Multi-Agent Reasoning via Automated Workflow Generation cites this paper.

Adaptive Multi-Agent Reasoning via Automated Workflow Generation Large Language Models Are Not Strong Abstract Reasoners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:10:28.200105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:10:28.200105Z digest=sha256:e02d0201ff088832502618ca5f1bb64135cec7d3e6affd123120e933278b27c7

Observation 7a465cd9-7fe6-402c-bdf2-9d091861f7c3 · inbound

Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning cites this paper.

Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning Large Language Models Are Not Strong Abstract Reasoners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T21:10:30.905697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:10:30.905697Z digest=sha256:3c6c244f915243ed9eff2b0d3337f52f839acce3ecccd1a5eb559b85100073bf

Observation c048c413-1338-4c46-8ad1-e97401805e32 · inbound

Analysis of Error Sources in LLM-based Hypothesis Search for Few-Shot Rule Induction cites this paper.

Analysis of Error Sources in LLM-based Hypothesis Search for Few-Shot Rule Induction Large Language Models Are Not Strong Abstract Reasoners

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T13:02:25.658020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:02:25.658020Z digest=sha256:c09c64fe6f0eb038c0dd7567edb025d22205fd12fdc553ae20ade4c45e3a750d

Observation aecfdd32-f831-4261-9a74-6e6c44aa3e31 · inbound

Gradient-Based Program Synthesis with Neurally Interpreted Languages cites this paper.

Gradient-Based Program Synthesis with Neurally Interpreted Languages Large Language Models Are Not Strong Abstract Reasoners

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:08.328649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:29:33.858344Z digest=sha256:ceed1bc85170497ef46800965063c1c64a037b638c68112da9e72920d79d6af2

Observation 5812a36e-d6aa-4ddd-b97e-396c810e9ddf · inbound

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform cites this paper.

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform Large Language Models Are Not Strong Abstract Reasoners

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:46.626588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:36:44.033596Z digest=sha256:9779cf686ea0ad639d0bbf7f536b33a7c92b2b1e875ba194f2c7fa5c19d381f1

Observation 2fed1131-5412-4b61-b831-7c61c8caf7b2 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Large Language Models Are Not Strong Abstract Reasoners

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.563473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:064393eb27d140266dd3b493f2c977618d5f24b5ab2fc3ad0ff79c8e55964833

Observation d1d5af42-3690-4818-9150-77adace763c3 · inbound

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models cites this paper.

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models Large Language Models Are Not Strong Abstract Reasoners

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.074083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T05:32:59.640335Z digest=sha256:58a0de381118cbc7a9a7538ac4db1b72c75bfdadbb80dd8a43f1a9e2f4cf5cce

Observation a77aed3c-ae71-4185-b50f-f8008370f086 · inbound

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models cites this paper.

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models Large Language Models Are Not Strong Abstract Reasoners

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:33:45.602884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T05:19:07.877346Z digest=sha256:305049c540cec1d7afe6334ab549fdede7b80ae9e3b71a4614a5792e9320e0da