Pith. sign in

Paper Citation Record · LEDGER

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models

As of 22 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2505.04914.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04914 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:20:08.568346Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 333aca7d-44e7-4a81-b141-ed8db403be5b · outbound

This paper cites Language models are few-shot learners,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Language models are few-shot learners,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.468639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.468639Z digest=sha256:56817f67ac3fdb6f1d326bd652aa10e00b9a3e793bda6c9eecf04208024d4f58

Observation 5e8f6177-913b-41d5-af9d-4105ac1346d8 · outbound

This paper cites Large language models are zero-shot reasoners,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Large language models are zero-shot reasoners,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.472690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.472690Z digest=sha256:962845f999a0f8c07869dc8a4079bd5fee77a714160498835c5ee387bed3aae4

Observation 4f4c8c94-beb7-4d88-9ab6-936d28ee658d · outbound

This paper cites Emergent abilities of large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Emergent abilities of large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.943339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:20:08.476360Z digest=sha256:7b1eeffc1a6cc08d2616ec20ebad8122c0f202e390cd28759b895b414db9b4e3

Observation 4f69adf3-6afa-410c-b08b-e55dce563efc · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.479874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.479874Z digest=sha256:f38d2390e7528ef017fc8d19d11c395ff3f55de9b9e2d4f5881772aee94be0ce

Observation 69417eef-736a-40ba-9acc-e855007ffeff · outbound

This paper cites Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.484192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.484192Z digest=sha256:c14b456f78ce444db3b4b1da673b209b341f18ecdb835b5b89496a17d9a9a9a2

Observation da2fc541-3f33-4fae-b90c-6d8a45b1345a · outbound

This paper cites SoK: Memorization in General-Purpose Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models SoK: Memorization in General-Purpose Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.488011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.488011Z digest=sha256:4d911143974f54323c5ec0dd6ed45ec60cf1aef24b940c562f1a65085cc2cead

Observation a8abbe21-7763-4344-bf70-28c043c00879 · outbound

This paper cites an unresolved cited work.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:20:08.929798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:20:08.492026Z digest=sha256:e021d88870480b1e7fa011562aa657cf715b89848a69189bd3014c86b60435aa

Observation 57a493b4-dac7-4c51-8c40-317b29909cb1 · outbound

This paper cites Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.495579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.495579Z digest=sha256:0bb1dff582baae06b887530f1483d03f2e278e20a7537d9ed71ee9cf34518589

Observation 518dd686-baa3-47c2-aa78-58e8a2ebe417 · outbound

This paper cites McCorduck, Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models McCorduck, Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.919204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:20:08.499331Z digest=sha256:4ae59578c3532e716638801d2b1d4e56c5c158d5e343b88754a57e1166aa7b04

Observation 043baa1d-48d3-4f8f-b2c0-ebe92d14104f · outbound

This paper cites On the Measure of Intelligence.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models On the Measure of Intelligence

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.503732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.503732Z digest=sha256:6d658d2ef6fe9e134caff3f93e9632a134a10832908364ab9c9109a726cf2d46

Observation c030bab5-6e7e-41c6-a4c0-6476e242bd4f · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.506938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.506938Z digest=sha256:39859efcb40bb2b0c12319c837c76aac59a2d611cd49995fba31c2d8127ae135

Observation 701eaa0a-b922-4945-b917-1ddf75e2539c · outbound

This paper cites A survey on evaluation of large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models A survey on evaluation of large language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.510370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.510370Z digest=sha256:213df26fc82f8fd43b4b9daf3457be9d9cda642a92e2bd7605d97a7a9ebba7ec

Observation b02b5006-0ba1-4891-82ec-71dd956c850f · outbound

This paper cites Challenges and applications of large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Challenges and applications of large language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.513721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.513721Z digest=sha256:826fd0a79041b82a736bebfd5e6904270f84d2e6fd8bc1e2cee43c4031c837bd

Observation ed7d9718-e454-42a8-b0e3-2ec90d0e0881 · outbound

This paper cites Llmrg: Improving recommendations through large language model reasoning graphs,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Llmrg: Improving recommendations through large language model reasoning graphs,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.902101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:20:08.520461Z digest=sha256:21316e230ec8a9ac21a8c91a0c52124d5c431a95550059f63455d5a7579e407f

Observation 14a1b7a0-5b6f-4688-9b4e-95105873ac88 · outbound

This paper cites Language models as agent models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Language models as agent models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.890811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:20:08.523819Z digest=sha256:51588ecf8d16c83e1eb61cddbb258fbbef3a780945cc9a580255611a14975957

Observation 9b408a06-8115-4f97-93e6-81635fc54f3f · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models ReAct: Synergizing reasoning and acting in language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.527247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.527247Z digest=sha256:36d41bc7e42e719603f8a405f331feb6462769519fa825b89618b8df7b5536ff

Observation 39ae713b-5ec7-4602-a424-c51cb3affc8f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.530551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.530551Z digest=sha256:ba4ce747dc502322ba16d81735450c493fdc04d81c178a14df0507595e9390aa

Observation 8965a074-44a9-4a7d-af14-3aedf77bf502 · outbound

This paper cites Few shot chain-of-thought driven reasoning to prompt LLMs for open ended medical question answering.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Few shot chain-of-thought driven reasoning to prompt LLMs for open ended medical question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.533705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.533705Z digest=sha256:9124d05b194f78c255e1aa4a072abc11c57826f777e813638b33401b3c248fa9

Observation c15e7fd7-0594-41a5-86a1-ea74a54a257d · outbound

This paper cites Automatic chain of thought prompting in large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Automatic chain of thought prompting in large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.868946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:20:08.537247Z digest=sha256:1b339e2c055a63e6255b5a069028d28971e1778d862a0245eb7b6336be15ebb1

Observation 7ee16372-ad1e-4c3a-99da-532159d07318 · outbound

This paper cites Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.540467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.540467Z digest=sha256:69037b88e77f09523ead3fbdd2a39859d4dce133e4a379033997b219a9730e93

Observation 587fec94-3293-4db4-aacf-97718806b476 · outbound

This paper cites Can neural networks understand monotonicity reasoning?.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Can neural networks understand monotonicity reasoning?

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.857001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:20:08.543976Z digest=sha256:59db8732cefded150a5fb4ba7c5d64ce57be8394782fb8a8e6aa2d8e0e169f85

Observation 5e84b47b-ea99-44c7-b5e1-bd7773c0a06a · outbound

This paper cites Idea: Enhancing the rule learning ability of large language model agent through induction, deduction, and abduction,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Idea: Enhancing the rule learning ability of large language model agent through induction, deduction, and abduction,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.547216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.547216Z digest=sha256:2997be53bbdf3a29b0f68acbc9bf5becd79abe38bfe5647a1231351210dd4de8

Observation 53d80a29-79cc-4682-ad6f-6a4a680e9708 · outbound

This paper cites Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.550424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.550424Z digest=sha256:b6c1054b5d9cd68c614a8b0e035466468dea778db5b35e882e64674b169cf77c

Observation 47da1ea2-99d4-407b-872c-ec3411d487bb · outbound

This paper cites A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.553997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.553997Z digest=sha256:8ecf64ef9ff745876e4f72c6a529f5cb20f4c587a11fe7cf332b7d613d8c72b6

Observation 6e691e3d-5e6e-48bd-b7a0-bd4c94adcb08 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.557571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.557571Z digest=sha256:56b7422f5fe7cf289ff4cd1b471375ebb1fb5bab0f8e3125d2fd97a551521c7e

Observation f6194edb-0887-4c0a-af76-0c35cb09657b · outbound

This paper cites Premise Order Matters in Reasoning with Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Premise Order Matters in Reasoning with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.561070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.561070Z digest=sha256:4aa17d75fdb7820ff8410ce7860c47c69fd8c80ac17ec6c465bb66a1df0d21d1

Observation 40ada801-bf78-409d-9207-dec4c021e8fc · outbound

This paper cites Large Language Models Are Not Strong Abstract Reasoners.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Large Language Models Are Not Strong Abstract Reasoners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.564736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.564736Z digest=sha256:8728550e72936a6c694ed6291ed11c305f44a2ba31a1c20c09d891bd9938b415

Observation 71c00e42-9e8d-455e-b729-955d5c97650f · outbound

This paper cites Inductive Biases for Deep Learning of Higher-Level Cognition.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Inductive Biases for Deep Learning of Higher-Level Cognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.568346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.568346Z digest=sha256:0671f0a515aa04b9d157048fe0dab0d2e77c3c0a8268863d21efbe10b67e83bd

Observation bfe22202-1c31-43a6-ae55-7d468a48431d · outbound

This paper cites Challenges and Applications of Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Challenges and Applications of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.517021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.517021Z digest=sha256:7d01394554a237e02596b3b6157a4ce0e8678339a6b919f9cec0074d290f0e42

Pith citing papers

No inbound Pith citation observations are available.