Pith. sign in

Paper Citation Record · LEDGER

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models

As of 17 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2505.04914.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04914 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:20:08.568346Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 333aca7d-44e7-4a81-b141-ed8db403be5b · outbound

This paper cites Language models are few-shot learners,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Language models are few-shot learners,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.468639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.468639Z digest=sha256:0ff42ccbe15a60cce72e391d9705096be258acb2f72821104477f9e0598547fd

Observation 5e8f6177-913b-41d5-af9d-4105ac1346d8 · outbound

This paper cites Large language models are zero-shot reasoners,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Large language models are zero-shot reasoners,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.472690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.472690Z digest=sha256:c4b9f6f14e09898d09d67436db88b8684d67bb139adff1952b2f5a7a71b92d39

Observation 4f4c8c94-beb7-4d88-9ab6-936d28ee658d · outbound

This paper cites Emergent abilities of large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Emergent abilities of large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.943339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:20:08.476360Z digest=sha256:e72d412706f2baa1adbefee1bb430807a36b8197a623d18fe48bad858241a448

Observation 4f69adf3-6afa-410c-b08b-e55dce563efc · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.479874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.479874Z digest=sha256:2ee19d01d2031b15285eac4847ce279c6abafe7afd4ca771500d26b2b2216811

Observation 69417eef-736a-40ba-9acc-e855007ffeff · outbound

This paper cites Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.484192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.484192Z digest=sha256:c7a297b31ac0db6d8a0efd367b0c8cd81d2e06cac0328f5be0ba6356962da7dd

Observation da2fc541-3f33-4fae-b90c-6d8a45b1345a · outbound

This paper cites SoK: Memorization in General-Purpose Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models SoK: Memorization in General-Purpose Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.488011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.488011Z digest=sha256:eb680c6ae14bbaafa8b5101a547b64b8abbd583af3f1cb0a8b45600784d7437c

Observation a8abbe21-7763-4344-bf70-28c043c00879 · outbound

This paper cites an unresolved cited work.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:20:08.929798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:20:08.492026Z digest=sha256:5d1cb2d93cc02fa8f68fd4c142c2fea2589714f74ac0609137cfa50a5d0477d3

Observation 57a493b4-dac7-4c51-8c40-317b29909cb1 · outbound

This paper cites Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.495579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.495579Z digest=sha256:5b1329160bfb9f93cb3323217fd747660d449885981fea96288593573fc07117

Observation 518dd686-baa3-47c2-aa78-58e8a2ebe417 · outbound

This paper cites McCorduck, Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models McCorduck, Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.919204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:20:08.499331Z digest=sha256:0826772c189ac0bb5e638d3823848a00a44cd7e4c7d7fe3ed460d581794e9956

Observation 043baa1d-48d3-4f8f-b2c0-ebe92d14104f · outbound

This paper cites On the Measure of Intelligence.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models On the Measure of Intelligence

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.503732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.503732Z digest=sha256:e258a71022821af5fbcd1e1e460bd29eea0adada733abdce139b3f95f9c187b9

Observation c030bab5-6e7e-41c6-a4c0-6476e242bd4f · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.506938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.506938Z digest=sha256:714837894ea9159b13b8924ba45400cd8c59fcf88612112915614d0ab7fd6f8e

Observation 701eaa0a-b922-4945-b917-1ddf75e2539c · outbound

This paper cites A survey on evaluation of large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models A survey on evaluation of large language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.510370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.510370Z digest=sha256:1d0ab51ae9657aa6b3f0ee0d4bd41c0a18d19c9d6815009d5cb76db7b390ec31

Observation b02b5006-0ba1-4891-82ec-71dd956c850f · outbound

This paper cites Challenges and applications of large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Challenges and applications of large language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.513721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.513721Z digest=sha256:c6bfdded447856692d731c9cc4c8af87bd4f2176b1ed124f40ad1d02dc4bb3c3

Observation ed7d9718-e454-42a8-b0e3-2ec90d0e0881 · outbound

This paper cites Llmrg: Improving recommendations through large language model reasoning graphs,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Llmrg: Improving recommendations through large language model reasoning graphs,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.902101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:20:08.520461Z digest=sha256:79ffb203bad1597029c2f642d403ed1698e142b51e1e60cd8aa3585bc3f98e12

Observation 14a1b7a0-5b6f-4688-9b4e-95105873ac88 · outbound

This paper cites Language models as agent models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Language models as agent models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.890811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:20:08.523819Z digest=sha256:94b9e711952077fa242dc2c87207adea965c2e4c0e14ba3ee03be347268012fa

Observation 9b408a06-8115-4f97-93e6-81635fc54f3f · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models ReAct: Synergizing reasoning and acting in language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.527247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.527247Z digest=sha256:c5e3074b702046bf0ad280e2e063acc9c3e8e0ac745f171e102b753c4cf31747

Observation 39ae713b-5ec7-4602-a424-c51cb3affc8f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.530551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.530551Z digest=sha256:1e2e42a08f02b582fdc6d34faa1c600e639be07b9e822863f7ed1a505a94d476

Observation 8965a074-44a9-4a7d-af14-3aedf77bf502 · outbound

This paper cites Few shot chain-of-thought driven reasoning to prompt LLMs for open ended medical question answering.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Few shot chain-of-thought driven reasoning to prompt LLMs for open ended medical question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.533705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.533705Z digest=sha256:699a2993b216e64e8a9b1fb4f3899a4cc0539654558a8697c8783ff2ea9f4f88

Observation c15e7fd7-0594-41a5-86a1-ea74a54a257d · outbound

This paper cites Automatic chain of thought prompting in large language models,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Automatic chain of thought prompting in large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.868946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:20:08.537247Z digest=sha256:de3cb82bc684ad8f3e86563073bd2cfe2baa6869fade2d9347187ba8607ede62

Observation 7ee16372-ad1e-4c3a-99da-532159d07318 · outbound

This paper cites Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.540467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.540467Z digest=sha256:dcd3bf3758c48bb8a4f41de67461ecc045d64a7549395dc77da47b4d89d997ca

Observation 587fec94-3293-4db4-aacf-97718806b476 · outbound

This paper cites Can neural networks understand monotonicity reasoning?.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Can neural networks understand monotonicity reasoning?

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:20:08.857001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:20:08.543976Z digest=sha256:339e173bc13c86113b4ee8ca1f71b13f5c11950e3125b6c20ef3ccf95be6b7eb

Observation 5e84b47b-ea99-44c7-b5e1-bd7773c0a06a · outbound

This paper cites Idea: Enhancing the rule learning ability of large language model agent through induction, deduction, and abduction,.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Idea: Enhancing the rule learning ability of large language model agent through induction, deduction, and abduction,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.547216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.547216Z digest=sha256:4d254f96bc270bf6c9cac16f6f1cbc222e526645349984afd4d49571365d5183

Observation 53d80a29-79cc-4682-ad6f-6a4a680e9708 · outbound

This paper cites Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.550424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.550424Z digest=sha256:3ddb2f262bc2d5792b4e6f4e4fe383ef797c5e6e9eb8792b3c11c499de6cdc0d

Observation 47da1ea2-99d4-407b-872c-ec3411d487bb · outbound

This paper cites A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.553997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.553997Z digest=sha256:224526ab7dce4d01158c3adc512fd550e288e265be80febcf595aea23878e28c

Observation 6e691e3d-5e6e-48bd-b7a0-bd4c94adcb08 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.557571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.557571Z digest=sha256:24ea517c354490292db1b7b2420dbf5141dfa2ae08bf7d8012b1d9c6ca3e0902

Observation f6194edb-0887-4c0a-af76-0c35cb09657b · outbound

This paper cites Premise Order Matters in Reasoning with Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Premise Order Matters in Reasoning with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.561070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.561070Z digest=sha256:21532a64ec495536206ceebbade06a083c3a8cb5c4ee1b8f5865ff4eeeadfea4

Observation 40ada801-bf78-409d-9207-dec4c021e8fc · outbound

This paper cites Large Language Models Are Not Strong Abstract Reasoners.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Large Language Models Are Not Strong Abstract Reasoners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.564736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.564736Z digest=sha256:c559cd462bee1e245e434aa31d3816161de0590868e8483cfb313442b2c95ccf

Observation 71c00e42-9e8d-455e-b729-955d5c97650f · outbound

This paper cites Inductive Biases for Deep Learning of Higher-Level Cognition.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Inductive Biases for Deep Learning of Higher-Level Cognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.568346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.568346Z digest=sha256:36b8fe2cec787d9f73877ba806e7bf8bbac496c0e4d85edb0b679c1f7a2df17a

Observation bfe22202-1c31-43a6-ae55-7d468a48431d · outbound

This paper cites Challenges and Applications of Large Language Models.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Challenges and Applications of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.517021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.517021Z digest=sha256:52d7a4c59b73e27e65edb3f328fd6a1b8548ee6bf1db0a97417c4bf8ef8569fa

Pith citing papers

No inbound Pith citation observations are available.