Pith. sign in

Paper Citation Record · LEDGER

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers

As of 8 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2506.09099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09099 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:04:55.962152Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:02:06.788056Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:46:25.975143Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8161e917-cecd-4e3a-a2a4-1c889c3a655e · outbound

This paper cites GPT-4 Technical Report.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.911995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.911995Z digest=sha256:0987ff9188ccb6af9aa1722a6430da7d684b5ae0e754405e4bfd3065bf26e2e3

Observation 57c1af3c-30f3-4ba1-8fa7-3fe26dac7beb · outbound

This paper cites 9, 2024); accessed May 19,.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers 9, 2024); accessed May 19,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:56.167851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:04:55.925247Z digest=sha256:9afba4667ffed0d12603fe4e8774ef1f746d383037db0b8e6938a957df53972c

Observation 04fcb939-cc24-4f4e-89c4-98b4e2b86369 · outbound

This paper cites doi: 10.1016/j.neunet.2024.106550.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers doi: 10.1016/j.neunet.2024.106550

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.932743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.932743Z digest=sha256:c4c91e1e3b58b7f4634007fcdf0d44adb3b1acf6833ecd9773d6f8f1ba22b23a

Observation d55ccfad-9b27-495b-b0a9-0810362e5e9e · outbound

This paper cites Training language models to follow instructions with human feedback.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Training language models to follow instructions with human feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.936614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.936614Z digest=sha256:82ef22eda08f94b05e0b202be5dd7b7cd26d94a4541a6ac77f39a899a0bd89e2

Observation e040c504-c48e-47f9-8502-6c11da144c51 · outbound

This paper cites Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.940338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.940338Z digest=sha256:1426ced441a8bca9e6cac503644a873d6410f806667a9df6ac4b7e5d617a9fc5

Observation 170e1a5b-513d-4023-b290-12eee9943dc2 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.944197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.944197Z digest=sha256:c270791703f11cf1c32bdf782f21074dcf41676e5e00e24141f61f6cedcd4d99

Observation d414a4b0-7877-4fb0-bf7d-2931d92c6792 · outbound

This paper cites Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.951254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.951254Z digest=sha256:fe811bfcaa413d42fa0bbf9f28579a87fc0e0d56d915a4f2c3a85b7ba9511c42

Observation 43e43de8-b55d-4d46-8ce2-d425768a8620 · outbound

This paper cites Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.954824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.954824Z digest=sha256:e9cf3a7dce1a8bccbb51e3efb1843175e182bfa63ea86fea1d2d1fc881325828

Observation 456e782f-fb2c-42de-9a69-8828c499f5c4 · outbound

This paper cites Exploring Memorization in Fine-tuned Language Models.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Exploring Memorization in Fine-tuned Language Models

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:04:55.958433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.958433Z digest=sha256:92b6d434d6eccae3c37c460c32bc1c78aa94e5e375bfa791c6e0646f609bde1b

Observation 48de51b3-0a53-43ee-91df-0e18c07519cb · outbound

This paper cites Goldilocks zone.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Goldilocks zone

Reference 14

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:04:56.157226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:04:55.962152Z digest=sha256:a714552d86e07e403a46daf4c23164ec69ffd0b5b3bb2b6778c0815afc7d269f

Observation 6083ae0c-f92c-4cd0-93e5-190a11b12064 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Gemini: A Family of Highly Capable Multimodal Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.948019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.948019Z digest=sha256:b602cff24e59e50fc082018771d24f05a9caae6a6929ae7b5f6af696a916f2a5

Observation f86a55e6-0981-4687-a4ba-da29f92be046 · outbound

This paper cites Towards Understanding Grokking: An Effective Theory of Representation Learning.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Towards Understanding Grokking: An Effective Theory of Representation Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.929001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.929001Z digest=sha256:c07bfda6418cf8cc8f37b060436443cd3d9a8a78b431eb54b74bfbc07e8660d4

Observation 7d4bf7da-300f-490b-b06e-29a15f30c65a · outbound

This paper cites GPT-4o System Card.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.921232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.921232Z digest=sha256:419b8992d4a7a2eda50ad09b285e46eac69bb13782f15b32bffca90f23609bcd

Observation a0b355b9-aa0d-4b4d-9c9b-0e98309b49f6 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.916932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.916932Z digest=sha256:91cd09dbfccebaa161bed1aaa05a24fdfbdf1992ea6cff72f8c8c02bf3816cbb

Pith citing papers

Observation 6a6da3b7-5808-4654-9ae7-d0e6edac71f7 · inbound

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities cites this paper.

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:25.982733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:02:06.788056Z digest=sha256:9c436896f696cefd3bd20c2d410cbb96ab8b1c7b93fe8544367cd06932f0c8d5