Pith. sign in

Paper Citation Record · LEDGER

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers

As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2506.09099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09099 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:04:55.962152Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:02:06.788056Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:46:25.975143Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8161e917-cecd-4e3a-a2a4-1c889c3a655e · outbound

This paper cites GPT-4 Technical Report.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.911995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.911995Z digest=sha256:98ce52ab71bfdd6393c614ddf367f6c992065b4e694cb3d0d3651d28f8f9aaf1

Observation 57c1af3c-30f3-4ba1-8fa7-3fe26dac7beb · outbound

This paper cites 9, 2024); accessed May 19,.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers 9, 2024); accessed May 19,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:56.167851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:55.925247Z digest=sha256:cc49368ffb7f032f24d8a46ce5db3c6818c5e4c5a3f4ef4e2a59adba9f9c1d3f

Observation 04fcb939-cc24-4f4e-89c4-98b4e2b86369 · outbound

This paper cites doi: 10.1016/j.neunet.2024.106550.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers doi: 10.1016/j.neunet.2024.106550

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.932743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.932743Z digest=sha256:c1747fb261995816ad4144a492d9409400dd1a24b71d319074b5034a4877dffe

Observation d55ccfad-9b27-495b-b0a9-0810362e5e9e · outbound

This paper cites Training language models to follow instructions with human feedback.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Training language models to follow instructions with human feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.936614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.936614Z digest=sha256:2c97ac6e1da658a15c4b4f9305b3a38136569323ea2db78c67c29f265de395b9

Observation e040c504-c48e-47f9-8502-6c11da144c51 · outbound

This paper cites Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.940338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.940338Z digest=sha256:ce2edad23b4c6a319ad49b9d2da5003025e11802a2a89a548e386f8ecccf08ed

Observation 170e1a5b-513d-4023-b290-12eee9943dc2 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.944197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.944197Z digest=sha256:288af1a602ce1bae655f396c889be1ebfdbac6774fde0a8b9573e2802c28c0d7

Observation d414a4b0-7877-4fb0-bf7d-2931d92c6792 · outbound

This paper cites Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.951254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.951254Z digest=sha256:211f62a7e207a69dee2f6a9e97be306574adbd5d49592d1d327e468faac7e7a6

Observation 43e43de8-b55d-4d46-8ce2-d425768a8620 · outbound

This paper cites Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.954824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.954824Z digest=sha256:c63b80d80a78118bde2c78f70313dc7564fa570b3e6d8456b08038394ee9d92b

Observation 456e782f-fb2c-42de-9a69-8828c499f5c4 · outbound

This paper cites Exploring Memorization in Fine-tuned Language Models.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Exploring Memorization in Fine-tuned Language Models

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:04:55.958433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.958433Z digest=sha256:52fbab2bbfe49850e4d05ecae9fd6eec37a898692b70b109a3a260c2c6a32402

Observation 48de51b3-0a53-43ee-91df-0e18c07519cb · outbound

This paper cites Goldilocks zone.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Goldilocks zone

Reference 14

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:04:56.157226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:55.962152Z digest=sha256:fa837ee8b0bd9c6f483c82de667ed77e9077e9b39c03601214d3b440aea6d3ea

Observation 6083ae0c-f92c-4cd0-93e5-190a11b12064 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Gemini: A Family of Highly Capable Multimodal Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.948019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.948019Z digest=sha256:adb28a92c78786f42da76c2625e6fa141dd9d3a6cd5de061c3a6572b27dc8e24

Observation f86a55e6-0981-4687-a4ba-da29f92be046 · outbound

This paper cites Towards Understanding Grokking: An Effective Theory of Representation Learning.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Towards Understanding Grokking: An Effective Theory of Representation Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.929001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.929001Z digest=sha256:0f95ae02655293e80ca36789da3bd8287dbdcb46b952e756fac137d5c6f0cec1

Observation 7d4bf7da-300f-490b-b06e-29a15f30c65a · outbound

This paper cites GPT-4o System Card.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.921232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.921232Z digest=sha256:06a21a4f39730a62d2efa83898dae149c0ec2777f24ee428f402a21c49833f26

Observation a0b355b9-aa0d-4b4d-9c9b-0e98309b49f6 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.916932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.916932Z digest=sha256:06edef2c2584232475cb526ceec4d5babd2b3bac80dc73f380f74b98bb7e5966

Pith citing papers

Observation 6a6da3b7-5808-4654-9ae7-d0e6edac71f7 · inbound

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities cites this paper.

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:25.982733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:02:06.788056Z digest=sha256:1f9201b65b925e788f0d2c6d48c9e7c2ab9632cc68ed520dbbeb758ea5d3e36a