Pith. sign in

Paper Citation Record · LEDGER

Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2407.14985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.14985 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:51:31.625209Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:56:57.463531Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f4666535-99b4-407c-adfa-d15ec3f23a44 · inbound

VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs cites this paper.

VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:31.625209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:51:31.625209Z digest=sha256:fd17f147f99b0f27025435f8b07b3a1c61ac7aed3cea53dd92a9534fd9ecc8a3

Observation c4437a5f-218f-44ee-bc9f-3254a8608198 · inbound

AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution cites this paper.

AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:35:56.673147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:35:56.673147Z digest=sha256:a138338665c2d20a3acfbc56dcec50cc5ddf470edf1ab28af8c48afbfdd45eaa

Observation 6dae89d5-9120-421f-b2a4-bdae562c361a · inbound

The Pitfalls of Memorization: When Memorization Hurts Generalization cites this paper.

The Pitfalls of Memorization: When Memorization Hurts Generalization Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:40.847859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:45:40.847859Z digest=sha256:fd4013e51e3e11303e991b824407c64d9cf8a9784998dddaaf20fea0739cc115

Observation 568eba03-a1d3-49e6-833c-25752633c68f · inbound

Attributing Culture-Conditioned Generations to Pretraining Corpora cites this paper.

Attributing Culture-Conditioned Generations to Pretraining Corpora Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:07:41.670904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T07:05:52.610830Z digest=sha256:af7ae756d3ea49bc2b30df5d4f1b0bf2403194e05330f84ae060ccf72c073bbd

Observation 33ad8a1f-2858-4c42-a594-14c44495db14 · inbound

Enhancing Generalization in Chain of Thought Reasoning for Smaller Models cites this paper.

Enhancing Generalization in Chain of Thought Reasoning for Smaller Models Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:42:22.712987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:42:22.712987Z digest=sha256:0c2ce9dbb6a94c75e7f05864169eee4882f90e3534aaca329c1d267a3d3b3649

Observation 38733b99-0b44-4409-9a65-7ecd40589275 · inbound

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training cites this paper.

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T21:31:30.245269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T21:31:30.202477Z digest=sha256:d1a5007a7e68ac87b5ff1183eb35157916ef891d71b95f16bef85913d2bd6f42

Observation 3fff36b1-6b09-4f6a-bd91-16d978ae4a92 · inbound

An Annotated Reading of 'The Singer of Tales' in the LLM Era cites this paper.

An Annotated Reading of 'The Singer of Tales' in the LLM Era Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T20:08:32.920085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:08:32.920085Z digest=sha256:17d7cffc17090b9d7cc7d54f2784f05072dae2340f843aaf636d41d616c56630

Observation 2c5e2e8c-d1f1-45c9-808b-d94d2edc3773 · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.095763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.095763Z digest=sha256:1e36b0509100a6f4285eb6ceb42060d25653c59ee8a2af33d70b51186be8effd

Observation 71656722-bb95-43e0-b6f6-202da549b4f6 · inbound

Counterfactual Influence as a Distributional Quantity cites this paper.

Counterfactual Influence as a Distributional Quantity Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:58.081705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:58.081705Z digest=sha256:6197b5673c091916394ea8effb71a7b16dc4259eb9e785f608bed47aea9c6133

Observation 20315f6b-4426-4237-b172-e0699b58a691 · inbound

Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs cites this paper.

Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:06.739410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:06.739410Z digest=sha256:fe751c92cd32f64bcd421dd7207582d9a07a47d9571f855dc236597bb34dc35b

Observation df5a10f9-4434-46e4-9ae7-06f36e703b51 · inbound

Adaptive Multi-Agent Reasoning via Automated Workflow Generation cites this paper.

Adaptive Multi-Agent Reasoning via Automated Workflow Generation Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:10:28.285063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:10:28.285063Z digest=sha256:f8ab6a0e178ea4b83e1555283bb077c2d20afc069996f96b0cfb8576729cc579

Observation 32107197-b353-4d20-97d1-32a1e1ba110a · inbound

Rethinking Memorization Measures and their Implications in Large Language Models cites this paper.

Rethinking Memorization Measures and their Implications in Large Language Models Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T15:54:35.942487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:54:35.942487Z digest=sha256:5de46a0cbf5937a2a273e2143d44e60dd341173358404b2d2b817abe6d14612a

Observation 112b370f-6934-434e-a8ea-3635e7008483 · inbound

Access Paths for Efficient Ordering with Large Language Models cites this paper.

Access Paths for Efficient Ordering with Large Language Models Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:34:24.173404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T22:32:23.352584Z digest=sha256:919293320b956ca03b3ead5eab70748270527c26b9714cd52467fb1c1797b5bd

Observation 0b00a481-05e9-450d-95bc-956650b6dd0e · inbound

Remembering Unequally: Global and Disciplinary Bias in LLM Reconstruction of Scholarly Coauthor Lists cites this paper.

Remembering Unequally: Global and Disciplinary Bias in LLM Reconstruction of Scholarly Coauthor Lists Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:50:38.090115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T01:48:49.289059Z digest=sha256:fe13164a6c960a7c10c34498c8d0efc49aedb9e0791f2407decef05ebab0043b

Observation 08fbc6f7-faa8-4e76-9379-ae702127fc99 · inbound

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs cites this paper.

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:57.465018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T01:40:53.284131Z digest=sha256:c25bf4446d4903b92d0ac279670adfe362171e6689646a569ad00ce2ac65ef80

Observation 10a4e715-c655-4ca4-aee8-55fcf3ac4764 · inbound

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training cites this paper.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:a7bc66a2979bca756765173f95e35c93a261175610a01f90aa6bd3385e01279c