Pith. sign in

Paper Citation Record · LEDGER

Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2407.14985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.14985 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:08:32.920085Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:56:57.463531Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 568eba03-a1d3-49e6-833c-25752633c68f · inbound

Attributing Culture-Conditioned Generations to Pretraining Corpora cites this paper.

Attributing Culture-Conditioned Generations to Pretraining Corpora Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:07:41.670904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T07:05:52.610830Z digest=sha256:8c1490aaf9f24437e5e52fe9c9dfe37a81c88c6a2292bea3c96a99750daa0f59

Observation 38733b99-0b44-4409-9a65-7ecd40589275 · inbound

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training cites this paper.

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T21:31:30.245269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T21:31:30.202477Z digest=sha256:946031fe33af6889d64bfad2e3f71e9d6cc5b5da7ed39a57dec2e4ca1a3b9f00

Observation 3fff36b1-6b09-4f6a-bd91-16d978ae4a92 · inbound

An Annotated Reading of 'The Singer of Tales' in the LLM Era cites this paper.

An Annotated Reading of 'The Singer of Tales' in the LLM Era Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T20:08:32.920085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:08:32.920085Z digest=sha256:8ce18f805671f26a1571dd8300554478bedc01c734f117930438a9df82f19b8a

Observation 2c5e2e8c-d1f1-45c9-808b-d94d2edc3773 · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.095763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.095763Z digest=sha256:68a28d2f562194319b18ad81acaef93b6108977aed1957df7cf4ca1f06f713ff

Observation 71656722-bb95-43e0-b6f6-202da549b4f6 · inbound

Counterfactual Influence as a Distributional Quantity cites this paper.

Counterfactual Influence as a Distributional Quantity Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:58.081705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:58.081705Z digest=sha256:b37477810bcc4ccd4f93ad7140b65fa5808c903056a74ce788725eba35ca6cc1

Observation 20315f6b-4426-4237-b172-e0699b58a691 · inbound

Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs cites this paper.

Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:06.739410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:06.739410Z digest=sha256:2d42af9853d5d4a8a2ec0b3b793d559cea37dfdb84e8cd6a75140ce0a8851094

Observation df5a10f9-4434-46e4-9ae7-06f36e703b51 · inbound

Adaptive Multi-Agent Reasoning via Automated Workflow Generation cites this paper.

Adaptive Multi-Agent Reasoning via Automated Workflow Generation Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:10:28.285063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:10:28.285063Z digest=sha256:a34c00680ef3ddeba424c34f947f6d881953dbfefbaa3a6512af7a679b3c92d7

Observation 32107197-b353-4d20-97d1-32a1e1ba110a · inbound

Rethinking Memorization Measures and their Implications in Large Language Models cites this paper.

Rethinking Memorization Measures and their Implications in Large Language Models Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T15:54:35.942487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:54:35.942487Z digest=sha256:1483bf5da2b63f9e85d6f1037fe4e7dbe1f2a210edce532ed02a67b945a0d0c8

Observation 112b370f-6934-434e-a8ea-3635e7008483 · inbound

Access Paths for Efficient Ordering with Large Language Models cites this paper.

Access Paths for Efficient Ordering with Large Language Models Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:34:24.173404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T22:32:23.352584Z digest=sha256:9f40b28e8f41abdd420acaa253ecc552093b971759a1b1afc8f1a6d6bef709c3

Observation 0b00a481-05e9-450d-95bc-956650b6dd0e · inbound

Remembering Unequally: Global and Disciplinary Bias in LLM Reconstruction of Scholarly Coauthor Lists cites this paper.

Remembering Unequally: Global and Disciplinary Bias in LLM Reconstruction of Scholarly Coauthor Lists Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:50:38.090115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T01:48:49.289059Z digest=sha256:a125273ba224cf22aa580bfe23bf66da16815b1eafbbf8304ec67a98c2289d7d

Observation 08fbc6f7-faa8-4e76-9379-ae702127fc99 · inbound

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs cites this paper.

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:57.465018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T01:40:53.284131Z digest=sha256:7873d1b6abd50dced4e980e4640a8ed6c899b02a77a85f6a407ccb296f3f2328

Observation 10a4e715-c655-4ca4-aee8-55fcf3ac4764 · inbound

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training cites this paper.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:119d8018347df8c519e0a1715ba6e2f91dc181619bfe72b36acf280c9654657d