Pith. sign in

Paper Citation Record · LEDGER

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training

As of 15 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.04969.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04969 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T10:45:46.618668Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a1d000a-3215-432f-be92-848e41b96587 · outbound

This paper cites Phi-4 Technical Report.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:47eaf18c5e3b258826317eea43124478c5f6eaa35c94c0838b4ef7a52c53f151

Observation 21cff844-1be4-4e10-bcde-fd3fc1f8428f · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:7a64c22e2f24b1f36bfb94a40fa02ab30698ab055fea334dcccc77502de62a2d

Observation 57b7f3bd-d63e-46e9-8f64-c6b2c8044616 · outbound

This paper cites Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:8e28b4d7396894772f995c9ef50f8ab9ebc1036bb51c6b067b357a39bfa91397

Observation ac4162bd-e12c-4948-a030-39ba601e5c8c · outbound

This paper cites Reformulation for Pretraining Data Augmentation.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Reformulation for Pretraining Data Augmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:04cb7ba95f0a544e91aab4bfaf8847c0817d033f60b1c8d13523ae9c21e1876e

Observation 0f37e195-1eca-4716-a8b1-1c6d23669b05 · outbound

This paper cites Scaling Laws and Interpretability of Learning from Repeated Data.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Scaling Laws and Interpretability of Learning from Repeated Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:fc3bf961a83c57de26caa709bc054f8ed22ae4fc3f35c862161268d1b3be5b37

Observation 0736a376-dae2-4317-820f-6f4f87747bc7 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Training Compute-Optimal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:b6cf73db23da789279ef86e1d4f278bec45c3b6b7f9b72263b08dd748e3fbc35

Observation e364a06c-7f2d-4cab-aeaf-bc5df9d406ab · outbound

This paper cites Ultra-Sparse Memory Network.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Ultra-Sparse Memory Network

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:941a4ccd0ef017f4ec3a720aa5a25631f2f715b24a6a96e31261c9740dc7b0b4

Observation 5b965e8a-d04d-415a-a662-6fe5946470a2 · outbound

This paper cites Ziyue Li, Chenrui Fan, and Tianyi Zhou.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Ziyue Li, Chenrui Fan, and Tianyi Zhou

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:11fc886d520f6bf0051f76c591c1c7f55b84581b06a1cfc8b01b24fbcf5114f0

Observation d0dfc0bd-c81c-49e7-a214-27aa69869332 · outbound

This paper cites Let's Verify Step by Step.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Let's Verify Step by Step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:5d120354864b0b9931d0ce1106a3fc973d8fee6320a497cb1af96175abbead19

Observation 7ddfe9df-2dac-499f-83e3-0f48e186728f · outbound

This paper cites How much do language models memorize?.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training How much do language models memorize?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:e7c3b55c25d13301d77db27fb56b7e187e3920524a1646d7e1d0198c88cd1667

Observation 016b0d27-4f1c-4fdf-9077-d0a313695566 · outbound

This paper cites Olmo 3.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Olmo 3

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:f73bc58330094ce30b7680d9eaa8a2480ac593aa6dad5f13db95c0927f6af631

Observation 5fb69b54-f28c-4bea-95ec-b9069f16e924 · outbound

This paper cites Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:88428ed50bbc31100761a23441bb990b17da0ee50ac1b22cf5837357898643cb

Observation 9d5df397-5493-4268-bdaa-aa093e709d94 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:9d2fa61eb338827c0e3cfea127ebaf9c0b40a5d443fe9eb1422b1b9e7067cf66

Observation 9f79dfb9-1d36-46f9-a59d-b50b33d7fb36 · outbound

This paper cites Galactica: A Large Language Model for Science.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Galactica: A Large Language Model for Science

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:98ac849e4b5ded07cb7ad97f2d958e3014f87fcc2b34cffd53db76142f806fe6

Observation c2f585ae-f956-40fb-ba97-e8247b0a8355 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Kimi K2: Open Agentic Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:ce197f4638149a19d581ca418f7f5571352bd3c4f14b4215f589cbd7a996b5fe

Observation 72e0f637-ddf4-4168-bb56-6415d1dfcb4d · outbound

This paper cites OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:33a8a56d3d11e31eb3d7cb3c0a57a2dd61a9934188de9829d0c139b2f8256361

Observation 10a4e715-c655-4ca4-aee8-55fcf3ac4764 · outbound

This paper cites Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:470b1ce2f1516ecd97e34a4abf0f628d5530fbdd9459fd1ef5ef7b5fdc9bf1ca

Observation d13ccbe8-d9a3-40c6-90d6-4c078dd88cd3 · outbound

This paper cites Larger datasets can be repeated more: A theoretical analysis of multi-epoch scaling in linear regression.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Larger datasets can be repeated more: A theoretical analysis of multi-epoch scaling in linear regression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:0689709c989753253366bb5e0ae5ae83bf63079f89d31fb741c7c78dbe8148c0

Observation 920c41c3-595b-4484-8567-053d0cb39d89 · outbound

This paper cites Nicolas Zucchet, Francesco d’Angelo, Andrew K Lampinen, and Stephanie CY Chan.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Nicolas Zucchet, Francesco d’Angelo, Andrew K Lampinen, and Stephanie CY Chan

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:5aef37457c6eb733c155260b4a70caa9b72650db559dc80fb027bc7cac50c4d6

Observation dff1143d-d849-420a-9631-c9316323ac43 · outbound

This paper cites Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:196ed9a29938362d893a7c67a7dd99edf7fb3ad3ebd6f192735a9485be7387b5

Pith citing papers

No inbound Pith citation observations are available.