Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T10:45:46.618668Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.04969.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T10:45:46.618668Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0a1d000a-3215-432f-be92-848e41b96587 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Phi-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21cff844-1be4-4e10-bcde-fd3fc1f8428f · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57b7f3bd-d63e-46e9-8f64-c6b2c8044616 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac4162bd-e12c-4948-a030-39ba601e5c8c · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Reformulation for Pretraining Data Augmentation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f37e195-1eca-4716-a8b1-1c6d23669b05 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Scaling Laws and Interpretability of Learning from Repeated Data
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0736a376-dae2-4317-820f-6f4f87747bc7 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Training Compute-Optimal Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e364a06c-7f2d-4cab-aeaf-bc5df9d406ab · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Ultra-Sparse Memory Network
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b965e8a-d04d-415a-a662-6fe5946470a2 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Ziyue Li, Chenrui Fan, and Tianyi Zhou
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0dfc0bd-c81c-49e7-a214-27aa69869332 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Let's Verify Step by Step
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ddfe9df-2dac-499f-83e3-0f48e186728f · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training How much do language models memorize?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 016b0d27-4f1c-4fdf-9077-d0a313695566 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Olmo 3
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb69b54-f28c-4bea-95ec-b9069f16e924 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5df397-5493-4268-bdaa-aa093e709d94 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f79dfb9-1d36-46f9-a59d-b50b33d7fb36 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Galactica: A Large Language Model for Science
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2f585ae-f956-40fb-ba97-e8247b0a8355 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Kimi K2: Open Agentic Intelligence
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72e0f637-ddf4-4168-bb56-6415d1dfcb4d · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10a4e715-c655-4ca4-aee8-55fcf3ac4764 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d13ccbe8-d9a3-40c6-90d6-4c078dd88cd3 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Larger datasets can be repeated more: A theoretical analysis of multi-epoch scaling in linear regression
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 920c41c3-595b-4484-8567-053d0cb39d89 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Nicolas Zucchet, Francesco d’Angelo, Andrew K Lampinen, and Stephanie CY Chan
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff1143d-d849-420a-9631-c9316323ac43 · outbound
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.