Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T17:01:36.218426Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2502.06042.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T17:01:36.218426Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T22:30:58.664303Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
48 of 48 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 0e539dbd-18d5-4e79-8178-391069d27b79 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling laws for generative mixed-modal language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 54015f12-73e8-4298-bc75-689f38c3d2a1 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Physics in Next-token Prediction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 348d53aa-d60c-4d10-bf07-05829460d3d8 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection An Empirical Study of Scaling Laws for Transfer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad549d4-7fac-4b5a-90c7-53a13cc61fbb · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Chinchilla Scaling: A replication attempt
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d0fa82-5ef3-490b-8ea0-13d01b547b18 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5824e7e-4bcb-4e98-be0e-ebfb8295506d · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0003f1-0c81-4ad5-8ba1-00c8381cdae5 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12aa82ea-4964-4565-8705-4c5e29328f9b · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection B., Werra, L
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd78c85f-f506-471a-ab91-a72768059e65 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection The elements of statistical learning: data mining, inference, and prediction, 2017
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2bf46535-cc36-49d7-abf5-1af770924f88 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Towards a Unified View of Parameter-Efficient Transfer Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df1de12e-323e-4169-a5d7-b9b9546a62ed · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Measuring massive multitask language understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 722dd3be-544e-49c2-891c-b678f9a43b81 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Transfer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d8cfee-df61-4bc6-8e13-6b387f11ea40 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Deep Learning Scaling is Predictable, Empirically
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 670b62d3-df1b-43f4-ac17-63c093676837 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Disentangling and Mitigating the Impact of Task Similarity for Continual Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3629529-3168-4d20-bb8d-ce8902bde973 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Training Compute-Optimal Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ffca25e-5a4d-4446-8e11-50e9128774da · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13153a14-8642-4f09-a584-d0d7da39aedc · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Parameter-Efficient Transfer Learning for NLP
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f570756-ec8c-491d-9c7f-683279943162 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection LoRA: Low-Rank Adaptation of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9eebac-00a9-4b75-9358-e17ecb98a7fd · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection L., Wang, C., Yao, Y., Zhao, C., Zhou, J., Cai, J., Zhai, Z., Ding, N., Jia, C., Zeng, G., dahai li, Liu, Z., and Sun, M
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32c3eda3-14ab-4ce6-820c-2666f8d054b3 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Simple and Scalable Strategies to Continually Pre-train Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1c71820-692a-4f1a-a23e-6be4f017bc7c · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling laws for downstream task performance of large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46cd8a44-6f4f-4b39-bed4-e6aae40e8a25 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Forgetting When Fine-Tuning Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb9e3c6f-cab7-4d45-be90-58871f1ca401 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f69888e6-bb5c-4108-8578-105fac66082d · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Neural Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b50a27-f704-4b8b-b279-6409a8d22899 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e551a4f-ba48-4fff-85fc-cd1a357abb79 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Improved fine-tuning by better leveraging pre-training data
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5bdfefc6-a351-409c-ac89-22579ccc7eac · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection and Hutter, F
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9fa8498c-dad9-43e8-a60c-a0e333f548b3 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection and Hutter, F
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0701b48-8b4b-4e01-86cc-847043c43dbd · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f4ec0a-15e4-410c-9400-07e39d7e2ef8 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22ad3363-1613-43d8-a6f3-8698700c4d9a · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Metaicl: Learning to learn in context
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b42c65b6-0a14-4b5f-a976-f198e5fb034a · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 275cdcbb-c633-49dd-8b0e-b6b879a9b018 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Training language models to follow instructions with human feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774c7d91-55e4-49be-8a31-36fbb86bdd82 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Resolving Discrepancies in Compute-Optimal Scaling of Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1acaf38-48d1-4533-b92d-d1ccf2ea4737 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Self-attention Does Not Need $O(n^2)$ Memory
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f9934a-f0b0-4695-ad68-967cbd97e01a · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Language models are unsupervised multitask learners
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e25533c-d209-44cc-a08c-28141256f423 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Multitask prompted training enables zero-shot task generalization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 642d8bb8-aa55-41cd-a887-240b8c4f5928 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 115c1904-e99d-429f-903e-b3d1bc5afdfe · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Sequence to Sequence Learning with Neural Networks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b9816fa-688c-4341-b4a2-73273cece9e3 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db0581da-92ba-4af8-90c0-2553e2f59445 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Law with Learning Rate Annealing
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 702def6f-d4b2-4b6f-9f4f-80b9658d61cf · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4c8af24-347c-4f88-8d86-0e579484ba57 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c8321d0-e366-4e90-b046-338df60caa46 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection W., Lester, B., Du, N., Dai, A
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec7e55ca-b143-495b-bae8-4fbbdeee3771 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection What makes a high-quality training dataset for large language models: A practitioners' perspective
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e58dd0b-5e36-44bd-a36e-76d72e08d32c · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection When scaling meets LLM finetuning: The effect of data, model and finetuning method
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da611d14-44c9-419b-bc38-a58e6281c029 · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Asymmetry in Low-Rank Adapters of Foundation Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 170695dc-26c0-4cb5-af04-f682e50a23ea · outbound
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection write newline
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a545c2e-3c8f-4a92-9ebf-6f75cc838343 · inbound
Diversity in Large Language Models under Supervised Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 978121c7-1d7d-4fd1-beeb-272e51c1da1b · inbound
Diversity in Large Language Models under Supervised Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 85bd389e-4357-4b8c-aa69-868cf41f28f9 · inbound
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 960bddf1-eedd-4e95-98d1-0465059c9ed9 · inbound
Scaling Laws for Mixture Pretraining Under Data Constraints Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0701f0ff-3b37-47eb-9a35-895ffbbcad32 · inbound
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e5085ee-2d3f-4315-a358-1eabbe3127eb · inbound
Knowledge Editing in Masked Diffusion Language Models Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 85a28b91-5f85-4d9c-af51-c4cff710587a · inbound
A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.