Pith. sign in

Paper Citation Record · LEDGER

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2502.06042.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06042 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:01:36.218426Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:30:58.664303Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0e539dbd-18d5-4e79-8178-391069d27b79 · outbound

This paper cites Scaling laws for generative mixed-modal language models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling laws for generative mixed-modal language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.874266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.045067Z digest=sha256:285fc8de6adefb33e7a69eccbb390b0dddd1417fc853f39d301b89198f645e1a

Observation 54015f12-73e8-4298-bc75-689f38c3d2a1 · outbound

This paper cites Physics in Next-token Prediction.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Physics in Next-token Prediction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-08T17:01:36.675840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.049342Z digest=sha256:0f2b7477dd6067d8860a58ac8184c8b1d41c994efceba3cec3f70bdb2c06b121

Observation 348d53aa-d60c-4d10-bf07-05829460d3d8 · outbound

This paper cites An Empirical Study of Scaling Laws for Transfer.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection An Empirical Study of Scaling Laws for Transfer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.054331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.054331Z digest=sha256:62c9763ceb0b7517bb6fe249f99f068a0f20143a0519435d773e8ca10622d432

Observation 5ad549d4-7fac-4b5a-90c7-53a13cc61fbb · outbound

This paper cites Chinchilla Scaling: A replication attempt.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Chinchilla Scaling: A replication attempt

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.058241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.058241Z digest=sha256:ced13057bc0880dc4dd1da7b6a2b02918184cc14fb415f73624cc5de61a19fad

Observation d3d0fa82-5ef3-490b-8ea0-13d01b547b18 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.062327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.062327Z digest=sha256:fb6dfd748036de341cd81a697646a31f849b005c5ba3ce8e3a118dd170f3f3dd

Observation f5824e7e-4bcb-4e98-be0e-ebfb8295506d · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.066170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.066170Z digest=sha256:e7cf126a6f6632fae959e96e8d084b300e159cc1520e64704442ae56c051cf04

Observation 6f0003f1-0c81-4ad5-8ba1-00c8381cdae5 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.070822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.070822Z digest=sha256:a251557338047c2a865e41e7bd65899336dae97810be6c7576c45395a5cd3c3c

Observation 12aa82ea-4964-4565-8705-4c5e29328f9b · outbound

This paper cites B., Werra, L.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection B., Werra, L

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.864014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.074685Z digest=sha256:9464ceed186b3921e94d91edfc381a78ddd7de3a10631979d17dc930863e8d0f

Observation dd78c85f-f506-471a-ab91-a72768059e65 · outbound

This paper cites The elements of statistical learning: data mining, inference, and prediction, 2017.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection The elements of statistical learning: data mining, inference, and prediction, 2017

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.854328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.078314Z digest=sha256:6fe0fc5b95e29e1ed607a09c49a8a5e268c1ab4ab472b4ac8b74dd01af1bbc9c

Observation 2bf46535-cc36-49d7-abf5-1af770924f88 · outbound

This paper cites Towards a Unified View of Parameter-Efficient Transfer Learning.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Towards a Unified View of Parameter-Efficient Transfer Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.082196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.082196Z digest=sha256:b956ba870bea7edda5bcfad6953ac1904dacf22672258097af000a2e795fe1af

Observation df1de12e-323e-4169-a5d7-b9b9546a62ed · outbound

This paper cites Measuring massive multitask language understanding.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Measuring massive multitask language understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.086743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.086743Z digest=sha256:734285a6995daaee6bbd6d4480e865a097dc41a65e8aa605a8ab080753681101

Observation 722dd3be-544e-49c2-891c-b678f9a43b81 · outbound

This paper cites Scaling Laws for Transfer.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.090967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.090967Z digest=sha256:fbb0204ad2c22a71fe9ef5d26662463bb5dd386c648bd8a52d2d92a6eaa06f2a

Observation 19d8cfee-df61-4bc6-8e13-6b387f11ea40 · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Deep Learning Scaling is Predictable, Empirically

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.095414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.095414Z digest=sha256:74f3d37efa0539ee2064a411cac52af3e218922ce606a1fec03c5f78dd616f41

Observation 670b62d3-df1b-43f4-ac17-63c093676837 · outbound

This paper cites Disentangling and Mitigating the Impact of Task Similarity for Continual Learning.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Disentangling and Mitigating the Impact of Task Similarity for Continual Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-08T17:01:36.579501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.099665Z digest=sha256:7e17560e1315c98b7b9af5c73b6abfeb6e9cdf5e942710fe59417742e7ea508e

Observation d3629529-3168-4d20-bb8d-ce8902bde973 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Training Compute-Optimal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.103707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.103707Z digest=sha256:f8c7002abb2e40b958145dc0f3c16645f031187b2b5c600287e98418e33dbeb5

Observation 7ffca25e-5a4d-4446-8e11-50e9128774da · outbound

This paper cites an unresolved cited work.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:01:36.836051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.107743Z digest=sha256:f59f884d782378be3d63e189b485f4deea87478abe2f5f11ecaef0657f4f7f90

Observation 13153a14-8642-4f09-a584-d0d7da39aedc · outbound

This paper cites Parameter-Efficient Transfer Learning for NLP.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Parameter-Efficient Transfer Learning for NLP

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.110759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.110759Z digest=sha256:3a1a427d8420fbd92d4a168cf030c43506fa6c95cf432db3d37b861b813e977e

Observation 2f570756-ec8c-491d-9c7f-683279943162 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.113930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.113930Z digest=sha256:d3723f6cb46340a0a0df5ddfc5c7e25a04ada151b9b6905cf1141a4e234dc551

Observation 5e9eebac-00a9-4b75-9358-e17ecb98a7fd · outbound

This paper cites L., Wang, C., Yao, Y., Zhao, C., Zhou, J., Cai, J., Zhai, Z., Ding, N., Jia, C., Zeng, G., dahai li, Liu, Z., and Sun, M.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection L., Wang, C., Yao, Y., Zhao, C., Zhou, J., Cai, J., Zhai, Z., Ding, N., Jia, C., Zeng, G., dahai li, Liu, Z., and Sun, M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.825506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.117451Z digest=sha256:e102f57d9a55d85be53f3357e1f4929960248ce360f19550fc851f622f531042

Observation 32c3eda3-14ab-4ce6-820c-2666f8d054b3 · outbound

This paper cites Simple and Scalable Strategies to Continually Pre-train Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Simple and Scalable Strategies to Continually Pre-train Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.120434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.120434Z digest=sha256:1630c5ea0b058c4eba7596516ce98179c8473109d20d3e9d2ec3fa76d735772e

Observation c1c71820-692a-4f1a-a23e-6be4f017bc7c · outbound

This paper cites Scaling laws for downstream task performance of large language models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling laws for downstream task performance of large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.123633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.123633Z digest=sha256:4a1852756eeb4f16ab62a10cc0e4aad55f377aa033b92f9fb3f6e9dd767d494b

Observation 46cd8a44-6f4f-4b39-bed4-e6aae40e8a25 · outbound

This paper cites Scaling Laws for Forgetting When Fine-Tuning Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Forgetting When Fine-Tuning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.126796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.126796Z digest=sha256:427c9da19beadac62402fd1a2901d128092ed1af4186635390ec83156ab9f153

Observation eb9e3c6f-cab7-4d45-be90-58871f1ca401 · outbound

This paper cites Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.130140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.130140Z digest=sha256:8d5f160d47778513ed32e19cd7ad33181dd9d5f1867ebc9d0e43cfec049398e4

Observation f69888e6-bb5c-4108-8578-105fac66082d · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.133857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.133857Z digest=sha256:997ae4c4a7b42483098b7faaaa181a4c7982ce54716f52b9638dcf5ca4c4d377

Observation d4b50a27-f704-4b8b-b279-6409a8d22899 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.137146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.137146Z digest=sha256:4561a6d7f6b7542497a8606fcc17dea026cab03156e62816376845103b6e07f4

Observation 9e551a4f-ba48-4fff-85fc-cd1a357abb79 · outbound

This paper cites Improved fine-tuning by better leveraging pre-training data.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Improved fine-tuning by better leveraging pre-training data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.814908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.140667Z digest=sha256:199d9a2c47101d83c726070fcde8851e0b3519f7d675e86ff5bfdab99f83181a

Observation 5bdfefc6-a351-409c-ac89-22579ccc7eac · outbound

This paper cites and Hutter, F.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection and Hutter, F

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.803295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.144198Z digest=sha256:462d806e3f2f80c606b7fed8f9ecedb3dab08b1e5ec116ec4611aa468ab42595

Observation 9fa8498c-dad9-43e8-a60c-a0e333f548b3 · outbound

This paper cites and Hutter, F.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection and Hutter, F

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.147470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.147470Z digest=sha256:07a6d953130663430d9230c99ba9dcfc836841b124ca5d61e1c6dbe7600a78ec

Observation b0701b48-8b4b-4e01-86cc-847043c43dbd · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.150551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.150551Z digest=sha256:0b40e2dfa4c91470a948426f11910861fb30922e7ea1c1fed581dbf8f0652da3

Observation 89f4ec0a-15e4-410c-9400-07e39d7e2ef8 · outbound

This paper cites LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.154116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.154116Z digest=sha256:67948fcaf7da594453edaa8e8d233c9fe9a293c84e09c8b862fea75ddc972037

Observation 22ad3363-1613-43d8-a6f3-8698700c4d9a · outbound

This paper cites Metaicl: Learning to learn in context.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Metaicl: Learning to learn in context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.786124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.157349Z digest=sha256:37b9f01b5c2afee64c1563ce6a83b0f5556ace95e3601881258f012e6961e7ac

Observation b42c65b6-0a14-4b5f-a976-f198e5fb034a · outbound

This paper cites an unresolved cited work.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.160335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.160335Z digest=sha256:2cccb0c589b9a8fa4a8362c919b571dca7a9fb2ae00cad4989ba6c19f0cf911f

Observation 275cdcbb-c633-49dd-8b0e-b6b879a9b018 · outbound

This paper cites Training language models to follow instructions with human feedback.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.163627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.163627Z digest=sha256:3f4420af43b404172da450e08cc1171154434ff80854a438759e32b5fb7855e3

Observation 774c7d91-55e4-49be-8a31-36fbb86bdd82 · outbound

This paper cites Resolving Discrepancies in Compute-Optimal Scaling of Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.166405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.166405Z digest=sha256:c25ee18f5b0b291654cc756aefbc562ea92f0ba0127240de2fa00ec17e87b639

Observation d1acaf38-48d1-4533-b92d-d1ccf2ea4737 · outbound

This paper cites Self-attention Does Not Need $O(n^2)$ Memory.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Self-attention Does Not Need $O(n^2)$ Memory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.169940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.169940Z digest=sha256:f1670a179e9b16dc99cc58d93a64b5bbce61b151bc9874934d4ddbf6b837bf08

Observation d3f9934a-f0b0-4695-ad68-967cbd97e01a · outbound

This paper cites Language models are unsupervised multitask learners.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Language models are unsupervised multitask learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.760912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.174011Z digest=sha256:702a1717ee9b4deecbee6beb8d17fafaa3ab60d62a040b754602c9d1f80cecfd

Observation 9e25533c-d209-44cc-a08c-28141256f423 · outbound

This paper cites Multitask prompted training enables zero-shot task generalization.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Multitask prompted training enables zero-shot task generalization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.749186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.177949Z digest=sha256:c3bef0baf27957b642f702b3099194411eee6bd1bd185cb58d8827ba4e376888

Observation 642d8bb8-aa55-41cd-a887-240b8c4f5928 · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.181915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.181915Z digest=sha256:142f74b5dc7444b525171e3c575c5fd6aefdb910c90b3f37ed3d186efdfb68c1

Observation 115c1904-e99d-429f-903e-b3d1bc5afdfe · outbound

This paper cites Sequence to Sequence Learning with Neural Networks.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Sequence to Sequence Learning with Neural Networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.186433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.186433Z digest=sha256:f71fef1ee5d277c8d386c7fe7a98616c6f4a708ce7b7557ea3f312f23261af15

Observation 3b9816fa-688c-4341-b4a2-73273cece9e3 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.738109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.190495Z digest=sha256:0f3d37c7156c7d38880acd4be60343ceb1c990ec03d3352cdefd921b706ffc33

Observation db0581da-92ba-4af8-90c0-2553e2f59445 · outbound

This paper cites Scaling Law with Learning Rate Annealing.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Law with Learning Rate Annealing

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.194110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.194110Z digest=sha256:6ab3f1802713ff5115b564ad7869a154fe408f84532be65fcc2d2726dc7b5104

Observation 702def6f-d4b2-4b6f-9f4f-80b9658d61cf · outbound

This paper cites When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.198148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.198148Z digest=sha256:ecbaf1d8d0e2eeb7e9289e2dbe3c08ca1a40554039dc9b9c82104256f44b0c3f

Observation e4c8af24-347c-4f88-8d86-0e579484ba57 · outbound

This paper cites an unresolved cited work.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:01:36.727455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.202404Z digest=sha256:b6ecfb26d66f21ed49f309bd727fbb80d3a7a62504d521116f7d73b96c383184

Observation 9c8321d0-e366-4e90-b046-338df60caa46 · outbound

This paper cites W., Lester, B., Du, N., Dai, A.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection W., Lester, B., Du, N., Dai, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.717023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.205762Z digest=sha256:d7e1695ccf932320a265ab4c0f787026cce6506bd83f830b0ed2580232e39c46

Observation ec7e55ca-b143-495b-bae8-4fbbdeee3771 · outbound

This paper cites What makes a high-quality training dataset for large language models: A practitioners' perspective.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection What makes a high-quality training dataset for large language models: A practitioners' perspective

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.705655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.208846Z digest=sha256:5abac2e0ef2862fdf497218d9aea2b81ce9c7a6cb77f02acf4ca9dab58b17ba5

Observation 8e58dd0b-5e36-44bd-a36e-76d72e08d32c · outbound

This paper cites When scaling meets LLM finetuning: The effect of data, model and finetuning method.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection When scaling meets LLM finetuning: The effect of data, model and finetuning method

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.693831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.211919Z digest=sha256:98d000d75f36fe0759fc44508d10e5335bbfad7374b4a13a72e1115d89b81a3a

Observation da611d14-44c9-419b-bc38-a58e6281c029 · outbound

This paper cites Asymmetry in Low-Rank Adapters of Foundation Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Asymmetry in Low-Rank Adapters of Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.214939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.214939Z digest=sha256:fccfe04b330fc229a6f6a666804b96ecd9b0b5192a37959581c6b05ed7b717a1

Observation 170695dc-26c0-4cb5-af04-f682e50a23ea · outbound

This paper cites write newline.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.218426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.218426Z digest=sha256:66f6872ca10c449449407035caef0bc7db34b1779f3bffe9c0ed113798ce032f

Pith citing papers

Observation 6a545c2e-3c8f-4a92-9ebf-6f75cc838343 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:37:32.124103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:bd74563848baa47fb239301a3df07f36d2cae4005934a19faa1fc6b87a3bb4c7

Observation 978121c7-1d7d-4fd1-beeb-272e51c1da1b · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.148664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:1971dfb94bdd1902d3f58a0777cb4b15c5fbde2d3c2f68ac731013d3049526bd

Observation 85bd389e-4357-4b8c-aa69-868cf41f28f9 · inbound

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning cites this paper.

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.617673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:45:27.673290Z digest=sha256:4437651c046dadcb43703a68fa0772c6eafb9cd6aa63932d087bf08b5a4fd459

Observation 960bddf1-eedd-4e95-98d1-0465059c9ed9 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.978072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:ae355d6329b76cf117032f912838819fb717638187386d3a3479af78df8d7358

Observation 0701f0ff-3b37-47eb-9a35-895ffbbcad32 · inbound

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay cites this paper.

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:01.970658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T23:15:45.174086Z digest=sha256:dd3cabfaed50b20b487df45c132b1b9f270b1ef9aa4627b7a3cb8cfdeb314089

Observation 0e5085ee-2d3f-4315-a358-1eabbe3127eb · inbound

Knowledge Editing in Masked Diffusion Language Models cites this paper.

Knowledge Editing in Masked Diffusion Language Models Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:36:29.935755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T09:51:42.811856Z digest=sha256:cb055abe1140a0fa36fb493e58fc6cd470c9fa2dac2f1b2b858e563fbf78855a

Observation 85a28b91-5f85-4d9c-af51-c4cff710587a · inbound

A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting cites this paper.

A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T22:30:58.664303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:30:58.664303Z digest=sha256:ef04e771e6d4d4e6996a1a6b38a19afc2b0e74fdee8744449d163f5add049cc8