Pith. sign in

REVIEW 1 cited by

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.06042 v2 pith:LJPVDZY5 submitted 2025-02-09 cs.LG cs.CL

classification cs.LGcs.CL
keywords datamodelpretrainingtargetfinetuningforgettingdomaininjecting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A widespread strategy to obtain a language model that performs well on a target domain is to finetune a pretrained model to perform unsupervised next-token prediction on data from that target domain. Finetuning presents two challenges: (i) if the amount of target data is limited, as in most practical applications, the model will quickly overfit, and (ii) the model will drift away from the original model, forgetting the pretraining data and the generic knowledge that comes with it. We aim to derive scaling laws that quantify these two phenomena for various target domains, amounts of available target data, and model scales. We measure the efficiency of injecting pretraining data into the finetuning data mixture to avoid forgetting and mitigate overfitting. A key practical takeaway from our study is that injecting as little as 1% of pretraining data in the finetuning data mixture prevents the model from forgetting the pretraining set.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A contextual-bandit correction layer with few-shot masked updates improves ML demand forecasts by 3.7–14.9% and cuts inventory costs in two retail datasets.

Pith tools