Pith. sign in

REVIEW 10 cited by

Investigating Continual Pretraining in Large Language Models: Insights and Implications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17400 v2 pith:HLR7UNPA submitted 2024-02-27 cs.CL

classification cs.CL
keywords continualmodelspretrainingdomainsllmsknowledgebetterlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Continual learning (CL) in large language models (LLMs) is an evolving domain that focuses on developing efficient and sustainable training strategies to adapt models to emerging knowledge and achieve robustness in dynamic environments. Our primary emphasis is on continual domain-adaptive pretraining, a process designed to equip LLMs with the ability to integrate new information from various domains while retaining previously learned knowledge. Since existing works concentrate mostly on continual fine-tuning for a limited selection of downstream tasks or training domains, we introduce a new benchmark designed to measure the adaptability of LLMs to changing pretraining data landscapes. We further examine the impact of model size on learning efficacy and forgetting, as well as how the progression and similarity of emerging domains affect the knowledge transfer within these models. Our findings uncover several key insights: (i) continual pretraining consistently improves <1.5B models studied in this work and is also superior to domain adaptation, (ii) larger models always achieve better perplexity than smaller ones when continually pretrained on the same corpus, (iii) smaller models are particularly sensitive to continual pretraining, showing the most significant rates of both learning and forgetting, (iv) continual pretraining boosts downstream task performance of GPT-2 family, (v) continual pretraining enables LLMs to specialize better when the sequence of domains shows semantic similarity while randomizing training domains leads to better transfer and final performance otherwise. We posit that our research establishes a new benchmark for CL in LLMs, providing a more realistic evaluation of knowledge retention and transfer across diverse domains.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Theoretical analysis of continual factual knowledge acquisition shows data replay stabilizes pretrained knowledge by shifting convergence dynamics while regularization only slows forgetting, leading to the STOC method...

  2. Shortcut Solutions Learned by Transformers Impair Continual Compositional Reasoning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    BERT learns shortcut solutions that impair generalization and forward transfer in continual LEGO, while ALBERT learns loop-like solutions for better performance, yet both fail at cross-experience composition, with ALB...

  3. Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Seven stance-making linguistic features strengthen pro-animal-welfare preference in fine-tuned Llama-3.2-1B and Mistral-7B; hedging and concreteness dilute it; first-person is null.

  4. Cortex-Inspired Continual Learning: Unsupervised Instantiation and Recovery of Functional Task Networks

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    FTN achieves near-zero forgetting on continual learning benchmarks by isolating task subnetworks via self-organizing binary masks generated through gradient descent, smoothing, and k-winner-take-all.

  5. Capacity-Aware Mixture Law Enables Efficient LLM Data Optimization

    cs.LG 2026-03 unverdicted novelty 6.0 of 10

    CAMEL is a scaling law capturing nonlinear model-size and mixture interactions to extrapolate optimal data mixtures for large LLMs from small-model experiments, reducing optimization cost by 50% and improving benchmar...

  6. Forward-Only Continual Learning

    cs.LG 2025-09 conditional novelty 6.0 of 10

    FoRo achieves strong continual learning accuracy and low forgetting on CIFAR-100, ImageNet-R, and CUB-200 using only forward updates, via CMA-ES prompt tuning and a recursive knowledge encoding matrix.

  7. The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Fine-tuned LoRA patches still work when applied to a language model that has gone through ten more rounds of continual pretraining, and the paper attributes this portability to near-orthogonality of high-dimensional p...

  8. SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

    cs.CL 2026-07 conditional novelty 5.0 of 10

    An Ascend-NPU training stack reaches 34.22% MFU on DeepSeek-V4-Pro, and a solver-verified CPT+SFT recipe raises OR benchmark averages to 71.81% (Flash) and 77.33% (Pro).

  9. Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    Assertive linguistic features in training data increase LLMs' pro-animal-welfare reasoning while hedged and sensory-description features decrease it.

  10. Threat Modelling using Domain-Adapted Language Models: Empirical Evaluation and Insights

    cs.CR 2026-05 unverdicted novelty 3.0 of 10

    Domain-adapted LLMs and SLMs do not consistently outperform general models on STRIDE threat classification for 5G, with decoding strategies and model scale affecting validity but gains remaining insufficient for reliable use.

Pith tools