Pith. sign in

REVIEW 6 cited by

Spike No More: Stabilizing the Pre-training of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.16903 v4 pith:SFCKLZ65 submitted 2023-12-28 cs.CL cs.AI

classification cs.CLcs.AI
keywords pre-traininglargespikeslanguagelossmodelsconditionsduring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Loss spikes often occur during pre-training of large language models. The spikes degrade the performance of large language models and sometimes ruin the pre-training. Since the pre-training needs a vast computational budget, we should avoid such spikes. Based on the assumption that the loss spike is caused by the sudden growth of the gradient norm, we explore factors to keep the gradient norm small through an analysis of the spectral norms of the Jacobian matrices for the sub-layers. Our findings suggest that stabilizing the pre-training process requires two conditions: small sub-layers and large shortcut. We conduct various experiments to empirically verify our theoretical analyses. Experimental results demonstrate that methods satisfying the conditions effectively prevent loss spikes during pre-training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions

    cs.LG 2026-07 conditional novelty 6.0 of 10

    HyenaND is an input-dependent, multi-dimensional convolution that runs in near-linear time and matches attention baselines on vision, medical, PDE, and genomics benchmarks.

  2. When Does Sparsity Mitigate the Curse of Depth in LLMs

    cs.CL 2026-03 conditional novelty 6.0 of 10

    Implicit and explicit sparsity reduce residual-stream variance and improve layer effectiveness metrics, enabling a depth-scaling recipe with about 4.6 points higher downstream accuracy.

  3. DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

    cs.DC 2026-01 conditional novelty 6.0 of 10

    DataStates-LLM hides LLM checkpointing overhead by lazily copying immutable weights during forward/backward passes and streaming heterogeneous shards to storage via composable state providers, cutting end-to-end train...

  4. GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

    cs.LG 2025-06 conditional novelty 6.0 of 10

    GPAS scales down intermediate activations while preserving backward gradients, reducing activation variance growth in Pre-LN transformers and improving pretraining convergence and downstream performance.

  5. Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Weight decay on scale-invariant weights creates a norm-dependent sharpness boundary; crossing it predicts loss spikes in normalized networks.

  6. SpanNorm: Reconciling Training Stability and Performance in Deep Transformers

    cs.CL 2026-01 conditional novelty 5.0 of 10

    SpanNorm—a block-level residual with PostNorm-style normalization—trains deeper transformers more stably and outperforms PreNorm and hybrid normalization on LM benchmarks.

Pith tools