Pith. sign in

Paper Citation Record · LEDGER

Spike No More: Stabilizing the Pre-training of Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2312.16903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.16903 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:18:13.667569Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:38:43.634840Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 17df7458-4181-46be-abcf-a8ad4fba1259 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:39:41.474681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:6887dbd6914ef710da4bd50887737bfabd6938dab85ba40bb7695ba148d19818

Observation 6184a1d1-ae4a-47ff-ae75-e7bbb2156631 · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:13.667569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:13.667569Z digest=sha256:2dab069c5df44f569c347e53cb10ebf66e9cc2247816c1b13d229f10ad904b35

Observation 992f5f56-b36b-417e-bdae-16dc89ca8156 · inbound

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling cites this paper.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.325852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.325852Z digest=sha256:01897f2a738407c79d281e49296b5ab7e63e60a7dc39cab1f090e0502733c96d

Observation b17893d6-0fc5-462c-8b2f-03367d4f9bed · inbound

Foundation Models for Discovery and Exploration in Chemical Space cites this paper.

Foundation Models for Discovery and Exploration in Chemical Space Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:52:25.010557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:52:10.848118Z digest=sha256:7a62cfee74b91f4bec90c741ab763dc0dbbf3e02ba8d6db591f6bc5a28d82fdf

Observation c63a0b42-d534-471a-9761-203a00acd1b7 · inbound

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers cites this paper.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.527546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.527546Z digest=sha256:6a6e76a712dbb69340066af2a7f1bb669767b37027ead385de2e4852044e14cf

Observation 092d188c-22f4-48ec-aecf-9283e1ad4ed6 · inbound

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers cites this paper.

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:37:53.414766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:37:53.414766Z digest=sha256:6bd7f2c93dfa753793bc01b034812e3db04b22a234bc3f2822c722a7094462e8

Observation ae562f90-4d5a-4e53-a27e-1f3a0d64a612 · inbound

When Does Sparsity Mitigate the Curse of Depth in LLMs cites this paper.

When Does Sparsity Mitigate the Curse of Depth in LLMs Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T20:29:33.439034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:29:33.439034Z digest=sha256:f963dee953608f52777f1dd3857a0c2ff309827629033c9096e785c4306a8cc2

Observation aef623dd-8646-42fb-b7cc-058c7ff0384b · inbound

Parcae: Scaling Laws For Stable Looped Language Models cites this paper.

Parcae: Scaling Laws For Stable Looped Language Models Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:01.089544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:33:04.442462Z digest=sha256:966abb4eed23d8e8f3c2f8726c488e5bea08bb802efa7cfdbde500da1f753065

Observation 48f2859d-28c4-47df-bc13-7f552cded153 · inbound

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates cites this paper.

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.440694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T13:34:09.129379Z digest=sha256:7c5d9b1881ff949bf08b56f3c48a1da36a96de7e42793727ad1971a2ca967afb

Observation e9ce5ca0-8ae6-4b13-a4dc-d9c285ad2854 · inbound

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training cites this paper.

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:42:36.035343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T18:53:47.187437Z digest=sha256:6c04ba49f6c70065a9b435f21090706f35865592e44019003751b9257e1aa2ce

Observation 500ad647-176e-40c1-9fcc-d05e13a516cb · inbound

Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding cites this paper.

Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:43.636171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-03T17:30:39.458521Z digest=sha256:be86839dcd5cb92abaa51a9b9373b3f7049ff38fe17d6aeeb86213f77cc8a6f3

Observation 39d98781-3a3c-4b92-afaa-a6f7a00e3b52 · inbound

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions cites this paper.

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:47.212034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:47.212034Z digest=sha256:1933c522d1d25697e8262b7f4aa6d281ef488e3156088af45226ab4af838efa0

Observation 948613b1-a859-4158-8167-07321c08220c · inbound

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions cites this paper.

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T02:08:20.137063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:08:20.137063Z digest=sha256:f4c138681887148d0462ef8254248b4cf64e5e01cd2757e2d1245d27d69d0eff

Observation 05c4c720-1162-4aa0-ad63-e86d1da21025 · inbound

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay cites this paper.

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T08:50:45.563563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:50:45.563563Z digest=sha256:f5a1271ebc1741a24655312b86b036b872c3bcd681fc5d1407f9c923b254e102