Pith. sign in

Paper Citation Record · LEDGER

Spike No More: Stabilizing the Pre-training of Large Language Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2312.16903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.16903 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:17:55.863452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:38:43.634840Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ebafdae7-2f71-4a6e-8c3a-b718eab9b0c1 · inbound

YuLan-Mini: An Open Data-efficient Language Model cites this paper.

YuLan-Mini: An Open Data-efficient Language Model Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T05:17:55.863452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:17:55.863452Z digest=sha256:7795f15e72eff5b2a4b0fb5b75db8c648ed198e13f38c6aaf80152214dd56036

Observation 524ec8bc-f2e0-4fec-8c91-549b02490826 · inbound

SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training cites this paper.

SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:25.911183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:56:25.911183Z digest=sha256:196266f215243d441880329dafd3b2dbd79b30530030675df39498aafa88861a

Observation 8a60c67d-ea8b-4711-8474-e9476ac3d168 · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.997471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.997471Z digest=sha256:97c4b610564bf9d5c8943cf441d84c3ed001065c7801cbdd9095ed551f933a56

Observation 17df7458-4181-46be-abcf-a8ad4fba1259 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:39:41.474681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:2a20a85c12963c3dd36bce73b0c9d75ee05fea71b6450b2b8454a18cbe1adab5

Observation 6184a1d1-ae4a-47ff-ae75-e7bbb2156631 · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:13.667569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:13.667569Z digest=sha256:4e4bb3c6caaf648696181bd0106bf2dbc1b89ead3590a0ac3e89f96ec9d567dd

Observation 992f5f56-b36b-417e-bdae-16dc89ca8156 · inbound

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling cites this paper.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.325852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.325852Z digest=sha256:5190549416ba4ebcd1aeb6eeda44c6f260d9f2b1dd693b3f3ee4d935efe463dc

Observation b17893d6-0fc5-462c-8b2f-03367d4f9bed · inbound

Foundation Models for Discovery and Exploration in Chemical Space cites this paper.

Foundation Models for Discovery and Exploration in Chemical Space Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:52:25.010557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T05:52:10.848118Z digest=sha256:21a52677b007dd1ea68f2eeebd1b011718f262ee3072f9d67bed91b9657ba32c

Observation c63a0b42-d534-471a-9761-203a00acd1b7 · inbound

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers cites this paper.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.527546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.527546Z digest=sha256:83af35b22aba8b8e1e43bc8f2f4e2018013700c922b7440f773c0d66a0867abb

Observation 092d188c-22f4-48ec-aecf-9283e1ad4ed6 · inbound

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers cites this paper.

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:37:53.414766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:37:53.414766Z digest=sha256:01198002a8ed5ac167841ac3f5f924f5980a3b3969cff81bc5f2eb2947e0ebc1

Observation ae562f90-4d5a-4e53-a27e-1f3a0d64a612 · inbound

When Does Sparsity Mitigate the Curse of Depth in LLMs cites this paper.

When Does Sparsity Mitigate the Curse of Depth in LLMs Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T20:29:33.439034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:29:33.439034Z digest=sha256:ee695f539b501e58a852a2cb16c167c26ba5935f24b41b50829be20c6ab5a173

Observation aef623dd-8646-42fb-b7cc-058c7ff0384b · inbound

Parcae: Scaling Laws For Stable Looped Language Models cites this paper.

Parcae: Scaling Laws For Stable Looped Language Models Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:01.089544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:33:04.442462Z digest=sha256:0deb6e6f27c4da03486165a9879b00d24a698ef916fe270e8fe74d27f3300ad3

Observation 48f2859d-28c4-47df-bc13-7f552cded153 · inbound

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates cites this paper.

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.440694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T13:34:09.129379Z digest=sha256:af7486d90997232afc01ba258871b251035b4e41297438943e9cde37693fdef7

Observation e9ce5ca0-8ae6-4b13-a4dc-d9c285ad2854 · inbound

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training cites this paper.

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:42:36.035343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T18:53:47.187437Z digest=sha256:ebc53ed82a8b3ce48843c6db39882d38569b1d7ca41a2947876d249bb568c59b

Observation 500ad647-176e-40c1-9fcc-d05e13a516cb · inbound

Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding cites this paper.

Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:43.636171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-03T17:30:39.458521Z digest=sha256:c1686cca742e2c2fa9c36afc8d1f8a133b3c8a4c280d5ae7ebb031065f95fdfc

Observation 39d98781-3a3c-4b92-afaa-a6f7a00e3b52 · inbound

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions cites this paper.

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:47.212034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:47.212034Z digest=sha256:94b5a24148d1629b55dae5a9d9fcd91f7f98dfbb1c24a4aa9fb016daf3a24767

Observation 948613b1-a859-4158-8167-07321c08220c · inbound

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions cites this paper.

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T02:08:20.137063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:08:20.137063Z digest=sha256:d9cfd2b489d485f9589aaaf312f4ef8c13270123e0346d8b8b26c36ef8e6f63a

Observation 05c4c720-1162-4aa0-ad63-e86d1da21025 · inbound

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay cites this paper.

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T08:50:45.563563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:50:45.563563Z digest=sha256:1bf3c966e852566ae6195f261379b300ec6fa1b872b9f0b5cf70678c483d0693