Pith. sign in

Paper Citation Record · LEDGER

The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2502.19002.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.19002 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:14:19.151865Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:06:44.940784Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 25877bc9-d626-4227-85e5-d8ab8e245406 · inbound

GradPower: Powering Gradients for Faster Language Model Pre-Training cites this paper.

GradPower: Powering Gradients for Faster Language Model Pre-Training The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:50:58.405641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T02:48:23.732350Z digest=sha256:954ff007edcfbe594da27cce724c876c6324798bc4d660d5689a6d33b1e82d7a

Observation 3f1e8b5a-c0d8-492c-813a-7a138b35fe2d · inbound

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws cites this paper.

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T04:14:19.151865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:14:19.151865Z digest=sha256:383d562b7c1abfc5932e3a07e7cd9d2aaed426c9ebe084dbdc8433bc781584e8

Observation a360dbec-4d41-4696-b2de-eebb94653473 · inbound

HTMuon: Improving Muon via Heavy-Tailed Spectral Correction cites this paper.

HTMuon: Improving Muon via Heavy-Tailed Spectral Correction The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:55:26.251896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T06:50:29.893126Z digest=sha256:bdd5abef672975cd6af248da47445a2ad2f5158d31ff8a1a081c68c98d9cc461

Observation c1711373-9d3f-429f-a3a8-4747c24c20ad · inbound

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio cites this paper.

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:04.840918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T15:39:51.611115Z digest=sha256:7968b18e31dc73850b12021e10a04558d910708681e9ebeba4fef45dda8e6aa6

Observation 3d071d1a-4f3c-4beb-a4fa-a2db5db6bb12 · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:44:42.616826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T07:44:16.677054Z digest=sha256:2f3504e3528956af848bedbcebd6759882f8549afc317cc8d2d46a00cc8e7d1f

Observation 8b49eecf-6cbd-4f43-9330-9c01c3b7f723 · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.616526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:7cb2d959796257ef6494139ad12ee773139baf29620447e96c998ed0ba299950

Observation d6858e8e-591d-4f87-a6ff-64e406de7cea · inbound

FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo cites this paper.

FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.017012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T15:49:18.685160Z digest=sha256:af4e02274d93c0b9a3ece246af9f16104b6386ac0e10ea40840014fc9c55e470

Observation 492aa194-f467-4031-9608-41f63f9e446c · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.942382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:c5ffc045438f9f595c1d74f3e69929e46d97ba66731ad8d27fba3a12f8d15c02