Pith. sign in

Paper Citation Record · LEDGER

Why Gradients Rapidly Increase Near the End of Training

As of 14 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 10 inbound Pith citation observations for arXiv:2506.02285.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02285 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:04.384284Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:27:56.628221Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:40:24.372091Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8667ed2e-7c56-4e37-a221-106fbab8bc69 · outbound

This paper cites Layer Normalization.

Why Gradients Rapidly Increase Near the End of Training Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:02.886552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:02.886552Z digest=sha256:e33f5534f90250a30a60fb0e4a076e9ae8889664bde2b58fea1313521d1a8046

Observation 7672aa3e-7fed-494e-ab7b-b28e33ba24e4 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:06.594928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:02.936151Z digest=sha256:74f8f1c169b58b00c5dbad7eb25d6e1dfb79b8be0429b587a8c86a059e32c4a3

Observation 5531a407-aaf0-4ace-8ac2-0c91d91a7ef6 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:06.390649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.009411Z digest=sha256:f2350e74a8f5880a768cc8706fa8b10744f26b6a528bce83e7b69f899f035050

Observation c6c93d21-9bbd-48e5-bd22-a63995c07b03 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:06.122567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.095688Z digest=sha256:54d4bb2180ab0bd4b0dff4eb700a4113cc5040216cfe30c5243ea2dc7b2d08b6

Observation 08db36d8-a0df-4c7c-8936-ef46e243c71d · outbound

This paper cites and Bottou, L.

Why Gradients Rapidly Increase Near the End of Training and Bottou, L

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:05.877245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.175874Z digest=sha256:ec01830afdc2cc8f67d9cef4899a58478deb5bb821e8edc293df56724cf67ad7

Observation 4b0b303b-89ab-4fd6-bd30-7809a85fbde7 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:05.627056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.230459Z digest=sha256:0e76aff77242c17191a8c78f356161da6ffdce1638d5b74f7235aa9af793336b

Observation 2be33ce0-a8e0-4f4e-a52d-2e5194feacd1 · outbound

This paper cites and Gower, R.

Why Gradients Rapidly Increase Near the End of Training and Gower, R

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.275141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.275141Z digest=sha256:c20567f307c6272d47661a7f9aa0e904c033774c8207603e965cee40123bd986

Observation 9a60bbc3-005f-40d8-9af8-f4a85fee0872 · outbound

This paper cites and Mishchenko, K.

Why Gradients Rapidly Increase Near the End of Training and Mishchenko, K

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.377669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.377669Z digest=sha256:e172a10740f0479deeb04b2bbadfcc8dcc6bbc95b799da1c8163cb03ab558c29

Observation 4f5789c7-82a3-4da5-974f-41a7fc2e7191 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:05.400559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.474851Z digest=sha256:021d262ff4d0946f89fcc8d290b4d101ae1719514bfc82468f4f0e3753ac2d9e

Observation 27c72ec4-65e5-4f5a-a0ea-94f42066bec4 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:05.334360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.579676Z digest=sha256:53683e2f42f0f65c244f5907da6cd89549c54f34afc7942093a0d11fe9314b9e

Observation ef705d8e-9ab2-4a36-a65c-b673f8576b72 · outbound

This paper cites and Szegedy, C.

Why Gradients Rapidly Increase Near the End of Training and Szegedy, C

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:05.238177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.676898Z digest=sha256:7a35f56be0c3efe52a7b1d39b7ed38fe19d7074c4caaed02a86a74210542eeee

Observation 9263c93b-63c6-4827-a537-8a53c4918d37 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.723875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.723875Z digest=sha256:dc1afc012d5f5eb15712ab5a15493853540982c337ab5e9fb33aa3cf4e003fb6

Observation 636ae290-6d76-4cf3-b12e-1e9760ad81d5 · outbound

This paper cites and Hutter, F.

Why Gradients Rapidly Increase Near the End of Training and Hutter, F

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.811301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.811301Z digest=sha256:2416c5caf8552dd4d5b179d61623615efb473786c026453c138b24fbf4797b26

Observation 4611d213-a767-43d1-a745-7f4e4bce64e4 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Why Gradients Rapidly Increase Near the End of Training Online Learning: A Modern Introduction Using Convex Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.877272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.877272Z digest=sha256:d1f5083560f4e7dab4d421480bd4c0535cfe3204c7f1b91679dd85d7e086589b

Observation 28337d32-f3c9-4675-ba61-592f81b7276a · outbound

This paper cites B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L.

Why Gradients Rapidly Increase Near the End of Training B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:05.093205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.915870Z digest=sha256:722a1ac011b3e599f29957c2a14c057666e047796234e1a61810899d15e6c0d6

Observation 66ba23e5-347c-45bc-af2c-85bb6040e140 · outbound

This paper cites C., and Fei-Fei, L.

Why Gradients Rapidly Increase Near the End of Training C., and Fei-Fei, L

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:04.951450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.994499Z digest=sha256:a68fa343a66a16ea7fc7977673670d49b2da2f5a78bbb535920255055a6f9a40

Observation 817e3174-5ca4-4ea4-81ba-0167e048b52c · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.884836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.050469Z digest=sha256:8e77749f89a1b35e104fd96371eccf8f5152200f4d89254d81937d9aaa3b1d75

Observation 7c43b120-6677-4a8d-b09a-739151b93b4c · outbound

This paper cites L2 Regularization versus Batch and Weight Normalization.

Why Gradients Rapidly Increase Near the End of Training L2 Regularization versus Batch and Weight Normalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.103673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:04.103673Z digest=sha256:3c66fba18b5f64b5e66071f3b843cc4013302c0409a2cffac5795b14792a6ab0

Observation 69a1258c-d2bc-4057-a133-b27204da1a16 · outbound

This paper cites Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization.

Why Gradients Rapidly Increase Near the End of Training Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.192151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:04.192151Z digest=sha256:11f0f53bdfe17d0d5b1938eebac507ab017b1fc206538935d5bc239486b49c0a

Observation 6b97a015-9c9b-499f-b25b-f1fbb9092572 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.764571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.246337Z digest=sha256:358449cdee102b772a9ebb7b9781ef6c161fc2e9bf1319f00879620e37a9996c

Observation 7db33698-d989-43f0-955a-82cd872ea666 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.653154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.332003Z digest=sha256:a79d41d824cd5d2dfac1b22cace3f290368dadb7b5fbd90e9acec2d2c36eb8d2

Observation 4e3c916a-4668-49dc-9a72-591cd53e0501 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.541396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.384284Z digest=sha256:9619baf241c113683fda72706bcb5b6755b203b1e3302ee819fa1b18ef41eea2

Pith citing papers

Observation 57c806d0-96f4-4a92-99f6-aedbce155fcc · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective Why Gradients Rapidly Increase Near the End of Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:53.777630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:53.777630Z digest=sha256:9ced6623c15037d34ee809dbf7793fe92642c643acc6ab261b1a6eb0db4f730c

Observation 574e5e0b-156e-4985-8126-dcf919846b31 · inbound

Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients cites this paper.

Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients Why Gradients Rapidly Increase Near the End of Training

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T19:10:48.092792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:10:48.092792Z digest=sha256:e8f31380f4333ad63d0b40e35da313179319162090358afe2a89739ed3ba710e

Observation 4f6d60b2-9096-4fc1-ae1f-5a2baabaab1e · inbound

Rethinking Language Model Scaling under Transferable Hypersphere Optimization cites this paper.

Rethinking Language Model Scaling under Transferable Hypersphere Optimization Why Gradients Rapidly Increase Near the End of Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:53:02.525835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T21:51:14.678941Z digest=sha256:6ede3703997fb1f95a945b72d205fae41f5e4884ca8e5285f6fdccec7b25c2b4

Observation f0d07f9d-136b-4fa3-8284-a26bcbc926ab · inbound

Broximal Alignment for Global Non-Convex Optimization cites this paper.

Broximal Alignment for Global Non-Convex Optimization Why Gradients Rapidly Increase Near the End of Training

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T13:00:24.674807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T12:57:52.201675Z digest=sha256:007f04562cdba4b7b6c0be6e29777f462ab93103d2a1d0293b74b980c93f9bab

Observation 366a7266-7668-49cf-a79c-f180b3c32392 · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training Why Gradients Rapidly Increase Near the End of Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:08.661283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:a992d3348c205fd64539f3d132b1022ddc51bc29462a2f8dd4a29bd2fd478f8c

Observation 2c37b115-b48a-4f6b-83e9-cfa5e1166495 · inbound

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less cites this paper.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Why Gradients Rapidly Increase Near the End of Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.465260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:d944a50588b8177eb4ec305153f240ef6ada0cb7a00f5cf0a22c6f2be781623b

Observation 41ab2342-e1d9-4e23-96d5-104b291b08d0 · inbound

Anytime Training with Schedule-Free Spectral Optimization cites this paper.

Anytime Training with Schedule-Free Spectral Optimization Why Gradients Rapidly Increase Near the End of Training

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:40:24.375625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T05:38:16.958574Z digest=sha256:cde0d0611939e8694fdef2c088ea98797f70fc418729ce1fbb8366ce4a3e8a9f

Observation 75168a28-59d3-473b-a6c4-d52b80b9be66 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Why Gradients Rapidly Increase Near the End of Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:ca7579b5ec9d1064118de9746a5ec9d527b02148a090504183def014f6f9c316

Observation cf743017-35ec-4b6e-a066-38f82f0bb173 · inbound

Scale Weight Decay and Train Better cites this paper.

Scale Weight Decay and Train Better Why Gradients Rapidly Increase Near the End of Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.954445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.954445Z digest=sha256:f3be84eb93919ab4398673e9ac786927e96a6ab3ba52a54c8e64ca418ad9a52a

Observation adbc6f64-beda-4a03-9c7f-03f92a182b1a · inbound

Full-bandwidth transformer cites this paper.

Full-bandwidth transformer Why Gradients Rapidly Increase Near the End of Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:56.628221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:27:56.628221Z digest=sha256:69290367473bb5b69bb4173cf72b98ccb36c23894cfc7b75390a5397f6e55240