Pith. sign in

Paper Citation Record · LEDGER

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

As of 9 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 3 inbound Pith citation observations for arXiv:2509.07972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07972 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:32:36.011814Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:38:58.891916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T20:15:04.523466Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4718808-baa3-4b27-ac4b-061ff6581a49 · outbound

This paper cites Qsgd: Communication- Efficient SGD via Gradient Quantization and Encoding.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Qsgd: Communication- Efficient SGD via Gradient Quantization and Encoding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:41.308989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:32.871158Z digest=sha256:58105948cdd7950635765ee04233e5b6ac6166c2cb587164edb391e2cf33bcc1

Observation dc6aff50-060d-44e7-9c7a-b86e3e5caf38 · outbound

This paper cites Duchi, Dylan J.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Duchi, Dylan J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.966682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:32.970569Z digest=sha256:b493a1c5dc8cb73f1b6acbf66a53763ffd9ec045ded930f8a7e576e913033bba

Observation ca417af7-faa8-406f-bd36-1b52e48d41ac · outbound

This paper cites Curtis, and Jorge Nocedal.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Curtis, and Jorge Nocedal

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.749870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.084868Z digest=sha256:a90a8dd86822e7d0b6a3f147792b4b2b7d264bc69dc3592572e71a15c33a2770

Observation ec9e8a2e-59b2-4cf0-993b-e3cbc9711f6b · outbound

This paper cites Convex optimization.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convex optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.151389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.151389Z digest=sha256:b0a88f5c23d8f6a58bbcd2283a34e1399cccbe797b32c2dc8aa26e592b37091b

Observation 9b97a3c6-abf5-4c56-9ca9-c68d227c8443 · outbound

This paper cites Gradient descent on neural networks typically occurs at the edge of stability.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Gradient descent on neural networks typically occurs at the edge of stability

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.409316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.230781Z digest=sha256:52bd905e92b2c33bb292a7ff3f72f2d16bfc555572f887a9153b0c4c71beba64

Observation 4ee23d18-6703-4cd9-ac07-1b29ba8e725c · outbound

This paper cites Robustness to unbounded smoothness of generalized signsgd.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Robustness to unbounded smoothness of generalized signsgd

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.183963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.317211Z digest=sha256:87dab6565f0e30ed6a9f315257da3b3f30c7ed8e3688c71ea8b7507fb0938575

Observation ff52d472-7ba2-4df3-b8ab-a87f8c73c2f9 · outbound

This paper cites A loss curvature perspective on training instabilities of deep learning models.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A loss curvature perspective on training instabilities of deep learning models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.910840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.412706Z digest=sha256:a595581039bb32d4a97f43384220d1ee27ad56b627233846568657e274d9b848

Observation 6ebaf8dc-8753-4f3b-817b-50221d9502ab · outbound

This paper cites A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.488985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.488985Z digest=sha256:33f621940483f5a9a958e5b71c9a3999de8de57c2773377b30884639c4f34bcc

Observation ff8a23f2-32d0-4b8f-b34d-cdcf5d9b0fcd · outbound

This paper cites A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.688153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.601486Z digest=sha256:c7f779cbc7848c1174586e69e3db2b4ab32f7dadde907203970a6455f8aa80b4

Observation cbcc7641-dd17-465f-87d8-f237f7ed0fbc · outbound

This paper cites SGD: General Analysis and Improved Rates.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence SGD: General Analysis and Improved Rates

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.743464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.743464Z digest=sha256:86ce8e4b3da4d46022a38b44b7c68e5f015539435da321a71acf22e24dea76c1

Observation d984b7e2-2db1-4594-be0e-ee657e728883 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.855290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.855290Z digest=sha256:94e22fd886ca44238c33244b30cfa43bc43168dfeaf7bbf932b82234e9408f29

Observation 7474a037-c114-41b5-a4b6-1033745c136a · outbound

This paper cites Deep residual learning for image recognition.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Deep residual learning for image recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.998114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.998114Z digest=sha256:33df4f0050b3f1e9dc45875f1422ab016717d05c1ade65fb13d6e7019e12804b

Observation 98a5cef4-ef28-45f2-9dd9-f922df3f1c56 · outbound

This paper cites Three factors influencing minima in SGD , 2018.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Three factors influencing minima in SGD , 2018

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.329972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.144952Z digest=sha256:f535b7571d989fe9470b1c3c164d4b4dcc66b23b7836202697fe232597e45395

Observation b47f25e8-7dcc-479d-8228-28461a978702 · outbound

This paper cites Why warmup the learning rate? underlying mechanisms and improvements.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Why warmup the learning rate? underlying mechanisms and improvements

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.070350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.271114Z digest=sha256:663f5e2b3172f4b25def518485e99725062062a504092d2327a8edc4c005ca21

Observation b7c5919c-8153-4c8a-b900-931a12fdc66c · outbound

This paper cites Better Theory for SGD in the Nonconvex World.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Better Theory for SGD in the Nonconvex World

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.824116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.414677Z digest=sha256:69af8308845fc1c3ba54d1e74e1910f786754717f5933be0ffff442b01749434

Observation afccdaec-29b4-4ae4-b533-bbf112fe3a6f · outbound

This paper cites Distributed learning with compressed gradients.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Distributed learning with compressed gradients

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:32:36.237186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.555629Z digest=sha256:cac8d761bd05cb1dd25505a01b15553b5d06ae973e545899985641b9ddd9c520

Observation d3180bf2-e1fc-4b2d-82f1-9378c2df9ebb · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Adam: A Method for Stochastic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:34.617678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:34.617678Z digest=sha256:9f1a9c46353be16fe940d6ccb65a300d5dc3f0217ce511672a2e2e7d02bc1fbe

Observation 07c72ed4-452d-47d7-8528-37717622a62c · outbound

This paper cites Analyzing & reducing the need for learning rate warmup in gpt training.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Analyzing & reducing the need for learning rate warmup in gpt training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.616868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.680021Z digest=sha256:84c9bafd92d78700833837554d4510e4629c2665ad7e4a90c2bae811a3e73194

Observation 314450eb-82f1-4de0-9886-746ee0df8559 · outbound

This paper cites Convex and Non -convex Optimization Under Generalized Smoothness.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convex and Non -convex Optimization Under Generalized Smoothness

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.378429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.742379Z digest=sha256:c44587187a0daec07ee22fe56c8dce42004e4854ff2ae3891143949aa83b0e9e

Observation 4ef6a172-b240-4d5d-8efd-4f75da586e6b · outbound

This paper cites Convergence of Adam Under Relaxed Assumptions.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convergence of Adam Under Relaxed Assumptions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.127455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.805621Z digest=sha256:efeb25cd044977a3d10de6e47bd6a47675192f7a28c536dcf6a26dc029424a6b

Observation 8f0ad54b-13d1-488d-9177-1fb185bedf73 · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence On the Variance of the Adaptive Learning Rate and Beyond

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.879480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.892504Z digest=sha256:ccbef250027228a08fa9815314adb94a7fa4193ef921a8c0245db9b7ca48653a

Observation 95096954-aea3-4c7a-9389-6f4fe541cb3c · outbound

This paper cites AdaGrad under Anisotropic Smoothness.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence AdaGrad under Anisotropic Smoothness

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:34.950133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:34.950133Z digest=sha256:716429de94c634eb4b0923a7ed23f8b004b71e2fb563efcf96a7e45104e55d19

Observation 5c9a1aa6-d925-4778-a765-04781f2f3c79 · outbound

This paper cites Revisiting the last-iterate convergence of stochastic gradient methods.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Revisiting the last-iterate convergence of stochastic gradient methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.018805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.018805Z digest=sha256:2b70ecef23874cbec8d616a22a2a3da9bca1d5739e820402bb6efcd2e6e33e43

Observation fc032b57-383b-4fde-906c-ae001ac6363c · outbound

This paper cites SGDR : Stochastic gradient descent with warm restarts.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence SGDR : Stochastic gradient descent with warm restarts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.082983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.082983Z digest=sha256:94f63507e00afc4c398d57ab2afea9893e8ac7b7f163c804ce255883f4f0b439

Observation 06ea6c50-4a2e-49d3-b6bc-fbbec7f9a9ad · outbound

This paper cites Adaptive Gradient Descent without Descent.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Adaptive Gradient Descent without Descent

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.598245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.141570Z digest=sha256:40d4685265ad4a31e26463311c6aa4fafa188581907e49f01e8bc0754f15e2b4

Observation 0d3c2c69-fdca-4b8f-afcf-3954453339b0 · outbound

This paper cites Lectures on convex optimization, volume 137.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Lectures on convex optimization, volume 137

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.205786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.205786Z digest=sha256:5eaf3dcc9658854c87a19ef125c10ab7ff7bc53cc42a21e9029a1467f9077dc6

Observation d1d7617d-4e8b-4f70-bdbf-b8be3886a8d8 · outbound

This paper cites Global Convergence and Stability of Stochastic Gradient Descent.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Global Convergence and Stability of Stochastic Gradient Descent

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.355620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.260121Z digest=sha256:c180a6ed117432ecddad87327ab7455bf65703ff8270cfcb6cc1138eeba196c5

Observation 2604dd10-241a-420a-9448-984bf0d78c36 · outbound

This paper cites Understanding Gradient Clipping In Incremental Gradient Methods.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Understanding Gradient Clipping In Incremental Gradient Methods

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.177582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.337374Z digest=sha256:92ff48209550e45e5d27cd66d8a82fb5c16553d606bed5730cc632472c52677b

Observation a48916cd-bc8a-49db-94b1-08c7f8f83388 · outbound

This paper cites Smith, Pieter-Jan Kindermans, and Quoc V.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Smith, Pieter-Jan Kindermans, and Quoc V

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.062095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.416029Z digest=sha256:ae2d623924a541f627fca093f1542f247067d1b9df0fc52439540dfde7727413

Observation cbb4a111-831c-45db-83f1-ce519472d5c2 · outbound

This paper cites An elementary approach to tight worst case complexity analysis of gradient based methods.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence An elementary approach to tight worst case complexity analysis of gradient based methods

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.979005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.483322Z digest=sha256:04b9e5e6e723e4054a875f1d3433edc109aaeb0d53bf5264faccf9ca6840976d

Observation b6f8f43b-f42d-48b6-99c4-75379aa3e0fd · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.560961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.560961Z digest=sha256:6dc6b9637238c2d26894151d66b33f8aca42b8cf542330ddd2c9a1c3d0b009ea

Observation 3b624295-8e1f-4767-99bf-3ffbf28c1104 · outbound

This paper cites Toward a Unified Theory of Gradient Descent under Generalized Smoothness.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Toward a Unified Theory of Gradient Descent under Generalized Smoothness

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.794213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.645274Z digest=sha256:98eaad4c4f5c60bac548b60600bcceafedf3b33f9f720ad17a0907758c3ef01b

Observation 55ec8de5-bbd5-4992-8004-caac4d780bb8 · outbound

This paper cites Attention is all you need.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Attention is all you need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.738963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.738963Z digest=sha256:51ef7d70d14d6ab52cd18c76944ee749f2a28f03c7ad2109614951176bd98713

Observation 3e405299-3eb1-4fca-9f5b-0bc3331cc7b8 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.801321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.801321Z digest=sha256:fd6d259fc03e02ebb18727444b1458aac3b0c6354fb83c887006ed959fa123c6

Observation 099ab271-6b10-4276-9f52-cfbd1bd47165 · outbound

This paper cites Improved analysis of clipping algorithms for non-convex optimization.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Improved analysis of clipping algorithms for non-convex optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.878410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.878410Z digest=sha256:2dd382954c5e7ffa1f18ae17eeb08f5ff8c519716ff6ebb25373687fde5a7f5e

Observation 34286f8d-6728-41e9-8beb-e90599e97273 · outbound

This paper cites Why Gradient Clipping Accelerates Training : A Theoretical Justification for Adaptivity.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Why Gradient Clipping Accelerates Training : A Theoretical Justification for Adaptivity

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.561393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.927559Z digest=sha256:d6f57b9ba96afc125f1352e54bfdf85c9cf38f70bc950324c7a08a8d283eaf77

Observation 4b4e9247-268c-4084-a859-25ca2fecfb87 · outbound

This paper cites On the convergence and improvement of stochastic normalized gradient descent.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence On the convergence and improvement of stochastic normalized gradient descent

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.404112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:32:36.011814Z digest=sha256:863ccb2cbe82e16877f02e7bd86b1d968dbc02f3c5fd0328fc5d03c611f382f8

Pith citing papers

Observation 89caaa2e-7bee-452a-9079-db1bb129da9e · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.891916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.891916Z digest=sha256:7188a8127121f0db023d7b8b4b442cbadc49e96a5bb7ae70e0fc0c5716e51724

Observation 1a80413b-21af-4eee-9d4f-a1b31709a47f · inbound

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow cites this paper.

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:26:25.972480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T10:57:04.360455Z digest=sha256:f5d75ede5fdbe014095abbe82de156f5c6f4fdfdd9cb460bcf127cbfacd8d17a

Observation 9c487ebe-6959-4395-86ea-d46ccef35102 · inbound

Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness cites this paper.

Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:15:04.525674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T20:13:36.440752Z digest=sha256:df7bbf9070700f313853a1c4a5a5bdf30d4cab3a606c5f0b66f885a5d936eb64