Pith. sign in

Paper Citation Record · LEDGER

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

As of 22 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 4 inbound Pith citation observations for arXiv:2509.07972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07972 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:32:36.011814Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:46:01.008817Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T20:15:04.523466Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4718808-baa3-4b27-ac4b-061ff6581a49 · outbound

This paper cites Qsgd: Communication- Efficient SGD via Gradient Quantization and Encoding.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Qsgd: Communication- Efficient SGD via Gradient Quantization and Encoding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:41.308989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:32.871158Z digest=sha256:6df8d003e47a56a60946eacb27bc2b06ee36cd85c7654da9750dce0584d1c018

Observation dc6aff50-060d-44e7-9c7a-b86e3e5caf38 · outbound

This paper cites Duchi, Dylan J.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Duchi, Dylan J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.966682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:32.970569Z digest=sha256:15e4b49478485dd9c4c79aaf453c1835bcefcde9cd1ce05652f823b22632465a

Observation ca417af7-faa8-406f-bd36-1b52e48d41ac · outbound

This paper cites Curtis, and Jorge Nocedal.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Curtis, and Jorge Nocedal

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.749870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.084868Z digest=sha256:5d27e39639580f6d27b6a0de6bc10dbafd13b49fb4c2d0a62e44bedc7a523e27

Observation ec9e8a2e-59b2-4cf0-993b-e3cbc9711f6b · outbound

This paper cites Convex optimization.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convex optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.151389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.151389Z digest=sha256:c92b71bc1d58a514016b92cce9641945c02ee050a0f998e6ba214f8d3b7f4128

Observation 9b97a3c6-abf5-4c56-9ca9-c68d227c8443 · outbound

This paper cites Gradient descent on neural networks typically occurs at the edge of stability.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Gradient descent on neural networks typically occurs at the edge of stability

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.409316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.230781Z digest=sha256:0f622894f166be4952afb1ad32ab84f3853a26d0da91d1dd292a59fd29750972

Observation 4ee23d18-6703-4cd9-ac07-1b29ba8e725c · outbound

This paper cites Robustness to unbounded smoothness of generalized signsgd.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Robustness to unbounded smoothness of generalized signsgd

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.183963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.317211Z digest=sha256:61b36b93db27c55d123b8d4adb51fce26fc84fb709bbfce21c2ef9c68874a68b

Observation ff52d472-7ba2-4df3-b8ab-a87f8c73c2f9 · outbound

This paper cites A loss curvature perspective on training instabilities of deep learning models.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A loss curvature perspective on training instabilities of deep learning models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.910840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.412706Z digest=sha256:d3da6a375842e3002efd7f34353f3f5e3f31ab36b1c6122f5821b947220f952c

Observation 6ebaf8dc-8753-4f3b-817b-50221d9502ab · outbound

This paper cites A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.488985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.488985Z digest=sha256:8de9ac0f2940f97514afa171b30bf23f4ba0c0cacc4e0e8a4ac5c09cb047cdc3

Observation ff8a23f2-32d0-4b8f-b34d-cdcf5d9b0fcd · outbound

This paper cites A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.688153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.601486Z digest=sha256:392ec30de71553b7e1b44573d15a6fc8a4f22291358148008c28b016121fe8a4

Observation cbcc7641-dd17-465f-87d8-f237f7ed0fbc · outbound

This paper cites SGD: General Analysis and Improved Rates.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence SGD: General Analysis and Improved Rates

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.743464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.743464Z digest=sha256:8460d8eb766de950462b56c0b9889e5930d4af76337a346a9d1a2f0b464cec7a

Observation d984b7e2-2db1-4594-be0e-ee657e728883 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.855290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.855290Z digest=sha256:65aab067d8c6c62ce2269be664be377f501d3eed89ad5b0df4fe9e051fa2e767

Observation 7474a037-c114-41b5-a4b6-1033745c136a · outbound

This paper cites Deep residual learning for image recognition.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Deep residual learning for image recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.998114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.998114Z digest=sha256:652f17bea7440532095689714b418b4c0ac801c171cf1273ca60df5e3423e476

Observation 98a5cef4-ef28-45f2-9dd9-f922df3f1c56 · outbound

This paper cites Three factors influencing minima in SGD , 2018.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Three factors influencing minima in SGD , 2018

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.329972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.144952Z digest=sha256:2ebc1494614055a67a803782efc09891667bbbc635bb749b5c482149455fa38b

Observation b47f25e8-7dcc-479d-8228-28461a978702 · outbound

This paper cites Why warmup the learning rate? underlying mechanisms and improvements.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Why warmup the learning rate? underlying mechanisms and improvements

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.070350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.271114Z digest=sha256:8ad21b9379cb1ae02df7adfb57e9d01635417bded650fbf558685282b436d9ed

Observation b7c5919c-8153-4c8a-b900-931a12fdc66c · outbound

This paper cites Better Theory for SGD in the Nonconvex World.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Better Theory for SGD in the Nonconvex World

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.824116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.414677Z digest=sha256:62d86d3d7db429ad0db2b4ceece522f448506b4ea8ef35dd97c2a82809ae7724

Observation afccdaec-29b4-4ae4-b533-bbf112fe3a6f · outbound

This paper cites Distributed learning with compressed gradients.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Distributed learning with compressed gradients

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:32:36.237186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.555629Z digest=sha256:6d29b618669f5262e29c9ebea9a497e05169657a5bfc3999d45e99bab755672e

Observation d3180bf2-e1fc-4b2d-82f1-9378c2df9ebb · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Adam: A Method for Stochastic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:34.617678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:34.617678Z digest=sha256:824b443e5123a74ec5894180c0e895d8394151ffeeb7dd5ceae50a74cee582f9

Observation 07c72ed4-452d-47d7-8528-37717622a62c · outbound

This paper cites Analyzing & reducing the need for learning rate warmup in gpt training.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Analyzing & reducing the need for learning rate warmup in gpt training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.616868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.680021Z digest=sha256:916ca59218583a5de3b0f4694d8112e667764a3deb12c516416aee12f2c24897

Observation 314450eb-82f1-4de0-9886-746ee0df8559 · outbound

This paper cites Convex and Non -convex Optimization Under Generalized Smoothness.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convex and Non -convex Optimization Under Generalized Smoothness

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.378429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.742379Z digest=sha256:9d44a94ff510401635978adbebf0c71c13e339cef5e1d41e040efa32990371e8

Observation 4ef6a172-b240-4d5d-8efd-4f75da586e6b · outbound

This paper cites Convergence of Adam Under Relaxed Assumptions.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convergence of Adam Under Relaxed Assumptions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.127455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.805621Z digest=sha256:628ee79645f601b3f57b44b0fefcc1f4231ce1818e1b66d97dab22a641cd719a

Observation 8f0ad54b-13d1-488d-9177-1fb185bedf73 · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence On the Variance of the Adaptive Learning Rate and Beyond

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.879480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.892504Z digest=sha256:440692c20487d39b6c724e746aab998ad2504bd334d7bb668dfc52034986ab64

Observation 95096954-aea3-4c7a-9389-6f4fe541cb3c · outbound

This paper cites AdaGrad under Anisotropic Smoothness.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence AdaGrad under Anisotropic Smoothness

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:34.950133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:34.950133Z digest=sha256:9a08118609b07713cb6a5dcd58b262a4dd2ff702c0350216719bda2a4f13c0d6

Observation 5c9a1aa6-d925-4778-a765-04781f2f3c79 · outbound

This paper cites Revisiting the last-iterate convergence of stochastic gradient methods.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Revisiting the last-iterate convergence of stochastic gradient methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.018805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.018805Z digest=sha256:1c1ea361dc4130af8bd3059960328900223880a671067e1785b93efb41b729c1

Observation fc032b57-383b-4fde-906c-ae001ac6363c · outbound

This paper cites SGDR : Stochastic gradient descent with warm restarts.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence SGDR : Stochastic gradient descent with warm restarts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.082983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.082983Z digest=sha256:445e8668128fb2a8c638066f38dcc860d714ac8e5d5d0652fb6b42ebdc9e16a9

Observation 06ea6c50-4a2e-49d3-b6bc-fbbec7f9a9ad · outbound

This paper cites Adaptive Gradient Descent without Descent.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Adaptive Gradient Descent without Descent

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.598245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.141570Z digest=sha256:426ff85b1585b84810d40da5b3496777ac9222e5dd2a61b9174df77ca68171cd

Observation 0d3c2c69-fdca-4b8f-afcf-3954453339b0 · outbound

This paper cites Lectures on convex optimization, volume 137.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Lectures on convex optimization, volume 137

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.205786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.205786Z digest=sha256:941773a8c5b7238add3fd8d7c61893b0b4f576fde54b52fc407cf2e749b44b51

Observation d1d7617d-4e8b-4f70-bdbf-b8be3886a8d8 · outbound

This paper cites Global Convergence and Stability of Stochastic Gradient Descent.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Global Convergence and Stability of Stochastic Gradient Descent

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.355620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.260121Z digest=sha256:6856e9a2a12588ac06dd21042661a75a78637c2d772f998f1c1f8647ddd23cc0

Observation 2604dd10-241a-420a-9448-984bf0d78c36 · outbound

This paper cites Understanding Gradient Clipping In Incremental Gradient Methods.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Understanding Gradient Clipping In Incremental Gradient Methods

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.177582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.337374Z digest=sha256:bd5061762687ddc2e7d083a8287102f68b95364d528bc6aa77ec4e4dceca27cb

Observation a48916cd-bc8a-49db-94b1-08c7f8f83388 · outbound

This paper cites Smith, Pieter-Jan Kindermans, and Quoc V.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Smith, Pieter-Jan Kindermans, and Quoc V

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.062095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.416029Z digest=sha256:a1f919e6ea0735b30c540a82b543804a612eb038ccee2aa5aea5ee016d2d6c49

Observation cbb4a111-831c-45db-83f1-ce519472d5c2 · outbound

This paper cites An elementary approach to tight worst case complexity analysis of gradient based methods.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence An elementary approach to tight worst case complexity analysis of gradient based methods

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.979005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.483322Z digest=sha256:4c7746527f0b63802b0ad61e6baeac4bef28efb59a8288c9c0f653f12ef12929

Observation b6f8f43b-f42d-48b6-99c4-75379aa3e0fd · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.560961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.560961Z digest=sha256:1eafdbd98dc9e707fe545276557cd4495855bde046fb69c1e53772a2182f9698

Observation 3b624295-8e1f-4767-99bf-3ffbf28c1104 · outbound

This paper cites Toward a Unified Theory of Gradient Descent under Generalized Smoothness.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Toward a Unified Theory of Gradient Descent under Generalized Smoothness

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.794213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.645274Z digest=sha256:012c88cb97ce43dacaf4bb9c99a16e92798cfbad450f35db4ab981c172bd3966

Observation 55ec8de5-bbd5-4992-8004-caac4d780bb8 · outbound

This paper cites Attention is all you need.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Attention is all you need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.738963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.738963Z digest=sha256:17d74da9b1740fa87b7f2cefb8c262d0e31c3c22614658ebb8fe811eb57061f1

Observation 3e405299-3eb1-4fca-9f5b-0bc3331cc7b8 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.801321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.801321Z digest=sha256:953bbacc7e78c394e3f06d1395d4dc7b39eda24c3121104fc6ffcc41e7683b33

Observation 099ab271-6b10-4276-9f52-cfbd1bd47165 · outbound

This paper cites Improved analysis of clipping algorithms for non-convex optimization.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Improved analysis of clipping algorithms for non-convex optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.878410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.878410Z digest=sha256:02d7f48f4b4980a68924a14bae9cf7b5997bf935e4582c3b3414dc2ee7ec6773

Observation 34286f8d-6728-41e9-8beb-e90599e97273 · outbound

This paper cites Why Gradient Clipping Accelerates Training : A Theoretical Justification for Adaptivity.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Why Gradient Clipping Accelerates Training : A Theoretical Justification for Adaptivity

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.561393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.927559Z digest=sha256:6aa1e7190ccefcd5dfefb9d62226a82a019955c9e0a3a3141c5b57fc2486331f

Observation 4b4e9247-268c-4084-a859-25ca2fecfb87 · outbound

This paper cites On the convergence and improvement of stochastic normalized gradient descent.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence On the convergence and improvement of stochastic normalized gradient descent

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.404112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-04T21:32:36.011814Z digest=sha256:e7c623bccd45f3f5a485f7a7c154895fa04d8420875af9d5ec1cbfa60224a6f2

Pith citing papers

Observation 89caaa2e-7bee-452a-9079-db1bb129da9e · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.891916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.891916Z digest=sha256:fcf90f3b901983e38c147443108884733973fb503881c208e3b2213c256e046f

Observation 1a80413b-21af-4eee-9d4f-a1b31709a47f · inbound

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow cites this paper.

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:26:25.972480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T10:57:04.360455Z digest=sha256:c2213c7bc57607e2e85d21028b89cd047293de561d6f88bb13f1f0782ff69a92

Observation 9c487ebe-6959-4395-86ea-d46ccef35102 · inbound

Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness cites this paper.

Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:15:04.525674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T20:13:36.440752Z digest=sha256:b22fb4f77bff39357604ff1057dc3822f8b8bef19f8d8498525f47579f6d4b34

Observation 22e57482-75f2-4b7e-b7ed-99c2f982a226 · inbound

A Few Accelerated Algorithms for Convex Optimization under $(H_0,H_1)$-Smoothness cites this paper.

A Few Accelerated Algorithms for Convex Optimization under $(H_0,H_1)$-Smoothness Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:46:01.008817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:46:01.008817Z digest=sha256:0ee58d490131b1635a989ad60c8c36cec744d072d45b32aa0ce2ddc79d5420e5