Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T21:32:36.011814Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 3 inbound Pith citation observations for arXiv:2509.07972.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T21:32:36.011814Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:38:58.891916Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T20:15:04.523466Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d4718808-baa3-4b27-ac4b-061ff6581a49 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Qsgd: Communication- Efficient SGD via Gradient Quantization and Encoding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dc6aff50-060d-44e7-9c7a-b86e3e5caf38 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Duchi, Dylan J
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ca417af7-faa8-406f-bd36-1b52e48d41ac · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Curtis, and Jorge Nocedal
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ec9e8a2e-59b2-4cf0-993b-e3cbc9711f6b · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convex optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b97a3c6-abf5-4c56-9ca9-c68d227c8443 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Gradient descent on neural networks typically occurs at the edge of stability
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4ee23d18-6703-4cd9-ac07-1b29ba8e725c · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Robustness to unbounded smoothness of generalized signsgd
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ff52d472-7ba2-4df3-b8ab-a87f8c73c2f9 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A loss curvature perspective on training instabilities of deep learning models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6ebaf8dc-8753-4f3b-817b-50221d9502ab · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff8a23f2-32d0-4b8f-b34d-cdcf5d9b0fcd · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cbcc7641-dd17-465f-87d8-f237f7ed0fbc · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence SGD: General Analysis and Improved Rates
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d984b7e2-2db1-4594-be0e-ee657e728883 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7474a037-c114-41b5-a4b6-1033745c136a · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Deep residual learning for image recognition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a5cef4-ef28-45f2-9dd9-f922df3f1c56 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Three factors influencing minima in SGD , 2018
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b47f25e8-7dcc-479d-8228-28461a978702 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Why warmup the learning rate? underlying mechanisms and improvements
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7c5919c-8153-4c8a-b900-931a12fdc66c · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Better Theory for SGD in the Nonconvex World
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation afccdaec-29b4-4ae4-b533-bbf112fe3a6f · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Distributed learning with compressed gradients
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d3180bf2-e1fc-4b2d-82f1-9378c2df9ebb · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Adam: A Method for Stochastic Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c72ed4-452d-47d7-8528-37717622a62c · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Analyzing & reducing the need for learning rate warmup in gpt training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 314450eb-82f1-4de0-9886-746ee0df8559 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convex and Non -convex Optimization Under Generalized Smoothness
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4ef6a172-b240-4d5d-8efd-4f75da586e6b · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convergence of Adam Under Relaxed Assumptions
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8f0ad54b-13d1-488d-9177-1fb185bedf73 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence On the Variance of the Adaptive Learning Rate and Beyond
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 95096954-aea3-4c7a-9389-6f4fe541cb3c · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence AdaGrad under Anisotropic Smoothness
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c9a1aa6-d925-4778-a765-04781f2f3c79 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Revisiting the last-iterate convergence of stochastic gradient methods
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc032b57-383b-4fde-906c-ae001ac6363c · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence SGDR : Stochastic gradient descent with warm restarts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06ea6c50-4a2e-49d3-b6bc-fbbec7f9a9ad · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Adaptive Gradient Descent without Descent
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d3c2c69-fdca-4b8f-afcf-3954453339b0 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Lectures on convex optimization, volume 137
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d7617d-4e8b-4f70-bdbf-b8be3886a8d8 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Global Convergence and Stability of Stochastic Gradient Descent
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2604dd10-241a-420a-9448-984bf0d78c36 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Understanding Gradient Clipping In Incremental Gradient Methods
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a48916cd-bc8a-49db-94b1-08c7f8f83388 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Smith, Pieter-Jan Kindermans, and Quoc V
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cbb4a111-831c-45db-83f1-ce519472d5c2 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence An elementary approach to tight worst case complexity analysis of gradient based methods
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b6f8f43b-f42d-48b6-99c4-75379aa3e0fd · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b624295-8e1f-4767-99bf-3ffbf28c1104 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Toward a Unified Theory of Gradient Descent under Generalized Smoothness
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 55ec8de5-bbd5-4992-8004-caac4d780bb8 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Attention is all you need
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e405299-3eb1-4fca-9f5b-0bc3331cc7b8 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 099ab271-6b10-4276-9f52-cfbd1bd47165 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Improved analysis of clipping algorithms for non-convex optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34286f8d-6728-41e9-8beb-e90599e97273 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Why Gradient Clipping Accelerates Training : A Theoretical Justification for Adaptivity
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b4e9247-268c-4084-a859-25ca2fecfb87 · outbound
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence On the convergence and improvement of stochastic normalized gradient descent
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 89caaa2e-7bee-452a-9079-db1bb129da9e · inbound
Why Do We Need Warm-up? A Theoretical Perspective Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a80413b-21af-4eee-9d4f-a1b31709a47f · inbound
A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9c487ebe-6959-4395-86ea-d46ccef35102 · inbound
Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.