Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:39:03.472216Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 2 inbound Pith citation observations for arXiv:2510.03164.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:39:03.472216Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T20:13:36.440752Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T20:15:04.536047Z
76 of 76 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 728eded1-397e-4d7f-8e45-3cc1e04f527b · outbound
Why Do We Need Warm-up? A Theoretical Perspective plainlm: Language model pretraining in pytorch
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a51332-e1c6-4389-b70f-fc7e06fd687a · outbound
Why Do We Need Warm-up? A Theoretical Perspective vision: Vision model pretraining in pytorch
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 431fc2e4-8aa5-48a6-ad74-3cd69f9691da · outbound
Why Do We Need Warm-up? A Theoretical Perspective Benefits of learning rate annealing for tuning-robustness in stochastic optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 149fb9d5-ab9c-4174-9e23-7fbde6b68020 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Layer Normalization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a889a417-0937-40e6-a32c-1016652dc328 · outbound
Why Do We Need Warm-up? A Theoretical Perspective A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74688c1-4fe4-4611-9377-81ff0ab82c92 · outbound
Why Do We Need Warm-up? A Theoretical Perspective On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c806d0-96f4-4a92-99f6-aedbce155fcc · outbound
Why Do We Need Warm-up? A Theoretical Perspective Why Gradients Rapidly Increase Near the End of Training
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1bde069-6bd4-47e7-9f02-85ea680d3b84 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Optimal Linear Decay Learning Rate Schedules and Further Refinements
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 184c32fa-1ef5-4c17-a20f-3776cb854a4e · outbound
Why Do We Need Warm-up? A Theoretical Perspective An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ed8a55-fe3a-47d2-8dbc-4f922557ea2f · outbound
Why Do We Need Warm-up? A Theoretical Perspective Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b39ca0-925c-4eb8-8faa-2bdc87902b32 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d1f73df-10a9-4d79-b227-7833a31e70a4 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Beyond uniform smoothness: A stopped analysis of adaptive sgd
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e70ac17-cdb1-4ac4-b931-0c8ea7a2d36d · outbound
Why Do We Need Warm-up? A Theoretical Perspective Accelerated stochastic optimization methods under quasar-convexity
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491f06aa-614d-407b-b324-f5fd86136d8c · outbound
Why Do We Need Warm-up? A Theoretical Perspective Convergence of Clipped SGD on Convex $(L_0,L_1)$-Smooth Functions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3373c31-9333-445c-982e-05503cf50e41 · outbound
Why Do We Need Warm-up? A Theoretical Perspective A Loss Curvature Perspective on Training Instability in Deep Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72115b05-738f-407a-ae44-a5cd7453e73a · outbound
Why Do We Need Warm-up? A Theoretical Perspective Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee57c98-a2a0-4b37-b089-570b82feb3a2 · outbound
Why Do We Need Warm-up? A Theoretical Perspective A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a24f7417-8014-4bcd-af6c-ac5c3a40258a · outbound
Why Do We Need Warm-up? A Theoretical Perspective Sgd for structured nonconvex functions: Learning rates, minibatching and interpolation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdc5061f-6659-4cd0-8063-56f34f4ec4bc · outbound
Why Do We Need Warm-up? A Theoretical Perspective Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f633f650-1203-4683-aafd-08491a57fc17 · outbound
Why Do We Need Warm-up? A Theoretical Perspective No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590dcdf8-1c8f-4e78-a4f7-8312681dabf5 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Scaling laws and compute-optimal training beyond fixed training durations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec30366-a3cc-4b14-a855-8cd81cc3c4a4 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Gradient descent learns linear dynamical systems
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9da8f203-1a1f-49b6-9256-473c352cb250 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Deep residual learning for image recognition
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75d76bc-65a8-42f5-b00a-7d9f6a864020 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Gaussian Error Linear Units (GELUs)
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 383d99fb-5cdd-4ed3-a254-cb89e9b486d4 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Near-optimal methods for minimizing star-convex functions and beyond
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a7c06d-9155-4ed8-87d9-c3264157ab62 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Training Compute-Optimal Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6269db09-27c3-495f-9c95-86dc0e42ccac · outbound
Why Do We Need Warm-up? A Theoretical Perspective An empirical analysis of compute-optimal large language model training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fb71426-0151-48e3-ba08-f7b28326a805 · outbound
Why Do We Need Warm-up? A Theoretical Perspective MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 036ce580-74e4-433c-a011-634ce01a186c · outbound
Why Do We Need Warm-up? A Theoretical Perspective Improving transformer optimization through better initialization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00671aad-5bd1-43cb-95ec-42f36a89434a · outbound
Why Do We Need Warm-up? A Theoretical Perspective Loss landscape characterization of neural networks without over-parametrization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d011ce7-8726-40bc-91bb-7a331b5e3f86 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Why warmup the learning rate? underlying mechanisms and improvements
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aef0ebda-5f64-4efa-aed9-2b6e38e4f6ad · outbound
Why Do We Need Warm-up? A Theoretical Perspective Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b36a278-8f3b-4563-8896-c93683abe061 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99514f4e-22b3-4f30-957e-afc94d8b474f · outbound
Why Do We Need Warm-up? A Theoretical Perspective Deep learning without poor local minima
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 116cd242-7c43-4402-9b22-ea8c50783d42 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Adam: A Method for Stochastic Optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 706da65f-d409-42cf-bd5b-c7c9ea81f64f · outbound
Why Do We Need Warm-up? A Theoretical Perspective An alternative view: When does sgd escape local minima? In International conference on machine learning, 2018
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1228a4f-78df-4eb7-9f23-4784716f46cf · outbound
Why Do We Need Warm-up? A Theoretical Perspective Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0c02409-bb40-46c1-b44c-3037fcfe5985 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Analyzing & reducing the need for learning rate warmup in gpt training
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b53decd7-bd22-4517-b052-90244089927c · outbound
Why Do We Need Warm-up? A Theoretical Perspective Convex and non-convex optimization under generalized smoothness
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 719634b7-1aa9-4edb-895a-08254fc9d44c · outbound
Why Do We Need Warm-up? A Theoretical Perspective Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21efe338-a1bc-4817-b4e5-3e35f32d45ad · outbound
Why Do We Need Warm-up? A Theoretical Perspective Aiming towards the minimizers: fast convergence of sgd for overparametrized problems
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a969c6c-cba2-4130-8c93-7308ff18d135 · outbound
Why Do We Need Warm-up? A Theoretical Perspective On the Variance of the Adaptive Learning Rate and Beyond
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89caaa2e-7bee-452a-9079-db1bb129da9e · outbound
Why Do We Need Warm-up? A Theoretical Perspective Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648c8223-3bce-474f-ad61-e981523edfb9 · outbound
Why Do We Need Warm-up? A Theoretical Perspective SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d91fec-cb0e-47cd-adb3-84a865af0838 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Decoupled Weight Decay Regularization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1bac448-bfbd-4ba4-b4da-9731b4c89095 · outbound
Why Do We Need Warm-up? A Theoretical Perspective The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffa64622-a0a7-40a3-a49a-eeea6c54b6bd · outbound
Why Do We Need Warm-up? A Theoretical Perspective Matrix differential calculus with applications to simple, hadamard, and kronecker products
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed497575-5314-43e2-91da-05d9b6c4f757 · outbound
Why Do We Need Warm-up? A Theoretical Perspective An Empirical Model of Large-Batch Training
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcbb8440-442b-4c34-9222-c176b5e1dabb · outbound
Why Do We Need Warm-up? A Theoretical Perspective The fineweb datasets: Decanting the web for the finest text data at scale
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75c04dd8-88c7-448a-bd6c-7b29f0a42f71 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Gradient methods for minimizing functionals
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b7c22af-6cb4-40cd-b65c-e90f2ac141be · outbound
Why Do We Need Warm-up? A Theoretical Perspective Language models are unsupervised multitask learners
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a62f5f-478d-46fb-89b6-07d1f2248832 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b0d5fe-f2e7-4478-b1b0-2521860b2175 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Stepping on the edge: Curvature aware learning rate tuners
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce80e345-2c2a-46c4-b820-4fbd67f88085 · outbound
Why Do We Need Warm-up? A Theoretical Perspective The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e558cf7f-ddfc-4c8a-a555-a88298d98e2f · outbound
Why Do We Need Warm-up? A Theoretical Perspective GLU Variants Improve Transformer
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5a15b6-fd2b-4aba-8e62-141cba87251f · outbound
Why Do We Need Warm-up? A Theoretical Perspective On the generalization benefit of noise in stochastic gradient descent
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ba07a3e-b380-47db-96df-167002843834 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Roformer: Enhanced transformer with rotary position embedding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217d4747-d2a8-4904-a697-80d9c57895c1 · outbound
Why Do We Need Warm-up? A Theoretical Perspective On the importance of initialization and momentum in deep learning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b7c4293-5cfe-4b91-af61-bbeaacf0670a · outbound
Why Do We Need Warm-up? A Theoretical Perspective Fast convergence in learning two-layer neural networks with separable data
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adeb38bf-a005-40ec-8257-9ae372c03ed9 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f68236-ec75-4826-868b-f8c0866214bc · outbound
Why Do We Need Warm-up? A Theoretical Perspective Empirical tests of optimization assumptions in deep learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80581cf-2114-4a12-88fb-cf5e581d546b · outbound
Why Do We Need Warm-up? A Theoretical Perspective Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1650146-ef7d-4966-b190-4e8aa5da6398 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Attention is all you need
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f300618-eb15-4eee-8bcf-2c4bf1534160 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Convergence of adagrad for non-convex objectives: Simple proofs and relaxed assumptions
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc431704-a053-41e8-993a-0db5dd316058 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125f1764-fac9-4207-85eb-4803f87867a0 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Small-scale proxies for large-scale Transformer training instabilities
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a98b8919-b85a-4f15-abff-ca0412f2a3df · outbound
Why Do We Need Warm-up? A Theoretical Perspective On the overlooked pitfalls of weight decay and how to mitigate them: A gradient-norm perspective
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9abf9a17-3192-4b9a-abf4-93183a4fca72 · outbound
Why Do We Need Warm-up? A Theoretical Perspective On layer normalization in the transformer architecture
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e455ef0-d364-4d76-a416-edcf5f517509 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Dive into deep learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2fe7e7d-00db-4cad-840f-f698d4c32a39 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Root mean square layer normalization
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdfc5bf-6ca6-49af-94f2-6e3622e0ceda · outbound
Why Do We Need Warm-up? A Theoretical Perspective Improved analysis of clipping algorithms for non-convex optimization
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c02199-f48d-4b06-a70f-52f3feac5a49 · outbound
Why Do We Need Warm-up? A Theoretical Perspective Why gradient clipping accelerates training: A theoretical justification for adaptivity
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d90af45d-98dc-4e8a-9d6f-b03a9e6d6e45 · outbound
Why Do We Need Warm-up? A Theoretical Perspective On the convergence and improvement of stochastic normalized gradient descent
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba8c39b4-7f2b-47c8-aa2b-6a096121e6db · outbound
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28aafa4e-8ded-42f3-a3f4-0ad38030173d · outbound
Why Do We Need Warm-up? A Theoretical Perspective Unresolved cited work
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c689fca2-05c8-4d5c-bb73-37bdf0933cbd · outbound
Why Do We Need Warm-up? A Theoretical Perspective Unresolved cited work
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e5ed69-2867-42ed-bbba-e29abcf80be0 · inbound
Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning with Real-Time Feedback Why Do We Need Warm-up? A Theoretical Perspective
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2ce85ae-8831-4397-8cca-f6e072d2affe · inbound
Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness Why Do We Need Warm-up? A Theoretical Perspective
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.