Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:55:18.298267Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2412.02153.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:55:18.298267Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-07T10:57:04.360455Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T09:26:25.980990Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d27cd9ed-7dde-4c26-8e3d-246a2fdad191 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization SIAM review, 60(2):223–311, 2018
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fbc0659b-0346-4dc8-bac3-3b6acd5abaa7 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adaptivesubgradientmethodsforonlinelearning and stochastic optimization.Journal of machine learning research, 12(7), 2011
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0b55d37c-ec4e-489f-a94f-f40e7caf39ea · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Neuralnetworksformachinelearning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7154ba05-afe1-4f39-92c0-adf66094f4ee · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adam: A Method for Stochastic Optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791857b2-973e-4f78-a31b-d11c903f21e4 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Why adam outperforms gradient descent on language models: A heavy-tailed class imbalance problem
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5d3d1554-d265-4b76-8d1e-8be221b1f8a0 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Toward Understanding Why Adam Converges Faster Than SGD for Transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ba7857-1527-425f-ab0b-7398f40742e2 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c5191b5a-a2da-49eb-80b0-249757cd9174 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Why Transformers Need Adam: A Hessian Perspective
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d40876-d56f-4502-a5ef-fc0265e9f73d · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adam can converge without any modification on update rules.Advances in neural information processing systems, 35:28386–28399, 2022
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0e254062-20e2-40e9-8364-d3a6bf27174b · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Ontheconvergenceofadamandbeyond
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b531cc5b-adf0-4d4e-9c55-05dde3af477c · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Onthevarianceoftheadaptivelearningrateandbeyond
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eeb74156-0b34-4583-9847-c698e58a1d5a · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adaptive gradient methods with dy- namic bound of learning rate
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2a2d0d84-fbfa-4931-a3ae-5280e971826b · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c21456-6732-4846-908e-81c9634db9b2 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Momentum is all you need for data-driven adaptive optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd2b5e40-05e0-4b00-bfc2-808a81e52afa · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Survey of optimization algo- rithms in modern neural networks.Mathematics, 11(11):2466, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4e825763-1ff0-458c-bf6f-5ebb7ecc4d23 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization No train no gain: Revisitingefficienttrainingalgorithmsfortransformer-basedlanguagemodels
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9503e7e6-637f-4e50-8912-3ff3dea5a0b7 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Attention is all you need.Advances in Neural Information Processing Systems, 2017
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 616b0594-0698-463c-b76b-efe36d10d86f · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Dissecting adam: The sign, magnitude and variance of stochastic gradients
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e3e569c9-d134-4e9a-91e2-950098a2c5d8 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 483bd74d-39e7-4e6f-ab3b-bc86c1cfd077 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fedc59b5-eaae-43be-8a59-e702ccc86ad0 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18c163d-93a1-47ce-b256-7de82ad31881 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization An image is worth 16x16 words: Transformers for image recognition at scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee428aa-ab62-43fa-afe1-524ba5a3c689 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Escaping the Big Data Paradigm with Compact Transformers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee7bd2b8-1511-4123-9a25-16db88075725 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Improving transformer op- timization through better initialization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d2bc7fba-6729-4177-b206-a07a6563fe18 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Scan and snap: Understanding trainingdynamicsandtokencompositionin1-layertransformer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3dea8294-5312-4dad-8477-2a38c972e11a · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the difficulty of training Recurrent Neural Networks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1632498-320c-4c21-ac5a-f62c91ea4c9d · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Visualizing the loss landscape of neural nets.Advances in neural information processing systems, 31, 2018
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e78a07-1177-471e-9b03-3a0ee96d91df · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the sdes and scal- ing rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e2030e20-49c8-4dec-9706-1db481a493ab · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b306e551-bd51-4afc-b145-a2dc601c30f5 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Understanding the difficulty of training deep feedforward neural networks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa438afe-3957-41d3-a43e-215acaba8e28 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3239ca3-1b01-4c0f-bc14-0c8ed465182f · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Deep residual learning for image recognition
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6e4a4f91-c95a-47b0-8350-dd3f958c4894 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aa2311f3-a0de-4ec4-9659-49bacba4de32 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Long short-term memory.Neural Computation, 9(8):1735–1780, 1997
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9ec2b5c5-28f0-4b26-8b76-093774850168 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization A stochastic approximation method.The annals of mathe- matical statistics, pages 400–407, 1951
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f29ab8a6-0f0f-4bc1-b165-631fff1c70a3 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Some methods of speeding up the convergence of iteration methods.Ussr computational mathematics and mathematical physics, 4(5):1–17, 1964
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae64e145-0cc5-4fda-b4f1-9bc6ba3983c6 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Decoupled weight decay regularization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b053159-668b-4b04-baf2-91b25ac8acf6 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Learningmultiplelayersoffeaturesfromtinyimages
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 767f1587-9603-4c95-a5d5-ead3f3c3b69e · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Imagenet large scale visual recognition challenge.International journal of computer vision, 115:211–252, 2015
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8485ca84-2abe-4b3b-848f-a44ac385636e · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Building a large annotated corpus of english: The penn treebank.Computational linguistics, 19(2):313–330, 1993
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623c7226-ab06-4426-9dfe-b0543400e1c1 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization fairseq: A fast, extensible toolkit for sequence modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fe6288ff-e5e5-494e-b28a-d770ca723e25 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Bleu: amethodforautomatic evaluation of machine translation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b8bfc735-fe4c-4d20-b7fd-e98bc15aedaf · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae5cb523-8896-457c-bfa7-14d0dbffc699 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Gans trained by a two time-scale update rule converge to a local nash equilibrium.Ad- vances in neural information processing systems, 30, 2017
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3873b934-4b51-4775-aa0d-001588530268 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Asymmetric valleys: Beyond sharp and flat local minima
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 36084f00-71b3-4382-ba2a-97d75361106a · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Towards the- oretically understanding why sgd generalizes better than adam in deep learning.Advances in Neural Information Processing Systems, 33:21285–21296, 2020
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9b683fd7-1aef-43d0-bb67-1032c1c28dfc · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the adequacy of untuned warmup for adaptive optimization
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4de66683-7cb4-4aa8-948d-f2f6eb1234f1 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Re- thinkingtheinceptionarchitectureforcomputervision
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7b7fe0e2-1d91-49d4-9b38-9a320a418488 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization SGDR: Stochastic gradient descent with warm restarts
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 018f5842-299b-4eae-9179-df7b4a65e0df · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Linear mode connectivity and the lottery ticket hypothesis
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 478ed16b-4bee-4527-bf8a-ceb2170032bf · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization The small denominator leads to excessively large initial updates, particularly when¯g is small orσ2 is large
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9c013d2f-01a8-4f6a-98c3-74560a32b667 · outbound
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization The model is trained with a length penalty of 1.0, a beam size of 5, and an initial warmup step size of10−7
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c147f485-ea7b-4125-8d91-2a028299233c · inbound
A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.