Pith. sign in

Paper Citation Record · LEDGER

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2412.02153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02153 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:55:18.298267Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T10:57:04.360455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:26:25.980990Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d27cd9ed-7dde-4c26-8e3d-246a2fdad191 · outbound

This paper cites SIAM review, 60(2):223–311, 2018.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization SIAM review, 60(2):223–311, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.875502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.090406Z digest=sha256:db7f2eb18ee7647cb887d4987757531d379079495925a0f5cfe6c57bf71a4631

Observation fbc0659b-0346-4dc8-bac3-3b6acd5abaa7 · outbound

This paper cites Adaptivesubgradientmethodsforonlinelearning and stochastic optimization.Journal of machine learning research, 12(7), 2011.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adaptivesubgradientmethodsforonlinelearning and stochastic optimization.Journal of machine learning research, 12(7), 2011

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.863774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.094726Z digest=sha256:bd7b213f5cef5bcf08859d69e0ae12dbaead48778ed7569f4be4ba6e2481dc0f

Observation 0b55d37c-ec4e-489f-a94f-f40e7caf39ea · outbound

This paper cites Neuralnetworksformachinelearning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Neuralnetworksformachinelearning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.851195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.098513Z digest=sha256:9825ba0057f4536a88295436efe8ffe2996c30b17c8324e603219ca1ea494c4a

Observation 7154ba05-afe1-4f39-92c0-adf66094f4ee · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adam: A Method for Stochastic Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.102499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.102499Z digest=sha256:36da6bf222d2bdcb3ae084a673814e81088192981093e340e4f68c6b5194edba

Observation 791857b2-973e-4f78-a31b-d11c903f21e4 · outbound

This paper cites Why adam outperforms gradient descent on language models: A heavy-tailed class imbalance problem.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Why adam outperforms gradient descent on language models: A heavy-tailed class imbalance problem

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.839261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.106926Z digest=sha256:0db3ae3e899a0bae3ef1d2751cd8ee0264187334fa45173f87805100c47360ca

Observation 5d3d1554-d265-4b76-8d1e-8be221b1f8a0 · outbound

This paper cites Toward Understanding Why Adam Converges Faster Than SGD for Transformers.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.115103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.115103Z digest=sha256:dd020abe090969c4c9754c060ba2024599bd2a6437e9f3f4a0ae0e71eb3b2e78

Observation 74ba7857-1527-425f-ab0b-7398f40742e2 · outbound

This paper cites an unresolved cited work.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:55:18.826280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.119808Z digest=sha256:1b63462ed9b8b36b6efa266d877a2983ee74e146663728d32791eabc9700d96d

Observation c5191b5a-a2da-49eb-80b0-249757cd9174 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Why Transformers Need Adam: A Hessian Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.123604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.123604Z digest=sha256:a0acb86dbc84ef9d0525f441d2cc70a8842adf1cdf66770a7dca90ee9055cc82

Observation 92d40876-d56f-4502-a5ef-fc0265e9f73d · outbound

This paper cites Adam can converge without any modification on update rules.Advances in neural information processing systems, 35:28386–28399, 2022.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adam can converge without any modification on update rules.Advances in neural information processing systems, 35:28386–28399, 2022

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.815678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.127601Z digest=sha256:3a4a2b6839fa7b262e9990bed96c92435e6490fae507a8c8a0de3f86e55dcdc4

Observation 0e254062-20e2-40e9-8364-d3a6bf27174b · outbound

This paper cites Ontheconvergenceofadamandbeyond.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Ontheconvergenceofadamandbeyond

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.804212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.131423Z digest=sha256:7f8a052b75462c03b38b80e04b7c2126b35d8a662968218538862b158a906ffb

Observation b531cc5b-adf0-4d4e-9c55-05dde3af477c · outbound

This paper cites Onthevarianceoftheadaptivelearningrateandbeyond.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Onthevarianceoftheadaptivelearningrateandbeyond

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.793186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.135397Z digest=sha256:4bef8363eab1d46bc340d1563a16acba5335a48150cb7e673f89cbfb0f779555

Observation eeb74156-0b34-4583-9847-c698e58a1d5a · outbound

This paper cites Adaptive gradient methods with dy- namic bound of learning rate.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adaptive gradient methods with dy- namic bound of learning rate

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.781935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.140039Z digest=sha256:5b7d5f358cb048798f6e497bf6bdfcea6b0634e295a6d066624a1173c527f3d5

Observation 2a2d0d84-fbfa-4931-a3ae-5280e971826b · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.143664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.143664Z digest=sha256:c8376352114e862503d1f79e974a2d7d52a2a28a4b3aabf080dd7c608b8159ea

Observation 60c21456-6732-4846-908e-81c9634db9b2 · outbound

This paper cites Momentum is all you need for data-driven adaptive optimization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Momentum is all you need for data-driven adaptive optimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.762047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.147621Z digest=sha256:218bee2d89e77b21007629886f798fa965074c29a43e70e773fd57b114e64954

Observation fd2b5e40-05e0-4b00-bfc2-808a81e52afa · outbound

This paper cites Survey of optimization algo- rithms in modern neural networks.Mathematics, 11(11):2466, 2023.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Survey of optimization algo- rithms in modern neural networks.Mathematics, 11(11):2466, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.749731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.151424Z digest=sha256:8729e74a7f5cc9b3e68c914932da57405831d710f85e131e53688d00e93c2c7f

Observation 4e825763-1ff0-458c-bf6f-5ebb7ecc4d23 · outbound

This paper cites No train no gain: Revisitingefficienttrainingalgorithmsfortransformer-basedlanguagemodels.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization No train no gain: Revisitingefficienttrainingalgorithmsfortransformer-basedlanguagemodels

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.733754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.155290Z digest=sha256:4d5ef3817762fa3b2859411c9c30f57ca7f235d9bee90fbc1972cc14ffe2e5f0

Observation 9503e7e6-637f-4e50-8912-3ff3dea5a0b7 · outbound

This paper cites Attention is all you need.Advances in Neural Information Processing Systems, 2017.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Attention is all you need.Advances in Neural Information Processing Systems, 2017

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.159144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.159144Z digest=sha256:64d8a4ec9ae6aa18a5e436d095c7ad1830e784cf7e252c5d32093ba28119533d

Observation 616b0594-0698-463c-b76b-efe36d10d86f · outbound

This paper cites Dissecting adam: The sign, magnitude and variance of stochastic gradients.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Dissecting adam: The sign, magnitude and variance of stochastic gradients

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.714484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.162590Z digest=sha256:93ef75d67cf0db66cf83c1eb547d801c86b7804b786cb5227d03e005e99c13ef

Observation e3e569c9-d134-4e9a-91e2-950098a2c5d8 · outbound

This paper cites Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.700866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.166576Z digest=sha256:e2ad44ad2ce6e5fa5b5373a05828238f8e999279abc1d65f0a0bef04d059d8dd

Observation 483bd74d-39e7-4e6f-ab3b-bc86c1cfd077 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.171126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.171126Z digest=sha256:f4cc6be98375a225e29d9e0dcb1694d59342ab96ae07383298401b0d406fd5fe

Observation fedc59b5-eaae-43be-8a59-e702ccc86ad0 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.174774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.174774Z digest=sha256:cc844d0f05e6a7d77d6c518904ad96601a68b417cb5812c0c8708b69195047eb

Observation c18c163d-93a1-47ce-b256-7de82ad31881 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization An image is worth 16x16 words: Transformers for image recognition at scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.178936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.178936Z digest=sha256:a97eb7e1b3e9cc9520e853d0a069970a7ac9d6ce622df469a0f0307fff432a97

Observation bee428aa-ab62-43fa-afe1-524ba5a3c689 · outbound

This paper cites Escaping the Big Data Paradigm with Compact Transformers.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Escaping the Big Data Paradigm with Compact Transformers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.183017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.183017Z digest=sha256:44e5447e8d59a1f7b0d088ba5b5a921fdde17a76ae8123522c54bfa10f756bd1

Observation ee7bd2b8-1511-4123-9a25-16db88075725 · outbound

This paper cites Improving transformer op- timization through better initialization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Improving transformer op- timization through better initialization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.681565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.187370Z digest=sha256:7d6a44453c30c48cd1df7ab5c1d84aa778209b009236873cd793e619b94a3d48

Observation d2bc7fba-6729-4177-b206-a07a6563fe18 · outbound

This paper cites Scan and snap: Understanding trainingdynamicsandtokencompositionin1-layertransformer.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Scan and snap: Understanding trainingdynamicsandtokencompositionin1-layertransformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.669067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.190823Z digest=sha256:a21722fdb5038b6cb60306be97156a4fc958097b08155d0a6b40995d5d74d47f

Observation 3dea8294-5312-4dad-8477-2a38c972e11a · outbound

This paper cites On the difficulty of training Recurrent Neural Networks.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the difficulty of training Recurrent Neural Networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.194230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.194230Z digest=sha256:750d78b28a2d19aaf1fce0f3ff871d4b04faab91c252a2c109dcf23255d61045

Observation a1632498-320c-4c21-ac5a-f62c91ea4c9d · outbound

This paper cites Visualizing the loss landscape of neural nets.Advances in neural information processing systems, 31, 2018.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Visualizing the loss landscape of neural nets.Advances in neural information processing systems, 31, 2018

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.198263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.198263Z digest=sha256:211ed7b84872e2300779e10efa8ca13662a2432eeebd8e0fd5414b8ed1206a96

Observation f9e78a07-1177-471e-9b03-3a0ee96d91df · outbound

This paper cites On the sdes and scal- ing rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the sdes and scal- ing rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.650192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.202867Z digest=sha256:36b12923b7ba71b425357d68a274393e5ea9578c99ce29adfb0faf1338c808f1

Observation e2030e20-49c8-4dec-9706-1db481a493ab · outbound

This paper cites an unresolved cited work.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:55:18.638428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.210430Z digest=sha256:a50abbcfdce00a5377f32d5d1012206988ff38eafe955518217c323d03b313bd

Observation b306e551-bd51-4afc-b145-a2dc601c30f5 · outbound

This paper cites Understanding the difficulty of training deep feedforward neural networks.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Understanding the difficulty of training deep feedforward neural networks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.214398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.214398Z digest=sha256:8a16c62ea3c31a183dc1084022a0cb148c3ff695b875beb5c1c53347dfbcb39f

Observation aa438afe-3957-41d3-a43e-215acaba8e28 · outbound

This paper cites On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.218036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.218036Z digest=sha256:1f98dd35a35f026aa52bca9d2b3ddbe2caf3fbb2f300bdf38adda73b86737dbf

Observation a3239ca3-1b01-4c0f-bc14-0c8ed465182f · outbound

This paper cites Deep residual learning for image recognition.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Deep residual learning for image recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.620554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.221996Z digest=sha256:623bb02d6faa75c8c7ec246e3d66e8094342bc91853ad24845373326b67f5b0d

Observation 6e4a4f91-c95a-47b0-8350-dd3f958c4894 · outbound

This paper cites Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.609138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.226209Z digest=sha256:899b268bd97df7f3071a8eaa0a85c65d8843bd7746c8c8b589c77be020dfd4f6

Observation aa2311f3-a0de-4ec4-9659-49bacba4de32 · outbound

This paper cites Long short-term memory.Neural Computation, 9(8):1735–1780, 1997.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Long short-term memory.Neural Computation, 9(8):1735–1780, 1997

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.597359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.230462Z digest=sha256:73660a293a0195ec97875979f8a464de15b0cdb22fb3072f5d900aa86e64526e

Observation 9ec2b5c5-28f0-4b26-8b76-093774850168 · outbound

This paper cites A stochastic approximation method.The annals of mathe- matical statistics, pages 400–407, 1951.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization A stochastic approximation method.The annals of mathe- matical statistics, pages 400–407, 1951

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.585088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.234921Z digest=sha256:914700e3c2ae5d0ed1fb32f8ada49a7689433bf23022ed478ca4db0c01de3e7f

Observation f29ab8a6-0f0f-4bc1-b165-631fff1c70a3 · outbound

This paper cites Some methods of speeding up the convergence of iteration methods.Ussr computational mathematics and mathematical physics, 4(5):1–17, 1964.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Some methods of speeding up the convergence of iteration methods.Ussr computational mathematics and mathematical physics, 4(5):1–17, 1964

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.239074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.239074Z digest=sha256:5011cb3658236d26b387ff149ae909be92c37b6e6276fc456f858b54bccf637f

Observation ae64e145-0cc5-4fda-b4f1-9bc6ba3983c6 · outbound

This paper cites Decoupled weight decay regularization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Decoupled weight decay regularization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.243371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.243371Z digest=sha256:dce2030b6e1a3bed6efa66be413adf4f9aff102bebe2ff2105d1ebd7e1af5b56

Observation 8b053159-668b-4b04-baf2-91b25ac8acf6 · outbound

This paper cites Learningmultiplelayersoffeaturesfromtinyimages.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Learningmultiplelayersoffeaturesfromtinyimages

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.557433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.246993Z digest=sha256:593034792aa29e292735a0be631de2193802f70023f274d9fcf780015ab0f2be

Observation 767f1587-9603-4c95-a5d5-ead3f3c3b69e · outbound

This paper cites Imagenet large scale visual recognition challenge.International journal of computer vision, 115:211–252, 2015.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Imagenet large scale visual recognition challenge.International journal of computer vision, 115:211–252, 2015

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.250663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.250663Z digest=sha256:c1c6e56c3407bd5a34bafa3d93e6540cbf9093d103152906fd706c98cd95bad4

Observation 8485ca84-2abe-4b3b-848f-a44ac385636e · outbound

This paper cites Building a large annotated corpus of english: The penn treebank.Computational linguistics, 19(2):313–330, 1993.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Building a large annotated corpus of english: The penn treebank.Computational linguistics, 19(2):313–330, 1993

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.254472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.254472Z digest=sha256:08daa0f305b8edf64a8b0dae499014867ce23a4450787191b6edae9e4438efec

Observation 623c7226-ab06-4426-9dfe-b0543400e1c1 · outbound

This paper cites fairseq: A fast, extensible toolkit for sequence modeling.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization fairseq: A fast, extensible toolkit for sequence modeling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.533290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.258037Z digest=sha256:945b59217a1ef642e9dd1a6405ee327aab9b5b40a81df7ba7115f23ea7d31a53

Observation fe6288ff-e5e5-494e-b28a-d770ca723e25 · outbound

This paper cites Bleu: amethodforautomatic evaluation of machine translation.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Bleu: amethodforautomatic evaluation of machine translation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.521952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.261515Z digest=sha256:9cc580c60341c4b2805cc683ef4da989768d36392a93d12d49d15cc3d38e4636

Observation b8bfc735-fe4c-4d20-b7fd-e98bc15aedaf · outbound

This paper cites Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.265028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.265028Z digest=sha256:a340bde1bf6cd0cf603d064a8042f7480246992ff4ce4a5d8f61bc586a63fee3

Observation ae5cb523-8896-457c-bfa7-14d0dbffc699 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Ad- vances in neural information processing systems, 30, 2017.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Gans trained by a two time-scale update rule converge to a local nash equilibrium.Ad- vances in neural information processing systems, 30, 2017

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.510175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.269334Z digest=sha256:f76e79491e7fe188c8bb7288d87ffe910f4bf2d0ae02fc580da3853babb41302

Observation 3873b934-4b51-4775-aa0d-001588530268 · outbound

This paper cites Asymmetric valleys: Beyond sharp and flat local minima.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Asymmetric valleys: Beyond sharp and flat local minima

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.499600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.272984Z digest=sha256:a8a568d078ebec690f550d712a9e16f6fb972086da0b992e3c6172a696174f14

Observation 36084f00-71b3-4382-ba2a-97d75361106a · outbound

This paper cites Towards the- oretically understanding why sgd generalizes better than adam in deep learning.Advances in Neural Information Processing Systems, 33:21285–21296, 2020.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Towards the- oretically understanding why sgd generalizes better than adam in deep learning.Advances in Neural Information Processing Systems, 33:21285–21296, 2020

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.487962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.276768Z digest=sha256:7970f43df00a30e3dd3c371a9633eb4da9b2573efdb2ea94c2ca6677b5af02ea

Observation 9b683fd7-1aef-43d0-bb67-1032c1c28dfc · outbound

This paper cites On the adequacy of untuned warmup for adaptive optimization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the adequacy of untuned warmup for adaptive optimization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.475374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.280225Z digest=sha256:05b71640548dd857d0525f749facc931a4247c668fb0447336e425662708f812

Observation 4de66683-7cb4-4aa8-948d-f2f6eb1234f1 · outbound

This paper cites Re- thinkingtheinceptionarchitectureforcomputervision.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Re- thinkingtheinceptionarchitectureforcomputervision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.464169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.283631Z digest=sha256:7cc17ae0d66782505cff934adcec83893cc946d15867ce61c4269546dd9353d6

Observation 7b7fe0e2-1d91-49d4-9b38-9a320a418488 · outbound

This paper cites SGDR: Stochastic gradient descent with warm restarts.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization SGDR: Stochastic gradient descent with warm restarts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.287122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.287122Z digest=sha256:63340f7bdd491a15ba8c1ccad22499545555ca78a81f650b72a852d1e5a0b002

Observation 018f5842-299b-4eae-9179-df7b4a65e0df · outbound

This paper cites Linear mode connectivity and the lottery ticket hypothesis.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Linear mode connectivity and the lottery ticket hypothesis

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.445825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.290543Z digest=sha256:f609387e00c689eab5f849269c3874310c32640a7311d66de7e6dc7d412a100e

Observation 478ed16b-4bee-4527-bf8a-ceb2170032bf · outbound

This paper cites The small denominator leads to excessively large initial updates, particularly when¯g is small orσ2 is large.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization The small denominator leads to excessively large initial updates, particularly when¯g is small orσ2 is large

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.434649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.294054Z digest=sha256:a2d05c69bb842be28c4601f2ac658a0e95f7a456e7c94233ec2d9fec9d084929

Observation 9c013d2f-01a8-4f6a-98c3-74560a32b667 · outbound

This paper cites The model is trained with a length penalty of 1.0, a beam size of 5, and an initial warmup step size of10−7.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization The model is trained with a length penalty of 1.0, a beam size of 5, and an initial warmup step size of10−7

Reference 52

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T23:55:18.423003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.298267Z digest=sha256:0c5edb6aa877d3f9faece6fc14a63ca1459a083ce9a899ffff52eb8bcd54376b

Pith citing papers

Observation c147f485-ea7b-4125-8d91-2a028299233c · inbound

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow cites this paper.

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow Revisiting the Initial Steps in Adaptive Gradient Descent Optimization

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:26:25.982708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T10:57:04.360455Z digest=sha256:da84ddaf3194d8faa5f32b9ac3e2618d7245b7f65bcd6c208173ce8af594a422