Pith. sign in

Paper Citation Record · LEDGER

Linear Convergence of Adaptive Stochastic Gradient Descent

As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:1908.10525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.10525 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:47:31.432012Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ab720da-0006-41a2-8b13-bdf23c017ba5 · outbound

This paper cites (2018): F (xk0−1) ≤ F (x0) + η2L 2 (1 + log( b2 k0−1 b2 0 )).

Linear Convergence of Adaptive Stochastic Gradient Descent (2018): F (xk0−1) ≤ F (x0) + η2L 2 (1 + log( b2 k0−1 b2 0 ))

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:47:31.747751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T10:47:31.421243Z digest=sha256:ef6c5aedc51bd312d7ad7059aa3ca89dd136c71b016e655b3aa30ea0cf228cf0

Observation 8eb866e5-a1fe-4b43-9a4c-f99af31a8853 · outbound

This paper cites However, after tuning η = 10000 in stochastic setting and η = 100 in batch setting, the convergence rate of AdaGrad-Norm is better again.

Linear Convergence of Adaptive Stochastic Gradient Descent However, after tuning η = 10000 in stochastic setting and η = 100 in batch setting, the convergence rate of AdaGrad-Norm is better again

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:47:31.731828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T10:47:31.426103Z digest=sha256:474f85472bfd6f2e5a87cddd5d277f85418171f346b18b7759ecc76b98b121b1

Observation 4b273839-6338-4422-83f8-afdb1f0a4925 · outbound

This paper cites an unresolved cited work.

Linear Convergence of Adaptive Stochastic Gradient Descent Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:47:31.715509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T10:47:31.432012Z digest=sha256:854f2cf4013531546d73760e492d5f00d26f54fda35cd105b547057bff8c95c6

Observation 4c03a4da-f398-43b0-a21e-c62c94452a9d · outbound

This paper cites A Sufficient Condition for Convergences of Adam and RMSProp.

Linear Convergence of Adaptive Stochastic Gradient Descent A Sufficient Condition for Convergences of Adam and RMSProp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.375228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.375228Z digest=sha256:68c07c9b448bf3a588af49606351541da5e0ec0a5f7c0158fe003def478cda3f

Observation 904dc547-dc3b-4335-a6f3-35741be8a6b3 · outbound

This paper cites AdaGrad stepsizes: Sharp convergence over nonconvex landscapes.

Linear Convergence of Adaptive Stochastic Gradient Descent AdaGrad stepsizes: Sharp convergence over nonconvex landscapes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.380155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.380155Z digest=sha256:2dd72098521acb7c51824fb25a3f8e53c8fab16f117b690f2876468ec515f20c

Observation be367608-7198-49ba-b812-13d0d7321d36 · outbound

This paper cites WNGrad: Learn the Learning Rate in Gradient Descent.

Linear Convergence of Adaptive Stochastic Gradient Descent WNGrad: Learn the Learning Rate in Gradient Descent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.385502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.385502Z digest=sha256:436828adb51574949b8771b7833f0db9c73326a3147c49f6ec7e06e5923da753

Observation f6db588b-1399-4e02-9f65-f0dcd41ac1df · outbound

This paper cites On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes.

Linear Convergence of Adaptive Stochastic Gradient Descent On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.390275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.390275Z digest=sha256:b1d9d3739bf6b5e9b38afbe1230318e5e988d43e67ee3ceed4496c461f7247a0

Observation ca916c6b-dc1e-4797-849b-39e7af5b864d · outbound

This paper cites Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron.

Linear Convergence of Adaptive Stochastic Gradient Descent Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.394878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.394878Z digest=sha256:69de029c43f638408e8cc9ef4d7f5a2919ab0c4736609516ddcab835e5ddb909

Observation d50a180f-be27-4024-9110-15c9ec01e24b · outbound

This paper cites Understanding deep learning requires rethinking generalization.

Linear Convergence of Adaptive Stochastic Gradient Descent Understanding deep learning requires rethinking generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.401412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.401412Z digest=sha256:10b088acd693e0feea94556244ab2f14078ae12c666950e26b8353104d1e3798

Observation 4fb55d86-6a6a-4bb4-af75-616871783ff7 · outbound

This paper cites Fast Convergence of Stochastic Gradient Descent under a Strong Growth Condition.

Linear Convergence of Adaptive Stochastic Gradient Descent Fast Convergence of Stochastic Gradient Descent under a Strong Growth Condition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.411341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.411341Z digest=sha256:6e404ad4da6fdf4ace4bd7f285ea10db09039328aac2922c65e7d1848b59ff2a

Observation 77a35abd-4fcb-4c89-b5a7-2933697ae1f4 · outbound

This paper cites Global Convergence of Adaptive Gradient Methods for An Over-parameterized Neural Network.

Linear Convergence of Adaptive Stochastic Gradient Descent Global Convergence of Adaptive Gradient Methods for An Over-parameterized Neural Network

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.416037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.416037Z digest=sha256:f3f3730554d7db3d5ceb0ad5b363c746666e53b1dfca2984367455d181c57a1c

Observation a0b0faa3-ae45-40ad-9c3b-1769a56a0136 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Linear Convergence of Adaptive Stochastic Gradient Descent Adam: A Method for Stochastic Optimization

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.359712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.359712Z digest=sha256:dacbaab8ca5cebcb74d5cf338a345caa8edfaef6b2f0124f13714682eca2daeb

Observation 0f4a679c-2d5e-4bbb-a109-ee8c53aa91ff · outbound

This paper cites Adaptive Bound Optimization for Online Convex Optimization.

Linear Convergence of Adaptive Stochastic Gradient Descent Adaptive Bound Optimization for Online Convex Optimization

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.354327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.354327Z digest=sha256:36e94154eea6dfc8423ff3757ed83a4b0d31ca560a22234d2c270cb378a513b7

Observation a279fb87-f15b-4112-b7d2-7bc0a0e96f4e · outbound

This paper cites Diagonal Rescaling For Neural Networks.

Linear Convergence of Adaptive Stochastic Gradient Descent Diagonal Rescaling For Neural Networks

Reference 2014

Resolution
verified exact
local_arxiv, observed 2026-08-14T10:47:31.631995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T10:47:31.364853Z digest=sha256:8212fef13c883378448d9450364b4d1c5ded23fb5bde32bc5405fdaa1aa8d753

Observation e54fa5fe-e7ce-4bf7-b501-f56eb3228447 · outbound

This paper cites A Convergence Theory for Deep Learning via Over-Parameterization.

Linear Convergence of Adaptive Stochastic Gradient Descent A Convergence Theory for Deep Learning via Over-Parameterization

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.343819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.343819Z digest=sha256:745ffc36c13748c07ace9ab209c81ba2eaf17ae7fc29872c0aad449944a7327a

Observation 9c02d46b-dbf4-4e69-9580-1d8f4f47814f · outbound

This paper cites On the Convergence of Adam and Beyond.

Linear Convergence of Adaptive Stochastic Gradient Descent On the Convergence of Adam and Beyond

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.370129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.370129Z digest=sha256:126de45fcba23d90f933ea862b91ad101f4c4b957e9452800620fc62a93cd244

Observation 2c219bbd-55e3-4a38-bbd7-6e3ab40b0eda · outbound

This paper cites Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks.

Linear Convergence of Adaptive Stochastic Gradient Descent Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.349191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.349191Z digest=sha256:96bcead1f638180ae90a6659646378937f807044be94aa28152af68fa36950d7

Observation e882a8e9-4bb5-4bd9-b6c9-6a8e06f19d78 · outbound

This paper cites On exponential convergence of SGD in non-convex over-parametrized learning.

Linear Convergence of Adaptive Stochastic Gradient Descent On exponential convergence of SGD in non-convex over-parametrized learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.406356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.406356Z digest=sha256:dfbca7c6cc11c713c95d9788d9aee8b3c9c7809b1eceb21d6fafc243b2f114d3

Pith citing papers

No inbound Pith citation observations are available.