Pith. sign in

Paper Citation Record · LEDGER

Linear Convergence of Adaptive Stochastic Gradient Descent

As of 22 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:1908.10525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.10525 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:47:31.432012Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ab720da-0006-41a2-8b13-bdf23c017ba5 · outbound

This paper cites (2018): F (xk0−1) ≤ F (x0) + η2L 2 (1 + log( b2 k0−1 b2 0 )).

Linear Convergence of Adaptive Stochastic Gradient Descent (2018): F (xk0−1) ≤ F (x0) + η2L 2 (1 + log( b2 k0−1 b2 0 ))

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:47:31.747751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T10:47:31.421243Z digest=sha256:dca3ab87770b158b1edddc209aa095b58e4a032c5d0c06d8bcbbfa3339c9290d

Observation 8eb866e5-a1fe-4b43-9a4c-f99af31a8853 · outbound

This paper cites However, after tuning η = 10000 in stochastic setting and η = 100 in batch setting, the convergence rate of AdaGrad-Norm is better again.

Linear Convergence of Adaptive Stochastic Gradient Descent However, after tuning η = 10000 in stochastic setting and η = 100 in batch setting, the convergence rate of AdaGrad-Norm is better again

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:47:31.731828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T10:47:31.426103Z digest=sha256:62292a782860ece16dddd0dd7c9dc1b0ffd3c66201f93ba0bdd10a222aa1ec16

Observation 4b273839-6338-4422-83f8-afdb1f0a4925 · outbound

This paper cites an unresolved cited work.

Linear Convergence of Adaptive Stochastic Gradient Descent Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:47:31.715509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T10:47:31.432012Z digest=sha256:1b05433c75640eaf94f231181419fd17707aee7f410ad8e80e9504c8d463b65a

Observation 4c03a4da-f398-43b0-a21e-c62c94452a9d · outbound

This paper cites A Sufficient Condition for Convergences of Adam and RMSProp.

Linear Convergence of Adaptive Stochastic Gradient Descent A Sufficient Condition for Convergences of Adam and RMSProp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.375228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.375228Z digest=sha256:13c608d7b3a8471a9e30bcb471ccae3eb0d425df394d043e3ecf06bd2c5326de

Observation 904dc547-dc3b-4335-a6f3-35741be8a6b3 · outbound

This paper cites AdaGrad stepsizes: Sharp convergence over nonconvex landscapes.

Linear Convergence of Adaptive Stochastic Gradient Descent AdaGrad stepsizes: Sharp convergence over nonconvex landscapes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.380155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.380155Z digest=sha256:411379cefad2d06cf378f664a3641e16f26d5a1b611b6716ca526b58be03771c

Observation be367608-7198-49ba-b812-13d0d7321d36 · outbound

This paper cites WNGrad: Learn the Learning Rate in Gradient Descent.

Linear Convergence of Adaptive Stochastic Gradient Descent WNGrad: Learn the Learning Rate in Gradient Descent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.385502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.385502Z digest=sha256:436828adb51574949b8771b7833f0db9c73326a3147c49f6ec7e06e5923da753

Observation f6db588b-1399-4e02-9f65-f0dcd41ac1df · outbound

This paper cites On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes.

Linear Convergence of Adaptive Stochastic Gradient Descent On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.390275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.390275Z digest=sha256:4ce1c77433660710fa32bd3623481fe0a38000be95d67efbb87e0e9943c7f315

Observation ca916c6b-dc1e-4797-849b-39e7af5b864d · outbound

This paper cites Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron.

Linear Convergence of Adaptive Stochastic Gradient Descent Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.394878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.394878Z digest=sha256:a029e1ff0cc689f9de82e6427ded71e23882819a5aed5e1aa49c373e290239fc

Observation d50a180f-be27-4024-9110-15c9ec01e24b · outbound

This paper cites Understanding deep learning requires rethinking generalization.

Linear Convergence of Adaptive Stochastic Gradient Descent Understanding deep learning requires rethinking generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.401412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.401412Z digest=sha256:10b088acd693e0feea94556244ab2f14078ae12c666950e26b8353104d1e3798

Observation 4fb55d86-6a6a-4bb4-af75-616871783ff7 · outbound

This paper cites Fast Convergence of Stochastic Gradient Descent under a Strong Growth Condition.

Linear Convergence of Adaptive Stochastic Gradient Descent Fast Convergence of Stochastic Gradient Descent under a Strong Growth Condition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.411341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.411341Z digest=sha256:397294cd7bd59535f6c42464a0f11e99c21f1a28b1bc5c16586b2c24be9a497b

Observation 77a35abd-4fcb-4c89-b5a7-2933697ae1f4 · outbound

This paper cites Global Convergence of Adaptive Gradient Methods for An Over-parameterized Neural Network.

Linear Convergence of Adaptive Stochastic Gradient Descent Global Convergence of Adaptive Gradient Methods for An Over-parameterized Neural Network

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.416037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.416037Z digest=sha256:9b9d2b682c9707ba7bfa813036faca6985f919a7ce1a7ad1b906f6434518ea88

Observation a0b0faa3-ae45-40ad-9c3b-1769a56a0136 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Linear Convergence of Adaptive Stochastic Gradient Descent Adam: A Method for Stochastic Optimization

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.359712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.359712Z digest=sha256:deef6d0233270f6b848959ae81b35c71d03310ad95278eafc5f25e5e89c4a16b

Observation 0f4a679c-2d5e-4bbb-a109-ee8c53aa91ff · outbound

This paper cites Adaptive Bound Optimization for Online Convex Optimization.

Linear Convergence of Adaptive Stochastic Gradient Descent Adaptive Bound Optimization for Online Convex Optimization

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.354327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.354327Z digest=sha256:1cabab872f719c718a15406537cae803997dbb88ea514a33aa39c946847f783b

Observation a279fb87-f15b-4112-b7d2-7bc0a0e96f4e · outbound

This paper cites Diagonal Rescaling For Neural Networks.

Linear Convergence of Adaptive Stochastic Gradient Descent Diagonal Rescaling For Neural Networks

Reference 2014

Resolution
verified exact
local_arxiv, observed 2026-08-14T10:47:31.631995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T10:47:31.364853Z digest=sha256:e1f8e6f432f6731492276f219be90945db8e7ed1e9524110e6509cb7e9047a78

Observation e54fa5fe-e7ce-4bf7-b501-f56eb3228447 · outbound

This paper cites A Convergence Theory for Deep Learning via Over-Parameterization.

Linear Convergence of Adaptive Stochastic Gradient Descent A Convergence Theory for Deep Learning via Over-Parameterization

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.343819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.343819Z digest=sha256:745ffc36c13748c07ace9ab209c81ba2eaf17ae7fc29872c0aad449944a7327a

Observation 9c02d46b-dbf4-4e69-9580-1d8f4f47814f · outbound

This paper cites On the Convergence of Adam and Beyond.

Linear Convergence of Adaptive Stochastic Gradient Descent On the Convergence of Adam and Beyond

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.370129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.370129Z digest=sha256:126de45fcba23d90f933ea862b91ad101f4c4b957e9452800620fc62a93cd244

Observation 2c219bbd-55e3-4a38-bbd7-6e3ab40b0eda · outbound

This paper cites Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks.

Linear Convergence of Adaptive Stochastic Gradient Descent Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.349191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.349191Z digest=sha256:96bcead1f638180ae90a6659646378937f807044be94aa28152af68fa36950d7

Observation e882a8e9-4bb5-4bd9-b6c9-6a8e06f19d78 · outbound

This paper cites On exponential convergence of SGD in non-convex over-parametrized learning.

Linear Convergence of Adaptive Stochastic Gradient Descent On exponential convergence of SGD in non-convex over-parametrized learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.406356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.406356Z digest=sha256:aa6854a5d1327b809cf0546906d6c52df0af84d88ff3c5e827d432978d55bd71

Pith citing papers

No inbound Pith citation observations are available.