Pith. sign in

Paper Citation Record · LEDGER

Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:1905.11286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1905.11286 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T22:10:49.683444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:16:22.222379Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a75aa2dc-c08f-49eb-9a41-737f2961125e · inbound

CTRL: A Conditional Transformer Language Model for Controllable Generation cites this paper.

CTRL: A Conditional Transformer Language Model for Controllable Generation Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:14:02.519341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:14:02.423030Z digest=sha256:ffc6205db08da305b12764eafa70436c08be4305e87df7015e4123b24ea6e63c

Observation bf703a2d-a6a5-4a75-ae44-fe4b49af4696 · inbound

Training Deep Learning Models with Norm-Constrained LMOs cites this paper.

Training Deep Learning Models with Norm-Constrained LMOs Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:22:37.188171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T21:22:36.870292Z digest=sha256:ed452af94e19dae74300eddbdb174f4356a6e9bd37f50eef0d935800da707172

Observation a31b7371-6603-4e10-aa48-60e3ed20d88c · inbound

A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad, Shampoo and Muo cites this paper.

A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad, Shampoo and Muo Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:26:59.851439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:19:59.335184Z digest=sha256:1b97865ad2e95ae233637165e73326ff220789fda6d840323f7ea78dba7d4d55

Observation 47681e5b-d854-4310-a243-3dfd4e0f68d8 · inbound

Rethinking Neural Network Learning Rates: A Stackelberg Perspective cites this paper.

Rethinking Neural Network Learning Rates: A Stackelberg Perspective Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:43:06.700212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T14:42:45.648114Z digest=sha256:c23ae80b39be08624e7df261e957bd5f0660bad6450acaf399f43458875dec14

Observation c06dbc12-ad35-4ca9-abdb-0ac8f30f8d81 · inbound

Rethinking Neural Network Learning Rates: A Stackelberg Perspective cites this paper.

Rethinking Neural Network Learning Rates: A Stackelberg Perspective Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:55:01.077020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T19:53:45.109784Z digest=sha256:d87b6c308a6d2a9ee8967d07e5b8950451d58262d31e5936b16a5f23f254a5c4

Observation cb54988f-78b3-45ff-95cc-11e8760ae3fa · inbound

Stochastic convergence of parallel asynchronous adaptive first-order methods cites this paper.

Stochastic convergence of parallel asynchronous adaptive first-order methods Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:22.224581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:27:48.433313Z digest=sha256:4b91ccb1b6fd2b9e6171f135c0144a22ea9a2f6a7739086f4dadafc65f9ab3ce

Observation b35f5a33-13e2-4dc9-bb9c-be8aef451204 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:b3b730f92a6c385610887266b0e6ad432d41a9ac39c9859780fe19cd668792ff