Pith. sign in

Paper Citation Record · LEDGER

A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2002.03495.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2002.03495 v14

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:38:49.975072Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:47:10.361461Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b3196e76-1da8-4d26-a43d-1c8e6edc023b · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:49.975072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:49.975072Z digest=sha256:3475e1d0276ee5c176c61a472bcbc2e6b0a831cac16f40eb3439f04185df8d85

Observation a73c2707-3f8c-4199-a448-fc8893299004 · inbound

Dimension-Free Saddle-Point Escape in Muon cites this paper.

Dimension-Free Saddle-Point Escape in Muon A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:26.388383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:38:25.480020Z digest=sha256:a954fc74c8cc8c26f91e0be50a55450f2d26582206c28fc2bec046c0ea678b1c

Observation 16e5f64e-e377-4fe1-9fde-d22be7d914d5 · inbound

Thermodynamic Irreversibility of Training Algorithms cites this paper.

Thermodynamic Irreversibility of Training Algorithms A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:41:04.270581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T04:38:51.773377Z digest=sha256:cfb075196165edb04d9ac16c5049a5fbef7f018f701a3527f780c008faf11f0b

Observation 01460dae-6bd1-42a3-be76-26a25132f30a · inbound

Second-Order Path Kernel Interpolation Formulas in Machine Learning cites this paper.

Second-Order Path Kernel Interpolation Formulas in Machine Learning A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:47:10.362665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T22:20:15.074286Z digest=sha256:4284a6557848676fe2498141d2fba5bcd347bdb73081efe1930e441cec3a6263

Observation c45f49d2-bdb6-4bdd-a42e-b6bafb9d8c7e · inbound

SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates cites this paper.

SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 199

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:05:45.536954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T01:05:17.842447Z digest=sha256:1d37508e61552ba48c4ad7b1a57b1d651b6b710a6b3da18c375df2bd11e110d7