Pith. sign in

Paper Citation Record · LEDGER

Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:1611.01838.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1611.01838 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:12:34.202549Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

115
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 696ae77a-636b-4c35-b96e-d146e6b23050 · inbound

On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima cites this paper.

On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:58:48.536037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:58:48.515002Z digest=sha256:f8930e2079114f53a9157377b4e76714528f0692435bb9f7bba55d986e197d84

Observation 4263f2cd-d578-43c0-9b34-4712a416ed17 · inbound

First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise cites this paper.

First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-25T18:56:08.858446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T18:55:33.003637Z digest=sha256:b9b653af6c6246f122971bdbaa4dc300e63c9d9d4841e2006e85dd89bbef1c7e

Observation b4861c65-5ff2-4fab-824f-f8895f6bac89 · inbound

A Stochastic Composite Gradient Method with Incremental Variance Reduction cites this paper.

A Stochastic Composite Gradient Method with Incremental Variance Reduction Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-25T17:06:07.855790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T17:02:34.602404Z digest=sha256:2b481e98b54701ac343267762b9cca522ff22aaa743bb165609a3a7ce7eb8c48

Observation 8b8b53d5-dfc7-4d7f-ac91-da82991aff26 · inbound

Gradient Noise Convolution (GNC): Smoothing Loss Function for Distributed Large-Batch SGD cites this paper.

Gradient Noise Convolution (GNC): Smoothing Loss Function for Distributed Large-Batch SGD Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T15:45:59.706527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T15:41:27.232807Z digest=sha256:928cb371f5caa8fa512054e625bb5db942db13e37bb641f7816b74b533262953

Observation a396cc71-c528-4659-ae75-c781d35beeb7 · inbound

Chaining Meets Chain Rule: Multilevel Entropic Regularization and Training of Neural Nets cites this paper.

Chaining Meets Chain Rule: Multilevel Entropic Regularization and Training of Neural Nets Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-25T15:45:59.309563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T15:45:38.743857Z digest=sha256:a27d3230f5c6ff75105d2c49280752fa50d60e44e9e96314b8c5051439715eee

Observation 5350fbc6-a279-4d89-8763-c266f739855b · inbound

Heavy-ball Algorithms Always Escape Saddle Points cites this paper.

Heavy-ball Algorithms Always Escape Saddle Points Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-24T17:59:46.450026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T17:59:27.851487Z digest=sha256:cb973ebbaf5f49a742c217e3ad0788e95e326aaac66a9a693dc60300ba6f8fbe

Observation fc57bbef-a897-4ccc-8d46-c5aabb156747 · inbound

Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization cites this paper.

Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T16:44:41.707991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T16:42:26.048100Z digest=sha256:5cfa6418983fea23de7bc786e24b8d1758c3075763666af1f1e1242f7ddac762

Observation 64bf8a70-56d8-42ff-a46e-5be0c598ff66 · inbound

Sharpness-Aware Minimization for Efficiently Improving Generalization cites this paper.

Sharpness-Aware Minimization for Efficiently Improving Generalization Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:13:53.441561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T20:13:53.366151Z digest=sha256:e086e02e01ef8cb131fac33941044a16f47f7520cdf35484e56a7e064b3853e1

Observation f7c5d364-69b5-4966-a16d-61c2f6e9cb2d · inbound

LSAM: Asynchronous Distributed Training with Landscape-Smoothed Sharpness-Aware Minimization cites this paper.

LSAM: Asynchronous Distributed Training with Landscape-Smoothed Sharpness-Aware Minimization Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:12:34.202549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:12:34.202549Z digest=sha256:2428fb2f5f9bdb03527443f39c4937bef330bae94786ebf15ab7659eb8c09b07

Observation 0b3dbd2e-2c4a-401d-a603-d0b7ac4b63fe · inbound

Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting cites this paper.

Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:42.376885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T18:33:26.638240Z digest=sha256:2128ae91f62ba9bdd133919c24e598d562118b6305325916fa3077e7c7ce56fd

Observation c6e11566-7fed-497c-bb22-3889b5267912 · inbound

Estimating Implicit Regularization in Deep Learning cites this paper.

Estimating Implicit Regularization in Deep Learning Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:26:13.034669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T15:54:21.139902Z digest=sha256:df01e39a2760e7770ca927755e35372664801e441d0bed7962066f10e0498aba

Observation bac4abc2-6be2-41ef-99a8-ab43288d6c9e · inbound

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing cites this paper.

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:18:59.541406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T20:15:44.030714Z digest=sha256:a356f388b54121f268efa311ce2918627e73ad34d4fe006a7334791936ff0384