Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:26:43.530125Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2501.09137.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:26:43.530125Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T13:20:54.303605Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T13:23:28.017867Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a3c0857e-9576-4883-b16a-138151270d33 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f48a3d-e7f8-4ddc-a62f-e1616fe57885 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks T., Suarez, F., and Zhang, Y
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5b2b8e0-c4fb-4cb4-b295-8567443c7197 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks A Convergence Analysis of Gradient Descent for Deep Linear Neural Networks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da037326-655f-489c-9d80-f24bc6feaf81 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks P., Selman, B., and Weinberger, K
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e59cfc15-b3c9-487d-88e6-bf090f02520f · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Optimization Methods for Large-Scale Machine Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d9d63a-5642-4d59-ba11-4f91cf4fb02d · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks and Bruna, J
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5ec7451e-75d1-4dd3-9dad-51763c0f20e3 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Z., and Talwalkar, A
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 98fce254-a364-4024-8a8d-cf374f5f1177 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Implicit Regularization of Discrete Gradient Dynamics in Linear Neural Networks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a66da8c3-f9f6-4ce4-8668-f04cfd0b3d25 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks and Schmidhuber, J
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 836d0d80-4c5a-441e-ba53-19e71e36882a · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks The Break-Even Point on Optimization Trajectories of Deep Neural Networks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb4d54e9-6740-4744-9a80-a8b7f500874f · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation be8bfd10-f67b-4c79-a432-26eb978412e4 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks B., and Müller, K.-R
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f64b7da-712d-495f-8161-69a21ad71df1 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks The large learning rate phase of deep learning: the catapult mechanism
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d8049ef-d133-44ef-920f-a71bdfe7ac22 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Towards explaining the regularization effect of initial large learning rate in training neural networks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4181a5dd-1f4d-4c31-a8c4-174f953c33b4 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks M., Rauhut, H., and Terstiege, U
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2a8ab38-6d47-4557-8a67-bc6e331773df · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks The Effect of Network Width on Stochastic Gradient Descent and Generalization : an Empirical Study
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca90ba2d-03ac-4964-967f-6c2b1c79ed06 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2dd077db-079a-4f7d-bebe-0efc2e47a981 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fef5958b-0a7e-4721-b896-35d878e34995 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks A Bayesian Perspective on Generalization and Stochastic Gradient Descent
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9308a71-ecfb-4aa5-bc96-0aba767f8dbf · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks D., and Vidal, R
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a26044fc-ce5c-4f30-bf6e-eeab91f6b694 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Large Learning Rate Tames Homogeneity : Convergence and Balancing Effect
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation edd11cbf-48dc-408a-9fe7-0f43af6fa797 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Three Mechanisms of Feature Learning in a Linear Network
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0952d60f-06b8-453c-b4c8-c883720e3fd0 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Linear convergence of gradient descent for finite width over-parametrized linear networks with general initialization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7aa214f3-1596-4b92-a4ec-3480c1750796 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks @esa (Ref
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e389e42-4f56-4eaa-bd43-7f8f4921c8b6 · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b494e2f5-df1c-419d-8b51-6f59a2a6230a · outbound
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks ,# (7),01444 '9=82<.342C 2! !22222222222222222222222222222222222222222222222222
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f1271a0b-d08e-4163-8536-08678091a4bd · inbound
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.