Pith. sign in

REVIEW 9 cited by

A Convergence Analysis of Gradient Descent for Deep Linear Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.02281 v3 pith:VPBOHLDU submitted 2018-10-04 cs.LG cs.NEstat.ML

classification cs.LGcs.NEstat.ML
keywords convergencelineardeepinitializationlossdescentdimensionsglobal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We analyze speed of convergence to global optimum for gradient descent training a deep linear neural network (parameterized as $x \mapsto W_N W_{N-1} \cdots W_1 x$) by minimizing the $\ell_2$ loss over whitened data. Convergence at a linear rate is guaranteed when the following hold: (i) dimensions of hidden layers are at least the minimum of the input and output dimensions; (ii) weight matrices at initialization are approximately balanced; and (iii) the initial loss is smaller than the loss of any rank-deficient solution. The assumptions on initialization (conditions (ii) and (iii)) are necessary, in the sense that violating any one of them may lead to convergence failure. Moreover, in the important case of output dimension 1, i.e. scalar regression, they are met, and thus convergence to global optimum holds, with constant probability under a random initialization scheme. Our results significantly extend previous analyses, e.g., of deep linear residual networks (Bartlett et al., 2018).

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. How are linear representations learned? Exact solutions to the dynamics of abstraction

    cs.LG 2026-07 conditional novelty 8.0 of 10

    Exact solutions show abstraction is set by input/target geometry, rises with depth, peaks under small init, and is attenuated by nonlinearities—improving LLM probes via GELU ablation.

  2. Differentiable Approximations for Distance Queries

    cs.CG 2026-07 conditional novelty 7.0 of 10

    A (1+ε)-approximate Euclidean distance function that is differentiable and returns gradients, using O(n/ε^(d/2)) space and O(log(n/ε)) query time.

  3. Monotonic Kolmogorov-Arnold Networks: A Theoretical and Empirical Study of Monotonicity as an Inductive Bias

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    MKAN adds hard unconstrained monotonicity to KANs via reparameterization and proves a size bound of at most 2N* for monotone equivalents of ball-partition feature extractors.

  4. The Implicit Bias of Depth: From Neural Collapse to Softmax Codes

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Depth induces an implicit low-rank bias in deep unconstrained feature models trained with unregularized multiclass cross-entropy, promoting softmax codes over neural collapse via more efficient norm propagation.

  5. EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models

    cs.AI 2026-04 unverdicted novelty 7.0 of 10

    EmergentBridge improves zero-shot cross-modal transfer for unpaired modality pairs by learning noisy bridge anchors and enforcing proxy alignment only in the orthogonal subspace to preserve existing anchor alignments.

  6. EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    EmergentBridge enhances zero-shot cross-modal performance on unpaired modalities by learning noisy bridge anchors from existing alignments and enforcing proxy alignment only in the orthogonal subspace to avoid gradien...

  7. Conservation Laws for Modern Neural Architectures

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Unified framework characterizes conservation laws for gradient flow in feedforward networks with GELU/SiLU/SwiGLU, multihead attention with positional encodings, and MoE models under various gating.

  8. Geodesics in the Deep Linear Network

    math.DG 2025-09 unverdicted novelty 5.0 of 10

    Derives ODEs and explicit solutions for geodesics between full-rank matrices in deep linear network geometry and shows that certain horizontal straight lines in the invariant balanced manifold remain geodesics under R...

  9. Intrinsic Strain-Driven Topological Evolution in SrRuO3 via Flexural Strain Engineering

    cond-mat.mtrl-sci 2025-08 unverdicted novelty 5.0 of 10

    The abstract reports a 21% anomalous Hall conductivity increase in flexurally strained SrRuO3, but the submitted full text belongs to a different machine learning paper.

Pith tools