Pith. sign in

REVIEW 12 cited by

Gradient Descent Maximizes the Margin of Homogeneous Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.05890 v4 pith:JK22FA35 submitted 2019-06-13 cs.LG cs.NEstat.ML

classification cs.LGcs.NEstat.ML
keywords gradientmarginnetworksdescenthomogeneousneuralresultsloss
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we study the implicit regularization of the gradient descent algorithm in homogeneous neural networks, including fully-connected and convolutional neural networks with ReLU or LeakyReLU activations. In particular, we study the gradient descent or gradient flow (i.e., gradient descent with infinitesimal step size) optimizing the logistic loss or cross-entropy loss of any homogeneous model (possibly non-smooth), and show that if the training loss decreases below a certain threshold, then we can define a smoothed version of the normalized margin which increases over time. We also formulate a natural constrained optimization problem related to margin maximization, and prove that both the normalized margin and its smoothed version converge to the objective value at a KKT point of the optimization problem. Our results generalize the previous results for logistic regression with one-layer or multi-layer linear networks, and provide more quantitative convergence results with weaker assumptions than previous results for homogeneous smooth neural networks. We conduct several experiments to justify our theoretical finding on MNIST and CIFAR-10 datasets. Finally, as margin is closely related to robustness, we discuss potential benefits of training longer for improving the robustness of the model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 57 citations worldwide. Full citation record

  1. Multiclass Loss Geometry Matters for Generalization of Gradient Descent in Separable Classification

    cs.LG 2025-05 accept novelty 8.0 of 10

    In separable multiclass classification, the risk of gradient descent scales as k^{2/p} for losses with ℓ_p-smooth templates, giving logarithmic k-dependence for p=∞ and linear k-dependence for p=2 (provably unavoidable).

  2. Querying Kernel Methods Suffices for Reconstructing their Training Data

    cs.LG 2025-05 conditional novelty 8.0 of 10

    Query-only access to kernel regression, SVM and KDE models suffices to reconstruct their exact training points, via a measure-theoretic proof and image experiments.

  3. A Counterexample to Fourier Alignment in Single-Neuron Modular Addition

    math.OC 2026-08 accept novelty 7.0 of 10

    For every prime p at least 5, a single ReLU neuron can, with positive Gaussian probability, train to a frozen or diverging state whose normalized Fourier spectrum flatly distributes energy, contradicting single-freque...

  4. A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning

    cs.LG 2026-02 conditional novelty 6.0 of 10

    In diagonal linear networks, fine-tuning generalization is governed by a tunable per-dimension penalty whose sparsity and pretraining dependence define four regimes and a trade-off between feature reuse and new-featur...

  5. Muon in Associative Memory Learning: Training Dynamics and Scaling Laws

    cs.LG 2026-02 conditional novelty 6.0 of 10

    In a linear softmax memory model, Muon equalizes learning across frequency tiers and gives exponential (noiseless) or T^{-2} (noisy power-law) convergence, versus polynomial or T^{-(1-1/β)} for gradient descent.

  6. Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Identity-bridge regularization, rephrased into an out-of-context reasoning form, yields ~40% reversal accuracy in a 1B LLM and provably fixes reversal in an idealized one-layer transformer.

  7. Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Adding identity supervision on bridge tokens enables out-of-distribution two-hop reasoning in simple transformers, with a nuclear-norm theory explaining the benefit.

  8. PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning

    math.OC 2025-05 conditional novelty 6.0 of 10

    PADAM runs K differently averaged Adam trajectories in parallel, selects the one with the smallest test error, and achieves the best optimization error in nearly all of 13 tested scientific machine learning problems w...

  9. Understanding Nonlinear Implicit Bias via Region Counts in Input Space

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Region count, the number of connected same-label regions along random input-space lines, correlates strongly with the generalization gap and is proposed as a reparameterization-invariant measure of nonlinear implicit bias.

  10. Neural Thermodynamic Laws for Large Language Model Training

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Under a river-valley model of the loss landscape, the paper shows that valley fluctuations behave like heat, with learning rate as temperature, and derives a 1/t optimal decay schedule.

  11. The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations

    cs.LG 2025-07 conditional novelty 5.0 of 10

    FACT is a first-order stationarity identity for weight matrices that matches or beats the Neural Feature Ansatz as a description of learned features at convergence.

  12. Theoretical Issues in Deep Networks: Approximation, Optimization and Generalization

    cs.LG 2019-08 conditional novelty 3.0 of 10

    A synthesis of approximation, optimization, and generalization theory arguing that gradient descent's implicit norm control on weight directions explains why overparameterized deep networks generalize.

Pith tools