Pith. sign in

REVIEW 11 cited by

Gradient Descent Maximizes the Margin of Homogeneous Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.05890 v4 pith:JK22FA35 submitted 2019-06-13 cs.LG cs.NEstat.ML

Gradient Descent Maximizes the Margin of Homogeneous Neural Networks

classification cs.LG cs.NEstat.ML
keywords gradientmarginnetworksdescenthomogeneousneuralresultsloss
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In this paper, we study the implicit regularization of the gradient descent algorithm in homogeneous neural networks, including fully-connected and convolutional neural networks with ReLU or LeakyReLU activations. In particular, we study the gradient descent or gradient flow (i.e., gradient descent with infinitesimal step size) optimizing the logistic loss or cross-entropy loss of any homogeneous model (possibly non-smooth), and show that if the training loss decreases below a certain threshold, then we can define a smoothed version of the normalized margin which increases over time. We also formulate a natural constrained optimization problem related to margin maximization, and prove that both the normalized margin and its smoothed version converge to the objective value at a KKT point of the optimization problem. Our results generalize the previous results for logistic regression with one-layer or multi-layer linear networks, and provide more quantitative convergence results with weaker assumptions than previous results for homogeneous smooth neural networks. We conduct several experiments to justify our theoretical finding on MNIST and CIFAR-10 datasets. Finally, as margin is closely related to robustness, we discuss potential benefits of training longer for improving the robustness of the model.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. What Does the Weight Norm Control in Grokking? Logit-Scale Mediation under Cross-Entropy

    cs.LG 2026-06 conditional novelty 7.0

    Grokking delay under cross-entropy is mediated primarily by logit scale and resulting softmax saturation, with weight norm acting only as an upstream handle that adds 1-2% beyond the scale.

  2. The Implicit Bias of Depth: From Neural Collapse to Softmax Codes

    cs.LG 2026-05 unverdicted novelty 7.0

    Depth induces an implicit low-rank bias in deep unconstrained feature models trained with unregularized multiclass cross-entropy, promoting softmax codes over neural collapse via more efficient norm propagation.

  3. Efficient Techniques for Data Reconstruction, with Finite-Width Recovery Guarantees

    cs.LG 2026-05 unverdicted novelty 7.0

    A unified data reconstruction attack achieves provable finite-width recovery in random feature networks and efficient subspace-based reconstruction for general models using weight changes.

  4. Implicit Bias in Deep Linear Discriminant Analysis

    cs.LG 2026-03 unverdicted novelty 7.0

    Gradient flow on deep diagonal linear LDA networks with balanced initialization converts additive updates to multiplicative updates, automatically conserving the (2/L) quasi-norm.

  5. Convergence of Continual Learning in Homogeneous Deep Networks

    cs.LG 2026-06 unverdicted novelty 6.0

    Continual classification in homogeneous models is sequential projections onto margin sets, with local linear convergence under regularity properties for random and cyclic tasks, extended to regression.

  6. A Theory on Flow Matching with Neural Networks

    cs.LG 2026-06 unverdicted novelty 6.0

    Establishes convergence guarantees for overparameterized 2-layer ReLU networks in flow matching, generalization bounds for the velocity-field objective, and Wasserstein guarantees for generated samples, using multi-ta...

  7. The Neural Tangent Kernel for Classification

    cs.LG 2026-05 unverdicted novelty 6.0

    Wide neural networks with cross-entropy loss maintain constant NTK under parameter regularization or non-degenerate targets, enabling linearized approximation and explicit NTK-based solution characterization.

  8. The Neural Tangent Kernel for Classification

    cs.LG 2026-05 unverdicted novelty 6.0

    Wide neural networks with cross-entropy loss remain in the lazy training regime under parameter-space regularization or non-degenerate targets, allowing explicit NTK-based solution characterization and uncertainty analysis.

  9. The Effect of Mini-Batch Noise on the Implicit Bias of Adam

    cs.LG 2026-02 unverdicted novelty 6.0

    Mini-batch noise reverses how Adam's β2 controls anti-regularization, making default momentum values suitable for small batches but requiring β1 closer to β2 for large batches to favor flatter minima.

  10. Prediction horizon shapes representations in predictive learning

    cs.LG 2025-11 unverdicted novelty 6.0

    Longer prediction horizons in predictive learning interact with model biases to recover the latent geometry of the task.

  11. Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

    cs.LG 2024-01 unverdicted novelty 6.0

    SPIN lets weak LLMs become strong by self-generating training data from previous model versions and training to prefer human-annotated responses over its own outputs, outperforming DPO even with extra GPT-4 data on be...