Pith. sign in

REVIEW 3 cited by

Loss landscapes and optimization in over-parameterized non-linear systems and neural networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.00307 v2 pith:4YKGQG43 submitted 2020-02-29 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords conditionoptimizationover-parameterizedsystemsnetworksneuralnon-lineardeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The success of deep learning is due, to a large extent, to the remarkable effectiveness of gradient-based optimization methods applied to large neural networks. The purpose of this work is to propose a modern view and a general mathematical framework for loss landscapes and efficient optimization in over-parameterized machine learning models and systems of non-linear equations, a setting that includes over-parameterized deep neural networks. Our starting observation is that optimization problems corresponding to such systems are generally not convex, even locally. We argue that instead they satisfy PL$^*$, a variant of the Polyak-Lojasiewicz condition on most (but not all) of the parameter space, which guarantees both the existence of solutions and efficient optimization by (stochastic) gradient descent (SGD/GD). The PL$^*$ condition of these systems is closely related to the condition number of the tangent kernel associated to a non-linear system showing how a PL$^*$-based non-linear theory parallels classical analyses of over-parameterized linear equations. We show that wide neural networks satisfy the PL$^*$ condition, which explains the (S)GD convergence to a global minimum. Finally we propose a relaxation of the PL$^*$ condition applicable to "almost" over-parameterized systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predator-Prey Model: Driven Hunt for Accelerated Grokking

    cs.NE 2025-09 conditional novelty 4.0 of 10

    A predator-prey two-agent optimizer accelerates the post-memorization grokking phase by tens to a hundred times in gradient calls on modular arithmetic and MNIST, but still requires standard pre-training to memorization.

  2. From Sublinear to Linear: Local Convergence in Finite-Width Networks via Locally Polyak-Lojasiewicz Regions

    stat.ML 2025-07 conditional novelty 4.0 of 10

    Local NTK positivity plus Lipschitz stability gives a local Polyak-Lojasiewicz constant lambda0 minus L_Theta times the region radius, yielding linear gradient descent convergence whenever the iterates stay in the LQC...

  3. Optimal and Diffusion Transports in Machine Learning

    math.OC 2025-12 accept novelty 1.0 of 10

    A survey showing how optimal transport, diffusion models, and transformer dynamics all fit into a common framework of time-evolving probability measures.

Pith tools