Pith. sign in

REVIEW 5 cited by

Deep Learning without Poor Local Minima

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1605.07110 v3 pith:SFEWRGXR submitted 2016-05-23 stat.ML cs.LGmath.OC

classification stat.MLcs.LGmath.OC
keywords deeplearningnetworkssaddledifficultlocalminimumpoint
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we prove a conjecture published in 1989 and also partially address an open problem announced at the Conference on Learning Theory (COLT) 2015. With no unrealistic assumption, we first prove the following statements for the squared loss function of deep linear neural networks with any depth and any widths: 1) the function is non-convex and non-concave, 2) every local minimum is a global minimum, 3) every critical point that is not a global minimum is a saddle point, and 4) there exist "bad" saddle points (where the Hessian has no negative eigenvalue) for the deeper networks (with more than three layers), whereas there is no bad saddle point for the shallow networks (with three layers). Moreover, for deep nonlinear neural networks, we prove the same four statements via a reduction to a deep linear model under the independence assumption adopted from recent work. As a result, we present an instance, for which we can answer the following question: how difficult is it to directly train a deep model in theory? It is more difficult than the classical machine learning models (because of the non-convexity), but not too difficult (because of the nonexistence of poor local minima). Furthermore, the mathematically proven existence of bad saddle points for deeper models would suggest a possible open problem. We note that even though we have advanced the theoretical foundations of deep learning and non-convex optimization, there is still a gap between theory and practice.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Absence of poor local minima in matrix product states

    quant-ph 2026-06 unverdicted novelty 7.0 of 10

    MPS energy landscapes lack poor local minima because gauge freedom induces overparametrization that concentrates local minima near the global minimum, with the local minimum distribution proven invariant under orthogo...

  2. Impact of Bottleneck Layers and Skip Connections on the Generalization of Linear Denoising Autoencoders

    stat.ML 2025-05 conditional novelty 7.0 of 10

    Two-layer linear denoising autoencoders show a bias-variance trade-off in bottleneck width, and skip connections reduce variance near the interpolation peak.

  3. Landscape analysis for shallow neural networks: Complete classification of critical points for cubic activation and affine target functions

    math.OC 2026-07 conditional novelty 6.0 of 10

    For cubic-activation shallow networks with affine targets, the squared-loss landscape has no local maxima; every critical point is a global minimizer, a rigid non-global local minimum, or a saddle, and zero loss is ac...

  4. Deep learning inference with the Event Horizon Telescope II. The Zingularity framework for Bayesian artificial neural networks

    astro-ph.IM 2025-06 conditional novelty 5.0 of 10

    Bayesian neural networks trained on synthetic EHT observations of Sgr A* and M87* recover spin and magnetic state well in cross-code tests, but give overconfident wrong estimates for temperature ratio and inclination ...

  5. Statistical Properties of Training & Generalization

    stat.ML 2026-06 unverdicted novelty 2.0 of 10

    Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.

Pith tools