Pith. sign in

REVIEW 4 cited by

Critical Points of Neural Networks: Analytical Forms and Landscape Properties

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1710.11205 v1 pith:AHXJIUB5 submitted 2017-10-30 stat.ML cs.LG

classification stat.MLcs.LG
keywords lossnetworkscriticalpointsanalyticalformsfunctionsminimum
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Due to the success of deep learning to solving a variety of challenging machine learning tasks, there is a rising interest in understanding loss functions for training neural networks from a theoretical aspect. Particularly, the properties of critical points and the landscape around them are of importance to determine the convergence performance of optimization algorithms. In this paper, we provide full (necessary and sufficient) characterization of the analytical forms for the critical points (as well as global minimizers) of the square loss functions for various neural networks. We show that the analytical forms of the critical points characterize the values of the corresponding loss functions as well as the necessary and sufficient conditions to achieve global minimum. Furthermore, we exploit the analytical forms of the critical points to characterize the landscape properties for the loss functions of these neural networks. One particular conclusion is that: The loss function of linear networks has no spurious local minimum, while the loss function of one-hidden-layer nonlinear networks with ReLU activation function does have local minimum that is not global minimum.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Impact of Bottleneck Layers and Skip Connections on the Generalization of Linear Denoising Autoencoders

    stat.ML 2025-05 conditional novelty 7.0 of 10

    Two-layer linear denoising autoencoders show a bias-variance trade-off in bottleneck width, and skip connections reduce variance near the interpolation peak.

  2. A Theory on Flow Matching with Neural Networks

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Establishes convergence guarantees for overparameterized 2-layer ReLU networks in flow matching, generalization bounds for the velocity-field objective, and Wasserstein guarantees for generated samples, using multi-ta...

  3. Critical Organization of Deep Neural Networks, and p-Adic Statistical Field Theories

    cs.LG 2026-01 reject novelty 6.0 of 10

    A p-adic integral-equation formulation of deep networks is shown to have a unique hidden state under a contraction condition; the claimed thermodynamic limit and infinite-state bifurcation are not proven.

  4. Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

    cs.LG 2024-01 unverdicted novelty 6.0 of 10

    SPIN lets weak LLMs become strong by self-generating training data from previous model versions and training to prefer human-annotated responses over its own outputs, outperforming DPO even with extra GPT-4 data on be...

Pith tools