Pith. sign in

REVIEW 3 cited by

The Difficulty of Training Sparse Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.10732 v3 pith:Z7P6IF6L submitted 2019-06-25 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords sparsepathsolutionssubspacedecreasingfindfoundgood
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We investigate the difficulties of training sparse neural networks and make new observations about optimization dynamics and the energy landscape within the sparse regime. Recent work of \citep{Gale2019, Liu2018} has shown that sparse ResNet-50 architectures trained on ImageNet-2012 dataset converge to solutions that are significantly worse than those found by pruning. We show that, despite the failure of optimizers, there is a linear path with a monotonically decreasing objective from the initialization to the "good" solution. Additionally, our attempts to find a decreasing objective path from "bad" solutions to the "good" ones in the sparse subspace fail. However, if we allow the path to traverse the dense subspace, then we consistently find a path between two solutions. These findings suggest traversing extra dimensions may be needed to escape stationary points found in the sparse subspace.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Preserving Deep Representations In One-Shot Pruning: A Hessian-Free Second-Order Optimization Framework

    cs.LG 2024-11 conditional novelty 6.0 of 10

    SNOWS prunes vision networks in one shot by optimizing a K-step nonlinear reconstruction objective with Hessian-free Newton steps, improving accuracy over layer-wise least-squares methods.

  2. On the rate of convergence of fully connected very deep neural network regression estimates

    stat.ML 2019-08 accept novelty 6.0 of 10

    Least squares estimators based on fully connected ReLU networks achieve dimension-free rates for hierarchical composition models, with either logarithmic depth and growing width or fixed width and very large depth.

  3. Implicit Deep Learning

    cs.LG 2019-08 conditional novelty 6.0 of 10

    Implicit deep learning replaces layered neural networks with a single fixed-point equation, enabling well-posedness conditions, robustness bounds, and new training algorithms.

Pith tools