Pith. sign in

REVIEW 1 cited by

Opening the Black Box: predicting the trainability of deep neural networks with reconstruction entropy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12916 v3 pith:ZOW36CVE submitted 2024-06-13 cs.LG cond-mat.dis-nnhep-thstat.ML

classification cs.LGcond-mat.dis-nnhep-thstat.ML
keywords networksneuralnetworkdeepmethodtrainingcascadedemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An important challenge in machine learning is to predict the initial conditions under which a given neural network will be trainable. We present a method for predicting the trainable regime in parameter space for deep feedforward neural networks (DNNs) based on reconstructing the input from subsequent activation layers via a cascade of single-layer auxiliary networks. We show that a single epoch of training of the shallow cascade networks is sufficient to predict the trainability of the deep feedforward network on a range of datasets (MNIST, CIFAR10, FashionMNIST, and white noise), thereby providing a significant reduction in overall training time. We achieve this by computing the relative entropy between reconstructed images and the original inputs, and show that this probe of information loss is sensitive to the phase behaviour of the network. We further demonstrate that this method generalizes to residual neural networks (ResNets) and convolutional neural networks (CNNs). Moreover, our method illustrates the network's decision making process by displaying the changes performed on the input data at each layer, which we demonstrate for both a DNN trained on MNIST and the vgg16 CNN trained on the ImageNet dataset. Our results provide a technique for significantly accelerating the training of large neural networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Spin Glass Characterization of Neural Networks

    cond-mat.dis-nn 2025-08 unverdicted novelty 6.0 of 10

    A Hopfield-type spin glass constructed from a feedforward network yields replica overlap statistics that serve as a per-instance descriptor of the network's generalization, capacity, and robustness.

Pith tools