Pith. sign in

REVIEW 1 cited by

Multi-layer Perceptron Trainability Explained via Variability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.08911 v3 pith:HHO5H5IV submitted 2021-05-19 cs.LG

classification cs.LG
keywords variabilitydeeptrainabilityfunctioncalledcorrelatedlearningmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the tremendous successes of deep neural networks (DNNs) in various applications, many fundamental aspects of deep learning remain incompletely understood, including DNN trainability. In a trainability study, one aims to discern what makes one DNN model easier to train than another under comparable conditions. In particular, our study focuses on multi-layer perceptron (MLP) models equipped with the same number of parameters. We introduce a new notion called variability to help explain the benefits of deep learning and the difficulties in training very deep MLPs. Simply put, variability of a neural network represents the richness of landscape patterns in the data space with respect to well-scaled random weights. We empirically show that variability is positively correlated to the number of activations and negatively correlated to a phenomenon called "Collapse to Constant", which is related but not identical to the well-known vanishing gradient phenomenon. Experiments on a small stylized model problem confirm that variability can indeed accurately predict MLP trainability. In addition, we demonstrate that, as an activation function in MLP models, the absolute value function can offer better variability than the popular ReLU function can.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transforming Chatbot Text: A Sequence-to-Sequence Approach

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Fine-tuned T5-small and BART can lower the accuracy of GPT-text classifiers, but retraining the classifiers on the transformed text restores detection accuracy.

Pith tools