Pith. sign in

REVIEW 3 cited by

The boundary of neural network trainability is fractal

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.06184 v1 pith:5VE5RS2K submitted 2024-02-09 cs.LG cs.NEnlin.CD

classification cs.LGcs.NEnlin.CD
keywords boundaryhyperparametersnetworkneuraldivergentfractalfunctioniterating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Some fractals -- for instance those associated with the Mandelbrot and quadratic Julia sets -- are computed by iterating a function, and identifying the boundary between hyperparameters for which the resulting series diverges or remains bounded. Neural network training similarly involves iterating an update function (e.g. repeated steps of gradient descent), can result in convergent or divergent behavior, and can be extremely sensitive to small changes in hyperparameters. Motivated by these similarities, we experimentally examine the boundary between neural network hyperparameters that lead to stable and divergent training. We find that this boundary is fractal over more than ten decades of scale in all tested configurations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse

    cs.LG 2026-07 accept novelty 7.0 of 10

    Noisy SGD in the mean-field regime forces wide multivariate ReLU networks to an effective width of at most 2P-1, yielding a continuous piecewise-affine predictor whose hyperplanes are non-redundant with respect to the...

  2. The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions

    cs.LG 2025-06 conditional novelty 7.0 of 10

    A single-weight perturbation at the very start of training makes otherwise identical neural networks diverge to different loss basins, and this sensitivity drops sharply within the first fraction of training.

  3. Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification

    cs.LG 2026-07 reject novelty 6.0 of 10

    KSSE maps frozen CNN features onto QC-LDPC graphs and does Nishimori-temperature spectral embedding, reporting 88.93% ImageNet-1K Top-1 under a transductive protocol with ~21M parameters.

Pith tools