Pith. sign in

REVIEW 1 cited by

Exact full-RSB SAT/UNSAT transition in infinitely wide two-layer neural networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.06717 v4 pith:3IRT7TCS submitted 2024-10-09 cond-mat.dis-nn cs.LGmath.PR

classification cond-mat.dis-nncs.LGmath.PR
keywords overlapstatesalgorithmscapacityexactfull-rsbmarginmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We analyze the problem of storing random pattern-label associations using two classes of continuous non-convex weights models, namely the perceptron with negative margin and an infinite-width two-layer neural network with non-overlapping receptive fields and generic activation function. Using a full-RSB ansatz we compute the exact value of the SAT/UNSAT transition. Furthermore, in the case of the negative perceptron we show that the overlap distribution of typical states displays an overlap gap (a disconnected support) in certain regions of the phase diagram defined by the value of the margin and the density of patterns to be stored. This implies that some recent theorems that ensure convergence of Approximate Message Passing (AMP) based algorithms to capacity are not applicable. Finally, we show that Gradient Descent is not able to reach the maximal capacity, irrespectively of the presence of an overlap gap for typical states. This finding, similarly to what occurs in binary weight models, suggests that gradient-based algorithms are biased towards highly atypical states, whose inaccessibility determines the algorithmic threshold.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep ReLU networks -- injectivity capacity upper bounds

    stat.ML 2024-12 reject novelty 6.0 of 10

    For deep ReLU networks with random Gaussian weights, the paper gives upper bounds on the layer expansion needed for injectivity and finds the expansion need saturates by four layers.

Pith tools