Pith. sign in

REVIEW 1 cited by

Overparameterized Neural Networks Implement Associative Memory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.12362 v2 pith:QEQN74TF submitted 2019-09-26 cs.LG stat.ML

classification cs.LGstat.ML
keywords mechanismmemoryoverparameterizedautoencodersautoencodingdataefficientencoding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Identifying computational mechanisms for memorization and retrieval of data is a long-standing problem at the intersection of machine learning and neuroscience. Our main finding is that standard overparameterized deep neural networks trained using standard optimization methods implement such a mechanism for real-valued data. Empirically, we show that: (1) overparameterized autoencoders store training samples as attractors, and thus, iterating the learned map leads to sample recovery; (2) the same mechanism allows for encoding sequences of examples, and serves as an even more efficient mechanism for memory than autoencoding. Theoretically, we prove that when trained on a single example, autoencoders store the example as an attractor. Lastly, by treating a sequence encoder as a composition of maps, we prove that sequence encoding provides a more efficient mechanism for memory than autoencoding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic and episodic memories in a predictive coding model of the neocortex

    cs.LG 2025-09 conditional novelty 4.0 of 10

    A predictive coding model of the neocortex recalls individual MNIST examples only when trained on a tiny batch; training on the full dataset preserves semantic reconstruction but degrades episodic recall.

Pith tools