Pith. sign in

REVIEW 1 cited by

The Persian Rug: solving toy models of superposition using large-scale symmetries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.12101 v2 pith:76DWL4H7 submitted 2024-10-15 cs.LG cond-mat.dis-nncs.AI

classification cs.LGcond-mat.dis-nncs.AI
keywords modelmodelsalgorithmdatalossweightsactivationfunction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a complete mechanistic description of the algorithm learned by a minimal non-linear sparse data autoencoder in the limit of large input dimension. The model, originally presented in arXiv:2209.10652, compresses sparse data vectors through a linear layer and decompresses using another linear layer followed by a ReLU activation. We notice that when the data is permutation symmetric (no input feature is privileged) large models reliably learn an algorithm that is sensitive to individual weights only through their large-scale statistics. For these models, the loss function becomes analytically tractable. Using this understanding, we give the explicit scalings of the loss at high sparsity, and show that the model is near-optimal among recently proposed architectures. In particular, changing or adding to the activation function any elementwise or filtering operation can at best improve the model's performance by a constant factor. Finally, we forward-engineer a model with the requisite symmetries and show that its loss precisely matches that of the trained models. Unlike the trained model weights, the low randomness in the artificial weights results in miraculous fractal structures resembling a Persian rug, to which the algorithm is oblivious. Our work contributes to neural network interpretability by introducing techniques for understanding the structure of autoencoders. Code to reproduce our results can be found at https://github.com/KfirD/PersianRug .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Lattice Representation Hypothesis of Large Language Models

    cs.AI 2026-03 unverdicted novelty 5.0 of 10

    LLMs are claimed to encode conceptual hierarchies as concept lattices from thresholded linear attribute directions, but the empirical support is in-sample and partly LLM-generated.

Pith tools