Derives tight upper and lower bounds on L2 loss in sparse autoencoders with power activations, tight in the very sparse regime.
The persian rug: solving toy models of superposition using large-scale symmetries.CoRR, abs/2410.12101
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
LLMs are claimed to encode conceptual hierarchies as concept lattices from thresholded linear attribute directions, but the empirical support is in-sample and partly LLM-generated.
citing papers explorer
-
Effects of sparsity and superposition on loss in simple autoencoders
Derives tight upper and lower bounds on L2 loss in sparse autoencoders with power activations, tight in the very sparse regime.
-
The Lattice Representation Hypothesis of Large Language Models
LLMs are claimed to encode conceptual hierarchies as concept lattices from thresholded linear attribute directions, but the empirical support is in-sample and partly LLM-generated.