REVIEW 1 cited by
On Implications of Scaling Laws on Feature Superposition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Using results from scaling laws, this theoretical note argues that the following two statements cannot be simultaneously true: 1. Superposition hypothesis where sparse features are linearly represented across a layer is a complete theory of feature representation. 2. Features are universal, meaning two models trained on the same data and achieving equal performance will learn identical features.
Forward citations
Cited by 1 Pith paper
-
SAFR: Neuron Redistribution for Interpretability
SAFR regularizes a transformer so that important tokens become monosemantic (one neuron per meaning) and correlated tokens share neurons, and it uses the accuracy drop after deleting high-capacity tokens as its interp...
Discussion (0). Continue with ORCID to comment.