Pith. sign in

REVIEW 1 cited by

On Implications of Scaling Laws on Feature Superposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.01459 v1 pith:CCRSY7UU submitted 2024-07-01 cs.LG cs.AI

classification cs.LGcs.AI
keywords featuresfeaturelawsscalingsuperpositionachievingacrossargues
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Using results from scaling laws, this theoretical note argues that the following two statements cannot be simultaneously true: 1. Superposition hypothesis where sparse features are linearly represented across a layer is a complete theory of feature representation. 2. Features are universal, meaning two models trained on the same data and achieving equal performance will learn identical features.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SAFR: Neuron Redistribution for Interpretability

    cs.LG 2025-01 conditional novelty 4.0 of 10

    SAFR regularizes a transformer so that important tokens become monosemantic (one neuron per meaning) and correlated tokens share neurons, and it uses the accuracy drop after deleting high-capacity tokens as its interp...

Pith tools