Sparse autoencoder probes of Pythia models show concept activations jump at roughly 410M parameters and during mid-training, while early-layer features re-emerge at the output layer.
Brown, Benjamin Mann, Nick Ryder, and et al
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models
Sparse autoencoder probes of Pythia models show concept activations jump at roughly 410M parameters and during mid-training, while early-layer features re-emerge at the output layer.