REVIEW 9 cited by
Polysemanticity and Capacity in Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Individual neurons in neural networks often represent a mixture of unrelated features. This phenomenon, called polysemanticity, can make interpreting neural networks more difficult and so we aim to understand its causes. We propose doing so through the lens of feature \emph{capacity}, which is the fractional dimension each feature consumes in the embedding space. We show that in a toy model the optimal capacity allocation tends to monosemantically represent the most important features, polysemantically represent less important features (in proportion to their impact on the loss), and entirely ignore the least important features. Polysemanticity is more prevalent when the inputs have higher kurtosis or sparsity and more prevalent in some architectures than others. Given an optimal allocation of capacity, we go on to study the geometry of the embedding space. We find a block-semi-orthogonal structure, with differing block sizes in different models, highlighting the impact of model architecture on the interpretability of its neurons.
Forward citations
Cited by 9 Pith papers
-
Adversarial Attacks Leverage Interference Between Features in Superposition
Superposition—packing more features than dimensions—is sufficient to create adversarial vulnerability, and attack directions and transferability are predictable from the resulting feature geometry.
-
Compressed Computation: Dense Circuits in a Toy Model of the Universal-AND Problem
A toy model learns a dense binary-weighted circuit for the Universal-AND problem, which outperforms prior sparse constructions at low sparsity.
-
Localizing Persona Representations in LLMs
Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.
-
Expand Neurons, Not Parameters
Fixed Parameter Expansion — duplicating neurons and partitioning their incoming weights into disjoint sparse sub-neurons at constant non-zero parameter count — reduces measured feature interference and improves classi...
-
How Causal Abstraction Underpins Computational Explanation
Computational implementation is analyzed as abstraction-under-translation in causal models, with representation and generalization as further constraints.
-
DCN^2: Interplay of Implicit Collision Weights and Explicit Cross Layers for Large-Scale Recommendation
DCN^2 augments DCNv2 with collision-weighted lookups, a dense-only cross layer, and an FFM-like similarity layer, and reports improved offline and online recommendation performance.
-
A Closer Look at Multimodal Representation Collapse
The authors argue that modality collapse is caused by polysemantic neurons entangling noisy and predictive features across modalities, and that knowledge distillation or explicit basis reallocation frees rank bottlene...
-
Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
LLM explanations split into local and mechanistic tracks; the paper argues they are trustworthy only if they pass causal and contrastive stress tests, adapt to the explainee, and satisfy eight trust principles.
-
Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning
Three architecture-level interventions, overlapping convolutional tokenization plus sequence pooling, variadic attention receptive fields, and flow-aware distillation, each improve efficiency or quality in vision and ...
Discussion (0). Sign in to comment.