Pith. sign in

REVIEW 9 cited by

Polysemanticity and Capacity in Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.01892 v4 pith:PRTWLIKR submitted 2022-10-04 cs.NE cs.AIcs.LG

classification cs.NEcs.AIcs.LG
keywords capacityfeaturesimportantnetworksneuralpolysemanticityrepresentallocation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Individual neurons in neural networks often represent a mixture of unrelated features. This phenomenon, called polysemanticity, can make interpreting neural networks more difficult and so we aim to understand its causes. We propose doing so through the lens of feature \emph{capacity}, which is the fractional dimension each feature consumes in the embedding space. We show that in a toy model the optimal capacity allocation tends to monosemantically represent the most important features, polysemantically represent less important features (in proportion to their impact on the loss), and entirely ignore the least important features. Polysemanticity is more prevalent when the inputs have higher kurtosis or sparsity and more prevalent in some architectures than others. Given an optimal allocation of capacity, we go on to study the geometry of the embedding space. We find a block-semi-orthogonal structure, with differing block sizes in different models, highlighting the impact of model architecture on the interpretability of its neurons.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adversarial Attacks Leverage Interference Between Features in Superposition

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Superposition—packing more features than dimensions—is sufficient to create adversarial vulnerability, and attack directions and transferability are predictable from the resulting feature geometry.

  2. Compressed Computation: Dense Circuits in a Toy Model of the Universal-AND Problem

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A toy model learns a dense binary-weighted circuit for the Universal-AND problem, which outperforms prior sparse constructions at low sparsity.

  3. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  4. Expand Neurons, Not Parameters

    cs.LG 2025-10 reject novelty 5.0 of 10

    Fixed Parameter Expansion — duplicating neurons and partitioning their incoming weights into disjoint sparse sub-neurons at constant non-zero parameter count — reduces measured feature interference and improves classi...

  5. How Causal Abstraction Underpins Computational Explanation

    cs.LG 2025-08 conditional novelty 5.0 of 10

    Computational implementation is analyzed as abstraction-under-translation in causal models, with representation and generalization as further constraints.

  6. DCN^2: Interplay of Implicit Collision Weights and Explicit Cross Layers for Large-Scale Recommendation

    cs.IR 2025-06 conditional novelty 5.0 of 10

    DCN^2 augments DCNv2 with collision-weighted lookups, a dense-only cross layer, and an FFM-like similarity layer, and reports improved offline and online recommendation performance.

  7. A Closer Look at Multimodal Representation Collapse

    cs.LG 2025-05 reject novelty 5.0 of 10

    The authors argue that modality collapse is caused by polysemantic neurons entangling noisy and predictive features across modalities, and that knowledge distillation or explicit basis reallocation frees rank bottlene...

  8. Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations

    cs.CL 2025-10 conditional novelty 4.0 of 10

    LLM explanations split into local and mechanistic tracks; the paper argues they are trustworthy only if they pass causal and contrastive stress tests, adapt to the explainee, and satisfy eight trust principles.

  9. Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning

    cs.CV 2025-07 conditional novelty 2.0 of 10

    Three architecture-level interventions, overlapping convolutional tokenization plus sequence pooling, variadic attention receptive fields, and flow-aware distillation, each improve efficiency or quality in vision and ...

Pith tools