Pith. sign in

REVIEW 1 cited by

PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.06453 v1 pith:36WJYWZH submitted 2024-04-09 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords neuronspurepolysemanticdeepfeaturefeaturesidentifyingmethod
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically and encode for multiple (unrelated) features, which renders their interpretation difficult. We present a method for disentangling polysemanticity of any Deep Neural Network by decomposing a polysemantic neuron into multiple monosemantic "virtual" neurons. This is achieved by identifying the relevant sub-graph ("circuit") for each "pure" feature. We demonstrate how our approach allows us to find and disentangle various polysemantic units of ResNet models trained on ImageNet. While evaluating feature visualizations using CLIP, our method effectively disentangles representations, improving upon methods based on neuron activations. Our code is available at https://github.com/maxdreyer/PURE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Guaranteed Optimal Compositional Explanations for Neurons

    cs.AI 2025-11 conditional novelty 7.0 of 10

    A best-first search with an admissible heuristic computes guaranteed-optimal compositional neuron explanations in practical time and shows beam search was suboptimal in 10–40% of cases with overlapping concepts.

Pith tools