Pith. sign in

REVIEW 4 cited by

Interpreting the Second-Order Effects of Neurons in CLIP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04341 v3 pith:DPE52BLB submitted 2024-06-06 cs.CV

classification cs.CV
keywords effectsneuronneuronsclipeffectsecond-orderanalyzingconcepts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We interpret the function of individual neurons in CLIP by automatically describing them using text. Analyzing the direct effects (i.e. the flow from a neuron through the residual stream to the output) or the indirect effects (overall contribution) fails to capture the neurons' function in CLIP. Therefore, we present the "second-order lens", analyzing the effect flowing from a neuron through the later attention heads, directly to the output. We find that these effects are highly selective: for each neuron, the effect is significant for <2% of the images. Moreover, each effect can be approximated by a single direction in the text-image space of CLIP. We describe neurons by decomposing these directions into sparse sets of text representations. The sets reveal polysemantic behavior - each neuron corresponds to multiple, often unrelated, concepts (e.g. ships and cars). Exploiting this neuron polysemy, we mass-produce "semantic" adversarial examples by generating images with concepts spuriously correlated to the incorrect class. Additionally, we use the second-order effects for zero-shot segmentation, outperforming previous methods. Our results indicate that an automated interpretation of neurons can be used for model deception and for introducing new model capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A new automated pipeline decomposes fMRI activity into components and labels them with visual concepts, claiming thousands of interpretable patterns across the human visual cortex.

  2. Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation

    cs.LG 2025-06 reject novelty 6.0 of 10

    A rotation-sensitivity hypothesis test plus Varimax rotation produces sparse concept dictionaries from CLIP embeddings and improves worst-group accuracy after spurious concept removal.

  3. Evaluating Neuron Explanations: A Unified Framework with Sanity Checks

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Most commonly used neuron explanation evaluation metrics fail two new sanity checks, and only Correlation, Cosine, AUPRC, F1-score, and IoU pass.

  4. LLMs can see and hear without any training

    cs.CV 2025-01 conditional novelty 5.0 of 10

    MILS uses an LLM plus a pretrained scorer in an iterative generate-score-refine loop to produce competitive zero-shot captions and improved text-to-image prompts without training.

Pith tools