Pith. sign in

REVIEW 2 cited by

Compositional Explanations of Neurons

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.14032 v2 pith:3R7CT2XU submitted 2020-06-24 cs.LG cs.AIcs.CLcs.CVstat.ML

classification cs.LGcs.AIcs.CLcs.CVstat.ML
keywords neuronscompositionalexplanationsbehaviorconceptsperformancecorrelateddetect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We describe a procedure for explaining neurons in deep representations by identifying compositional logical concepts that closely approximate neuron behavior. Compared to prior work that uses atomic labels as explanations, analyzing neurons compositionally allows us to more precisely and expressively characterize their behavior. We use this procedure to answer several questions on interpretability in models for vision and natural language processing. First, we examine the kinds of abstractions learned by neurons. In image classification, we find that many neurons learn highly abstract but semantically coherent visual concepts, while other polysemantic neurons detect multiple unrelated features; in natural language inference (NLI), neurons learn shallow lexical heuristics from dataset biases. Second, we see whether compositional explanations give us insight into model performance: vision neurons that detect human-interpretable concepts are positively correlated with task performance, while NLI neurons that fire for shallow heuristics are negatively correlated with task performance. Finally, we show how compositional explanations provide an accessible way for end users to produce simple "copy-paste" adversarial examples that change model behavior in predictable ways.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 51 citations worldwide. Full citation record

  1. Evaluating SAE interpretability without explanations

    cs.LG 2025-07 conditional novelty 5.0 of 10

    SAE latent interpretability can be scored directly from activation examples via intruder detection and embedding clustering, with LLM scores correlating strongly with human scores.

  2. Quantum-Inspired Differentiable Integral Neural Networks (QIDINNs): A Feynman-Based Architecture for Continuous Learning Over Streaming Data

    cs.SE 2025-06 reject novelty 2.0 of 10

    QIDINNs define parameter updates as kernel-weighted integrals of past gradients, essentially continuous-time momentum, and claim superior streaming learning without providing the backpropagation cost they claim to avoid.

Pith tools