Pith. sign in

REVIEW 4 cited by

Explaining Explainability: Recommendations for Effective Use of Concept Activation Vectors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03713 v2 pith:RPQYIR2N submitted 2024-04-04 cs.LG cs.AIcs.CVcs.HC

Explaining Explainability: Recommendations for Effective Use of Concept Activation Vectors

classification cs.LG cs.AIcs.CVcs.HC
keywords conceptconceptscavsdatasetpropertiesrecommendationsactivationdesigned
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Concept-based explanations translate the internal representations of deep learning models into a language that humans are familiar with: concepts. One popular method for finding concepts is Concept Activation Vectors (CAVs), which are learnt using a probe dataset of concept exemplars. In this work, we investigate three properties of CAVs: (1) inconsistency across layers, (2) entanglement with other concepts, and (3) spatial dependency. Each property provides both challenges and opportunities in interpreting models. We introduce tools designed to detect the presence of these properties, provide insight into how each property can lead to misleading explanations, and provide recommendations to mitigate their impact. To demonstrate practical applications, we apply our recommendations to a melanoma classification task, showing how entanglement can lead to uninterpretable results and that the choice of negative probe set can have a substantial impact on the meaning of a CAV. Further, we show that understanding these properties can be used to our advantage. For example, we introduce spatially dependent CAVs to test if a model is translation invariant with respect to a specific concept and class. Our experiments are performed on natural images (ImageNet), skin lesions (ISIC 2019), and a new synthetic dataset, Elements. Elements is designed to capture a known ground truth relationship between concepts and classes. We release this dataset to facilitate further research in understanding and evaluating interpretability methods.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. $\alpha$-TCAV: A Unified Framework for Testing with Concept Activation Vectors

    stat.ML 2026-05 unverdicted novelty 7.0

    α-TCAV replaces TCAV's hard indicator with a tunable smooth function to create a unified probabilistic framework with lower variance and guidance for parameter choice or Bayes-optimal scoring.

  2. E-TCAV: Formalizing Penultimate Proxies for Efficient Concept Based Interpretability

    cs.AI 2026-05 unverdicted novelty 6.0

    E-TCAV shows that penultimate-layer proxies yield TCAV scores agreeing with earlier layers while delivering linear speedups in computation.

  3. The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail

    cs.LG 2025-12 conditional novelty 6.0

    Reliable concept presence in transformers is concentrated in the extreme high-activation tail of in-concept tokens; thresholding that tail improves concept detection and localization.

  4. GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability

    cs.CV 2025-08 conditional novelty 5.0

    A cross-layer fusion framework that reduces variance in TCAV concept-attribution scores, at the cost of some concept-signal drift toward 0.5.