Pith. sign in

REVIEW 30 cited by

Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.11279 v5 pith:KXRTGZHV submitted 2017-11-30 stat.ML

classification stat.ML
keywords cavsclassificationconceptimageinternalstatetestingactivation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The interpretation of deep learning models is a challenge due to their size, complexity, and often opaque internal state. In addition, many systems, such as image classifiers, operate on low-level features rather than high-level concepts. To address these challenges, we introduce Concept Activation Vectors (CAVs), which provide an interpretation of a neural net's internal state in terms of human-friendly concepts. The key idea is to view the high-dimensional internal state of a neural net as an aid, not an obstacle. We show how to use CAVs as part of a technique, Testing with CAVs (TCAV), that uses directional derivatives to quantify the degree to which a user-defined concept is important to a classification result--for example, how sensitive a prediction of "zebra" is to the presence of stripes. Using the domain of image classification as a testing ground, we describe how CAVs may be used to explore hypotheses and generate insights for a standard image classification network as well as a medical application.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 30 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 479 citations worldwide. Full citation record

  1. When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Tensor similarity is a symmetry-invariant metric that measures functional equivalence between tensor-based networks using a recursive algorithm for cross-layer mechanisms.

  2. From Mechanistic to Compositional Interpretability

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    The paper introduces compositional interpretability as a category-theoretic framework that casts mechanistic explanations as commuting syntactic-semantic mappings optimized under faithfulness and complexity constraint...

  3. Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Alignment on Llama 3 reduces explicit bias but amplifies implicit bias, because aligned models no longer represent 'black' and 'white' as racial concepts in ambiguous contexts.

  4. Unifying machine learning and quantum chemistry -- a deep neural network for molecular wavefunctions

    physics.chem-ph 2019-06 unverdicted novelty 7.0 of 10

    Deep neural network predicts molecular wavefunctions in atomic orbital basis from which quantum properties are derived at force-field efficiency.

  5. Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

    astro-ph.HE 2026-08 conditional novelty 6.0 of 10

    Sparse autoencoders reveal clean-brightness and auxiliary-activity latents in PolarBERT's CLS representation that are causally inert for direction reconstruction but causally used by an uncertainty head, which improve...

  6. CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A sparse dictionary learned from an ECG foundation model's embeddings recovers interpretable cardiac concepts—PVCs, atrial fibrillation, bundle branch blocks, ST/T-wave segments—and transfers to an external dataset wi...

  7. CLIF: Concept-Level Influence Functions for Transparent Bottleneck Models

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    CLIF applies influence functions to pinpoint influential training samples and key concepts in Concept Bottleneck Models, enabling data debugging and behavioral insights on CEBaB and Yelp datasets.

  8. Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    A latent mediation framework with sparse autoencoders enables non-additive token-level influence attribution in LLMs by learning orthogonal features and back-propagating attributions.

  9. UNBOX: Unveiling Black-box visual models with Natural-language

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    UNBOX recovers interpretable text concepts that maximally activate classes in black-box vision models by recasting activation maximization as semantic search with LLMs and diffusion models.

  10. Learning Encoding-Decoding Direction Pairs to Unveil Concepts of Influence in Deep Vision Networks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    An unsupervised method, EDDP, jointly learns encoding-decoding direction pairs for concepts in CNN latent spaces, recovering interpretable and influential concepts without labels, validated on synthetic and real data.

  11. Concept Based Explanations and Class Contrasting

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A concept-based explanation method filters hidden-layer activations by attribution, decomposes them with NMF, and validates by recombining concept crops, plus a class-contrast variant using a linear classifier.

  12. Concept Boundary Vectors

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Concept boundary vectors are derived from the boundary between latent concept clusters, and the paper reports they capture semantic relationships better than concept activation vectors.

  13. Latent Fact-Checking: Detecting Misinformation through Activation Engineering

    cs.LG 2026-08 conditional novelty 5.0 of 10

    A projection onto a contrastive falsehood direction in frozen LLM activations, followed by a small MLP, outperforms zero-shot and few-shot prompting on LIAR and FACTors and supports the claim that truthfulness is line...

  14. Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A Bayesian-network monitor built on calibrated LLM log-probabilities gives workflow-level uncertainty scores for an actuarial multi-agent system, reproducing baseline RMSE but not clearly separating normal from pertur...

  15. Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    Sparse autoencoders with supervised feature selection achieve near-LoRA performance on the AxBench steering benchmark and identify causal features.

  16. From Features to Actions: Explainability in Traditional and Agentic AI Systems

    cs.AI 2026-02 conditional novelty 5.0 of 10

    Attribution explanations that work for static classifiers do not diagnose failures in multi-step AI agents; trace-grounded rubric evaluation does, with state-tracking inconsistency 2.7x more common in failed agent runs.

  17. A Concept-based approach to Voice Disorder Detection

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Concept bottleneck and concept embedding models, trained on clinical concepts extracted from patient notes by a large language model, detect voice pathology from audio almost as accurately as an end-to-end transformer.

  18. Wanting to Be Understood Explains the Meta-Problem of Consciousness

    q-bio.NC 2025-06 conditional novelty 5.0 of 10

    A social motivation to be understood, combined with the severe bandwidth limit of language, explains why conscious experience feels ineffable and why the hard problem of consciousness persists.

  19. Towards a Science of Causal Interpretability in Deep Learning for Software Engineering

    cs.SE 2025-05 conditional novelty 5.0 of 10

    The dissertation presents docode, a causal interpretability method for neural code models, and uses a case study to show that some correlations between code properties and model performance are confounded rather than causal.

  20. "All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Encoder models trained on noisy classroom ratings look super-human under standard concordance metrics, but generalizability, disattenuation, and hierarchical rater analyses show the apparent advantage is partly spurio...

  21. Generative Counterfactual Introspection for Explainable Deep Learning

    cs.LG 2019-07 unverdicted novelty 5.0 of 10

    A generative-model-driven introspection method produces counterfactual image edits to explain deep neural network predictions on MNIST and CelebA.

  22. Towards explainable decision support using hybrid neural models for logistic terminal automation

    cs.AI 2025-09 unverdicted novelty 4.0 of 10

    The paper proposes a three-stage Interpretable Neural System Dynamics pipeline for interpretable-by-design decision support in intermodal logistics, but provides no validation.

  23. Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs

    cs.LG 2025-07 reject novelty 4.0 of 10

    BAGEL trains per-layer logistic-regression probes on CLIP-defined concepts and compares per-class concept probabilities with dataset-level concept frequencies, visualizing the alignment in a knowledge graph.

  24. Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning

    cs.AI 2025-02 unverdicted novelty 4.0 of 10

    A perspective paper that advocates direct, post hoc interpretability for multi-agent deep reinforcement learning and offers a taxonomy of where those methods might apply.

  25. Explaining Model Overfitting in CNNs via GMM Clustering

    cs.LG 2024-12 reject novelty 4.0 of 10

    A CNN filter whose feature-map GMM clustering has small outlier clusters is called an anomaly filter and is claimed to indicate model overfitting, but the supporting experiments are weakly consistent.

  26. New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing

    cs.CL 2024-11 conditional novelty 4.0 of 10

    The thesis shows that randomly masking input tokens during fine-tuning makes post-hoc explanations of NLP models consistently faithful under an erasure-based faithfulness metric.

  27. FastCAV: Efficient Computation of Concept Activation Vectors for Explaining Deep Neural Networks

    cs.LG 2025-05 conditional novelty 3.0 of 10

    Concept activation vectors can be computed as the normalized difference between concept-mean and global-mean activations, giving a 46.4x average speedup over SVM-based CAVs with comparable quality.

  28. Explainable Artificial Intelligence Techniques for Interpretation of Food Models: a Review

    cs.AI 2025-04 unverdicted novelty 3.0 of 10

    A survey proposing a taxonomy of XAI techniques for food quality research organized by data types and explanation methods.

  29. Unexplainability and Incomprehensibility of Artificial Intelligence

    cs.CY 2019-06 unverdicted novelty 3.0 of 10

    Advanced AI systems are unexplainable in full and produce explanations that humans cannot comprehend.

  30. Platonic Projection Structures: Operator-Induced Observability in Representation Learning

    cs.LG 2026-07 reject novelty 2.0 of 10

    The paper introduces 'Platonic Projection Structures,' a reformulation of standard PSD operator theory applied to representation learning, with experiments that verify definitions rather than test predictions.

Pith tools