REVIEW 7 cited by
Explaining Classifiers with Causal Concept Effect (CaCE)
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
How can we understand classification decisions made by deep neural networks? Many existing explainability methods rely solely on correlations and fail to account for confounding, which may result in potentially misleading explanations. To overcome this problem, we define the Causal Concept Effect (CaCE) as the causal effect of (the presence or absence of) a human-interpretable concept on a deep neural net's predictions. We show that the CaCE measure can avoid errors stemming from confounding. Estimating CaCE is difficult in situations where we cannot easily simulate the do-operator. To mitigate this problem, we use a generative model, specifically a Variational AutoEncoder (VAE), to measure VAE-CaCE. In an extensive experimental analysis, we show that the VAE-CaCE is able to estimate the true concept causal effect, compared to baselines for a number of datasets including high dimensional images.
Forward citations
Cited by 7 Pith papers
-
TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models
TimeSAE trains a sparse autoencoder with counterfactual and consistency losses to explain black-box time series predictions, claiming better faithfulness and out-of-distribution robustness than eight baselines.
-
Local Label-Informed Feature Transfer for Generating Ground-Truth Medical Images: A Comparison of GAN- and Diffusion-Based Approaches
LLIFT generates semi-synthetic brain MRIs with user-placed lesion-like patches using weak labels, though its image-level FID results do not show the patch itself is pathology-realistic.
-
ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series
ConceptCF generates time-series counterfactuals by evolving interpretable concepts (scale, bias, frequency bands) with a genetic algorithm, achieving top-tier or near-top-tier scores on validity, proximity, sparsity, ...
-
GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability
A cross-layer fusion framework that reduces variance in TCAV concept-attribution scores, at the cost of some concept-signal drift toward 0.5.
-
A Concept-based approach to Voice Disorder Detection
Concept bottleneck and concept embedding models, trained on clinical concepts extracted from patient notes by a large language model, detect voice pathology from audio almost as accurately as an end-to-end transformer.
-
Towards explainable decision support using hybrid neural models for logistic terminal automation
The paper proposes a three-stage Interpretable Neural System Dynamics pipeline for interpretable-by-design decision support in intermodal logistics, but provides no validation.
-
Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey
A broad modality-aware survey of explainable AI methods for biomedical imaging, covering heatmap, concept, text, and latent-space approaches plus tools, metrics, and vision-language models.
Discussion (0). Sign in to comment.