Pith. sign in

REVIEW 9 cited by

Explaining Classifiers with Causal Concept Effect (CaCE)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.07165 v2 pith:RMUNGK3J submitted 2019-07-16 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords cacecausalconcepteffectconfoundingdeepmeasureneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How can we understand classification decisions made by deep neural networks? Many existing explainability methods rely solely on correlations and fail to account for confounding, which may result in potentially misleading explanations. To overcome this problem, we define the Causal Concept Effect (CaCE) as the causal effect of (the presence or absence of) a human-interpretable concept on a deep neural net's predictions. We show that the CaCE measure can avoid errors stemming from confounding. Estimating CaCE is difficult in situations where we cannot easily simulate the do-operator. To mitigate this problem, we use a generative model, specifically a Variational AutoEncoder (VAE), to measure VAE-CaCE. In an extensive experimental analysis, we show that the VAE-CaCE is able to estimate the true concept causal effect, compared to baselines for a number of datasets including high dimensional images.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 94 citations worldwide. Full citation record

  1. TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models

    cs.LG 2026-01 conditional novelty 6.0 of 10

    TimeSAE trains a sparse autoencoder with counterfactual and consistency losses to explain black-box time series predictions, claiming better faithfulness and out-of-distribution robustness than eight baselines.

  2. Do you see what I see? An Ambiguous Optical Illusion Dataset exposing limitations of Explainable AI

    cs.CV 2025-05 reject novelty 6.0 of 10

    Ambivision is a new dataset of AI-generated animal optical illusions, and the paper claims that adding gaze and eye annotations as visible image features improves classification while exposing limits of pixel-based ex...

  3. Enhancing the Comprehensibility of Text Explanations via Unsupervised Concept Discovery

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ECO-Concept uses slot attention plus LLM-based comprehensibility feedback to learn explainable text concepts without concept annotations.

  4. Local Label-Informed Feature Transfer for Generating Ground-Truth Medical Images: A Comparison of GAN- and Diffusion-Based Approaches

    cs.CV 2026-07 conditional novelty 5.0 of 10

    LLIFT generates semi-synthetic brain MRIs with user-placed lesion-like patches using weak labels, though its image-level FID results do not show the patch itself is pathology-realistic.

  5. ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series

    cs.LG 2026-07 conditional novelty 5.0 of 10

    ConceptCF generates time-series counterfactuals by evolving interpretable concepts (scale, bias, frequency bands) with a genetic algorithm, achieving top-tier or near-top-tier scores on validity, proximity, sparsity, ...

  6. GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A cross-layer fusion framework that reduces variance in TCAV concept-attribution scores, at the cost of some concept-signal drift toward 0.5.

  7. A Concept-based approach to Voice Disorder Detection

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Concept bottleneck and concept embedding models, trained on clinical concepts extracted from patient notes by a large language model, detect voice pathology from audio almost as accurately as an end-to-end transformer.

  8. Towards explainable decision support using hybrid neural models for logistic terminal automation

    cs.AI 2025-09 unverdicted novelty 4.0 of 10

    The paper proposes a three-stage Interpretable Neural System Dynamics pipeline for interpretable-by-design decision support in intermodal logistics, but provides no validation.

  9. Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A broad modality-aware survey of explainable AI methods for biomedical imaging, covering heatmap, concept, text, and latent-space approaches plus tools, metrics, and vision-language models.

Pith tools