Pith. sign in

REVIEW 1 cited by

Explaining Image Classifiers Using Contrastive Counterfactuals in Generative Latent Spaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.05257 v1 pith:L5MRP5WJ submitted 2022-06-10 cs.CV cs.AI

classification cs.CVcs.AI
keywords classifierscausalclassifiercontrastivecounterfactualexplanationsgenerativeimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite their high accuracies, modern complex image classifiers cannot be trusted for sensitive tasks due to their unknown decision-making process and potential biases. Counterfactual explanations are very effective in providing transparency for these black-box algorithms. Nevertheless, generating counterfactuals that can have a consistent impact on classifier outputs and yet expose interpretable feature changes is a very challenging task. We introduce a novel method to generate causal and yet interpretable counterfactual explanations for image classifiers using pretrained generative models without any re-training or conditioning. The generative models in this technique are not bound to be trained on the same data as the target classifier. We use this framework to obtain contrastive and causal sufficiency and necessity scores as global explanations for black-box classifiers. On the task of face attribute classification, we show how different attributes influence the classifier output by providing both causal and contrastive feature attributions, and the corresponding counterfactual images.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Faithful Counterfactual Visual Explanations (FCVE)

    cs.CV 2025-01 reject novelty 4.0 of 10

    FCVE uses a decoder to turn modified convolutional filters into visual counterfactual explanations for MNIST and Fashion-MNIST classifiers.

Pith tools