Pith. sign in

REVIEW 1 cited by

ViT-CX: Causal Explanation of Vision Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.03064 v3 pith:XUMDLMBB submitted 2022-11-06 cs.CV cs.AI

classification cs.CVcs.AI
keywords vit-cxvitscausalexplanationbetterembeddingsmapsmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the popularity of Vision Transformers (ViTs) and eXplainable AI (XAI), only a few explanation methods have been designed specially for ViTs thus far. They mostly use attention weights of the [CLS] token on patch embeddings and often produce unsatisfactory saliency maps. This paper proposes a novel method for explaining ViTs called ViT-CX. It is based on patch embeddings, rather than attentions paid to them, and their causal impacts on the model output. Other characteristics of ViTs such as causal overdetermination are also considered in the design of ViT-CX. The empirical results show that ViT-CX produces more meaningful saliency maps and does a better job revealing all important evidence for the predictions than previous methods. The explanation generated by ViT-CX also shows significantly better faithfulness to the model. The codes and appendix are available at https://github.com/vaynexie/CausalX-ViT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Head Explainer: A General Framework to Improve Explainability in CNNs and Transformers

    cs.CV 2025-01 reject novelty 3.0 of 10

    MHEX inserts attention-gated deep-supervision heads into ResNet and BERT and derives saliency maps from the product of the head weights, claiming better accuracy and more detailed explanations.

Pith tools