Pith. sign in

REVIEW 4 cited by

Transformer Interpretability Beyond Attention Visualization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.09838 v2 pith:QM7HG4GO submitted 2020-12-17 cs.CV

classification cs.CV
keywords attentionclassificationexistinglayersmethodsrelevancytransformermethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-attention techniques, and specifically Transformers, are dominating the field of text processing and are becoming increasingly popular in computer vision classification tasks. In order to visualize the parts of the image that led to a certain classification, existing methods either rely on the obtained attention maps or employ heuristic propagation along the attention graph. In this work, we propose a novel way to compute relevancy for Transformer networks. The method assigns local relevance based on the Deep Taylor Decomposition principle and then propagates these relevancy scores through the layers. This propagation involves attention layers and skip connections, which challenge existing methods. Our solution is based on a specific formulation that is shown to maintain the total relevancy across layers. We benchmark our method on very recent visual Transformer networks, as well as on a text classification problem, and demonstrate a clear advantage over the existing explainability methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformers as Unrolled Inference in Probabilistic Laplacian Eigenmaps: An Interpretation and Potential Improvements

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Transformers can be viewed as unrolled inference in a probabilistic Laplacian Eigenmaps model, and replacing the attention matrix by attention minus identity improves validation performance.

  2. Learning to Explain: Prototype-Based Surrogate Models for LLM Classification

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A sentence-level prototype surrogate is trained to imitate LLM predictions and its prototype matches are used as explanations; the paper claims state-of-the-art faithfulness with moderate evidence.

  3. Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A method and case study for visualizing cross-attention maps in Wan video diffusion transformers, showing token-region alignment over time and their use as artistic material.

  4. Interpretable by AI Mother Tongue: Native Symbolic Reasoning in Neural Models

    cs.CL 2025-08 reject novelty 4.0 of 10

    A VQ-based gated Transformer trained on AG News produces symbol traces that are supposed to be interpretable, but accuracy is low (50.94% then 47.32%) and central results are unreported.

Pith tools