REVIEW 4 major objections 3 minor 107 references
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
T0 review · 4 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Pruning and rescaling backward gradients restores an exact-output-completeness property in Transformers, improving every gradient attribution method tested.
desk verdict A genuinely useful gradient-surgery method for ViT attributions with broad empirical gains, but the 'universally enhances' claim is contradicted by its own tables and the completeness theory is a conservation identity for a surrogate gradient, not a faithfulness guarantee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is FullGrad-completeness (FG-completeness), the identity $f(x) = J_x f \cdot x + \sum_i J_{b_i} f \cdot b_i$ that forces attributions to sum exactly to the model output. The argument is carried by two backward-pass manipulation operators defined on top of an unchanged forward pass: the constant operator $[\cdot]_{\mathrm{cst}}$, which zeroes the gradient while keeping the value, and SwapBackward, which evaluates one function but propagates the Jacobian of another. Theorems 1–3 show that locally affine functions, compositions, and sums are FG-complete, which explains why CNNs inherit the property, and Propositions 1–3 show which Transformer operations break it; Theorems 4–5 then give the repair rules, scale product branches by coefficients summing to one or prune a non-FG branch. The practical Libra modules instantiate these rules: attention contributes gradients only through the value branch, LayerNorm treats its denominator as constant, gated activations discard the gate's gradient, and SwiGLU splits its backward gradient equally between the two branches.
What would settle it
Take a Transformer fine-tuned on ImageNet and compute attributions on the paper's own zebra/elephant COCO images with only the attention-gradient pruning applied and LayerNorm gradients untouched: if the pruned attention scores, rather than LayerNorm, encode the co-occurring-class distinction, the maps will localize the wrong animal. Conversely, the theory predicts that a Transformer whose normalizations are replaced by affine scaling needs no Libra LayerNorm fix, so a faithfulness gain there would disprove FG-completeness as the operative mechanism.
Extended reading notes
Core claim
LibraGrad claims that the reason gradient-based explanations underperform on Transformers is a set of non-locally-affine operations — attention softmax scores, gated activations such as GELU and SiLU, SwiGLU self-gating, and LayerNorm — that violate FullGrad-completeness, the identity $f(x) = J_x f \cdot x + \sum_i J_{b_i} f \cdot b_i$ that CNNs satisfy automatically. The paper proves that naive element-wise multiplication over-counts attributions by a factor of two, that division makes FullGrad vanish identically, and that LayerNorm's denominator drives FullGrad to zero as $\varepsilon \to 0$. It then proves that scaling the two Jacobian branches of a product by coefficients summing to one, or pruning one branch entirely by treating it as constant in the backward pass, restores the identity (Theorems 4 and 5, Corollary 2), and that applying these fixes to attention, gated activations, self-gating, and LayerNorm yields a Transformer that is FG-complete (Corollary 3). Empirically, Libra versions of Input×Grad, GradCAM+, HiResCAM, XGradCAM+, FullGrad+, AttCAT, GenAtt, and TokenTM beat their unmodified counterparts on faithfulness, completeness error, and segmentation alignment on nearly every model and dataset, with Libra FullGrad+ at 0.0 completeness error on all tested models and the top average segmentation AP.
Load-bearing premise
The load-bearing premise is that making the gradient-sensitive parts of a Transformer behave as constants in the backward pass, which guarantees attributions sum to the output, is enough to make attributions faithful, and that the discarded gradient signal through attention weights and LayerNorm denominators carries no essential explanation.
Editorial extensions
If this is right
- Once gradients are rebalanced, general-purpose gradient methods such as FullGrad+ outperform attention-based and Transformer-specific attribution methods on faithfulness and segmentation metrics, suggesting that specialized explainers for Transformers are unnecessary.
- Libra FullGrad+ reaches 0.0 completeness error on every large model tested and raises average segmentation AP from 43.4 to 67.9, placing it above every attention-based method in the comparison.
- LibraGrad imposes no forward-pass modification and no extra computational or memory overhead, so it can be dropped into any existing gradient pipeline.
- The method transfers to the attention-free MLP-Mixer architecture, supporting the claim that gradient imbalance, not attention per se, is the root cause of attribution failures.
- The ablation study attributes most of the gain to fixing LayerNorm, with attention and self-gating secondary and bias terms negligible, identifying the highest-value repair for practitioners.
Reading between the lines
- The argument suggests a general design rule for future architectures: whenever a new layer introduces multiplicative gating or normalization, the backward pass should be audited against FG-completeness before any attribution method is built, and the Libra recipes could be applied preemptively.
- A testable prediction follows: on a Transformer variant whose normalizations are already affine, the gain from the Libra LayerNorm fix should shrink, cleanly separating the normalization effect from the attention effect.
- The pruning choice is a modeling decision, and the scaling coefficients of Theorem 4 could be tuned to interpolate between full gradients and pruned gradients, trading exact completeness against retaining some attention-gradient signal; the paper does not explore this.
- If FG-completeness is the operative mechanism, faithfulness gains should correlate with measured completeness error across architectures and checkpoints, a correlation the paper reports only indirectly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LibraGrad, a post-hoc attribution enhancement for vision transformers that modifies backward gradients by zeroing gradients through attention softmax scores and LayerNorm denominators, pruning non-linear gates, and scaling self-gating branches. It claims that these changes restore FullGrad-completeness, that this is crucial for faithful interpretability, and that LibraGrad universally enhances a wide range of gradient-based attribution methods across faithfulness, completeness error, and segmentation metrics. The experiments cover eight architectures, four model sizes, and four datasets, with code released.
Significance. If the universal-improvement claim held, the contribution would be significant: it offers a cheap, drop-in gradient surgery applicable to IG, FullGrad+, GradCAM+, and other methods, and it poses a clear hypothesis about why gradients fail on transformers. The paper's theoretical toolbox (constant operator, SwapBackward, Theorems 4 and 5) is clean and reusable, and the empirical sweep is unusually broad, spanning multiple architectures, model sizes, datasets, and metric families. However, as detailed below, the theoretical guarantee applies to modified backward Jacobians rather than to the model's true Jacobians, and the 'universal' claim is contradicted by the paper's own tables. The method remains a plausible strong heuristic; its significance depends on an honest reframing and on reporting the negative cases.
major comments (4)
- [§3.4, Definition 1, Corollary 3] FG-completeness is stated for the true Jacobians of f, but LibraGrad's definitions replace those Jacobians with modified backward values. For example, Libra-LayerNorm keeps the forward value but zeroes the gradient through the denominator, and Libra-Attention zeroes gradients through softmax(QK^T); these are not the Jacobians of the forward function. Corollary 3 therefore certifies that a modified backward pass satisfies f(x)=J^Libra_x f·x+Σ J^Libra_bi f·bi, which is enforced by construction rather than derived from the model. Table 4's CE=0 (Libra FullGrad) is an implementation sanity check of that custom backward pass, as Section B.1 acknowledges, and it cannot by itself support the abstract's 'theoretically grounded' faithfulness claim. The divergence is concrete: Proposition 3 shows the true JxLN·x tends to zero as epsilon approaches zero, while the Libra backward pass returns LN(x). The paper should either prove a formal connection between modified-Jacobian completeness and the faithfulness metrics or explicitly reframe the contribution as a heuristic gradient reweighting with empirical support.
- [Tables 2 and 3; §4.2] The claim of universal enhancement across all metrics is not supported by the paper's own results. In Table 2 (MIF accuracy on ViT-B), Libra Input×Grad drops on MURA from 25.5 to 21.6, and Libra TokenTM is unchanged at 28.0. In Table 3 (Segmentation AP), Libra GradCAM+ drops on SigLIP-L from 44.3±0.4 to 41.7±0.3 and on DeiT3-H from 60.3±0.4 to 46.7±0.4; similar decreases appear in other cells. Section 4.2's statement that 'LibraGrad universally enhances gradient-based attribution methods across all tested models, architectures, and datasets' is therefore false as written. The authors should present a systematic account of cases with no gain or a loss, and the abstract and conclusion should be revised to describe large but not universal gains.
- [§3.4–§3.5] The choice of scaling versus pruning coefficients is underdetermined by the theory and not justified empirically. Theorem 4 allows any a,b with a+b=1, and the paper fixes a=b=1/2 for self-gating while setting a=0 (pruning) for attention and LayerNorm; Theorem 5 justifies pruning only when the surviving branch is FG-complete. The decision to zero Q and K gradients rather than scale them is presented as a design choice in Section 3.5 with no ablation or sensitivity analysis over a,b or over pruning versus scaling for any component. Since these choices determine the attribution maps, the paper should at least report an ablation over the free coefficients and over scaling alternatives for the attention path.
- [Table 5] The ablation's interaction pattern complicates the 'LayerNorm is the most significant factor' interpretation. Removing LayerNorm alone drops MIF predicted accuracy from 71.7 to 49.9, but removing both attention and LayerNorm gives 61.2, which is higher than removing LayerNorm alone; similar nonmonotonicity appears in the GT-accuracy column (65.2 vs 49.9 vs 63.6). The text in Section 4.2 asserts the LayerNorm effect is the most significant without discussing this interaction. The authors should explain the non-additivity or soften the causal attribution.
minor comments (3)
- [§4.1] Table 2 reports that standard deviations were bounded by 0.1 and omitted, but it does not state the number of random seeds or images used for each cell; a single sentence on replication would help.
- [Table 4 caption] The Completeness Error table excludes attention-based methods 'as their incompleteness is evident'; this should be stated as a scope restriction in the main text so that the comparison is not read as a complete ranking.
- [§3.5] The expression 'Libra-Attention(Q, K, V) = [softmax(QK^T)]_cst · V' has ambiguous operator precedence; adding parentheses would clarify that the constant operator applies to the whole softmax argument.
Circularity Check
No significant circularity: the only self-referential element is the acknowledged by-construction FG-completeness of Libra's modified backward pass; the main empirical claims rest on external metrics.
-
self definitional
[Definition 1 (Section 3); Section 3.5; Corollary 3; Table 4; Appendix B.1]
"Definition 1: "A function f is FullGrad-complete (or FG-complete) if, for all x ∈ Rn, f (x) = Jxf · x + X i Jbi f · bi" ... Section 3.5: "Libra-Attention(Q, K, V) = [softmax(QK T )]cst. · V", "Libra-LayerNorm(x) = x − µ / [√(σ2 + ε)]cst." ... Corollary 3: "A Transformer architecture attains FG-completeness when all non-linear components—specifically its attention mechanisms, activation functions, self-gating operations, and LayerNorms—are replaced with their Libra counterparts.""
Definition 1 defines FG-completeness via true Jacobians Jxf=∂f/∂x and Jbi f. LibraGrad replaces these by definition: [·]cst. has Jx[y]cst.=0, and Libra-Attention / Libra-LayerNorm detach softmax(QK^T) and the LayerNorm denominator. Corollary 3 thus holds because the Libra backward pass was constructed so f(x)=J^Libra_x f·x+Σ J^Libra_bi f·b_i is satisfied by construction, not because the model's true gradients satisfy it. Table 4's CE=0.0 for Libra FullGrad is an implementation check of the custom backward pass; the paper calls CE 'just a sanity check' (Appendix B.1), so the self-reference is acknowledged and the main faithfulness/segmentation claims do not depend on it.
full rationale
The paper's main claims are empirical: faithfulness deletions, ImageNet-S segmentation AP, and qualitative CLIP/co-occurring-class comparisons are all external benchmark measurements of the produced attribution maps, not functions of the LibraGrad definition itself. The FG-completeness theorem is a genuine mathematical statement but only about the modified backward graph, which makes the CE=0 table a by-construction sanity check rather than evidence about the original model's gradients; the paper labels it as such, so this is a minor acknowledged self-reference rather than a load-bearing circularity. No load-bearing self-citation is present: FullGrad+ (ref. [49], by overlapping authors) is used as a baseline/component, not as the argument for LibraGrad's improvement. The paper's own tables contain counterexamples to the word 'universally' (e.g., MURA Input×Grad 25.5→21.6 in Table 2; SigLIP-L GradCAM+ 44.3→41.7 in Table 3) and the link from modified-backward FG-completeness to faithfulness is asserted rather than proven, but those are correctness/completeness concerns, not circularity.
Assumptions & free parameters
free parameters (1)
- Self-gate gradient scaling coefficient (a=b) =
0.5
assumptions (4)
- domain assumption FG-completeness with respect to modified backward Jacobians is a critical property for attribution faithfulness.
- domain assumption Modifying backward gradients while leaving the forward pass unchanged yields valid attributions of the original model.
- ad hoc to paper Pruning attention score gradients (Q and K) is preferable to scaling them.
- standard math Element-wise product and quotient Jacobians follow the standard chain rule for the specified branch structure.
Cite this review
Pith. "Pith review of LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions." pith.science (2026). https://pith.science/paper/WUONYQSD
@misc{pith2026241116760,
author = {Pith},
title = {Pith review of: LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions},
year = {2026},
howpublished = {\url{https://pith.science/paper/WUONYQSD}},
note = {Machine review of arXiv:2411.16760}
}
read the original abstract
Why do gradient-based explanations struggle with Transformers, and how can we improve them? We identify gradient flow imbalances in Transformers that violate FullGrad-completeness, a critical property for attribution faithfulness that CNNs naturally possess. To address this issue, we introduce LibraGrad -- a theoretically grounded post-hoc approach that corrects gradient imbalances through pruning and scaling of backward paths, without changing the forward pass or adding computational overhead. We evaluate LibraGrad using three metric families: Faithfulness, which quantifies prediction changes under perturbations of the most and least relevant features; Completeness Error, which measures attribution conservation relative to model outputs; and Segmentation AP, which assesses alignment with human perception. Extensive experiments across 8 architectures, 4 model sizes, and 4 datasets show that LibraGrad universally enhances gradient-based methods, outperforming existing white-box methods -- including Transformer-specific approaches -- across all metrics. We demonstrate superior qualitative results through two complementary evaluations: precise text-prompted region highlighting on CLIP models and accurate class discrimination between co-occurring animals on ImageNet-finetuned models -- two settings on which existing methods often struggle. LibraGrad is effective even on the attention-free MLP-Mixer architecture, indicating potential for extension to other modern architectures. Our code is freely available at https://github.com/NightMachinery/LibraGrad.
Figures
Reference graph
Works this paper leans on
-
[1]
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema. Quantifying attention flow in transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 4190–4197, Online, 2020. Association for Computa- tional Linguistics. 2, 114
2020
-
[2]
AttnLRP: Attention- aware layer-wise relevance propagation for transformers
Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebas- tian Lapuschkin, and Wojciech Samek. AttnLRP: Attention- aware layer-wise relevance propagation for transformers. In Proceedings of the 41st International Conference on Ma- chine Learning, pages 135–168. PMLR, 2024. 2, 4, 114
2024
-
[3]
XAI for trans- formers: Better explanations through conservative propaga- tion
Ameen Ali, Thomas Schnake, Oliver Eberle, Gr ´egoire Mon- tavon, Klaus-Robert M ¨uller, and Lior Wolf. XAI for trans- formers: Better explanations through conservative propaga- tion. In Proceedings of the 39th International Conference on Machine Learning, pages 435–451. PMLR, 2022. 2, 114
2022
-
[4]
Marco Ancona, Enea Ceolini, Cengiz ¨Oztireli, and Markus H. Gross. Towards better understanding of gradient- based attribution methods for deep neural networks. In In- ternational Conference on Learning Representations , 2017. 2, 113
2017
-
[5]
Anders, David Neumann, Talmaj Marinc, Wojciech Samek, Klaus-Robert M ¨uller, and Sebastian La- puschkin
Christopher J. Anders, David Neumann, Talmaj Marinc, Wojciech Samek, Klaus-Robert M ¨uller, and Sebastian La- puschkin. Xai for analyzing and unlearning spurious corre- lations in imagenet. 2020. 113
2020
-
[6]
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Gr ´egoire Montavon, Frederick Klauschen, Klaus-Robert M ¨uller, and Wojciech Samek. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE, 10, 2015. 2
2015
-
[7]
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. ArXiv, abs/2106.08254, 2021. 5
arXiv 2021
-
[8]
Tenenbaum, and Boris Katz
Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Joshua B. Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Neural Information Processing Systems, 2019. 6
2019
Show all 107 references
-
[9]
Ecqx: Explainability-driven quantization for low-bit and sparse dnns
Daniel Becking, Maximilian Dreyer, Wojciech Samek, Karsten M ¨uller, and Sebastian Lapuschkin. Ecqx: Explainability-driven quantization for low-bit and sparse dnns. ArXiv, abs/2109.04236, 2021. 113
2021 arXiv
-
[10]
H’enaff, Alexander Kolesnikov, Xi- aohua Zhai, and A ¨aron van den Oord
Lucas Beyer, Olivier J. H’enaff, Alexander Kolesnikov, Xi- aohua Zhai, and A ¨aron van den Oord. Are we done with imagenet? ArXiv, abs/2006.07159, 2020. 6
2006 arXiv
-
[11]
Alabdulmohsin, and Filip Pavetic
Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim M. Alabdulmohsin, and Filip Pavetic. Flexivit: One model for all patch sizes. 2023 IEEE/CVF Conference on Computer Vision and Pat- te...
2023
-
[12]
Layer-wise rel- evance propagation for deep neural network architectures
Alexander Binder, Sebastian Bach, Gr ´egoire Montavon, Klaus-Robert M¨uller, and Wojciech Samek. Layer-wise rel- evance propagation for deep neural network architectures
-
[13]
De- coupling pixel flipping and occlusion strategy for consistent xai benchmarks
Stefan Bl ¨ucher, Johanna Vielhaben, and Nils Strodthoff. De- coupling pixel flipping and occlusion strategy for consistent xai benchmarks. Transactions on Machine Learning Re- search, 2024. 5, 11
2024
-
[14]
On iden- tifiability in transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. On iden- tifiability in transformers. In International Conference on Learning Representations, 2020. 114
2020
-
[15]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv’e J’egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9630–9640, 2021. 2
2021
-
[16]
Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers
Hila Chefer, Shir Gur, and Lior Wolf. Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV) , pages 397–406, 2021. 2, 114
2021
-
[17]
Transformer inter- pretability beyond attention visualization
Hila Chefer, Shir Gur, and Lior Wolf. Transformer inter- pretability beyond attention visualization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 782–791, 2021. 2, 6, 12, 114
2021
-
[18]
Optimizing relevance maps of vision transformers improves robustness
Hila Chefer, Idan Schwartz, and Lior Wolf. Optimizing relevance maps of vision transformers improves robustness. ArXiv, abs/2206.01161, 2022. 113
2022 arXiv
-
[19]
Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models
Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models. ArXiv, abs/2301.13826, 2023. 113
2023 arXiv
-
[20]
Generative pre- training from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Hee- woo Jun, David Luan, and Ilya Sutskever. Generative pre- training from pixels. In Proceedings of the 37th Interna- tional Conference on Machine Learning , pages 1691–1703. PMLR, 2020. 5, 11
2020
-
[21]
Training deep nets with sublinear memory cost, 2016
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost, 2016. 4
2016
-
[22]
Learning to estimate shapley values with vision transformers
Ian Covert, Chanwoo Kim, and Su-In Lee. Learning to estimate shapley values with vision transformers. ArXiv, abs/2206.05282, 2022. 6, 12
2022 arXiv
-
[23]
Atman: Understanding transformer predictions through memory efficient attention manipulation
Mayukh Deb, Bj ¨orn Deiseroth, Samuel Weinbach, Patrick Schramowski, and Kristian Kersting. Atman: Understanding transformer predictions through memory efficient attention manipulation. CoRR, abs/2301.08110, 2023. 115
2023 arXiv
-
[24]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. 5, 6, 11
2009
-
[25]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[26]
Use hirescam in- stead of grad-cam for faithful explanations of convolutional neural networks
Rachel Lea Draelos and Lawrence Carin. Use hirescam in- stead of grad-cam for faithful explanations of convolutional neural networks. 2020. 2, 113
2020
-
[27]
Explain to not forget: Defending against catas- trophic forgetting with xai
Sami Ede, Serop Baghdadlian, Leander Weber, An Thai Nguyen, Dario Zanca, Wojciech Samek, and Sebastian La- 9 puschkin. Explain to not forget: Defending against catas- trophic forgetting with xai. In International Cross-Domain Conference on Machine Learning and Knowledge Extrac...
2022
-
[28]
Eva: Exploring the limits of masked visual representation learning at scale
Yuxin Fang, Wen Wang, Binhui Xie, Quan-Sen Sun, Ledell Yu Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva: Exploring the limits of masked visual representation learning at scale. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages ...
2023
-
[29]
Eva-02: A visual representation for neon genesis
Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xin- long Wang, and Yue Cao. Eva-02: A visual representation for neon genesis. ArXiv, abs/2303.11331, 2023. 5
2023 arXiv
-
[30]
Adaptive token sampling for efficient vision transformers
Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari, Sunando Sengupta, Hamid Reza Vaezi Joze, Eric Sommerlade, Hamed Pirsiavash, and Juergen Gall. Adaptive token sampling for efficient vision transformers. In European Conference on Computer Vision,
-
[31]
Craft: Concept recursive activation factoriza- tion for explainability
Thomas Fel, Agustin Picard, Louis B ´ethune, Thibaut Boissin, David Vigouroux, Julien Colin, R’emi Cadene, and Thomas Serre. Craft: Concept recursive activation factoriza- tion for explainability. 2023 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pa...
2023
-
[32]
G ´allego, and Marta R
Javier Ferrando, Gerard I. G ´allego, and Marta R. Costa- juss`a. Measuring the mixing of contextual information in the transformer. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages 8698–8714, Abu Dhabi, United Arab Emirates, 20...
2022
-
[33]
Axiom-based grad-cam: To- wards accurate visualization and explanation of cnns
Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yulan Guo, Yinghui Gao, and Biao Li. Axiom-based grad-cam: To- wards accurate visualization and explanation of cnns. ArXiv, abs/2008.02312, 2020. 2, 113
2008 arXiv
-
[34]
Shangqi Gao, Zhong-Yu Li, Ming-Hsuan Yang, Mingg-Ming Cheng, Junwei Han, and Philip H. S. Torr. Large-scale un- supervised semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence , 45:7457–7476,
-
[35]
Self-attention at- tribution: Interpreting information interactions inside trans- former
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. Self-attention at- tribution: Interpreting information interactions inside trans- former. In AAAI Conference on Artificial Intelligence, 2020. 2
2020
-
[36]
Benchmarking neu- ral network robustness to common corruptions and perturba- tions
Dan Hendrycks and Thomas Dietterich. Benchmarking neu- ral network robustness to common corruptions and perturba- tions. Proceedings of the International Conference on Learn- ing Representations, 2019. 6
2019
-
[37]
The many faces of robustness: A critical analysis of out-of-distribution generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kada- vath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. ICC...
2021
-
[38]
Natural adversarial examples.CVPR,
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Stein- hardt, and Dawn Song. Natural adversarial examples.CVPR,
-
[39]
Transferable ad- versarial attack based on integrated gradients
Yi Huang and Adams Wai-Kin Kong. Transferable ad- versarial attack based on integrated gradients. ArXiv, abs/2205.13152, 2022. 113
2022 arXiv
-
[40]
Ex- plaining convolutional neural networks using softmax gradi- ent layer-wise relevance propagation
Brian Kenji Iwana, Ryohei Kuroki, and Seiichi Uchida. Ex- plaining convolutional neural networks using softmax gradi- ent layer-wise relevance propagation. 2019 IEEE/CVF In- ternational Conference on Computer Vision Workshop (IC- CVW), pages 4176–4185, 2019. 13
2019
-
[41]
Layercam: Exploring hierarchical class activation maps for localization
Peng-Tao Jiang, Chang-Bin Zhang, Qibin Hou, Ming-Ming Cheng, and Yunchao Wei. Layercam: Exploring hierarchical class activation maps for localization. IEEE Transactions on Image Processing, 30:5875–5888, 2021. 2
2021
-
[42]
Dense text-to-image generation with attention modulation
Yunji Kim, Jiyoung Lee, Jin-Hwa Kim, Jung-Woo Ha, and Jun-Yan Zhu. Dense text-to-image generation with attention modulation. ArXiv, abs/2308.12964, 2023. 113
2023 arXiv
-
[43]
Investigating the influence of noise and distractors on the interpretation of neural networks
Pieter-Jan Kindermans, Kristof Sch ¨utt, Klaus-Robert M¨uller, and Sven D ¨ahne. Investigating the influence of noise and distractors on the interpretation of neural networks. CoRR, abs/1611.07270, 2016. 113
2016 arXiv
-
[44]
Attention is not only a weight: Analyzing trans- formers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Ken- taro Inui. Attention is not only a weight: Analyzing trans- formers with vector norms. In Proceedings of the 2020 Con- ference on Empirical Methods in Natural Language Process- ing (EMNLP), pages 7057–7075, Online, 2020....
2020
-
[45]
Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Ken- taro Inui. Incorporating Residual and Normalization Layers into Analysis of Masked Language Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4547–4568, Online and P...
2021
-
[46]
Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014. 7, 13
2014
-
[47]
Towards faithful model explanation in nlp: A survey
QING LYU, Marianna Apidianaki, and Chris Callison- Burch. Towards faithful model explanation in nlp: A survey. ArXiv, abs/2209.11326, 2022. 1, 113
2022 arXiv
-
[48]
Andreas Madsen, Siva Reddy, and A. P. Sarath Chandar. Post-hoc interpretability for neural nlp: A survey. ACM Computing Surveys, 55:1 – 42, 2021. 1, 113
2021
-
[49]
SkipPLUS: Skip the first few layers to better explain vision transform- ers
Faridoun Mehri, Mohsen Fayyaz, Mahdieh Soleymani Baghshah, and Mohammad Taher Pilehvar. SkipPLUS: Skip the first few layers to better explain vision transform- ers. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 204–215,
2024
-
[50]
GlobEnc: Quantifying global token attribution by incorporating the whole encoder layer in transformers
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, and Mohammad Taher Pilehvar. GlobEnc: Quantifying global token attribution by incorporating the whole encoder layer in transformers. In Proceedings of the 2022 Confer- ence of the North American Chapter of the Association f...
2022
-
[51]
Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar
A. Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar. Adapler: Speeding up inference by adaptive length reduction. In Annual Meeting of the Association for Compu- tational Linguistics, 2022. 113
2022
-
[52]
De- compX: Explaining transformers decisions by propagating token decomposition
Ali Modarressi, Mohsen Fayyaz, Ehsan Aghazadeh, Yadol- lah Yaghoobzadeh, and Mohammad Taher Pilehvar. De- compX: Explaining transformers decisions by propagating token decomposition. In Proceedings of the 61st An- nual Meeting of the Association for Computational Linguis- tics...
2023
-
[53]
Ex- plaining nonlinear classification decisions with deep taylor decomposition
Gr ´egoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert M ¨uller. Ex- plaining nonlinear classification decisions with deep taylor decomposition. Pattern Recogn., 65(C):211–222, 2017. 114
2017
-
[54]
Comparing automatic and human evaluation of local explanations for text classification
Dong Nguyen. Comparing automatic and human evaluation of local explanations for text classification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan- guage Technologies, Volume 1 (Long Papers), pag...
2018
-
[55]
Decompose-and-compose: A composi- tional approach to mitigating spurious correlation
Fahimeh Hosseini Noohdani, Parsa Hosseini, Arian Yaz- dan Parast, Hamidreza Yaghoubi Araghi, and Mahdieh So- leymani Baghshah. Decompose-and-compose: A composi- tional approach to mitigating spurious correlation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision a...
2024
-
[56]
Making sense of dependence: Efficient black-box explanations using dependence measure
Paul Novello, Thomas Fel, and David Vigouroux. Making sense of dependence: Efficient black-box explanations using dependence measure. ArXiv, abs/2206.06219, 2022. 115
2022 arXiv
-
[57]
No token left be- hind: Explainability-aided image classification and genera- tion
Roni Paiss, Hila Chefer, and Lior Wolf. No token left be- hind: Explainability-aided image classification and genera- tion. In Computer Vision – ECCV 2022 , pages 334–350, Cham, 2022. Springer Nature Switzerland. 113
2022
-
[58]
Parkhi, Andrea Vedaldi, Andrew Zisserman, and C
Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V . Jawahar. Cats and dogs. InIEEE Conference on Com- puter Vision and Pattern Recognition, 2012. 6
2012
-
[59]
Beit v2: Masked image modeling with vector-quantized visual tokenizers
Zhiliang Peng, Li Dong, Hangbo Bao, Qixiang Ye, and Furu Wei. Beit v2: Masked image modeling with vector-quantized visual tokenizers. ArXiv, abs/2208.06366, 2022. 5
2022 arXiv
-
[60]
Rise: Random- ized input sampling for explanation of black-box models
Vitali Petsiuk, Abir Das, and Kate Saenko. Rise: Random- ized input sampling for explanation of black-box models. ArXiv, abs/1806.07421, 2018. 115
2018 arXiv
-
[61]
AttCAT: Explaining transformers via attentive class activation tokens
Yao Qiang, Deng Pan, Chengyin Li, Xin Li, Rhongho Jang, and Dongxiao Zhu. AttCAT: Explaining transformers via attentive class activation tokens. In Advances in Neural In- formation Processing Systems, 2022. 2, 114
2022
-
[62]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...
2021
-
[63]
Irvin, Aarti Bagul, Daisy Yi Ding, Tony Duan, Hershel Mehta, Brandon Yang, Kaylie Zhu, Dillon Laird, Robyn L
Pranav Rajpurkar, Jeremy A. Irvin, Aarti Bagul, Daisy Yi Ding, Tony Duan, Hershel Mehta, Brandon Yang, Kaylie Zhu, Dillon Laird, Robyn L. Ball, C. Langlotz, Katie S. Sh- panskaya, Matthew P. Lungren, and A. Ng. Mura dataset: Towards radiologist-level abnormality detection in m...
2017 arXiv
-
[64]
Do imagenet classifiers generalize to im- agenet? In International Conference on Machine Learning,
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to im- agenet? In International Conference on Machine Learning,
-
[65]
why should i trust you?
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. “why should i trust you?”: Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD Interna- tional Conference on Knowledge Discovery and Data Min- ing, 2016. 115
2016
-
[66]
Anders, and Klaus-Robert M ¨uller
Wojciech Samek, Gr ´egoire Montavon, Sebastian La- puschkin, Christopher J. Anders, and Klaus-Robert M ¨uller. Explaining deep neural networks and beyond: A review of methods and applications. Proceedings of the IEEE , 109: 247–278, 2021. 1, 113
2021
-
[67]
Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In 2017 IEEE International Conference on Computer Vision (ICCV) , pages 618–626,
2017
-
[68]
Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Dhruv Batra, and Devi Parikh
Ramprasaath R. Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Dhruv Batra, and Devi Parikh. Taking a hint: Leverag- ing explanations to make vision and language models more grounded. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 2591–2600, 2019. 113
2019
-
[69]
Noam M. Shazeer. Glu variants improve transformer. ArXiv, abs/2002.05202, 2020. 3
2002 arXiv
-
[70]
Pami: partition input and aggregate outputs for model inter- pretation
Wei Shi, Wentao Zhang, Weishi Zheng, and Ruixuan Wang. Pami: partition input and aggregate outputs for model inter- pretation. ArXiv, abs/2302.03318, 2023. 115
2023 arXiv
-
[71]
Not just a black box: Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje. Not just a black box: Learning important features through propagating activation differences. ArXiv, abs/1605.01713, 2016. 2, 113
2016 arXiv
-
[72]
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. Learning important features through propagating activation differences. In International Conference on Machine Learn- ing, 2017. 2, 113
2017
-
[73]
Deep inside convolutional networks: Visualising image clas- sification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: Visualising image clas- sification models and saliency maps. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Workshop Track P...
2014
-
[74]
Riedmiller
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin A. Riedmiller. Striving for simplicity: The all convolutional net. CoRR, abs/1412.6806, 2014. 113
2014 arXiv
-
[75]
Full-gradient represen- tation for neural network visualization
Suraj Srinivas and Franc ¸ois Fleuret. Full-gradient represen- tation for neural network visualization. In Neural Informa- tion Processing Systems, 2019. 2, 4, 113
2019
-
[76]
Eva-clip: Improved training techniques for clip at scale
Quan Sun, Yuxin Fang, Ledell Yu Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. ArXiv, abs/2303.15389, 2023. 5 11
2023 arXiv
-
[77]
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In Proceedings of the 34th In- ternational Conference on Machine Learning , pages 3319–
-
[78]
Imagenet-hard: The hard- est images remaining from a study of the power of zoom and spatial biases in image classification
Mohammad Reza Taesiri, Giang Nguyen, Sarra Habchi, Cor- Paul Bezemer, and Anh Nguyen. Imagenet-hard: The hard- est images remaining from a study of the power of zoom and spatial biases in image classification. 2023. 6
2023
-
[79]
Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lu- cas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy
Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lu- cas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. Mlp-mixer: An all-mlp architecture for vision. In Neural Information Processing Systems, 2021. 5
2021
-
[80]
Train- ing data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv’e J’egou. Train- ing data-efficient image transformers & distillation through attention. ArXiv, abs/2012.12877, 2020. 5
2012 arXiv
-
[81]
Deit iii: Revenge of the vit
Hugo Touvron, Matthieu Cord, and Herv’e J’egou. Deit iii: Revenge of the vit. In European Conference on Computer Vision, 2022. 5
2022
-
[82]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neu- ral Information Processing Systems. Curran Associates, Inc.,
-
[83]
Analyzing multi-head self-attention: Spe- cialized heads do the heavy lifting, the rest can be pruned
Elena V oita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self-attention: Spe- cialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 57...
2019
-
[84]
Learning robust global representations by penalizing local predictive power
Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. In Advances in Neural Information Processing Systems, pages 10506–10518, 2019. 6
2019
-
[85]
Score-cam: Score-weighted visual explanations for convolutional neural networks
Haofan Wang, Zifan Wang, Mengnan Du, Fan Yang, Zi- jian Zhang, Sirui Ding, Piotr (Peter) Mardziel, and Xia Hu. Score-cam: Score-weighted visual explanations for convolutional neural networks. 2020 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition Work- shops (CV...
2020
-
[86]
Beyond explaining: Opportunities and challenges of xai-based model improvement.Inf
Leander Weber, Sebastian Lapuschkin, Alexander Binder, and Wojciech Samek. Beyond explaining: Opportunities and challenges of xai-based model improvement.Inf. Fusion, 92: 154–176, 2022. 113
2022
-
[87]
Token transformation matters: Towards faithful post-hoc ex- planation for vision transformer
Junyi Wu, Bin Duan, Weitai Kang, Hao Tang, and Yan Yan. Token transformation matters: Towards faithful post-hoc ex- planation for vision transformer. 2024 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 10926–10935, 2024. 2, 5, 6, 11, 12, 114
2024
-
[88]
Lyu, and Yu-Wing Tai
Weibin Wu, Yuxin Su, Xixian Chen, Shenglin Zhao, Ir- win King, Michael R. Lyu, and Yu-Wing Tai. Boost- ing the transferability of adversarial samples via attention. 2020 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 1158–1167, 2020. 113
2020
-
[89]
Vit-cx: Causal explanation of vision transformers
Weiyan Xie, Xiao hui Li, Caleb Chen Cao, and Nevin L.Zhang. Vit-cx: Causal explanation of vision transformers. In International Joint Conference on Artificial Intelligence ,
-
[90]
Puyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane ling Wang, and Michael I. Jordan. Ml-loo: Detecting adversarial exam- ples with feature attribution. In AAAI Conference on Artifi- cial Intelligence, 2019. 113
2019
-
[91]
Mambaout: Do we really need mamba for vision? arXiv preprint arXiv:2405.07992,
Weihao Yu and Xinchao Wang. Mambaout: Do we really need mamba for vision? arXiv preprint arXiv:2405.07992,
-
[92]
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. ArXiv, abs/2303.15343, 2023. 5
2023 arXiv
-
[93]
Lin, Jonathan Brandt, Xiaohui Shen, and Stan Sclaroff
Jianming Zhang, Zhe L. Lin, Jonathan Brandt, Xiaohui Shen, and Stan Sclaroff. Top-down neural attention by excitation backprop. International Journal of Computer Vision , 126: 1084–1102, 2016. 113
2016
-
[94]
Jianping Zhang, Weibin Wu, Jen tse Huang, Yizhan Huang, Wenxuan Wang, Yuxin Su, and Michael R. Lyu. Improv- ing adversarial transferability via neuron attribution-based attacks. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14973–14982, 2022...
2022
-
[97]
Via duality: SwapBackward (f, g)(x) = f (x).detach() + (g(x) − g(x).detach())
-
[98]
Remark 3 (Computational Efficiency)
Via custom backward: Define an autograd.Function that returns f (x) in forward and propagates gradients as if it were g(x) in backward Both implementations yield identical gradients, though the latter may be more computationally efficient, while the former may be easier to imp...
-
[99]
⊙ f2(x) where [·]cst
When a = 0 , yielding f (x) = [ f1(x)]cst. ⊙ f2(x) where [·]cst. is the constant operator that zeroes gradients, f is FG- complete if f2 is FG-complete
-
[100]
By symmetry, when b = 0, f is FG-complete if f1 is FG-complete. 6 Proof. Let a = 0 (thus b = 1). If f2 is FG-complete: [diag(f1(x)) · Jxf2] · x + X i [diag(f1(x)) · Jbi f2] · bi = diag(f1(x)) · (Jxf2 · x + X i Jbi f2 · bi) = diag(f1(x)) · f2(x) = f1(x) ⊙ f2(x) = f (x) proving ...
-
[101]
Centering: y = x − µ1, where 1 is the vector of ones
-
[102]
Zebra” and “African Elephant
Scaling: z = y/s, where s = √ σ2 + ε The Jacobian of centering is: (Jxy)ij = δij − 1 N which gives (Jxy · x)i = xi − µ = yi. The Jacobian of scaling is: (Jyz)ij = δij s − yiyj N s3 By the chain rule: JxLN · x = Jyz · Jxy · x = Jyz · y Computing (Jyz · y)i: (Jyz · y)i = NX j=1 ...
-
[103]
Negative Value Removal: We first apply ReLU to remove negative attribution scores, as we focus on positive feature contributions
-
[104]
We then scale the values by dividing by this robust maximum
Robust Scaling: Rather than using absolute maximum values which can be sensitive to outliers, we compute the 99th percentile of the attribution scores. We then scale the values by dividing by this robust maximum
-
[105]
Spatial Upsampling: The token-level attribution map is upsampled to the original image resolution using bicubic inter- polation
-
[106]
Range Normalization: Finally, we clamp values to [0, 1]. 13 C. Qualitative Results Following the evaluation protocol in Appendix B.4, we present a comprehensive qualitative analysis below. C.1. Text-Prompted Qualitative Examples on EV A2-CLIP-Large Our first evaluation scenari...
-
[107]
GlobEnc & ALTI
extends AttIN to also incorporate the residual connec- tions. GlobEnc & ALTI. AttIN assumes that tokens retain their original identity. As each self-attention module mixes all the tokens, this assumption might not necessarily hold. Us- ing gradient-based techniques, Brunner et...
-
[2024]
2, 5, 6, 11, 12, 13, 113
-
[3328]
2, 4, 6, 113
PMLR, 2017. 2, 4, 6, 113
2017
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.