Pith. sign in

REVIEW 4 major objections 3 minor 107 references

LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions

T0 review · 4 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Pruning and rescaling backward gradients restores an exact-output-completeness property in Transformers, improving every gradient attribution method tested.

desk verdict A genuinely useful gradient-surgery method for ViT attributions with broad empirical gains, but the 'universally enhances' claim is contradicted by its own tables and the completeness theory is a conservation identity for a surrogate gradient, not a faithfulness guarantee. read the letter →

arxiv 2411.16760 v1 pith:WUONYQSD submitted 2024-11-24 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords LibraGradgradientattributionTransformerinterpretabilityFullGrad-completenessfaithfulnessmetricsLayerNormattentionattributionspost-hocexplanation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Gradient-based attribution methods that work well on CNNs fail on Vision Transformers, and this paper says the reason is not attention itself but unbalanced gradient flow in the backward pass. It formalizes a property, FullGrad-completeness, that makes attributions decompose a model's output exactly, shows that classical CNNs have it naturally while Transformer components break it, and introduces LibraGrad, a set of backward-pass modifications that prune and rescale gradients by treating attention weights and LayerNorm denominators as constants and splitting self-gating gradients. Across eight architectures, four model sizes, and four datasets, the modified gradients improve every gradient-based attribution method tested, including hybrids designed for Transformers, and the general-purpose Libra FullGrad+ becomes the best overall method. The package requires no forward-pass changes and no extra compute, which matters because faithful explanations of Transformer decisions are a prerequisite for deploying these models in high-stakes settings.

What carries the argument

The central object is FullGrad-completeness (FG-completeness), the identity $f(x) = J_x f \cdot x + \sum_i J_{b_i} f \cdot b_i$ that forces attributions to sum exactly to the model output. The argument is carried by two backward-pass manipulation operators defined on top of an unchanged forward pass: the constant operator $[\cdot]_{\mathrm{cst}}$, which zeroes the gradient while keeping the value, and SwapBackward, which evaluates one function but propagates the Jacobian of another. Theorems 1–3 show that locally affine functions, compositions, and sums are FG-complete, which explains why CNNs inherit the property, and Propositions 1–3 show which Transformer operations break it; Theorems 4–5 then give the repair rules, scale product branches by coefficients summing to one or prune a non-FG branch. The practical Libra modules instantiate these rules: attention contributes gradients only through the value branch, LayerNorm treats its denominator as constant, gated activations discard the gate's gradient, and SwiGLU splits its backward gradient equally between the two branches.

What would settle it

Take a Transformer fine-tuned on ImageNet and compute attributions on the paper's own zebra/elephant COCO images with only the attention-gradient pruning applied and LayerNorm gradients untouched: if the pruned attention scores, rather than LayerNorm, encode the co-occurring-class distinction, the maps will localize the wrong animal. Conversely, the theory predicts that a Transformer whose normalizations are replaced by affine scaling needs no Libra LayerNorm fix, so a faithfulness gain there would disprove FG-completeness as the operative mechanism.

Watch

Extended reading notes

Core claim

LibraGrad claims that the reason gradient-based explanations underperform on Transformers is a set of non-locally-affine operations — attention softmax scores, gated activations such as GELU and SiLU, SwiGLU self-gating, and LayerNorm — that violate FullGrad-completeness, the identity $f(x) = J_x f \cdot x + \sum_i J_{b_i} f \cdot b_i$ that CNNs satisfy automatically. The paper proves that naive element-wise multiplication over-counts attributions by a factor of two, that division makes FullGrad vanish identically, and that LayerNorm's denominator drives FullGrad to zero as $\varepsilon \to 0$. It then proves that scaling the two Jacobian branches of a product by coefficients summing to one, or pruning one branch entirely by treating it as constant in the backward pass, restores the identity (Theorems 4 and 5, Corollary 2), and that applying these fixes to attention, gated activations, self-gating, and LayerNorm yields a Transformer that is FG-complete (Corollary 3). Empirically, Libra versions of Input×Grad, GradCAM+, HiResCAM, XGradCAM+, FullGrad+, AttCAT, GenAtt, and TokenTM beat their unmodified counterparts on faithfulness, completeness error, and segmentation alignment on nearly every model and dataset, with Libra FullGrad+ at 0.0 completeness error on all tested models and the top average segmentation AP.

Load-bearing premise

The load-bearing premise is that making the gradient-sensitive parts of a Transformer behave as constants in the backward pass, which guarantees attributions sum to the output, is enough to make attributions faithful, and that the discarded gradient signal through attention weights and LayerNorm denominators carries no essential explanation.

Editorial extensions

If this is right

  • Once gradients are rebalanced, general-purpose gradient methods such as FullGrad+ outperform attention-based and Transformer-specific attribution methods on faithfulness and segmentation metrics, suggesting that specialized explainers for Transformers are unnecessary.
  • Libra FullGrad+ reaches 0.0 completeness error on every large model tested and raises average segmentation AP from 43.4 to 67.9, placing it above every attention-based method in the comparison.
  • LibraGrad imposes no forward-pass modification and no extra computational or memory overhead, so it can be dropped into any existing gradient pipeline.
  • The method transfers to the attention-free MLP-Mixer architecture, supporting the claim that gradient imbalance, not attention per se, is the root cause of attribution failures.
  • The ablation study attributes most of the gain to fixing LayerNorm, with attention and self-gating secondary and bias terms negligible, identifying the highest-value repair for practitioners.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The argument suggests a general design rule for future architectures: whenever a new layer introduces multiplicative gating or normalization, the backward pass should be audited against FG-completeness before any attribution method is built, and the Libra recipes could be applied preemptively.
  • A testable prediction follows: on a Transformer variant whose normalizations are already affine, the gain from the Libra LayerNorm fix should shrink, cleanly separating the normalization effect from the attention effect.
  • The pruning choice is a modeling decision, and the scaling coefficients of Theorem 4 could be tuned to interpolate between full gradients and pruned gradients, trading exact completeness against retaining some attention-gradient signal; the paper does not explore this.
  • If FG-completeness is the operative mechanism, faithfulness gains should correlate with measured completeness error across architectures and checkpoints, a correlation the paper reports only indirectly.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper proposes LibraGrad, a post-hoc attribution enhancement for vision transformers that modifies backward gradients by zeroing gradients through attention softmax scores and LayerNorm denominators, pruning non-linear gates, and scaling self-gating branches. It claims that these changes restore FullGrad-completeness, that this is crucial for faithful interpretability, and that LibraGrad universally enhances a wide range of gradient-based attribution methods across faithfulness, completeness error, and segmentation metrics. The experiments cover eight architectures, four model sizes, and four datasets, with code released.

Significance. If the universal-improvement claim held, the contribution would be significant: it offers a cheap, drop-in gradient surgery applicable to IG, FullGrad+, GradCAM+, and other methods, and it poses a clear hypothesis about why gradients fail on transformers. The paper's theoretical toolbox (constant operator, SwapBackward, Theorems 4 and 5) is clean and reusable, and the empirical sweep is unusually broad, spanning multiple architectures, model sizes, datasets, and metric families. However, as detailed below, the theoretical guarantee applies to modified backward Jacobians rather than to the model's true Jacobians, and the 'universal' claim is contradicted by the paper's own tables. The method remains a plausible strong heuristic; its significance depends on an honest reframing and on reporting the negative cases.

major comments (4)
  1. [§3.4, Definition 1, Corollary 3] FG-completeness is stated for the true Jacobians of f, but LibraGrad's definitions replace those Jacobians with modified backward values. For example, Libra-LayerNorm keeps the forward value but zeroes the gradient through the denominator, and Libra-Attention zeroes gradients through softmax(QK^T); these are not the Jacobians of the forward function. Corollary 3 therefore certifies that a modified backward pass satisfies f(x)=J^Libra_x f·x+Σ J^Libra_bi f·bi, which is enforced by construction rather than derived from the model. Table 4's CE=0 (Libra FullGrad) is an implementation sanity check of that custom backward pass, as Section B.1 acknowledges, and it cannot by itself support the abstract's 'theoretically grounded' faithfulness claim. The divergence is concrete: Proposition 3 shows the true JxLN·x tends to zero as epsilon approaches zero, while the Libra backward pass returns LN(x). The paper should either prove a formal connection between modified-Jacobian completeness and the faithfulness metrics or explicitly reframe the contribution as a heuristic gradient reweighting with empirical support.
  2. [Tables 2 and 3; §4.2] The claim of universal enhancement across all metrics is not supported by the paper's own results. In Table 2 (MIF accuracy on ViT-B), Libra Input×Grad drops on MURA from 25.5 to 21.6, and Libra TokenTM is unchanged at 28.0. In Table 3 (Segmentation AP), Libra GradCAM+ drops on SigLIP-L from 44.3±0.4 to 41.7±0.3 and on DeiT3-H from 60.3±0.4 to 46.7±0.4; similar decreases appear in other cells. Section 4.2's statement that 'LibraGrad universally enhances gradient-based attribution methods across all tested models, architectures, and datasets' is therefore false as written. The authors should present a systematic account of cases with no gain or a loss, and the abstract and conclusion should be revised to describe large but not universal gains.
  3. [§3.4–§3.5] The choice of scaling versus pruning coefficients is underdetermined by the theory and not justified empirically. Theorem 4 allows any a,b with a+b=1, and the paper fixes a=b=1/2 for self-gating while setting a=0 (pruning) for attention and LayerNorm; Theorem 5 justifies pruning only when the surviving branch is FG-complete. The decision to zero Q and K gradients rather than scale them is presented as a design choice in Section 3.5 with no ablation or sensitivity analysis over a,b or over pruning versus scaling for any component. Since these choices determine the attribution maps, the paper should at least report an ablation over the free coefficients and over scaling alternatives for the attention path.
  4. [Table 5] The ablation's interaction pattern complicates the 'LayerNorm is the most significant factor' interpretation. Removing LayerNorm alone drops MIF predicted accuracy from 71.7 to 49.9, but removing both attention and LayerNorm gives 61.2, which is higher than removing LayerNorm alone; similar nonmonotonicity appears in the GT-accuracy column (65.2 vs 49.9 vs 63.6). The text in Section 4.2 asserts the LayerNorm effect is the most significant without discussing this interaction. The authors should explain the non-additivity or soften the causal attribution.
minor comments (3)
  1. [§4.1] Table 2 reports that standard deviations were bounded by 0.1 and omitted, but it does not state the number of random seeds or images used for each cell; a single sentence on replication would help.
  2. [Table 4 caption] The Completeness Error table excludes attention-based methods 'as their incompleteness is evident'; this should be stated as a scope restriction in the main text so that the comparison is not read as a complete ranking.
  3. [§3.5] The expression 'Libra-Attention(Q, K, V) = [softmax(QK^T)]_cst · V' has ambiguous operator precedence; adding parentheses would clarify that the constant operator applies to the whole softmax argument.

Circularity Check

1 steps flagged · score 2.0 of 10

No significant circularity: the only self-referential element is the acknowledged by-construction FG-completeness of Libra's modified backward pass; the main empirical claims rest on external metrics.

  1. self definitional [Definition 1 (Section 3); Section 3.5; Corollary 3; Table 4; Appendix B.1]
    "Definition 1: "A function f is FullGrad-complete (or FG-complete) if, for all x ∈ Rn, f (x) = Jxf · x + X i Jbi f · bi" ... Section 3.5: "Libra-Attention(Q, K, V) = [softmax(QK T )]cst. · V", "Libra-LayerNorm(x) = x − µ / [√(σ2 + ε)]cst." ... Corollary 3: "A Transformer architecture attains FG-completeness when all non-linear components—specifically its attention mechanisms, activation functions, self-gating operations, and LayerNorms—are replaced with their Libra counterparts.""

    Definition 1 defines FG-completeness via true Jacobians Jxf=∂f/∂x and Jbi f. LibraGrad replaces these by definition: [·]cst. has Jx[y]cst.=0, and Libra-Attention / Libra-LayerNorm detach softmax(QK^T) and the LayerNorm denominator. Corollary 3 thus holds because the Libra backward pass was constructed so f(x)=J^Libra_x f·x+Σ J^Libra_bi f·b_i is satisfied by construction, not because the model's true gradients satisfy it. Table 4's CE=0.0 for Libra FullGrad is an implementation check of the custom backward pass; the paper calls CE 'just a sanity check' (Appendix B.1), so the self-reference is acknowledged and the main faithfulness/segmentation claims do not depend on it.

full rationale

The paper's main claims are empirical: faithfulness deletions, ImageNet-S segmentation AP, and qualitative CLIP/co-occurring-class comparisons are all external benchmark measurements of the produced attribution maps, not functions of the LibraGrad definition itself. The FG-completeness theorem is a genuine mathematical statement but only about the modified backward graph, which makes the CE=0 table a by-construction sanity check rather than evidence about the original model's gradients; the paper labels it as such, so this is a minor acknowledged self-reference rather than a load-bearing circularity. No load-bearing self-citation is present: FullGrad+ (ref. [49], by overlapping authors) is used as a baseline/component, not as the argument for LibraGrad's improvement. The paper's own tables contain counterexamples to the word 'universally' (e.g., MURA Input×Grad 25.5→21.6 in Table 2; SigLIP-L GradCAM+ 44.3→41.7 in Table 3) and the link from modified-backward FG-completeness to faithfulness is asserted rather than proven, but those are correctness/completeness concerns, not circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the assumption that modified-backward attributions are valid explanations, and on the choice to prune rather than scale certain gradient paths. There is one numeric hand-picked coefficient (1/2) and two structural design choices that are not ablated. No new physical entities are introduced.

free parameters (1)
  • Self-gate gradient scaling coefficient (a=b) = 0.5
    In Libra-SelfGate, both branches are scaled by 1/2. Any a+b=1 preserves FG-completeness per Theorem 4, but the attribution scores change with a; the choice of 0.5 is not derived from first principles or empirically ablated against other values.
assumptions (4)
  • domain assumption FG-completeness with respect to modified backward Jacobians is a critical property for attribution faithfulness.
    Stated in Section 3 ('This is crucial for faithful interpretability'); used to motivate LibraGrad, but no theorem links FG-completeness to the faithfulness metrics. The paper relies on empirical evidence for this link.
  • domain assumption Modifying backward gradients while leaving the forward pass unchanged yields valid attributions of the original model.
    Section 3.4 introduces the Constant Operator and SwapBackward, implicitly assuming that explanations computed with non-true gradients still explain the forward function. This is the load-bearing premise behind all Libra operations.
  • ad hoc to paper Pruning attention score gradients (Q and K) is preferable to scaling them.
    Section 3.5 defines Libra-Attention with softmax treated as constant. The pruning choice is motivated by softmax's non-FG-completeness, but alternative scaling of both paths (Theorem 4) is not empirically compared, leaving a design decision untested.
  • standard math Element-wise product and quotient Jacobians follow the standard chain rule for the specified branch structure.
    Propositions 1, 2 and Theorem 4 rely on standard differential calculus for products and quotients; this is uncontroversial background.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions." pith.science (2026). https://pith.science/paper/WUONYQSD

@misc{pith2026241116760,
  author       = {Pith},
  title        = {Pith review of: LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WUONYQSD}},
  note         = {Machine review of arXiv:2411.16760}
}
read the original abstract

Why do gradient-based explanations struggle with Transformers, and how can we improve them? We identify gradient flow imbalances in Transformers that violate FullGrad-completeness, a critical property for attribution faithfulness that CNNs naturally possess. To address this issue, we introduce LibraGrad -- a theoretically grounded post-hoc approach that corrects gradient imbalances through pruning and scaling of backward paths, without changing the forward pass or adding computational overhead. We evaluate LibraGrad using three metric families: Faithfulness, which quantifies prediction changes under perturbations of the most and least relevant features; Completeness Error, which measures attribution conservation relative to model outputs; and Segmentation AP, which assesses alignment with human perception. Extensive experiments across 8 architectures, 4 model sizes, and 4 datasets show that LibraGrad universally enhances gradient-based methods, outperforming existing white-box methods -- including Transformer-specific approaches -- across all metrics. We demonstrate superior qualitative results through two complementary evaluations: precise text-prompted region highlighting on CLIP models and accurate class discrimination between co-occurring animals on ImageNet-finetuned models -- two settings on which existing methods often struggle. LibraGrad is effective even on the attention-free MLP-Mixer architecture, indicating potential for extension to other modern architectures. Our code is freely available at https://github.com/NightMachinery/LibraGrad.

Figures

Figures reproduced from arXiv: 2411.16760 by the authors.

Figure 1
Figure 1. Qualitative comparison on EVA2-CLIP-Large. Our pro [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Cross-method comparison of class discriminativity on ViT-B. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

107 extracted references · 56 canonical work pages

  1. [1]

    Quantifying attention flow in transformers

    Samira Abnar and Willem Zuidema. Quantifying attention flow in transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 4190–4197, Online, 2020. Association for Computa- tional Linguistics. 2, 114

  2. [2]

    AttnLRP: Attention- aware layer-wise relevance propagation for transformers

    Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebas- tian Lapuschkin, and Wojciech Samek. AttnLRP: Attention- aware layer-wise relevance propagation for transformers. In Proceedings of the 41st International Conference on Ma- chine Learning, pages 135–168. PMLR, 2024. 2, 4, 114

  3. [3]

    XAI for trans- formers: Better explanations through conservative propaga- tion

    Ameen Ali, Thomas Schnake, Oliver Eberle, Gr ´egoire Mon- tavon, Klaus-Robert M ¨uller, and Lior Wolf. XAI for trans- formers: Better explanations through conservative propaga- tion. In Proceedings of the 39th International Conference on Machine Learning, pages 435–451. PMLR, 2022. 2, 114

  4. [4]

    Marco Ancona, Enea Ceolini, Cengiz ¨Oztireli, and Markus H. Gross. Towards better understanding of gradient- based attribution methods for deep neural networks. In In- ternational Conference on Learning Representations , 2017. 2, 113

  5. [5]

    Anders, David Neumann, Talmaj Marinc, Wojciech Samek, Klaus-Robert M ¨uller, and Sebastian La- puschkin

    Christopher J. Anders, David Neumann, Talmaj Marinc, Wojciech Samek, Klaus-Robert M ¨uller, and Sebastian La- puschkin. Xai for analyzing and unlearning spurious corre- lations in imagenet. 2020. 113

  6. [6]

    On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation

    Sebastian Bach, Alexander Binder, Gr ´egoire Montavon, Frederick Klauschen, Klaus-Robert M ¨uller, and Wojciech Samek. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE, 10, 2015. 2

  7. [7]

    Beit: Bert pre-training of image transformers

    Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. ArXiv, abs/2106.08254, 2021. 5

  8. [8]

    Tenenbaum, and Boris Katz

    Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Joshua B. Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Neural Information Processing Systems, 2019. 6

Show all 107 references
  1. [9]

    Ecqx: Explainability-driven quantization for low-bit and sparse dnns

    Daniel Becking, Maximilian Dreyer, Wojciech Samek, Karsten M ¨uller, and Sebastian Lapuschkin. Ecqx: Explainability-driven quantization for low-bit and sparse dnns. ArXiv, abs/2109.04236, 2021. 113

  2. [10]

    H’enaff, Alexander Kolesnikov, Xi- aohua Zhai, and A ¨aron van den Oord

    Lucas Beyer, Olivier J. H’enaff, Alexander Kolesnikov, Xi- aohua Zhai, and A ¨aron van den Oord. Are we done with imagenet? ArXiv, abs/2006.07159, 2020. 6

  3. [11]

    Alabdulmohsin, and Filip Pavetic

    Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim M. Alabdulmohsin, and Filip Pavetic. Flexivit: One model for all patch sizes. 2023 IEEE/CVF Conference on Computer Vision and Pat- te...

  4. [12]

    Layer-wise rel- evance propagation for deep neural network architectures

    Alexander Binder, Sebastian Bach, Gr ´egoire Montavon, Klaus-Robert M¨uller, and Wojciech Samek. Layer-wise rel- evance propagation for deep neural network architectures

  5. [13]

    De- coupling pixel flipping and occlusion strategy for consistent xai benchmarks

    Stefan Bl ¨ucher, Johanna Vielhaben, and Nils Strodthoff. De- coupling pixel flipping and occlusion strategy for consistent xai benchmarks. Transactions on Machine Learning Re- search, 2024. 5, 11

  6. [14]

    On iden- tifiability in transformers

    Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. On iden- tifiability in transformers. In International Conference on Learning Representations, 2020. 114

  7. [15]

    Emerg- ing properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv’e J’egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9630–9640, 2021. 2

  8. [16]

    Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

    Hila Chefer, Shir Gur, and Lior Wolf. Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV) , pages 397–406, 2021. 2, 114

  9. [17]

    Transformer inter- pretability beyond attention visualization

    Hila Chefer, Shir Gur, and Lior Wolf. Transformer inter- pretability beyond attention visualization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 782–791, 2021. 2, 6, 12, 114

  10. [18]

    Optimizing relevance maps of vision transformers improves robustness

    Hila Chefer, Idan Schwartz, and Lior Wolf. Optimizing relevance maps of vision transformers improves robustness. ArXiv, abs/2206.01161, 2022. 113

  11. [19]

    Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

    Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models. ArXiv, abs/2301.13826, 2023. 113

  12. [20]

    Generative pre- training from pixels

    Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Hee- woo Jun, David Luan, and Ilya Sutskever. Generative pre- training from pixels. In Proceedings of the 37th Interna- tional Conference on Machine Learning , pages 1691–1703. PMLR, 2020. 5, 11

  13. [21]

    Training deep nets with sublinear memory cost, 2016

    Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost, 2016. 4

  14. [22]

    Learning to estimate shapley values with vision transformers

    Ian Covert, Chanwoo Kim, and Su-In Lee. Learning to estimate shapley values with vision transformers. ArXiv, abs/2206.05282, 2022. 6, 12

  15. [23]

    Atman: Understanding transformer predictions through memory efficient attention manipulation

    Mayukh Deb, Bj ¨orn Deiseroth, Samuel Weinbach, Patrick Schramowski, and Kristian Kersting. Atman: Understanding transformer predictions through memory efficient attention manipulation. CoRR, abs/2301.08110, 2023. 115

  16. [24]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. 5, 6, 11

  17. [25]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...

  18. [26]

    Use hirescam in- stead of grad-cam for faithful explanations of convolutional neural networks

    Rachel Lea Draelos and Lawrence Carin. Use hirescam in- stead of grad-cam for faithful explanations of convolutional neural networks. 2020. 2, 113

  19. [27]

    Explain to not forget: Defending against catas- trophic forgetting with xai

    Sami Ede, Serop Baghdadlian, Leander Weber, An Thai Nguyen, Dario Zanca, Wojciech Samek, and Sebastian La- 9 puschkin. Explain to not forget: Defending against catas- trophic forgetting with xai. In International Cross-Domain Conference on Machine Learning and Knowledge Extrac...

  20. [28]

    Eva: Exploring the limits of masked visual representation learning at scale

    Yuxin Fang, Wen Wang, Binhui Xie, Quan-Sen Sun, Ledell Yu Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva: Exploring the limits of masked visual representation learning at scale. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages ...

  21. [29]

    Eva-02: A visual representation for neon genesis

    Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xin- long Wang, and Yue Cao. Eva-02: A visual representation for neon genesis. ArXiv, abs/2303.11331, 2023. 5

  22. [30]

    Adaptive token sampling for efficient vision transformers

    Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari, Sunando Sengupta, Hamid Reza Vaezi Joze, Eric Sommerlade, Hamed Pirsiavash, and Juergen Gall. Adaptive token sampling for efficient vision transformers. In European Conference on Computer Vision,

  23. [31]

    Craft: Concept recursive activation factoriza- tion for explainability

    Thomas Fel, Agustin Picard, Louis B ´ethune, Thibaut Boissin, David Vigouroux, Julien Colin, R’emi Cadene, and Thomas Serre. Craft: Concept recursive activation factoriza- tion for explainability. 2023 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pa...

  24. [32]

    G ´allego, and Marta R

    Javier Ferrando, Gerard I. G ´allego, and Marta R. Costa- juss`a. Measuring the mixing of contextual information in the transformer. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages 8698–8714, Abu Dhabi, United Arab Emirates, 20...

  25. [33]

    Axiom-based grad-cam: To- wards accurate visualization and explanation of cnns

    Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yulan Guo, Yinghui Gao, and Biao Li. Axiom-based grad-cam: To- wards accurate visualization and explanation of cnns. ArXiv, abs/2008.02312, 2020. 2, 113

  26. [34]

    Shangqi Gao, Zhong-Yu Li, Ming-Hsuan Yang, Mingg-Ming Cheng, Junwei Han, and Philip H. S. Torr. Large-scale un- supervised semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence , 45:7457–7476,

  27. [35]

    Self-attention at- tribution: Interpreting information interactions inside trans- former

    Yaru Hao, Li Dong, Furu Wei, and Ke Xu. Self-attention at- tribution: Interpreting information interactions inside trans- former. In AAAI Conference on Artificial Intelligence, 2020. 2

  28. [36]

    Benchmarking neu- ral network robustness to common corruptions and perturba- tions

    Dan Hendrycks and Thomas Dietterich. Benchmarking neu- ral network robustness to common corruptions and perturba- tions. Proceedings of the International Conference on Learn- ing Representations, 2019. 6

  29. [37]

    The many faces of robustness: A critical analysis of out-of-distribution generalization

    Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kada- vath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. ICC...

  30. [38]

    Natural adversarial examples.CVPR,

    Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Stein- hardt, and Dawn Song. Natural adversarial examples.CVPR,

  31. [39]

    Transferable ad- versarial attack based on integrated gradients

    Yi Huang and Adams Wai-Kin Kong. Transferable ad- versarial attack based on integrated gradients. ArXiv, abs/2205.13152, 2022. 113

  32. [40]

    Ex- plaining convolutional neural networks using softmax gradi- ent layer-wise relevance propagation

    Brian Kenji Iwana, Ryohei Kuroki, and Seiichi Uchida. Ex- plaining convolutional neural networks using softmax gradi- ent layer-wise relevance propagation. 2019 IEEE/CVF In- ternational Conference on Computer Vision Workshop (IC- CVW), pages 4176–4185, 2019. 13

  33. [41]

    Layercam: Exploring hierarchical class activation maps for localization

    Peng-Tao Jiang, Chang-Bin Zhang, Qibin Hou, Ming-Ming Cheng, and Yunchao Wei. Layercam: Exploring hierarchical class activation maps for localization. IEEE Transactions on Image Processing, 30:5875–5888, 2021. 2

  34. [42]

    Dense text-to-image generation with attention modulation

    Yunji Kim, Jiyoung Lee, Jin-Hwa Kim, Jung-Woo Ha, and Jun-Yan Zhu. Dense text-to-image generation with attention modulation. ArXiv, abs/2308.12964, 2023. 113

  35. [43]

    Investigating the influence of noise and distractors on the interpretation of neural networks

    Pieter-Jan Kindermans, Kristof Sch ¨utt, Klaus-Robert M¨uller, and Sven D ¨ahne. Investigating the influence of noise and distractors on the interpretation of neural networks. CoRR, abs/1611.07270, 2016. 113

  36. [44]

    Attention is not only a weight: Analyzing trans- formers with vector norms

    Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Ken- taro Inui. Attention is not only a weight: Analyzing trans- formers with vector norms. In Proceedings of the 2020 Con- ference on Empirical Methods in Natural Language Process- ing (EMNLP), pages 7057–7075, Online, 2020....

  37. [45]

    Incorporating Residual and Normalization Layers into Analysis of Masked Language Models

    Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Ken- taro Inui. Incorporating Residual and Normalization Layers into Analysis of Masked Language Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4547–4568, Online and P...

  38. [46]

    Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

    Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014. 7, 13

  39. [47]

    Towards faithful model explanation in nlp: A survey

    QING LYU, Marianna Apidianaki, and Chris Callison- Burch. Towards faithful model explanation in nlp: A survey. ArXiv, abs/2209.11326, 2022. 1, 113

  40. [48]

    Andreas Madsen, Siva Reddy, and A. P. Sarath Chandar. Post-hoc interpretability for neural nlp: A survey. ACM Computing Surveys, 55:1 – 42, 2021. 1, 113

  41. [49]

    SkipPLUS: Skip the first few layers to better explain vision transform- ers

    Faridoun Mehri, Mohsen Fayyaz, Mahdieh Soleymani Baghshah, and Mohammad Taher Pilehvar. SkipPLUS: Skip the first few layers to better explain vision transform- ers. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 204–215,

  42. [50]

    GlobEnc: Quantifying global token attribution by incorporating the whole encoder layer in transformers

    Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, and Mohammad Taher Pilehvar. GlobEnc: Quantifying global token attribution by incorporating the whole encoder layer in transformers. In Proceedings of the 2022 Confer- ence of the North American Chapter of the Association f...

  43. [51]

    Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar

    A. Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar. Adapler: Speeding up inference by adaptive length reduction. In Annual Meeting of the Association for Compu- tational Linguistics, 2022. 113

  44. [52]

    De- compX: Explaining transformers decisions by propagating token decomposition

    Ali Modarressi, Mohsen Fayyaz, Ehsan Aghazadeh, Yadol- lah Yaghoobzadeh, and Mohammad Taher Pilehvar. De- compX: Explaining transformers decisions by propagating token decomposition. In Proceedings of the 61st An- nual Meeting of the Association for Computational Linguis- tics...

  45. [53]

    Ex- plaining nonlinear classification decisions with deep taylor decomposition

    Gr ´egoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert M ¨uller. Ex- plaining nonlinear classification decisions with deep taylor decomposition. Pattern Recogn., 65(C):211–222, 2017. 114

  46. [54]

    Comparing automatic and human evaluation of local explanations for text classification

    Dong Nguyen. Comparing automatic and human evaluation of local explanations for text classification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan- guage Technologies, Volume 1 (Long Papers), pag...

  47. [55]

    Decompose-and-compose: A composi- tional approach to mitigating spurious correlation

    Fahimeh Hosseini Noohdani, Parsa Hosseini, Arian Yaz- dan Parast, Hamidreza Yaghoubi Araghi, and Mahdieh So- leymani Baghshah. Decompose-and-compose: A composi- tional approach to mitigating spurious correlation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision a...

  48. [56]

    Making sense of dependence: Efficient black-box explanations using dependence measure

    Paul Novello, Thomas Fel, and David Vigouroux. Making sense of dependence: Efficient black-box explanations using dependence measure. ArXiv, abs/2206.06219, 2022. 115

  49. [57]

    No token left be- hind: Explainability-aided image classification and genera- tion

    Roni Paiss, Hila Chefer, and Lior Wolf. No token left be- hind: Explainability-aided image classification and genera- tion. In Computer Vision – ECCV 2022 , pages 334–350, Cham, 2022. Springer Nature Switzerland. 113

  50. [58]

    Parkhi, Andrea Vedaldi, Andrew Zisserman, and C

    Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V . Jawahar. Cats and dogs. InIEEE Conference on Com- puter Vision and Pattern Recognition, 2012. 6

  51. [59]

    Beit v2: Masked image modeling with vector-quantized visual tokenizers

    Zhiliang Peng, Li Dong, Hangbo Bao, Qixiang Ye, and Furu Wei. Beit v2: Masked image modeling with vector-quantized visual tokenizers. ArXiv, abs/2208.06366, 2022. 5

  52. [60]

    Rise: Random- ized input sampling for explanation of black-box models

    Vitali Petsiuk, Abir Das, and Kate Saenko. Rise: Random- ized input sampling for explanation of black-box models. ArXiv, abs/1806.07421, 2018. 115

  53. [61]

    AttCAT: Explaining transformers via attentive class activation tokens

    Yao Qiang, Deng Pan, Chengyin Li, Xin Li, Rhongho Jang, and Dongxiao Zhu. AttCAT: Explaining transformers via attentive class activation tokens. In Advances in Neural In- formation Processing Systems, 2022. 2, 114

  54. [62]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...

  55. [63]

    Irvin, Aarti Bagul, Daisy Yi Ding, Tony Duan, Hershel Mehta, Brandon Yang, Kaylie Zhu, Dillon Laird, Robyn L

    Pranav Rajpurkar, Jeremy A. Irvin, Aarti Bagul, Daisy Yi Ding, Tony Duan, Hershel Mehta, Brandon Yang, Kaylie Zhu, Dillon Laird, Robyn L. Ball, C. Langlotz, Katie S. Sh- panskaya, Matthew P. Lungren, and A. Ng. Mura dataset: Towards radiologist-level abnormality detection in m...

  56. [64]

    Do imagenet classifiers generalize to im- agenet? In International Conference on Machine Learning,

    Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to im- agenet? In International Conference on Machine Learning,

  57. [65]

    why should i trust you?

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. “why should i trust you?”: Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD Interna- tional Conference on Knowledge Discovery and Data Min- ing, 2016. 115

  58. [66]

    Anders, and Klaus-Robert M ¨uller

    Wojciech Samek, Gr ´egoire Montavon, Sebastian La- puschkin, Christopher J. Anders, and Klaus-Robert M ¨uller. Explaining deep neural networks and beyond: A review of methods and applications. Proceedings of the IEEE , 109: 247–278, 2021. 1, 113

  59. [67]

    Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra

    Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In 2017 IEEE International Conference on Computer Vision (ICCV) , pages 618–626,

  60. [68]

    Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Dhruv Batra, and Devi Parikh

    Ramprasaath R. Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Dhruv Batra, and Devi Parikh. Taking a hint: Leverag- ing explanations to make vision and language models more grounded. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 2591–2600, 2019. 113

  61. [69]

    Noam M. Shazeer. Glu variants improve transformer. ArXiv, abs/2002.05202, 2020. 3

  62. [70]

    Pami: partition input and aggregate outputs for model inter- pretation

    Wei Shi, Wentao Zhang, Weishi Zheng, and Ruixuan Wang. Pami: partition input and aggregate outputs for model inter- pretation. ArXiv, abs/2302.03318, 2023. 115

  63. [71]

    Not just a black box: Learning important features through propagating activation differences

    Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje. Not just a black box: Learning important features through propagating activation differences. ArXiv, abs/1605.01713, 2016. 2, 113

  64. [72]

    Learning important features through propagating activation differences

    Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. Learning important features through propagating activation differences. In International Conference on Machine Learn- ing, 2017. 2, 113

  65. [73]

    Deep inside convolutional networks: Visualising image clas- sification models and saliency maps

    Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: Visualising image clas- sification models and saliency maps. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Workshop Track P...

  66. [74]

    Riedmiller

    Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin A. Riedmiller. Striving for simplicity: The all convolutional net. CoRR, abs/1412.6806, 2014. 113

  67. [75]

    Full-gradient represen- tation for neural network visualization

    Suraj Srinivas and Franc ¸ois Fleuret. Full-gradient represen- tation for neural network visualization. In Neural Informa- tion Processing Systems, 2019. 2, 4, 113

  68. [76]

    Eva-clip: Improved training techniques for clip at scale

    Quan Sun, Yuxin Fang, Ledell Yu Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. ArXiv, abs/2303.15389, 2023. 5 11

  69. [77]

    Axiomatic attribution for deep networks

    Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In Proceedings of the 34th In- ternational Conference on Machine Learning , pages 3319–

  70. [78]

    Imagenet-hard: The hard- est images remaining from a study of the power of zoom and spatial biases in image classification

    Mohammad Reza Taesiri, Giang Nguyen, Sarra Habchi, Cor- Paul Bezemer, and Anh Nguyen. Imagenet-hard: The hard- est images remaining from a study of the power of zoom and spatial biases in image classification. 2023. 6

  71. [79]

    Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lu- cas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy

    Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lu- cas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. Mlp-mixer: An all-mlp architecture for vision. In Neural Information Processing Systems, 2021. 5

  72. [80]

    Train- ing data-efficient image transformers & distillation through attention

    Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv’e J’egou. Train- ing data-efficient image transformers & distillation through attention. ArXiv, abs/2012.12877, 2020. 5

  73. [81]

    Deit iii: Revenge of the vit

    Hugo Touvron, Matthieu Cord, and Herv’e J’egou. Deit iii: Revenge of the vit. In European Conference on Computer Vision, 2022. 5

  74. [82]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neu- ral Information Processing Systems. Curran Associates, Inc.,

  75. [83]

    Analyzing multi-head self-attention: Spe- cialized heads do the heavy lifting, the rest can be pruned

    Elena V oita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self-attention: Spe- cialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 57...

  76. [84]

    Learning robust global representations by penalizing local predictive power

    Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. In Advances in Neural Information Processing Systems, pages 10506–10518, 2019. 6

  77. [85]

    Score-cam: Score-weighted visual explanations for convolutional neural networks

    Haofan Wang, Zifan Wang, Mengnan Du, Fan Yang, Zi- jian Zhang, Sirui Ding, Piotr (Peter) Mardziel, and Xia Hu. Score-cam: Score-weighted visual explanations for convolutional neural networks. 2020 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition Work- shops (CV...

  78. [86]

    Beyond explaining: Opportunities and challenges of xai-based model improvement.Inf

    Leander Weber, Sebastian Lapuschkin, Alexander Binder, and Wojciech Samek. Beyond explaining: Opportunities and challenges of xai-based model improvement.Inf. Fusion, 92: 154–176, 2022. 113

  79. [87]

    Token transformation matters: Towards faithful post-hoc ex- planation for vision transformer

    Junyi Wu, Bin Duan, Weitai Kang, Hao Tang, and Yan Yan. Token transformation matters: Towards faithful post-hoc ex- planation for vision transformer. 2024 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 10926–10935, 2024. 2, 5, 6, 11, 12, 114

  80. [88]

    Lyu, and Yu-Wing Tai

    Weibin Wu, Yuxin Su, Xixian Chen, Shenglin Zhao, Ir- win King, Michael R. Lyu, and Yu-Wing Tai. Boost- ing the transferability of adversarial samples via attention. 2020 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 1158–1167, 2020. 113

  81. [89]

    Vit-cx: Causal explanation of vision transformers

    Weiyan Xie, Xiao hui Li, Caleb Chen Cao, and Nevin L.Zhang. Vit-cx: Causal explanation of vision transformers. In International Joint Conference on Artificial Intelligence ,

  82. [90]

    Puyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane ling Wang, and Michael I. Jordan. Ml-loo: Detecting adversarial exam- ples with feature attribution. In AAAI Conference on Artifi- cial Intelligence, 2019. 113

  83. [91]

    Mambaout: Do we really need mamba for vision? arXiv preprint arXiv:2405.07992,

    Weihao Yu and Xinchao Wang. Mambaout: Do we really need mamba for vision? arXiv preprint arXiv:2405.07992,

  84. [92]

    Sigmoid loss for language image pre-training

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. ArXiv, abs/2303.15343, 2023. 5

  85. [93]

    Lin, Jonathan Brandt, Xiaohui Shen, and Stan Sclaroff

    Jianming Zhang, Zhe L. Lin, Jonathan Brandt, Xiaohui Shen, and Stan Sclaroff. Top-down neural attention by excitation backprop. International Journal of Computer Vision , 126: 1084–1102, 2016. 113

  86. [94]

    Jianping Zhang, Weibin Wu, Jen tse Huang, Yizhan Huang, Wenxuan Wang, Yuxin Su, and Michael R. Lyu. Improv- ing adversarial transferability via neuron attribution-based attacks. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14973–14982, 2022...

  87. [97]

    Via duality: SwapBackward (f, g)(x) = f (x).detach() + (g(x) − g(x).detach())

  88. [98]

    Remark 3 (Computational Efficiency)

    Via custom backward: Define an autograd.Function that returns f (x) in forward and propagates gradients as if it were g(x) in backward Both implementations yield identical gradients, though the latter may be more computationally efficient, while the former may be easier to imp...

  89. [99]

    ⊙ f2(x) where [·]cst

    When a = 0 , yielding f (x) = [ f1(x)]cst. ⊙ f2(x) where [·]cst. is the constant operator that zeroes gradients, f is FG- complete if f2 is FG-complete

  90. [100]

    By symmetry, when b = 0, f is FG-complete if f1 is FG-complete. 6 Proof. Let a = 0 (thus b = 1). If f2 is FG-complete: [diag(f1(x)) · Jxf2] · x + X i [diag(f1(x)) · Jbi f2] · bi = diag(f1(x)) · (Jxf2 · x + X i Jbi f2 · bi) = diag(f1(x)) · f2(x) = f1(x) ⊙ f2(x) = f (x) proving ...

  91. [101]

    Centering: y = x − µ1, where 1 is the vector of ones

  92. [102]

    Zebra” and “African Elephant

    Scaling: z = y/s, where s = √ σ2 + ε The Jacobian of centering is: (Jxy)ij = δij − 1 N which gives (Jxy · x)i = xi − µ = yi. The Jacobian of scaling is: (Jyz)ij = δij s − yiyj N s3 By the chain rule: JxLN · x = Jyz · Jxy · x = Jyz · y Computing (Jyz · y)i: (Jyz · y)i = NX j=1 ...

  93. [103]

    Negative Value Removal: We first apply ReLU to remove negative attribution scores, as we focus on positive feature contributions

  94. [104]

    We then scale the values by dividing by this robust maximum

    Robust Scaling: Rather than using absolute maximum values which can be sensitive to outliers, we compute the 99th percentile of the attribution scores. We then scale the values by dividing by this robust maximum

  95. [105]

    Spatial Upsampling: The token-level attribution map is upsampled to the original image resolution using bicubic inter- polation

  96. [106]

    Range Normalization: Finally, we clamp values to [0, 1]. 13 C. Qualitative Results Following the evaluation protocol in Appendix B.4, we present a comprehensive qualitative analysis below. C.1. Text-Prompted Qualitative Examples on EV A2-CLIP-Large Our first evaluation scenari...

  97. [107]

    GlobEnc & ALTI

    extends AttIN to also incorporate the residual connec- tions. GlobEnc & ALTI. AttIN assumes that tokens retain their original identity. As each self-attention module mixes all the tokens, this assumption might not necessarily hold. Us- ing gradient-based techniques, Brunner et...

  98. [2024]

    2, 5, 6, 11, 12, 13, 113

  99. [3328]

    2, 4, 6, 113

    PMLR, 2017. 2, 4, 6, 113

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.