Pith. sign in

REVIEW 42 cited by

Quantifying Attention Flow in Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.00928 v2 pith:K5W5IRX2 submitted 2020-05-02 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords attentionflowinformationinputtokensmethodsweightsquantifying
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the Transformer model, "self-attention" combines information from attended embeddings into the representation of the focal embedding in the next layer. Thus, across layers of the Transformer, information originating from different tokens gets increasingly mixed. This makes attention weights unreliable as explanations probes. In this paper, we consider the problem of quantifying this flow of information through self-attention. We propose two methods for approximating the attention to input tokens given attention weights, attention rollout and attention flow, as post hoc methods when we use attention weights as the relative relevance of the input tokens. We show that these methods give complementary views on the flow of information, and compared to raw attention, both yield higher correlations with importance scores of input tokens obtained using an ablation method and input gradients.

Discussion (0). Sign in to comment.

Forward citations

Cited by 42 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator

    cond-mat.dis-nn 2026-07 conditional novelty 7.0 of 10

    Resolvent analysis of trained causal attention shows sinks act as transient dampers, routing heads carry excess Kreiss reserve, and eigenvalue depth predictions fail by 7–11 orders of magnitude.

  2. Architecture-Aware Explanation Auditing for Industrial Visual Inspection

    cs.LG 2026-05 conditional novelty 7.0 of 10

    Explanation faithfulness for deep classifiers on wafer maps is highest when the explainer matches the model's native readout structure, with ViT-Tiny plus Attention Rollout achieving lower Deletion AUC than mismatched...

  3. Architecture-Aware Explanation Auditing for Industrial Visual Inspection

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    An audit protocol on wafer maps finds that ViT-Tiny with Attention Rollout achieves better deletion faithfulness than other models and explainers, with readout structure as the key factor and RISE outperforming native...

  4. Trustworthiness in Retrieval-Augmented Generation Systems: A Survey

    cs.IR 2024-09 unverdicted novelty 7.0 of 10

    Introduces Trust-RAG Compass framework and TRC Bench benchmark to assess RAG trustworthiness across factuality, robustness, fairness, transparency, accountability, and privacy, with evaluations showing performance gap...

  5. A Generalist Agent

    cs.AI 2022-05 accept novelty 7.0 of 10

    Gato is a multi-modal, multi-task, multi-embodiment generalist policy using one transformer network to handle text, vision, games, and robotics tasks.

  6. AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model

    astro-ph.GA 2026-07 conditional novelty 6.0 of 10

    An uncertainty-aware transformer reconstructs masked AGN broad lines and spectral halves with 4-16% flux errors and beats eleven purpose-built Lyα-reconstruction algorithms on a blind benchmark.

  7. When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    STS is a two-stage pruning framework that decouples structural diversity via repulsion sampling from semantic filtering via cross-attention to reduce redundancy in visual tokens for VLMs.

  8. Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Contribution Weights combine attention, value magnitude, and directional alignment to measure token influence more faithfully than attention alone, and show attention sinks actively suppress information via a convex s...

  9. AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    AnchorDiff performs training-free concept grounding in multi-modal diffusion transformers by anchor selection followed by graph propagation on attention-derived graphs, reducing concept leakage on a new multi-concept dataset.

  10. Architecture-Aware Explanation Auditing for Industrial Visual Inspection

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    The paper proposes an architecture-aware explanation audit protocol demonstrating that perturbation-based faithfulness is bounded by structural compatibility between explainer and model readout rather than architectur...

  11. From Clever Hans to Scientific Discovery: Interpreting EEG Foundational Transformers with LRP

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    LRP on EEG transformers reveals Clever Hans artifacts in motor imagery tasks and a recurring central electrode cluster as a candidate sensorimotor signature of arousal.

  12. Feature-level Interaction Explanations in Multimodal Transformers

    cs.LG 2026-03 conditional novelty 6.0 of 10

    FL-I2MoE separates unique, synergistic, and redundant cross-modal evidence at the token/patch level and uses SII and redundancy-gap scores to rank pairs whose removal degrades performance more than random masking.

  13. Explainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer Models

    cs.CL 2026-02 unverdicted novelty 6.0 of 10

    CA-LIG is a unified hierarchical attribution method that computes layer-wise Integrated Gradients fused with class-specific attention gradients to generate signed, context-sensitive explanations for transformer models.

  14. Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Replacing square spectrogram patches with full-frequency temporal patches plus patch-aligned masking improves audio classification accuracy and reduces compute for Transformer and Mamba models.

  15. DSS-Prompt: Dynamic-Static Synergistic Prompting for Few-Shot Class-Incremental Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DSS-Prompt combines static prompts with instance-aware dynamic prompts generated from BLIP multi-modal features to achieve state-of-the-art few-shot class-incremental learning on four benchmarks without incremental training.

  16. Towards White-Box Deep Wireless Sensing

    cs.LG 2025-07 conditional novelty 6.0 of 10

    RF-CRATE derives a fully complex-valued white-box transformer for RF sensing from the sparse rate reduction principle and shows it matches black-box baselines across five datasets.

  17. Self-Guided Masked Autoencoder

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A Masked Autoencoder that masks the object cluster found by its own early patch-clustering signal learns better representations than random masking, with no external labels or models.

  18. CytoSAE: Interpretable Cell Embeddings for Hematology

    cs.CV 2025-07 conditional novelty 6.0 of 10

    CytoSAE learns sparse, expert-validated morphological concepts from blood-cell images that generalize across datasets and can classify AML subtypes at patient level with F1 0.83.

  19. VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A perturbation computed from early attention and value matrices can make LLaVA, Instruct-BLIP, and BLIP2-T5 fail to detect objects inside a specified image region while keeping the rest of the image usable.

  20. GraphPINE: Graph Importance Propagation for Interpretable Drug Response Prediction

    cs.LG 2025-04 unverdicted novelty 6.0 of 10

    GraphPINE is a GNN architecture that initializes node importance from prior knowledge graphs and propagates updates via an importance propagation layer for interpretable drug response prediction on over 5,000 genes an...

  21. Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers

    cs.CV 2026-07 conditional novelty 5.0 of 10

    FDT adds foveation and binary fixation modules to DeiT so multi-scale tokens are selected dynamically in one pass, improving ImageNet100 accuracy, MACs, and robustness without adversarial training.

  22. HRVConformer: Neonatal Hypoxic-Ischemic Encephalopathy Classification from the Heart Rate signals

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    HRVConformer, a convolution-Transformer hybrid, classifies HIE from raw HR signals with 83.23% AUC on a held-out 215-hour expert-annotated test set, outperforming Transformer, ResNet50 and FCN baselines.

  23. Learning Quantifiable Visual Explanations Without Ground-Truth

    cs.AI 2026-05 unverdicted novelty 5.0 of 10

    A perturbation-based metric for XAI quality that formalizes sufficiency and necessity, paired with an adapter trained via differentiable supervision to generate causal explanations on black-box models.

  24. SAIL: Structure-Aware Interpretable Learning for Anatomy-Aligned Post-hoc Explanations in OCT

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    SAIL integrates anatomical priors at the representation level with semantic features via fusion to produce more anatomically aligned attribution maps in OCT without altering existing explainability techniques.

  25. Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    HETA is a new attribution framework for decoder-only LLMs that combines semantic transition vectors, Hessian-based sensitivity scores, and KL divergence to produce more faithful and human-aligned token attributions th...

  26. From Features to Actions: Explainability in Traditional and Agentic AI Systems

    cs.AI 2026-02 conditional novelty 5.0 of 10

    Attribution explanations that work for static classifiers do not diagnose failures in multi-step AI agents; trace-grounded rubric evaluation does, with state-tracking inconsistency 2.7x more common in failed agent runs.

  27. Robust Representation Learning in Masked Autoencoders

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Masked Autoencoders build class-separable representations across depth and keep their embeddings directionally stable under blur and occlusion, which tracks their robust classification.

  28. Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A frozen 2D vision foundation model with lightweight LoRA adapters and attention-based slice fusion achieves state-of-the-art 3D medical image classification across 12 tasks with about 1M trainable parameters per task.

  29. An Autoencoder and Vision Transformer-based Interpretability Analysis of the Differences in Automated Staging of Second and Third Molars

    cs.CV 2025-09 conditional novelty 5.0 of 10

    An autoencoder-plus-ViT pipeline improves dental staging accuracy and uses latent space, reconstructions, and attention maps to attribute the weaker third-molar performance to high intra-class data variability.

  30. Attention Maps in 3D Shape Classification for Dental Stage Estimation with Class Node Graph Attention Networks

    cs.CV 2025-09 conditional novelty 5.0 of 10

    CGAT, a graph attention network with a CLS node, achieves 0.76 weighted F1 on Demirjian stage classification of 3D third-molar meshes and generates attention maps that highlight roots and furcation regions.

  31. Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A method and case study for visualizing cross-attention maps in Wan video diffusion transformers, showing token-region alignment over time and their use as artistic material.

  32. Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A systematic review of 55 papers finds explainability for multimodal attention-based models is dominated by attention-weight visualizations, while evaluation remains mostly qualitative and non-standardized.

  33. User Experience Estimation in Human-Robot Interaction Via Multi-Instance Learning of Multimodal Social Signals

    cs.RO 2025-07 reject novelty 5.0 of 10

    A multimodal Transformer with multi-instance learning estimates user experience (UX) questionnaire ratings from facial expressions and voice during human-robot interaction, reporting accuracy above third-party human raters.

  34. Fair-FLIP: Fair Deepfake Detection with Fairness-Oriented Final Layer Input Prioritising

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Fair-FLIP improves fairness parity in deepfake detection by reweighting final-layer features based on between-ethnicity variance, with negligible accuracy loss.

  35. Saccade Attention Networks: Using Transfer Learning of Attention to Reduce Network Sizes

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    Saccade Attention Networks use transfer learning to sparsify transformer inputs to attended image features, cutting calculations by ~80% with comparable results.

  36. Safer Skin Lesion Classification with Global Class Activation Probability Map Evaluation and SafeML

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A pixel-level argmax over per-class Grad-CAM maps, combined with a selective predictor, is proposed to detect unreliable skin lesion classifications.

  37. Generalizable Federated Learning using Client Adaptive Focal Modulation

    cs.CV 2025-08 reject novelty 4.0 of 10

    The abstract describes AdaptFED, a claimed federated learning method, but the full text is an unrelated graph theory paper, so the claimed results are absent from the submission.

  38. Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding

    cs.AI 2026-05 unverdicted novelty 3.0 of 10

    Advanced language representations shape LLMs' schemas to improve knowledge activation and problem-solving.

  39. Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models

    cs.LG 2025-08 unverdicted novelty 3.0 of 10

    A systematic literature review of explainability in multimodal attention models finds most studies focus on vision-language tasks with attention-based explanations, but evaluation methods lack consistency and modality...

  40. Learning from Limited and Imperfect Data

    cs.LG 2025-07 unverdicted novelty 3.0 of 10

    A doctoral thesis compiling nine peer-reviewed papers on long-tailed image generation, long-tailed recognition, semi-supervised learning, and domain adaptation.

  41. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

    cs.AI 2026-06 unverdicted novelty 2.0 of 10

    A comprehensive reference book organizing existing techniques for agentic AI systems across LLM substrate, reasoning, agent design patterns, inter-agent coordination, and production deployment.

  42. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

    cs.AI 2026-06 unverdicted novelty 1.0 of 10

    A survey-style reference book mapping the full agentic-AI stack from transformer internals to production deployment, with no new research result.

Pith tools