GMAR weights ViT attention heads by the magnitude of their gradients to the predicted class and feeds them into attention rollout, improving four interpretability metrics over standard rollout in one fine-tuned setting.
Experimental Setup We utilize the Google ViT-Large-Patch16-224 model [23, 24] as the base architecture, fine-tuning it on the Tiny-ImageNet dataset [25]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
GMAR weights ViT attention heads by the magnitude of their gradients to the predicted class and feeds them into attention rollout, improving four interpretability metrics over standard rollout in one fine-tuned setting.