REVIEW 6 cited by
Explainability of Vision Transformers: A Comprehensive Review and New Perspectives
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Transformers have had a significant impact on natural language processing and have recently demonstrated their potential in computer vision. They have shown promising results over convolution neural networks in fundamental computer vision tasks. However, the scientific community has not fully grasped the inner workings of vision transformers, nor the basis for their decision-making, which underscores the importance of explainability methods. Understanding how these models arrive at their decisions not only improves their performance but also builds trust in AI systems. This study explores different explainability methods proposed for visual transformers and presents a taxonomy for organizing them according to their motivations, structures, and application scenarios. In addition, it provides a comprehensive review of evaluation criteria that can be used for comparing explanation results, as well as explainability tools and frameworks. Finally, the paper highlights essential but unexplored aspects that can enhance the explainability of visual transformers, and promising research directions are suggested for future investment.
Forward citations
Cited by 6 Pith papers
-
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
PhyCheck is a 69,825-pair video QA benchmark that tests and improves Video-LLMs' ability to judge whether events obey physical laws, with fine-grained evidence questions and a context-sensitivity pilot.
-
Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
Ablation experiments on DETR, DDETR, and DINO show a clear resilience gradient, with DINO's static content queries becoming largely expendable after training.
-
X-SiT: Inherently Interpretable Surface Vision Transformers for Dementia Diagnosis
A surface vision transformer diagnoses dementia by weighing how each cortical patch resembles a corresponding prototype from real training brains, with competitive accuracy and clinically aligned explanations.
-
Interpretable Classification of Levantine Ceramic Thin Sections via Neural Networks
Transfer-learned ResNet18 and ViT models classify thin-section images of Levantine ceramics into ten petrographic fabrics with up to 92% accuracy, and saliency maps show the models focus on mineral inclusions.
-
An Explainable Transformer Model for Alzheimer's Disease Detection Using Retinal Imaging
Retformer, a from-scratch transformer with CNN patch embedding, rotary position encoding, grouped query attention, and SwiGLU, reports 92% OCT and 94% fundus accuracy for Alzheimer's detection, surpassing several pret...
-
PiPViT: Patch-based Visual Interpretable Prototypes for Retinal Image Analysis
PiPViT combines vision transformers and prototype learning to classify retinal OCT scans while showing the spatial extent of the biomarker that drove the decision.
Discussion (0). Continue with ORCID to comment.