ASCENT-ViT aligns ViT image representations with human-annotated concepts using multi-scale features and deformable attention, improving both accuracy and concept localization.
The following convolution blocks consist of a single convolution layer of kernel size=3 and stride=2, as well as the batch norm and ReLU activations
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers
ASCENT-ViT aligns ViT image representations with human-annotated concepts using multi-scale features and deformable attention, improving both accuracy and concept localization.