Adding the position embedding to the final layer normalization, called MPVG, improves vision transformer top-1 accuracy by 0.27 to 1.37 percentage points across several GAP-based models and tasks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Maximizing the Position Embedding for Vision Transformers with Global Average Pooling
Adding the position embedding to the final layer normalization, called MPVG, improves vision transformer top-1 accuracy by 0.27 to 1.37 percentage points across several GAP-based models and tasks.