Using the mean register-token embedding together with the CLS token in frozen DINOv2 backbones improves ImageNet out-of-distribution accuracy by about 2-4% and anomaly rejection FPR by about 2-3% over CLS plus mean-patch baselines.
Vision transformers need registers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Leveraging Registers in Vision Transformers for Robust Adaptation
Using the mean register-token embedding together with the CLS token in frozen DINOv2 backbones improves ImageNet out-of-distribution accuracy by about 2-4% and anomaly rejection FPR by about 2-3% over CLS plus mean-patch baselines.