Under symmetric label noise, the larger ViTl32 model consistently outperforms smaller and higher-token-count variants in both accuracy and calibration, while Swin transformers lag behind.
Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Balancing Accuracy, Calibration, and Efficiency in Active Learning with Vision Transformers Under Label Noise
Under symmetric label noise, the larger ViTl32 model consistently outperforms smaller and higher-token-count variants in both accuracy and calibration, while Swin transformers lag behind.