REVIEW 3 cited by
Better plain ViT baselines for ImageNet-1k
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
It is commonly accepted that the Vision Transformer model requires sophisticated regularization techniques to excel at ImageNet-1k scale data. Surprisingly, we find this is not the case and standard data augmentation is sufficient. This note presents a few minor modifications to the original Vision Transformer (ViT) vanilla training setting that dramatically improve the performance of plain ViT models. Notably, 90 epochs of training surpass 76% top-1 accuracy in under seven hours on a TPUv3-8, similar to the classic ResNet50 baseline, and 300 epochs of training reach 80% in less than one day.
Forward citations
Cited by 3 Pith papers
-
Lilith: Backdoor Generalization under Training-Inference Trigger Shift
A single poisoned training trigger can create a backdoor that fires for a whole family of unseen inference-time triggers, provided the variants preserve the anchor's representation geometry.
-
Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm
A spectral-norm inner perturbation combined with a Muon outer update achieves the best ImageNet validation accuracy among the compared SAM variants on both a ViT-Small/16 and a ResNet-50.
-
Fine-Grained Image Recognition from Scratch with Teacher-Guided Data Augmentation
TGDA trains fine-grained-recognition students from scratch using attention maps and soft labels from a fine-tuned, pretrained teacher, reporting large gains on three benchmarks.
Discussion (0). Sign in to comment.