Pith. sign in

REVIEW 3 cited by

Better plain ViT baselines for ImageNet-1k

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.01580 v1 pith:TGEPQTNF submitted 2022-05-03 cs.CV

classification cs.CV
keywords trainingdataepochsimagenet-1kplaintransformervisionaccepted
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It is commonly accepted that the Vision Transformer model requires sophisticated regularization techniques to excel at ImageNet-1k scale data. Surprisingly, we find this is not the case and standard data augmentation is sufficient. This note presents a few minor modifications to the original Vision Transformer (ViT) vanilla training setting that dramatically improve the performance of plain ViT models. Notably, 90 epochs of training surpass 76% top-1 accuracy in under seven hours on a TPUv3-8, similar to the classic ResNet50 baseline, and 300 epochs of training reach 80% in less than one day.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lilith: Backdoor Generalization under Training-Inference Trigger Shift

    cs.CR 2026-07 conditional novelty 7.0 of 10

    A single poisoned training trigger can create a backdoor that fires for a whole family of unseen inference-time triggers, provided the variants preserve the anchor's representation geometry.

  2. Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A spectral-norm inner perturbation combined with a Muon outer update achieves the best ImageNet validation accuracy among the compared SAM variants on both a ViT-Small/16 and a ResNet-50.

  3. Fine-Grained Image Recognition from Scratch with Teacher-Guided Data Augmentation

    cs.CV 2025-07 reject novelty 6.0 of 10

    TGDA trains fine-grained-recognition students from scratch using attention maps and soft labels from a fine-tuned, pretrained teacher, reporting large gains on three benchmarks.

Pith tools