Pith. sign in

REVIEW 3 cited by

DeiT III: Revenge of the ViT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.07118 v1 pith:4JMK5KMT submitted 2022-04-14 cs.CV

classification cs.CV
keywords recentpre-trainingprocedureself-supervisedtrainingarchitectureslearningpriors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A Vision Transformer (ViT) is a simple neural architecture amenable to serve several computer vision tasks. It has limited built-in architectural priors, in contrast to more recent architectures that incorporate priors either about the input data or of specific tasks. Recent works show that ViTs benefit from self-supervised pre-training, in particular BerT-like pre-training like BeiT. In this paper, we revisit the supervised training of ViTs. Our procedure builds upon and simplifies a recipe introduced for training ResNet-50. It includes a new simple data-augmentation procedure with only 3 augmentations, closer to the practice in self-supervised learning. Our evaluations on Image classification (ImageNet-1k with and without pre-training on ImageNet-21k), transfer learning and semantic segmentation show that our procedure outperforms by a large margin previous fully supervised training recipes for ViT. It also reveals that the performance of our ViT trained with supervision is comparable to that of more recent architectures. Our results could serve as better baselines for recent self-supervised approaches demonstrated on ViT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Journey Operators for Structured Multi-Axis Composition

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Path-independent multi-axis composition requires commuting axis generators, and under toral axioms the only compatible scores are block-wise rotations—motivating value-side RoPE (JoFormer).

  2. Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings

    cs.CV 2025-07 conditional novelty 4.0 of 10

    LoRA fine-tuning of FetalCLIP achieves F1 0.757 for fetal ultrasound frame-quality classification, and a thresholded segmentation variant reaches F1 0.771.

  3. DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding

    cs.CV 2025-06 reject novelty 4.0 of 10

    DeepTraverse is a weight-tied residual network plus squeeze-and-excitation attention, framed as depth-first search, with claimed efficiency gains that rest on a questionable ImageNet subset comparison.

Pith tools