REVIEW 4 cited by
DeiT III: Revenge of the ViT
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A Vision Transformer (ViT) is a simple neural architecture amenable to serve several computer vision tasks. It has limited built-in architectural priors, in contrast to more recent architectures that incorporate priors either about the input data or of specific tasks. Recent works show that ViTs benefit from self-supervised pre-training, in particular BerT-like pre-training like BeiT. In this paper, we revisit the supervised training of ViTs. Our procedure builds upon and simplifies a recipe introduced for training ResNet-50. It includes a new simple data-augmentation procedure with only 3 augmentations, closer to the practice in self-supervised learning. Our evaluations on Image classification (ImageNet-1k with and without pre-training on ImageNet-21k), transfer learning and semantic segmentation show that our procedure outperforms by a large margin previous fully supervised training recipes for ViT. It also reveals that the performance of our ViT trained with supervision is comparable to that of more recent architectures. Our results could serve as better baselines for recent self-supervised approaches demonstrated on ViT.
Forward citations
Cited by 4 Pith papers
-
Journey Operators for Structured Multi-Axis Composition
Path-independent multi-axis composition requires commuting axis generators, and under toral axioms the only compatible scores are block-wise rotations—motivating value-side RoPE (JoFormer).
-
EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion
EVM-Fusion combines three feature pathways and a learnable iterative fusion block to classify medical images across nine classes, reporting 94.79% accuracy on a composite public dataset.
-
Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings
LoRA fine-tuning of FetalCLIP achieves F1 0.757 for fetal ultrasound frame-quality classification, and a thresholded segmentation variant reaches F1 0.771.
-
DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding
DeepTraverse is a weight-tied residual network plus squeeze-and-excitation attention, framed as depth-first search, with claimed efficiency gains that rest on a questionable ImageNet subset comparison.
Discussion (0). Sign in to comment.