Pith. sign in

REVIEW 9 cited by

When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.01548 v3 pith:YFMVAVKX submitted 2021-06-03 cs.CV cs.LG

classification cs.CVcs.LG
keywords datavitsaugmentationsmodelspre-trainingstrongvisionaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision Transformers (ViTs) and MLPs signal further efforts on replacing hand-wired features or inductive biases with general-purpose neural architectures. Existing works empower the models by massive data, such as large-scale pre-training and/or repeated strong data augmentations, and still report optimization-related problems (e.g., sensitivity to initialization and learning rates). Hence, this paper investigates ViTs and MLP-Mixers from the lens of loss geometry, intending to improve the models' data efficiency at training and generalization at inference. Visualization and Hessian reveal extremely sharp local minima of converged models. By promoting smoothness with a recently proposed sharpness-aware optimizer, we substantially improve the accuracy and robustness of ViTs and MLP-Mixers on various tasks spanning supervised, adversarial, contrastive, and transfer learning (e.g., +5.3\% and +11.0\% top-1 accuracy on ImageNet for ViT-B/16 and Mixer-B/16, respectively, with the simple Inception-style preprocessing). We show that the improved smoothness attributes to sparser active neurons in the first few layers. The resultant ViTs outperform ResNets of similar size and throughput when trained from scratch on ImageNet without large-scale pre-training or strong data augmentations. Model checkpoints are available at \url{https://github.com/google-research/vision_transformer}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 103 citations worldwide. Full citation record

  1. Flat Minima and Generalization: Insights from Stochastic Convex Optimization

    cs.LG 2025-11 conditional novelty 7.0 of 10

    In smooth stochastic convex optimization, flat empirical minima can incur constant population risk while sharp minima generalize optimally, and sharpness-aware algorithms can converge to such bad flat minima.

  2. Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A spectral-norm inner perturbation combined with a Muon outer update achieves the best ImageNet validation accuracy among the compared SAM variants on both a ViT-Small/16 and a ResNet-50.

  3. Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MANO replaces quadratic self-attention with multiscale windowed attention over progressively downsampled grids, keeping complexity linear and reporting competitive accuracy on vision and PDE benchmarks.

  4. ReMem: Mutual Information-Aware Fine-tuning of Pretrained Vision Transformers for Effective Knowledge Distillation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Downgrading the top MLP blocks of a fine-tuned ViT and using large-radius SAM during fine-tuning improves knowledge distillation to small students by preserving mutual information between inputs and teacher outputs.

  5. TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TDVE-DB and TDVE-Assessor deliver a large MOS-annotated benchmark and a Qwen2.5-VL-based assessor for text-driven video editing quality.

  6. Underwater Monocular Metric Depth Estimation: Real-World Benchmarks and Synthetic Fine-Tuning with Vision Foundation Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Fine-tuning Depth Anything V2 on physics-based synthetic underwater versions of Hypersim improves metric depth accuracy on real underwater benchmarks like FLSea and SQUID, though one AbsRel number worsens slightly.

  7. Enhancing Wireless Networks for IoT with Large Vision Models: Foundations and Applications

    cs.NI 2025-08 conditional novelty 4.0 of 10

    The paper surveys LVM applications in wireless and reports a case study where progressive fine-tuning of pretrained LVMs gives more robust joint beamforming and positioning than a from-scratch CNN.

  8. Attributing Data for Sharpness-Aware Minimization

    cs.LG 2025-07 reject novelty 4.0 of 10

    SAM-HIF and SAM-GIF are proposed as data attribution scores for SAM-trained models, but SAM-GIF is TracIn with SAM gradients and SAM-HIF's derivation contains a load-bearing error.

  9. Learning from Limited and Imperfect Data

    cs.LG 2025-07 unverdicted novelty 3.0 of 10

    A doctoral thesis compiling nine peer-reviewed papers on long-tailed image generation, long-tailed recognition, semi-supervised learning, and domain adaptation.

Pith tools