REVIEW 4 cited by
Understanding Why ViT Trains Badly on Small Datasets: An Intuitive Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision transformer (ViT) is an attention neural network architecture that is shown to be effective for computer vision tasks. However, compared to ResNet-18 with a similar number of parameters, ViT has a significantly lower evaluation accuracy when trained on small datasets. To facilitate studies in related fields, we provide a visual intuition to help understand why it is the case. We first compare the performance of the two models and confirm that ViT has less accuracy than ResNet-18 when trained on small datasets. We then interpret the results by showing attention map visualization for ViT and feature map visualization for ResNet-18. The difference is further analyzed through a representation similarity perspective. We conclude that the representation of ViT trained on small datasets is hugely different from ViT trained on large datasets, which may be the reason why the performance drops a lot on small datasets.
Forward citations
Cited by 4 Pith papers
-
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
A warm-up phase of ordinary federated training followed by zeroth-order forward-pass-only updates lets low-resource clients participate in federated pre-training from random initialization.
-
eMamba: Efficient Acceleration Framework for Mamba Models in Edge Computing
An end-to-end Mamba edge accelerator using hardware-friendly approximations, INT8 quantization, and NAS achieves 4.95x-5.62x lower latency and 1.63x-19.9x smaller models than ViT/CNN baselines.
-
Efficient optimization of expensive black-box simulators via marginal means, with application to neutrino detector design
A new estimator, BOMM, uses marginal mean functions to propose optimizer candidates beyond evaluated simulator runs, with proven consistency and improved high-dimensional rates under an additive model.
-
Self-Supervised Ultrasound-Video Segmentation with Feature Prediction and 3D Localised Loss
A 3D relative-localisation auxiliary loss improves V-JEPA pre-training for cardiac ultrasound video segmentation, with larger gains in low-label regimes.
Discussion (0). Continue with ORCID to comment.