REVIEW 9 cited by
ResNet strikes back: An improved training procedure in timm
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The influential Residual Networks designed by He et al. remain the gold-standard architecture in numerous scientific publications. They typically serve as the default architecture in studies, or as baselines when new architectures are proposed. Yet there has been significant progress on best practices for training neural networks since the inception of the ResNet architecture in 2015. Novel optimization & data-augmentation have increased the effectiveness of the training recipes. In this paper, we re-evaluate the performance of the vanilla ResNet-50 when trained with a procedure that integrates such advances. We share competitive training settings and pre-trained models in the timm open-source library, with the hope that they will serve as better baselines for future work. For instance, with our more demanding training setting, a vanilla ResNet-50 reaches 80.4% top-1 accuracy at resolution 224x224 on ImageNet-val without extra data or distillation. We also report the performance achieved with popular models with our training procedure.
Forward citations
Cited by 9 Pith papers
-
Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification in the Foundation-Model Era
At foundation-model scale, vehicle Re-ID no longer benefits from multi-branch or cross-backbone fusion: a tuned single backbone with exact re-ranking matches or beats fused multi-branch systems, with fusion gains boun...
-
On the Reliability of Cue Conflict and Beyond
Stylized cue-conflict bias scores are confounded by impure cues, imbalance, ratio metrics and restricted labels; REFINED-BIAS supplies pure balanced stimuli and full-label MRR sensitivity for reliable diagnosis.
-
Model Parallelism With Subnetwork Data Parallelism
Training each GPU on a fixed overlapping subnetwork and averaging shared parameters cuts per-device memory by up to 60 percent without exchanging activations, matching DDP accuracy under FLOP-matched budgets.
-
Robustness of Sensitivity Evaluations for Gravitational Wave Detection Algorithms
AresGW model 1's injection detection count at a false-alarm rate of 1/month varies with noise dataset by up to 39% coefficient of variation, while sensitive distance varies by only a few percent.
-
Expandable Residual Approximation for Knowledge Distillation
A new knowledge distillation method decomposes the teacher-student feature gap into multiple residual steps and reports improved accuracy on ImageNet and COCO.
-
Low-latency vision transformers via large-scale multi-head attention
Attention heads in compact vision transformers each recognize small label subsets with little noise, which the authors exploit for diverse ensembles and low-latency hybrid architectures on CIFAR-100.
-
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
An open-source image-generation family shows that agentic prompt rewriting and a stronger text encoder can lift quality to near closed-source levels with only 208.62M images and about $400K of training compute.
-
Automated Multi-label Classification of Eleven Retinal Diseases: A Benchmark of Modern Architectures and a Meta-Ensemble on a Large Synthetic Dataset
A benchmark of six neural networks plus a stacked ensemble shows that training on synthetic fundus images transfers to real retinal disease classification, with macro-AUC up to 0.88 externally.
-
Vision Transformers for End-to-End Quark-Gluon Jet Classification from Calorimeter Images
Vision transformer hybrids achieve the highest F1 and ROC-AUC among the tested models on CMS Open Data quark-gluon jet images, but the headline claim that ViT-based models beat CNNs is not supported by the standalone ...
Discussion (0). Sign in to comment.