Pith. sign in

REVIEW 4 major objections 5 minor 63 references

Self-supervised Benchmark Lottery on ImageNet: Do Marginal Improvements Translate to Improvements on Similar Datasets?

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper finds that top-1 accuracy on ImageNet validation is a weak predictor of accuracy on similar datasets, with model rankings nearly inverting on style-shifted variants.

desk verdict The rank inversion on OOD ImageNet variants is likely real, but the per-framework linear probe confound keeps it from being fully convincing. read the letter →

arxiv 2501.15431 v1 pith:KSMYFEJT submitted 2025-01-26 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords self-supervisedlearningImageNetvariantsbenchmarklotterymodelevaluationdistributionshiftlinearprobingrepresentationout-of-distributiongeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the small top-1 accuracy gains that self-supervised learning (SSL) frameworks report on ImageNet's validation set actually indicate better models. To answer, the authors evaluate twelve popular SSL frameworks on five ImageNet-derived datasets (ReaL, v2, Rendition, Sketch, and Adversarial) and find that rankings change substantially when the data shifts. DINO and SwAV, which sit first and second on ImageNet validation, fall to the middle of the field on Rendition and Sketch, while MoCo and Barlow Twins climb. The paper argues that benchmarking on ImageNet alone hides useful properties of SSL models, and it proposes two aggregate metrics that combine accuracy across the variants into a single ranking. If the finding holds, a model's ImageNet validation accuracy should not be treated as a reliable predictor of its performance on similar but shifted data.

What carries the argument

The key machinery is the set of five ImageNet variant datasets used as evaluation probes, together with two aggregate accuracy metrics. The variants are ReaL (multi-label relabeled validation images), v2 (a re-collected near-copy of the validation set), Rendition (stylized and altered real images), Sketch (line drawings), and Adversarial (naturally misleading images). Each variant exposes a different failure mode of the SSL feature extractors, which are all ResNet-50 backbones taken from the frameworks' official repositories and evaluated through a linear probe trained on ImageNet. The two aggregate measures—a weighted average with weights proportional to dataset size, and a geometric mean of accuracies—are introduced to turn the per-dataset results into a single ranking; the geometric mean is deliberately 'pessimistic' because it amplifies changes on low-scoring datasets such as Adversarial.

What would settle it

Re-run the linear probing stage for DINO, SwAV, MoCo, and Barlow Twins with multiple random seeds and learning-rate schedules on Rendition and Sketch; if the 4–5 percentage point gaps between MoCo and the others shrink or overlap, the reported rank inversion would not be statistically reliable, but if the gaps persist across seeds, the paper's conclusion holds.

Watch

Extended reading notes

Core claim

The central discovery is that top-1 linear accuracy on ImageNet validation is a weak predictor of accuracy on datasets that resemble ImageNet in class structure but differ in style, labeling, or distribution. The paper reports Pearson correlations of r=0.99 with ImageNet ReaL and v2, but only about r=0.6 with Rendition, Sketch, and Adversarial. On Rendition and Sketch, the two models that lead the validation ranking, DINO and SwAV, are displaced by MoCo and BYOL, with MoCo ahead of the next model by about four to four and a half percentage points. The paper interprets this as evidence that marginal improvements on ImageNet validation can be an artifact of the benchmark rather than a genuine improvement in the learned representation, and it concludes that a suite of ImageNet variants is needed to compare SSL frameworks fairly.

Load-bearing premise

Each self-supervised framework is represented by exactly one official checkpoint and one linear probe trained with the original paper's hyperparameters, and these single measurements are treated as exact when models are ranked.

Editorial extensions

If this is right

  • ImageNet validation accuracy alone should not be used to compare self-supervised frameworks.
  • Models such as MoCo and Barlow Twins are likely underrated relative to their ImageNet ranking, while DINO and SwAV are likely overrated for out-of-distribution use.
  • Benchmarking on a suite of ImageNet variants, as the proposed aggregate metrics do, gives a more complete picture of a model's strengths and weaknesses.
  • The absence of a consistent winner across SSL methodologies suggests that claims about one methodology being generally superior are unjustified.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural follow-up study could train variants of MoCo and BYOL with and without specific augmentations or loss terms to identify what makes them resilient to sketch and rendition inputs.
  • Since the validation differences among DINO, SwAV, and MoCo are only a few tenths of a percent, a multi-seed linear-probe re-evaluation could establish whether the reported inversion is statistically stable or rests on a single checkpoint.
  • The aggregate metrics treat all variants as equally important; a deployment-aware evaluation would weight each variant according to the target application's expected data distribution.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper evaluates twelve self-supervised learning frameworks (ResNet-50 backbones pretrained on ImageNet) by training a linear probe on the ImageNet training set and measuring top-1 accuracy on ImageNet validation and five variants: ReaL, v2, Rendition, Sketch, and Adversarial. The authors report that while DINO and Swav lead on ImageNet validation, they drop sharply on Rendition and Sketch, where MoCo and Barlow Twins improve relative to their validation ranks. They also propose weighted-average and geometric-mean aggregate metrics to combine performance across variants. The central claim is that ImageNet validation accuracy is an incomplete and sometimes misleading benchmark for SSL models.

Significance. The study addresses an important and timely question: whether marginal gains on ImageNet validation imply improvements on shifted ImageNet-like distributions. Its strengths are that it uses publicly released official checkpoints and standard public datasets, does not fit any free parameters, and covers a broad spread of SSL families. If the headline rank inversion on Rendition/Sketch is confirmed under a controlled evaluation protocol, the result would be a valuable caution against single-benchmark SSL evaluation. The main weakness is that the linear probe protocol is taken from each framework's original paper, so the comparison is of backbone-plus-recommended-probe systems rather than of frozen representations alone, and the paper does not quantify uncertainty. These issues are fixable but currently leave the central representation-level claim not fully established.

major comments (4)
  1. [Section III-B] The central comparison confounds representation quality with linear-probe protocol. The paper states that each backbone is evaluated with the SGD routine outlined in the respective paper, so the probes differ across frameworks in optimizer, epochs, learning-rate schedule, augmentation, weight decay, and possibly input resolution. The large gaps in Table II on Rendition (MoCo 24.4 vs DINO 18.9) and Sketch (26.1 vs 19.2) could therefore reflect probe robustness to distribution shift rather than the quality of the frozen representations. Please rerun at least the four headline models (DINO, Swav, MoCo, Barlow, and ideally all twelve) with a fixed linear-probe protocol and confirm that the rank inversion persists.
  2. [Section IV-A, Table III] The rank-change evidence on ReaL and v2 is not distinguishable from noise. On ReaL, DINO 81.4, Swav 81.3, and MoCo 81.3; on v2, DINO 61.4, Swav 61.3, and MoCo 61.1. With 50,000 and 10,000 test images, single measurements have approximate binomial standard errors of about 0.17% and 0.49%, respectively, and the reported differences are 0.1-0.3%. The paper uses these close rankings to argue that marginal improvements on ImageNet may be misleading, but unrepeated single measurements cannot support rank claims at this scale. Provide confidence intervals or repeated-seed measurements and restrict rank-change claims to differences above the noise floor.
  3. [Section IV-C] The paper explicitly acknowledges that it does not quantify uncertainty in the accuracy on individual datasets, but it then constructs aggregate metrics and rank tables directly from those single measurements. The geometric-mean ranks in Table IV are particularly sensitive to the ImageNet Adversarial column, where all accuracies are below 3.4% and absolute differences are 0.1-1.9%; OBoW's rank-1 geometric mean is driven by its 3.3% on Adversarial. The aggregate ranking conclusions should be supported by a sensitivity analysis, for example a bootstrap over images or over repeated probes, before statements such as 'MoCo takes the first spot' are made.
  4. [Reproducibility] The paper releases no code and no exact evaluation configuration; it only refers the reader to the 'Linear Evaluation' sections of the original SSL papers. For an empirical benchmark paper whose entire contribution is a comparative measurement, this makes the central result not independently testable. Please release the exact linear-probe command lines, dataset preprocessing, and evaluation harness, or a containerized version of the evaluation.
minor comments (5)
  1. [Fig. 1 caption] The caption contains a typo: 'Barlow TWins' should be 'Barlow Twins'.
  2. [Section IV-A] The text 'natural adversarial adversarial examples' contains a duplicated word; it should be 'natural adversarial examples'.
  3. [Fig. 3] The x-axis label 'ImageNet Real Accuracy' is inconsistent with the dataset name 'ReaL' used elsewhere; please unify the spelling.
  4. [Section IV-B] The reported Pearson correlation coefficients (r=0.99 for ReaL and v2, r=0.63-0.67 for Rendition, Sketch, and Adversarial) are based on only 12 models with no confidence intervals; a bootstrap interval would help the reader judge the strength of the difference.
  5. [Table I] For the ImageNet Adversarial row, the 'Image per class' entry of ~37 is described as approximate; please clarify whether this is the mean or median, given that the dataset has 7,500 images across 200 classes.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper is an empirical benchmark study anchored to externally published checkpoints and public datasets, with no fitted parameter renamed as a prediction and no load-bearing self-citation.

full rationale

The paper's central claim is that ImageNet validation accuracy does not reliably predict SSL model performance on similar out-of-distribution datasets. This claim is supported by direct measurements: top-1 accuracies of twelve SSL backbones on five public ImageNet variant datasets, using pretrained ResNet-50 checkpoints taken as-is from the respective framework repositories. No parameter in the paper is fitted to the reported accuracies; the aggregate metrics (weighted average and geometric mean) are deterministic functions of the measured accuracies, not fitted quantities. The linear probes are trained with each framework's own published SGD recipe, which introduces an experimental-protocol concern about whether rankings reflect backbone quality or probe robustness, but this is a correctness risk, not circularity: the probe recipe is not defined in terms of the target result and is not derived from the paper's own conclusions. Self-citations appear only as contextual references (e.g., a prior survey by the authors and a prior study on adversarial attack evaluation) and do not carry any load-bearing argument. The paper also explicitly acknowledges a limitation—that quantifying uncertainty in dataset accuracies is not attempted—which further confirms that no hidden fitted input is being recycled as a prediction. There is no equation in which an input is defined in terms of an output, no fitted parameter renamed as a prediction, and no uniqueness theorem or ansatz imported from the authors' prior work. The benchmark-lottery framing is borrowed from an external, non-self-cited source and is used only as motivation. Therefore, the derivation chain is self-contained and empirically anchored, and the circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No numeric free parameters are fitted; the weighted-average weights are dataset sizes fixed by the data, and the aggregate metrics are parameter-free. The central comparison rests on five domain assumptions about checkpoint faithfulness, linear probe stability, metric comparability, aggregate meaningfulness, and statistical power. No invented entities are introduced.

assumptions (5)
  • domain assumption Official repository pretrained ResNet-50 checkpoints are faithful implementations of each SSL framework.
    Section III-B: 'we take the parameters of ResNet-50 backbone networks as-is, without making any changes from the respective repositories.' The cross-model comparison assumes each checkpoint fairly represents the framework.
  • domain assumption A single linear probe per framework, trained with the hyperparameters of each original paper, gives stable and comparable accuracy estimates.
    Section III-B: 'we adhere to the SGD training routines outlined in the respective papers.' No seeds, repeats, or confidence intervals are reported.
  • domain assumption Top-1 accuracy against multi-label ReaL labels is comparable to top-1 accuracy on single-label ImageNet validation.
    Section III-A describes ReaL as multi-label where any correct label counts; Table II lists both under 'Top-1 accuracy' without noting the protocol difference in the table.
  • domain assumption Weighted average and geometric mean across five heterogeneous datasets produce meaningful unified rankings.
    Section IV-C defines the metrics without external validation; the geometric mean is strongly affected by near-chance adversarial accuracies, e.g., OBoW becomes rank 1 due to 3.3% on Adversarial.
  • domain assumption Pearson correlation over twelve models is sufficient to infer whether ImageNet validation accuracy predicts variant accuracy.
    Section IV-B reports r=0.99 for ReaL/v2 and r=0.63 to 0.67 for Rendition, Sketch, and Adversarial with only 12 points; no p-values or standard errors are given.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Self-supervised Benchmark Lottery on ImageNet: Do Marginal Improvements Translate to Improvements on Similar Datasets?." pith.science (2026). https://pith.science/paper/KSMYFEJT

@misc{pith2026250115431,
  author       = {Pith},
  title        = {Pith review of: Self-supervised Benchmark Lottery on ImageNet: Do Marginal Improvements Translate to Improvements on Similar Datasets?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KSMYFEJT}},
  note         = {Machine review of arXiv:2501.15431}
}
read the original abstract

Machine learning (ML) research strongly relies on benchmarks in order to determine the relative effectiveness of newly proposed models. Recently, a number of prominent research effort argued that a number of models that improve the state-of-the-art by a small margin tend to do so by winning what they call a "benchmark lottery". An important benchmark in the field of machine learning and computer vision is the ImageNet where newly proposed models are often showcased based on their performance on this dataset. Given the large number of self-supervised learning (SSL) frameworks that has been proposed in the past couple of years each coming with marginal improvements on the ImageNet dataset, in this work, we evaluate whether those marginal improvements on ImageNet translate to improvements on similar datasets or not. To do so, we investigate twelve popular SSL frameworks on five ImageNet variants and discover that models that seem to perform well on ImageNet may experience significant performance declines on similar datasets. Specifically, state-of-the-art frameworks such as DINO and Swav, which are praised for their performance, exhibit substantial drops in performance while MoCo and Barlow Twins displays comparatively good results. As a result, we argue that otherwise good and desirable properties of models remain hidden when benchmarking is only performed on the ImageNet validation set, making us call for more adequate benchmarking. To avoid the "benchmark lottery" on ImageNet and to ensure a fair benchmarking process, we investigate the usage of a unified metric that takes into account the performance of models on other ImageNet variant datasets.

Figures

Figures reproduced from arXiv: 2501.15431 by the authors.

Figure 1
Figure 1. (top) Top-1 accuracy of SSL models plotted for variants [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Example images from ImageNet variants. positive pairs, while contrasting negative pairs [37]. Given a batch of n images [x1, . . . , xn], SimCLR generates an aug￾mented versions of each image, thus resulting with 2n images B = [x1, x (a) 1 , . . . , xn, x (a) n ] and minimizes the InfoNCE loss for each positive pair as follows: LCON := − log exp(sim(Bi , Bj )/τ ) P2n k=1 1{k̸=i} exp(sim(Bi , Bk)/τ ) . (2) Above, 1{k… view at source ↗
Figure 4
Figure 4. Aggregate measures of accuracy for each model across [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figures from the paper (1 more)
Figure 3
Figure 3. Figure 3: Comparison of top-1 linear accuracy on ImageNet [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

63 extracted references · 51 canonical work pages

  1. [1]

    ImageNet Classification with Deep Convolutional Neural Networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in Advances in Neural Information Processing Systems , 2012

  2. [2]

    Very Deep Convolutional Networks For Large-Scale Image Recognition,

    K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks For Large-Scale Image Recognition,” International Conference on Learning Representations, 2015

  3. [3]

    Aggregated Residual Transformations for Deep Neural Networks,

    S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated Residual Transformations for Deep Neural Networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pp. 1492– 1500, 2017

  4. [4]

    Data Labeling: An Empirical Investigation Into Industrial Challenges and Mitigation Strategies,

    T. Fredriksson, D. I. Mattos, J. Bosch, and H. H. Olsson, “Data Labeling: An Empirical Investigation Into Industrial Challenges and Mitigation Strategies,” in International Conference on Product-Focused Software Process Improvement, 2020

  5. [5]

    Generalizing From a Few Examples: A Survey on Few-shot Learning,

    Y . Wang, Q. Yao, J. T. Kwok, and L. M. Ni, “Generalizing From a Few Examples: A Survey on Few-shot Learning,” ACM Computing Surveys, 2020

  6. [6]

    A Survey of Transfer Learning,

    K. Weiss, T. M. Khoshgoftaar, and D. Wang, “A Survey of Transfer Learning,” Journal of Big data , 2016

  7. [7]

    Automatic Differentiation in PyTorch,

    A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic Differentiation in PyTorch,” 2017

  8. [8]

    ImageNet Large Scale Visual Recognition Challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision , vol. 115, no. 3, pp. 211–252, 2015

Show all 63 references
  1. [9]

    A Simple Frame- work for Contrastive Learning of Visual Representations,

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A Simple Frame- work for Contrastive Learning of Visual Representations,” in Interna- tional Conference on Machine Learning , 2020

  2. [10]

    Bootstrap your own latent-a new approach to self-supervised learning,

    J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al. , “Bootstrap your own latent-a new approach to self-supervised learning,” Advances in Neural Information Processing Systems , 2020

  3. [11]

    Momentum Contrast for Unsupervised Visual Representation Learning,

    K. He, H. Fan, Y . Wu, S. Xie, and R. Girshick, “Momentum Contrast for Unsupervised Visual Representation Learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020

  4. [12]

    Vicreg: Variance-Invariance- Covariance Regularization for Self-supervised Learning,

    A. Bardes, J. Ponce, and Y . LeCun, “Vicreg: Variance-Invariance- Covariance Regularization for Self-supervised Learning,” arXiv preprint arXiv:2105.04906, 2021

  5. [13]

    Barlow Twins: Self-supervised Learning via Redundancy Reduction,

    J. Zbontar, L. Jing, I. Misra, Y . LeCun, and S. Deny, “Barlow Twins: Self-supervised Learning via Redundancy Reduction,” in International Conference on Machine Learning , 2021

  6. [14]

    Prototypical Contrastive Learn- ing of Unsupervised Representations,

    J. Li, P. Zhou, C. Xiong, and S. Hoi, “Prototypical Contrastive Learn- ing of Unsupervised Representations,” in International Conference on Learning Representations, 2021

  7. [15]

    Unsupervised learning of visual features by contrasting cluster assign- ments,

    M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin, “Unsupervised learning of visual features by contrasting cluster assign- ments,” Advances in Neural Information Processing Systems , 2020

  8. [16]

    Self-supervised Learning of Pretext- invariant Representations,

    I. Misra and L. v. d. Maaten, “Self-supervised Learning of Pretext- invariant Representations,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020

  9. [17]

    A Cookbook of Self-Supervised Learn- ing,

    R. Balestriero, M. Ibrahim, V . Sobal, A. Morcos, S. Shekhar, T. Gold- stein, F. Bordes, A. Bardes, G. Mialon, Y . Tian, A. Schwarzschild, A. G. Wilson, J. Geiping, Q. Garrido, P. Fernandez, A. Bar, H. Pirsiavash, Y . LeCun, and M. Goldblum, “A Cookbook of Self-Supervised Lear...

  10. [18]

    Know Your Self- supervised Learning: A Survey on Image-based Generative and Discrim- inative Training,

    U. Ozbulak, H. J. Lee, B. Boga, E. T. Anzaku, H.-m. Park, A. V . Messem, W. D. Neve, and J. Vankerschaver, “Know Your Self- supervised Learning: A Survey on Image-based Generative and Discrim- inative Training,” Transactions on Machine Learning Research , 2023

  11. [19]

    A metric learning reality check,

    K. Musgrave, S. Belongie, and S.-N. Lim, “A metric learning reality check,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXV 16, pp. 681– 699, Springer, 2020

  12. [20]

    The benchmark lottery,

    M. Dehghani, Y . Tay, A. A. Gritsenko, Z. Zhao, N. Houlsby, F. Diaz, D. Metzler, and O. Vinyals, “The benchmark lottery,” arXiv preprint arXiv:2107.07002, 2021

  13. [21]

    An Empirical Study of Training Self- supervised Vision Transformers,

    X. Chen, S. Xie, and K. He, “An Empirical Study of Training Self- supervised Vision Transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021

  14. [22]

    Self-organizing neural network that discovers surfaces in random-dot stereograms,

    S. Becker and G. E. Hinton, “Self-organizing neural network that discovers surfaces in random-dot stereograms,” Nature, 1992

  15. [23]

    Learning classification with unlabeled data,

    V . R. de Sa, “Learning classification with unlabeled data,” Advances in Neural Information Processing Systems , 1994

  16. [24]

    Colorful Image Colorization,

    R. Zhang, P. Isola, and A. A. Efros, “Colorful Image Colorization,” in European Conference on Computer Vision , 2016

  17. [25]

    Learning Representations for Automatic Colorization,

    G. Larsson, M. Maire, and G. Shakhnarovich, “Learning Representations for Automatic Colorization,” in European Conference on Computer Vision, 2016

  18. [26]

    Photo-realistic single image super-resolution using a generative adversarial network,

    C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, et al. , “Photo-realistic single image super-resolution using a generative adversarial network,” in Proceedings of the IEEE Conference on Computer Vision and Pattern ...

  19. [27]

    Context encoders: Feature learning by inpainting,

    D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” inProceedings of the IEEE conference on computer vision and pattern recognition , pp. 2536– 2544, 2016

  20. [28]

    Unsupervised Repre- sentation Learning by Predicting Image Rotations,

    S. Gidaris, P. Singh, and N. Komodakis, “Unsupervised Repre- sentation Learning by Predicting Image Rotations,” arXiv preprint arXiv:1803.07728, 2018

  21. [29]

    Unsupervised Visual Repre- sentation Learning by Context Prediction,

    C. Doersch, A. Gupta, and A. A. Efros, “Unsupervised Visual Repre- sentation Learning by Context Prediction,” in Proceedings of the IEEE International Conference on Computer Vision , 2015

  22. [30]

    Split-brain Autoencoders: Unsu- pervised Learning by Cross-channel Prediction,

    R. Zhang, P. Isola, and A. A. Efros, “Split-brain Autoencoders: Unsu- pervised Learning by Cross-channel Prediction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017

  23. [31]

    Deep Clustering for Unsupervised Learning of Visual Features,

    M. Caron, P. Bojanowski, A. Joulin, and M. Douze, “Deep Clustering for Unsupervised Learning of Visual Features,” in European Conference on Computer Vision , 2018

  24. [32]

    Self-labelling via Si- multaneous Clustering and Representation Learning,

    Y . M. Asano, C. Rupprecht, and A. Vedaldi, “Self-labelling via Si- multaneous Clustering and Representation Learning,” arXiv preprint arXiv:1911.05371, 2019

  25. [33]

    Obow: Online bag-of-visual-words generation for self-supervised learn- ing,

    S. Gidaris, A. Bursuc, G. Puy, N. Komodakis, M. Cord, and P. Perez, “Obow: Online bag-of-visual-words generation for self-supervised learn- ing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021

  26. [34]

    Exploring Simple Siamese Representation Learn- ing,

    X. Chen and K. He, “Exploring Simple Siamese Representation Learn- ing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021

  27. [35]

    Emerging properties in self-supervised vision transformers,

    M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021

  28. [36]

    Billion-scale similarity search with gpus,

    J. Johnson, M. Douze, and H. Jégou, “Billion-scale similarity search with gpus,” IEEE Transactions on Big Data , 2019

  29. [37]

    Representation Learning with Contrastive Predictive Coding,

    A. v. d. Oord, Y . Li, and O. Vinyals, “Representation Learning with Contrastive Predictive Coding,” arXiv preprint arXiv:1807.03748, 2018

  30. [38]

    Learning Representations by Predicting Bags of Visual Words,

    S. Gidaris, A. Bursuc, N. Komodakis, P. Pérez, and M. Cord, “Learning Representations by Predicting Bags of Visual Words,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020

  31. [39]

    Gradient-Based Learning Applied To Document Recognition,

    Y . LeCun, L. Bottou, Y . Bengio, and P. Haffner, “Gradient-Based Learning Applied To Document Recognition,” Proceedings of the IEEE, 1998

  32. [40]

    Microsoft Coco: Common Objects In Con- text,

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft Coco: Common Objects In Con- text,” in Proceedings of the IEEE European Conference on Computer Vision, pp. 740–755, Springer, 2014

  33. [41]

    A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark,

    X. Zhai, J. Puigcerver, A. Kolesnikov, P. Ruyssen, C. Riquelme, M. Lu- cic, J. Djolonga, A. S. Pinto, M. Neumann, A. Dosovitskiy, et al. , “A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark,” arXiv preprint arXiv:1910.04867 , 2019

  34. [42]

    Scaling and Benchmark- ing Self-supervised Visual Representation Learning,

    P. Goyal, D. Mahajan, A. Gupta, and I. Misra, “Scaling and Benchmark- ing Self-supervised Visual Representation Learning,” in Proceedings of the ieee/cvf International Conference on computer vision , 2019

  35. [43]

    Do Better Imagenet Models Transfer Better?,

    S. Kornblith, J. Shlens, and Q. V . Le, “Do Better Imagenet Models Transfer Better?,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019

  36. [44]

    How Well Do Self- supervised Models Transfer?,

    L. Ericsson, H. Gouk, and T. M. Hospedales, “How Well Do Self- supervised Models Transfer?,” in Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition , 2021

  37. [45]

    Natural Adversarial Examples,

    D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song, “Natural Adversarial Examples,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021

  38. [46]

    Learning Robust Global Representations by Penalizing Local Predictive Power,

    H. Wang, S. Ge, Z. Lipton, and E. P. Xing, “Learning Robust Global Representations by Penalizing Local Predictive Power,” Advances in Neural Information Processing Systems , 2019

  39. [47]

    The many faces of robustness: A critical analysis of out-of-distribution generalization,

    D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo, et al. , “The many faces of robustness: A critical analysis of out-of-distribution generalization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021

  40. [48]

    Improving Robustness Against Common Corruptions by Covariate Shift Adaptation,

    S. Schneider, E. Rusak, L. Eck, O. Bringmann, W. Brendel, and M. Bethge, “Improving Robustness Against Common Corruptions by Covariate Shift Adaptation,” Advances in Neural Information Processing Systems, vol. 33, pp. 11539–11551, 2020

  41. [49]

    Do imagenet classifiers generalize to imagenet?,

    B. Recht, R. Roelofs, L. Schmidt, and V . Shankar, “Do imagenet classifiers generalize to imagenet?,” in International Conference on Machine Learning, 2019

  42. [50]

    Impact of ImageNet Model Selection on Domain Adaptation,

    Y . Zhang and B. D. Davison, “Impact of ImageNet Model Selection on Domain Adaptation,” in Proceedings of the IEEE/CVF Winter Confer- ence on Applications of Computer Vision Workshops , 2020

  43. [51]

    Semi- supervised Models are Strong Unsupervised Domain Adaptation Learn- ers,

    Y . Zhang, H. Zhang, B. Deng, S. Li, K. Jia, and L. Zhang, “Semi- supervised Models are Strong Unsupervised Domain Adaptation Learn- ers,” arXiv preprint arXiv:2106.00417 , 2021

  44. [52]

    A Baseline for Detecting Misclassified and Out-of-distribution Examples in Neural Networks,

    D. Hendrycks and K. Gimpel, “A Baseline for Detecting Misclassified and Out-of-distribution Examples in Neural Networks,” arXiv preprint arXiv:1610.02136, 2016

  45. [53]

    Contrastive Training for Improved Out-of-distribution Detection,

    J. Winkens, R. Bunel, A. G. Roy, R. Stanforth, V . Natarajan, J. R. Ledsam, P. MacWilliams, P. Kohli, A. Karthikesalingam, S. Kohl, et al., “Contrastive Training for Improved Out-of-distribution Detection,”arXiv preprint arXiv:2007.05566, 2020

  46. [54]

    Evaluating adversarial attacks on imagenet: A reality check on misclassification classes,

    U. Ozbulak, M. Pintor, A. Van Messem, and W. De Neve, “Evaluating adversarial attacks on imagenet: A reality check on misclassification classes,” arXiv preprint arXiv:2111.11056 , 2021

  47. [55]

    Deep Residual Learning For Image Recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning For Image Recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016

  48. [56]

    Going Deeper With Convolutions,

    C. Szegedy, W. Liu, Y . Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V . Vanhoucke, and A. Rabinovich, “Going Deeper With Convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1–9, 2015

  49. [57]

    Confident Learning: Estimating Uncertainty in Dataset Labels,

    C. Northcutt, L. Jiang, and I. Chuang, “Confident Learning: Estimating Uncertainty in Dataset Labels,” Journal of Artificial Intelligence Re- search, 2021

  50. [58]

    Selective brain damage: Measuring the disparate impact of model pruning,

    S. Hooker, Y . Dauphin, A. Courville, and A. Frome, “Selective brain damage: Measuring the disparate impact of model pruning,” 2019

  51. [59]

    Are We Done with Imagenet?,

    L. Beyer, O. J. Hénaff, A. Kolesnikov, X. Zhai, and A. v. d. Oord, “Are We Done with Imagenet?,” arXiv preprint arXiv:2006.07159 , 2020

  52. [60]

    ImageNet-trained CNNs are Biased Towards Texture; Increasing Shape Bias Improves Accuracy and Robustness,

    R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel, “ImageNet-trained CNNs are Biased Towards Texture; Increasing Shape Bias Improves Accuracy and Robustness,” arXiv preprint arXiv:1811.12231, 2018

  53. [61]

    Explaining and Harnessing Adversarial Examples,

    I. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and Harnessing Adversarial Examples,” International Conference on Learning Repre- sentations, 2015

  54. [62]

    Adversarial Examples In The Physical World,

    A. Kurakin, I. Goodfellow, and S. Bengio, “Adversarial Examples In The Physical World,”Workshop Track, International Conference on Learning Representations, 2016

  55. [63]

    Borenstein, L

    M. Borenstein, L. V . Hedges, J. P. T. Higgins, and H. R. Rothstein, Introduction to Meta-Analysis . Wiley, 2009

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.