Pith. sign in

REVIEW 72 references

Intra-class Patch Swap for Self-Distillation

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read An intra-class patch swap augmentation plus instance-to-instance KL distillation lets a single network train itself and beat several teacher-based and self-distillation baselines on image tasks.

desk verdict Novel self-distillation augmentation with a genuinely new idea, but the main CIFAR100 tables do not match the described protocol, so the headline claims need a correction before I would trust them. read the letter →

arxiv 2505.14124 v1 pith:TDW4GAI3 submitted 2025-05-20 cs.CV

classification cs.CV
keywords augmentationdistillationself-distillationintra-classsingleadditionalapproachesarchitectural
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Most knowledge distillation needs a large, pre-trained 'teacher' model to guide a small 'student' model. This paper removes the teacher: a single network plays both roles. The trick is a new data augmentation. For each pair of training images that share the same class label (say, two photos of otters), the method randomly cuts patches from one image and pastes them into the other, and vice versa. One resulting image usually keeps a strong, easy-to-recognize part (the otter's head), while the other loses it. The network then sees two images of the same class with different difficulty levels. The training loss has two parts: the normal cross-entropy loss with the true label, and a KL-divergence loss that pushes the network to produce similar prediction distributions for the two swapped images. The easy image acts as a weak teacher for the hard one.

The authors test this on CIFAR-100, ImageNet, fine-grained bird and dog datasets, semantic segmentation on PASCAL VOC and Cityscapes, and object detection on PASCAL VOC. They report gains of 1.3 to 3.4 points in top-1 accuracy over plain training, and better performance than several self-distillation and teacher-based distillation baselines. They also report improved robustness to adversarial attacks and better-calibrated predictions.

There are caveats. The main CIFAR-100 tables appear to use a progressive swap schedule that is not described in the method section, and the corruption-table averages do not match the per-corruption numbers. The ImageNet results are single runs. Because of these reporting problems, the exact size of the improvement is uncertain, but the core idea is simple and plausible.

Extended reading notes

Core claim

"Extensive experiments across image classification, semantic segmentation, and object detection show that our method consistently outperforms both existing self-distillation baselines and conventional teacher-based KD approaches." The paper also claims that "the success of self-distillation could hinge on the design of the augmentation itself."

Load-bearing premise

The paper depends on the assumption that the accuracy numbers in the main tables were produced by the method exactly as described in Section 4.2 (fixed pr=0.5, 4x4 patches, gamma=alpha=1.0). Cross-checking Table 3 against Table 15 shows that the CIFAR100 results match the progressive-pr variant labeled 'bpr', not the constant-pr variant 'pbar' that the text claims is used in all experiments. If the main tables actually use the undisclosed schedule, then the central performance claims, as stated, are not reproducible and the reported gains may partly reflect a tuned schedule rather than the patch-swap augmentation itself.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The central claim rests on hyperparameters tuned on CIFAR-100 (patch size, swap probability, temperature, loss weights) and on three domain assumptions about the effect of patch swapping. No new physical entities are introduced. The hidden use of the progressive schedule is a reporting issue rather than an added parameter, but it is counted in the free-parameter entry for pr.

free parameters (4)
  • Patch size (grid) = 4x4 (also 2x2, 3x3 explored)
    Chosen empirically; Figure 10 shows performance varies with patch size; 4x4 generally best. Affects the granularity of the swap and the confidence gap.
  • Swap probability pr = 0.5 constant or progressive 0.1 to 0.5
    Sections 3.2 and 4.5.4. Main CIFAR100 tables appear to use the progressive schedule, though the method states pr=0.5. Directly controls how often the augmentation is applied.
  • Temperature T = 4
    Section 3.3: 'We empirically found that T=4 leads to the best performance.' Used in the softmax for KL divergence.
  • Loss weights gamma and alpha = 1.0 (or 0.5 for ResNet34/50/101, VGG16)
    Section 4.2 says gamma=alpha=1.0, but smaller values are used for higher-capacity networks. Tuned per architecture.
assumptions (3)
  • domain assumption Swapping patches between same-class images preserves the validity of the original class label for both images.
    The method's loss uses the original one-hot labels for both swapped inputs (Eq. 1). If a swapped patch changes the true class, the cross-entropy loss would be noisy. The paper argues intra-class swap preserves object integrity but provides no verification.
  • domain assumption The KL divergence between the two swapped outputs provides a useful teaching signal (simulated teacher-student).
    Eq. 2 uses symmetric KL divergence as the distillation loss. The paper provides gradient-norm analyses in Section 3.4 but no formal proof that this objective improves generalization.
  • domain assumption Random patch selection creates a meaningful confidence gap between the two swapped samples.
    The mechanism requires one image to have a strong discriminative part and the other to lose it. Section 5 admits this is not guaranteed, since no attention is used to select informative patches.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Intra-class Patch Swap for Self-Distillation." pith.science (2026). https://pith.science/paper/TDW4GAI3

@misc{pith2026250514124,
  author       = {Pith},
  title        = {Pith review of: Intra-class Patch Swap for Self-Distillation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TDW4GAI3}},
  note         = {Machine review of arXiv:2505.14124}
}
read the original abstract

Knowledge distillation (KD) is a valuable technique for compressing large deep learning models into smaller, edge-suitable networks. However, conventional KD frameworks rely on pre-trained high-capacity teacher networks, which introduce significant challenges such as increased memory/storage requirements, additional training costs, and ambiguity in selecting an appropriate teacher for a given student model. Although a teacher-free distillation (self-distillation) has emerged as a promising alternative, many existing approaches still rely on architectural modifications or complex training procedures, which limit their generality and efficiency. To address these limitations, we propose a novel framework based on teacher-free distillation that operates using a single student network without any auxiliary components, architectural modifications, or additional learnable parameters. Our approach is built on a simple yet highly effective augmentation, called intra-class patch swap augmentation. This augmentation simulates a teacher-student dynamic within a single model by generating pairs of intra-class samples with varying confidence levels, and then applying instance-to-instance distillation to align their predictive distributions. Our method is conceptually simple, model-agnostic, and easy to implement, requiring only a single augmentation function. Extensive experiments across image classification, semantic segmentation, and object detection show that our method consistently outperforms both existing self-distillation baselines and conventional teacher-based KD approaches. These results suggest that the success of self-distillation could hinge on the design of the augmentation itself. Our codes are available at https://github.com/hchoi71/Intra-class-Patch-Swap.

Figures

Figures reproduced from arXiv: 2505.14124 by the authors.

Figure 1
Figure 1. Three different distillation mechanisms. Teacher-to-Student and Student-to-Student require multiple networks, while self-distillation [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Top: Illustrations of design choices when learning the single network (ResNet-18) on CIFAR100. Instead of matching function directly (Hard Label, CutMix), our method tries to match the outputs (called relaxed knowledge) coming from two swapped inputs while we still use cross-entropy loss with the hard label. Bottom: Averaged top-1 probabilities of the target label from the test data. The blue and orange graphs indic… view at source ↗
Figure 3
Figure 3. Overall framework of the proposed method. The process of generating new training inputs by exchanging patches between positive [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Magnitude of gradients in each layer from ResNet18. Patch swap, as used in our method, can indirectly alleviate the vanishing gradient [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: The averaged L1 norm of the gradient. Patch swap keeps relatively high values after 150 epochs, which suggests that the model continues [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Semantic segmentation results obtained from the RUGD dataset. The baseline method, utilizing Hard Labels, struggles to accurately [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: Three instances showcasing detection results of the baseline and the proposed method with SSD+ResNet50 on the VOC2007 test set. [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Predictive distributions on a misclassified example from CIFAR100. Our method assigns top-5 probability classes within the people [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: The activation maps generated by the Grad-CAM of the three examples of ImageNet validation set. [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: Best Top-1 accuracy for all combinations in the hyperparameters with ResNet18 on CIFAR100. [PITH_FULL_IMAGE:figures/full_fig_p021_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

72 extracted references · 64 canonical work pages

  1. [1]

    Hinton, O

    G. Hinton, O. Vinyals, J. Dean, Distilling the knowledge in a neural network, in: Proceedings of the NIPS Deep Learning and Representation Learning Workshop, 2015. URL http://arxiv.org/abs/1503.02531

  2. [2]

    Zagoruyko, N

    S. Zagoruyko, N. Komodakis, Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer, in: Proceedings of the International Conference on Learning Representations, 2017

  3. [3]

    F. Tung, G. Mori, Similarity-preserving knowledge distillation, in: Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 1365–1374. 22

  4. [4]

    Romero, N

    A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, Y . Bengio, Fitnets: Hints for thin deep nets, in: Proceedings of the International Conference on Learning Representations, 2015

  5. [5]

    C. Tan, J. Liu, Improving knowledge distillation with a customized teacher, IEEE Transactions on Neural Networks and Learning Systems (2022)

  6. [6]

    Zhang, T

    Y . Zhang, T. Xiang, T. M. Hospedales, H. Lu, Deep mutual learning, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 4320–4328

  7. [7]

    L. Yuan, F. E. Tay, G. Li, T. Wang, J. Feng, Revisiting knowledge distillation via label smoothing regularization, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 3903–3911

  8. [8]

    Zhang, J

    L. Zhang, J. Song, A. Gao, J. Chen, C. Bao, K. Ma, Be your own teacher: Improve the performance of convolutional neural networks via self distillation, in: Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 3713–3722

Show all 72 references
  1. [9]

    Moslemi, A

    A. Moslemi, A. Briskina, Z. Dang, J. Li, A survey on knowledge distillation: Recent advancements, Machine Learning with Applications (2024) 100605

  2. [10]

    Chaudhari, A

    P. Chaudhari, A. Choromanska, S. Soatto, Y . LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, R. Zecchina, Entropy-sgd: Biasing gradient descent into wide valleys, Journal of Statistical Mechanics: Theory and Experiment 2019 (12) (2019) 124018

  3. [11]

    Pereyra, G

    G. Pereyra, G. Tucker, J. Chorowski, L. Kaiser, G. E. Hinton, Regularizing neural networks by penalizing confident output distributions, in: Proceedings of the International Conference on Learning Representations Workshop, 2017. URL https://openreview.net/forum?id=HyhbYrGYe

  4. [12]

    N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, P. T. P. Tang, On large-batch training for deep learning: Generalization gap and sharp minima, in: Proceedings of the International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=H1oyRlYgg

  5. [13]

    Kumar Singh, Y

    K. Kumar Singh, Y . Jae Lee, Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization, in: Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 3524–3533

  6. [14]

    Zhang, M

    H. Zhang, M. Cisse, Y . N. Dauphin, D. Lopez-Paz, mixup: Beyond empirical risk minimization, in: Proceedings of the International Confer- ence on Learning Representations, 2018

  7. [15]

    S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, Y . Yoo, Cutmix: Regularization strategy to train strong classifiers with localizable features, in: Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 6023–6032

  8. [16]

    H. Choi, E. S. Jeon, A. Shukla, P. Turaga, Understanding the role of mixup in knowledge distillation: An empirical study, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2319–2328

  9. [17]

    W. Park, D. Kim, Y . Lu, M. Cho, Relational knowledge distillation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 3967–3976

  10. [18]

    Zhang, P.-T

    C.-B. Zhang, P.-T. Jiang, Q. Hou, Y . Wei, Q. Han, Z. Li, M.-M. Cheng, Delving deep into label smoothing, IEEE Transactions on Image Processing 30 (2021) 5984–5996

  11. [19]

    Xu, C.-L

    T.-B. Xu, C.-L. Liu, Data-distortion guided self-distillation for deep neural networks, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 33, 2019, pp. 5565–5572

  12. [20]

    Y . Shen, L. Xu, Y . Yang, Y . Li, Y . Guo, Self-distillation from the last mini-batch for consistency regularization, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 11943–11952

  13. [21]

    S. Yun, J. Park, K. Lee, J. Shin, Regularizing class-wise predictions via self-knowledge distillation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 13876–13885

  14. [22]

    C. Yang, Z. An, H. Zhou, L. Cai, X. Zhi, J. Wu, Y . Xu, Q. Zhang, Mixskd: Self-knowledge distillation from mixup for image recognition, in: European Conference on Computer Vision, Springer, 2022, pp. 534–551

  15. [23]

    W. Cui, S. Yan, Isotonic data augmentation for knowledge distillation, in: Z.-H. Zhou (Ed.), Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI), International Joint Conferences on Artificial Intelligence Organization, 2021, pp. 2314–...

  16. [24]

    T. Wang, L. Yuan, X. Zhang, J. Feng, Distilling object detectors with fine-grained feature imitation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4933–4942

  17. [25]

    D. Wang, D. Wen, J. Liu, W. Tao, T.-W. Chen, K. Osa, M. Kato, Fully supervised and guided distillation for one-stage detectors, in: Proceedings of the Asian Conference on Computer Vision (ACCV), 2020, pp. 171–188

  18. [26]

    Zhang, S

    L. Zhang, S. Huang, W. Liu, Intra-class part swapping for fine-grained image classification, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 3209–3218

  19. [27]

    J. Gou, B. Yu, S. J. Maybank, D. Tao, Knowledge distillation: A survey, International Journal of Computer Vision (IJCV) 129 (6) (2021) 1789–1819

  20. [28]

    T. Wen, S. Lai, X. Qian, Preparing lessons: Improve knowledge distillation with better supervision, Neurocomputing 454 (2021) 25–33

  21. [29]

    E. S. Jeon, H. Choi, A. Shukla, P. Turaga, Leveraging angular distributions for improved knowledge distillation, Neurocomputing 518 (2023) 466–481

  22. [30]

    K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778

  23. [31]

    Krizhevsky, G

    A. Krizhevsky, G. Hinton, et al., Learning multiple layers of features from tiny images, Technical report (2009). URL https://www.cs.toronto.edu/˜kriz/cifar.html

  24. [32]

    Simonyan, A

    K. Simonyan, A. Zisserman, Very deep convolutional networks for large-scale image recognition, in: Proceedings of the International Con- ference on Learning Representations, 2015. URL http://arxiv.org/abs/1409.1556

  25. [33]

    N. Ma, X. Zhang, H.-T. Zheng, J. Sun, Shufflenet v2: Practical guidelines for efficient cnn architecture design, in: Proceedings of the European Conference on Computer Vision, 2018, pp. 116–131

  26. [34]

    Szegedy, V

    C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, Z. Wojna, Rethinking the inception architecture for computer vision, in: Proceedings of the 23 IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 2818–2826

  27. [35]

    L. Xie, J. Wang, Z. Wei, M. Wang, Q. Tian, Disturblabel: Regularizing cnn on the loss layer, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4753–4762

  28. [36]

    X. Deng, Y . Xiao, B. Long, Z. Zhang, Reducing flipping errors in deep neural networks, in: Proceedings of the AAAI Conference on Artificial Intelligence, 2022

  29. [37]

    Y . Tian, D. Krishnan, P. Isola, Contrastive representation distillation, in: Proceedings of the International Conference on Learning Represen- tations, 2019

  30. [38]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255. URL https://image-net.org/download.php

  31. [39]

    C. Wah, S. Branson, P. Welinder, P. Perona, S. Belongie, The caltech-ucsd birds-200-2011 dataset, Technical Report CNS-TR-2011-001 (2011). URL https://www.vision.caltech.edu/datasets/cub_200_2011/

  32. [40]

    Khosla, N

    A. Khosla, N. Jayadevaprakash, B. Yao, F.-F. Li, Novel dataset for fine-grained image categorization: Stanford dogs, in: Proceedings of the CVPR workshop on fine-grained visual categorization, V ol. 2, Citeseer, 2011. URL http://vision.stanford.edu/aditya86/ImageNetDogs/

  33. [41]

    Sandler, A

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, L.-C. Chen, Mobilenetv2: Inverted residuals and linear bottlenecks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 4510–4520

  34. [42]

    Li, C.-Y

    Y . Li, C.-Y . Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, C. Feichtenhofer, Mvitv2: Improved multiscale vision transformers for classifica- tion and detection, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4804–4814

  35. [43]

    Liang, L

    J. Liang, L. Li, Z. Bing, B. Zhao, Y . Tang, B. Lin, H. Fan, Efficient one pass self-distillation with zipf’s label smoothing, in: Proceedings of the European Conference on Computer Vision, 2022, pp. 104–119

  36. [44]

    DeVries, G

    T. DeVries, G. W. Taylor, Improved regularization of convolutional neural networks with cutout, arXiv preprint arXiv:1708.04552 (2017)

  37. [45]

    L.-C. Chen, Y . Zhu, G. Papandreou, F. Schroff, H. Adam, Encoder-decoder with atrous separable convolution for semantic image segmenta- tion, in: Proceedings of the European Conference on Computer Vision, 2018, pp. 801–818

  38. [46]

    Everingham, L

    M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, A. Zisserman, The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results. URL http://www.pascal-network.org/challenges/VOC/voc2012/workshop/index.html

  39. [47]

    Cordts, M

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, B. Schiele, The cityscapes dataset for semantic urban scene understanding, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 3213–3223. URL http...

  40. [48]

    Wigness, S

    M. Wigness, S. Eum, J. G. Rogers, D. Han, H. Kwon, A rugd dataset for autonomous navigation and visual perception in unstructured outdoor environments, in: Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, 2019, pp. 5000–5007. URL http://r...

  41. [49]

    T. Guan, D. Kothandaraman, R. Chandra, A. J. Sathyamoorthy, K. Weerakoon, D. Manocha, Ga-nav: Efficient terrain segmentation for robot navigation in unstructured outdoor environments, IEEE Robotics and Automation Letters 7 (3) (2022) 8138–8145

  42. [50]

    H. Zhao, J. Shi, X. Qi, X. Wang, J. Jia, Pyramid scene parsing network, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2881–2890

  43. [51]

    J. Long, E. Shelhamer, T. Darrell, Fully convolutional networks for semantic segmentation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3431–3440

  44. [52]

    J. Fu, J. Liu, H. Tian, Y . Li, Y . Bao, Z. Fang, H. Lu, Dual attention network for scene segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3146–3154

  45. [53]

    Y . Yuan, X. Chen, J. Wang, Object-contextual representations for semantic segmentation, in: Proceedings of the European conference on computer vision (ECCV), 2020, pp. 173–190

  46. [54]

    H. Zhao, Y . Zhang, S. Liu, J. Shi, C. C. Loy, D. Lin, J. Jia, Psanet: Point-wise spatial attention network for scene parsing, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 267–283

  47. [55]

    C. Yu, C. Gao, J. Wang, G. Yu, C. Shen, N. Sang, Bisenet v2: Bilateral network with guided aggregation for real-time semantic segmentation, International Journal of Computer Vision 129 (2021) 3051–3068

  48. [56]

    T. Wu, S. Tang, R. Zhang, J. Cao, Y . Zhang, Cgnet: A light-weight context guided network for semantic segmentation, IEEE Transactions on Image Processing 30 (2020) 1169–1179

  49. [57]

    R. P. Poudel, S. Liwicki, R. Cipolla, Fast-scnn: Fast semantic segmentation network, arXiv preprint arXiv:1902.04502 (2019)

  50. [58]

    H. Wu, J. Zhang, K. Huang, K. Liang, Y . Yu, Fastfcn: Rethinking dilated convolution in the backbone for semantic segmentation, arXiv preprint arXiv:1903.11816 (2019)

  51. [59]

    Zheng, J

    S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y . Wang, Y . Fu, J. Feng, T. Xiang, P. H. Torr, et al., Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recogni- tio...

  52. [60]

    Ranftl, A

    R. Ranftl, A. Bochkovskiy, V . Koltun, Vision transformers for dense prediction, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 12179–12188

  53. [61]

    E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, P. Luo, Segformer: Simple and efficient design for semantic segmentation with transformers, Advances in Neural Information Processing Systems 34 (2021) 12077–12090

  54. [62]

    Strudel, R

    R. Strudel, R. Garcia, I. Laptev, C. Schmid, Segmenter: Transformer for semantic segmentation, in: Proceedings of the IEEE/CVF interna- tional conference on computer vision, 2021, pp. 7262–7272

  55. [63]

    W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y . Fu, A. C. Berg, Ssd: Single shot multibox detector, in: Proceedings of the European Conference on Computer Vision, 2016, pp. 21–37

  56. [64]

    Everingham, L

    M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, A. Zisserman, The PASCAL Visual Object Classes Challenge 2007 (VOC2007) 24 Results. URL http://www.pascal-network.org/challenges/VOC/voc2007/workshop/index.html

  57. [65]

    X. Yuan, P. He, Q. Zhu, X. Li, Adversarial examples: Attacks and defenses for deep learning, IEEE transactions on neural networks and learning systems 30 (9) (2019) 2805–2824

  58. [66]

    Geifman, G

    Y . Geifman, G. Uziel, R. El-Yaniv, Bias-reduced uncertainty estimation for deep neural classifiers, in: Proceedings of the International Conference on Learning Representations, 2018

  59. [67]

    M. P. Naeini, G. Cooper, M. Hauskrecht, Obtaining well calibrated probabilities using bayesian binning, in: Proceedings of the AAAI Conference on Artificial Intelligence, 2015

  60. [68]

    Gneiting, A

    T. Gneiting, A. E. Raftery, Strictly proper scoring rules, prediction, and estimation, Journal of the American statistical Association 102 (477) (2007) 359–378

  61. [69]

    Y . Wang, X. Ma, Z. Chen, Y . Luo, J. Yi, J. Bailey, Symmetric cross entropy for robust learning with noisy labels, in: Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 322–330

  62. [70]

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, Grad-cam: Visual explanations from deep networks via gradient- based localization, in: Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 618–626

  63. [71]

    Bengio, J

    Y . Bengio, J. Louradour, R. Collobert, J. Weston, Curriculum learning, in: Proceedings of the 26th annual international conference on machine learning, 2009, pp. 41–48

  64. [72]

    C. Finn, P. Abbeel, S. Levine, Model-agnostic meta-learning for fast adaptation of deep networks, in: International conference on machine learning, PMLR, 2017, pp. 1126–1135. 25

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.