Pith. sign in

REVIEW 5 major objections 6 minor 109 references

LayerMix: Enhanced Data Augmentation through Fractal Integration for Robust Deep Learning

T0 review · 5 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read LayerMix mixes training images with grayscale fractals and claims simultaneous gains in robustness, calibration, and consistency.

desk verdict Plausible augmentation recipe whose own tables undercut its headline claims; worth a serious referee only if the authors fix baselines, add error bars, and tone down the abstract. read the letter →

arxiv 2501.04861 v2 pith:APU72CPE submitted 2025-01-08 cs.CV

classification cs.CV
keywords dataaugmentationfractalscorruptionrobustnessadversarialmodelcalibrationpredictionconsistencyimageclassificationlabel-preservingmixing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces LayerMix, a data augmentation pipeline that blends training images with grayscale fractal patterns in three layers of increasing diversity. Its claim is that this structured mixing improves clean accuracy and, at the same time, several machine-learning safety metrics: robustness to natural corruptions and renditions, adversarial robustness, calibration, and prediction consistency. On CIFAR-100 the reported mean corruption error drops from 50.0 (baseline) to 30.3, and calibration error from 31.2 to 5.9; on ImageNet-200 the paper reports the first augmentation strategy in its comparison to beat the baseline on every safety metric at once. The reason these gains would matter is that they suggest robustness can be obtained from the training pipeline itself, without extra data, adversarial training, or new architectures.

What carries the argument

The central object is the LayerMix pipeline itself: a three-layer graph in which every Aug block shares one randomly sampled augmentation function (the covariance structure), every Blend block independently samples from a reweighted mixture of arithmetic mean, geometric mean, pixel mixing, and element mixing, and the deepest layer mixes in a grayscale fractal. The load-bearing identity is the auto-covariance expression of Eq. (5), $\mathrm{Cov}(X_i,X_j)=\mathbb{E}_k[\sigma^2_{ki}]+\mathbb{E}_k[\mu^2_{ki}]-\mathbb{E}_k[\mu_{ki}]^2$ on the diagonal and $\mathbb{E}_k[\mu_{ki}\mu_{kj}]-\mathbb{E}_k[\mu_{ki}]\mathbb{E}_k[\mu_{kj}]$ off the diagonal, which couples the pipeline's diversity to low-order statistics of the transformations. This lets the authors argue that stage-wise covariance removes redundant diversity without removing useful signal, while the grayscale fractals supply label-preserving structural complexity that does not impose new labels the way MixUp-style interpolation does.

What would settle it

Re-run the CIFAR-100 and ImageNet-200 comparisons with IPMix using its original consistency loss and the hyperparameters its authors recommend, then check whether its mean corruption error and RMS calibration error meet or beat LayerMix's reported values (30.3 and 5.9 on CIFAR-100; 29.77 mCE on ImageNet-200). If they do, the claim that LayerMix achieves comprehensive Pareto improvements over baselines is a comparison artifact.

Watch

Extended reading notes

Core claim

The paper's central claim is that LayerMix generates "semantically consistent" synthetic samples by correlating the augmentation stages instead of letting them act independently. The pipeline draws one random transformation $f_k$, applies it to the original image, blends two differently augmented copies, and then blends that result with a random grayscale fractal; the final training image is chosen uniformly among the three layers. Mathematically, the marginal distribution becomes $p_{\mathrm{layermix}}(x)=\mathbb{E}_k[\prod_n f_k(x_n)]$ rather than the IID form $\prod_n \mathbb{E}_k[f_k(x_n)]$, which produces nonzero off-diagonal covariance between stages, and the paper shows that this lets the method cut redundant diversity while raising the augmentation magnitude. The empirical assertion is that on CIFAR-10, CIFAR-100, ImageNet-200 and ImageNet-1K this pipeline beats the baselines and the PixMix and IPMix pipelines on corruption robustness, prediction consistency, and calibration while staying competitive on clean accuracy and adversarial error.

Load-bearing premise

The load-bearing premise is that the comparisons against PixMix and IPMix are fair: the paper trains IPMix without the extra consistency loss its authors normally use, and fixes both baselines to the same blending and mixing hyperparameters, so the small reported margins could vanish if the baselines were run with their own settings.

Editorial extensions

If this is right

  • Models trained with LayerMix no longer need extra data or adversarial training to become substantially more robust to natural corruptions: the reported CIFAR-100-C mCE drops from 50.0 to 30.3.
  • Calibration improves enough that confidence estimates become usable under distribution shift, with the reported CIFAR-100 RMS calibration error falling from 31.2 to 5.9.
  • The gains transfer across architectures (WRN-40-4, WRN-28-10, ResNet-18, ResNeXt-29, ResNet-50), so the method is not tied to one model family.
  • Combining LayerMix with the JSD consistency loss and longer training lowers errors further, which means training budget and augmentation design interact rather than being independent choices.
  • On ImageNet-200 the paper reports the first augmentation strategy in its comparison to beat the cropping-and-flipping baseline on every tested safety metric, so practitioners would not have to trade accuracy for robustness.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A follow-up comparison not reported in the paper would run IPMix with its original consistency loss and tuned hyperparameters, which would test whether the gains come from the pipeline structure itself or from dropping that loss term.
  • If stage covariance is the real driver, injecting the same shared-transform covariance into other pipelines such as AugMix or CutMix should reproduce part of the gain; that experiment is not in the paper.
  • Because the paper uses grayscale fractals, a natural extension is to check whether the benefit comes from fractal structure or simply from color-free high-frequency noise; replacing fractals with random sinusoidal or wavelet textures would separate those two hypotheses.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper introduces LayerMix, a label-preserving data augmentation pipeline that combines correlated augmentation stages, grayscale fractal images, and a reweighted mixture of blending methods, building directly on PixMix and IPMix. The authors derive a closed-form auto-covariance expression for the pipeline's joint input-output distribution (Eqs. (3)-(5)), propose a three-sample pipeline with controlled diversity, and evaluate on CIFAR-10, CIFAR-100, ImageNet-200, and ImageNet-1K across clean accuracy, corruption robustness, consistency, adversarial robustness, and calibration. The central claims are that LayerMix generates semantically consistent synthetic samples, achieves superior classification accuracy, and is the first augmentation strategy with comprehensive Pareto improvements over baselines across diverse safety criteria.

Significance. If the empirical claims were fully supported, LayerMix would be a practically useful augmentation method because it targets multiple safety metrics simultaneously and the authors release code and training metadata. The covariance calculation in Eqs. (3)-(5) is a clean mathematical identity and is a useful formalization of the proposed pipeline structure. The paper also runs a broad benchmark matrix, including multiple architectures and corruption benchmarks, which is valuable for the community. However, the advertised superiority and Pareto-improvement claims are directly undermined by the paper's own tables, and the comparison protocol for prior methods is not demonstrated to be faithful to those methods' intended configurations.

major comments (5)
  1. [Section 4.2.3, Table 5] Table 5 contradicts the paper's central Pareto claim: when IPMix is run with its own JSD loss and grayscale fractals, IPMix achieves mCE 28.29 versus LayerMix 28.89 and mFP 4.57 versus 5.30. Thus LayerMix is not a comprehensive Pareto improvement over the strongest prior fractal method, and the abstract's assertion of 'superior performance in classification accuracy' and Section 4.3.2's claim of 'comprehensive Pareto improvements over baseline measurements across diverse safety criteria' are not supported by the reported experiments.
  2. [Section 4.2.1 and Section 4.3.1, Tables 3, 4, 9] The baseline comparisons are not demonstrated to use the prior methods' recommended settings: IPMix is trained without its JSD loss 'for a fair comparison,' and fixed values beta=3, k=3 (CIFAR) and beta=4, k=4, m=1 (ImageNet) are assigned to all PixMix and IPMix runs. Because Table 5 shows that IPMix with its original JSD loss outperforms LayerMix on corruption robustness and consistency, the headline margins in Tables 3, 4, and 9 may be configuration artifacts rather than properties of the LayerMix pipeline. The authors should either run the baselines at their published settings or explicitly justify every deviation and show that the conclusions are robust to those changes.
  3. [Section 3.1, Eqs. (3)-(5)] The covariance derivation is algebraically correct but is disconnected from the empirical claims: no bound, mechanism, or experiment is provided that links the covariance structure of the augmentation pipeline to improved accuracy or robustness. The sentence 'This approach decreases sample diversity without hindering performance' is an assertion, not a consequence of the derivation. Since pipeline covariance is presented as a primary contribution, the paper should either provide a theoretical link to generalization or demonstrate empirically, through an ablation that contrasts correlated stages with IID stages under matched magnitude, that the covariance structure causes the reported gains.
  4. [Tables 4 and 9] The claim of superior clean accuracy is contradicted by the paper's own numbers: on CIFAR-100 with WRN-40-4, LayerMix has clean error 20.7 while PixMix and IPMix both have 20.4, and on ImageNet-1K, IPMix has clean error 22.51 versus LayerMix's 23.57 (m=8, beta=3) or 23.53 (m=1, beta=4). Thus the abstract's statement that experiments demonstrate 'superior performance in classification accuracy' is inaccurate; at best, LayerMix is comparable or slightly worse on clean accuracy in these settings.
  5. [Section 4.2.2, Table 6] The reported gains over baselines are small relative to the hyperparameter spread shown in Table 6: across LayerMix's own magnitude-blending combinations, clean error varies with mean 20.66 and standard deviation 0.15, and mCE varies with mean 30.70 and standard deviation 0.37. The margins in Table 4 (e.g., 20.7 versus 20.4 clean error; 30.3 versus 30.8 mCE) are within this range, so the paper should report error bars or multiple seeds, and should clarify whether the selected m=8, beta=3 configuration is meaningfully better than other configurations when compared against the baselines.
minor comments (6)
  1. [Abstract and Section 4.2.2] The abstract claims 'superior performance in classification accuracy' without qualification, but the detailed tables show that LayerMix is often second-best in clean accuracy; the wording should be revised to match the actual results.
  2. [Section 4.2.1] The text says 'magnitude = 8and blending ratio = 3' with missing spaces, and later 'we used β = 3and k = 3'; these typos should be corrected.
  3. [Section 4.2.3, Table 5 caption] The caption uses '• •' and '• • •' to denote colored and grayscale fractals, but the symbols are not explained in the caption or the body text; please define them explicitly.
  4. [Section 4.3.2, Table 8] The text states 'LayerMix represents the first augmentation strategy to achieve comprehensive Pareto improvements over baseline measurements across diverse safety criteria,' but Table 5 already shows a counterexample within the paper; this sentence should be removed or carefully qualified.
  5. [Various] There are several typographical errors, including 'inclunding' (Section 4), 'Imagnet-200-C' (Section 4.3.2), and 'CIFAR-10- C' (Section 4.1.1); these should be corrected in a final pass.
  6. [Section 3.4, Code Block 1] The pseudocode uses 'random.randint(3)' which in Python returns 0, 1, or 2, but the text describes three samples; the intended range and the uniform selection over samples 1, 2, and 3 should be made explicit to avoid off-by-one ambiguity.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation found: LayerMix's covariance analysis is a mathematical identity, and its empirical comparisons rest on external baselines rather than self-citation.

full rationale

The paper contains no load-bearing circular step. The main analytic claim, the auto-covariance computation in Section 3.1, is a direct algebraic identity relating Cov(X_i, X_j) to moments of the transformation distributions; it does not presuppose the paper's empirical conclusions about clean accuracy or robustness. The blending weights in Table 1 and hyperparameter choices (m=8, beta=3) are motivated by ablation studies and a hyperparameter scan on the same benchmarks, which is in-sample tuning rather than a fitted parameter later renamed as a prediction, so it does not qualify as circularity under the stated standards. Comparisons to PixMix and IPMix rely on the authors' re-implementations, including the deliberate omission of IPMix's JSD loss for 'fair comparison'; this is a measurement and fairness concern, not a circularity of derivation, and the paper's own Table 5 even shows IPMix with JSD loss outperforming LayerMix on some safety metrics. No uniqueness theorem is imported from the authors' prior work, and no self-citation is load-bearing. The paper's limitations section candidly identifies unresolved issues such as fractal availability and scalability. Therefore, despite possible correctness and benchmarking objections, no specific reduction of a claimed result to its own inputs is present.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

There are no new physical entities. The central claim rests on four tuned values, plus three design assumptions that are motivated by observed improvements rather than derived from first principles. The covariance analysis in Section 3.1 is a calculation, not a proof that these choices improve generalization.

free parameters (4)
  • augmentation magnitude m = 8 (CIFAR and ImageNet LayerMix runs)
    Controls the strength of the transformations in Table 2; set by the authors in Sections 4.2.1 and 4.3.1 and scanned in Table 6. Performance varies with m, so the reported results depend on this choice.
  • blending ratio beta = 3 for LayerMix
    Controls the conic combination in arithmetic and geometric blending; set in Sections 4.2.1 and 4.3.1. Table 6 shows beta=5 can give lower clean error, but beta=3 is used for standardization.
  • blending method probabilities = Arithmetic 33.3%, geometric 33.3%, pixel 16.6%, element 16.6%
    Chosen 'through extensive ablation studies' in Section 3.3, but no ablation data are shown. This distribution directly controls the generated samples and is tuned to favor clean accuracy while retaining robustness.
  • comparison hyperparameters for PixMix and IPMix = CIFAR: beta=3, k=3; ImageNet: beta=4, k=4, magnitude=1
    These settings are assigned by the authors to reproduce the baselines in Section 4.2.1 and Section 4.3.1. If they understate or alter the original methods, the fairness of the comparison and the size of the reported gains are affected.
assumptions (4)
  • ad hoc to paper Applying the same transformation at every augmentation stage reduces diversity without harming performance, and this structure improves generalization.
    Section 3.1 concludes with 'This approach decreases sample diversity without hindering performance,' but no derivation or ablation links the covariance structure to test accuracy.
  • domain assumption Fractal images are label-preserving and unlikely to cause manifold intrusion.
    Adopted from PixMix and the manifold intrusion literature cited in Section 2 and Section B.3; the paper does not re-derive or test this premise.
  • domain assumption Gray-scale fractals retain useful structural complexity while removing color information, and this improves augmentation quality.
    Stated in Section 3.2 with empirical support in Table 5, but the mechanism is not analyzed and the effect is not isolated from other pipeline changes.
  • domain assumption Single training runs without reported seeds or confidence intervals are treated as reliable point estimates.
    All main tables report one number per method; small margins such as 0.1 points on CIFAR-10 in Table 3 could be within seed noise.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LayerMix: Enhanced Data Augmentation through Fractal Integration for Robust Deep Learning." pith.science (2026). https://pith.science/paper/APU72CPE

@misc{pith2026250104861,
  author       = {Pith},
  title        = {Pith review of: LayerMix: Enhanced Data Augmentation through Fractal Integration for Robust Deep Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/APU72CPE}},
  note         = {Machine review of arXiv:2501.04861}
}
read the original abstract

Deep learning models have demonstrated remarkable performance across various computer vision tasks, yet their vulnerability to distribution shifts remains a critical challenge. Despite sophisticated neural network architectures, existing models often struggle to maintain consistent performance when confronted with Out-of-Distribution (OOD) samples, including natural corruptions, adversarial perturbations, and anomalous patterns. We introduce LayerMix, an innovative data augmentation approach that systematically enhances model robustness through structured fractal-based image synthesis. By meticulously integrating structural complexity into training datasets, our method generates semantically consistent synthetic samples that significantly improve neural network generalization capabilities. Unlike traditional augmentation techniques that rely on random transformations, LayerMix employs a structured mixing pipeline that preserves original image semantics while introducing controlled variability. Extensive experiments across multiple benchmark datasets, including CIFAR-10, CIFAR-100, ImageNet-200, and ImageNet-1K demonstrate LayerMixs superior performance in classification accuracy and substantially enhances critical Machine Learning (ML) safety metrics, including resilience to natural image corruptions, robustness against adversarial attacks, improved model calibration and enhanced prediction consistency. LayerMix represents a significant advancement toward developing more reliable and adaptable artificial intelligence systems by addressing the fundamental challenges of deep learning generalization. The code is available at https://github.com/ahmadmughees/layermix.

Figures

Figures reproduced from arXiv: 2501.04861 by the authors.

Figure 1
Figure 1. Comparison of different augmentation pipelines. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Complete Pipeline of LayerMix. The resulting image produced by the LayerMix pipeline is uniformly selected from [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Corruption Error values for various methods on CIFAR-100-C. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Individual Corruption Error (CE) values for various methods on ImageNet-200-C [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Corruption Error (CE) values for various methods on CIFAR-100- [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: Samples of grayscale fractals used in LayerMix. [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: Samples generated from the 3 different layers of LayerMix. [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: Samples generated from LayerMix [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: GradCam [97] class activation map visualizations from a ResNet-50 model across 5 randomly selected ImageNet samples for various augmentation pipelines. The models evaluated in [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

109 extracted references · 70 canonical work pages

  1. [1]

    Autoaugment: Learning augmentation strategies from data,

    E. D. Cubuk, B. Zoph, D. Mane, V . Vasudevan, and Q. V . Le, “Autoaugment: Learning augmentation strategies from data,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2019

  2. [2]

    Randaugment: Practical automated data augmentation with a reduced search space,

    E. D. Cubuk, B. Zoph, J. Shlens, and Q. V . Le, “Randaugment: Practical automated data augmentation with a reduced search space,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , June 2020

  3. [3]

    Random erasing data augmentation,

    Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang, “Random erasing data augmentation,” in AAAI Conference on Artificial Intelligence, 2017

  4. [4]

    Zero-shot image classification based on deep feature extraction,

    X. Wang, C. Chen, Y. Cheng, and Z. J. Wang, “Zero-shot image classification based on deep feature extraction,” IEEE Transactions on Cognitive and Developmental Systems, vol. 10, pp. 432–444, 2018

  5. [5]

    LiT: Zero-shot transfer with locked-image text tuning,

    X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer, “LiT: Zero-shot transfer with locked-image text tuning,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 18 102–18 112, 2021

  6. [6]

    Open-vocabulary object detection via vision and language knowledge distillation,

    X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui, “Open-vocabulary object detection via vision and language knowledge distillation,” in International Conference on Learning Representations, 2021

  7. [7]

    VoxNet: A 3D convolutional neural network for real-time object recognition,

    D. Maturana and S. A. Scherer, “VoxNet: A 3D convolutional neural network for real-time object recognition,” 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 922–928, 2015

  8. [8]

    PIXOR: Real-time 3D object detection from point clouds,

    B. Yang, W. Luo, and R. Urtasun, “PIXOR: Real-time 3D object detection from point clouds,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7652–7660, 2018

Show all 109 references
  1. [9]

    Voxel transformer for 3D object detection,

    J. Mao, Y. Xue, M. Niu, H. Bai, J. Feng, X. Liang, H. Xu, and C. Xu, “Voxel transformer for 3D object detection,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3144–3153, 2021

  2. [10]

    NormFace: L2 hypersphere embedding for face verification,

    F. Wang, X. Xiang, J. Cheng, and A. L. Yuille, “NormFace: L2 hypersphere embedding for face verification,” Proceedings of the 25th ACM international conference on Multimedia, 2017

  3. [11]

    CosFace: Large margin cosine loss for deep face recognition,

    H. Wang, Y. Wang, Z. Zhou, X. Ji, Z. Li, D. Gong, J. Zhou, and W. Liu, “CosFace: Large margin cosine loss for deep face recognition,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5265–5274, 2018

  4. [12]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778

  5. [13]

    Wide Residual Networks,

    S. Zagoruyko and N. Komodakis, “Wide Residual Networks,” in British Machine Vision Conference 2016 . York, France: British Machine Vision Association, Jan. 2016. [Online]. Available: https://enpc.hal.science/hal-01832503

  6. [14]

    Inception-v4, inception-resnet and the impact of residual connections on learning,

    C. Szegedy, S. Ioffe, V . Vanhoucke, and A. A. Alemi, “Inception-v4, inception-resnet and the impact of residual connections on learning,” in Thirty-First AAAI Conference on Artificial Intelligence , 2017

  7. [15]

    EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,

    M. Tan and Q. Le, “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” in Proceedings of the 36th International Conference on Machine Learning. PMLR, May 2019, pp. 6105–6114

  8. [16]

    YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information,

    C.-Y. Wang, I.-H. Yeh, and H.-Y. M. Liao, “YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information,” Feb. 2024

  9. [17]

    Efficientdet: Scalable and efficient object detection,

    M. Tan, R. Pang, and Q. V . Le, “Efficientdet: Scalable and efficient object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020

  10. [18]

    Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137–1149, Jun. 2017

  11. [19]

    Mask R-CNN,

    K. He, G. Gkioxari, P . Dollar, and R. Girshick, “Mask R-CNN,” in 2017 IEEE International Conference on Computer Vision (ICCV) . IEEE, Oct. 2017, pp. 2980–2988

  12. [20]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P . Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015, N. Navab, J. Hornegger, W. M. Wells, and A. F. Frangi, Eds. Cham: Springer International Pu...

  13. [21]

    SAM 2: Segment Anything in Images and Videos,

    N. Ravi, V . Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R ¨adle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V . Alwala, N. Carion, C.-Y. Wu, R. Girshick, P . Doll´ar, and C. Feichtenhofer, “SAM 2: Segment Anything in Images and Videos,” Oct. 2024

  14. [22]

    Do ImageNet classifiers generalize to ImageNet?

    B. Recht, R. Roelofs, L. Schmidt, and V . Shankar, “Do ImageNet classifiers generalize to ImageNet?” in International Conference on Machine Learning, 2019

  15. [23]

    Robustbench: a standardized adversarial robustness benchmark,

    F. Croce, M. Andriushchenko, V . Sehwag, E. Debenedetti, N. Flammarion, M. Chiang, P . Mittal, and M. Hein, “Robustbench: a standardized adversarial robustness benchmark,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round ...

  16. [24]

    Unsolved problems in ML safety,

    D. Hendrycks, N. Carlini, J. Schulman, and J. Steinhardt, “Unsolved problems in ML safety,” ArXiv, vol. abs/2109.13916, 2021

  17. [25]

    Benchmarking Neural Network Robustness to Common Corruptions and Perturbations,

    D. Hendrycks and T. Dietterich, “Benchmarking Neural Network Robustness to Common Corruptions and Perturbations,” in International Conference on Learning Representations, 2018

  18. [26]

    On interaction between augmentations and corruptions in natural corruption robustness,

    E. Mintun, A. Kirillov, and S. Xie, “On interaction between augmentations and corruptions in natural corruption robustness,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P . Liang, and J. W. Vaughan, Eds., vol. 34. Curran Assoc...

  19. [27]

    On detecting adversarial perturbations,

    J. H. Metzen, T. Genewein, V . Fischer, and B. Bischoff, “On detecting adversarial perturbations,” in International Conference on Learning Representations, 2017. [Online]. Available: https://openreview.net/forum?id=SJzCSf9xg

  20. [28]

    Network anomaly detection using LSTM based autoencoder,

    M. S. Elsayed, N.-A. Le-Khac, S. Dev, and A. D. Jurcut, “Network anomaly detection using LSTM based autoencoder,” Proceedings of the 16th ACM Symposium on QoS and Security for Wireless and Mobile Networks , 2020

  21. [29]

    KeepAugment: A simple information-preserving data augmentation approach,

    C. Gong, D. Wang, M. Li, V . Chandra, and Q. Liu, “KeepAugment: A simple information-preserving data augmentation approach,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 1055–1064, 2020

  22. [30]

    Defending against image corruptions through adversarial augmentations,

    D. A. Calian, F. Stimberg, O. Wiles, S.-A. Rebuffi, A. Gy ¨orgy, T. A. Mann, and S. Gowal, “Defending against image corruptions through adversarial augmentations,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum...

  23. [31]

    Puzzle mix: Exploiting saliency and local statistics for optimal mixup,

    J.-H. Kim, W. Choo, and H. O. Song, “Puzzle mix: Exploiting saliency and local statistics for optimal mixup,” in Proceedings of the 37th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, H. D. III and A. Singh, Eds., vol. 119. PMLR, ...

  24. [32]

    Evaluating model robustness and stability to dataset shift,

    A. Subbaswamy, R. J. Adams, and S. Saria, “Evaluating model robustness and stability to dataset shift,” in International Conference on Artificial Intelligence and Statistics, 2021

  25. [33]

    A Comprehensive Survey on Data Augmentation,

    Z. Wang, P . Wang, K. Liu, P . Wang, Y. Fu, C.-T. Lu, C. C. Aggarwal, J. Pei, and Y. Zhou, “A Comprehensive Survey on Data Augmentation,” May 2024

  26. [34]

    Data augmentation: A comprehensive survey of modern approaches,

    A. Mumuni and F. Mumuni, “Data augmentation: A comprehensive survey of modern approaches,” Array, vol. 16, p. 100258, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S2590005622000911

  27. [35]

    Fast AutoAugment,

    S. Lim, I. Kim, T. Kim, C. Kim, and S. Kim, “Fast AutoAugment,” in Neural Information Processing Systems, 2019

  28. [36]

    AutoAugment: Learning augmentation strategies from data,

    E. D. Cubuk, B. Zoph, D. Man ´e, V . Vasudevan, and Q. V . Le, “AutoAugment: Learning augmentation strategies from data,”2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 113–123, 2019

  29. [37]

    mixup: Beyond empirical risk minimization,

    H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in International Conference on Learning Representations, 2018. [Online]. Available: https://openreview.net/forum?id=r1Ddp1-Rb

  30. [38]

    RICAP: Random image cropping and patching data augmentation for deep cnns,

    R. Takahashi, T. Matsubara, and K. Uehara, “RICAP: Random image cropping and patching data augmentation for deep cnns,” in Asian Conference on Machine Learning, 2018

  31. [39]

    Co-mixup: Saliency guided joint mixup with supermodular diversity,

    J. Kim, W. Choo, H. Jeong, and H. O. Song, “Co-mixup: Saliency guided joint mixup with supermodular diversity,” in International Conference on Learning Representations, 2021. [Online]. Available: https://openreview.net/forum?id=gvxJzw8kW4b

  32. [40]

    PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures,

    D. Hendrycks, A. Zou, M. Mazeika, L. Tang, B. Li, D. Song, and J. Steinhardt, “PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 783–16 792

  33. [41]

    IPMix: Label-preserving data augmentation method for training robust classifiers,

    Z. Huang, X. Bao, N. Zhang, Q. Zhang, X. Tu, B. Wu, and X. Yang, “IPMix: Label-preserving data augmentation method for training robust classifiers,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA:...

  34. [42]

    DiffuseMix: Label-Preserving Data Augmentation with Diffusion Models,

    K. Islam, M. Z. Zaheer, A. Mahmood, and K. Nandakumar, “DiffuseMix: Label-Preserving Data Augmentation with Diffusion Models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 27 621–27 630

  35. [43]

    GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing,

    K. Islam, M. Z. Zaheer, A. Mahmood, K. Nandakumar, and N. Akhtar, “GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing,” Dec. 2024

  36. [44]

    Measures of complexity: A nonexhaustive list,

    S. Lloyd, “Measures of complexity: A nonexhaustive list,” IEEE Control Systems Magazine, vol. 21, no. 4, pp. 7–8, Aug. 2001

  37. [45]

    Pre-training without natural images,

    H. Kataoka, K. Okayasu, A. Matsumoto, E. Yamagata, R. Yamada, N. Inoue, A. Nakamura, and Y. Satoh, “Pre-training without natural images,” in Proceedings of the Asian Conference on Computer Vision (ACCV) , November 2020

  38. [46]

    Can Vision Transformers Learn without Natural Images?

    K. Nakashima, H. Kataoka, A. Matsumoto, K. Iwata, N. Inoue, and Y. Satoh, “Can Vision Transformers Learn without Natural Images?” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 2, pp. 1990–1998, Jun. 2022

  39. [47]

    MixUp as Locally Linear Out-of-Manifold Regularization,

    H. Guo, Y. Mao, and R. Zhang, “MixUp as Locally Linear Out-of-Manifold Regularization,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, pp. 3714–3722, Jul. 2019

  40. [48]

    Local Mixup: Interpolation of closest input signals to prevent manifold intrusion,

    R. Baena, L. Drumetz, and V . Gripon, “Local Mixup: Interpolation of closest input signals to prevent manifold intrusion,” Signal Processing, vol. 219, p. 109395, Jun. 2024

  41. [49]

    Natural adversarial examples,

    D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. X. Song, “Natural adversarial examples,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 15 257–15 266, 2019

  42. [50]

    The many faces of robustness: A critical analysis of out-of-distribution generalization,

    D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. L. Zhu, S. Parajuli, M. Guo, D. X. Song, J. Steinhardt, and J. Gilmer, “The many faces of robustness: A critical analysis of out-of-distribution generalization,” 2021 IEEE/CVF International Conferen...

  43. [51]

    Gradient-based learning applied to document recognition,

    Y. LeCun, L. Bottou, Y. Bengio, and P . Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998

  44. [52]

    Imagenet classification with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems, 2012, pp. 1097–1105

  45. [53]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V . Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garn...

  46. [54]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” Jun. 2021

  47. [55]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 012–10 022

  48. [56]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P . Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020

  49. [57]

    Deep unsupervised learning using nonequilibrium thermodynamics,

    J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in International Conference on Machine Learning. PMLR, 2015, pp. 2256–2265

  50. [58]

    Deep double descent: Where bigger models and more data hurt,

    P . Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever, “Deep double descent: Where bigger models and more data hurt,” Journal of Statistical Mechanics: Theory and Experiment , vol. 2021, no. 12, p. 124003, 2021

  51. [59]

    Affinity and diversity: Quantifying mechanisms of data augmentation,

    R. Gontijo-Lopes, S. J. Smullin, E. D. Cubuk, and E. Dyer, “Affinity and diversity: Quantifying mechanisms of data augmentation,” arXiv preprint arXiv:2002.08973, 2020

  52. [60]

    Investigating the effectiveness of data augmentation from similarity and diversity: An empirical study,

    S. Yang, S. Guo, J. Zhao, and F. Shen, “Investigating the effectiveness of data augmentation from similarity and diversity: An empirical study,” Pattern Recognition, vol. 148, p. 110204, 2024

  53. [61]

    Manifold mixup: Better representations by interpolating hidden states,

    V . Verma, A. Lamb, C. Beckham, A. Najafi, I. Mitliagkas, D. Lopez-Paz, and Y. Bengio, “Manifold mixup: Better representations by interpolating hidden states,” in International Conference on Machine Learning, 2018

  54. [62]

    Vicinal risk minimization,

    O. Chapelle, J. Weston, L. Bottou, and V . Vapnik, “Vicinal risk minimization,” Advances in neural information processing systems, vol. 13, 2000

  55. [63]

    AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty,

    D. Hendrycks*, N. Mu*, E. D. Cubuk, B. Zoph, J. Gilmer, and B. Lakshminarayanan, “AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty,” in International Conference on Learning Representations, Sep. 2019

  56. [64]

    Adversarial attacks on medical machine learning,

    S. G. Finlayson, J. D. Bowers, J. Ito, J. L. Zittrain, A. L. Beam, and I. S. Kohane, “Adversarial attacks on medical machine learning,” Science, vol. 363, no. 6433, pp. 1287–1289, 2019. [Online]. Available: https://www.science.org/doi/abs/10.1126/science.aaw4399

  57. [65]

    SoK: Security and privacy in machine learning,

    N. Papernot, P . Mcdaniel, A. Sinha, and M. P . Wellman, “SoK: Security and privacy in machine learning,” 2018 IEEE European Symposium on Security and Privacy (EuroS&P), pp. 399–414, 2018

  58. [66]

    Applying network analysis to explore the global scientific literature on food security,

    L. Skaf, E. Buonocore, S. Dumontet, R. Capone, and P . P . Franzese, “Applying network analysis to explore the global scientific literature on food security,” Ecol. Informatics, vol. 56, p. 101062, 2020

  59. [67]

    EMMA: End-to-End Multimodal Model for Autonomous Driving,

    J.-J. Hwang, R. Xu, H. Lin, W.-C. Hung, J. Ji, K. Choi, D. Huang, T. He, P . Covington, B. Sapp, Y. Zhou, J. Guo, D. Anguelov, and M. Tan, “EMMA: End-to-End Multimodal Model for Autonomous Driving,” Nov. 2024. 17

  60. [68]

    Improving Agent Behaviors with RL Fine-Tuning for Autonomous Driving,

    Z. Peng, W. Luo, Y. Lu, T. Shen, C. Gulino, A. Seff, and J. Fu, “Improving Agent Behaviors with RL Fine-Tuning for Autonomous Driving,” in Computer Vision – ECCV 2024 , A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds. Cham: Springer Nature Switze...

  61. [69]

    Language is not all you need: Aligning perception with language models,

    S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra, Q. Liu, K. Aggarwal, Z. Chi, N. Bjorck, V . Chaudhary, S. Som, X. SONG, and F. Wei, “Language is not all you need: Aligning perception with language models,” in Advances in Neural I...

  62. [70]

    Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,

    Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,” in Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran...

  63. [71]

    Gpt-4 technical report,

    OpenAI, J. Achiam, S. Adler et al., “Gpt-4 technical report,” 2024. [Online]. Available: https://arxiv.org/abs/2303.08774

  64. [72]

    LLaMA: Open and Efficient Foundation Language Models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “LLaMA: Open and Efficient Foundation Language Models,” Feb. 2023

  65. [73]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255

  66. [74]

    Boosting adversarial attacks with momentum,

    Y. Dong, F. Liao, T. Pang, H. Su, J. Zhu, X. Hu, and J. Li, “Boosting adversarial attacks with momentum,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9185–9193, 2017

  67. [75]

    Feature denoising for improving adversarial robustness,

    C. Xie, Y. Wu, L. van der Maaten, A. L. Yuille, and K. He, “Feature denoising for improving adversarial robustness,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 501–509, 2018

  68. [76]

    Adversarial examples improve image recognition,

    C. Xie, M. Tan, B. Gong, J. Wang, A. L. Yuille, and Q. V . Le, “Adversarial examples improve image recognition,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 816–825, 2019

  69. [77]

    Visda-2021 competition: Universal domain adaptation to improve performance on out-of-distribution data,

    D. Bashkirova, D. Hendrycks, D. Kim, H. Liao, S. Mishra, C. Rajagopalan, K. Saenko, K. Saito, B. U. Tayyab, P . Teterwak, and B. Usman, “Visda-2021 competition: Universal domain adaptation to improve performance on out-of-distribution data,” in Proceedings of the NeurIPS 2021 ...

  70. [78]

    A unifying review of deep and shallow anomaly detection,

    L. Ruff, J. R. Kauffmann, R. A. Vandermeulen, G. Montavon, W. Samek, M. Kloft, T. G. Dietterich, and K.-R. M ¨uller, “A unifying review of deep and shallow anomaly detection,” Proceedings of the IEEE, vol. 109, no. 5, pp. 756–795, 2021

  71. [79]

    A fourier perspective on model robustness in computer vision,

    D. Yin, R. Gontijo Lopes, J. Shlens, E. D. Cubuk, and J. Gilmer, “A fourier perspective on model robustness in computer vision,” in Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch ´e-Buc, E. Fox, and R. Garnett, Eds., vo...

  72. [80]

    Distributionally robust neural networks,

    S. Sagawa*, P . W. Koh*, T. B. Hashimoto, and P . Liang, “Distributionally robust neural networks,” in International Conference on Learning Representations, 2020. [Online]. Available: https://openreview.net/forum?id=ryxGuJrFvS

  73. [81]

    Wilds: A benchmark of in-the-wild distribution shifts,

    P . W. Koh, S. Sagawa, H. Marklund, S. M. Xie, M. Zhang, A. Balsubramani, W. Hu, M. Yasunaga, R. L. Phillips, I. Gao, T. Lee, E. David, I. Stavness, W. Guo, B. Earnshaw, I. Haque, S. M. Beery, J. Leskovec, A. Kundaje, E. Pierson, S. Levine, C. Finn, and P . Liang, “Wilds: A be...

  74. [82]

    Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness

    R. Geirhos, P . Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel, “Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness.” in International Conference on Learning Representations , 2019. [Online]. Available: ht...

  75. [83]

    On calibration of modern neural networks,

    C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proceedings of the 34th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70. PMLR, 06–11 Aug 2017, pp...

  76. [84]

    Accurate uncertainties for deep learning using calibrated regression,

    V . Kuleshov, N. Fenner, and S. Ermon, “Accurate uncertainties for deep learning using calibrated regression,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, J. Dy and A. Krause, Eds., vol. 80. PMLR, 10–...

  77. [85]

    Can you trust your model's uncertainty? evaluating predictive uncertainty under dataset shift,

    Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. Dillon, B. Lakshminarayanan, and J. Snoek, “Can you trust your model's uncertainty? evaluating predictive uncertainty under dataset shift,” in Advances in Neural Information Processing Systems , H. Wallach, H. L...

  78. [86]

    Improving fractal pre-training,

    C. Anderson and R. Farrell, “Improving fractal pre-training,” 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pp. 2412–2421, 2021

  79. [87]

    Learning multiple layers of features from tiny images,

    A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” 2009

  80. [88]

    Imagenet large scale visual recognition challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, and M. Bernstein, “Imagenet large scale visual recognition challenge,” International Journal of Computer Vision, vol. 115, no. 3, pp. 211–252, 2015

  81. [89]

    Some terminology and notation in information theory,

    I. J. Good, “Some terminology and notation in information theory,” Proceedings of the IEE-Part C: Monographs , vol. 103, no. 3, pp. 200–204, 1956

  82. [90]

    CutMix: Regularization strategy to train strong classifiers with localizable features,

    S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. J. Yoo, “CutMix: Regularization strategy to train strong classifiers with localizable features,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6022–6031, 2019

  83. [91]

    Towards deep learning models resistant to adversarial attacks,

    A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in International Conference on Learning Representations, 2018. [Online]. Available: https://openreview.net/forum?id=rJzIBfZAb

  84. [92]

    Deep anomaly detection with outlier exposure,

    D. Hendrycks, M. Mazeika, and T. Dietterich, “Deep anomaly detection with outlier exposure,” in International Conference on Learning Representations, 2019

  85. [93]

    Posterior calibration and exploratory analysis for natural language processing models,

    K. Nguyen and B. O’Connor, “Posterior calibration and exploratory analysis for natural language processing models,” in EMNLP. Lisbon, Portugal: Association for Computational Linguistics, 2015, pp. 1587–1598. [Online]. Available: https: //www.aclweb.org/anthology/D15-1182

  86. [94]

    SGDR: Stochastic gradient descent with warm restarts,

    I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” in International Conference on Learning Representations ,

  87. [95]

    Aggregated residual transformations for deep neural networks,

    S. Xie, R. Girshick, P . Dollar, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017

  88. [96]

    Tiny imagenet visual recognition challenge,

    Y. Le and X. Yang, “Tiny imagenet visual recognition challenge,” CS 231N, vol. 7, no. 7, p. 3, 2015

  89. [97]

    Grad-cam: Visual explanations from deep networks via gradient-based localization,

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 618–626. 18

  90. [98]

    Adversarial autoaugment,

    X. Zhang, Q. Wang, J. Zhang, and Z. Zhong, “Adversarial autoaugment,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=ByxdUySKvS

  91. [99]

    TrivialAugment: Tuning-free yet state-of-the-art data augmentation,

    S. G. M ¨uller and F. Hutter, “TrivialAugment: Tuning-free yet state-of-the-art data augmentation,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 754–762, 2021

  92. [100]

    Improved image augmentation for convolutional neural networks by copyout and copypairing,

    P . May, “Improved image augmentation for convolutional neural networks by copyout and copypairing,” 2020. [Online]. Available: https://openreview.net/forum?id=rkxWpCNKvS

  93. [101]

    Saliencymix: A saliency guided data augmentation strategy for better regularization,

    A. F. M. S. Uddin, M. S. Monira, W. Shin, T. Chung, and S.-H. Bae, “Saliencymix: A saliency guided data augmentation strategy for better regularization,” in International Conference on Learning Representations , 2021. [Online]. Available: https: //openreview.net/forum?id=-M0QkvBGTTq

  94. [102]

    AutoMix: Unveiling the power of mixup for stronger classifiers,

    Z. Liu, S. Li, D. Wu, Z. Chen, L. Wu, J. Guo, and S. Z. Li, “AutoMix: Unveiling the power of mixup for stronger classifiers,” in European Conference on Computer Vision, 2021

  95. [103]

    TokenMix: Rethinking image mixing for data augmentation in vision transformers,

    J. Liu, B. Liu, H. Zhou, H. Li, and Y. Liu, “TokenMix: Rethinking image mixing for data augmentation in vision transformers,” in European Conference on Computer Vision, 2022

  96. [104]

    SmoothMix: A Simple Yet Effective Data Augmentation to Train Robust Classifiers,

    J.-H. Lee, M. Z. Zaheer, M. Astrid, and S.-I. Lee, “SmoothMix: A Simple Yet Effective Data Augmentation to Train Robust Classifiers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , 2020, pp. 756–757

  97. [105]

    Fmix: Enhancing mixed sample data augmentation,

    E. Harris, A. Marcu, M. Painter, M. Niranjan, A. Prugel-Bennett, and J. Hare, “Fmix: Enhancing mixed sample data augmentation,” 2021. [Online]. Available: https://openreview.net/forum?id=oev4KdikGjy

  98. [106]

    CutPaste: Self-supervised learning for anomaly detection and localization,

    C.-L. Li, K. Sohn, J. Yoon, and T. Pfister, “CutPaste: Self-supervised learning for anomaly detection and localization,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9659–9669, 2021

  99. [107]

    Nonaccidental properties underlie human categorization of complex natural scenes,

    D. B. Walther and D. Shen, “Nonaccidental properties underlie human categorization of complex natural scenes,” Psychological Science , vol. 25, no. 4, pp. 851–860, 2014, pMID: 24474725. [Online]. Available: https://doi.org/10.1177/0956797613512662 APPENDIX A MORE SAMPLES . Fig...

  100. [176]

    PMLR, 06–14 Dec 2022, pp. 66–79. [Online]. Available: https://proceedings.mlr.press/v176/bashkirova22a.html

  101. [2017]

    Available: https://openreview.net/forum?id=Skq89Scxx

    [Online]. Available: https://openreview.net/forum?id=Skq89Scxx

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.