REVIEW 1 major objections 5 minor 83 references
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
T0 review · 1 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Occluding the input regions a model relies on most during training improves generalization and robustness.
desk verdict The image experiments show a small but credible gain from LRP-guided occlusion; the point-cloud experiments are undermined by a masking-equation flaw that makes β not do what the paper claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the normalized LRP relevance map, $R_{\mathrm{norm}}$, turned into a dropout probability distribution over input features. For 2D images, the mask is a rectangle of Random-Erasing-style dimensions (Eq. 5) centered on the single pixel with maximum normalized relevance (Eq. 4), with the dataset mean as the replacement value; for point clouds, the mask thresholds a convex combination of a uniform random vector and the normalized relevance vector, controlled by hyperparameters $\alpha$ (random vs. attribution-guided) and $\beta$ (total fraction removed). This mechanism adapts during training: as the model changes which features it relies on, the attribution map changes, so the augmentation always targets the current over-relied-upon features.
What would settle it
Train with the RelDrop pipeline but replace the LRP map with a random mask drawn from the same marginal statistics as the real LRP maps (same rectangle sizes and center distribution); if test accuracy, zero-shot transfer, and point-removal robustness match RelDrop's reported gains, the attribution signal is not load-bearing. A more direct test: shuffle the LRP maps across training samples so each image receives another image's attribution, which breaks the sample-specific localization while preserving map statistics.
Extended reading notes
Core claim
The paper's central claim is that occlusion-based data augmentation is more effective when the occluded region is chosen by attribution rather than at random. RelDrop builds a binary mask from a normalized LRP relevance map, occluding a rectangular block centered on the most relevant pixel for images, or replacing the most relevant points with the origin for point clouds, with hyperparameters that mix in random masking to avoid destroying all informative signal. This informed dropout yields consistent test-accuracy improvements over Random Erasing on every model and dataset tested, roughly doubling RE's improvement over the unaugmented baseline, and produces models whose attributions spread over more channels in deeper layers and whose zero-shot performance on distribution-shifted ImageNet variants improves. The authors interpret this as evidence that RelDrop counteracts overfitting by forcing reliance on a wider set of features.
Load-bearing premise
The load-bearing premise is that the normalized LRP relevance map correctly identifies the features the model has overfitted on, so that masking the most relevant region targets the problematic regularities; if attributions are noisy, point to background, or disagree with the features actually driving the prediction, RelDrop degrades toward random occlusion or distorts training.
Editorial extensions
If this is right
- RelDrop can replace Random Erasing as a drop-in regularizer: it needs no architectural change and consistently improves test accuracy, with average margins over baseline of 0.64%, 0.89%, and 0.93% on CIFAR-10, CIFAR-100, and ImageNet-1k respectively.
- RelDrop-trained ImageNet models improve zero-shot accuracy on ImageNet-R, ImageNet-A, and ImageNet-O by average relative margins of 2.88%, 8.83%, and 0.63% over baseline, indicating more transferable features.
- RelDrop-trained PointNet++ models are more robust to ordered point removal, retaining accuracy substantially longer than baseline as points are progressively set to the origin.
- The method costs one extra backward pass per batch (2–2.5× compute), and the authors recommend removing only a small fraction of points ($\beta=0.15$) to preserve discriminative signal.
Reading between the lines
- If the mechanism is forcing feature diversity, combining RelDrop with complementary augmentations such as CutMix or Mixup could yield additive gains; this is a direct test of the proposed mechanism.
- The attribution-quality sensitivity shown by the $\varepsilon$ ablation suggests that cheaper or noisier attribution estimators may not preserve RelDrop's benefits, so an approximate-attribution variant should be validated before use at scale.
- The point cloud formulation suggests a general recipe for any modality with removable input features; tabular and graph data, where feature masking is natural, are unexplored testbeds.
- The higher Relevance Rank Accuracy (RRA) on ImageNet-S implies less reliance on background; a causal test would be to finetune with RelDrop on a dataset with deliberately spurious background cues and measure whether background reliance drops.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RelDrop, a data augmentation technique that uses LRP attributions to mask the most relevant input regions during training. The method is evaluated on 2D image classification (CIFAR-10/100, ImageNet-1k with zero-shot transfer to ImageNet-R/A/O) and 3D point cloud classification (ModelNet40, ShapeNet) using ResNet and PointNet++ architectures. The authors report consistent accuracy improvements over a no-augmentation baseline and Random Erasing, improved mean Relevance Rank Accuracy, and greater robustness to point removal. An additional backward pass per batch is required, and the method introduces hyperparameters specific to the attribution method and to the balance between random and attribution-guided masking.
Significance. If the claims hold, RelDrop is a simple, computationally modest regularization technique that adds to the growing body of XAI-guided training methods. The paper's strengths include a clearly specified algorithm, public code, multi-dataset and multi-architecture experiments, an evaluation against ground-truth segmentation masks for RRA, and an explicit limitations section. However, the point-cloud experiments contain a confound in the definition of the masking hyperparameter β that affects the validity of the claimed robustness and generalization results for that domain. As the point-cloud results are a distinct, load-bearing pillar of the paper's central claim, the current evidence is insufficient to support the abstract's blanket conclusions.
major comments (1)
- [Section 3.2, Eq. (9); Table 2; Figure 4] The parameter β, defined as 'the overall fraction of points replaced', does not actually control the masked fraction when 0<α<1. In Eq. (9), a point is masked when α v_i + (1−α) R^3d_norm(i) ≥ (1−β). For α=0.5 and R^3d_norm roughly uniform on [0,1], the expected masked fraction is P(v+R ≥ 2(1−β)): for β=0.15 this is (0.3)^2/2 ≈ 4.5%, not 15%, and for β=0.85 it is 1 − (0.3)^2/2 ≈ 95.5%. The row α=0.5, β=0.15 in Table 2 therefore removes about one third as many points as the α=1.0, β=0.15 random baseline, while α=0.5, β=0.85 removes more points than the corresponding random baseline. The conclusion that α=0.5 is superior to α=1.0 may reflect differences in effective occlusion strength rather than a benefit of mixing random and relevance-guided masking, and the sharp degradation at α=0.5, β=0.85 is explained by near-total removal. Please reparameterize the mask so that the expected masked fraction equals β for all α (e.g., threshold at the (1−β)-quantile of αv+(1−α)R_norm, or explicitly sample a β fraction of points), rerun the point-cloud experiments and Figure 4, and report the empirical masked fractions for each configuration.
minor comments (5)
- [Table 1 caption] The sentence 'RelDrop improves upon the estimated model generalization ability in all investigated settings compared to RE and the baseline and increases Mean RRA, except for Mean RRA' is garbled; from the text of Section 4.1.1, 'except for ResNet101' appears to be intended.
- [Figure A.1 caption] The Center panel is described with the same parameters as the Left panel (α=0.5, β=0.15); it should probably be α=0.5, β=0.5, and the phrase 'and the,' before 'Right' contains a typo.
- [A.1 vs Table 3] Section A.1 states that LRP uses ε=1e−6 unless stated otherwise, but Table 3 lists ε=0.001 for the ResNet experiments; please reconcile these values.
- [Algorithm A.1] The lines 'I^{2d*} ← − I^{2d}' appear to contain a spurious minus sign, and the boundary check only ensures xcen+W_O ≤ W and ycen+H_O ≤ H without checking the lower bounds xcen−W_O/2 ≥ 0 and ycen−H_O/2 ≥ 0; please clarify how out-of-bounds rectangles are handled.
- [Section 4.1.1, Table 1] ImageNet-1k and zero-shot results are reported for a single randomly chosen seed, and several absolute gains are very small (e.g., ResNet-18 ImageNet-O 15.66→15.67); please provide multi-seed estimates or a statistical assessment to support the consistency claims, or explicitly temper the conclusions.
Circularity Check
No significant circularity: RelDrop's claimed gains are independent empirical outcomes, not constructed from its inputs.
full rationale
RelDrop's training mask (Eqs. 1-9) is computed from LRP attributions and a random vector; these are inputs to the augmentation, not the claimed outputs. The claimed outcomes—test accuracy, zero-shot accuracy, RRA, and point-flipping robustness—are all measured on held-out data or external ground-truth masks (e.g., ImageNet-S segmentation) after training with a standard classification loss. Attributions are recomputed each batch from the current model rather than fitted once to the evaluation metric, so no parameter is fitted and then renamed as a prediction. Although several cited tools and metrics (LRP, zennit, Quantus, RRA) share authors with this paper, the RRA evaluation is anchored to external ground-truth segmentation masks, and the LRP choice is motivated by prior empirical success rather than by a uniqueness or authority argument. The point-cloud hyperparameter discrepancy noted by the skeptic (Eq. 9 does not make beta the actual masked fraction when alpha<1) is a potential experimental confound, not a circularity: Table 2's comparisons may be at unequal perturbation strengths, but the reported accuracies are nevertheless independently measured outcomes. No load-bearing step reduces, by the paper's own equations, a claimed prediction to its own input.
Assumptions & free parameters
free parameters (6)
- p (image dropout probability) =
0.5
- S_low, S_high (erasing area range) =
S_high=0.4; S_low not explicitly stated in main text
- r_low (aspect ratio lower bound) =
0.3
- epsilon (LRP stabilizer) =
0.001
- alpha (point cloud random vs attribution balance) =
0.5 (best)
- beta (fraction of points replaced) =
0.15 (best)
assumptions (3)
- domain assumption LRP relevance maps faithfully identify the input features that the model currently relies on for its prediction.
- domain assumption The specific LRP rule composite (flat-rule first layer, epsilon/z+ for other layers, batchnorm canonization) produces attributions suitable for guiding augmentation.
- domain assumption Occluding the most relevant input features during training forces the model to learn a broader set of informative features without harming convergence.
Cite this review
Pith. "Pith review of Relevance-driven Input Dropout: an Explanation-guided Regularization Technique." pith.science (2026). https://pith.science/paper/NQMZAD64
@misc{pith2026250521595,
author = {Pith},
title = {Pith review of: Relevance-driven Input Dropout: an Explanation-guided Regularization Technique},
year = {2026},
howpublished = {\url{https://pith.science/paper/NQMZAD64}},
note = {Machine review of arXiv:2505.21595}
}
read the original abstract
Overfitting is a well-known issue extending even to state-of-the-art (SOTA) Machine Learning (ML) models, resulting in reduced generalization, and a significant train-test performance gap. Mitigation measures include a combination of dropout, data augmentation, weight decay, and other regularization techniques. Among the various data augmentation strategies, occlusion is a prominent technique that typically focuses on randomly masking regions of the input during training. Most of the existing literature emphasizes randomness in selecting and modifying the input features instead of regions that strongly influence model decisions. We propose Relevance-driven Input Dropout (RelDrop), a novel data augmentation method which selectively occludes the most relevant regions of the input, nudging the model to use other important features in the prediction process, thus improving model generalization through informed regularization. We further conduct qualitative and quantitative analyses to study how Relevance-driven Input Dropout (RelDrop) affects model decision-making. Through a series of experiments on benchmark datasets, we demonstrate that our approach improves robustness towards occlusion, results in models utilizing more features within the region of interest, and boosts inference time generalization performance. Our code is available at https://github.com/Shreyas-Gururaj/LRP_Relevance_Dropout.
Figures
Reference graph
Works this paper leans on
-
[1]
From attribution maps to human-understandable explanations through concept relevance propagation.Nature Machine Intelligence, 5(9):1006–1019, 2023
Reduan Achtibat, Maximilian Dreyer, Ilona Eisenbraun, Sebastian Bosse, Thomas Wiegand, Wojciech Samek, and Sebastian Lapuschkin. From attribution maps to human-understandable explanations through concept relevance propagation.Nature Machine Intelligence, 5(9):1006–1019, 2023
2023
-
[2]
Attnlrp: Attention-aware layer-wise relevance propagation for transformers
Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebastian Lapuschkin, and Wojciech Samek. Attnlrp: Attention-aware layer-wise relevance propagation for transformers. InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024
2024
-
[3]
Anders, David Neumann, Wojciech Samek, Klaus-Robert Müller, and Sebastian Lapuschkin
Christopher J. Anders, David Neumann, Wojciech Samek, Klaus-Robert Müller, and Sebastian Lapuschkin. Software for dataset-wide xai: From local explanations to global insights with Zennit, CoRelAy, and ViRelAy.CoRR, abs/2106.13200, 2021
arXiv 2021
-
[4]
CLEVR-XAI: A benchmark dataset for the ground truth evaluation of neural network explanations.Inf
Leila Arras, Ahmed Osman, and Wojciech Samek. CLEVR-XAI: A benchmark dataset for the ground truth evaluation of neural network explanations.Inf. Fusion, 81:14–40, 2022
2022
-
[5]
Lei Jimmy Ba and Brendan J. Frey. Adaptive dropout for training deep neural networks. In Christopher J. C. Burges, Léon Bottou, Zoubin Ghahramani, and Kilian Q. Weinberger (eds.),Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems
-
[6]
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation.PLoS ONE, 10(7):e0130140, 2015
work page 2015
-
[7]
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert Müller. How to explain individual classification decisions.Journal of Machine Learning Research, 11: 1803–1831, 2010
work page 2010
-
[8]
Network dissection: Quanti- fying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. Network dissection: Quanti- fying interpretability of deep visual representations. In2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, pp. 3319–3327. IEEE Computer Society, 2017
work page 2017
Show all 83 references
-
[9]
Ecq x: Explainability-driven quantization for low-bit and sparse dnns
Daniel Becking, Maximilian Dreyer, Wojciech Samek, Karsten Müller, and Sebastian Lapuschkin. Ecq x: Explainability-driven quantization for low-bit and sparse dnns. In Andreas Holzinger, Randy Goebel, Ruth Fong, Taesup Moon, Klaus-Robert Müller, and Wojciech Samek (eds.),xxAI -...
2020
-
[10]
Calmon, and Himabindu Lakkaraju
Usha Bhalla, Alex Oesterling, Suraj Srinivas, Flávio P. Calmon, and Himabindu Lakkaraju. Interpreting CLIP with sparse linear concept embeddings (splice). In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (eds.),...
2024
-
[11]
Gunn, Alexander Hammers, David Alexander Dickie, Maria del C
Christopher Bowles, Liang Chen, Ricardo Guerrero, Paul Bentley, Roger N. Gunn, Alexander Hammers, David Alexander Dickie, Maria del C. Valdés Hernández, Joanna M. Wardlaw, and Daniel Rueckert. GAN augmentation: Augmenting training data using generative adversarial networks.CoR...
-
[12]
Artificial intelligence in medicine: today and tomorrow.Frontiers in Medicine, 7:509744, 2020
Giovanni Briganti and Olivier Le Moine. Artificial intelligence in medicine: today and tomorrow.Frontiers in Medicine, 7:509744, 2020
2020
-
[13]
Roberts, and Chris C
Alexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J. Roberts, and Chris C. Holmes. Explicit regularisation in gaussian noise injections. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (eds.),Advances in Neural Informat...
2020
-
[14]
Chang, Thomas A
Angel X. Chang, Thomas A. Funkhouser, Leonidas J. Guibas, Pat Hanrahan, Qi-Xing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. Shapenet: An information-rich 3d model repository.CoRR, abs/1512.03012, 2015. 11
2015 arXiv
-
[15]
Cubuk, Barret Zoph, Dandelion Mané, Vijay Vasudevan, and Quoc V
Ekin D. Cubuk, Barret Zoph, Dandelion Mané, Vijay Vasudevan, and Quoc V . Le. Autoaugment: Learning augmentation strategies from data. InIEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pp. 113–123. Computer Vision F...
2019
-
[16]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA, pp. 248–255. IEEE C...
2009
-
[17]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...
2021
-
[18]
Mechanistic understanding and validation of large AI models with semanticlens.CoRR, abs/2501.05398, 2025
Maximilian Dreyer, Jim Berend, Tobias Labarta, Johanna Vielhaben, Thomas Wiegand, Sebastian La- puschkin, and Wojciech Samek. Mechanistic understanding and validation of large AI models with semanticlens.CoRR, abs/2501.05398, 2025
2025 arXiv
-
[19]
Explain to not forget: Defending against catastrophic forgetting with XAI
Sami Ede, Serop Baghdadlian, Leander Weber, An Nguyen, Dario Zanca, Wojciech Samek, and Sebastian Lapuschkin. Explain to not forget: Defending against catastrophic forgetting with XAI. In Andreas Holzinger, Peter Kieseberg, A Min Tjoa, and Edgar R. Weippl (eds.),Machine Learni...
2022
-
[20]
Fong and Andrea Vedaldi
Ruth C. Fong and Andrea Vedaldi. Interpretable explanations of black boxes by meaningful perturbation. InIEEE International Conference on Computer Vision, ICCV 2017, pp. 3449–3457. IEEE Computer Society, 2017
2017
-
[21]
Shanghua Gao, Zhong-Yu Li, Ming-Hsuan Yang, Ming-Ming Cheng, Junwei Han, and Philip H. S. Torr. Large-scale unsupervised semantic segmentation.IEEE Trans. Pattern Anal. Mach. Intell., 45(6): 7457–7476, 2023
2023
-
[22]
The missing curve detectors of inceptionv1: Applying sparse autoencoders to inceptionv1 early vision.CoRR, abs/2406.03662, 2024
Liv Gorton. The missing curve detectors of inceptionv1: Applying sparse autoencoders to inceptionv1 early vision.CoRR, abs/2406.03662, 2024
2024 arXiv
-
[23]
Martin, and Shi-Min Hu
Meng-Hao Guo, Junxiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R. Martin, and Shi-Min Hu. PCT: point cloud transformer.Comput. Vis. Media, 7(2):187–199, 2021
2021
-
[24]
Pruning by explaining revisited: Optimizing attribution methods to prune cnns and transformers.CoRR, abs/2408.12568, 2024
Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Reduan Achtibat, Thomas Wiegand, Wojciech Samek, and Sebastian Lapuschkin. Pruning by explaining revisited: Optimizing attribution methods to prune cnns and transformers.CoRR, abs/2408.12568, 2024
2024 arXiv
-
[25]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, pp. 770–778. IEEE Computer Society, 2016
2016
-
[26]
Anna Hedström, Leander Weber, Daniel Krakowczyk, Dilyara Bareeva, Franz Motzkus, Wojciech Samek, Sebastian Lapuschkin, and Marina M.-C. Höhne. Quantus: An explainable AI toolkit for responsible evaluation of neural network explanations and beyond.J. Mach. Learn. Res., 24:34:1–...
2023
-
[27]
The many faces of robustness: A critical analysis of out-of-distribution generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. In202...
2021
-
[28]
Natural adversarial examples
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. InIEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pp. 15262–15271. Computer Vision Foundation / IEEE, 2021
2021
-
[29]
Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Im- proving neural networks by preventing co-adaptation of feature detectors.CoRR, abs/1207.0580, 2012
2012 arXiv
-
[30]
Summit: Scaling deep learning interpretability by visualizing activation and attribution summarizations.IEEE Trans
Fred Hohman, Haekyu Park, Caleb Robinson, and Duen Horng (Polo) Chau. Summit: Scaling deep learning interpretability by visualizing activation and attribution summarizations.IEEE Trans. Vis. Comput. Graph., 26(1):1096–1106, 2020
2020
-
[31]
Architecture disentanglement for deep neural networks
Jie Hu, Liujuan Cao, Tong Tong, Qixiang Ye, Shengchuan Zhang, Ke Li, Feiyue Huang, Ling Shao, and Rongrong Ji. Architecture disentanglement for deep neural networks. pp. 652–661, 2021. 12
2021
-
[32]
Sparse autoencoders find highly interpretable features in language models
Robert Huben, Hoagy Cunningham, Logan Riggs, Aidan Ewart, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024
2024
-
[33]
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Francis R. Bach and David M. Blei (eds.),Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 Ju...
2015
-
[34]
Patchshuffle regularization.CoRR, abs/1707.07103, 2017
Guoliang Kang, Xuanyi Dong, Liang Zheng, and Yi Yang. Patchshuffle regularization.CoRR, abs/1707.07103, 2017
2017 arXiv
-
[35]
Cai, James Wexler, Fernanda B
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie J. Cai, James Wexler, Fernanda B. Viégas, and Rory Sayres. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCA V). InProceedings of the 35th International Conference on Machin...
2018
-
[36]
Dropout as data augmentation.CoRR, abs/1506.08700, 2015
Kishore Reddy Konda, Xavier Bouthillier, Roland Memisevic, and Pascal Vincent. Dropout as data augmentation.CoRR, abs/1506.08700, 2015
2015 arXiv
-
[37]
Derpanis, and Pavel Tokmakov
Matthew Kowal, Achal Dave, Rares Ambrus, Adrien Gaidon, Konstantinos G. Derpanis, and Pavel Tokmakov. Understanding video transformers via universal concept discovery. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 20...
2024
-
[38]
Krizhevsky and G
A. Krizhevsky and G. Hinton. Learning multiple layers of features from tiny images.Master’s thesis, Department of Computer Science, University of Toronto, 2009
2009
-
[39]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classification with deep convolutional neural networks. In Peter L. Bartlett, Fernando C. N. Pereira, Christopher J. C. Burges, Léon Bottou, and Kilian Q. Weinberger (eds.),Advances in Neural Information Process...
2012
-
[40]
Anders Krogh and John A. Hertz. A simple weight decay can improve generalization. In John E. Moody, Stephen Jose Hanson, and Richard Lippmann (eds.),Advances in Neural Information Processing Systems 4, [NIPS Conference, Denver, Colorado, USA, December 2-5, 1991], pp. 950–957. ...
1991
-
[41]
Improvement in deep networks for optimization using explainable artificial intelligence
Jin Ha Lee, Ik hee Shin, Sang gu Jeong, Seung-Ik Lee, Muhamamad Zaigham Zaheer, and Beom-Su Seo. Improvement in deep networks for optimization using explainable artificial intelligence. In2019 International Conference on Information and Communication Technology Convergence, IC...
2019
-
[42]
R-drop: Regularized dropout for neural networks
Xiaobo Liang, Lijun Wu, Juntao Li, Yue Wang, Qi Meng, Tao Qin, Wei Chen, Min Zhang, and Tie-Yan Liu. R-drop: Regularized dropout for neural networks. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (eds.),Advances in Neura...
2021
-
[43]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019
2019
-
[44]
Lundberg and Su-In Lee
Scott M. Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V . N. Vishwanathan, and Roman Garnett (eds.),Advances in Neural Information Processing Systems 30: Annua...
2017
-
[45]
Explaining nonlinear classification decisions with deep taylor decomposition.Pattern Recognition, 65: 211–222, 2017
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller. Explaining nonlinear classification decisions with deep taylor decomposition.Pattern Recognition, 65: 211–222, 2017
2017
-
[46]
Layer-wise relevance propagation: An overview
Grégoire Montavon, Alexander Binder, Sebastian Lapuschkin, Wojciech Samek, and Klaus-Robert Müller. Layer-wise relevance propagation: An overview. In Wojciech Samek, Grégoire Montavon, Andrea Vedaldi, Lars Kai Hansen, and Klaus-Robert Müller (eds.),Explainable AI: Interpreting...
2019
-
[47]
Measurably stronger explanation reliability via model canonization
Franz Motzkus, Leander Weber, and Sebastian Lapuschkin. Measurably stronger explanation reliability via model canonization. In2022 IEEE International Conference on Image Processing, ICIP 2022, Bordeaux, France, 16-19 October 2022, pp. 516–520. IEEE, 2022. 13
2022
-
[48]
Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton. When does label smoothing help? In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (eds.),Advances in Neural Information Processing Systems 32: Annual Conference...
2019
-
[49]
xAI-GAN: Enhancing Generative Adversarial Networks via Explainable AI Systems.CoRR, abs/2002.10438, 2020
Vineel Nagisetty, Laura Graves, Joseph Scott, and Vijay Ganesh. xAI-GAN: Enhancing Generative Adversarial Networks via Explainable AI Systems.CoRR, abs/2002.10438, 2020
2002 arXiv
-
[50]
Regularizing deep neural networks by noise: Its interpretation and optimization
Hyeonwoo Noh, Tackgeun You, Jonghwan Mun, and Bohyung Han. Regularizing deep neural networks by noise: Its interpretation and optimization. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V . N. Vishwanathan, and Roman Garnett (eds.),Advanc...
2017
-
[51]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leik...
2022
-
[52]
Charles Ruizhongtai Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pp. 77–85. IEEE ...
2017
-
[53]
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J. Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V . N. Vishwanathan, and Roman Garnett (eds.),Adv...
2017
-
[54]
why should I trust you?
Marco T. Ribeiro, Sameer Singh, and Carlos Guestrin. "why should I trust you?": Explaining the predictions of any classifier. InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, SIGKDD 2016, pp. 1135–1144. ACM, 2016
2016
-
[55]
Utilizing explainable AI for quantization and pruning of deep neural networks.CoRR, abs/2008.09072, 2020
Muhammad Sabih, Frank Hannig, and Jürgen Teich. Utilizing explainable AI for quantization and pruning of deep neural networks.CoRR, abs/2008.09072, 2020
2008 arXiv
-
[56]
Evaluating the visualization of what a deep neural network has learned.IEEE Trans
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, and Klaus-Robert Müller. Evaluating the visualization of what a deep neural network has learned.IEEE Trans. Neural Networks Learn. Syst., 28(11):2660–2673, 2017
2017
-
[57]
Learning important features through propa- gating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. Learning important features through propa- gating activation differences. InProceedings of the 34th International Conference on Machine Learning, ICML 2017, volume 70 ofProceedings of Machine Learning Research, pp. 3145–3...
2017
-
[58]
Hide-and-seek: Forcing a network to be meticulous for weakly- supervised object and action localization
Krishna Kumar Singh and Yong Jae Lee. Hide-and-seek: Forcing a network to be meticulous for weakly- supervised object and action localization. InIEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pp. 3544–3553. IEEE Computer Society, 2017
2017
-
[59]
Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting.J. Mach. Learn. Res., 15(1):1929–1958, 2014
1929
-
[60]
Explaining prediction models and individual predictions with feature contributions.Knowledge and Information Systems, 41(3):647–665, 2014
Erik Štrumbelj and Igor Kononenko. Explaining prediction models and individual predictions with feature contributions.Knowledge and Information Systems, 41(3):647–665, 2014
2014
-
[61]
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. InProceedings of the 34th International Conference on Machine Learning, ICML 2017, volume 70 ofProceedings of Machine Learning Research, pp. 3319–3328. PMLR, 2017
2017
-
[62]
Multi-dimensional concept discovery (MCD): A unifying framework with completeness guarantees.Trans
Johanna Vielhaben, Stefan Bluecher, and Nils Strodthoff. Multi-dimensional concept discovery (MCD): A unifying framework with completeness guarantees.Trans. Mach. Learn. Res., 2023, 2023
2023
-
[63]
Analyzing multi-head self- attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena V oita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self- attention: Specialized heads do the heavy lifting, the rest can be pruned. In Anna Korhonen, David R. Traum, and Lluís Màrquez (eds.),Proceedings of the 57th Conference of the ...
2019
-
[64]
Zeiler, Sixin Zhang, Yann LeCun, and Rob Fergus
Li Wan, Matthew D. Zeiler, Sixin Zhang, Yann LeCun, and Rob Fergus. Regularization of neural networks using dropconnect. InProceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013, volume 28 ofJMLR Workshop and Conferen...
2013
-
[65]
Sarma, Michael M
Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E. Sarma, Michael M. Bronstein, and Justin M. Solomon. Dynamic graph CNN for learning on point clouds.ACM Trans. Graph., 38(5):146:1–146:12, 2019
2019
-
[66]
Efficient and flexible neural network training through layer-wise feedback propagation.CoRR, abs/2308.12053, 2025
Leander Weber, Jim Berend, Moritz Weckbecker, Alexander Binder, Thomas Wiegand, Wojciech Samek, and Sebastian Lapuschkin. Efficient and flexible neural network training through layer-wise feedback propagation.CoRR, abs/2308.12053, 2025
2025 arXiv
-
[67]
Wei and Kai Zou
Jason W. Wei and Kai Zou. EDA: easy data augmentation techniques for boosting performance on text classification tasks. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (eds.),Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and th...
2019
-
[68]
Pytorch image models
Ross Wightman. Pytorch image models. https://github.com/rwightman/pytorch-image-models , 2019
2019
-
[69]
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3d shapenets: A deep representation for volumetric shapes. InIEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pp. 1912–19...
2015
-
[70]
Disturblabel: Regularizing CNN on the loss layer
Lingxi Xie, Jingdong Wang, Zhen Wei, Meng Wang, and Qi Tian. Disturblabel: Regularizing CNN on the loss layer. In2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV , USA, June 27-30, 2016, pp. 4753–4762. IEEE Computer Society, 2016
2016
-
[71]
Pointnet/pointnet2pytorch
Xu Yan. Pointnet/pointnet2pytorch. https: // github. com/ yanx27/ Pointnet_ Pointnet2_ pytorch, 2019
2019
-
[72]
AD-DROP: attribution-driven dropout for robust language model fine-tuning
Tao Yang, Jinghao Deng, Xiaojun Quan, Qifan Wang, and Shaoliang Nie. AD-DROP: attribution-driven dropout for robust language model fine-tuning. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh (eds.),Advances in Neural Information Processing System...
2022
-
[73]
Pruning by explaining: A novel criterion for deep neural network pruning.Pattern Recognition, 115:107899, 2021
Seul-Ki Yeom, Philipp Seegerer, Sebastian Lapuschkin, Alexander Binder, Simon Wiedemann, Klaus- Robert Müller, and Wojciech Samek. Pruning by explaining: A novel criterion for deep neural network pruning.Pattern Recognition, 115:107899, 2021
2021
-
[74]
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh, Youngjoon Yoo, and Junsuk Choe. Cutmix: Regularization strategy to train strong classifiers with localizable features. In2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October...
2019
-
[75]
Noise injection-based regularization for point cloud processing.CoRR, abs/2103.15027, 2021
Xiao Zang, Yi Xie, Siyu Liao, Jie Chen, and Bo Yuan. Noise injection-based regularization for point cloud processing.CoRR, abs/2103.15027, 2021
2021 arXiv
-
[76]
Zeiler and Rob Fergus
Matthew D. Zeiler and Rob Fergus. Stochastic pooling for regularization of deep convolutional neu- ral networks. In Yoshua Bengio and Yann LeCun (eds.),1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Conference Trac...
2013
-
[77]
Dauphin, and David Lopez-Paz
Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018
2018
-
[78]
Equivalence between dropout and data augmentation: A mathematical check.Neural Networks, 115:82–89, 2019
Dazhi Zhao, Guozhu Yu, Peng Xu, and Maokang Luo. Equivalence between dropout and data augmentation: A mathematical check.Neural Networks, 115:82–89, 2019
2019
-
[79]
Random erasing data augmentation
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. InThe Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AA...
2020
-
[80]
Zintgraf, Taco S
Luisa M. Zintgraf, Taco S. Cohen, Tameem Adel, and Max Welling. Visualizing deep neural network decisions: Prediction difference analysis. In5th International Conference on Learning Representations, ICLR 2017. OpenReview.net, 2017
2017
-
[81]
Regularization and variable selection via the elastic net.Journal of the Royal Statistical Society Series B: Statistical Methodology, 67(2):301–320, 03 2005
Hui Zou and Trevor Hastie. Regularization and variable selection via the elastic net.Journal of the Royal Statistical Society Series B: Statistical Methodology, 67(2):301–320, 03 2005. ISSN 1369-7412
2005
-
[82]
RE data augmentation
Andrea Zunino, Sarah Adel Bargal, Pietro Morerio, Jianming Zhang, Stan Sclaroff, and Vittorio Murino. Excitation dropout: Encouraging plasticity in deep neural networks.Int. J. Comput. Vis., 129(4):1139–1152, 2021. A Technical Appendix A.1 Details on Attribution Computation Fo...
2021
-
[2013]
3084– 3092, 2013
Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, United States, pp. 3084– 3092, 2013
2013
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.