REVIEW 4 major objections 9 minor 57 references
Towards Interpretable Adversarial Examples via Sparse Adversarial Attack
T0 review · 4 major / 9 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that the NP-hard L0 sparse-attack problem can be approximated by learning a mask over a fixed I-FGSM perturbation, and that this yields faster, sparser, more transferable attacks whose perturbations reveal two…
desk verdict A fast, genuinely sparse attack worth knowing, buried under an unsupported near-optimality claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the mask $w$ applied element-wise to a one-time I-FGSM perturbation $\delta$, so the search space is the support of that fixed dense attack. To make the discrete $\ell^0$ penalty trainable, the Heaviside function $H(\pi(w))$ is approximated by the Gaussian $q_a(x)$ from Eq. (12), whose derivative provides a surrogate for the Dirac delta and lets gradient descent decide which pixels to keep. The tailored ReLU $\pi(w - \tau/\epsilon)$ shifts weights down before thresholding, which Proposition 1 shows can only reduce the $\ell^0$ norm, while Proposition 2's cap $\Omega = \min(x/\epsilon, (1-x)/\epsilon)$ guarantees the output stays a valid image. The algorithm alternates gradient steps on the adversarial loss plus $\lambda \sum_j H(\pi(w_j - \tau/\epsilon))$ and then hard-thresholds the final weights to produce the sparse perturbation $\delta^* = \Omega \odot H(\pi(w - \tau/\epsilon)) \odot \delta$.
What would settle it
On a small dataset, compute the exact minimal $\ell^0$ attack by exhaustive search over pixel subsets and compare it with the mask learned on the fixed I-FGSM $\delta$; finding any image where the exact solution uses a pixel with $\delta_j = 0$ or an opposite sign would refute the near-optimality claim.
Extended reading notes
Core claim
The paper's central claim is that the NP-hard sparse-attack problem can be approximated by a masked optimization: minimize $\|w\|_0$ subject to $f_\theta(x + w \odot \delta) = y_{\text{adv}}$, where $\delta$ is a fixed perturbation from I-FGSM and $w$ is a learned mask. Because the $\ell^0$ count of nonzero mask entries is non-differentiable, the authors replace the Heaviside step function with a zero-centered Gaussian surrogate $q_a(x) = \frac{1}{|a|\sqrt{\pi}} \exp(-(x/a)^2)$ that converges to the Dirac delta as $a \to 0$, turning the objective into a smooth loss plus a sparsity penalty. A shifted ReLU $\pi(w - \tau/\epsilon)$ provably reduces the number of nonzero pixels, and the box constraint is enforced by $\Omega = \min(x/\epsilon, (1-x)/\epsilon)$, which guarantees $x + w' \odot \delta \in [0,1]^d$. The authors report that this single optimization beats ten sparse-attack baselines in sparsity (an average of 57 ImageNet pixels for non-targeted attacks), computation time (seconds versus minutes), transferability across VGG and ResNet models, and fooling rate on adversarially trained models. They further claim that the resulting minimal perturbations reveal two interpretable noise types, 'obscuring noise' and 'leading noise', which respectively hide the true class's features and add the target class's features.
Load-bearing premise
The load-bearing premise is that one fixed I-FGSM perturbation $\delta$ already contains every pixel and sign direction the optimal sparse attack will ever need; any pixel where $\delta$ is zero or has the wrong sign is permanently excluded from the search.
Editorial extensions
If this is right
- Sparse attacks become fast enough (seconds on ImageNet) to serve as a routine robustness benchmark for CNNs.
- Near-100% non-targeted fooling can be reached with about 57 ImageNet pixels on standard classifiers, and the examples transfer across VGG and ResNet families better than existing sparse attacks.
- The attack keeps meaningful fooling rates on adversarially trained models, suggesting sparse perturbations remain a threat that robustness training does not fully close.
- The discovered 'obscuring noise' and 'leading noise' give a concrete, visual vocabulary for how minimal pixel changes redirect a classifier's attention.
Reading between the lines
- A natural extension the paper does not explore is to run the mask optimization over several I-FGSM initializations (different $\epsilon$, random restarts, or targeted variants) and check whether sparsity improves, since the single fixed $\delta$ bounds the search space.
- If the two-noise taxonomy is right, it makes a testable causal prediction: deleting only the 'obscuring' pixels should restore the true class's saliency, and deleting only the 'leading' pixels should remove the targeted misprediction; the released code makes this ablation straightforward.
- Because the method is so cheap, it could plausibly be inserted into adversarial training as a regularizer that forces models to be robust to sparse changes, although the paper does not claim this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a sparse adversarial attack for image classifiers. The method first computes a dense initial perturbation δ with I-FGSM (Eqs. (3)-(4)), then optimizes a continuous weight w, masked through ReLU and a Heaviside threshold, such that the perturbed image x + π(w − τ/ε)⊙δ misclassifies to the adversarial label y_adv while the number of active pixels is minimized (Eqs. (6)-(10)). Differentiability is addressed by approximating the Heaviside derivative with a Gaussian q_a (Eqs. (11)-(18)), sparsity is increased by a shifted ReLU threshold τ/ε (Eq. (19)), and box-validity is enforced by the per-pixel bound Ω = min(x/ε, (1−x)/ε) (Eq. (22)). The authors state that solving the masked problem Eq. (6) yields near-optimal solutions to the original L0 problem Eq. (5), and they prove two propositions: sparsity monotonicity of the thresholded mask (Proposition 1) and box-validity of the output (Proposition 2). Experiments on MNIST, CIFAR-10, and ImageNet compare against ten sparse-attack baselines and report superior sparsity, computational cost, transferability, and fooling rates on robust models, together with a Grad-CAM-based interpretation of perturbations as 'obscuring noise' and 'leading noise'. Code is released.
Significance. The empirical core of the paper is strong and appears genuine. The authors release their code, evaluate against ten sparse-attack baselines across three datasets and several architectures (VGG-16/19, ResNet-50/101/152), report ablations over the three hyperparameters a, λ, τ, and test against adversarially trained models (PGD-AT, Fast-AT). The main empirical claims — that masking a dense I-FGSM perturbation with a learned thresholded weight yields very sparse, fast, and transferable attacks — are falsifiable and supported by the tables (e.g., Table 1: 57 ImageNet pixels at 100% fooling rate; Table 4: 4.6 s versus 87.2 s for BruSLeAttack on VGG-16). The interpretability analysis in Section 5.3 is qualitative but suggestive. If the claims survive rescoping, the paper would provide a practically useful sparse-attack baseline and a starting point for interpreting the spatial structure of minimal perturbations. The theoretical contribution is limited: Propositions 1 and 2 are elementary, and the near-optimality assertion relative to Eq. (5) is unsupported, so the paper's value lies primarily in its empirical findings rather than in new optimization theory.
major comments (4)
- [§4.4, Eq. (6)] The sentence following Eq. (6) claims that solving the masked problem 'yields near-optimal solutions for Eq. (5), i.e., with the fewest perturbed pixels,' but no proof or argument is given, and the claim is false in general. The feasible outputs of Eq. (6) are exactly masks of the fixed I-FGSM perturbation δ: since δ is computed once and never updated (Algorithm 1), any optimal sparse perturbation of Eq. (5) that requires a pixel with δ_j = 0, a sign opposite to δ_j, or a magnitude larger than ε is unreachable. Propositions 1 and 2 only establish that the final output is sparser than the initial mask and is box-valid; they say nothing about proximity to the minimum of Eq. (5), and no proposition guarantees that the thresholded output δ* = Ω⊙H(π(w−τ/ε))⊙δ even satisfies f(x+δ*) = y_adv. Since the abstract and Section 1 invoke 'theoretical performance guarantees,' the paper must either prove a formal guarantee under explicit assumptions on the quality of the I-FGSM initialization or rescope the claim to 'sparse successful masks of one fixed dense perturbation,' ideally adding a sanity check on MNIST where near-exhaustive L0 search can estimate the true optimum for a few examples.
- [§4.4, Eq. (18) and Algorithm 1 (Appendix B)] The reparameterization claimed to make the L0 term tractable is not reflected in Algorithm 1. After Eq. (18), the paper states dH/dx ≈ q_a(x), but the pseudocode computes J_adv = J(f(x+w⊙δ), y_adv) + λ Σ_j H(w_j) and then updates w with ∇_w J_adv without any appearance of q_a or a. Since H is a step function, the term λ Σ_j H(w_j) has zero gradient almost everywhere; as written, the sparsity penalty cannot influence the optimization, and the actual sparsification in the algorithm comes from the per-iteration thresholding w := π(w − τ/ε). Either the gradient computation in the implementation uses the surrogate q_a (in which case the pseudocode and the text around Eq. (18) should state this explicitly, including how a is chosen or annealed), or the method does not use the surrogate (in which case the claims that the technique 'approximates the NP-hard l0 optimization problem' and makes it 'computationally tractable' should be revised to describe a projected or thresholded gradient method).
- [Appendix A.2, Eq. (35)] The proof of Proposition 2 begins by assuming δ_j ∈ {ε, 0, −ε}, which is the output of a single FGSM step but not of the I-FGSM initialization used in Algorithm 1: after T iterations of Eq. (3) or Eq. (4), δ_j is an integer multiple of α bounded in absolute value by ε and is generically not equal to ±ε. The box-validity statement is still true for any δ with |δ_j| ≤ ε and Ω_j = min(x_j/ε, (1−x_j)/ε), but the proof needs to treat the general case x_j + Ω_j δ_j ∈ [x_j − Ω_j ε, x_j + Ω_j ε] instead of the two extreme sign cases; Eq. (35) should be removed or corrected.
- [§5.1 and Algorithm 1] The manuscript reports non-targeted attack results in Tables 1 and 2, but Algorithm 1 is written only for the targeted formulation: it minimizes the cross-entropy J(f(x + w⊙δ), y_adv), and Section 5.1 describes the targeted choice of y_adv as the least-likely class. For non-targeted attacks the paper does not state which label is used in J, what sign appears in front of the loss term, or whether the I-FGSM initialization follows Eq. (3) or Eq. (4). Because the sign of the gradient and the choice of y_adv drastically change the behavior of the algorithm, this specification gap should be closed for the non-targeted experiments to be reproducible.
minor comments (9)
- [§2, §5.2, Appendix C.2] The text contains repeated typos, including 'constriant' (twice in Section 2), 'SpareFool' (Section 5.2), 're-writed', 'Obivisly', and 'the internet' (Appendix C.2); these should be corrected.
- [Appendix C.1 vs. §2] Appendix C.1 cites 'SparseFool [7]', but the main text introduces SparseFool as reference [35]; the citation numbering should be made consistent.
- [§4.4, Eq. (12)] The function q_a is introduced as a 'zero-centered normal distribution' without stating its variance; writing q_a(x) = N(x; 0, a²/2) would make the normalizing constant 1/(|a|√π) self-evident and the connection to the Dirac delta clearer.
- [Appendix A.1] The first step of the proof of Proposition 1 ('δ appears in all three terms, so we only need to prove...') is not a valid reduction in general; although the claimed inequalities are true, the product with δ requires a short case analysis on the support of δ.
- [Algorithm 1] The variable w is overwritten by π(w − τ/ε) at the start of every iteration and then updated by a gradient step with respect to ∇_w J_adv, so it is unclear whether gradients are taken with respect to the pre- or post-threshold value; distinct symbols (e.g., w̃ for the thresholded value) would remove the ambiguity.
- [§5.5, Table 3] The reported five-trial means have no error bars or ranges; several adjacent entries (e.g., 86.3 vs. 84.2 for ResNet-101-generated attacks on VGG-19) are close, so standard errors or per-trial ranges should be reported.
- [§5.3] The 'obscuring noise' and 'leading noise' categories are assigned by visual inspection of one or two images using hand-drawn boxes; the paper should state explicitly that these are post-hoc interpretations and could quantify the spatial overlap of perturbed pixels with the relevant Grad-CAM regions to support the proposed mechanism.
- [§4.4, Eq. (7)] The term 'Lagrangian relaxation' is imprecise: Eq. (7) is a regularized objective that informally trades off the adversarial loss against L0 sparsity, not the Lagrangian of the constrained problem Eq. (6), whose constraint is not differentiable.
- [Theorem 1, §4.4] The proof of Theorem 1 establishes pointwise convergence of q_a to the Dirac delta; since the use in Eq. (18) is as a surrogate for the distributional derivative of H, a remark that q_a has unit integral and converges in the distributional sense, together with a comment on the bias of the surrogate gradient for finite a, would make the argument complete.
Circularity Check
No significant circularity: the empirical evaluations are independent benchmarks and the theoretical results are elementary statements about the method's own quantities, not disguised fits.
full rationale
The paper's central derivation replaces the L0 attack (Eq. 5) by a masked optimization over a fixed I-FGSM perturbation (Eq. 6), and then asserts without proof that Eq. (6) is near-optimal for Eq. (5). This is a genuine completeness and soundness gap: the output is by construction a sparse mask of one fixed dense perturbation, so pixels outside the support of that perturbation, or with sign opposite to it, are never available, and the paper's Propositions 1 and 2 establish only monotone sparsity reduction and box validity, not proximity to the optimum of Eq. (5). However, that gap is not circularity. Theorem 1 is a standard mollifier convergence fact with a proof given in the paper; Proposition 1 is an elementary inequality comparing three L0 counts; Proposition 2 is an elementary box-validity bound using the assumed sign structure of I-FGSM. None of these results is assumed from or equivalent to the empirical claims. The experimental comparisons in Tables 1 through 4 are run against external, independently published sparse attacks such as C&W, SparseFool, GreedyFool, Sparse-RS, and BruSLeAttack, with hyperparameters tuned by the authors' own ablations rather than fitted to the benchmark outcomes. No fitted parameter is renamed as a prediction, and no load-bearing uniqueness or ansatz is imported from the authors' prior work. The interpretability finding of obscuring noise versus leading noise is a post-hoc visualization category, not a derivation masquerading as a first-principles result. The unsupported near-optimality statement is a correctness risk that deserves a formal approximation bound, but it does not make the paper's derivation circular.
Assumptions & free parameters
free parameters (3)
- lambda (λ) =
1e-2 for MNIST and CIFAR-10, 1e-3 for ImageNet
- a =
0.1
- tau (τ) =
0.30
assumptions (5)
- standard math The zero-centered Gaussian q_a(x) converges to the Dirac delta in the distributional sense as a approaches 0.
- domain assumption Different DNNs learn similar decision boundaries, so dense attack perturbations transfer to other models.
- ad hoc to paper The fixed initial perturbation δ from I-FGSM contains all directions needed for the optimal sparse attack.
- ad hoc to paper Replacing the derivative of the Heaviside function with the Gaussian q_a(x) is a valid optimization surrogate.
- domain assumption Hard-thresholding the optimized weights with H(π(w - τ/eps)) preserves the adversarial property of the perturbation.
invented entities (2)
-
Obscuring noise
-
Leading noise
Cite this review
Pith. "Pith review of Towards Interpretable Adversarial Examples via Sparse Adversarial Attack." pith.science (2026). https://pith.science/paper/NPJVBC2S
@misc{pith2026250617250,
author = {Pith},
title = {Pith review of: Towards Interpretable Adversarial Examples via Sparse Adversarial Attack},
year = {2026},
howpublished = {\url{https://pith.science/paper/NPJVBC2S}},
note = {Machine review of arXiv:2506.17250}
}
read the original abstract
Sparse attacks are to optimize the magnitude of adversarial perturbations for fooling deep neural networks (DNNs) involving only a few perturbed pixels (i.e., under the l0 constraint), suitable for interpreting the vulnerability of DNNs. However, existing solutions fail to yield interpretable adversarial examples due to their poor sparsity. Worse still, they often struggle with heavy computational overhead, poor transferability, and weak attack strength. In this paper, we aim to develop a sparse attack for understanding the vulnerability of CNNs by minimizing the magnitude of initial perturbations under the l0 constraint, to overcome the existing drawbacks while achieving a fast, transferable, and strong attack to DNNs. In particular, a novel and theoretical sound parameterization technique is introduced to approximate the NP-hard l0 optimization problem, making directly optimizing sparse perturbations computationally feasible. Besides, a novel loss function is designed to augment initial perturbations by maximizing the adversary property and minimizing the number of perturbed pixels simultaneously. Extensive experiments are conducted to demonstrate that our approach, with theoretical performance guarantees, outperforms state-of-the-art sparse attacks in terms of computational overhead, transferability, and attack strength, expecting to serve as a benchmark for evaluating the robustness of DNNs. In addition, theoretical and empirical results validate that our approach yields sparser adversarial examples, empowering us to discover two categories of noises, i.e., "obscuring noise" and "leading noise", which will help interpret how adversarial perturbation misleads the classifiers into incorrect predictions. Our code is available at https://github.com/fudong03/SparseAttack.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Ahsan Ali, Xiaolong Ma, Syed Zawad, Paarijaat Aditya, Istemi Ekin Akkus, Ruichuan Chen, Lei Yang, and Feng Yan. Enabling scalable and adaptive machine learning training via serverless computing on public cloud.Performance Evaluation, 167:102451, 2025
work page 2025
-
[2]
Nicholas Carlini and David A. Wagner. Towards evaluating the robustness of neural net- works. InIEEE Symposium on Security and Privacy, 2017
work page 2017
-
[3]
Jianbo Chen, Michael I. Jordan, and Martin J. Wainwright. Hopskipjumpattack: A query- efficient decision-based attack. InIEEE Symposium on Security and Privacy, 2020, 2020
work page 2020
-
[4]
More data can expand the general- ization gap between adversarially robust and standard models
Lin Chen, Yifei Min, Mingrui Zhang, and Amin Karbasi. More data can expand the general- ization gap between adversarially robust and standard models. InInternational Conference on Machine Learning (ICML), 2020
work page 2020
-
[5]
Tiankuo Chu, Fudong Lin, Shubo Wang, Jason Jiang, Wiley Jia-Wei Gong, Xu Yuan, and Liyun Wang. Bonemet: An open large-scale multi-modal murine dataset for breast cancer bone metastasis diagnosis and prognosis. InThe Thirteenth International Conference on Learning Representations (ICLR), 2025. 16 F. Lin, J. Lou, et al
work page 2025
-
[6]
Singh, Nicolas Flammarion, and Matthias Hein
Francesco Croce, Maksym Andriushchenko, Naman D. Singh, Nicolas Flammarion, and Matthias Hein. Sparse-rs: A versatile framework for query-efficient sparse black-box ad- versarial attacks. InAAAI, 2022
work page 2022
-
[7]
Sparse and imperceivable adversarial attacks
Francesco Croce and Matthias Hein. Sparse and imperceivable adversarial attacks. InICCV, 2019
work page 2019
-
[8]
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. InConference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies (NAACL-HLT), 2019
work page 2019
Show all 57 references
-
[9]
Greedyfool: Distortion-aware sparse adversarial attack
Xiaoyi Dong, Dongdong Chen, Jianmin Bao, Chuan Qin, Lu Yuan, Weiming Zhang, Nenghai Yu, and Dong Chen. Greedyfool: Distortion-aware sparse adversarial attack. InNeurIPS, 2020
2020
-
[10]
Boosting adversarial attacks with momentum
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial attacks with momentum. InConference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[11]
Evading defenses to transferable ad- versarial examples by translation-invariant attacks
Yinpeng Dong, Tianyu Pang, Hang Su, and Jun Zhu. Evading defenses to transferable ad- versarial examples by translation-invariant attacks. InConference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[12]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...
2021
-
[13]
Robust physical-world attacks on deep learning visual classification
Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. Robust physical-world attacks on deep learning visual classification. InCVPR, 2018
2018
-
[14]
Sparse adversarial attack via perturbation factorization
Yanbo Fan, Baoyuan Wu, Tuanhui Li, Yong Zhang, Mingyang Li, Zhifeng Li, and Yujiu Yang. Sparse adversarial attack via perturbation factorization. InECCV, 2020
2020
-
[15]
Goodfellow, Jonathon Shlens, and Christian Szegedy
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing ad- versarial examples. InInternational Conference on Learning Representations (ICLR), 2015
2015
-
[16]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016
2016
-
[17]
Interpretable minority synthesis for imbalanced classification
Yi He, Fudong Lin, Xu Yuan, and Nian-Feng Tzeng. Interpretable minority synthesis for imbalanced classification. InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI), pages 2542–2548, 2021
2021
-
[18]
Highly accurate protein structure prediction with alphafold.Nature, 596(7873):583–589, 2021
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ron- neberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold.Nature, 596(7873):583–589, 2021
2021
-
[19]
Physgan: Generating physical-world- resilient adversarial examples for autonomous driving
Zelun Kong, Junfeng Guo, Ang Li, and Cong Liu. Physgan: Generating physical-world- resilient adversarial examples for autonomous driving. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020
2020
-
[20]
Learning multiple layers of features from tiny images.Technical Report, University of Toronto, 2009
Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images.Technical Report, University of Toronto, 2009
2009
-
[21]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. InAdvances in Neural Information Processing Systems (NIPS), 2012
2012
-
[22]
Goodfellow, and Samy Bengio
Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. Adversarial examples in the physical world. InInternational Conference on Learning Representations (ICLR), 2017
2017
-
[23]
Goodfellow, and Samy Bengio
Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. Adversarial machine learning at scale. In5th International Conference on Learning Representations (ICLR), 2017. Towards Interpretable Adversarial Examples via Sparse Adversarial Attack 17
2017
-
[24]
Inan, Sahar Abdelnabi, Janardhan Kulkarni, Lukas Wutschitz, Reza Shokri, Christopher G
Guangchen Lan, Huseyin A. Inan, Sahar Abdelnabi, Janardhan Kulkarni, Lukas Wutschitz, Reza Shokri, Christopher G. Brinton, and Robert Sim. Contextual integrity in LLMs via reasoning and reinforcement learning.arXiv preprint arXiv:2506.04245, 2025
2025
-
[25]
MNIST handwritten digit database
Yann LeCun and Corinna Cortes. MNIST handwritten digit database. 2010
2010
-
[26]
Mmst-vit: Climate change-aware crop yield prediction via multi-modal spatial-temporal vi- sion transformer
Fudong Lin, Summer Crawford, Kaleb Guillot, Yihe Zhang, Yan Chen, Xu Yuan, et al. Mmst-vit: Climate change-aware crop yield prediction via multi-modal spatial-temporal vi- sion transformer. InIEEE/CVF International Conference on Computer Vision (ICCV), pages 5751–5761, 2023
2023
-
[27]
Towards robust vision trans- former via masked adaptive ensemble
Fudong Lin, Jiadong Lou, Xu Yuan, and Nian-Feng Tzeng. Towards robust vision trans- former via masked adaptive ensemble. InProceedings of the 33rd ACM International Con- ference on Information and Knowledge Management (CIKM), pages 1389–1399, 2024
2024
-
[28]
Comprehensive transformer-based model architecture for real-world storm predic- tion
Fudong Lin, Xu Yuan, Yihe Zhang, Purushottam Sigdel, Li Chen, Lu Peng, and Nian-Feng Tzeng. Comprehensive transformer-based model architecture for real-world storm predic- tion. InJoint European Conference on Machine Learning and Knowledge Discovery in Databases (ECML-PKDD), p...
2023
-
[29]
Delving into transferable adversarial examples and black-box attacks
Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. Delving into transferable adversarial examples and black-box attacks. InICLR, 2017
2017
-
[30]
Learning sparse neural networks through l_0 regularization
Christos Louizos, Max Welling, and Diederik P Kingma. Learning sparse neural networks through l_0 regularization. InICLR, 2018
2018
-
[31]
Malletrain: Deep neural networks training on unfillable supercomputer nodes
Xiaolong Ma, Feng Yan, Lei Yang, Ian Foster, Michael E Papka, Zhengchun Liu, and Ra- jkumar Kettimuthu. Malletrain: Deep neural networks training on unfillable supercomputer nodes. InProceedings of the 15th ACM/SPEC International Conference on Performance Engineering, pages 19...
2024
-
[32]
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. InInternational Con- ference on Learning Representations (ICLR), 2018
2018
-
[33]
The curious case of adversarially robust models: More data can help, double descend, or hurt generalization
Yifei Min, Lin Chen, and Amin Karbasi. The curious case of adversarially robust models: More data can help, double descend, or hurt generalization. InUncertainty in Artificial Intelligence, 2021
2021
-
[34]
Transferable structural sparse adver- sarial attack via exact group sparsity training
Di Ming, Peng Ren, Yunlong Wang, and Xin Feng. Transferable structural sparse adver- sarial attack via exact group sparsity training. InComputer Vision and Pattern Recognition (CVPR), pages 24696–24705, 2024
2024
-
[35]
Sparsefool: A few pixels make a big difference
Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard. Sparsefool: A few pixels make a big difference. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[36]
Deepfool: A sim- ple and accurate method to fool deep neural networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard. Deepfool: A sim- ple and accurate method to fool deep neural networks. InConference on Computer Vision and Pattern Recognition (CVPR), 2016
2016
-
[37]
McDaniel, Somesh Jha, Matt Fredrikson, Z
Nicolas Papernot, Patrick D. McDaniel, Somesh Jha, Matt Fredrikson, Z. Berkay Celik, and Ananthram Swami. The limitations of deep learning in adversarial settings. InEuropean Symposium on Security and Privacy (EuroS&P), 2016
2016
-
[38]
Fast minimum-norm adver- sarial attacks through adaptive norm constraints
Maura Pintor, Fabio Roli, Wieland Brendel, and Battista Biggio. Fast minimum-norm adver- sarial attacks through adaptive norm constraints. InNeural Information Processing Systems (NeurIPS), pages 20052–20062, 2021
2021
-
[39]
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences.Proceedings of the National Academy of Sciences, 118(15), 2021
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences.Proceedings of the National A...
2021
-
[40]
Berg, and 18 F
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhi- heng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and 18 F. Lin, J. Lou, et al. Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge.International Jour...
2015
-
[41]
Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient- based localization. InICCV, 2017
2017
-
[42]
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. InInternational Conference on Learning Representations (ICLR), 2015
2015
-
[43]
One pixel attack for fooling deep neural networks.IEEE Trans
Jiawei Su, Danilo Vasconcellos Vargas, and Kouichi Sakurai. One pixel attack for fooling deep neural networks.IEEE Trans. Evol. Comput., 2019
2019
-
[44]
Hybrid batch attacks: Finding black- box adversarial examples with limited queries
Fnu Suya, Jianfeng Chi, David Evans, and Yuan Tian. Hybrid batch attacks: Finding black- box adversarial examples with limited queries. InUSENIX, 2020
2020
-
[45]
Goodfellow, and Rob Fergus
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. Intriguing properties of neural networks. InInternational Conference on Learning Representations (ICLR), 2014
2014
-
[46]
Fooling automated surveillance cam- eras: Adversarial patches to attack person detection
Simen Thys, Wiebe Van Ranst, and Toon Goedemé. Fooling automated surveillance cam- eras: Adversarial patches to attack person detection. InIEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2019
2019
-
[47]
Goodfellow, Dan Boneh, and Patrick D
Florian Tramèr, Alexey Kurakin, Nicolas Papernot, Ian J. Goodfellow, Dan Boneh, and Patrick D. McDaniel. Ensemble adversarial training: Attacks and defenses. InInternational Conference on Learning Representations (ICLR), 2018
2018
-
[48]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems (NIPS), 2017
2017
-
[49]
Brusleattack: a query-efficient score- based black-box sparse adversarial attack
Viet Quoc V o, Ehsan Abbasnejad, and Damith Ranasinghe. Brusleattack: a query-efficient score- based black-box sparse adversarial attack. InInternational Conference on Learning Representations(ICLR), 2024
2024
-
[50]
Heaviside step function, 2024
Wikipedia. Heaviside step function, 2024. [Online; accessed 01-January-2024]
2024
-
[51]
Black-box sparse adversarial attack via multi-objective optimisation
Phoenix Neale Williams and Ke Li. Black-box sparse adversarial attack via multi-objective optimisation. InComputer Vision and Pattern Recognition (CVPR), 2023
2023
-
[52]
Zico Kolter
Eric Wong, Leslie Rice, and J. Zico Kolter. Fast is better than free: Revisiting adversarial training. InInternational Conference on Learning Representations (ICLR), 2020
2020
-
[53]
Cihang Xie, Zhishuai Zhang, Yuyin Zhou, Song Bai, Jianyu Wang, Zhou Ren, and Alan L. Yuille. Improving transferability of adversarial examples with input diversity. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[54]
Structured adversarial attack: Towards general implementation and better interpretability
Kaidi Xu, Sijia Liu, Pu Zhao, Pin-Yu Chen, Huan Zhang, Quanfu Fan, Deniz Erdogmus, Yanzhi Wang, and Xue Lin. Structured adversarial attack: Towards general implementation and better interpretability. InICLR, 2019
2019
-
[55]
Fedcust: Offloading hyperparameter customization for federated learning.Perfor- mance Evaluation, 167:102450, 2025
Syed Zawad, Xiaolong Ma, Jun Yi, Cheng Li, Minjia Zhang, Lei Yang, Feng Yan, and Yux- iong He. Fedcust: Offloading hyperparameter customization for federated learning.Perfor- mance Evaluation, 167:102450, 2025
2025
-
[56]
Zeiler and Rob Fergus
Matthew D. Zeiler and Rob Fergus. Visualizing and understanding convolutional networks. InEuropean Conference on Computer Vision, 2014
2014
-
[57]
pelican" ∥δ∥0 = 248 (b) GreedyFool. “goldfish
Mingkang Zhu, Tianlong Chen, and Zhangyang Wang. Sparse and imperceptible adversarial attack via a homotopy algorithm. InICML, 2021. Towards Interpretable Adversarial Examples via Sparse Adversarial Attack 19 Outline This document supplements the main paper in three aspects. F...
2021
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.