REVIEW 3 major objections 5 minor 48 references
NAPPure recovers clean images from blur, patch, and distortion attacks by jointly maximizing the likelihood of the clean image and the unknown attack parameters under a known transformation model.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 09:37 UTC pith:3TELTHJA
load-bearing objection Plausible and well-specified extension of diffusion purification to non-additive perturbations, but the headline robustness numbers rest on non-adaptive attacks and per-attack tuning. the 3 major comments →
NAPPure: Adversarial Purification for Robust Image Classification under Non-Additive Perturbations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that adversarial images produced by a known-but-parameterized transformation can be purified by solving a joint optimization that disentangles clean content from perturbation parameters. A pretrained diffusion model scores image likelihood; a chosen prior pulls ε toward the identity element of the transformation; a reconstruction term forces f(x, ε) to match the observed image. Alternating gradient updates on x and ε yield a purified image that classifies correctly. The paper shows this on blur, center-patch occlusion, and flow-field distortion attacks, with robust accuracy gains of 25+ points over standard purification on GTSRB, and it demonstrates the same objective de
What carries the argument
The combined loss in Eq. (7): an ELBO-style image-likelihood term from a diffusion denoiser, a perturbation prior −ϕ(ε), and a reconstruction term ∥x_adv − f(x, ε)∥². x and ε are updated alternately with gradient-based optimization. For transformations that are non-differentiable in ε (e.g., patch location and size), a trained auxiliary network substitutes for f. The identity element ε₀ anchors the prior so clean inputs stay unchanged, and the prior plus reconstruction terms prevent semantic drift.
Load-bearing premise
The defender must know the exact functional form of the transformation f and its parameter domain; all experiments assume exact knowledge of blur kernel size, patch location, and flow-field structure.
What would settle it
Configure the defense for 5×5 blur kernels and attack with a 9×9 kernel, or place the patch off-center rather than at the image middle; if the robust accuracy collapses to the no-defense baseline, the joint optimization fails under structural misspecification of f.
If this is right
- Any attack whose transformation f is known in advance can be handled by the same template; the three implemented types (blur, patch, flow) serve as configurations, not limits.
- The framework is a compatible extension of standard purification: when f is additive, the objective collapses to pure image-likelihood maximization, matching existing methods.
- Composite attacks are handled by composing simple transformations with learned interpolation weights, giving robustness even when the exact attack type is unknown within a known set.
- Robustness degrades gracefully with mismatched attack parameters inside the assumed class, since the parameter search adapts (e.g., patch size and kernel size generalization).
Where Pith is reading between the lines
- A consequence the paper leaves implicit: the same joint-estimation objective applies to general inverse problems where the forward model is known — deblurring, inpainting, geometric correction — without an adversary, so NAPPure doubles as a parameter-agnostic restoration method.
- One testable extension is replacing the hand-set potential ϕ(ε) with a learned parameter distribution, which could reduce sensitivity to the λ weights and improve behavior on out-of-domain parameters.
- Because the method assumes exact knowledge of f, a natural attack vector is a transformation that lies outside the defender's assumed class (e.g., an off-center patch or a larger blur kernel than configured); quantifying this gap would clarify the practical security boundary.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes NAPPure, an adversarial purification framework for non-additive perturbations. It models the generation of an adversarial image as x_adv = f(x, ε) with a known transformation f and unknown perturbation parameter ε, and recovers the clean image by jointly optimizing an image-likelihood term (a diffusion ELBO from EDM), a perturbation prior φ(ε), and a reconstruction constraint ||x_adv − f(x, ε)||² (Eq. 7). Implementations are given for blur, occlusion, flow-field distortion, and composite transformations. Experiments on GTSRB, CIFAR-10, and ImageNet report large robust-accuracy gains over DiffPure, LM, and AT, e.g., 73.93% vs 43.2% average robust accuracy on GTSRB. A degeneration argument in §4.3 shows that the additive case reduces to the LM purification objective.
Significance. If the reported robustness held under a threat model in which the adversary knows the defense, NAPPure would be a meaningful extension of diffusion-based purification to a broad class of parametric corruptions. The objective in Eq. (7) is a clean and well-specified posterior decomposition; the use of an auxiliary model for non-differentiable transformations is sensible; and the degeneration proof in §4.3 is a useful formal connection. However, the central empirical claim is currently supported only under non-adaptive attacks and per-attack hyper-parameter tuning, so the significance is conditional on addressing these evaluation issues.
major comments (3)
- [§5.1, Eq. (1); Tables 1/2] The attack objective optimizes L(c(f(x,ε)), y) with respect to the raw classifier c only; no gradient is propagated through the NAPPure purification map g (Alg. 1). Thus the robust accuracy in Tables 1/2/7 is measured against attacks that do not know the defense, and the statement 'white-box adversarial examples for each defense method' is not supported by the protocol. This matters specifically for Eq. (7): the reconstruction term ||x_adv−f(x,ε)||² is a consistency constraint that an adaptive adversary can exploit through the non-uniqueness of f(x,ε)=x_adv; the §4.3 degeneration argument covers only additive f. The central robustness claim is not established until adaptive attacks (e.g., BPDA or unrolled optimization through Alg. 1) are evaluated.
- [§5.1, Table 8] λ1 and λ2 are grid-searched separately for each attack type and dataset on adversarial validation examples generated for the same attack configuration, and the number of iterations T also varies widely (e.g., 20 for CIFAR-10 Additive vs 500 for non-additive). The reported robust accuracy is therefore the best-case result for the exact evaluated attack configuration, not the performance of a fixed defense. Since the paper claims a general purification framework, the evaluation should report results with a single hyper-parameter setting across attack types, or justify why per-attack tuning is an appropriate threat model. Appendix A5 tests only small parameter variations (patch size, kernel size) and does not address this selection issue.
- [§3.1, App. A5] The framework requires exact knowledge of f and its parameter domain Ω; all main experiments instantiate this at the exact attack configuration (blur kernel size, patch centered at a fixed location, flow-field smoothing). This assumption is load-bearing for Eq. (7): if the assumed f is structurally wrong (e.g., off-center patch, different blur kernel shape, or an unmodeled component), the reconstruction term is a wrong constraint and no recovery guarantee is given. Appendix A5 tests only parameter variation inside the assumed model (3×3 vs 5×5 blur, 5×5–9×9 patches) and does not test structural misspecification. The paper should either demonstrate graceful degradation under a mismatched f or explicitly state this as a scope limitation in the threat model.
minor comments (5)
- [References] Reference formatting is inconsistent: ref. [4] has a garbled author string ('X. Yangand C. Q. Duan'), and several references contain broken spacing. A full bibliography pass is needed.
- [Tables 1/2] The results are point estimates over 512 images with no standard deviation or multiple-seed reporting; given the stochastic sampling of n and σ in Alg. 1, error bars are important for evaluating the robustness gains.
- [Appendix] The appendix tables are numbered 'Table 7/Table 8' while the main text refers to 'Tab. 7'/'Tab. 8'; please unify the table-labeling style.
- [§4.2 vs Alg. 1] The text around Eq. (4) says σ ∼ p_data(σ), while Alg. 1 line 3 samples σ ∼ U(0.4,0.6). Please align these specifications.
- [General] No code release is mentioned. For a method with many hyper-parameters (λ1, λ2, T, learning rates), providing code or a detailed reproducibility package would strengthen the paper.
Circularity Check
No significant circularity: the derivation is a direct posterior decomposition, and the self-referential elements are non-load-bearing.
full rationale
The derivation chain (Eqs. 2-7) is a Bayes decomposition: log p(x, ε | x_adv) = log p(x) + log p(ε) + log p(x_adv | x, ε) - log p(x_adv), with p(x) implemented by an external EDM/ELBO term (Eq. 4), p(ε) by an explicit user-specified potential (Eq. 5), and p(x_adv | x, ε) by the known transformation f (Eq. 6). None of these terms is defined in terms of the reported robust accuracy; the purified image is the output of an alternating optimization, not a re-display of a fitted quantity. The additive degeneration (Eqs. 9-10) is a genuine special-case reduction to LM, not an import of the conclusion. Self-citations [44] and [45] overlap with the author list, but they are cited only for general background ('adversarial purification has been proved effective', 'AT is less effective on unseen attacks') alongside external references [4] and [27], so they are not load-bearing. The per-attack grid search of lambda1/lambda2 is performed on adversarial examples from a separate validation set, so reporting test robust accuracy is hyperparameter selection, not a fitted prediction. The main limitations — exact knowledge of f (Sec. 3.1), fixed attack configurations, and the App. A5 caveat that generalization holds only while the attack kernel size does not exceed the defense kernel size — affect scope and threat model, but they do not make the derivation circular.
Axiom & Free-Parameter Ledger
free parameters (4)
- λ1 (perturbation prior weight) =
0.001–0.1 depending on attack/dataset (Table 8)
- λ2 (reconstruction weight) =
1–10 (Table 8)
- T (purification iterations) =
20–500 depending on setting (Table 8)
- σ sampling interval U(0.4,0.6) =
0.4–0.6
axioms (5)
- domain assumption The pretrained EDM diffusion model provides a valid approximation to log p(x) for natural images of the dataset.
- domain assumption The transformation function f is known in advance and the parameter domain Ω keeps semantics unchanged.
- domain assumption The reconstruction loss assumes x_adv ~ N(f(x,ε), σ²I) as a relaxation of the hard constraint.
- domain assumption The ELBO single-sample estimator with one n and σ per iteration is a sufficient optimizer for purification.
- domain assumption The potential function φ(ε) and identity element ε0 are chosen to represent the perturbation prior.
read the original abstract
Adversarial purification has achieved great success in combating adversarial image perturbations, which are usually assumed to be additive. However, non-additive adversarial perturbations such as blur, occlusion, and distortion are also common in the real world. Under such perturbations, existing adversarial purification methods are much less effective since they are designed to fit the additive nature. In this paper, we propose an extended adversarial purification framework named NAPPure, which can further handle non-additive perturbations. Specifically, we first establish the generation process of an adversarial image, and then disentangle the underlying clean image and perturbation parameters through likelihood maximization. Experiments on GTSRB and CIFAR-10 datasets show that NAPPure significantly boosts the robustness of image classification models against non-additive perturbations.
Figures
Reference graph
Works this paper leans on
-
[1]
Andriushchenko, F
M. Andriushchenko, F. Croce, N. Flammarion, and M. Hein. Square attack: a query-efficient black-box adversarial attack via random search. InEuropean conference on computer vision, pages 484–501. Springer, 2020. 2
2020
-
[2]
T. B. Brown, D. Mané, A. Roy, M. Abadi, and J. Gilmer. Adversarial patch.arXiv preprint arXiv:1712.09665, 2017. 1, 2
Pith/arXiv arXiv 2017
-
[3]
Chakraborty, M
A. Chakraborty, M. Alam, V . Dey, A. Chattopadhyay, and D. Mukhopadhyay. A survey on adversarial attacks and de- fences.CAAI Transactions on Intelligence Technology, 6(1): 25–45, 2021. 2
2021
-
[4]
H. R. Chen, Y . P. Dong, Z. Y . Wang, X. Yangand C. Q. Duan, H. Su, and J. Zhu. Robust classification via a single diffusion model.arXiv preprint arXiv:2305.15241, 2023. 1, 2, 3, 4, 6, 7, 8, 12
Pith/arXiv arXiv 2023
-
[5]
J. W. Chen and X. X. Wei. Defending adversarial patches via joint region localizing and inpainting.arXiv preprint arXiv:2307.14242, 2023. 2
Pith/arXiv arXiv 2023
-
[6]
X. J. Chu, L. Y . Chen, C. P. Chen, and X. Lu. Improving im- age restoration by revisiting global information aggregation. InEuropean Conference on Computer Vision, pages 53–71. Springer, 2022. 2
2022
-
[7]
Croce and M
F. Croce and M. Hein. Reliable evaluation of adversar- ial robustness with an ensemble of diverse parameter-free attacks. InInternational conference on machine learning, pages 2206–2216. PMLR, 2020. 6
2020
-
[8]
J. H. Dong, S. Moosavi-Dezfooli, J. H. Lai, and X. H. Xie. The enemy of my enemy is my friend: Exploring inverse ad- versaries for improving adversarial training. InProceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 24678–24687, 2023. 2
2023
-
[9]
Eykholt, I
K. Eykholt, I. Evtimov, E. Fernandes, B. Li, A. Rahmati, C. W. Xiao, A. Prakash, T. Kohno, and D. Song. Robust physical-world attacks on deep learning visual classification. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 1625–1634, 2018. 1, 2, 6, 7, 8
2018
-
[10]
Generative adversarial nets.Advances in neural information processing systems, 27, 2014
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets.Advances in neural information processing systems, 27, 2014. 3
2014
-
[11]
I. J. Goodfellow, J. Shlens, and C. Szegedy. Explain- ing and harnessing adversarial examples.arXiv preprint arXiv:1412.6572, 2014. 1, 2
Pith/arXiv arXiv 2014
-
[12]
Q. Guo, F. Juefei-Xu, X. F. Xie, L. Ma, J. Wang, W. Feng, and Y . Liu. Abba: Saliency-regularized motion-based adver- sarial blur attack.arXiv preprint arXiv:2002.03500, 2020. 1, 2, 3
Pith/arXiv arXiv 2002
-
[13]
Q. Guo, Z. Y . Cheng andF. Juefei-Xu, L. Ma, X. F. Xie, Y . Liu, and J. J. Zhao. Learning to adversarially blur visual ob- ject tracking. InProceedings of the IEEE/CVF international conference on computer vision, pages 10839–10848, 2021. 2
2021
-
[14]
Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020. 3
2020
-
[15]
S. Y . Huang, F. Ye, Z. C. Huang, W. Li, T. Q. Huang, and L. Q. Huang. Patchbreaker: defending against adversarial attacks by cutting-inpainting patches and joint adversarial training.Applied Intelligence, 54(21):10819–10832, 2024. 2
2024
-
[16]
D. Kalaria, A. Hazra, and P. P. Chakrabarti. Towards ad- versarial purification using denoising autoencoders.arXiv preprint arXiv:2208.13838, 2022. 2
Pith/arXiv arXiv 2022
-
[17]
Kanbak, S
C. Kanbak, S. M. Moosavi-Dezfooli, and P. Frossard. Geo- metric robustness of deep networks: analysis and improve- ment. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4441–4449, 2018. 1, 2
2018
-
[18]
Karras, M
T. Karras, M. Aittala, T. Aila, and S. Laine. Elucidating the design space of diffusion-based generative models.Ad- vances in neural information processing systems, 35:26565– 26577, 2022. 4, 6, 11
2022
-
[19]
Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114, 2013
Diederik P Kingma. Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114, 2013. 3
Pith/arXiv arXiv 2013
-
[20]
Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980,
Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980,
-
[21]
Krizhevsky, G
A. Krizhevsky, G. Hinton, et al. Learning multiple layers of features from tiny images. 2009. 6
2009
-
[22]
A. Kurakin, I. Goodfellow, and S. Bengio. Adversarial ma- chine learning at scale.arXiv preprint arXiv:1611.01236,
-
[23]
Defense-vae: A fast and accurate defense against adversarial attacks
Xiang Li and Shihao Ji. Defense-vae: A fast and accurate defense against adversarial attacks. InJoint European Con- ference on Machine Learning and Knowledge Discovery in Databases, pages 191–207. Springer, 2019. 3
2019
-
[24]
X. Liu, H. R. Yang, Z. W. Liu, L. H. Song, H. Li, and Y . R. Chen. Dpatch: An adversarial patch attack on object detec- tors.arXiv preprint arXiv:1806.02299, 2018. 1, 2
Pith/arXiv arXiv 2018
-
[25]
M ˛ adry, A
A. M ˛ adry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial attacks.stat, 1050(9), 2017. 2, 6, 7
2017
-
[26]
J. X. Mi, X. D. Wang, L. F. Zhou, and K. Cheng. Adver- sarial examples based on object detection tasks: A survey. Neurocomputing, 519:114–126, 2023. 2
2023
-
[27]
W. L. Nie, B. D. Guo, Y . J. Huang, C. W. Xiao, A. Vahdat, and A. Anandkumar. Diffusion models for adversarial purifi- cation.arXiv preprint arXiv:2205.07460, 2022. 1, 2, 3, 6, 7, 8, 11, 12
Pith/arXiv arXiv 2022
-
[28]
S. Rao, D. Stutz, and B. Schiele. Adversarial training against location-optimized adversarial patches. InEuropean confer- ence on computer vision, pages 429–448. Springer, 2020. 1, 2, 3
2020
-
[29]
Ren and T
H. Ren and T. Huang. Adversarial example attacks in the physical world. InMachine Learning for Cyber Security: Third International Conference, ML4CS 2020, Guangzhou, China, October 8–10, 2020, Proceedings, Part II 3, pages 572–582. Springer, 2020. 2
2020
-
[30]
P Samangouei. Defense-gan: protecting classifiers against adversarial attacks using generative models.arXiv preprint arXiv:1805.06605, 2018. 3
Pith/arXiv arXiv 2018
-
[31]
Mimicd- iffusion: Purifying adversarial perturbation via mimicking clean diffusion model
Kaiyu Song, Hanjiang Lai, Yan Pan, and Jian Yin. Mimicd- iffusion: Purifying adversarial perturbation via mimicking clean diffusion model. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 24665–24674, 2024. 2, 3
2024
-
[32]
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions.arXiv preprint arXiv:2011.13456, 2020. 3
Pith/arXiv arXiv 2011
-
[33]
Stallkamp, M
J. Stallkamp, M. Schlipsing, J. Salmen, and C. Igel. Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition.Neural networks, 32:323–332,
-
[34]
C. Szegedy. Intriguing properties of neural networks.arXiv preprint arXiv:1312.6199, 2013. 1
Pith/arXiv arXiv 2013
-
[35]
L. Tang, N. Ruiz, Q. H. Chu, Y . Z., A. Holynski, D. E. Jacobs, B. Hariharan, Y . Pritch, N. Wadhwa, K. Aberman, et al. Realfill: Reference-driven generation for authentic im- age completion.ACM Transactions on Graphics (TOG), 43 (4):1–12, 2024. 2
2024
-
[36]
Guided diffusion model for adversarial purification
Jinyi Wang, Zhaoyang Lyu, Dahua Lin, Bo Dai, and Hongfei Fu. Guided diffusion model for adversarial purification. arXiv preprint arXiv:2205.14969, 2022. 3
Pith/arXiv arXiv 2022
-
[37]
J. Wang, X. L. Liu, J. Hu, D. H. Wang, S. Y . Wu, T. S. Jiang, Y . F. Guo, A. S. Liu, and J. T. Zhou. Adversarial examples in the physical world: A survey.arXiv preprint arXiv:2311.01473, 2023. 1
Pith/arXiv arXiv 2023
-
[38]
H. Xu, Y . Ma, H. C. Liu, D. Deb, H. Liu, J. L. Tang, and A. K. Jain. Adversarial attacks and defenses in images, graphs and text: A review.International journal of automation and com- puting, 17:151–178, 2020. 1
2020
-
[39]
Y . Y . Yan, W. Q. Ren, Y . F. Guo, R. Wang, and X. C. Cao. Image deblurring via extreme channels prior. InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 4003–4011, 2017. 2
2017
-
[40]
R. A. Yeh, C. Chen, L. T. Yian, A. G. Schwing, M. Hasegawa-Johnson, and M. N. Do. Semantic image inpaint- ing with deep generative models. InProceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 5485–5493, 2017. 2
2017
-
[41]
Y . Yu, H. J. Lee, H. Lee, and Y . M. Ro. Defending person detection against adversarial patch attack by using universal defensive frame.IEEE Transactions on Image Processing, 31:6976–6990, 2022. 2
2022
-
[42]
Wide residual net- works.arXiv preprint arXiv:1605.07146, 2016
Sergey Zagoruyko and Nikos Komodakis. Wide residual net- works.arXiv preprint arXiv:1605.07146, 2016. 6, 7
Pith/arXiv arXiv 2016
- [43]
-
[44]
M. K . Zhang, K. Bi, W. Chen, Q. R. Chen, J. F. Guo, and X. Q. Cheng. Causaldiff: Causality-inspired disentan- glement via diffusion model for adversarial defense.arXiv preprint arXiv:2410.23091, 2024. 2
Pith/arXiv arXiv 2024
-
[45]
M. K. Zhang, J. N. Li, W. Chen, J. F. Guo, and X. Q. Cheng. Classifier guidance enhances diffusion-based adversarial pu- rification by preserving predictive information. InECAI 2024, pages 2234–2241. IOS Press, 2024. 1
2024
-
[46]
Y . H. Zhang, W. J. Ruan, F. Wang, and X. W. Huang. Gen- eralizing universal adversarial attacks beyond additive per- turbations. In2020 IEEE International Conference on Data Mining (ICDM), pages 1412–1417. IEEE, 2020. 1, 2, 3
2020
-
[47]
M. N. Zhao, L. H. Zhang, J. W. Ye, H. C. Lu, B. C. Yin, and X. C. Wang. Adversarial training: A survey.arXiv preprint arXiv:2410.15042, 2024. 2
Pith/arXiv arXiv 2024
-
[48]
Y . Z. Zhu, K. Zhang, J. Y . Liang, J. Z. Cao, B. H. Wen, R. Timofte, and G. L. Van. Denoising diffusion mod- els for plug-and-play image restoration. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1219–1229, 2023. 2 Appendix A1. Additional Experimental Results on ImageNet To further verify the scalability of NA...
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.