Pith. sign in

REVIEW 3 major objections 5 minor 74 references

IDATA: Scalable Invertible Diffusion for Unrestricted Adversarial Transfer Attack

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read IDATA: memory-efficient invertible diffusion attacks achieve deeper trajectories at constant memory and better black-box transferability.

desk verdict Useful memory trick for diffusion attacks, but the LFCM ablation contradicts the abstract and the invertibility claim needs verification. read the letter →

arxiv 2608.08734 v1 pith:V2RAI6IF submitted 2026-08-09 cs.CV

classification cs.CV
keywords adversarialattacktransferabilityinvertiblediffusionconstant-memorybackpropagationlow-frequencyperturbationdiscretewavelettransformblack-boxrobustnessunrestrictedexample
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that diffusion-based unrestricted adversarial transfer attacks can escape two bottlenecks at once: the memory cost of backpropagating through many denoising steps, and the tendency of full-latent perturbations to overfit high-frequency, surrogate-specific details. By making the denoising trajectory exactly invertible, IDATA backpropagates through arbitrarily deep trajectories at constant memory, and by restricting perturbations to low-frequency wavelet components it improves cross-model transferability while keeping images visually close to the originals. The paper reports consistent gains over prior diffusion-based attacks across CNN, transformer, and MLP classifiers, under purification defenses, and on fine-grained datasets. A sympathetic reader would care because it turns a scalability limitation into a practical knob: deeper attack optimization no longer demands proportionally more GPU memory.

What carries the argument

The Invertible Diffusion Module (IDM) duplicates the clean image into a primary and an auxiliary latent branch and applies two coupled affine transformations per step: a feature mixing that blends the branches with parameter $p$, followed by scheduled noise addition using coefficients $a_t, b_t$ and the shared noise estimator $\epsilon_\theta$. Because each transformation is sequentially parameterized, the reverse process recovers all intermediate states in closed form from the final latent pair, so gradients flow through the entire trajectory without storing activations. The Low-Frequency Constraint Module (LFCM) concatenates the two branch latents, applies a Haar Discrete Wavelet Transform, adds the learnable perturbation only to the low-frequency approximation $\mathbf{Y}_L$, and reconstructs the latent via the inverse wavelet transform before resuming denoising.

What would settle it

Run a clean image through the forward IDM to step $\tau$ and then through the reverse IDM back to $t=0$, decode, and measure LPIPS/PSNR against the original; additionally, measure peak GPU memory at increasing diffusion depths (e.g., 10, 20, 50 steps) to check whether memory actually stays constant.

Watch

Extended reading notes

Core claim

IDATA's central claim is that an adversarial latent perturbation can be optimized along a deep diffusion trajectory with O(1) activation memory by using an invertible diffusion module whose forward and reverse steps are exact algebraic inverses, and that constraining the perturbation to the low-frequency subspace of intermediate latents (via a discrete wavelet transform) yields superior transferability and imperceptibility compared with full-latent perturbation. The reported ablation shows peak GPU memory dropping from 37.9 GB to 13.1 GB when the invertible module is enabled, while average attack success stays level, and adding the low-frequency constraint improves LPIPS from 0.162 to 0.143 in the same setting. Across normally trained models, defended models, and two fine-grained datasets, IDATA reports the best or near-best average attack success rate among diffusion-based baselines.

Load-bearing premise

The memory-efficiency and attack-quality claims rest on the assumption that the reverse denoising steps are the exact algebraic inverse of the forward steps using the same noise estimator, with no approximation from classifier-free guidance, VAE encoding or decoding, or numerical precision.

Editorial extensions

If this is right

  • Deep trajectory optimization becomes practical: the reported ablation shows peak GPU memory falling from 37.9 GB to 13.1 GB when IDM is enabled, with no loss in attack success.
  • Restricting perturbations to low-frequency latent components is reported to improve both visual imperceptibility (LPIPS 0.162 to 0.143 in the ablation) and transferability compared with full-latent perturbation.
  • IDATA reports the highest average attack success rate on normally trained models across CNN, transformer, and MLP surrogates, and the smallest average success-rate drop under purification defenses such as DiffPure.
  • On CUB-200-2011 and Stanford Cars, IDATA reports the best or near-best transferability across three surrogate settings while keeping LPIPS and FID competitive with or better than diffusion-based baselines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If exact invertibility holds to numerical precision, the same on-demand reconstruction trick could apply to other trajectory-based latent optimizations—classifier guidance, style transfer, or inverse problems—where memory currently limits depth.
  • The frequency decomposition result suggests that intermediate diffusion latents have timestep-dependent frequency semantics, and that steering perturbations toward low-frequency components may be a general recipe for transferability, testable by ablating frequency bands at different timesteps.
  • A reader cannot yet verify the exact-invertibility assumption because the paper does not report a clean-image reconstruction error for the forward-and-reverse IDM chain; that single number would directly test whether the reconstructed states used in backpropagation are faithful.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes IDATA, a diffusion-based unrestricted adversarial transfer attack combining an Invertible Diffusion Module (IDM) and a Low-Frequency Constraint Module (LFCM). IDM reformulates forward diffusion and reverse denoising as coupled invertible transformations, claiming constant-memory backpropagation by reconstructing intermediate states on demand. LFCM applies a Haar DWT to intermediate latents and injects the adversarial perturbation only into the low-frequency components, which the authors argue improves transferability and imperceptibility. The method is evaluated on an ImageNet-compatible benchmark, CUB-200-2011, and Stanford Cars against CNNs, Transformers, MLPs, and defended models, reporting state-of-the-art attack success rates, lower LPIPS/FID, and reduced GPU memory compared with diffusion-based baselines.

Significance. If the central claims hold, IDATA would be a meaningful step for diffusion-based attacks: an O(1)-memory trajectory optimization would enable substantially deeper adversarial optimization than existing methods, and the empirical gains in imperceptibility are consistent across several tables. The paper includes code, compares against external baselines, and provides ablations. However, two load-bearing issues currently undermine the claims: the ablation in Table 4 contradicts the abstract's assertion that LFCM improves transferability, and the stated use of different classifier-free guidance scales during inversion and reverse denoising breaks the exact invertibility on which the O(1)-memory and reconstruction arguments depend. These are internal consistency problems, not disagreements with external consensus, and they must be resolved before the main contributions can be accepted.

major comments (3)
  1. [Abstract, Sec. 4.2, Table 4] The ablation in Table 4 directly contradicts the claim that LFCM improves transferability. Without IDM, adding LFCM reduces AVG from 67.1 to 64.4; with IDM, adding LFCM reduces AVG from 67.3 to 65.0. In both rows LPIPS improves (0.162 to 0.143 and 0.159 to 0.138), so LFCM behaves as an imperceptibility regularizer, not a transferability enhancer. Yet the abstract states that LFCM restricts perturbations to low-frequency subspaces "thereby improving transferability," and Sec. 4.2 says the design "improves cross-model transferability." This is a load-bearing internal inconsistency. The authors must either reposition LFCM as a pure imperceptibility/quality module and revise the abstract accordingly, or provide a controlled comparison under the same protocol as Table 1 that shows a transferability benefit. As written, the full IDATA (65.0) is worse in transferability than its own base pipeline without LFCM (67.1).
  2. [Sec. 4.1 and Sec. 5.1 (Implementation Details)] The O(1)-memory claim relies on exact algebraic invertibility between the forward IDM equations (1)-(2) and the reverse equations (5)-(6). However, Sec. 5.1 states "The guidance scale is 0 during inversion and 1 during reverse denoising." If epsilon_theta is evaluated with different classifier-free guidance scales in the forward and reverse passes, then the reverse transformation is not the inverse of the forward transformation, so reconstructing z_t from z_{t-1} via Eq. (1)-(2) will not recover the actual intermediate state used in the reverse pass. This breaks the on-demand reconstruction procedure that is the basis of the constant-memory backpropagation claim, and it also means gradients are computed through a trajectory that is not the one actually optimized. The paper reports no clean-image reconstruction error for the forward-then-reverse IDM chain, so the exactness assumption is unverified. Please either use the same guidance scale in both directions, provide a reconstruction-error measurement (e.g., max/mean absolute error in latent or pixel space) for the forward IDM followed by IDM^{-1}, and explain how exact inverses are maintained under CFG, or revise the memory-efficiency claim accordingly.
  3. [Sec. 5.1, Sec. 5.3, Fig. 6] Several key hyperparameters—perturbation timestep tau, step size eta, perceptual weight lambda, mixing weight p, and guidance scale g—are selected using the same evaluation benchmarks and surrogate models on which the final results are reported (Fig. 1 and Fig. 6). No held-out validation split is used. In addition, the paper reports no error bars or multiple-seed variance for any table. Given that the margins in Table 1 are often only 1-3 percentage points (e.g., IDATA 62.7 vs. DiffAttack 60.9 for Res-50), the SOTA claim could be an artifact of tuning on the test set. Please report results over at least 3 seeds with standard deviations and, if possible, tune hyperparameters on a separate validation split or demonstrate that the chosen values are not overfit to the evaluation benchmark.
minor comments (5)
  1. [Table 1] The Clean row contains an apparent typo: "3 6.3" should likely be "3.6" or "36.3"; please correct the formatting.
  2. [Table 3] In the Stanford Cars DiffAttack row, "16.20.095" is missing a space or separator between the FID and LPIPS values; please fix the typesetting.
  3. [Table 4] The AVG column in Table 4 is not defined: it is unclear which surrogate model, target set, and evaluation protocol produce these numbers, making it difficult to reconcile with Table 1. Please specify the protocol (e.g., Mix-B surrogate, ImageNet-compatible dataset, same as Table 1).
  4. [Sec. 5.3 (Perturbation Timestep)] The text says "we enforce a strict perceptual budget LPIPS <= 0.14" for the tau ablation, but Fig. 1 appears to show some IDATA points with LPIPS values above 0.14 (e.g., 0.17). Please clarify whether the budget is enforced during optimization or only used as a reporting filter.
  5. [Sec. 4.1 (Difference from EDICT)] The paragraph says IDM is "fundamentally redesigned for adversarial optimization rather than faithful reconstruction," yet the O(1)-memory argument depends on faithful reconstruction of intermediate states. Please clarify how these two statements are reconciled.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: IDM's memory claim follows from explicitly exhibited inverse equations, and LFCM/empirical results are not predicated on their own outputs.

full rationale

The paper's central claims are either structural or empirical, not circular. The O(1)-memory property is argued from the explicit algebraic inverse: Eq. (5)-(6) invert Eq. (1)-(2) step-by-step, so the backpropagation-through-reconstruction claim is a direct consequence of the displayed equations rather than a fitted result. The LFCM uses a defined DWT/IDWT decomposition and injects the perturbation into the low-frequency branch; whether that improves transferability is an empirical question, and Table 4's AVG numbers (67.1->64.4 without IDM; 67.3->65.0 with IDM) actually undercut the paper's own transferability claim. That is an internal-consistency/correctness problem, not a circularity, because LFCM's output is not defined in terms of the headline AVG metric. Self-citations in the related-work section (e.g., references to the authors' earlier invertible-network papers) are background context and are not load-bearing for the IDATA derivation; EDICT, the actual source of the invertible-diffusion idea, is cited as external prior work and explicitly acknowledged. Hyperparameter selection (tau, p, lambda, eta, g) on the evaluation benchmark is ordinary tuning bias, not a case where a fitted parameter is renamed as a prediction. No step in the derivation chain reduces to its own inputs by construction.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central claims rest on standard diffusion and DDIM machinery plus two domain assumptions: exact invertibility of the coupled denoiser steps, and semantic transferability of low-frequency latent components. Hyperparameters tau, p, lambda, eta, and g are tuned on the evaluation benchmarks, which is the main non-derivation input. No new physical or model entities are invented.

free parameters (5)
  • Perturbation timestep tau = 20
    Selected in Sec. 5.3 (Fig. 1) to maximize AVG under an LPIPS budget on the evaluation set; not derived from theory.
  • Mixing weight p = 0.93
    Set by sensitivity analysis in Sec. 5.3 (Fig. 6c) to balance AVG and LPIPS.
  • Perceptual loss weight lambda = 2
    Set by sensitivity analysis in Sec. 5.3 (Fig. 6b); AVG peaks at lambda=2 in the tested range.
  • Step size eta = 0.01
    Set by sensitivity analysis in Sec. 5.3 (Fig. 6a); larger eta increases LPIPS without proportional AVG gains.
  • Guidance scale g = 1
    Set by sensitivity analysis in Sec. 5.3 (Fig. 6d); larger g raises LPIPS.
assumptions (4)
  • domain assumption The DDIM deterministic sampling equations with coefficients a_t and b_t define the denoising trajectory, and the Stable Diffusion VAE encoder and decoder are fixed and differentiable.
    Sec. 3 uses DDIM and Stable Diffusion v2.0 as the generative backbone; the attack inherits all assumptions of that pipeline.
  • standard math The coupled transformations in Eq. (1)-(2) are exactly invertible under the same noise estimator epsilon_theta, enabling exact state reconstruction for backpropagation.
    Algebraic invertibility by construction, but it assumes the same epsilon network and coefficients are used in forward and reverse. The paper does not empirically verify reconstruction fidelity.
  • domain assumption Low-frequency components of diffusion latents are semantically stable and more transferable across architectures, while high-frequency components are surrogate-specific artifacts.
    Motivated by Fig. 3 and prior pixel-space frequency studies (Refs. 14, 50, 65-67); not proved for latent diffusion and partly contradicted by the Table 4 ablation.
  • domain assumption The surrogate classifier and the diffusion model share a label space, and ground-truth label text conditioning is available to the attacker.
    Sec. 5.1 conditions on the ground-truth label name; this is an attack-knowledge assumption that standard black-box transfer attacks do not always satisfy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of IDATA: Scalable Invertible Diffusion for Unrestricted Adversarial Transfer Attack." pith.science (2026). https://pith.science/paper/V2RAI6IF

@misc{pith2026260808734,
  author       = {Pith},
  title        = {Pith review of: IDATA: Scalable Invertible Diffusion for Unrestricted Adversarial Transfer Attack},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V2RAI6IF}},
  note         = {Machine review of arXiv:2608.08734}
}
read the original abstract

Unrestricted adversarial transfer attacks are important for evaluating the black-box robustness of deep visual models. Diffusion-based attacks have shown promising transferability and visual imperceptibility by optimizing adversarial perturbations along denoising trajectories in latent space. However, existing methods are limited by two challenges: memory-intensive multistep backpropagation and frequency-agnostic perturbation over intermediate latents. To address these issues, we propose IDATA, a memory-efficient diffusion framework for unrestricted adversarial transfer attack. IDATA consists of two key components: an Invertible Diffusion Module (IDM) and a Low-Frequency Constraint Module (LFCM). Specifically, IDM reformulates adversarial optimization over diffusion trajectories as an invertible process, enabling constant-memory backpropagation through on-demand reconstruction of intermediate states instead of storing the full denoising chain. Moreover, LFCM leverages Discrete Wavelet Transform (DWT) to decompose latent variables into low- and high-frequency components, restricting perturbations to semantically stable low-frequency subspaces, thereby improving transferability while preserving visual imperceptibility. Extensive experiments on multiple benchmarks and diverse model architectures demonstrate that IDATA consistently outperforms state-of-the-art baselines in attack success rate, memory efficiency, and visual imperceptibility. These results suggest that IDATA is a promising tool for black-box robustness evaluation of deep visual models. Code is available at https://github.com/colourful-pan/IDATA.

Figures

Figures reproduced from arXiv: 2608.08734 by the authors.

Figure 1
Figure 1. Comparison of DiffAttack [5], VENOM [28], ACA [6], and IDATA at different diffusion depths. Circle size represents peak GPU memory usage (GB), while color inten￾sity indicates LPIPS (darker means higher). The “value/value” label denotes GPU memory usage and LPIPS, respectively. First, trajectory-level optimization is prohibitively mem￾ory intensive. Current methods [5, 6, 28, 47, 62] require backprop￾agation through… view at source ↗
Figure 2
Figure 2. Overview of Invertible Diffusion Adversarial Transfer Attack (IDATA) framework. It consists of two core modules: the [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Reconstructions from high- and low-frequency [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Visual comparisons among different adversarial attacks. Please zoom in for a better view. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Comparison of adversarial examples generated by [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: A quantitative study on parameter settings: step [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

74 extracted references · 65 canonical work pages

  1. [1]

    Qianyue Bao, Fang Liu, Yang Liu, Licheng Jiao, Xu Liu, and Lingling Li. 2022. Hi- erarchical Scene Normality-Binding Modeling for Anomaly Detection in Surveil- lance Videos. InProceedings of the 30th ACM international conference on multime- dia. 6103–6112

  2. [2]

    Anand Bhattad, Min Jin Chong, Kaizhao Liang, Bo Li, and David A Forsyth. 2020. Unrestricted Adversarial Examples via Semantic Manipulation. InInternational Conference on Learning Representations. https://openreview.net/forum?id=Sye_ OgHFwH

  3. [3]

    Bin Chen, Yan Feng, Tao Dai, Jiawang Bai, Yong Jiang, Shu-Tao Xia, and Xuan Wang. 2022. Adversarial Examples Generation for Deep Product Quantization Networks on Image Retrieval.IEEE Transactions on Pattern Analysis and Machine Intelligence45, 2 (2022), 1388–1404

  4. [4]

    Bin Chen, Zhenyu Zhang, Weiqi Li, Chen Zhao, Jiwen Yu, Shijie Zhao, Jie Chen, and Jian Zhang. 2025. Invertible Diffusion Models for Compressed Sensing.IEEE Transactions on Pattern Analysis and Machine Intelligence(2025)

  5. [5]

    Jianqi Chen, Hao Chen, Keyan Chen, Yilan Zhang, Zhengxia Zou, and Zhenwei Shi. 2025. Diffusion Models for Imperceptible and Transferable Adversarial Attack.IEEE Transactions on Pattern Analysis and Machine Intelligence47, 2 (2025), 961–977. doi:10.1109/TPAMI.2024.3480519

  6. [6]

    Zhaoyu Chen, Bo Li, Shuang Wu, Kaixun Jiang, Shouhong Ding, and Wenqiang Zhang. 2023. Content-based Unrestricted Adversarial Attack.Advances in Neural Information Processing Systems36 (2023), 51719–51733

  7. [7]

    Zihan Chen, Tianrui Liu, Jun-Jie Huang, Wentao Zhao, Xing Bi, and Meng Wang

  8. [8]

    Zihan Chen, Ziyue Wang, Jun-Jie Huang, Wentao Zhao, Xiao Liu, and Dejian Guan. 2023. Imperceptible Adversarial Attack via Invertible Neural Networks. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 414–424

Show all 74 references
  1. [9]

    Xuelong Dai, Kaisheng Liang, and Bin Xiao. 2024. Advdiff: Generating Unre- stricted Adversarial Examples using Diffusion Models. InEuropean Conference on Computer Vision. Springer, 93–109

  2. [10]

    Zeyu Dai, Shengcai Liu, Rui He, Jiahao Wu, Ning Lu, Wenqi Fan, Qing Li, and Ke Tang. 2025. SemDiff: Generating Natural Unrestricted Adversarial Exam- ples via Semantic Attributes Optimization in Diffusion Models.arXiv preprint arXiv:2504.11923(2025)

  3. [11]

    Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. 2017. Density estimation using Real NVP. InInternational Conference on Learning Representations. https: //openreview.net/forum?id=HkpbnH9lx

  4. [12]

    Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. 2018. Boosting Adversarial Attacks with Momentum. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 9185–9193

  5. [13]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recogn...

  6. [14]

    Mingyuan Fan, Cen Chen, Chengyu Wang, and Jun Huang. 2025. Exploiting Pre- trained Models and Low-Frequency Preference for Cost-effective Transfer-based Attack.ACM Transactions on Knowledge Discovery from Data19, 2 (2025), 1–18

  7. [15]

    Yan Gan, Chengqian Wu, Deqiang Ouyang, Song Tang, Mao Ye, and Tao Xiang

  8. [16]

    Lianli Gao, Qilong Zhang, Jingkuan Song, Xianglong Liu, and Heng Tao Shen

  9. [17]

    LESEP: Boosting Adversarial Transferability via Latent Encoding and Semantic Embedding Perturbations.IEEE Transactions on Circuits and Systems for Video Technology(2024)

  10. [18]

    Xingshuo Han, Guowen Xu, Yuan Zhou, Xuehuan Yang, Jiwei Li, and Tianwei Zhang. 2022. Physical Backdoor Attacks to Lane Detection Systems in Au- tonomous Driving. InProceedings of the 30th ACM International Conference on Multimedia. 2957–2968

  11. [19]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 770–778

  12. [20]

    Goodfellow, Jonathon Shlens, and Christian Szegedy

    Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and Harnessing Adversarial Examples. InInternational Conference on Learning Repre- sentations. https://mlanthology.org/iclr/2015/goodfellow2015iclr-explaining/

  13. [21]

    Jun-Jie Huang, Zihan Chen, Tianrui Liu, Wentao Zhao, Xin Deng, Xinwang Liu, Meng Wang, and Pier Luigi Dragotti. 2025. SMILENet: Unleashing Extra- large Capacity Image Steganography via a Synergistic Mosaic Invertible Hiding Network.arXiv preprint arXiv:2503.05118(2025)

  14. [22]

    Jun-Jie Huang and Pier Luigi Dragotti. 2021. LINN: Lifting Inspired Invertible Neural Network for Image Denoising. In2021 29th European Signal Processing Conference (EUSIPCO). IEEE, 636–640

  15. [23]

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans Trained by a Two Time-scale Update Rule Converge to a Local Nash Equilibrium.Advances in Neural Information Processing Systems 30 (2017)

  16. [24]

    Jun-Jie Huang, Ziyue Wang, Tianrui Liu, Wenhan Luo, Zihan Chen, Wentao Zhao, and Meng Wang. 2024. DeMPAA: Deployable Multi-mini-patch Adversarial Attack for Remote Sensing Image Classification.IEEE Transactions on Geoscience and Remote Sensing62 (2024), 1–13

  17. [25]

    Durk P Kingma and Prafulla Dhariwal. 2018. Glow: Generative Flow with In- vertible 1x1 Convolutions.Advances in Neural Information Processing Systems31 (2018)

  18. [26]

    Jun-Jie Huang and Pier Luigi Dragotti. 2022. WINNet: Wavelet-Inspired Invertible Network for Image Denoising.IEEE Transactions on Image Processing31 (2022), 4377–4392

  19. [27]

    Alexey Kurakin, Ian Goodfellow, Samy Bengio, Yinpeng Dong, Fangzhou Liao, Ming Liang, Tianyu Pang, Jun Zhu, Xiaolin Hu, Cihang Xie, et al. 2018. Adver- sarial Attacks and Defences Competition. InThe NIPS’17 Competition: Building Intelligent Systems. Springer, 195–231

  20. [28]

    Hui Kuurila-Zhang, Haoyu Chen, and Guoying Zhao. 2025. VENOM: Text- driven Unrestricted Adversarial Example Generation with Diffusion Models. arXiv:2501.07922 [cs.CV] https://arxiv.org/abs/2501.07922

  21. [29]

    Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 2013. 3d Object Represen- tations for Fine-Grained Categorization. InProceedings of the IEEE International Conference on Computer Vision Workshops. 554–561

  22. [30]

    Baishuang Li, Siyi Mo, Wenming Yang, Guijin Wang, and Qingmin Liao. 2023. ERR-Net: Facial Expression Removal and Recognition Network with Residual Image.IEEE Transactions on Biometrics, Behavior, and Identity Science5, 4 (2023), 425–434

  23. [31]

    Tengjiao Li, Maosen Li, Yanhua Yang, and Cheng Deng. 2023. Frequency Domain Regularization for Iterative Adversarial Attacks.Pattern Recognition134 (2023), 109075

  24. [32]

    Cassidy Laidlaw and Soheil Feizi. 2019. Functional Adversarial Attacks. In Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Cur- ran Associates, Inc. https://proceedings.neurip...

  25. [33]

    Yanchun Li, Zemin Li, Long Huang, Lingzhi Hu, Li Zeng, and Dongsu Shen. 2025. Adversarial Purification with One-step Guided Diffusion Model.Neural Networks (2025), 107877

  26. [34]

    Jiabei Liu, Weiming Zhuang, Yuanyuan Liu, Yonggang Wen, Jun Huang, and Wei Lin. 2025. Personalized Federated Mutual Learning for Unsupervised Camera- Aware Person Re-Identification.ACM Transactions on Multimedia Computing, Communications and Applications20, 12 (2025), 1–19

  27. [35]

    Yuanbo Li, Cong Hu, Tianyang Xu, and Xiaojun Wu. 2023. When Diffusion Model Meets with Adversarial Attack: Generating Transferable Adversarial Examples on Face Recognition. InInternational Conference on Image and Graphics. Springer, 243–255

  28. [36]

    Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022. A Convnet for the 2020s. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11976–11986

  29. [37]

    Yuyang Long, Qilong Zhang, Boheng Zeng, Lianli Gao, Xianglong Liu, Jian Zhang, and Jingkuan Song. 2022. Frequency Domain Model Augmentation for Adversarial Attack. InEuropean Conference on Computer Vision. Springer, 549–566

  30. [38]

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. InProceedings of the IEEE/CVF International Conference on Computer Vision. 10012–10022

  31. [39]

    Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards Deep Learning Models Resistant to Adversar- ial Attacks. InInternational Conference on Learning Representations. https: //mlanthology.org/iclr/2018/madry2018iclr-deep/

  32. [40]

    Stephane G Mallat. 2002. A Theory for Multiresolution Signal Decomposition: the Wavelet Representation.IEEE Transactions on Pattern Analysis and Machine Intelligence11, 7 (2002), 674–693

  33. [41]

    Cheng Luo, Qinliang Lin, Weicheng Xie, Bizhu Wu, Jinheng Xie, and Linlin Shen. 2022. Frequency-driven Imperceptible Adversarial Attack on Semantic Similarity. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15315–15324

  34. [42]

    Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. 2017. Universal Adversarial Perturbations. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1765–1773

  35. [43]

    Muzammal Naseer, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Fatih Porikli. 2020. A Self-supervised Approach for Adversarial Robustness. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 262–271

  36. [44]

    Yuxi Mi, Yuge Huang, Jiazhen Ji, Hongquan Liu, Xingkun Xu, Shouhong Ding, and Shuigeng Zhou. 2022. Duetface: Collaborative Privacy-preserving Face Recognition via Channel Splitting in the Frequency Domain. InProceedings of the 30th ACM International Conference on Multimedia. 6...

  37. [45]

    Yi Pan, Jun-Jie Huang, Zihan Chen, Wentao Zhao, and Ziyue Wang. 2024. SVASTIN: Sparse Video Adversarial Attack via Spatio-Temporal Invertible Neural Networks. In2024 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6

  38. [46]

    Zhengzhao Pan, Hua Chen, and Xiaogang Zhang. 2025. DiffAdvMAP: Flexible Diffusion-Based Framework for Generating Natural Unrestricted Adversarial Examples. InForty-second International Conference on Machine Learning. https: //openreview.net/forum?id=q2s4DLsegO

  39. [47]

    Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, and An- imashree Anandkumar. 2022. Diffusion Models for Adversarial Purification. InProceedings of the 39th International Conference on Machine Learning (Pro- ceedings of Machine Learning Research, Vol. 162), Kam...

  40. [48]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 10684–10695

  41. [49]

    Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang- Chieh Chen. 2018. Mobilenetv2: Inverted Residuals and Linear Bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4510–4520

  42. [50]

    Zihao Pan, Weibin Wu, Yuhang Cao, and Zibin Zheng. 2024. SCA: Highly Ef- ficient Semantic-Consistent Unrestricted Adversarial Attack.arXiv preprint arXiv:2410.02240(2024)

  43. [51]

    Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Net- works for Large-Scale Image Recognition. InInternational Conference on Learning Representations. https://mlanthology.org/iclr/2015/simonyan2015iclr-very/

  44. [52]

    Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the Inception Architecture for Computer Vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2818–2826

  45. [53]

    Brubaker

    Yash Sharma, Gavin Weiguang Ding, and Marcus A. Brubaker. 2019. On the Effec- tiveness of Low Frequency Perturbations. InProceedings of the 28th International Joint Conference on Artificial Intelligence(Macao, China)(IJCAI’19). AAAI Press, 3389–3396

  46. [54]

    Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. 2021. Training Data-efficient Image Transformers & Distillation through Attention. InInternational Conference on Machine Learning. PMLR, 10347–10357

  47. [55]

    Florian Tramèr, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel. 2018. Ensemble Adversarial Training: Attacks and Defenses. InInternational Conference on Learning Representations. https: //openreview.net/forum?id=rkZvSe-RZ

  48. [56]

    Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. 2021. Mlp-mixer: An all-mlp Architecture for Vision.Advances in Neural Information Processing Systems34 ...

  49. [57]

    Bram Wallace, Akash Gokul, Stefano Ermon, and Nikhil Naik. 2023. End-to-end Diffusion Latent Optimization Improves Classifier Guidance. InProceedings of the IEEE/CVF International Conference on Computer Vision. 7280–7290

  50. [58]

    Bram Wallace, Akash Gokul, and Nikhil Naik. 2023. EDICT: Exact Diffusion Inversion via Coupled Transformations. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22532–22541

  51. [59]

    Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie

  52. [60]

    Ziyue Wang, Jun-Jie Huang, Tianrui Liu, Zihan Chen, Wentao Zhao, Xiao Liu, Yi Pan, and Lin Liu. 2023. Multi-patch Adversarial Attack for Remote Sensing Image Classification. InAsia-Pacific Web (APWeb) and Web-Age Information Management (W AIM) Joint International Conference on...

  53. [61]

    Cihang Xie, Zhishuai Zhang, Yuyin Zhou, Song Bai, Jianyu Wang, Zhou Ren, and Alan L Yuille. 2019. Improving Transferability of Adversarial Examples with Input Diversity. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2730–2739

  54. [62]

    Wenzhuo Xu, Kai Chen, Ziyi Gao, Zhipeng Wei, Jingjing Chen, and Yu-Gang Jiang

  55. [63]

    Jinyi Wang, Zhaoyang Lyu, Dahua Lin, Bo Dai, and Hongfei Fu. 2022. Guided Diffusion Model for Adversarial Purification.arXiv preprint arXiv:2205.14969 (2022)

  56. [64]

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang

  57. [65]

    Changfei Zhao, Xinyang Deng, and Wen Jiang. 2024. Improving Adversarial Transferability through Frequency Enhanced Momentum.Inf. Sci.665, C (April 2024), 14 pages. doi:10.1016/j.ins.2024.120409

  58. [66]

    Hegui Zhu, Yuchen Ren, Chong Liu, Xiaoyan Sui, and Libo Zhang. 2024. Frequency-based Methods for Improving the Imperceptibility and Transferability of Adversarial Examples.Applied Soft Computing150 (2024), 111088

  59. [67]

    InProceedings of the 32nd ACM International Conference on Multimedia

    Highly Transferable Diffusion-based Unrestricted Adversarial Attack on Pre-trained Vision-language Models. InProceedings of the 32nd ACM International Conference on Multimedia. 748–757

  60. [68]

    Hao Ye, Mengshi Qi, Zhaohong Liu, Liang Liu, and Huadong Ma. 2025. Safedriverag: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-augmented Generation. InProceedings of the 33rd ACM International Conference on Multimedia. 11170–11178

  61. [73]

    Jiang Zhu, Lingping Tan, Yanchun Li, Shujuan Tian, Jianqi Li, and Yaonan Wang

  62. [2011]

    The Caltech-ucsd Birds-200-2011 Dataset. (2011)

  63. [2018]

    InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 586–595

  64. [2020]

    InEuropean Confer- ence on Computer Vision

    Patch-wise Attack for Fooling Deep Neural Network. InEuropean Confer- ence on Computer Vision. Springer, 307–322

  65. [2024]

    InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    Invertible Mosaic Image Hiding Network for Very Large Capacity Image Steganography. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 4520–4524

  66. [2025]

    Guided Adversarial Attack in the Low-Frequency Space.IEEE Transactions on Multimedia(2025)

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.