Pith. sign in

REVIEW 4 major objections 6 minor 147 references

I Stolenly Swear That I Am Up to (No) Good: Design and Evaluation of Model Stealing Attacks

T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper argues that model stealing attacks on image classifiers are not standardised and provides a threat model and comparison framework to make attacks with the same attacker knowledge directly comparable.

desk verdict A useful first systematization of model stealing attacks; the headline comparability statistic rests on subjective re-coding that needs to be released and stress-tested. read the letter →

arxiv 2508.21654 v1 pith:AKSUD4U6 submitted 2025-08-29 cs.CR cs.LG

classification cs.CRcs.LG
keywords modelstealingextractionthreatevaluationmethodologysubstitutetrainingimageclassificationattackcomparabilityquerybudget
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Model stealing attacks let an adversary clone a machine-learning model by querying it and training a substitute. This paper argues that the field cannot measure progress because attack evaluations are not standardised: papers differ in what the attacker knows, what outputs the target model returns, how many queries are allowed, and which metrics are reported. It builds a detailed threat model for substitute-training attacks on image classifiers and a comparison framework that divides attacks into 24 attacker-knowledge segments, so that only attacks in the same segment are fairly comparable. Applying this to 47 prior works, it finds that only a small fraction of attacks can actually be compared with each other. If the field adopts the proposed best practices, new attacks can be placed in a specific segment and evaluated against the right baselines, making state-of-the-art claims testable.

What carries the argument

The central object is the threat model and the comparison framework built on it. The threat model decomposes the attacker into knowledge (data type, output type, architecture match, and pre-trained-model use), capabilities (query budget), and goals (accuracy, fidelity, transferability). The comparison framework turns the knowledge axis into a 24-segment diagram: one split for same versus different architecture, and within each side, circular sectors for data-free, non-problem-domain, problem-domain, and original data, each divided by labels, probabilities, or explanations. The rule is that only attacks in the same segment can be fairly compared, and query counts, ideally normalised per targe

What would settle it

Re-derive the paper's Table 1 classification of the 47 attacks from their experimental sections with two independent annotators. If the annotators disagree on the attacker-data category or the output type for more than about a fifth of the papers, the segment counts and the claim that only a small fraction of attacks are comparable would not be robust. A simpler check: verify whether the largest segment in Figure 1 still contains at most 11 papers when ambiguous papers are excluded instead of placed into every possible segment.

Watch

Extended reading notes

Core claim

The paper's central claim is that comparability of model stealing attacks is only meaningful when attacks share the same attacker knowledge. It defines that knowledge along three axes: what data the attacker has (original data, problem-domain data, non-problem-domain data, or none), what the target model returns (labels, confidence scores, or explanations and gradients), and whether the substitute architecture is the same as or different from the target's. It maps 47 substitute-training attacks on image classifiers onto these axes, producing 24 possible knowledge segments. It then shows that the largest segment contains at most 11 papers, a quarter of segments have no prior work, and a third

Load-bearing premise

The comparison statistics assume that a paper's threat model can be reliably reconstructed from its experimental setup, even when the paper never stated a threat model.

Editorial extensions

If this is right

  • New attacks should report accuracy and fidelity on the same test data as baselines drawn from the same threat-model segment; otherwise they are not comparable.
  • Attack efficiency should be reported not only as total query count but also relative to the target model's training-set size, which reveals that many data-free attacks need over 100 queries per training sample.
  • Authors should state whether the target, substitute, and any auxiliary models are pre-trained, because pre-training materially changes the attacker's knowledge and is often omitted.
  • Transferability scores from different papers cannot be compared until the adversarial perturbation method and its strength are standardised.
  • Attacks with very low query budgets, under 1,000 and even under 100 queries, are understudied despite being the most practical real-world setting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the framework were adopted as a community standard, the 24 segments could become the basis of a public leaderboard where ranking happens strictly within a segment, giving the field a concrete artefact for measuring progress.
  • The paper leaves pre-trained model knowledge out of its Table 1 classification but argues it matters; adding it as a fourth knowledge axis would split the 24 segments further and likely reduce comparability even more.
  • The paper's choice to count held-out parts of the target's original dataset as 'original data' is one of several plausible definitions; if the field adopts a stricter definition, some attacks would move to the problem-domain segment and change which baselines are fair.
  • The framework's logic transfers to other domains, but the paper only gestures at that transfer; adapting the comparison to tasks like text or graph learning would require redefining what counts as task accuracy and fidelity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper argues that model stealing attack evaluations are not standardised and proposes the first comprehensive threat model and comparison framework for substitute-training attacks against image classifiers. The authors define attacker knowledge (data type, output type, architecture knowledge, pre-training), capabilities (query budget), and goals (accuracy/fidelity/transferability); classify 47 prior papers into Table 1 and a 24-segment diagram; and use the resulting coverage to claim that only a small fraction of prior attacks are comparable, that a quarter of the segments are empty, and that several configurations are thinly studied. They further analyse dataset/architecture/query usage in prior experiments, derive best practices (R1–R3), and list open research questions.

Significance. If the classification is reliable, the framework would be a genuinely useful community resource: it provides a common vocabulary, a structured way to select comparable baselines, and concrete reporting recommendations. The paper is unusually transparent about its coding ambiguities and about missing information in prior work, and it offers falsifiable observations about segment coverage and experimental-setup frequencies. The open questions are thoughtful and practical. The main risk is that the headline quantitative claims rest on a subjective re-coding that is neither released nor stress-tested; as currently presented, the evidence is not yet sufficient to support the central 'only a small fraction of prior attacks can be compared' claim at the strength asserted.

major comments (4)
  1. [§3.4.1, Table 1, Figure 1] The central quantitative claim ('only a small fraction of prior attacks can be compared', 'largest segment at most 11 papers', 'a quarter of segments empty') is computed from assignments that the authors themselves describe as derived from experimental setups with several reasonable alternative resolutions. Section 3.4.1 explicitly lists unresolved ambiguities for original vs problem-domain data, problem-domain vs non-problem-domain data, and the data-free definition, and Section 3.4 states that 'there are also reasonable arguments for other solutions.' Because most prior papers did not define threat models, this is a subjective re-coding. No coding artifact or per-paper annotation is released, so the reader cannot audit the assignments. Please provide a sensitivity analysis under the alternative coding rules (varying the disputed thresholds and the handling of unknown architecture/outpu
  2. [§3.3, Table 1 (Metrics column)] Table 1's Metrics column appears to list 'AFT' for every analysed paper. If correct, the table contradicts the text in §3.4.3 ('no metric was reported in all papers') and in R2.2 ('none of the metrics was reported for every previous attack'), and it removes the evidence that goal/metric reporting is insufficient. Please correct the column to reflect which metrics each paper actually reports, or revise the textual claims. As printed, this is an internal inconsistency in evidence used to motivate R2.2.
  3. [§4.2, Figure 1] The comparability framework partitions attacks only by attacker knowledge (data/output/architecture), while attacker capabilities (query budget) and goals are excluded from the segment definition. §4.1 argues that capabilities and goals align with evaluation and can be 'easily adapted' by reporting the right numbers, but this conflates comparability of reported numbers with comparability of attack difficulty. Two attacks in the same knowledge segment but with query budgets of 100 and 10^6 are not directly comparable without an agreed query-budget or efficiency curve. Please clarify under which conditions 'same segment' is sufficient (e.g., same or matched query budget, same test data and metric definitions) or refine the partition.
  4. [§3, corpus construction] The paper does not report how the 47-paper corpus was constructed: no search venues, inclusion/exclusion criteria, time window, or screening process are given. The coverage statistics and the 'first comprehensive' claim are therefore claims about an unstated sample. Please specify the corpus construction methodology or, at minimum, provide the full list with explicit selection rules; otherwise the generalisation from 47 papers to 'the field' is not reproducible.
minor comments (6)
  1. [Eq. (3)] The transferability definition contains a typo: the consequent should be f(x'_i) != f(x_i), not f(x'_i) != f(x'_i). As written, the right-hand side is a tautological inequality that is always false, making the indicator identically zero.
  2. [§3.2] 'more than 1,000,0000 queries' should read 'more than 1,000,000 queries'.
  3. [Figure 1 and §4.3] The shading terminology is inconsistent: §4.2 says empty segments are filled in dark grey, while §4.3(ii) refers to 'the grey segments' as having no prior work. Please clarify the legend and distinguish unknown-assignment grey from empty dark-grey.
  4. [Table 1 / Figure 1 counting unit] Several papers appear in multiple segments (e.g., [14](1)/(2), [97], [103]). The text explains this, but the counting unit for the 'largest segment' statistic should be stated explicitly: segment entries are paper-segment incidences, not unique papers, and the claim 'at most 11 papers' needs to define how duplicates and unknown-assignment entries are counted.
  5. [§1 vs §4.3] The contribution bullet says 'a quarter of configurations have only been studied in one or two works', while §4.3 says a third have at most one publication and more than half have at most three publications. These numbers overlap but are not expressed consistently; please reconcile.
  6. [§6.1, §6.2] Several forward references to Section 7 appear before that section is introduced (e.g., 'as discussed in Section 7.1'). Consider renumbering or adding cross-reference placeholders.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the comparability framework and statistics are newly derived from a 47-paper corpus; self-citations provide context but are not load-bearing.

full rationale

The paper makes no fitted-parameter prediction and contains no derivation chain whose output equals its input. The central empirical claim — that only a small fraction of prior substitute-training attacks are comparable — is computed by coding 47 papers into Table 1 and Figure 1 according to attacker knowledge, capabilities, and goals. That coding is disclosed as subjective ('The threat models reported in Table 1 are primarily derived by us from the experimental setup of the papers', Section 3) and the paper itself flags the ambiguity of its boundary decisions (Section 3.4.1: 'The way we resolve raised questions can only be taken as suggestions, as there are also reasonable arguments for other solutions'). This is a reproducibility/sensitivity limitation, not a circular reduction: the categories are not defined in terms of the conclusion, and a different defensible coding would not make the claim true by construction. The threat-model dimensions are sourced externally (Biggio & Roli [8]) and from Jagielski et al. [36]; the authors' own survey [61] is used for goal terminology, side-channel scope, and the efficiency-score suggestion, but these are contextual definitions rather than load-bearing premises. No uniqueness theorem is imported, no ansatz is smuggled via citation, and no known result is merely renamed. The novelty claims ('first comprehensive threat model', 'first framework') are positioning statements, not derived results. Hence the derivation is self-contained for circularity purposes; any concern about the unreleased coding belongs to correctness/reproducibility, not circularity.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical or mathematical entities. Its load-bearing choices are hand-chosen category boundaries (query bins, data-type threshold, 24-cell partition) and domain assumptions about what prior papers' experimental setups reveal. The absence of a released coding dataset makes these choices hard to audit externally.

free parameters (3)
  • Query budget bins = <10k, 10k-100k, 100k-1m, >1m
    Hand-chosen grouping in Section 3.2 to categorize attacker capabilities; all query-related statistics and comparability claims depend on these bins.
  • Data-type classification threshold = >50% nPD data => nPD classification
    In Section 3.4.1 the authors resolve ambiguous datasets by classifying all works with more than half non-problem-domain samples as nPD. This heuristic affects the Table 1 categories and Figure 1.
  • Threat model segment partition = 4 data types x 3 output types x 2 architecture settings = 24 cells
    The authors choose these dimensions and values for the comparison space; the statement that only a small fraction of attacks are comparable depends on this particular partition.
assumptions (4)
  • domain assumption The 47 reviewed papers are an adequate, representative sample of substitute-training model stealing attacks on image classifiers.
    Section 2.1 selects this group without a formal systematic search protocol; all statistics and coverage claims in Figure 1 inherit this sampling choice.
  • domain assumption Experimental setups in prior papers reveal the true threat model even when the papers do not state one.
    Section 3 derives threat models primarily from experimental setups. If setups are not faithful to the intended attack, the categorization collapses.
  • domain assumption Accuracy and fidelity are unequivocally defined and sufficient for comparing effectiveness.
    Section 4.1 asserts this while excluding transferability due to measurement variability; the framework's comparability claims rely on accuracy and fidelity being comparable across papers.
  • domain assumption The threat model aspects from Biggio & Roli [8] and the goal categories from Jagielski et al. [36] and Oliynyk et al. [61] apply to model stealing.
    The framework is built directly on these external taxonomies rather than being derived from first principles within the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of I Stolenly Swear That I Am Up to (No) Good: Design and Evaluation of Model Stealing Attacks." pith.science (2026). https://pith.science/paper/AKSUD4U6

@misc{pith2026250821654,
  author       = {Pith},
  title        = {Pith review of: I Stolenly Swear That I Am Up to (No) Good: Design and Evaluation of Model Stealing Attacks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AKSUD4U6}},
  note         = {Machine review of arXiv:2508.21654}
}
read the original abstract

Model stealing attacks endanger the confidentiality of machine learning models offered as a service. Although these models are kept secret, a malicious party can query a model to label data samples and train their own substitute model, violating intellectual property. While novel attacks in the field are continually being published, their design and evaluations are not standardised, making it challenging to compare prior works and assess progress in the field. This paper is the first to address this gap by providing recommendations for designing and evaluating model stealing attacks. To this end, we study the largest group of attacks that rely on training a substitute model -- those attacking image classification models. We propose the first comprehensive threat model and develop a framework for attack comparison. Further, we analyse attack setups from related works to understand which tasks and models have been studied the most. Based on our findings, we present best practices for attack development before, during, and beyond experiments and derive an extensive list of open research questions regarding the evaluation of model stealing attacks. Our findings and recommendations also transfer to other problem domains, hence establishing the first generic evaluation methodology for model stealing attacks.

Figures

Figures reproduced from arXiv: 2508.21654 by the authors.

Figure 1
Figure 1. Model stealing attacks against image classifiers categorised accordingly to the attacker’s knowledge. [PITH_FULL_IMAGE:figures/full_fig_p012_1.png] view at source ↗
Figure 2
Figure 2. Dataset statistics. disclose a lot of information about the target model and the original data. However, attacks with a stronger assumption are more useful for estimating the worst-case scenario when testing defences. Designing new attacks in the lower part of the diagram is, therefore, more beneficial for defence development. (v) Complementing the previous point, we notice that the area of attacks that assume the a… view at source ↗
Figure 3
Figure 3. Target and substitute model statistics. 5.2 Target Model Another important factor besides the dataset is the target model architecture; statistics of usage of popular architectures are shown in Figure 3a. As with datasets, if several target models were trained in one paper, they were counted multiple times. The most popular choices are ResNet34 (20), AlexNet (9), LeNet (9), ResNet18 (9), VGG16 (6), ResNet50 (6), and… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Number of queries reported in papers targeting models trained on various datasets. Some of the papers contribute with [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 5
Figure 5. Figure 5: Queries per target model’s training data sample, plotted against the total number of queries. The grey area highlights attacks [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

147 extracted references · 50 canonical work pages

  1. [1]

    Real Attackers Don’t Compute Gradients

    Giovanni Apruzzese, Hyrum S. Anderson, Savino Dambra, David Freeman, Fabio Pierazzi, and Kevin Roundy. 2023. “Real Attackers Don’t Compute Gradients”: Bridging the Gap Between Adversarial ML Research and Practice. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE, Raleigh, NC, USA, 339–364. doi:10.1109/SaTML54575.2023.00031

  2. [2]

    Armstrong, Alistair Moffat, William Webber, and Justin Zobel

    Timothy G. Armstrong, Alistair Moffat, William Webber, and Justin Zobel. 2009. Improvements That Don’t Add up: Ad-Hoc Retrieval Results since 1998. In Proceedings of the 18th ACM Conference on Information and Knowledge Management . ACM, Hong Kong China, 601–610. doi:10.1145/ 1645953.1646031

  3. [3]

    Buse Gul Atli, Sebastian Szyller, Mika Juuti, Samuel Marchal, and N. Asokan. 2020. Extraction of Complex DNN Models: Real Threat or Boogeyman? In Engineering Dependable and Secure Machine Learning Systems, Onn Shehory, Eitan Farchi, and Guy Barash (Eds.). Vol. 1272. Springer International Publishing, Cham, 42–57. doi:10.1007/978-3-030-62144-5_4

  4. [4]

    Antonio Barbalau, Adrian Cosma, Radu Tudor Ionescu, and Marius Popescu. 2020. Black-Box Ripper: Copying Black-Box Models Using Generative Evolutionary Algorithms. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 20120–20129

  5. [5]

    Lejla Batina, Shivam Bhasin, Dirmanto Jap, and Stjepan Picek. 2019. CSI NN: Reverse Engineering of Neural Network Architectures through Electromagnetic Side Channel. In 28th USENIX Security Symposium (USENIX Security 19) . USENIX Association, Santa Clara, CA, 515–532

  6. [6]

    James Beetham, Navid Kardan, Ajmal Mian, and Mubarak Shah. 2023. Dual Student Networks for Data-Free Model Stealing. In Proceedings of the 11th International Conference on Learning Representations (ICLR)

  7. [7]

    Lukas Bieringer, Kevin Paeth, Jochen Stängler, Andreas Wespi, Alexandre Alahi, and Kathrin Grosse. 2025. Position: A Taxonomy for Reporting and Describing AI Security Incidents. arXiv preprint (2025). arXiv:2412.14855

  8. [8]

    Battista Biggio and Fabio Roli. 2018. Wild Patterns: Ten Years after the Rise of Adversarial Machine Learning. Pattern Recognition 84 (Dec. 2018), 317–331. doi:10.1016/j.patcog.2018.07.023

Show all 147 references
  1. [9]

    Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and Alexey Kurakin. 2019. On Evaluating Adversarial Robustness. arXiv preprint (2019). doi:10.48550/arXiv.1902.06705 arXiv:1902.06705 26 Oliynyk et al

  2. [10]

    Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. 2017. ZOO: Zeroth Order Optimization Based Black-box Attacks to Deep Neural Networks without Training Substitute Models. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security . ACM, ...

  3. [11]

    Sizhe Chen, Zhehao Huang, Qinghua Tao, and Xiaolin Huang. 2024. Query Attack by Multi-Identity Surrogates. IEEE Trans. Artif. Intell. 5, 2 (Feb. 2024), 684–697. doi:10.1109/TAI.2023.3257276

  4. [12]

    Yanjiao Chen, Rui Guan, Xueluan Gong, Jianshuo Dong, and Meng Xue. 2023. D-DAE: Defense-Penetrating Model Extraction Attacks. In 2023 IEEE Symposium on Security and Privacy (SP) . IEEE, San Francisco, CA, USA, 382–399. doi:10.1109/SP46215.2023.10179406

  5. [13]

    Moser, Alina Oprea, Battista Biggio, Marcello Pelillo, and Fabio Roli

    Antonio Emanuele Cinà, Kathrin Grosse, Ambra Demontis, Sebastiano Vascon, Werner Zellinger, Bernhard A. Moser, Alina Oprea, Battista Biggio, Marcello Pelillo, and Fabio Roli. 2023. Wild Patterns Reloaded: A Survey of Machine Learning Security against Training Data Poisoning. A...

  6. [14]

    Berriel, Claudine Badue, Alberto F

    Jacson Rodrigues Correia-Silva, Rodrigo F. Berriel, Claudine Badue, Alberto F. De Souza, and Thiago Oliveira-Santos. 2018. Copycat CNN: Stealing Knowledge by Persuading Confession with Random Non-Labeled Data. In 2018 International Joint Conference on Neural Networks (IJCNN) ....

  7. [15]

    Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. 2021. RobustBench: A Standardized Adversarial Robustness Benchmark. In Thirty-Fifth Conference on Neural Information Processing Syste...

  8. [16]

    2020-07-13/2020-07-18

    Francesco Croce and Matthias Hein. 2020-07-13/2020-07-18. Reliable Evaluation of Adversarial Robustness with an Ensemble of Diverse Parameter- Free Attacks. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. ...

  9. [17]

    David DeFazio and Arti Ramesh. 2020. Adversarial Model Extraction on Graph Neural Networks. In Proceedings of the AAAI Workshop on Deep Learning on Graphs: Methodologies and Applications (DLGMA) . New York, NY, USA

  10. [18]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image Is Worth 16x16 Words: Transformers for Image Recognit...

  11. [19]

    Vijay Rao, and Valentina E

    Vasisht Duddu, Debasis Samanta, D. Vijay Rao, and Valentina E. Balas. 2019. Stealing Neural Networks via Timing Side Channels. arXiv preprint (2019). arXiv:1812.11720

  12. [20]

    Manuel Fernández-Delgado, Eva Cernadas, Senén Barro, and Dinani Amorim. 2014. Do We Need Hundreds of Classifiers to Solve Real World Classification Problems? Journal of Machine Learning Research 15, 1 (2014), 3133–3181

  13. [21]

    Katarina Foss-Solbrekk. 2021. Three Routes to Protecting AI Systems and Their Algorithms under IP Law: The Good, the Bad and the Ugly.Journal of Intellectual Property Law & Practice 16, 3 (2021), 247–258

  14. [22]

    Xueluan Gong, Yanjiao Chen, Wenbin Yang, Guanghao Mei, and Qian Wang. 2021. InverseNet: Augmenting Model Extraction Attacks with Training Data Inversion. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence . International Joint Conferences...

  15. [23]

    Goodfellow, Jonathon Shlens, and Christian Szegedy

    Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and Harnessing Adversarial Examples. In Proceedings of the 3rd International Conference on Learning Representations (ICLR)

  16. [24]

    Maybank, and Dacheng Tao

    Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dacheng Tao. 2021. Knowledge Distillation: A Survey. Int J Comput Vis 129, 6 (June 2021), 1789–1819. doi:10.1007/s11263-021-01453-z

  17. [25]

    Kathrin Grosse and Alexandre Alahi. 2024. A qualitative AI security risk assessment of autonomous vehicles. Transportation Research Part C: Emerging Technologies 169 (2024), 104797

  18. [26]

    Besold, and Alexandre M

    Kathrin Grosse, Lukas Bieringer, Tarek R. Besold, and Alexandre M. Alahi. 2024. Towards More Practical Threat Models in Artificial Intelligence Security. In Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24) . USENIX Association, San Diego, CA, USA, 4891–4908

  19. [27]

    Besold, Battista Biggio, and Katharina Krombholz

    Kathrin Grosse, Lukas Bieringer, Tarek R. Besold, Battista Biggio, and Katharina Krombholz. 2023. Machine Learning Security in Industry: A Quantitative Survey. IEEE Transactions on Information Forensics and Security 18 (2023), 1749–1762. doi:10.1109/TIFS.2023.3241234

  20. [28]

    Jindong Gu, Xiaojun Jia, Pau de Jorge, Wenqain Yu, Xinwei Liu, Avery Ma, Yuan Xun, Anjun Hu, Ashkan Khakzar, Zhijiang Li, Xiaochun Cao, and Philip Torr. 2023. A Survey on Transferability of Adversarial Examples across Deep Neural Networks. arXiv preprint (2023). arXiv:2310.17626

  21. [29]

    Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. 2024. The False Promise of Imitating Proprietary Language Models. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR)

  22. [30]

    Laura Gustafson, Chloe Rolland, Nikhila Ravi, Quentin Duval, Aaron Adcock, Cheng-Yang Fu, Melissa Hall, and Candace Ross. 2023. FACET: Fairness in Computer Vision Evaluation Benchmark. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Paris, France, 2...

  23. [31]

    Xinlei He, Jinyuan Jia, Michael Backes, Neil Zhenqiang Gong, and Yang Zhang. 2021. Stealing Links from Graph Neural Networks. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21) . USENIX Association, Virtual, Online, 2669–2686

  24. [32]

    Yingzhe He, Guozhu Meng, Kai Chen, Xingbo Hu, and Jinwen He. 2021. DRMI: A Dataset Reduction Technology Based on Mutual Information for Black-Box Attacks. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21) . USENIX Association, Vancouver, B.C., Canada, 1901–1918

  25. [33]

    Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowledge in a Neural Network. arXiv preprint (2015). arXiv:1503.02531 I Stolenly Swear That I Am Up to (No) Good: Design and Evaluation of Model Stealing Attacks 27

  26. [34]

    Vlad Hondru and Radu Tudor Ionescu. 2025. Towards Few-Call Model Stealing via Active Self-Paced Knowledge Distillation and Diffusion-Based Image Generation. Artificial Intelligence Review 58, 254 (2025). doi:10.1007/s10462-025-11184-z

  27. [35]

    Chi Hong, Jiyue Huang, Robert Birke, and Lydia Y. Chen. 2023. Exploring and Exploiting Data-Free Model Stealing. In Machine Learning and Knowledge Discovery in Databases: Research Track , Danai Koutra, Claudia Plant, Manuel Gomez Rodriguez, Elena Baralis, and Francesco Bonchi ...

  28. [36]

    Matthew Jagielski, Nicholas Carlini, David Berthelot, Alex Kurakin, and Nicolas Papernot. 2020. High Accuracy and High Fidelity Extraction of Neural Networks. In Proceedings of the 29th USENIX Security Symposium (USENIX Security 20) . USENIX Association, 1345–1362

  29. [37]

    Balasubramanian, and Balaji Krishnamurthy

    Surgan Jandial, Yash Khasbage, Arghya Pal, Vineeth N. Balasubramanian, and Balaji Krishnamurthy. 2022. Distilling the Undistillable: Learning from a Nasty Teacher. In Computer Vision – ECCV 2022 , Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal...

  30. [38]

    Akshit Jindal, Vikram Goyal, Saket Anand, and Chetan Arora. 2024. Army of Thieves: Enhancing Black-Box Model Extraction via Ensemble Based Sample Selection. In 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). IEEE, Waikoloa, HI, USA, 3811–3820. doi:1...

  31. [39]

    Mika Juuti, Sebastian Szyller, Samuel Marchal, and N. Asokan. 2019. PRADA: Protecting Against DNN Model Stealing Attacks. In 2019 IEEE European Symposium on Security and Privacy (EuroS&P) . IEEE, Stockholm, Sweden, 512–527. doi:10.1109/EuroSP.2019.00044

  32. [40]

    Sanjay Kariyappa, Atul Prakash, and Moinuddin K Qureshi. 2021. MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient Estimation. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Nashville, TN, USA, 13809–13818. doi:10.1109/CVPR4...

  33. [41]

    Pratik Karmakar and Debabrota Basu. 2023. Marich: A Query-Efficient Distributionally Equivalent Model Extraction Attack Using Public Data. In Proceedings of the 37th International Conference on Neural Information Processing Systems (Nips ’23) . Curran Associates Inc., New Orle...

  34. [42]

    Kacem Khaled, Gabriela Nicolescu, and Felipe Gohring De Magalhaes. 2022. Careful What You Wish For: On the Extraction of Adversarially Trained Models. In 2022 19th Annual International Conference on Privacy, Security & Trust (PST) . IEEE, Fredericton, NB, Canada, 1–10. doi:10....

  35. [43]

    Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. 2020. Big Transfer (BiT): General Visual Representation Learning. In Computer Vision – ECCV 2020 , Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm...

  36. [44]

    Parikh, Nicolas Papernot, and Mohit Iyyer

    Kalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot, and Mohit Iyyer. 2020. Thieves on Sesame Street! Model Extraction of BERT-based Apis. In Proceedings of the 8th International Conference on Learning Representations (ICLR)

  37. [45]

    Souvik Kundu, Qirui Sun, Yao Fu, Massoud Pedram, and Peter A. Beerel. 2021. Analyzing the Confidentiality of Undistillable Teachers in Knowledge Distillation. In Proceedings of the 35th International Conference on Neural Information Processing Systems (Nips ’21) . Curran Assoc...

  38. [46]

    Goodfellow, and Samy Bengio

    Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. 2017. Adversarial Machine Learning at Scale. InProceedings of the 5th International Conference on Learning Representations (ICLR)

  39. [47]

    Pan Li, Peizhuo Lv, Kai Chen, Shengzhi Zhang, Yuling Cai, and Fan Xiang. 2025. A Model Stealing Attack Against Multi-Exit Networks. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, Hyderabad, India, 1–5. doi:10.110...

  40. [48]

    Siyuan Liang, Aishan Liu, Jiawei Liang, Longkang Li, Yang Bai, and Xiaochun Cao. 2022. Imitated Detectors: Stealing Knowledge of Black-box Object Detectors. In Proceedings of the 30th ACM International Conference on Multimedia . ACM, Lisboa Portugal, 4839–4847. doi:10.1145/350...

  41. [49]

    Zijun Lin, Ke Xu, Chengfang Fang, Huadi Zheng, Aneez Ahmed Jaheezuddin, and Jie Shi. 2023. QUDA: Query-Limited Data-Free Model Extraction. In Proceedings of the ACM Asia Conference on Computer and Communications Security . ACM, Melbourne VIC Australia, 913–924. doi:10.1145/357...

  42. [50]

    Yupei Liu, Jinyuan Jia, Hongbin Liu, and Neil Zhenqiang Gong. 2022. StolenEncoder: Stealing Pre-Trained Encoders in Self-Supervised Learning. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security (Ccs ’22) . Association for Computing Machiner...

  43. [51]

    Yang Liu, Ji Luo, Yi Yang, Xuan Wang, Mehdi Gheisari, and Feng Luo. 2023. ShrewdAttack: Low Cost High Accuracy Model Extraction. Entropy 25, 2 (Feb. 2023), 282. doi:10.3390/e25020282

  44. [52]

    Yiyong Liu, Rui Wen, Michael Backes, and Yang Zhang. 2024. Efficient Data-Free Model Stealing with Label Diversity. arXiv preprint (2024). arXiv:2404.00108

  45. [53]

    Yugeng Liu, Rui Wen, Xinlei He, Ahmed Salem, Zhikun Zhang, Michael Backes, Emiliano De Cristofaro, Mario Fritz, and Yang Zhang. 2022. ML-doctor: Holistic Risk Assessment of Inference Attacks against Machine Learning Models. In Proceedings of the 31st USENIX Security Symposium ...

  46. [54]

    Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Zheng Xin Yong, Suhas Kotha, Yi Zeng, Weiyan Shi, Xianjun Yang, Reid Southen, Alexander Robey, Patrick Chao, Diyi Yang, Ruoxi Jia, Daniel Kang,...

  47. [55]

    Haoyu Ma, Tianlong Chen, Ting-Kuei Hu, Chenyu You, Xiaohui Xie, and Zhangyang Wang. 2021. Undistillable: Making a Nasty Teacher That CANNOT Teach Students. In Proceedings of the 9th International Conference on Learning Representations (ICLR)

  48. [56]

    Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards Deep Learning Models Resistant to Adversarial Attacks. In Proceedings of the 6th International Conference on Learning Representations (ICLR)

  49. [57]

    Dragan, and Moritz Hardt

    Smitha Milli, Ludwig Schmidt, Anca D. Dragan, and Moritz Hardt. 2019. Model Reconstruction from Model Explanations. In Proceedings of the Conference on Fairness, Accountability, and Transparency . ACM, Atlanta GA USA, 1–9. doi:10.1145/3287560.3287562

  50. [58]

    Takayuki Miura, Satoshi Hasegawa, and Toshiki Shibahara. 2021. MEGEX: Data-Free Model Extraction Attack against Gradient-Based Explainable AI. doi:10.48550/ARXIV.2107.08909

  51. [59]

    Netanyahu

    Itay Mosafi, Eli Omid David, and Nathan S. Netanyahu. 2019. Stealing Knowledge from Protected Deep Neural Networks Using Composite Unlabeled Data. In 2019 International Joint Conference on Neural Networks (IJCNN) . IEEE, Budapest, Hungary, 1–8. doi:10.1109/IJCNN.2019.8851798

  52. [60]

    Rina Okada, Zen Ishikura, Toshiki Shibahara, and Satoshi Hasegawa. 2020. Special-Purpose Model Extraction Attacks: Stealing Coarse Model with Fewer Queries. In 2020 IEEE 19th International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom) . I...

  53. [61]

    Daryna Oliynyk, Rudolf Mayer, and Andreas Rauber. 2023. I Know What You Trained Last Summer: A Survey on Stealing Machine Learning Models and Defences. ACM Comput. Surv. 55, 14s (Dec. 2023), 1–41. doi:10.1145/3595292

  54. [62]

    Tribhuvanesh Orekondy, Bernt Schiele, and Mario Fritz. 2019. Knockoff Nets: Stealing Functionality of Black-Box Models. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Long Beach, CA, USA, 4949–4958. doi:10.1109/CVPR.2019.00509

  55. [63]

    Soham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade, Shirish Shevade, and Vinod Ganapathy. 2020. ActiveThief: Model Extraction Using Active Learning and Unannotated Public Data. AAAI 34, 01 (April 2020), 865–872. doi:10.1609/aaai.v34i01.5432

  56. [64]

    Shevade, and Vinod Ganapathy

    Soham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade, Shirish K. Shevade, and Vinod Ganapathy. 2019. A Framework for the Extraction of Deep Neural Networks by Leveraging Public Data. arXiv preprint (2019). arXiv:1905.09165

  57. [65]

    David Pape, Sina Däubener, Thorsten Eisenhofer, Antonio Emanuele Cinà, and Lea Schönherr. 2023. On the Limitations of Model Stealing with Uncertainty Quantification Models. In Proceedings of the ICML 2023 Workshop on Adversarial Machine Learning Frontiers (AdvML-frontiers)

  58. [66]

    Nicolas Papernot, Patrick McDaniel, and Ian Goodfellow. 2016. Transferability in Machine Learning: From Phenomena to Black-Box Attacks Using Adversarial Samples. arXiv preprint (2016). arXiv:1605.07277

  59. [67]

    Berkay Celik, and Ananthram Swami

    Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z. Berkay Celik, and Ananthram Swami. 2017. Practical Black-Box Attacks against Machine Learning. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security . ACM, Abu Dhabi United A...

  60. [68]

    Berkay Celik, and Ananthram Swami

    Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z. Berkay Celik, and Ananthram Swami. 2016. The Limitations of Deep Learning in Adversarial Settings. In 2016 IEEE European Symposium on Security and Privacy (EuroS&P) . IEEE, Saarbrucken, 372–387. doi:10.1109/Eu...

  61. [69]

    Li Pengcheng, Jinfeng Yi, and Lijun Zhang. 2018. Query-Efficient Black-Box Attack by Active Learning. In 2018 IEEE International Conference on Data Mining (ICDM). IEEE, Singapore, 1200–1205. doi:10.1109/ICDM.2018.00159

  62. [70]

    Maura Pintor, Luca Demetrio, Angelo Sotgiu, Ambra Demontis, Nicholas Carlini, Battista Biggio, and Fabio Roli. 2022. Indicators of Attack Failure: Debugging and Improving Optimization of Adversarial Examples. In Proceedings of the 36th International Conference on Neural Inform...

  63. [71]

    Jonas Rauber, Wieland Brendel, and Matthias Bethge. 2017. Foolbox: A Python Toolbox to Benchmark the Robustness of Machine Learning Models. In Reliable Machine Learning in the Wild Workshop, 34th International Conference on Machine Learning (ICML)

  64. [72]

    Nicholas Roberts, Vinay Uday Prabhu, and Matthew McAteer. 2019. Model Weight Theft with Just Noise Inputs: The Curious Case of the Petulant Attacker. In Proceedings of the ICML 2019 Workshop on Security and Privacy of Machine Learning . Long Beach, CA, USA

  65. [73]

    Jonathan Rosenthal, Eric Enouen, Hung Viet Pham, and Lin Tan. 2023. DisGUIDE: Disagreement-Guided Data-Free Model Extraction. AAAI 37, 8 (June 2023), 9614–9622. doi:10.1609/aaai.v37i8.26150

  66. [75]

    Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra

    Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2020. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. International Journal of Computer Vision 128, 2 (2020), 336–359. doi:10.1007/ s...

  67. [77]

    Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014. Intriguing Properties of Neural Networks. In Proceedings of the 2nd International Conference on Learning Representations (ICLR)

  68. [78]

    Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 97) . PMLR, Long Beach, California, USA, ...

  69. [79]

    Te Juin Lester Tan and Reza Shokri. 2020. Bypassing Backdoor Detection Algorithms in Deep Learning. In 2020 IEEE European Symposium on Security and Privacy (EuroS&P) . IEEE, Genoa, Italy, 175–183. doi:10.1109/EuroSP48549.2020.00019 I Stolenly Swear That I Am Up to (No) Good: D...

  70. [80]

    Reiter, and Thomas Ristenpart

    Florian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter, and Thomas Ristenpart. 2016. Stealing Machine Learning Models via Prediction Apis. In Proceedings of the 25th USENIX Security Symposium . USENIX Association, Austin, TX, USA, 601–618

  71. [81]

    Walls, and Nicolas Papernot

    Jean-Baptiste Truong, Pratyush Maini, Robert J. Walls, and Nicolas Papernot. 2021. Data-Free Model Extraction. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Nashville, TN, USA, 4769–4778. doi:10.1109/CVPR46437.2021.00474

  72. [82]

    Yixu Wang, Jie Li, Hong Liu, Yan Wang, Yongjian Wu, Feiyue Huang, and Rongrong Ji. 2022. Black-Box Dissector: Towards Erasing-Based Hard-Label Model Stealing Attack. In Computer Vision – ECCV 2022 , Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and T...

  73. [83]

    Yixu Wang and Xianming Lin. 2022. Enhance Model Stealing Attack via Label Refining. In2022 7th International Conference on Intelligent Computing and Signal Processing (ICSP) . IEEE, Xi’an, China, 1040–1043. doi:10.1109/ICSP54964.2022.9778562

  74. [84]

    Yang Wang, Biao Qian, Haipeng Liu, Yong Rui, and Meng Wang. 2024. Unpacking the Gap Box Against Data-Free Knowledge Distillation. IEEE Trans. Pattern Anal. Mach. Intell. 46, 9 (Sept. 2024), 6280–6291. doi:10.1109/TPAMI.2024.3379505

  75. [85]

    Zi Wang. 2021. Zero-Shot Knowledge Distillation from a Decision-Based Black-Box Model. In Proceedings of the 38th International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 139) . PMLR, Virtual, 10675–10685

  76. [86]

    Baoyuan Wu, Hongrui Chen, Mingda Zhang, Zihao Zhu, Shaokui Wei, Danni Yuan, and Chao Shen. 2022. BackdoorBench: A Comprehensive Benchmark of Backdoor Learning. In Proceedings of the 36th International Conference on Neural Information Processing Systems (Nips ’22) . Curran Asso...

  77. [87]

    Yun Xiang, Zhuangzhi Chen, Zuohui Chen, Zebin Fang, Haiyang Hao, Jinyin Chen, Yi Liu, Zhefu Wu, Qi Xuan, and Xiaoniu Yang. 2020. Open DNN Box by Power Side-Channel Attack. IEEE Trans. Circuits Syst. II 67, 11 (Nov. 2020), 2717–2721. doi:10.1109/TCSII.2020.2973007

  78. [88]

    Yi Xie, Mengdie Huang, Xiaoyu Zhang, Changyu Dong, Willy Susilo, and Xiaofeng Chen. 2022. GAME: Generative-Based Adaptive Model Extraction Attack. In Computer Security – ESORICS 2022 , Vijayalakshmi Atluri, Roberto Di Pietro, Christian D. Jensen, and Weizhi Meng (Eds.). Vol. 1...

  79. [89]

    Anli Yan, Ruitao Hou, Xiaozhang Liu, Hongyang Yan, Teng Huang, and Xianmin Wang. 2022. Towards Explainable Model Extraction Attacks.Int J of Intelligent Sys 37, 11 (Nov. 2022), 9936–9956. doi:10.1002/int.23022

  80. [90]

    Anli Yan, Ruitao Hou, Hongyang Yan, and Xiaozhang Liu. 2023. Explanation-Based Data-Free Model Extraction Attacks. World Wide Web 26, 5 (Sept. 2023), 3081–3092. doi:10.1007/s11280-023-01150-6

  81. [91]

    Anli Yan, Teng Huang, Lishan Ke, Xiaozhang Liu, Qi Chen, and Changyu Dong. 2023. Explanation Leaks: Explanation-guided Model Extraction Attacks. Information Sciences 632 (June 2023), 269–284. doi:10.1016/j.ins.2023.03.020

  82. [92]

    Anli Yan, Hongyang Yan, Li Hu, Xiaozhang Liu, and Teng Huang. 2023. Holistic Implicit Factor Evaluation of Model Extraction Attacks. IEEE Trans. Dependable and Secure Comput. 20, 6 (Nov. 2023), 4678–4689. doi:10.1109/TDSC.2022.3231271

  83. [93]

    Fletcher, and Josep Torrellas

    Mengjia Yan, Christopher W. Fletcher, and Josep Torrellas. 2020. Cache Telepathy: Leveraging Shared Resource Attacks to Learn DNN Architectures. In Proceedings of the 29th USENIX Security Symposium (USENIX Security 20) . USENIX Association, Virtual, 1187–1204

  84. [94]

    Enneng Yang, Zhenyi Wang, Li Shen, Nan Yin, Tongliang Liu, Guibing Guo, Xingwei Wang, and Dacheng Tao. 2023. Continual Learning from a Stream of Apis. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 12 (2023), 10684743. doi:10.1109/TPAMI.2024.3460871

  85. [95]

    Panpan Yang, Qinglong Wu, and Xinming Zhang. 2023. Efficient Model Extraction by Data Set Stealing, Balancing, and Filtering. IEEE Internet Things J. 10, 24 (Dec. 2023), 22717–22725. doi:10.1109/JIOT.2023.3304345

  86. [96]

    Wenbin Yang, Xueluan Gong, Yanjiao Chen, Qian Wang, and Jianshuo Dong. 2024. SwiftTheft: A Time-Efficient Model Extraction Attack Framework Against Cloud-Based Deep Neural Networks. Chinese J. Elect. 33, 1 (Jan. 2024), 90–100. doi:10.23919/cje.2022.00.377

  87. [97]

    Honggang Yu, Kaichen Yang, Teng Zhang, Yun-Yun Tsai, Tsung-Yi Ho, and Yier Jin. 2020. CloudLeak: Large-Scale Deep Learning Models Stealing Through Adversarial Examples. In Proceedings 2020 Network and Distributed System Security Symposium . Internet Society, San Diego, CA. doi...

  88. [98]

    Xiaoyong Yuan, Leah Ding, Lan Zhang, Xiaolin Li, and Dapeng Oliver Wu. 2022. ES Attack: Model Stealing Against Deep Neural Networks Without Data Hurdles. IEEE Trans. Emerg. Top. Comput. Intell. 6, 5 (Oct. 2022), 1258–1270. doi:10.1109/TETCI.2022.3147508

  89. [99]

    Zhenrui Yue, Zhankui He, Huimin Zeng, and Julian McAuley. 2021. Black-Box Attacks on Sequential Recommenders via Data-Free Model Extraction. In Fifteenth ACM Conference on Recommender Systems . ACM, Amsterdam Netherlands, 44–54. doi:10.1145/3460231.3474275

  90. [100]

    Jie Zhang, Bo Li, Jianghe Xu, Shuang Wu, Shouhong Ding, Lei Zhang, and Chao Wu. 2022. Towards Efficient Data Free Blackbox Adversarial Attack. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, New Orleans, LA, USA, 15094–15104. doi:10.1109/ ...

  91. [101]

    Qifan Zhang, Junjie Shen, Mingtian Tan, Zhe Zhou, Zhou Li, Qi Alfred Chen, and Haipeng Zhang. 2022. Play the Imitation Game: Model Extraction Attack against Autonomous Driving Localization. In Proceedings of the 38th Annual Computer Security Applications Conference . ACM, Aust...

  92. [102]

    Xinyi Zhang, Chengfang Fang, and Jie Shi. 2021. Thief, Beware of What Get You There: Towards Understanding Model Extraction Attack. arXiv preprint (2021). arXiv:2104.05921

  93. [103]

    Shiqian Zhao, Kangjie Chen, Meng Hao, Jian Zhang, Guowen Xu, Hongwei Li, and Tianwei Zhang. 2023. Extracting Cloud-Based Model with Prior Knowledge. arXiv preprint (2023). arXiv:2306.04192

  94. [104]

    Mingyi Zhou, Jing Wu, Yipeng Liu, Shuaicheng Liu, and Ce Zhu. 2020. DaST: Data-Free Substitute Training for Adversarial Attacks. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Seattle, WA, USA, 231–240. doi:10.1109/CVPR42600.2020.00031 30...

  95. [105]

    CIFAR-10, SVHN WSL, WideResNet-28-2 ResNet-v2-50, ResNet-v2-200, N/A

  96. [106]

    GTSRB, MNIST N/A Custom

  97. [107]

    AR Face, BU3DFE, JAFFE, MMI, RaFD, CIFAR- 10 N/A VGG-16

  98. [108]

    CIFAR10, GTSRB, MNIST Custom Custom

  99. [109]

    GTSRB, MNIST Custom Custom

  100. [110]

    Caltech256, CUB-200-2011, Diabetic5, Indoor67ResNet-34, VGG-16 AlexNet, DenseNet-161, ResNet-18, ResNet-34, ResNet-50, VGG-16

  101. [111]

    Caltech256, CIFAR-10, CUB-200-2011, Diabetic5, Indoor67 ResNet-34 ResNet-34

  102. [112]

    CIFAR-10, FMNIST, MNIST N/A, some ResNet Custom, VGG-19

  103. [113]

    CIFAR-10, FMNIST, GTSRB, MNIST Custom Custom

  104. [114]

    CIFAR-10 Custom VGG-16

  105. [115]

    CIFAR-10, KMNIST, MNIST, SVHN LeNet, ResNet-18, ResNet-34 LeNet, ResNet-18, ResNet-34

  106. [116]

    CIFAR-10, MNIST Custom, MLogReg, ResNet-18, VGG-11 Custom, MLogReg, ResNet-18, VGG-11

  107. [117]

    CIFAR-10, FMNIST, GTSRB, SVHN LeNet, ResNet-20 WideResNet-22

  108. [118]

    FMNIST, KMNIST, MNIST, notMNIST Custom Custom

  109. [119]

    CIFAR-10, FMNIST, 10 Monkey Species AlexNet, LeNet, ResNet-18, VGG-16 half-AlexNet, half-LeNet, ResNet-18, VGG-16

  110. [120]

    GTSRB, VGG Flower AlexNet, ResNet-50, VGG-19, VGG-Face ResNet-50, VGG-19, VGG-Face

  111. [121]

    CIFAR-10, GTSRB, MNIST Custom Custom, ResNet-18

  112. [124]

    CIFAR-10, CIFAR-100 AlexNet, ResNet-18, ResNet-34 half-AlexNet, ResNet-18

  113. [125]

    Caltech256, ImageNet, FMNIST AlexNet, Custom, LeNet, ResNet-34, ResNet-50 AlexNet, Custom, LeNet, ResNet-18, ResNet-34, ResNet-50

  114. [126]

    Caltech256, CIFAR-10, CUB-200-2011, SVHN ResNet-34 N/A

  115. [127]

    Caltech256, CIFAR-10, CUB-200-2011, SVHN ResNet-34 some DenseNet, ResNet-18, ResNet-34, ResNet- 50, VGG-16

  116. [128]

    CIFAR-10, CIFAR-100, FMNIST, MNIST DenseNet-161, ResNet-50, VGG-19 Custom

  117. [129]

    BelgiumTSC, MNIST AlexNet, LeNet half-AlexNet, half-LeNet, ResNet-18, VGG-16

  118. [130]

    CIFAR-10, FMNIST, GTSRB, ImageNette, MNISTLeNet, ResNet-34, VGG-16 LeNet, ResNet-34, VGG-16

  119. [131]

    CIFAR-10, ImageNet, MNIST Custom, Inception-V3, LeNet, ResNet-18, ResNet- 152 Custom, Inception-V3, LeNet, ResNet-18, ResNet- 152

  120. [132]

    CelebA, FMNIST, STL-10, UTKFace AlexNet, Custom, ResNet-18, VGG-19, Xception AlexNet, Custom, ResNet-18, VGG-19, Xception

  121. [133]

    CIFAR-10, CIFAR-100 ResNet-18, ResNet-34 ResNet-18

  122. [134]

    CIFAR-10, CIFAR-100, FMNIST, MNIST DenseNet-121, DenseNet-161, DenseNet-169, DenseNet-201, Inception-V1, Inception-V2, Inception-V3, ResNet-18, ResNet-34, ResNet-50, ResNet-101, ResNet-152, VGG-11, VGG-13, VGG-16, VGG-19 DenseNet-121, DenseNet-161, DenseNet-169, DenseNet-201, ...

  123. [135]

    CIFAR-10, CIFAR-100, FMNIST, MNIST, SVHN, TinyImageNet N/A, Custom, ResNet-34 N/A, Custom, ResNet-18

  124. [136]

    CIFAR-10, SVHN ResNet-34 ResNet-18

  125. [137]

    CIFAR-10, CIFAR-100, FMNIST, MNIST DenseNet-161, ResNet-50, VGG-19 N/A

  126. [138]

    CIFAR-10, FMNIST, MNIST AlexNet, LeNet half-AlexNet, half-LeNet

  127. [139]

    CIFAR-10, FMNIST AlexNet, LeNet, ResNet-18, VGG-11 AlexNet, ResNet-18, VGG-11

  128. [140]

    FMNIST, Intel-Image N/A SqueezeNet

  129. [141]

    CIFAR-10, SVHN N/A, ResNet-152 Custom, Inception-V3, ResNet-152

  130. [142]

    CelebA, CIFAR-10, SVHN ResNet-34 ResNet-18

  131. [143]

    CIFAR-10, Food-101 AlexNet, ResNet-50 half-AlexNet, ResNet-18

  132. [144]

    CIFAR-10, MNIST, SVHN Custom, ResNet-34 LeNet, VGG-16

  133. [145]

    CIFAR-10, Flower-17, GTSRB, STL-10 VGG-13 ResNet-50

  134. [146]

    CIFAR-10, MNIST Custom, some ResNet Custom, ResNet-18

  135. [147]

    Caltech256, CIFAR-10, CIFAR-100, CUB-200- 2011 ResNet-34 AlexNet, DenseNet-121, EfficientNet-B2, MobileNet-V3, ResNet-34

  136. [148]

    CIFAR-10, CIFAR-100, FMNIST, GTSRB, MNIST, SVHN ResNet-34 ResNet-18

  137. [149]

    CIFAR-10, FMNIST, MNIST, SVHN ResNet-34, VGG-16 VGG-11

  138. [150]

    CIFAR-10, GTSRB, VGG Flower Custom Custom

  139. [151]

    CIFAR-10, FMNIST, GTSRB, SVHN some MobileNet, some ResNet, some VGG some MobileNet, some ResNet, some VGG

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.