{"id":"69a67e78-d1b0-4e39-9fb7-827ce6fe76f9","arxiv_id":"2508.21654","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A systematization of 47 model stealing papers that introduces a threat model, a comparison framework, and evaluation best practices for substitute-model attacks.","lead":"This paper reviews dozens of model stealing attacks on image classifiers and proposes a shared threat model and comparison framework so different attacks can be fairly compared. It also lists best practices and open research questions for how model stealing research should be evaluated.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'small fraction comparable' result rests on an unreleased, subjective re-coding of 47 papers; Section 3.4 admits the categories have reasonable alternatives, so Figure 1's segment counts need a sensitivity check before the central claim is secure.","rationale":"The reader identified the reconstruction of prior threat models as the weakest assumption; this is exactly where the paper's central contribution is most exposed. The paper is honest about the ambiguity, but honesty does not remove the load: every downstream artifact (Figure 1, the comparability statistics, the baseline-selection rule R2.1) is a deterministic function of the contested codings. My proposed check would settle whether the headline conclusion survives definitional and annotator variation. Because the reader already set CONDITIONAL, my stress-test does not move the verdict; it adds a specific condition for lifting the condition: release the coding and demonstrate robustness. I found no independent reason to reject the framework; the threat model and best-practice sections are useful and the paper's self-identified limitations are appropriately scoped. The only other candidate concern — the transferability definition in Equation (3) appears to have a typo (the right-hand side reads f(x'_i) != f(x'_i)) — is worth fixing but is not load-bearing for the central comparability claim. Therefore UNCHANGED.","tokens_in":30754,"tokens_out":7741,"duration_ms":91283,"concrete_test":"Release the per-paper coding with raw evidence, and have two independent annotators re-code the 47 papers using the Section 3 taxonomy; report Cohen's kappa per dimension. Then recompute Figure 1 segment sizes under the alternative definitions in Section 3.4.1 (exact-training-data vs same-distribution 'original'; generator-trained-on-nPD as data-free vs nPD; and all grey/uncertain placements). If the largest segment grows by more than ~3 papers relative to the reported 11, or if the number of empty/one-paper segments changes by more than 3, the 'small fraction comparable' claim is not robust to coding choices.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim — that only a small fraction of prior model-stealing attacks are comparable — is computed from Table 1/Figure 1, whose threat-model assignments are 'primarily derived by us' from experimental setups (Section 3), because few papers defined threat models. The paper itself lists unresolved ambiguities in Section 3.4.1: e.g., 'original data' can mean exact training data or same-distribution data; data-free can exclude real data only from substitute training or from the whole attack; nPD/PD boundaries are decided by a 'more than 50%' rule. These are judgment calls with 'reasonable arguments for other solutions,' and no coding artifact is released. Since the framework's recommendation to compare only within a segment inherits every coding decision, an error in even a handful of assignments can change the 'largest segment <= 11 papers' statistic and the quarter-empty segments observation. This is a load-bearing reproducibility/correctness risk, not a stylistic issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that model stealing attack evaluations are not standardised and proposes the first comprehensive threat model and comparison framework for substitute-training attacks against image classifiers. The authors define attacker knowledge (data type, output type, architecture knowledge, pre-training), capabilities (query budget), and goals (accuracy/fidelity/transferability); classify 47 prior papers into Table 1 and a 24-segment diagram; and use the resulting coverage to claim that only a small fraction of prior attacks are comparable, that a quarter of the segments are empty, and that several configurations are thinly studied. They further analyse dataset/architecture/query usage in prior experiments, derive best practices (R1–R3), and list open research questions.","tokens_in":30993,"tokens_out":8452,"duration_ms":97191,"significance":"If the classification is reliable, the framework would be a genuinely useful community resource: it provides a common vocabulary, a structured way to select comparable baselines, and concrete reporting recommendations. The paper is unusually transparent about its coding ambiguities and about missing information in prior work, and it offers falsifiable observations about segment coverage and experimental-setup frequencies. The open questions are thoughtful and practical. The main risk is that the headline quantitative claims rest on a subjective re-coding that is neither released nor stress-tested; as currently presented, the evidence is not yet sufficient to support the central 'only a small fraction of prior attacks can be compared' claim at the strength asserted.","major_comments":[{"comment":"The central quantitative claim ('only a small fraction of prior attacks can be compared', 'largest segment at most 11 papers', 'a quarter of segments empty') is computed from assignments that the authors themselves describe as derived from experimental setups with several reasonable alternative resolutions. Section 3.4.1 explicitly lists unresolved ambiguities for original vs problem-domain data, problem-domain vs non-problem-domain data, and the data-free definition, and Section 3.4 states that 'there are also reasonable arguments for other solutions.' Because most prior papers did not define threat models, this is a subjective re-coding. No coding artifact or per-paper annotation is released, so the reader cannot audit the assignments. Please provide a sensitivity analysis under the alternative coding rules (varying the disputed thresholds and the handling of unknown architecture/outpu","section":"§3.4.1, Table 1, Figure 1"},{"comment":"Table 1's Metrics column appears to list 'AFT' for every analysed paper. If correct, the table contradicts the text in §3.4.3 ('no metric was reported in all papers') and in R2.2 ('none of the metrics was reported for every previous attack'), and it removes the evidence that goal/metric reporting is insufficient. Please correct the column to reflect which metrics each paper actually reports, or revise the textual claims. As printed, this is an internal inconsistency in evidence used to motivate R2.2.","section":"§3.3, Table 1 (Metrics column)"},{"comment":"The comparability framework partitions attacks only by attacker knowledge (data/output/architecture), while attacker capabilities (query budget) and goals are excluded from the segment definition. §4.1 argues that capabilities and goals align with evaluation and can be 'easily adapted' by reporting the right numbers, but this conflates comparability of reported numbers with comparability of attack difficulty. Two attacks in the same knowledge segment but with query budgets of 100 and 10^6 are not directly comparable without an agreed query-budget or efficiency curve. Please clarify under which conditions 'same segment' is sufficient (e.g., same or matched query budget, same test data and metric definitions) or refine the partition.","section":"§4.2, Figure 1"},{"comment":"The paper does not report how the 47-paper corpus was constructed: no search venues, inclusion/exclusion criteria, time window, or screening process are given. The coverage statistics and the 'first comprehensive' claim are therefore claims about an unstated sample. Please specify the corpus construction methodology or, at minimum, provide the full list with explicit selection rules; otherwise the generalisation from 47 papers to 'the field' is not reproducible.","section":"§3, corpus construction"}],"minor_comments":[{"comment":"The transferability definition contains a typo: the consequent should be f(x'_i) != f(x_i), not f(x'_i) != f(x'_i). As written, the right-hand side is a tautological inequality that is always false, making the indicator identically zero.","section":"Eq. (3)"},{"comment":"'more than 1,000,0000 queries' should read 'more than 1,000,000 queries'.","section":"§3.2"},{"comment":"The shading terminology is inconsistent: §4.2 says empty segments are filled in dark grey, while §4.3(ii) refers to 'the grey segments' as having no prior work. Please clarify the legend and distinguish unknown-assignment grey from empty dark-grey.","section":"Figure 1 and §4.3"},{"comment":"Several papers appear in multiple segments (e.g., [14](1)/(2), [97], [103]). The text explains this, but the counting unit for the 'largest segment' statistic should be stated explicitly: segment entries are paper-segment incidences, not unique papers, and the claim 'at most 11 papers' needs to define how duplicates and unknown-assignment entries are counted.","section":"Table 1 / Figure 1 counting unit"},{"comment":"The contribution bullet says 'a quarter of configurations have only been studied in one or two works', while §4.3 says a third have at most one publication and more than half have at most three publications. These numbers overlap but are not expressed consistently; please reconcile.","section":"§1 vs §4.3"},{"comment":"Several forward references to Section 7 appear before that section is introduced (e.g., 'as discussed in Section 7.1'). Consider renumbering or adding cross-reference placeholders.","section":"§6.1, §6.2"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on the authors' prior survey [61] for goal definitions and terminology; the new contribution should clearly delineate what is novel beyond that survey. The repeated 'first' claims in the abstract, introduction, and conclusion should be moderated unless the corpus methodology and coding robustness are addressed. The paper is a good fit for a security venue, and the topic is timely; the major revision should focus on making the empirical classification auditable and on tightening the comparability conditions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is the first real attempt at a threat model and comparison framework for substitute-training model stealing attacks, and it does a genuine service. The 24-cell taxonomy, the categorization of 47 papers, the analysis of experimental setups, and the list of best practices and open questions are all worth having. The central observation—that only a small fraction of prior attacks are comparable under a consistent threat model—is important if it holds.\n\nThe paper is transparent about the hard part. Section 3 states that the threat models in Table 1 are primarily derived by the authors from the experimental setups, since only a minority of papers defined them. Section 3.4 lists the ambiguities and admits that reasonable alternative choices exist for 'original' vs 'problem-domain' data, for 'data-free', and for query-budget boundaries. That honesty is to the authors' credit.\n\nBut it also means the main claim is built on a manual coding exercise that is not publicly available. No coding artifact is released, and there is no sensitivity analysis to show how Figure 1 changes under the alternative definitions the paper itself enumerates. The stress-test note is right: a handful of reclassifications could change the 'largest segment ≤ 11 papers' observation and the 'quarter of segments empty' claim. This is not a stylistic concern; it is a reproducibility risk on the paper's central empirical result.\n\nWhat holds up: the framework's logic is sound, the recommendations (compare only within a threat-model segment, report accuracy and fidelity, run ablations, specify pre-trained model usage) are sensible and follow from the analysis. The statistics on datasets and architectures, while partly inflated by papers with many configurations (which the paper acknowledges), are useful. The transferability measurement table makes the lack of standardization concrete.\n\nWho this is for: anyone working on model stealing attacks or defenses, as a way to position new work and pick baselines. It deserves a serious referee. The referee should ask for three things before acceptance: the per-paper coding data, a sensitivity analysis that re-codes ambiguous cases under the alternative definitions the paper itself lists, and a check of whether the headline comparative statistics survive. If the authors can do that, this becomes a standard reference. As it stands, I would not rely on the specific counts until that work is done.","headline":"A useful first systematization of model stealing attacks; the headline comparability statistic rests on subjective re-coding that needs to be released and stress-tested.","tokens_in":31473,"tokens_out":2576,"would_cite":true,"duration_ms":29639,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that model stealing attacks on image classifiers are not standardised and provides a threat model and comparison framework to make attacks with the same attacker knowledge directly comparable.","keywords":["model stealing","model extraction","threat model","evaluation methodology","substitute model training","image classification","attack comparability","query budget"],"falsifier":"Re-derive the paper's Table 1 classification of the 47 attacks from their experimental sections with two independent annotators. If the annotators disagree on the attacker-data category or the output type for more than about a fifth of the papers, the segment counts and the claim that only a small fraction of attacks are comparable would not be robust. A simpler check: verify whether the largest segment in Figure 1 still contains at most 11 papers when ambiguous papers are excluded instead of placed into every possible segment.","tokens_in":30659,"feed_emoji":"🕵️","tokens_out":7086,"duration_ms":75676,"temperature":0.7,"pith_summary":"Model stealing attacks let an adversary clone a machine-learning model by querying it and training a substitute. This paper argues that the field cannot measure progress because attack evaluations are not standardised: papers differ in what the attacker knows, what outputs the target model returns, how many queries are allowed, and which metrics are reported. It builds a detailed threat model for substitute-training attacks on image classifiers and a comparison framework that divides attacks into 24 attacker-knowledge segments, so that only attacks in the same segment are fairly comparable. Applying this to 47 prior works, it finds that only a small fraction of attacks can actually be compared with each other. If the field adopts the proposed best practices, new attacks can be placed in a specific segment and evaluated against the right baselines, making state-of-the-art claims testable.","feed_headline":"Only a fraction of model stealing attacks are comparable","feed_subtitle":"A threat model for substitute-training attacks maps 47 prior works and shows why most differ too much to compare.","key_machinery":"The central object is the threat model and the comparison framework built on it. The threat model decomposes the attacker into knowledge (data type, output type, architecture match, and pre-trained-model use), capabilities (query budget), and goals (accuracy, fidelity, transferability). The comparison framework turns the knowledge axis into a 24-segment diagram: one split for same versus different architecture, and within each side, circular sectors for data-free, non-problem-domain, problem-domain, and original data, each divided by labels, probabilities, or explanations. The rule is that only attacks in the same segment can be fairly compared, and query counts, ideally normalised per targe","core_discovery":"The paper's central claim is that comparability of model stealing attacks is only meaningful when attacks share the same attacker knowledge. It defines that knowledge along three axes: what data the attacker has (original data, problem-domain data, non-problem-domain data, or none), what the target model returns (labels, confidence scores, or explanations and gradients), and whether the substitute architecture is the same as or different from the target's. It maps 47 substitute-training attacks on image classifiers onto these axes, producing 24 possible knowledge segments. It then shows that the largest segment contains at most 11 papers, a quarter of segments have no prior work, and a third","pith_inferences":["If the framework were adopted as a community standard, the 24 segments could become the basis of a public leaderboard where ranking happens strictly within a segment, giving the field a concrete artefact for measuring progress.","The paper leaves pre-trained model knowledge out of its Table 1 classification but argues it matters; adding it as a fourth knowledge axis would split the 24 segments further and likely reduce comparability even more.","The paper's choice to count held-out parts of the target's original dataset as 'original data' is one of several plausible definitions; if the field adopts a stricter definition, some attacks would move to the problem-domain segment and change which baselines are fair.","The framework's logic transfers to other domains, but the paper only gestures at that transfer; adapting the comparison to tasks like text or graph learning would require redefining what counts as task accuracy and fidelity."],"forward_implications":["New attacks should report accuracy and fidelity on the same test data as baselines drawn from the same threat-model segment; otherwise they are not comparable.","Attack efficiency should be reported not only as total query count but also relative to the target model's training-set size, which reveals that many data-free attacks need over 100 queries per training sample.","Authors should state whether the target, substitute, and any auxiliary models are pre-trained, because pre-training materially changes the attacker's knowledge and is often omitted.","Transferability scores from different papers cannot be compared until the adversarial perturbation method and its strength are standardised.","Attacks with very low query budgets, under 1,000 and even under 100 queries, are understudied despite being the most practical real-world setting."],"supporting_citations":[{"why":"Supplies the original demonstration that a model offered as a prediction API can be cloned by querying it and training a substitute.","marker":"[80]"},{"why":"Supplies the prior survey of model stealing and defences, including the attacker-goal categories and efficiency score that this paper extends.","marker":"[61]"},{"why":"Defines functionally equivalent, fidelity, and task-accuracy extraction, which becomes the basis for the accuracy, fidelity, and transferability metrics.","marker":"[36]"},{"why":"Provides the knowledge-capabilities-goals vocabulary used to structure the threat model.","marker":"[8]"},{"why":"Provides the model for evaluation guidelines and adaptive-attack testing that the paper adapts into best practices for model stealing.","marker":"[9]"},{"why":"Provides real-world data on API query limits and deployment constraints used to question the practicality of million-query attacks.","marker":"[26]"},{"why":"Shows that fine-tuned imitation of a large proprietary model does not fully reproduce it, supporting the paper's concern about generalisation to complex models.","marker":"[29]"},{"why":"Shows that probability outputs and pre-trained substitutes improve extraction, supporting the recommendation to report output type and pre-training.","marker":"[3]"},{"why":"Shows that gradient-based explanations give better substitute models than probabilities, used in the discussion of output knowledge.","marker":"[58]"}],"fun_headline_variants":["Model stealing attacks: why most can't be compared","New threat model maps 47 model stealing attacks","Attacker knowledge key to comparing theft attacks","Most model stealing attacks differ too much to compare","Standardizing model stealing attack evaluations"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The comparison statistics assume that a paper's threat model can be reliably reconstructed from its experimental setup, even when the paper never stated a threat model.","fun_headline_variants_meta":{"raw":{"variants":["Model stealing attacks: why most can't be compared","New threat model maps 47 model stealing attacks","Attacker knowledge key to comparing theft attacks","Most model stealing attacks differ too much to compare","Standardizing model stealing attack evaluations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1135,"prompt_tokens":717,"completion_tokens":418,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":461,"completion_tokens_details":{"reasoning_tokens":350}},"tokens_in":461,"tokens_out":418,"duration_ms":4745,"temperature":1.0,"reasoning_tokens":350,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:03:59.581427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-derive the paper's Table 1 classification of the 47 attacks from their experimental sections with two independent annotators. If the annotators disagree on the attacker-data category or the output type for more than about a fifth of the papers, the segment counts and the claim that only a small fraction of attacks are comparable would not be robust. A simpler check: verify whether the largest segment in Figure 1 still contains at most 11 papers when ambiguous papers are excluded instead of placed into every possible segment.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that probability outputs and pre-trained substitutes improve extraction, supporting the recommendation to report output type and pre-training."},{"cited_title":"MEGEX: Data-Free Model Extraction Attack against Gradient-Based Explainable AI","cited_arxiv_id":"2107.08909","evidence_quote":"Shows that gradient-based explanations give better substitute models than probabilities, used in the discussion of output knowledge."}],"review_version":1}