{"id":"b91d92a7-94ae-4668-af17-2c5bc3dcae64","arxiv_id":"2508.15031","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper classifies model extraction attacks and defenses into attack, defense, and computing environment categories and surveys their current state.","lead":"This preprint surveys model extraction attacks and defenses, organizing the literature into a taxonomy of attack mechanisms, defense strategies, and computing environments. It is a reference-style review with an online repository, not a new attack or defense result.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Taxonomy categories overlap: query-based and data-driven attacks are not mutually exclusive, undermining the systematic-taxonomy claim.","rationale":"The reader's weakest assumption focused on the absence of a systematic literature-selection protocol, which is a valid concern about coverage. My concern is different: even if the coverage were perfect, the proposed taxonomy—the paper's central contribution—does not provide a mutually exclusive classification. The definitions in Sec. 4.1.1 and Sec. 4.2 overlap, and the same landmark papers (Knockoff Nets, ActiveThief, MAZE) appear under multiple categories. This is an internal-consistency issue, not a disagreement with outside consensus. It directly affects the 'first unified and comprehensive framework' claim: a framework whose categories overlap cannot serve as a systematic map of the field, and readers relying on Fig. 3 may misplace attacks. The survey is nevertheless useful as a structured literature review, so I do not recommend rejection. The conditional verdict already captures the need for verification; my concrete test would settle whether the taxonomy is actually coherent. Hence UNCHANGED.","tokens_in":41758,"tokens_out":3469,"duration_ms":40122,"concrete_test":"Construct the full list of papers cited in Fig. 3 under Query-based Attacks and Data-driven Attacks (including subsections). For each paper, check whether it is cited in both categories. If any paper appears in two sibling categories (e.g., [155], [158], [96]), the taxonomy is not a partition. A stronger test: ask two annotators to assign each of the roughly 60 papers in Fig. 3 to exactly one attack category using only the definitions in Sec. 4, and measure inter-annotator agreement (Cohen's kappa). If kappa < 0.8, the classification is not reproducible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is a 'novel and comprehensive taxonomy' (Sec. 1) that 'systematically classifies MEAs based on attack mechanisms' (Sec. 3). This taxonomy is not internally consistent: its attack categories are not mutually exclusive. Sec. 4.1.1 defines query-based attacks as techniques that 'rely on systematically querying a target model'; Sec. 4.2 defines data-driven attacks as 'characterized by their use of data to query and replicate target models.' Since any black-box extraction attack must query the target, 'query-based' is not a discriminative category. Concretely, Knockoff Nets [155] is cited as a canonical substitute-model-training attack in Sec. 4.1.2 and again as the exemplar of problem-domain data-driven attacks in Sec. 4.2.1; MAZE [96] appears under Data-free Attacks in Sec. 4.2.3 and also under gradient-estimation in Sec. 4.4.3; ActiveThief [158] is used in both 4.1.2 and 4.2.1. Thus the same papers populate multiple sibling categories. The taxonomy also mixes orthogonal axes (information channel vs. data availability vs. modality). This undermines the claim of offering a 'unified and comprehensive framework' and, more importantly, makes Fig. 3 a less reliable navigation aid than claimed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a survey of model extraction attacks (MEAs) and defenses, organized around a proposed taxonomy that classifies attacks by mechanism, defenses by strategy, and research by computing environment (cloud, edge, federated learning). It reviews attack families (query-based, data-driven, side-channel, gradient-based, and modality-specific), defense families (detection, ownership verification, prevention, and integrated defenses), evaluation metrics, real-world application scenarios (finance, healthcare, autonomous vehicles, cybersecurity), and future research directions. The paper claims to provide 'the first unified and comprehensive framework' for MEAs and maintains a continuously updated online repository of related literature. It does not present new empirical measurements; its contribution is a synthesis and systematization of existing work.","tokens_in":42029,"tokens_out":4743,"duration_ms":55426,"significance":"If the taxonomy were internally consistent and the coverage demonstrably systematic, this survey would be a valuable reference for researchers, practitioners, and policymakers. The paper has real strengths: it covers recent developments in LLM, GNN, and edge/federated extraction; it discusses both attacks and defenses across multiple computing paradigms; it includes a structured evaluation-metrics section; and the online repository is a useful community resource. The breadth of cited work is impressive. However, the paper's central value proposition—the unified taxonomy—is currently undermined by overlapping and inconsistently applied category definitions, and the 'comprehensive' claim is not supported by a transparent selection methodology. These issues are load-bearing for a survey whose main purpose is to organize the field.","major_comments":[{"comment":"The five attack categories are not mutually exclusive, which undermines the central claim of a systematic taxonomy. Query-based attacks (Sec 4.1.1) are defined by 'systematically querying a target model,' while data-driven attacks (Sec 4.2) are defined as 'use of data to query and replicate target models.' Since every black-box attack must query, the first two categories do not partition the space. Concretely, Knockoff Nets [155] appears under substitute model training (Sec 4.1.2) and again under problem-domain data-driven attacks (Sec 4.2.1); MAZE [96] appears under data-free attacks (Sec 4.2.3) and again under gradient estimation (Sec 4.4.3); ActiveThief [158] appears in both Sec 4.1.2 and Sec 4.2.1. Additionally, 'attacks on other data modalities' (Sec 4.5) is not an attack mechanism but a data type. The taxonomy mixes orthogonal axes (information channel, data availability, data moda","section":"Section 3, Fig. 3; Sections 4.1.1, 4.2, 4.4.3"},{"comment":"The paper claims to offer 'the first unified and comprehensive framework' and to provide a 'comprehensive and up-to-date overview,' but no systematic literature search protocol is presented. There is no description of databases searched, keywords, inclusion/exclusion criteria, time window, or screening process. The reader cannot assess whether the papers selected for Figure 3 are representative or whether important works are omitted. This is a load-bearing issue for a survey whose central claim is comprehensiveness. Please add a methodology subsection and discuss limitations of coverage.","section":"Sections 1 and 3"},{"comment":"The same defense name 'ModelGuard' is characterized as an information-theoretic monitoring method in Sec 5.1.1 and as an output-perturbation method in Sec 5.2.2, citing [198] in the latter. This internal inconsistency suggests that the taxonomic criteria are applied differently across sections. Combined with the overlapping attack categories, this weakens the paper's stated goal of providing a consistent classification. Please unify definitions and citation contexts.","section":"Sections 5.1.1 and 5.2.2"}],"minor_comments":[{"comment":"The sentence 'ESAs extend beyond model parameters to extract training hyperparameters, an attack method that targets the fundamental training configuration of machine learning models.' appears twice in the same paragraph. Please remove the duplicate.","section":"Section 4.1.3"},{"comment":"The text contains a missing citation placeholder: '? ] proposed a method to prevent weight stealing by obfuscating the network structure.' This reference must be supplied before publication.","section":"Section 5.3.2"},{"comment":"Typo: 'Caculate the proportion' should be 'Calculate the proportion.'","section":"Section 7.2"},{"comment":"In the autonomous vehicles paragraph, the sentence ends abruptly with 'often involving edge computing [136].' The thought appears incomplete; please revise.","section":"Section 8"}],"recommendation":"major_revision","confidential_remarks":"The 'first unified and comprehensive framework' claim is supported by a self-authored comparison table (Table 1) and by frequent citations to the corresponding author's own methods (e.g., MISLEADER in Sec 6.1, plus related works such as Atom and CEGA). This is not improper, but the novelty claim would be more persuasive if benchmarked against independent taxonomies and if the survey acknowledged potential self-citation bias. The duplicated sentence and missing reference placeholder suggest the manuscript needs a careful proofreading pass. I recommend major revision rather than rejection because the survey has clear utility once the taxonomy is made coherent and the methodology is documented."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The survey gives you a good map of a messy field: attack families, defenses, evaluation metrics, and the cloud/edge/federated dimensions are all covered, and the GitHub repository is a practical resource. If you want a quick orientation to model extraction, this is one of the better entry points right now.\n\nThe real value is organizational. The paper collects recent work across modalities (text, vision, graph), distinguishes computing environments, and structures defenses into detection, ownership verification, prevention, and integrated strategies. The evaluation metrics section is a handy reference. That is genuine contribution.\n\nThe soft spots are real and mostly in the taxonomy. Attack categories overlap in the way the stress test flags: query-based and data-driven are not mutually exclusive, and the same papers appear under siblings (e.g., Knockoff Nets under substitute model training and problem-domain attacks; MAZE under data-free and gradient estimation). That undermines the claim of a systematic, unified framework. The 'first' claim is also shaky, since prior surveys exist, and the comparison table is self-assessed. There is no reproducible search protocol or inclusion criteria. The duplicated sentence in Section 4.1.3 and the missing placeholder in 5.3.2 indicate polishing needed.\n\nThese are fixable, but they are not cosmetic. The taxonomy needs either clearer separation along one axis or an honest statement that it is a multi-axis organization rather than a partition. The comprehensiveness claim should be softened or supported with a methodology. The survey is honest in places—it describes limitations of defenses and does not invent experimental results—but the framing overreaches.\n\nSerious referee? Yes. A survey covering this breadth with an online repository is worth the referee time, but the verdict should be 'major revision' until the taxonomy is sharpened and the contribution claims are made defensible. I would bring it to a reading group as an example of both useful synthesis and how overclaimed structure can mislead.","headline":"A useful, current survey of model extraction, but the 'novel taxonomy' claims are overstated and the attack categories overlap; worth publishing after revision.","tokens_in":42496,"tokens_out":940,"would_cite":false,"duration_ms":15566,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a unified taxonomy that classifies model extraction attacks by mechanism, defenses by strategy, and both across cloud, edge, and federated environments, claiming to be the first framework to combine all three dimensions.","keywords":["model extraction","model stealing","MLaaS security","taxonomy","adversarial machine learning","defense strategies","side-channel attacks","federated learning"],"falsifier":"A systematic literature search using the same keywords with explicit inclusion criteria could count published extraction attacks and defenses that do not fit any leaf of Figure 3; if a material number of well-established methods require a new branch on the taxonomy, the comprehensiveness claim fails. A simpler check: the paper claims to be the first to combine attack mechanisms, defense strategies, and computing environments, so locating any earlier survey that already integrates all three axes would also falsify the novelty claim.","tokens_in":41660,"feed_emoji":"🔐","tokens_out":6690,"duration_ms":74604,"temperature":0.7,"pith_summary":"This survey tries to establish that the scattered literature on model extraction attacks and defenses fits into one unified map. It classifies attacks by the channel through which information leaks—query-based, data-driven, side-channel, gradient-based, or modality-specific—and defenses by strategy: detection, ownership verification, prevention, and integrated protection. It adds a third axis for the computing environment, arguing that cloud, edge, and federated settings each shape both how models get stolen and how they can realistically be protected. If the map holds, researchers gain a shared vocabulary, practitioners gain a way to pick defenses by threat type, and open problems such as the utility–security trade-off become visible as coherent research targets.","feed_headline":"One taxonomy maps model extraction attacks and defenses","feed_subtitle":"Survey groups attacks by mechanism, defenses by strategy, and threats by cloud, edge, and federated settings to guide defenders.","key_machinery":"The load-bearing object is the three-axis taxonomy shown in Figure 3. Its attack axis is organized by the information channel through which model knowledge leaks, progressing from explicit query–response probing to implicit side-channel leakage. Its defense axis is organized by the timing and logic of protection: detecting attacks during querying, verifying ownership after the fact, preventing extraction up front, or combining multiple measures. Its environment axis separates cloud, edge, and federated deployment, because each context changes what attackers can access and what defenders can afford. The taxonomy does the argument's work by making every surveyed paper comparable along the same","core_discovery":"The paper's central claim is that model extraction is not an unstructured collection of tricks but a field with a discoverable structure. It proposes a novel taxonomy with three axes: attack mechanism, defense approach, and computing environment. On the attack side it distinguishes query-based attacks, data-driven attacks, side-channel attacks, gradient-based attacks, and attacks on specific data modalities (text, vision, graph). On the defense side it distinguishes attack detection, ownership verification, attack prevention, and integrated or compositional defenses. On the environment axis it separates cloud computing, edge computing, and federated learning. The paper further claims to be t","pith_inferences":["Because the paper treats attack channels as separate but notes that edge attacks need to combine side-channel and query information, the taxonomy points toward hybrid cross-channel attacks as a likely blind spot in current defenses.","The survey does not state a systematic search protocol, inclusion criteria, or coverage window, so its comprehensiveness should be read as a synthesis of the authors' selected literature rather than an exhaustive census; a formal meta-analysis with explicit criteria could test whether the taxonomy's leaves cover the whole field.","The same three-axis structure could be extended to emerging model families—multimodal vision-language systems, diffusion models, and agentic LLMs—that the paper mentions only in passing; applying the taxonomy to those families would be a direct test of its generality.","If the taxonomy were adopted as the organizing scheme for the authors' continuously updated online repository, category imbalance would become visible at a glance, showing which attack–defense combinations are still underexplored."],"forward_implications":["A new attack or defense can be positioned within the taxonomy, making it straightforward to compare against prior work in the same leaf and to identify neighboring categories that lack protection.","Practitioners can match defensive mechanisms to concrete threat categories: monitoring for query floods, watermarking and fingerprinting for ownership disputes, perturbation and access control for prevention, and integrated frameworks for high-stakes deployments.","The paper's proposed metrics—extraction accuracy, fidelity, efficiency, transferability, and parameter similarity for attacks; defense success rate, utility–security trade-off, query detectability, and robustness to adaptive attacks for defenses—give a common evaluation vocabulary to a field that has lacked one.","The explicit future directions, including certified defense guarantees, standardized benchmarks, and cross-environment integration, become concrete research programs rather than vague calls for more work.","The taxonomy highlights the utility–security trade-off as a structural feature of defenses, meaning that any practical protection must be evaluated not just by attack failure but by how much legitimate accuracy and latency are sacrificed."],"supporting_citations":[{"why":"A prior survey of model extraction attack methodologies that the paper positions its unified framework against.","marker":"[62]"},{"why":"Another prior survey that catalogs attacks; the paper contrasts its integrated attack–defense taxonomy with this catalog-only approach.","marker":"[153]"},{"why":"Knockoff Nets, the foundational substitute-model training attack, anchors the query-based branch of the taxonomy.","marker":"[155]"},{"why":"MAZE, a data-free model stealing attack using zeroth-order gradient estimation, anchors the data-driven attack branch.","marker":"[96]"},{"why":"CSI NN, an electromagnetic side-channel attack that reverse-engineers neural architectures, anchors the side-channel attack branch.","marker":"[18]"},{"why":"D-DAE, a defense-penetrating gradient-based extraction attack, anchors the gradient-based attack branch.","marker":"[41]"},{"why":"PRADA, a query-pattern-based detector, anchors the monitoring-based attack detection defense branch.","marker":"[95]"},{"why":"DAWN, a dynamic adversarial watermarking framework, anchors the ownership-verification defense branch.","marker":"[193]"},{"why":"Adaptive misinformation, which gives wrong answers to out-of-distribution queries, anchors the output-perturbation prevention branch.","marker":"[98]"},{"why":"A model extraction warning system for MLaaS anchors the access-control prevention branch and the cloud computing environment discussion.","marker":"[100]"}],"fun_headline_variants":["Taxonomy categorizes model extraction attacks and defenses in three axes","Survey maps model extraction attacks across cloud, edge, and federated setups","Model extraction attacks decoded: a taxonomy for defense","Utility vs security: taxonomy for model extraction defenses"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The claim that the taxonomy is unified and comprehensive depends on the assumption that the papers the authors selected and placed in Figure 3 are representative of the whole model extraction literature, but the paper gives no explicit search protocol, inclusion criteria, or coverage dates to establish that representativeness.","fun_headline_variants_meta":{"raw":{"variants":["Taxonomy categorizes model extraction attacks and defenses in three axes","Survey maps model extraction attacks across cloud, edge, and federated setups","Model extraction attacks decoded: a taxonomy for defense","Utility vs security: taxonomy for model extraction defenses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000837,"raw_usage":{"total_tokens":3495,"prompt_tokens":764,"completion_tokens":2731,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":2664}},"tokens_in":508,"tokens_out":2731,"duration_ms":21641,"temperature":1.0,"reasoning_tokens":2664,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:07:56.650471+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search using the same keywords with explicit inclusion criteria could count published extraction attacks and defenses that do not fit any leaf of Figure 3; if a material number of well-established methods require a new branch on the taxonomy, the comprehensiveness claim fails. A simpler check: the paper claims to be the first to combine attack mechanisms, defense strategies, and computing environments, so locating any earlier survey that already integrates all three axes would also falsify the novelty claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Knockoff Nets, the foundational substitute-model training attack, anchors the query-based branch of the taxonomy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"DAWN, a dynamic adversarial watermarking framework, anchors the ownership-verification defense branch."}],"review_version":1}