{"id":"dccdc20c-6a31-4851-aaa5-f97bed559978","arxiv_id":"2412.13021","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A simple baseline that flags a suspect model whenever it repeats a victim model's mistakes performs as well as or better than complex state-of-the-art model fingerprints on existing benchmarks, exposing those benchmarks as too easy.","lead":"The paper tests whether simple methods can detect stolen AI models as well as complex fingerprinting schemes, and finds that a basic baseline which checks agreement on the victim model's mistakes matches or beats published state-of-the-art methods on current benchmarks. It then organizes all fingerprinting approaches into a query, representation, and detection framework and explains why current test benchmarks are too easy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'on par with state-of-the-art' claim is tested only against four reimplemented fingerprints, leaving omitted SOTA methods (e.g., DeepJudge, MetaV) untested; if one of those beats AKH, the conclusion that current benchmarks are solved collapses.","rationale":"The reader's verdict is CONDITIONAL, and the paper's empirical core is credible: the AKH baseline is simple, the evaluation is runnable, and the QuRD decomposition is a useful lens. However, the central claim that a simple baseline matches 'existing state-of-the-art fingerprints' is a negative empirical statement, and negative statements of this kind are only as strong as the comparison set. The paper compares only four reimplemented schemes and leaves other published methods untested. This is not an internal inconsistency, and it is not an attack on the authors' integrity; it is a completeness risk. If DeepJudge or MetaV were to achieve substantially higher TPR@5% on the same benchmark pairs, the paper's strongest conclusion—that current benchmarks are either non-discriminative or solved by the baseline—would no longer hold; the field would still benefit from the QuRD framework and the baseline, but the headline would need narrowing. The proposed test directly settles this: running the omitted methods on the released benchmark pairs and metric would either refute the concern or confirm it. Until that comparison is made, CONDITIONAL is the appropriate verdict: the framework and empirical analysis are valuable, but the central claim's generality depends on an untested comparison set.","tokens_in":14079,"tokens_out":10224,"duration_ms":104965,"concrete_test":"Using the released ModelReuse and SACBench model pairs and the paper's evaluation code, run at least two omitted fingerprints with available implementations—DeepJudge and MetaV—at both their originally reported settings and the paper's 100-query budget, computing TPR@5% as in Eqs. (10)-(11). If any omitted method exceeds AKH's TPR@5% by more than 0.1 at either setting, the 'on par with state-of-the-art' claim and the 'benchmarks solved' conclusion require qualification or revision; if none does, the omission concern is refuted and the central claim is strengthened.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central negative result—that AKH matches or beats state-of-the-art fingerprints and thus current benchmarks are 'solved'—rests on an empirical comparison whose completeness is load-bearing. Section 'Evaluation Setup' limits the comparison to reimplementations of IPGuard, ModelDiff, SAC, and ZestOfLIME, yet the abstract and Figure 1 generalize to 'existing state-of-the-art fingerprints.' Table 1 itself lists prominent omitted methods with available or describable implementations: DeepJudge, FCAE, FUAP, MetaV, FBI, SSF, ModelGiF, TAFA, and AFA. If any omitted fingerprint achieves materially higher TPR@5% on the same ModelReuse/SACBench pairs, the inference that complex schemes add no measurable value on these benchmarks is false. The roughly 100 QuRD mixtures do not fully mitigate this gap because they are constructed only from the components of the four chosen schemes and exclude detection paradigms such as learned meta-verifiers (MetaV) and mutual-information comparisons (FBI). The reader's weakest assumption about Proposition 1 is a secondary, theoretical concern; the more immediate threat to the central claim is that the baseline has not actually been pitted against the full state-of-the-art set.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses model fingerprinting for image classifiers: deciding whether a suspected model h' is a stolen copy of a victim model h. Its first contribution is the Anna Karenina Heuristic (AKH), which samples inputs that the victim model misclassifies and flags the suspect whenever it reproduces the victim's (wrong) labels on those inputs; Proposition 1 gives a one-sided-error property-test bound for this rule. The authors report that AKH matches or beats four reimplemented state-of-the-art fingerprints (IPGuard, ModelDiff, SAC, ZestOfLIME) on ModelReuse (Flower102, SDog120) and SACBench (CIFAR10) at TPR@5%, conclude that current benchmarks are largely solved or non-discriminative, and support this conclusion with a Query-Representation-Detection (QuRD) decomposition that yields roughly 100 new fingerprint variants, a per-task breakdown of detection performance (Table 3), and a benchmark-difficulty analysis based on conditioned Hamming distance. The paper closes with a released toolbox and recommendations for harder benchmarks.","tokens_in":14298,"tokens_out":12980,"duration_ms":113624,"significance":"If the empirical claims hold, the paper delivers a field-relevant negative result: on the two most-used model-fingerprinting benchmarks, the simple AKH baseline is competitive with substantially more complex fingerprints, and the hard remaining subtask is label-only model extraction rather than model leak or probit extraction. The QuRD decomposition is a genuinely useful organizing device that makes the design space and its unexplored combinations explicit, and the task-disaggregated results of Table 3 together with the conditioned-Hamming-distance diagnostic of Figure 3 are concrete tools the community can reuse. The paper is also commendable for open-sourcing the toolbox, reporting five seeded runs, and stating its limitations (adaptive adversaries, non-image modalities) explicitly. The main caveat is that the negative result is a claim about state-of-the-art fingerprints as a class, so its strength depends on the completeness of the reimplemented comparison set, which is the subject of major comment 1.","major_comments":[{"comment":"The central negative result—that AKH performs on par with 'existing state-of-the-art fingerprints' and that ModelReuse and SACBench are therefore 'either not discriminative or solved'—is established only against four reimplemented schemes (IPGuard, ModelDiff, SAC, ZestOfLIME; see Evaluation Setup). This scoping conflicts with the abstract and with the paper's own Table 1, which catalogs at least nine additional existing fingerprints (DeepJudge, FCAE, FUAP, MetaV, FBI, SSF, ModelGiF, TAFA, AFA) that are never evaluated. Because the roughly 100 QuRD mixtures are assembled solely from components of the four reimplemented schemes, they cannot cover detection paradigms such as learned meta-verifiers (MetaV) or mutual-information comparisons (FBI). If any omitted fingerprint attains a materially higher TPR@5% on the same positive/negative pairs, the inference that complex schemes add no measurable value on these benchmarks is false. I request either that the omitted schemes with available or describable implementations be evaluated, or that all claims be explicitly restricted to the four reimplemented fingerprints.","section":"Evaluation Setup / Table 1 / Figure 1"},{"comment":"Contribution 1 overstates the theoretical support for AKH. Eq. (1) is tautological (identical models agree on every input) and does not address the robustness regime h' approximately equal to h. Eq. (2) lower-bounds the true-negative rate of a single-query test by (delta - (1 - alpha')) / (1 - alpha), which is vacuous whenever the bound is negative or small, so it does not establish the separation between positive and negative pairs reported in Figures 1, 3, and 4 and Table 3; it only supports the heuristic 'when to expect gains' reading that the text itself partially acknowledges. Moreover, the experimental AKH is a multi-query score thresholded at FPR = 5%, whereas Proposition 1 analyzes a one-shot test and an unconditional true-negative rate; no argument connects delta_C or the agreement count on misclassified inputs to the TPR@5% statistic used throughout the evaluation. The theory should either be extended to the thresholded multi-query protocol or be presented as heuristic motivation rather than as a guarantee.","section":"Proposition 1 / Eq. (2) / Algorithm 1"}],"minor_comments":[{"comment":"In the appendix proof, the under/over-brace annotations in the displayed equation are ambiguous (the 'delta', '<= 1 - alpha'', and '1 - alpha' labels appear attached to the wrong sub-expressions); please reformat the derivation so that each annotation is unambiguously tied to its intended sub-expression.","section":"Proof of Proposition 1 (appendix)"},{"comment":"The procedure for fixing the detection threshold at FPR <= 5% is not fully specified: the paper explains how TPR and FPR are aggregated (Eqs. 10-11) but not whether the threshold is calibrated on a separate pool of negative pairs or on the same U(h) used for evaluation; this should be stated explicitly since TPR@5% is the paper's headline metric.","section":"Fingerprint evaluation"},{"comment":"The explanation for the TPR@5% drop of ModelDiff and SAC between 100 and 400 queries is explicitly speculative ('We believe that when the number of query points is increased, the self-correlation increases'); a small diagnostic plot of the positive/negative distance gap versus query budget would turn this into an evidence-based claim.","section":"Comparing apples to apples (Figure 4)"},{"comment":"The encoding of model access (four text decorations) and representation type (three text emphases) in Table 1 is extremely difficult to read; using separate columns or a legend matrix for access and representation would make the taxonomy actually usable.","section":"Table 1"},{"comment":"The property-testing statement defines the two cases as h = h' versus h != h' with probability thresholds of 2/3, but the surrounding requirements (Robustness) and the experiments concern h' approximately equal to h under extraction, fine-tuning, and pruning; aligning the formal definition with the approximate-copy regime would improve precision.","section":"Problem setting"},{"comment":"In the sentence 'these two objectives differ in difficulty ... but they also differ greatly in the efforts the adversary has to consent to in order to reach the same accuracy', the phrase 'consent to' appears to be a translation artifact and should be rephrased (e.g., 'the effort the adversary must expend').","section":"The majority of benchmarked tasks are solved"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope well and the code release is a genuine asset. My recommendation of major_revision is driven by the mismatch between the abstract's universal claim ('existing state-of-the-art fingerprints') and the four-scheme evaluation; I would not insist on reimplementing all nine omitted methods, but at least those with readily available code (e.g., FBI, and the describable MetaV meta-verifier) should be added, or the claims should be explicitly scoped. Proposition 1 should be repositioned as a motivating bound rather than a guarantee, since it does not cover the thresholded TPR@5% protocol."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful, thought-provoking paper, and the AKH baseline deserves to become a standard point of comparison. But the claim that it performs on par with existing state-of-the-art fingerprints is only established against four reimplemented methods, and that scope matters.\n\nWhat is genuinely new: the Anna Karenina Heuristic is simple, query the suspected model on inputs the victim misclassifies and flag on label agreement. It is a simplification of negative-sampling ideas already present in SAC and SSF, but naming it and testing it as a baseline is overdue. The QuRD decomposition (Query, Representation, Detection) is a clean way to organize the design space, and the ~100 mixtures plus the open-source toolbox are a real service to the community. The per-task breakdown in Table 3 is also valuable; it shows most model-leak tasks are solved, while model extraction remains the hard subproblem.\n\nSoft spots. The biggest is the generality gap. The abstract says \"existing state-of-the-art fingerprints,\" but the evaluation only includes IPGuard, ModelDiff, SAC, and ZestOfLIME. Table 1 itself lists DeepJudge, MetaV, FBI, SSF, ModelGiF, TAFA, and AFA as prior work, several with available code. If any of those beats AKH on the same pairs, the conclusion that the benchmarks are solved by a simple baseline does not hold as stated. The authors should either add those comparisons or carefully justify the exclusion. Proposition 1 is honest but weak: the lower bound on the true-negative rate can be negative or small, and it does not directly bound TPR@5% as evaluated. That is not fatal for a baseline paper, but the \"first theoretical analysis of a fingerprinting scheme\" claim is likely overbroad given earlier analyses of watermarking and property testing. The figures would also be easier to read with error bars; the table reports them, so this is a presentation gap, not missing evidence.\n\nThe central message stands well enough on its own: on these two benchmarks, a trivial label-agreement test is competitive with much more elaborate fingerprints. That is worth saying and worth publishing after the comparison is broadened. The stress-test concern is legitimate and should be addressed head-on.\n\nWho it is for: anyone building or evaluating model stealing detection. I would bring it to reading group and would cite the QuRD taxonomy. It deserves a serious referee, with the omitted-methods comparison as the main required revision.","headline":"Useful baseline and taxonomy, but the 'on par with state-of-the-art' headline only covers four reimplemented fingerprints, and that scope gap should be fixed before the paper ships.","tokens_in":14866,"tokens_out":2429,"would_cite":true,"duration_ms":23855,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a simple label-agreement baseline, the Anna Karenina Heuristic (AKH), matches or beats state-of-the-art model fingerprinting schemes on the ModelReuse and SACBench benchmarks, and that current benchmarks therefore do…","keywords":["model fingerprinting","model stealing detection","Anna Karenina heuristic","negative sampling","benchmark evaluation","property testing","deep neural networks"],"falsifier":"Build a benchmark whose negative pairs are genuinely unrelated models that share systematic errors — for example, the same architecture trained independently on the same dataset with different random seeds — and measure AKH's false-positive rate; if it exceeds 5%, the negative-sampling premise fails. Alternatively, fine-tune a stolen model on a few dozen of the victim's misclassified inputs so those errors disappear, and check whether AKH's true-positive rate collapses below the published fingerprints.","tokens_in":13867,"feed_emoji":"🕵️","tokens_out":7067,"duration_ms":57681,"temperature":0.7,"pith_summary":"Model fingerprinting aims to detect when a deployed machine learning model has been stolen by checking whether a suspected model behaves like the victim's. The paper's central claim is that the simplest conceivable test — sample an input the victim model misclassifies, then flag the suspect as stolen if it makes the same mistake — matches or outperforms four published fingerprinting schemes (IPGuard, ModelDiff, SAC, ZestOfLIME) on the ModelReuse and SACBench benchmarks. If correct, this means the benchmarks are not measuring hard detection: most stealing/obfuscation tasks they contain are already solved by almost any method, and the baseline solves the remaining model-extraction tasks as well as the complex schemes. The paper then introduces the QuRD decomposition (Query, Representation, Detection) to explain the result, generates roughly one hundred previously unexplored scheme combinations, and proposes metrics for building harder, more representative benchmarks. A sympathetic reader would take away a call for simple baselines and harder benchmarks in model stealing detection research.","feed_headline":"Simple mistake-matching test outdoes complex fingerprints","feed_subtitle":"Repeating the victim's mistakes marks a model as stolen — as well as complex tools do, exposing benchmarks that are too easy.","key_machinery":"Two named mechanisms carry the argument. The Anna Karenina Heuristic (AKH) is the baseline: a label-agreement test performed on the victim's own misclassifications (negative sampling), whose guarantee is Proposition 1's lower bound on the true-negative rate; repeating the one-query test with majority vote drives the false-negative rate down exponentially. The Query, Representation, Detection (QuRD) decomposition is the structural tool: it splits any fingerprint into a query sampler (uniform, adversarial, negative, or subsampling), a representation of the collected outputs (raw labels/logits, pairwise, or listwise correlation), and a detection rule (distance threshold or learned classifier). The paper reimplements IPGuard, ModelDiff, SAC and ZestOfLIME under QuRD, mixes their components to form roughly one hundred new schemes, and uses the decomposition to show that negative sampling consistently matches or beats adversarial sampling while using fewer queries.","core_discovery":"On the paper's own terms, the contribution is the Anna Karenina Heuristic (AKH): a property test that draws one input $x \\sim \\mathcal{D}$ such that the victim model errs, $h(x) \\neq c(x)$, and returns 'stolen' iff the suspected model agrees with the victim on that input, $h'(x) = h(x)$. Proposition 1 shows AKH has one-sided error: if $h' = h$ it always flags, and if $h' \\neq h$ its true-negative rate is $\\delta_C \\geq (\\delta - (1 - \\alpha')) / (1 - \\alpha)$, where $\\delta$ is the Hamming distance between the models and $\\alpha, \\alpha'$ their accuracies. Empirically, the authors report TPR@5% for AKH that is at least as good as IPGuard, ModelDiff, SAC and ZestOfLIME on ModelReuse (Flower102 and SDog120) and SACBench, and better on Flower102. Because this baseline needs only label query access and no gradient computation, the result is presented as an evaluation artifact: the benchmarks are either non-discriminative or already solved, and the QuRD framework is offered as a systematic way to design and compare the next generation of fingerprinting schemes and benchmarks.","pith_inferences":["Inference: AKH's success depends on unrelated models disagreeing with the victim on the victim's mistakes; a natural stress test is to build negative pairs from models trained independently on the same data with different seeds, which would share many systematic errors and could drive AKH's false-positive rate above 5%.","Inference: The paper's argument implies that adaptive adversaries could defeat AKH by fine-tuning a stolen model on a few of the victim's misclassified inputs so that it no longer reproduces those errors; measuring this drop in true-positive rate would quantify the baseline's robustness ceiling.","Inference: The query-budget plateau suggests a design heuristic beyond the paper's scope: future fingerprinting schemes should target the 50–100 query regime rather than the thousands-of-queries regime, since queries beyond that range add cost without detection gain.","Inference: A logit-extension of AKH — comparing full prediction vectors on negative inputs instead of only argmax labels — is a natural untested variant that could keep the baseline strong even when stolen models disagree on the hard-label choice but share soft prediction structure."],"forward_implications":["New fingerprinting papers should be required to compare against a negative-sampling label baseline such as AKH; without it, a reported win over prior art may be a win over nothing.","On the tasks the paper classifies as solved (model leak with identical, quantized, fine-tuned, transferred or pruned weights), a simple label test suffices, and complex fingerprint computations add no measurable detection power at TPR@5%.","Benchmark design should separate per-task results rather than report aggregated scores, because SACBench's low diversity in positive/negative pair generation overestimates fingerprint performance.","For these benchmarks there is an optimal query budget around 50–100 queries, independent of the fingerprinting scheme; extra queries do not help and can even hurt pairwise/listwise representations."],"supporting_citations":[{"why":"Supplies the ModelReuse benchmark and the ModelDiff fingerprint that AKH is compared against.","marker":"(Li et al. 2021)"},{"why":"Supplies the SACBench benchmark and the SAC fingerprint based on negative sampling.","marker":"(Guan, Liang, and He 2022)"},{"why":"IPGuard, an adversarial-sampling fingerprint used as a comparison baseline and modified in QuRD experiments.","marker":"(Cao, Jia, and Gong 2021)"},{"why":"ZestOfLIME, a subsampling-based fingerprint used as a comparison baseline.","marker":"(Jia et al. 2022)"},{"why":"Provides the property-testing formalism that frames the fingerprinting task and justifies repeating the AKH test for a majority vote.","marker":"(Goldreich 2017)"}],"fun_headline_variants":["Simple mistake-matching test rivals complex fingerprints","One-error test matches state-of-the-art model fingerprints","Trivial test for stolen models exposes easy benchmarks","QuRD framework reveals 100 untried fingerprint schemes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the assumption that a stolen model reproduces the victim's mistakes on the selected hard inputs, while an unrelated model disagrees with the victim often enough on those inputs; the paper's own lower bound for the true-negative rate can be close to zero, so this separation is not guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Simple mistake-matching test rivals complex fingerprints","One-error test matches state-of-the-art model fingerprints","Trivial test for stolen models exposes easy benchmarks","QuRD framework reveals 100 untried fingerprint schemes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000293,"raw_usage":{"total_tokens":1754,"prompt_tokens":1038,"completion_tokens":716,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":654,"completion_tokens_details":{"reasoning_tokens":656}},"tokens_in":654,"tokens_out":716,"duration_ms":7829,"temperature":1.0,"reasoning_tokens":656,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:32:14.163268+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a benchmark whose negative pairs are genuinely unrelated models that share systematic errors — for example, the same architecture trained independently on the same dataset with different random seeds — and measure AKH's false-positive rate; if it exceeds 5%, the negative-sampling premise fails. Alternatively, fine-tune a stolen model on a few dozen of the victim's misclassified inputs so those errors disappear, and check whether AKH's true-positive rate collapses below the published fingerprints.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SACBench benchmark and the SAC fingerprint based on negative sampling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"IPGuard, an adversarial-sampling fingerprint used as a comparison baseline and modified in QuRD experiments."},{"cited_title":"S.; and Papernot, N","cited_arxiv_id":null,"evidence_quote":"ZestOfLIME, a subsampling-based fingerprint used as a comparison baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the property-testing formalism that frames the fingerprinting task and justifies repeating the AKH test for a majority vote."}],"review_version":1}