{"id":"463c38c6-e846-48c5-bfb7-e72283dfa2fe","arxiv_id":"2607.07414","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MIH shows perfect pairwise precision but only 0.36/0.44 per-wallet precision/recall on real CASP ground truth, with near-total failure for some services and dominance by one large entity.","lead":"The multi-input heuristic for Bitcoin address clustering looks strong on labeled pairs but fails badly for many services once full clusters and entity-level metrics are used. Law enforcement must treat MIH clusters as metric- and entity-dependent leads, not reliable attribution.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The strongest claim is an empirical observation about metric- and entity-dependence on the given ground truth, not a universal reliability number. Tables 2–4 and the leave-one-out analysis supply direct evidence for that observation. The reader correctly notes that the seven-service set is limited and non-public, which justifies CONDITIONAL rather than unconditional ACCEPT; that limitation is already stated by the authors and does not falsify the reported numbers or the legal caution that follows from them. No internal inconsistency, calculation error, or unacknowledged assumption that would reverse the claim was found. Therefore the reader’s verdict and confidence remain appropriate; no adjustment is required.","tokens_in":22323,"tokens_out":478,"duration_ms":4784,"concrete_test":"Re-run the nine-metric framework (public mih-toolkit) on any second independent labeled CASP or wallet set of comparable size; if the same qualitative pattern appears—high labeled-domain pairwise precision, substantially lower full-cluster per-wallet scores, and large entity-level variance—the claim is reinforced; if the pattern collapses, the present results are idiosyncratic to this seven-service set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper’s central claim is carefully scoped: MIH reliability is metric-dependent and entity-dependent on this ground-truth set of seven European CASPs (cutoff block 795357). Pairwise scores on the labeled domain (precision 1.0, recall 0.71) are driven by service 1 and ignore unlabeled contamination; full-cluster per-wallet metrics fall to 0.36/0.44; entity-level results show near-total failure for several services (Table 3); leave-one-out confirms service-1 dominance (Table 4). These statements are directly supported by the reported numbers and do not require the seven services to be a representative sample of all Bitcoin entities. The authors already flag the limited entity set, non-public data, unfiltered CoinJoins, and static snapshot in §5.3 and Ethical Considerations. The reader’s weakest-assumption concern (representativeness for general LE conclusions) is therefore a scope caveat the paper itself acknowledges rather than a hidden load-bearing flaw that would overturn the stated claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper evaluates the multi-input heuristic (MIH) for Bitcoin address clustering against ground-truth address-to-entity mappings from seven European crypto-asset service providers reported under statutory obligations (cutoff block 795357). It reimplements nine metrics from prior work (pairwise precision/recall/F1, per-wallet precision/recall/F1, NMI, aNMI, AER) in a reusable framework, applies them to the full unfiltered transaction corpus, and reports both dataset-level and entity-level results. On the labeled domain, pairwise precision is 1.0 and recall 0.71 (F1 0.83), but full-cluster per-wallet metrics fall to precision 0.36 / recall 0.44 / F1 0.27, with NMI 0.41, aNMI 0.36 and AER 0.51. Entity-level scores (Table 3) and leave-one-out analysis (Table 4) show that dataset-level pairwise performance is dominated by one large service and that several services are essentially unrecovered. The authors interpret these metric- and entity-dependent findings for three law-enforcement use cases under German and U.S. procedure and recommend treating MIH clusters as investigative leads rather than definitive attribution.","tokens_in":22491,"tokens_out":1122,"duration_ms":8947,"significance":"If the reported numbers hold, the paper supplies the first systematic MIH evaluation on recent, legally mandated, non-heuristic ground truth and unifies the fragmented metric landscape of prior studies. The reusable open evaluation framework, explicit separation of labeled-domain pairwise scores from full-cluster per-wallet scores, and entity-level / leave-one-out analyses are concrete methodological contributions that future clustering evaluations can reuse. The legal discussion usefully maps metric families onto asymmetric false-positive / false-negative costs in suspicion, seizure, and trial settings. Even with a small entity set, the demonstrated sensitivity of headline scores to metric choice and to a single dominant service is a result that both researchers and practitioners should take into account.","major_comments":[{"comment":"Section 3.1 and Figure 1: the ground-truth set comprises only seven services with extreme size imbalance (service 1 holds ~84% of labeled addresses; services 4, 6, 7 are near-singletons). Table 3 and Table 4 correctly expose the resulting heterogeneity and service-1 dominance, yet the abstract and conclusion still frame the findings as guidance for law-enforcement use of MIH in general. The Limitations section already flags the small entity set; the manuscript should more tightly scope every general claim (including the abstract) to “these seven CASPs” and treat broader LE recommendations as provisional until additional entity types are evaluated.","section":null},{"comment":"Section 3.1 / 3.2: the authors deliberately leave CoinJoin / mixing transactions unfiltered “for comparability with prior work.” Because such transactions systematically violate the MIH co-spend assumption, the low per-wallet precision (0.36) and high AER for several services may partly reflect contamination rather than pure co-spend failure. A short sensitivity experiment that re-runs the nine metrics after a standard CoinJoin filter (or at least reports the fraction of multi-input transactions that match known CoinJoin patterns) would clarify how much of the reported degradation is attributable to known assumption violations versus genuine entity fragmentation.","section":null}],"minor_comments":[{"comment":"Table 2 caption and §4.1: the per-wallet F1 of 0.27 is correctly described as the macro-average of entity-level F1s, not the harmonic mean of the dataset-level precision and recall; a parenthetical reminder in the table itself would prevent misreading.","section":null},{"comment":"Figure 1 uses a log-scale share axis and labels services only by index; adding absolute address counts (or a second panel) would make the imbalance immediately quantitative for readers.","section":null},{"comment":"Section 5.2 is long relative to the empirical core; a short summary table mapping each legal use-case to the most relevant metric family (e.g., per-wallet precision for seizure risk) would improve accessibility for non-legal readers.","section":null},{"comment":"Typographical / consistency: “absolutley” (p. 2), occasional spacing around decimals (“0 .71”), and mixed “forfeiture/ confiscation” hyphenation should be cleaned.","section":null},{"comment":"Open-science statement: the anonymous repository link is welcome; once de-anonymized, a DOI or permanent archive citation would strengthen long-term reproducibility claims.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The non-public ground truth is a genuine strength (legally mandated CASP reports) but also a practical barrier to independent verification. The open metric framework partially mitigates this; editors may wish to confirm that the authors can share the labeled-domain contingency tables or synthetic surrogates under the existing agreement so that the numerical claims remain checkable. Scope is appropriate for a security / forensics venue; the dual legal-technical framing is a plus rather than a mismatch."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: on this CASP ground truth, MIH looks clean if you only score labeled pairs (precision 1.0, recall 0.71), but full-cluster per-wallet numbers drop to 0.36/0.44, one big service drives the pairwise scores, and several services nearly vanish. That gap is the paper.\n\nWhat is new is the combination of a reusable nine-metric reimplementation (Nick, Cazabet, Gong) plus legally mandated European CASP address maps, plus entity-level and leave-one-out tables. Prior MIH checks used old, simulated, or vulnerability-derived labels and never put all nine metrics on the same footing. The experimental choices are transparent: full corpus to height 795357, no CoinJoin filter for comparability, clear labeled-domain vs full-cluster definitions, macro-averaging, and the leave-one-out that shows service 1 owns the pairwise recall. Code is shipped; the legal discussion of suspicion, seizure, and Daubert-style use is concrete rather than hand-wavy.\n\nSoft spots are real but already flagged. Seven services, heavily skewed, non-public labels, static snapshot, unfiltered mixers. Those limit how far you can generalize to individuals or privacy-heavy actors; they do not undercut the scoped claim that reliability is metric- and entity-dependent on this data. The stress-test is right: representativeness is a scope caveat, not a hidden circularity. Math and citations look ordinary and solid; no free parameters, no invented entities.\n\nThis is for people who build or rely on blockchain forensics pipelines and for anyone writing about evidentiary standards. I would bring it to reading group, cite the metric gap and the entity heterogeneity, and send it to peer review. It deserves a serious referee even if the ground-truth set stays small.","headline":"Solid multi-metric MIH evaluation on rare legal ground truth; the pairwise-vs-full-cluster gap and entity failures are real and carefully scoped.","tokens_in":23164,"tokens_out":480,"would_cite":true,"duration_ms":4961,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"The multi-input heuristic for Bitcoin address clustering looks strong only on labeled addresses; full-cluster and entity-level metrics show precision and recall collapse and some services fail almost completely.","keywords":["crypto currency forensics","multi-input heuristic","address clustering evaluation","Bitcoin","blockchain forensics","law enforcement"],"falsifier":"An independent ground-truth collection of comparable size and diversity that yields uniformly high per-wallet precision and recall for every entity (including the small and medium ones) under the same unfiltered multi-input pipeline would overturn the claim of metric- and entity-dependent unreliability.","tokens_in":23205,"feed_emoji":"₿","tokens_out":933,"duration_ms":18597,"temperature":0.7,"pith_summary":"This paper tests the multi-input heuristic (MIH)—the standard rule that groups all input addresses of a Bitcoin transaction as belonging to one controller—against real address-to-entity ground truth obtained from European crypto-asset service providers under legal reporting duties. On the labeled addresses alone the method looks solid: it never merges different reported services and recovers same-service pairs with recall 0.71. Once the full clusters are examined, however, precision and recall fall to 0.36 and 0.44 because unlabeled addresses flood the clusters and many entities are only partially recovered. Performance also varies sharply by service: one large provider is reconstructed well while several others are recovered almost not at all. When these clusters are used to form criminal suspicion, seize assets, or support trial evidence, prosecutors and judges therefore need to treat reliability as both metric-dependent and entity-dependent.","feed_headline":"Bitcoin address clustering fails for many real services","feed_subtitle":"Full-cluster precision drops to 0.36; one large service hides near-total failures elsewhere.","key_machinery":"The multi-input heuristic (MIH), which transitively merges all co-spent input addresses into one cluster, scored by a unified re-implementation of nine metrics (pairwise and per-wallet precision/recall/F1, NMI, aNMI, AER) on legally mandated address-to-entity sets.","core_discovery":"When the multi-input heuristic is evaluated on verified ground-truth mappings from seven European crypto service providers, pairwise metrics restricted to reported addresses give perfect precision and moderate recall, yet metrics that assess the full clusters yield precision 0.36 and recall 0.44; entity-level scores further show near-complete failure for several services. Dataset-level averages are dominated by a single large, well-clustered service, so the heuristic cannot be treated as uniformly reliable for investigative or evidentiary use.","pith_inferences":["Commercial black-box forensic tools that rely on MIH as a core component likely inherit the same label-restricted versus full-cluster gap, so claims of rare false positives may not survive full-cluster scrutiny.","Extending ground truth beyond regulated service providers to mixers, darknet markets or individual wallets would probably enlarge the observed failure modes.","A single law-enforcement-oriented metric that explicitly weights false-positive contamination more heavily for seizure decisions could replace the current heterogeneous suite of nine scores.","Because wallet software and multi-party spending patterns continue to evolve, the same evaluation framework should be re-run periodically on fresh ground-truth snapshots."],"forward_implications":["Prosecutors and judges must treat MIH clusters as investigative leads rather than definitive attribution when forming suspicion or ordering preliminary asset seizure.","Future evaluations of clustering heuristics must report entity-level distributions, not only dataset averages, because averages can mask total failure on specific targets.","Pairwise scores computed only on labeled address pairs are insufficient for operational risk assessment; full-cluster purity measures are required.","Courts that admit clustering evidence under reliability standards need service-specific error figures rather than a single global score.","Combinations of MIH with other heuristics still require per-entity validation against independent ground truth before being used for seizure or trial."],"fun_headline_variants":["Bitcoin multi-input clustering drops to 0.36 precision on full clusters","MIH recovers services poorly: full-cluster recall only 0.44","One large service masks near-total MIH failures for others","Entity-level Bitcoin clustering fails for several verified services","Pairwise MIH looks strong but full clusters reveal major gaps"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The seven European crypto service providers that must report their controlled addresses are representative enough of the entities law enforcement actually targets for the reliability conclusions to generalize.","fun_headline_variants_meta":{"raw":{"variants":["Bitcoin multi-input clustering drops to 0.36 precision on full clusters","MIH recovers services poorly: full-cluster recall only 0.44","One large service masks near-total MIH failures for others","Entity-level Bitcoin clustering fails for several verified services","Pairwise MIH looks strong but full clusters reveal major gaps"]},"model":"grok-4.5","effort":"low","cost_usd":0.003208,"raw_usage":{"total_tokens":1140,"prompt_tokens":812,"num_sources_used":0,"completion_tokens":92,"cost_in_usd_ticks":32080000,"prompt_tokens_details":{"text_tokens":812,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":236,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":812,"tokens_out":92,"duration_ms":2708,"temperature":1.0,"reasoning_tokens":236,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T06:42:34.000229+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"An independent ground-truth collection of comparable size and diversity that yields uniformly high per-wallet precision and recall for every entity (including the small and medium ones) under the same unfiltered multi-input pipeline would overturn the claim of metric- and entity-dependent unreliability.","supporting_citations":[],"review_version":2}