{"id":"8f21c356-5ea5-4a0a-8a17-b4987c861425","arxiv_id":"2504.13573","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A measurement of 220K Ethereum NFT collections identifies 8,019 name-squatting collections imitating 654 popular projects, with $59.26M in associated mint fees and creator royalties.","lead":"This study measures how scammers imitate popular NFT projects by creating copycat collections with similar names and images, finding 8,019 such collections on Ethereum. It estimates these schemes moved about $59 million from over 670,000 accounts, which gives marketplaces and users concrete warning signals.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated precision of the Stage IV false-positive filter: the 8,019 collections, 670K victims, and $59.26M all rest on five generic heuristics with no ground-truth evaluation, and Section 7 admits false positives cannot be eliminated.","rationale":"The paper is a valuable first systematic measurement with transparent data collection and a reproducible pipeline; the qualitative findings, that name-similar collections exist and cluster into recognizable tactics, are credible. However, the precision of the detection stage is unsupported: the five heuristics in Section 3.3.4 are not cybersquatting-specific, no precision or recall is reported, and the 'manual review' is described as an automated majority rule while also being claimed to ensure 'ground truth-like accuracy.' Section 7's admission that false positives cannot be eliminated is both honest and damaging to the exact counts. The reader's weakest assumption is the same concern I would flag, so I agree with the CONDITIONAL verdict: the existence and taxonomy claims are likely sound, but the exact 8,019 / 670K / $59.26M numbers should be treated as upper bounds until precision and deception are measured. The proposed audit directly addresses this: if annotators confirm high precision and a large share of sampled victims had prior exposure to the official collection, the concern is resolved; if not, the paper requires revision or a reframing of its headline impact claims.","tokens_in":24108,"tokens_out":7043,"duration_ms":66133,"concrete_test":"Precision audit with blinded annotators: randomly sample 300 of the 8,019 flagged collections; have two independent annotators (not involved in the study) classify each as (a) deliberate name imitation of the target project, (b) legitimate derivative/fan project, or (c) unrelated/coincidental name, using only public on-chain history, metadata, and the official project's relationship. Compute precision and Cohen's kappa. Additionally, for a random sample of 100 'victim' addresses from top flagged collections, check whether the address had any prior on-chain or social interaction with the official target collection (held official tokens, followed official Twitter, visited official site) before its first payment; if either precision or the prior-contact fraction falls below a pre-registered threshold (e.g., 90% and 50%), reframe the headline figures as upper bounds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3.4's Stage IV filter is the load-bearing component: a candidate collection is counted as cybersquatting if it meets four of five heuristics (price collapse, transfer collapse, social silence, malicious external label, image similarity). These heuristics are generic indicators of an inactive or failed NFT project; they do not demonstrate that the name was chosen to deceive, and the threshold-based 'manual review' has no labeled ground truth, no inter-annotator agreement, and no reported precision. The paper itself concedes 'false positives cannot be eliminated' (Section 7), and examples such as 'Doodles Flipped' or 'Metaverse Cool Cats' (Tables 7) could plausibly be legitimate derivatives or fan projects. Because every downstream number, 8,019 collections, 670,817 victim addresses, and $59.26M in financial exploitation, is computed from this list, even a modest false-positive rate shrinks all headline figures. Independently, the victim definition in Section 2 and Section 6.2 counts every minter and buyer as a 'victim' without evidence of deception, so the financial impact figure is gross revenue from all contributors, not demonstrated exploitation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a measurement study of cybersquatting NFT collections on Ethereum. The authors collect on-chain data for 220,918 NFT contracts and 245M transfer events up to September 2023, build a keyword corpus using domain-squatting tools (URLCrazy, URLInsane, DNSTwist), match those keywords against collection names, and apply a five-heuristic false-positive filter to identify 8,019 cybersquatting collections targeting 654 of 996 popular projects. They characterize seven naming tactics, analyze collection metadata, social media, content theft, and phishing links, cluster creators into 794 possible scam campaigns, and estimate 670,817 victim addresses and $59.26 million in financial exploitation from mint fees and creator earnings. The paper concludes with mitigation proposals and an ethical disclosure that 2,572 of the identified collections were removed by OpenSea after reporting.","tokens_in":24395,"tokens_out":6385,"duration_ms":58653,"significance":"If the headline numbers are validated, this would be the first large-scale, systematic measurement of NFT cybersquatting and a useful reference for marketplace defenses. The paper's strengths are its transparent on-chain data collection from a Geth node, the public data release, the comparison against the prior Levenshtein-distance approach, the identification of several concrete scam-campaign clusters, and the phishing-link analysis with external corroboration from VirusTotal and SlowMist. However, the quantitative claims—8,019 collections, 670K victims, and $59.26 million—all rest on an unvalidated heuristic filter and on an equating of payment with victimization. The paper's contribution is best assessed as a reproducible measurement pipeline and a qualitative taxonomy; the precision of its headline figures is currently not established.","major_comments":[{"comment":"The Stage IV false-positive filter is the load-bearing component of the headline counts, but its precision is not measured. A collection is counted as cybersquatting when it meets four of five heuristics (price collapse, transfer collapse, social silence, malicious external label, image similarity), which are generic signs of an inactive or failed project and do not by themselves establish deceptive intent. No labeled ground-truth set, inter-annotator agreement, or precision/recall numbers are reported, and Section 7 concedes that false positives cannot be eliminated. Because the 8,019 collections, 670,817 victim addresses, and $59.26 million are all computed from this list, even a modest false-positive rate changes every headline figure. Please provide a validation study—for example, a random sample of flagged and non-flagged candidates reviewed by multiple annotators, and a precision estimate using the 2,572 collections that OpenSea removed after disclosure as a partial ground truth—and report the resulting ranges for all downstream quantities.","section":"§3.3.4 and §7"},{"comment":"The paper defines victims as 'entities that have direct financial contributions to scammers' and counts every minter and secondary-market buyer as a victim. This equates payment with deception: no evidence is presented that each address was misled, and legitimate purchases of derivative or fan collections (which the Stage IV filter may include) would be counted as losses. Correspondingly, the $59.26 million figure in §6.3 is gross revenue received by the flagged contracts, not demonstrated financial exploitation and not net of any resale proceeds. Please relabel these quantities as 'payer addresses' and 'gross funds received,' and provide a sensitivity analysis that excludes borderline cases (e.g., collections with no malicious label and no phishing link) to show how the victim and profit totals change.","section":"§2 and §6.2"},{"comment":"The seven-tactic distribution in Table 3 is generated from the same domain-squatting tools (URLCrazy, URLInsane, DNSTwist) used to build the keyword corpus, so the relative frequencies are conditional on the generator's variation space rather than an independent discovery about attacker behavior. For instance, the predominance of combination squatting (67.22%) and the absence of other possible tactics may reflect the fact that these tools do not generate those variations. The Section 7 limitation that 'these names are randomly generated' does not quantify this. Please state this conditioning explicitly when presenting the taxonomy as an answer to RQ1, and ideally report the generator coverage by comparing with a Levenshtein-based or manually curated set of additional squatting variants.","section":"§3.3.2 and §4.2"}],"minor_comments":[{"comment":"The workflow is described as a three-stage method, but the heading in §3.3.4 is labeled 'Stage IV'; please align the numbering.","section":"§3.3"},{"comment":"Table 3's ERC-721 row total is 6,494, while §4.1 states 6,495; the identical-name percentage is reported as 8.76% in Table 3 and 8.77% in the text. These should be reconciled.","section":"Table 3 and §4.1"},{"comment":"Table 1 reports 245,377,798 total transfer events, while §3.2.1 sums to 219,114,287 + 27,548,181 = 246,662,468; similarly, Table 1 reports 98,390,236 market trades while §3.2.2 states 97,902,053. Please reconcile the discrepancies.","section":"Table 1 and §3.2"},{"comment":"The ETH-USD conversion used to obtain $59.26 million from 21.6K ETH is not stated; please report the exchange rate and date, and note that the conversion is time-varying.","section":"§6.3"},{"comment":"The mutation-based squatting subtotal (1,925, or 24.0%) is discussed in the text but not printed in Table 3; adding a subtotal row would help readers verify the arithmetic.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper has solid measurement infrastructure and a useful public dataset, but the headline figures are likely to be cited independently of their caveats. I would require the precision validation and the reframing of victim/profit language before accepting the paper as a measurement study. The internal numeric inconsistencies (Table 1 and Table 3) suggest that a careful proofreading pass is also needed before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, this is the first systematic measurement of NFT collection name-squatting, and the core existence and prevalence claims hold up. Second, the headline numbers — 8,019 collections, 670K victims, $59.26M — are softer than the abstract implies, because the detection pipeline's precision is never measured and the victim definition counts every payer. I'd send it to review, but I'd insist the authors reframe the quantitative claims as upper bounds and add ground-truth validation.\n\nWhat is actually new: the seven-tactic taxonomy, the campaign clustering (794 clusters), the victim and profit accounting, and the comparison against Das et al.'s Levenshtein approach. The data collection is transparent (220K contracts, full transfer events, market trades, metadata) and the repo is promised. The finding that most squatting collections are short-lived, low-supply, and combination-squatted is robust and useful. The contrast with Das et al. — 424 of 7,316 non-identical cases caught, 85% of their false positives — is a genuine, reproducible improvement.\n\nSoft spots, in proportion. The Stage IV filter (four-of-five heuristics: price collapse, transfer collapse, social silence, malicious label, image similarity) is load-bearing and has no measured precision. Those heuristics mostly detect inactivity or failure, not deceptive intent. The paper admits false positives cannot be eliminated but does not quantify them. Tables 7 and 8 include projects like Doodles Flipped and Metaverse Cool Cats that could be legitimate derivatives or fan projects. Since every downstream number flows from this list, even 10–20% false positives would shrink all headline figures materially. The victim definition (§2, §6.2) counts every minter and buyer as a victim without evidence of deception; the $59.26M is gross revenue from all contributors, not demonstrated exploitation. That should be labeled an upper bound. There are also small internal inconsistencies (6,495+1,525 vs 8,019; 8.76% vs 8.77%) that suggest a hasty final pass, but they are minor.\n\nThe circularity concern from the stress-test note is real but partial: the seven-tactic distribution is partly an artifact of the generators' variation space, so tactic frequencies should be described as bounded by the generators' coverage. That said, the existence of each tactic is still evidenced by concrete examples and on-chain behavior, so I would not call the taxonomy circular in a damaging way.\n\nWho this is for: NFT marketplaces, security vendors, and anyone studying online impersonation. A serious referee should engage with it, but the revision needs precision evaluation (manual labels with inter-annotator agreement, or a held-out set) and a rewritten abstract that separates detected squatters from confirmed victims.\n\nMy call: accept for peer review, with major revision expected.","headline":"A credible first systematic measurement of NFT name-squatting, but the headline victim and profit numbers are unvalidated upper bounds due to an unmeasured-precision filter.","tokens_in":24887,"tokens_out":704,"would_cite":true,"duration_ms":8160,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A measurement of 220K Ethereum NFT collections finds 8,019 that imitate 654 popular projects and pulled $59.26 million from over 670,000 victims.","keywords":["cybersquatting","NFT scams","Ethereum","name squatting","counterfeit NFTs","phishing","measurement study"],"falsifier":"Take a random sample of the 8,019 flagged collections, check whether the official project or the marketplace confirmed them as fraud (through removals, spam flags, or victim reports), and see whether users who paid actually believed they were buying the official item. If a large share of flags are legitimate fan projects or abandoned derivatives, the 670K victim count and the $59.26M figure are overestimates.","tokens_in":23913,"feed_emoji":"🎭","tokens_out":7778,"duration_ms":67890,"temperature":0.7,"pith_summary":"The paper claims that cybersquatting—registering names that imitate a well-known brand in order to deceive buyers—has moved from web domains into NFT collections on Ethereum. Using 220K NFT collections and more than 150M tokens, it identifies 8,019 cybersquatting collections aimed at 654 popular projects, and argues they extracted about $59.26 million from over 670,000 victims through mint fees and creator royalties. If true, this means name-based NFT counterfeiting is not anecdotal but a structured economy with at least seven naming tactics and hundreds of organized scam campaigns. The result matters because NFT marketplaces, investors, and creators all rely on collection names as the main way to find and trust assets.","feed_headline":"8,019 NFT collections impersonate top projects, study finds","feed_subtitle":"Scan of 220K Ethereum collections ties $59.26M in scam profits and 670K victims to name mimicking.","key_machinery":"The carrying object is a four-stage detection pipeline adapted from domain-squatting practice. Stage I selects the top 996 NFT collections by market cap as targets; Stage II feeds their names into three domain-squatting name generators to synthesize a keyword corpus of identical, combination, and mutation variants; Stage III exact- and partial-matches those keywords against all 220K collection names; Stage IV filters false positives by discarding same-team derivatives and pre-official deployments, then keeps a candidate only if at least four of five signals fire: floor price down more than 90% for 30 days, monthly transfers down more than 90% for two months, no social-media activity after the last on-chain event, a malicious external label, or perceptual-hash image similarity (DHash) below 5. The same pipeline produces the seven-tactic taxonomy and the victim/profit accounting, so the paper's headline numbers all flow through this heuristic gate.","core_discovery":"The paper claims that NFT cybersquatting on Ethereum is widespread and organized. From 220,918 NFT collections and 245M transfer events, it identifies 8,019 cybersquatting collections that target 654 of the 996 most valuable NFT projects; these involve 5.5M+ tokens and 1,679,896 transfer events. It further claims that scammers use seven naming tactics—identical name replication, combination squatting, and six mutation variants (character insertion, character omission, case substitution, misspelling substitution, homoglyph, and homophone)—with combination squatting accounting for 67.22% of cases. On the actor side, the paper reports 6,411 distinct scammer addresses, 794 scam campaigns found by clustering shared external links, creator addresses, and exchange deposit addresses, and 670,817 victim addresses. Finally, it claims that 2,255 of these collections were profitable, generating 21.6K ETH (about $59.26 million) from mint fees and creator earnings.","pith_inferences":["Beyond the paper, the same name-generation and signal-filter design should transfer to NFT ecosystems on other blockchains and to name-based Web3 identifiers such as decentralized name registries, where the economic incentive to imitate a well-known name is identical.","Beyond the paper, the $59.26M total counts only mint fees and creator royalties; wallet-draining phishing launched from the associated websites could add losses that this measurement does not capture.","Beyond the paper, the marketplace takedown response noted in the paper (2,572 of 8,019 flagged collections removed by the time of writing) can serve as independent ground truth: a follow-up precision study could check how many flags the platform confirmed, which would test the heuristic gate.","Beyond the paper, since 'NFT', 'official', 'by', and 'collection' are the most frequent combination-squatting suffixes, a simple registration-time blocklist of brand name plus these tokens would plausibly stop most new squats before they mint, a mitigation the paper does not explicitly propose."],"forward_implications":["Marketplaces can run the name-variant matching and the five-signal filter at listing time, because the manual review step in the paper is described as fully automatable, making real-time blocking of new squatting collections feasible.","Name-similarity detectors that rely only on string edit distance (Levenshtein distance) will keep missing most NFT squatting: the paper finds that such a method catches only 424 of 7,316 non-identical cases, so detection needs explicit mutation and combination patterns.","Enforcement directed at the few biggest collections would have outsized effect: 10% of profitable cybersquatting collections capture 84.64% of total profit, and 20% of mint-fee-earning collections capture over 90% of mint revenue.","Because 81.53% of flagged collections stay active for one day or less, any practical defense has to act within hours of deployment, not after weeks of review.","The 794 identified scam campaigns mean takedowns should target groups of collections linked by shared creator addresses, external websites, or exchange deposit addresses, rather than individual contracts in isolation."],"supporting_citations":[{"why":"Supplies the prior NFT-security baseline that used Levenshtein distance to find counterfeit NFT names; the paper compares its detection coverage against this method.","marker":"[10]"},{"why":"Supplies the counterfeit ERC-20 token detection patterns and false-positive filtering inspiration that the NFT pipeline adapts.","marker":"[27]"},{"why":"Supplies the NFT rug-pull detection methodology, threshold choices, and on-chain/off-chain data conventions used for filtering and profit measurement.","marker":"[28]"},{"why":"Supplies the method for collecting secondary-market trade events from major NFT marketplaces, which underlies the profit calculation.","marker":"[33]"},{"why":"One of the domain-squatting name generators used to synthesize the cybersquatting keyword corpus.","marker":"[34]"},{"why":"One of the domain-squatting name generators used to synthesize the cybersquatting keyword corpus.","marker":"[35]"},{"why":"One of the domain-squatting name generators used to synthesize the cybersquatting keyword corpus.","marker":"[36]"},{"why":"Supplies the mobile-app squatting classification and filtering approach that informs the NFT naming-tactic taxonomy and false-positive rules.","marker":"[40]"},{"why":"Supplies the deposit-address clustering heuristic used to link NFT creators and identify the 794 scam campaigns.","marker":"[53]"}],"fun_headline_variants":["NFT name mimicry: 8,019 fake collections net $59M","Cybersquatting hits NFTs: 8,019 fakes, $59M in harm","Web3 cybersquatting: 8,019 NFT clones, 670K victims","8,019 NFT squats target top projects, study finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole estimate depends on the assumption that a similarly named collection showing four of five warning signs—price collapse, transfer collapse, social-media silence, a malicious label, or near-identical images—was actually created to deceive people, and the paper never measures how often that label is wrong.","fun_headline_variants_meta":{"raw":{"variants":["NFT name mimicry: 8,019 fake collections net $59M","Cybersquatting hits NFTs: 8,019 fakes, $59M in harm","Web3 cybersquatting: 8,019 NFT clones, 670K victims","8,019 NFT squats target top projects, study finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000618,"raw_usage":{"total_tokens":2872,"prompt_tokens":952,"completion_tokens":1920,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":1831}},"tokens_in":568,"tokens_out":1920,"duration_ms":12986,"temperature":1.0,"reasoning_tokens":1831,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:05:31.355876+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the 8,019 flagged collections, check whether the official project or the marketplace confirmed them as fraud (through removals, spam flags, or victim reports), and see whether users who paid actually believed they were buying the official item. If a large share of flags are legitimate fan projects or abandoned derivatives, the 670K victim count and the $59.26M figure are overestimates.","supporting_citations":[{"cited_title":"Under- standing security issues in the nft ecosystem,","cited_arxiv_id":null,"evidence_quote":"Supplies the prior NFT-security baseline that used Levenshtein distance to find counterfeit NFT names; the paper compares its detection coverage against this method."},{"cited_title":"Tracking counterfeit cryptocurrency end-to-end,","cited_arxiv_id":null,"evidence_quote":"Supplies the counterfeit ERC-20 token detection patterns and false-positive filtering inspiration that the NFT pipeline adapts."},{"cited_title":"Miracle or mirage? a measurement study of nft rug pulls,","cited_arxiv_id":null,"evidence_quote":"Supplies the NFT rug-pull detection methodology, threshold choices, and on-chain/off-chain data conventions used for filtering and profit measurement."},{"cited_title":"Unveiling the paradox of nft prosperity,","cited_arxiv_id":null,"evidence_quote":"Supplies the method for collecting secondary-market trade events from major NFT marketplaces, which underlies the profit calculation."},{"cited_title":"Urlcrazy,","cited_arxiv_id":null,"evidence_quote":"One of the domain-squatting name generators used to synthesize the cybersquatting keyword corpus."},{"cited_title":"Urlinsane,","cited_arxiv_id":null,"evidence_quote":"One of the domain-squatting name generators used to synthesize the cybersquatting keyword corpus."},{"cited_title":"Dnstwist,","cited_arxiv_id":null,"evidence_quote":"One of the domain-squatting name generators used to synthesize the cybersquatting keyword corpus."},{"cited_title":"Mobile app squatting,","cited_arxiv_id":null,"evidence_quote":"Supplies the mobile-app squatting classification and filtering approach that informs the NFT naming-tactic taxonomy and false-positive rules."},{"cited_title":"Address clustering heuristics for ethereum,","cited_arxiv_id":null,"evidence_quote":"Supplies the deposit-address clustering heuristic used to link NFT creators and identify the 794 scam campaigns."}],"review_version":1}