{"id":"6554626e-4e88-4a0d-9244-aea308cfa1bd","arxiv_id":"2506.04307","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"IPFS pinning services and gateways allow anonymous, persistent hosting of malicious content, and live monitoring found denylisted phishing and malware CIDs advertised by major pinning providers.","lead":"This paper tested whether IPFS pinning services and public gateways stop people from anonymously uploading harmful files. It found they mostly do not: temporary email accounts worked, malware-like files were accepted, and cached content stayed online for days.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gateway conclusion ignores known Bad Bits enforcement: experiment never tests denylisted CIDs, so 'lack mechanisms' is unsupported.","rationale":"The paper is a useful measurement study with real experiments: anonymous account creation, AV-flagged uploads over Tor, gateway cache-retention measurements, and IPNI scanning. Those observations are not in question. The issue is the inference from them to the abstract's universal claim. Because the conclusion is asserted as an absence ('lack mechanisms'), a single documented counterexample—Protocol Labs gateways enforcing Bad Bits—is enough to falsify the universal as stated. The authors appear aware of this counterexample (cited in §3.1, stated in §4.4 and §6) yet designed the gateway experiment so that it could not observe it: the four test files per gateway are new CIDs not present in any denylist. Consequently the paper's strongest claim is broader than its evidence. This does not invalidate the measurement results, but it does require that the verdict remain conditional: either the claim is re-scoped to 'the tested files were not blocked' or the denylist-enforcing gateways are excluded from the universal. The reader identified a related weakness (small test set and inference of no scanning), but the existing-mechanism contradiction is more direct and more damaging to the central claim; hence partial agreement. A single targeted experiment can settle whether the concern lands, without rejecting the paper's empirical contributions.","tokens_in":12316,"tokens_out":5393,"duration_ms":52590,"concrete_test":"Independently select a set of CIDs currently listed in the Bad Bits Denylist, including the three retrieved phishing CIDs from §4.4 if still listed, plus a controlled file added to the denylist. Request each through ipfs.io, gateway.pinata.cloud, infura-ipfs.io, flk-ipfs.xyz, and 4everland.io, with non-denylisted control CIDs. Record HTTP status and served content. If ipfs.io returns a block/error while others serve the content, the correct conclusion is that some gateways have and enforce restriction mechanisms, contradicting the abstract's universal 'lack mechanisms' claim; if all serve the denylisted CIDs, the gateway half of the claim is supported for those services.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a universal negative: 'pinning services and public gateways lack mechanisms to assess or restrict the propagation of malicious content.' For gateways, the paper itself provides direct evidence against that universal. Section 3.1 cites [35] and says 'some public gateways follow blocking mechanisms'; §4.4 states that the Bad Bits Denylist 'is enforced on the public gateways operated by Protocol Labs'; §6 repeats that this policy is enforced only on Protocol Labs gateways. The gateway experiment in §4.2, however, uploads four freshly created files per gateway, none of which is on any denylist, and tests cache retention for three days. It never fetches a CID present in Bad Bits through ipfs.io, which the authors identify as enforcing the denylist. Thus cache persistence of arbitrary files cannot establish that gateways 'lack mechanisms' to restrict malicious content: a denylisted phishing page could be blocked at the HTTP layer even if its bytes remain cached. The most the experiment shows is that these five gateways did not block these particular non-denylisted files. The conclusion should be re-scoped to the tested content, or the denylist path must be probed directly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how malicious actors can abuse the IPFS ecosystem for anonymous, persistent content hosting. It evaluates the registration/KYC requirements of five pinning services (Pinata, Filebase, Fleek, Web3.Storage, 4EVERLAND), uploads a simulated AV-flagged executable and a legacy WannaCry sample to test content filtering, and runs a multi-day caching experiment across five public gateways. It additionally mines IPNI advertisements from three pinning providers and reports five CIDs matching the Bad Bits Denylist, three of which were retrieved and identified as phishing or malicious content. Based on these experiments, the abstract and Section 6 claim that pinning services and public gateways lack mechanisms to assess or restrict the propagation of malicious content, allowing attackers to ensure persistent availability while masking their identities. The paper also sketches a double-extortion scenario and discusses countermeasures.","tokens_in":12447,"tokens_out":2969,"duration_ms":31063,"significance":"If the central claim were established at the claimed generality, the paper would be a useful, timely empirical contribution: it would quantify real abuse paths in a widely deployed decentralized storage system and would provide concrete evidence for the security community and for IPFS ecosystem operators. The authors supply reproducible code, describe their ethical safeguards in unusual detail, and combine active measurements with observational data from real provider advertisements. However, the paper's headline conclusion is a universal negative about content-screening and blocking mechanisms, and the evidence presented is too narrow to support that universal claim. The contribution remains valuable if re-scoped to the specific services, files, and gateways tested, and if the gateway claim is tested against the denylist mechanism that the paper itself acknowledges exists.","major_comments":[{"comment":"The gateway experiment tests only four freshly created, non-denylisted files per gateway and never fetches a CID that is on the Bad Bits Denylist through a gateway that enforces it. The paper itself states in §3.1 that 'some public gateways follow blocking mechanisms' and in §6 that the Bad Bits Denylist 'is enforced only on gateways managed by Protocol Labs.' Cache persistence of arbitrary files therefore cannot establish that gateways 'lack mechanisms to assess or restrict the propagation of malicious content,' because a denylisted phishing page could be blocked at the HTTP layer even while its bytes remain cached. The claim should be re-scoped to 'the tested gateways did not block these particular non-list-listed files,' or the experiment should directly probe denylisted CIDs through ipfs.io and the other gateways.","section":"§4.2, §3.1, §6"},{"comment":"The inference that pinning services apply no content scanning rests on one simulated executable and one legacy WannaCry sample per service. Acceptance of these files shows that these particular samples passed, but it does not establish the absence of scanning mechanisms in general: providers may scan and miss these samples, or may scan only certain file types, or may apply post-upload checks that were not observed. The authors should either broaden the test set (more malware families, benign-but-suspicious files, known phish pages, varied file types) or soften the conclusion from 'lack mechanisms' to 'accepted the tested malicious content without detectable screening.'","section":"§4.1, RQ2"},{"comment":"The inference of 'organized exploitation' from five denylisted CIDs advertised by three pinning services over 24 hours lacks a baseline or statistical framing. Without knowing how many denylisted CIDs appear in comparable non-malicious or random advertisement streams, or how often benign users accidentally pin blocked content, overlap across providers cannot be distinguished from chance. The authors should either provide a baseline comparison, or present the five CIDs as anecdotal evidence rather than as a demonstration of organized abuse.","section":"§4.4, RQ4"},{"comment":"The caching experiment is a single un-replicated run over approximately three days on five gateways, with no information about gateway load, cache size, or request patterns from other users during the experiment. Moreover, the paper's proposed attack of 'sustaining availability by periodically sending requests' is not directly tested: the experiment measures retention after fixed intervals (1, 6, 12, 24 hours) and does not measure whether periodic refresh requests actually prevent eviction. The authors should either replicate the experiment across multiple runs, test the refresh mechanism explicitly, or describe the current results as preliminary observations rather than a validated attack.","section":"§4.2, Figure 3"}],"minor_comments":[{"comment":"The column header 'KYC T emp Mail F ree Registered DMCA' is difficult to read due to spacing; please reformat as separate columns or use clear line breaks.","section":"Table 1"},{"comment":"The description of standardizing CID formats for comparison with the Bad Bits Denylist omits the concrete conversion method (e.g., CIDv0 to CIDv1, case normalization); please specify the exact procedure so the matching can be reproduced.","section":"§4.4"},{"comment":"The sentence 'a malicious actor can circumvent it by simply choosing an alternative chunking size when adding the file to IPFS (RQ5)' appears to address content filtering, not RQ5, which concerns gateway caching abuse; the research-question label seems misplaced.","section":"§6, RQ5"},{"comment":"The phrase 'the corresponding URLs have been siphoned' is unclear; presumably the authors mean the malware sample URLs have been sanitized or taken down, but the wording should be revised.","section":"Ethical considerations"},{"comment":"The axis label 'ratio ν5 per hour' is undefined; please define ν and explain how the ratio is computed and smoothed over time.","section":"Figure 3"},{"comment":"The statement that 'the recorded IP address differed from our actual address' would be more convincing if the paper described the method used to confirm the observed IP address (e.g., a WHOIS or IP-echo check from the service's perspective).","section":"§4.1, Tor tests"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim is broader than its evidence, but the underlying measurements are sound enough to be the basis of a revised version. The gateway claim in particular should be re-scoped or tested against the Bad Bits Denylist, since the paper already concedes that Protocol Labs gateways enforce that list. I see no issue with the paper's fit for a security venue; the main risk is that the abstract and Section 6 overstate a universal negative that the experiments do not establish."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful empirical study of IPFS pinning services and gateway caching, with released code, and the IPNI-to-denylist matching is a nice contribution. The abstract and §4.4 push the conclusions further than the data supports, and the gateway claim in particular contradicts the paper's own statements about badbits enforcement.\n\nWhat's new: the KYC test across Pinata, Filebase, Fleek, Web3.Storage, and 4EVERLAND showing temporary email and wallet registration is feasible; the gateway cache persistence experiment; and the 24-hour IPNI enumeration that found five denylisted CIDs advertised by major pinning services. Code is on GitHub. Those are concrete, reproducible results.\n\nWhere it softens: first, the universal \"lack mechanisms to assess or restrict propagation\" is not established for gateways. The paper itself says (§3.1) some public gateways follow blocking mechanisms and (§4.4/§6) that the Bad Bits Denylist is enforced on Protocol Labs gateways. The gateway experiment uploaded four freshly created files per gateway, none denylisted, and measured cache retention for three days. That shows those gateways did not block those particular files; it does not show they lack mechanisms to restrict malicious content. The claim should be re-scoped to the tested content, or the experiment should probe the denylisted-CID path directly. Second, the KYC and scanning tests are thin: one or a few files per service, one simulated malware package plus one WannaCry sample. That supports \"these samples were accepted\" but not \"no content scanning exists.\" Third, the \"organized exploitation\" inference from five CIDs is a stretch without a baseline; coincidental overlap across two or three services is plausible. Fourth, withholding which gateways evicted files is defensible ethically, but it reduces reproducibility.\n\nNone of these are load-bearing for the whole paper: the core observation that anonymous pinning upload with AV-flagged content is currently possible is solid. But the abstract and conclusions should be tightened. I'd send this to peer review and ask for a re-scoped paper with more experimental detail.","headline":"A useful measurement study with real code, but the headline claim overreaches for gateways given the paper's own admission that Bad Bits is enforced on Protocol Labs gateways.","tokens_in":13060,"tokens_out":1977,"would_cite":true,"duration_ms":21678,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that IPFS pinning services and public gateways provide no effective checks against malicious content, letting attackers anonymously upload and keep malware and phishing pages online.","keywords":["IPFS","Anonymity","Pinning Services","Public Gateways","Malware Distribution","KYC Bypass","Bad Bits Denylist","Web3 Security"],"falsifier":"Upload a diverse batch of known-malicious files, for example several dozen VirusTotal-flagged executables and current malware families, to each pinning service and monitor whether any are rejected, quarantined, or reported; if a substantial share are blocked, the claim that providers lack content-assessment mechanisms collapses. Similarly, repeat the gateway-cache experiment with continuous high-volume traffic and compare eviction times; if files are evicted within hours under realistic load, the public-gateway-as-pseudo-pinning attack would not sustain persistence as claimed.","tokens_in":12042,"feed_emoji":"🦠","tokens_out":5946,"duration_ms":53432,"temperature":0.7,"pith_summary":"This paper tries to establish that IPFS, a decentralized file-sharing system, currently offers a practical anonymous hosting channel for malicious content. The authors show that popular pinning services accept accounts from temporary email addresses or cryptocurrency wallets, and accept both a simulated antivirus-flagged executable and a legacy WannaCry sample without any content check. They also show that public gateways cache uploaded files and can keep them reachable for days after the uploader disconnects, turning the cache into a de facto pinning service. Finally, they report live evidence that CIDs listed on the Bad Bits denylist, including phishing pages, were advertised by major pinning services within a single 24-hour window. If these findings hold, IPFS and its ecosystem currently give attackers persistent and hard-to-takedown content availability with weak attribution.","feed_headline":"IPFS pinning services and gateways host malware unchecked","feed_subtitle":"Disposable emails or a crypto wallet suffice to keep AV-flagged files online for days, with no attribution.","key_machinery":"The argument runs on the lifecycle of a CID (content identifier) in IPFS: content-addressed files are pinned by remote providers, then served through HTTP gateways that cache by least-recently-used (LRU) policy. The paper combines four instruments: temporary-email and crypto-wallet registration to bypass KYC; VirusTotal-flagged but benign executables and a legacy WannaCry sample to test content screening; repeated HTTP requests to public gateways to keep files in their LRU caches; and the ipni-cli indexer to collect CIDs from providers' advertisements, matched against the Bad Bits Denylist. The Bad Bits Denylist is the named mechanism Protocol Labs maintains to block undesirable content, but it is enforced only on Protocol Labs gateways and advisory elsewhere; the paper also notes that changing the chunking size when adding a file produces a different CID that bypasses the list.","core_discovery":"The central claim is that the combination of weak KYC practices at pinning services and the caching behavior of public gateways lets an attacker distribute malicious content on IPFS persistently and anonymously. Registration tests found that Pinata and Fleek accepted the first temporary email, Filebase accepted on the fourth attempt, 4EVERLAND required only a crypto wallet, and Web3.Storage required a payment account for uploads; Tor access preserved the attacker's IP anonymity. Upload tests found that every tested service accepted the simulated malware and the WannaCry sample and made them immediately retrievable through public gateways, indicating no content scanning or restriction. Gateway tests found that after one hour, six hours, twelve hours, or twenty-four hours of requested popularity, files remained cached beyond sixteen hours on three of five gateways and the authors could keep them alive by periodic requests. A 24-hour IPNI-based collection recovered over two million CIDs advertised by Pinata, Filebase, and Fleek, five of which were on the Bad Bits denylist, with three retrievable samples identified as phishing or malicious content.","pith_inferences":["The 24-hour indexer snapshot likely undercounts abuse, because ipni random samples from the most recent advertisement and only three providers were monitored; a longer or wider collection would probably find more denylisted and novel malicious CIDs.","The same gateway-cache trick should transfer to other content-addressable storage systems with HTTP gateways and LRU caches, so the finding is a general property of the bridge design rather than an IPFS implementation detail.","A concrete next test would be uploading a larger corpus of known-malicious files to see whether any provider rejects them, which would reveal whether the observed acceptance is universal or sample-specific.","If regulators require gateway-level filtering or pinning-services KYC, the protocol's immutable CIDs mean the burden moves to access points, suggesting reputation systems or decentralized flagging as the only scalable responses."],"forward_implications":["An attacker can create anonymous accounts on major pinning services with a disposable email or a fresh crypto wallet and upload malicious binaries that remain pinned and immediately accessible through public gateways.","Even without a pinning service, an attacker can upload from a short-lived node, generate traffic to fill public gateway caches, and then refresh the caches periodically to keep the content online.","DMCA-compliant removal by one pinning service does not delete content from IPFS; the file stays available through other gateways and nodes, which matters for double-extortion ransomware scenarios.","The Bad Bits Denylist provides only partial mitigation because not all gateways enforce it and because re-chunking content yields a new CID that escapes the list.","Stricter KYC, content scanning, and universal denylist enforcement at gateways would reduce, but not eliminate, this abuse because the underlying protocol has no deletion mechanism."],"supporting_citations":[{"why":"ipni-cli is the tool used to collect provider-advertised CIDs from pinning services for the real-world evidence.","marker":"[3]"},{"why":"Defines IPFS content-addressed design, CIDs, and the DHT provider-record mechanism that the attacks build on.","marker":"[9]"},{"why":"Tor is the tested anonymity layer that lets an attacker upload to pinning services without exposing their real IP address.","marker":"[14]"},{"why":"Prior work on ransomware as a service using IPFS supplies the abuse context that motivates the double-extortion scenario.","marker":"[22]"},{"why":"Documents earlier botnet activity on IPFS, supporting the claim that attackers can use botnets to generate gateway cache traffic.","marker":"[31]"},{"why":"Bad Bits Denylist is the reference list used to identify malicious CIDs among those advertised by pinning services.","marker":"[33]"},{"why":"Prior content-moderation study that shows how alternative chunking sizes generate new CIDs and bypass denylists.","marker":"[35]"},{"why":"Design and evaluation of IPFS that explains Bitswap, gateway caching, and the LRU strategy behind the public gateway attack.","marker":"[38]"}],"fun_headline_variants":["One throwaway email gets malware onto IPFS, study finds","IPFS pinning services have no malware filters, tests show","Anonymous malware persists on IPFS for days via gateways","Weak KYC lets attackers spread malware on IPFS anonymously","IPFS gateway caching keeps malicious files live past 16 hours"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that pinning services and gateways have no content-screening or restriction mechanisms rests on a small test set: one simulated AV-flagged executable per service, one legacy WannaCry sample, five gateways, and a 24-hour observation window; if providers actually scan but miss these particular samples, or if gateway caches behave differently under realistic traffic, the general claim would not follow.","fun_headline_variants_meta":{"raw":{"variants":["One throwaway email gets malware onto IPFS, study finds","IPFS pinning services have no malware filters, tests show","Anonymous malware persists on IPFS for days via gateways","Weak KYC lets attackers spread malware on IPFS anonymously","IPFS gateway caching keeps malicious files live past 16 hours"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00064,"raw_usage":{"total_tokens":2910,"prompt_tokens":870,"completion_tokens":2040,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":1956}},"tokens_in":486,"tokens_out":2040,"duration_ms":15502,"temperature":1.0,"reasoning_tokens":1956,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:45:07.293207+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Upload a diverse batch of known-malicious files, for example several dozen VirusTotal-flagged executables and current malware families, to each pinning service and monitor whether any are rejected, quarantined, or reported; if a substantial share are blocked, the claim that providers lack content-assessment mechanisms collapses. Similarly, repeat the gateway-cache experiment with continuous high-volume traffic and compare eviction times; if files are evicted within hours under realistic load, the public-gateway-as-pseudo-pinning attack would not sustain persistence as claimed.","supporting_citations":[{"cited_title":"https://github.com/ipni/ipni-cli","cited_arxiv_id":null,"evidence_quote":"ipni-cli is the tool used to collect provider-advertised CIDs from pinning services for the real-world evidence."},{"cited_title":"Tor: The second-generation onion router","cited_arxiv_id":null,"evidence_quote":"Tor is the tested anonymity layer that lets an attacker upload to pinning services without exposing their real IP address."},{"cited_title":"Ransomware as a service using smart contracts and IPFS","cited_arxiv_id":null,"evidence_quote":"Prior work on ransomware as a service using IPFS supplies the abuse context that motivates the double-extortion scenario."},{"cited_title":"Looking Into the Eye of the Interplanetary Storm, 2020","cited_arxiv_id":null,"evidence_quote":"Documents earlier botnet activity on IPFS, supporting the claim that attackers can use botnets to generate gateway cache traffic."},{"cited_title":"Bad Bits Denylist","cited_arxiv_id":null,"evidence_quote":"Bad Bits Denylist is the reference list used to identify malicious CIDs among those advertised by pinning services."},{"cited_title":"Guardians of the galaxy: Content moderation in the interplanetary file system","cited_arxiv_id":null,"evidence_quote":"Prior content-moderation study that shows how alternative chunking sizes generate new CIDs and bypass denylists."},{"cited_title":"Design and evaluation of IPFS: a storage layer for the decentralized web","cited_arxiv_id":null,"evidence_quote":"Design and evaluation of IPFS that explains Bitswap, gateway caching, and the LRU strategy behind the public gateway attack."}],"review_version":1}