{"id":"65dbafda-b514-4a75-b9f1-aa2f0e6ea15a","arxiv_id":"2411.15091","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Most professional artists lack the awareness and ability to use robots.txt, many hosting platforms do not let them edit it, and a substantial share of AI assistant crawlers ignore the protocol.","lead":"This paper measures how well today's tools let artists keep AI companies from copying their online work. It finds most artists have never heard of the main tool, robots.txt, many cannot change it on their hosting service, and some AI crawlers ignore it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 59% awareness figure and 20/23 non-compliant crawler count anchor the abstract, but the survey's snowball sample and leading robots.txt description make the 'strong demand' generalization conditional on re-analysis without the priming prompt.","rationale":"The paper is a measurement-heavy study with two distinct pillars: large-scale robots.txt/active-blocking measurements and a 203-artist survey. The measurement pillar is strong: the Common Crawl longitudinal analysis is validated against the Internet Archive and a fresh crawl, the crawler-compliance test is a direct experiment on the authors' own sites, and the Cloudflare grey-box evaluation uses ground-truth accounts. The survey pillar is the weakest link in the central claim, because the paper's headline generalization about content creators' awareness and demand rests on a snowball convenience sample and a survey instrument with a leading robots.txt description. The reader's weakest_assumption identifies exactly this: representativeness of the sample and priming from the robots.txt description. I agree with that assessment. My concrete test is the natural follow-up: a control-condition re-run or re-analysis that removes the 'over 90% of artists don't realize' and 'quick win' language, plus a quota-based sampling check. The crawler-compliance result (20/23) is also worth reporting with uncertainty bounds and a reproducible crawler list, but it is a secondary concern because the passive and active measurements are direct and the result is directionally robust: a majority of third-party assistant crawlers do not fetch robots.txt. Overall, the paper's central claim is plausible and well-supported by the measurement portions, but the survey-based awareness/demand numbers are conditional on sample and instrument validity, so the reader's CONDITIONAL verdict is appropriate.","tokens_in":33152,"tokens_out":2554,"duration_ms":20029,"concrete_test":"Re-run or re-analyze the survey with a control condition that removes the leading 'over 90% of artists don't realize...' and 'quick win' framing from the robots.txt description before asking Q26 (adoption intention); if the 75% intention-to-adopt figure drops materially (e.g., below 60%) or if the 59% awareness figure changes when measured on a quota-based sample rather than a snowball sample, the abstract's 'strong demand' and 'critical hurdles' claims need to be re-scaled. Additionally, report the full list of 23 third-party crawlers and the inclusion criteria for GPT apps so the 20/23 result can be reproduced.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that 'strong demand for tools like robots.txt' among content creators 'is significantly constrained by critical hurdles in technical awareness, agency in deploying them, and limited efficacy against unresponsive crawlers.' The load-bearing empirical statements are (a) 59% of surveyed professional artists had never heard of robots.txt, and (b) 20 of 23 tested third-party AI assistant crawlers do not fetch robots.txt and therefore do not respect it. The crawler measurement (Section 5) is a direct, falsifiable test on the authors' own websites and is internally consistent; the main soft spot is the survey. The authors recruited 203 artists via social circles, professional Discord channels, and snowball distribution (Section 4.1), and Appendix D.1 shows that the robots.txt description shown before the adoption questions asserts 'over 90% of artists don't realize...' and calls robots.txt an 'easy way' and 'quick win.' That framing is a leading prompt that can inflate the 75% intention-to-adopt figure. The 59% awareness figure is less susceptible to this particular priming because it is asked before the description, but the sample itself is a convenience snowball sample concentrated in North America (109/203) and heavily skewed toward illustration (163/203), concept art, and digital 2D artists; it is not shown to be representative of professional visual artists generally. The abstract and title generalize to 'content creators,' yet the artist-side numbers come from this one non-probability sample. A second, narrower concern: the 20/23 non-compliant crawler count is presented without uncertainty bounds, and the active measurement only covers 23 third-party assistant crawlers reachable through GPT apps, so the headline 'most AI assistant crawlers do not respect robots.txt' is sensitive to how that list was compiled. These are correctness-risk concerns, not internal inconsistencies, and the measurement portion of the paper is generally careful.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies whether the web-control mechanisms available to individual content creators — robots.txt, noai meta tags, and active blocking via reverse proxies — are known to creators, available to them through their hosting platforms, and effective against AI crawlers. The paper combines four complementary measurements: (1) a longitudinal analysis of robots.txt adoption across 40,455 domains that appear in the Tranco top-100k in every month from October 2022 to October 2024, using Common Crawl snapshots cross-validated against the Internet Archive; (2) a user study of 203 artists, plus a measurement of 1,182 artist websites and the control that their eight most common hosting providers expose; (3) passive and active tests of whether 24 AI-related user agents respect robots.txt on two sites under the authors' control; and (4) an evaluation of active blocking, including an inferred behavior model of Cloudflare's 'Block AI Bots' feature and an estimate of its deployment across the Tranco top-10k. The headline findings are that 59% of surveyed artists had never heard of robots.txt, that most hosting platforms do not let artists edit it, that most large AI companies' data crawlers respect robots.txt while 20 of 23 triggered third-party AI assistant crawlers do not fetch it, and that active blocking is more enforceable but sparsely deployed and incomplete in coverage.","tokens_in":33388,"tokens_out":17263,"duration_ms":151175,"significance":"If the measurements hold, this is the most complete empirical map to date of the gap between what individual content creators want from anti-AI-crawling tools and what the current Web actually provides, and it is likely to become the reference point for the state of AI-crawler blocking circa 2024-2025. The network-measurement components are genuinely strong: the robots.txt parser is validated against RFC 9309 edge cases, the Common Crawl longitudinal data is cross-checked against the Internet Archive and an independent crawl, the crawler-compliance tests are direct and falsifiable (20 of 23 triggered third-party assistant crawlers never fetched robots.txt), and the Cloudflare setting is inferred through paired grey-box tests with the authors' own account as ground truth. The paper also ships public code and data via a GitHub repository. There are no fitted parameters and the central claims do not reduce to an input, so circularity is not a concern.","major_comments":[{"comment":"The 75% intention-to-adopt figure is load-bearing for the abstract's 'strong demand for tools like robots.txt,' but it is measured by Q26, which is asked immediately after a description that asserts 'over 90% of artists don't realize they can use a simple tool called robots.txt,' calls the tool 'an easy way for artists to protect their work,' and calls adding the file 'a quick win' (Appendix D.1). This is a leading prompt rather than neutral information, so 75% plausibly overstates the adoption intent that a neutral framing would elicit; Q27's preamble ('most companies respect it') raises the same concern for the 77% distrust figure. Please (i) report the un-primed Q22/Q23 results (the 97% desire-for-a-button result) side by side with the primed Q26 result, (ii) add a sensitivity analysis, e.g., re-scoring only the 'very likely' responses, and (iii) add a priming caveat to the Limitations section and hedge the abstract accordingly.","section":"§4.2, Appendix D.1 (and Abstract)"},{"comment":"The 59% awareness figure, which anchors the title and abstract, is estimated from a snowball convenience sample: 203 artists recruited through the authors' social circles, professional Discord channels, and participant referrals (Section 4.1), with 109 of 203 respondents based in North America (89 in the US) and heavy concentration in illustration (163), digital 2D (143), and character/creature design (99) (Appendix D.2). The awareness question (Q24) is asked before the robots.txt description, so it is not affected by the priming in the previous comment, but the sample's representativeness is still unestablished, and recruiting through AI-activism-adjacent professional networks could plausibly bias awareness and sentiment in unquantified ways. I credit the bogus-item attention check ('nearest diffusion tree') and the geographic caveat in Section 7, but the paper should add a confidence interval for the 59% estimate, an explicit discussion of the direction and plausible magnitude of selection bias, and should scope the abstract's generalization to the surveyed population.","section":"§4.1, §7, Appendix D.2 (and Abstract)"},{"comment":"The '20 of 23' third-party crawler result is an existence proof for a large class of non-compliant assistant crawlers, and the measurement design (triggered fetches to sites under the authors' control) is sound. However, Section 4.2 generalizes this result to 'the majority of AI assistant crawlers do not respect robots.txt,' while the 23 crawlers were obtained by prompting the top-5k GPT-store apps listed on a third-party directory with two specific prompts ('Start action, fetch page: [url]' and 'Get web page content: [url]'). That sampling frame is a self-selected slice of the assistant-crawler ecosystem, likely dominated by small gateway services, and it is not demonstrably representative of all AI assistant crawlers. Please state this sampling frame explicitly in Section 5.2.2, and soften the class-level generalization in Section 4.2 and the abstract to 'the 23 third-party assistant crawlers we were able to trigger.'","section":"§5.2.2 and §4.2"}],"minor_comments":[{"comment":"The abstract refers to '203 professional artists,' but only 136 of 203 respondents (67%) self-identify as professional in Q1; the remaining third earn income from art but declined the label, so please make the wording precise, e.g., '203 artists, of whom 67% self-identify as professional.'","section":"Abstract and §4.1"},{"comment":"The 100% '% Disallow AI' value for Carbonmade reflects the provider's default robots.txt file (which disallows GPTBot and CCBot) rather than artist behavior; the text explains this, but the table should carry an explicit footnote so the column is not read as a measure of artist agency.","section":"Table 2"},{"comment":"'A significant majority (185, 93%)' appears inconsistent as written: 185/203 is 91%, so the 93% is presumably computed over the 'over 97%' subset who expressed a desire to block; please state the denominator explicitly.","section":"§4.2"},{"comment":"The noai/noimageai check (17 and 16 sites among the Tranco top-10k of October 2024) is reported without methodological detail; a sentence describing the fetch procedure and the measurement date would make this reproducible.","section":"§2.2"},{"comment":"The criterion for identifying the 2,018 (20%) Tranco top-10k sites hosted on Cloudflare is not described; please state the detection method (e.g., DNS, cdn-cgi endpoints, or response headers).","section":"§6.3"},{"comment":"The caption mixes per-period semantics ('removed restrictions... in each time period') with a cumulative curve ('explicitly allowed'); please clarify the semantics in the caption or legend so the two curves are not misread as comparable.","section":"Figure 4"},{"comment":"The two test sites share one IP address, but the paper does not state which robots.txt configuration (the wildcard site or the per-agent site) the active GPT-app fetches targeted; since the per-agent file disallows only the 24 listed agents, this distinction matters for interpreting the 20/23 result, so please state it explicitly.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the network-measurement core of this paper is strong and I would expect it to clear the IMC bar once the survey-based generalizations are tightened, because the abstract's 'strong demand' and 'technical awareness' headlines rest on a convenience-snowball sample with a leading prompt (Appendix D.1). If the authors do not add the sensitivity analysis and hedging described in my major comments, the version of record would overstate the artist-side findings. One transparency point worth having the authors address: the author team includes the creators of Glaze/Nightshade and the survey recruits through artist communities adjacent to AI-protection activism; this is a plausible source of selection bias that the paper should discuss on its own terms. Finally, the paper's claim that the parser used in prior work [70] has a roughly 10% error rate is a serious public accusation; the paper states the authors were notified and the issue was corrected, but I would ask the editor to confirm this is adequately substantiated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead this paper if you work on web content controls or AI data governance. The headline is that the network-measurement half is careful and the artist-survey half is suggestive but shaky. The paper does several distinct things: a two-year longitudinal analysis of robots.txt across 40k stable domains, a 203-artist survey, a hosting-provider control audit, direct tests of which AI crawlers respect robots.txt, and a grey-box evaluation of Cloudflare's Block AI Bots. The longitudinal analysis is well-validated: they use Google's parser, cross-check Common Crawl against the Internet Archive, and find real trends (an initial surge in disallow rules, then some reversals tied to licensing deals). The crawler compliance test is a direct, falsifiable measurement on the authors' own sites; finding that most big-company data crawlers respect robots.txt while 20 of 23 third-party AI assistant crawlers do not is a concrete, useful result. The Cloudflare inference method is also careful: paired grey-box tests against ground truth from their own dashboard.\n\nThe soft spot is the user study. The 59% awareness figure and the 75% adoption intent anchor the abstract, but the sample is a convenience snowball: 203 artists, mostly North American, mostly illustrators/digital artists, recruited via social circles and Discord. That alone would make the numbers descriptive, not generalizable. Worse, Appendix D.1 shows that participants who had not heard of robots.txt were shown a description stating 'over 90% of artists don't realize...' and calling robots.txt an 'easy way' and 'quick win' before being asked about adoption intent. That is a leading prompt, and it likely inflates the 75% figure. The authors note the convenience sample in their Limitations section, but they do not address the priming.\n\nThe 20/23 crawler count is based on a list from GPTStore; the paper acknowledges it is limited to apps reachable through ChatGPT's store, so 'most AI assistant crawlers' should be read as 'most third-party assistant crawlers we could trigger.' That is a minor overstatement, not a flaw in the measurement itself.\n\nVerdict: the paper deserves a serious referee. The measurement components are reproducible (code is on GitHub), the results are useful, and the artist perspective is a genuine gap in the literature. I would ask for a revision that either removes the leading description from the survey or re-analyzes the adoption-intent numbers without it, and that adds explicit caveats about the sample. With those fixes, I would be comfortable citing it for the crawler-compliance and Cloudflare results.\n\nRecommendation: engage with it. Send it to review.","headline":"A careful measurement study of AI crawler controls whose artist survey is the soft spot; the network results are solid and worth citing.","tokens_in":34068,"tokens_out":4192,"would_cite":true,"duration_ms":40357,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the existing voluntary web controls for AI crawling fail individual artists in three ways: most surveyed artists have never heard of robots.txt, most hosting platforms give them no way to edit it, and most…","keywords":["robots.txt","AI crawlers","content creators","visual artists","web content control","active blocking","generative AI training data","user study"],"falsifier":"Re-run the artist survey with a probability-sampled panel of working artists and a neutral description of robots.txt that does not call it an 'easy way' or 'quick win'; if awareness exceeds, say, 80% or a majority of artists report being able to edit robots.txt through their host, the paper's headline gap collapses. Independently, instrument a fresh set of 100 AI-assistant endpoints over three months and count how many fetch robots.txt before requesting content; if most fetch and obey, the 20-of-23 non-compliance result is a snapshot, not a stable property.","tokens_in":32923,"feed_emoji":"🤖","tokens_out":5947,"duration_ms":56709,"temperature":0.7,"pith_summary":"The paper tries to establish that the web's principal tool for refusing AI crawling—robots.txt—does not currently work for the individual creators who need it most. On the evidence presented, most professional visual artists have never heard of robots.txt (59% of 203 surveyed), most of the hosting platforms artists actually use give them no way to edit it, and a majority of the third-party AI assistant crawlers tested (20 of 23) do not even fetch the file. The consequence would be that the current voluntary, honor-based system leaves artists with little real protection, and that stronger or default-on protections need to be built into platforms. This matters because the same measurement shows strong demand: once told about robots.txt, 75% of previously unaware artists said they would adopt it.","feed_headline":"Most artists can't use robots.txt; AI assistants ignore it","feed_subtitle":"A 203-artist survey and six months of crawl tests show the web's main opt-out is invisible, locked away, or ignored.","key_machinery":"The load-bearing object is robots.txt, the Robots Exclusion Protocol (RFC 9309): a plain-text file placed at a site's root that names user agents (e.g., GPTBot, ChatGPT-User) and disallows or allows paths. It is an honor-based signal, not an enforceable access control, so its value depends on each crawler's willingness to fetch and obey it. The paper's machinery for testing that willingness is a pair of researcher-controlled honey-pot websites (one disallowing all crawlers, one disallowing each AI user agent individually) whose server logs reveal which crawlers fetch robots.txt and then still request content, plus active triggering of ChatGPT-store GPT apps to force third-party assistant crawlers to visit. For the artist-side claims, the machinery is a 203-respondent survey with an embedded attention-check term and a DNS-level census of 1,182 artist sites to see which hosting providers expose robots.txt controls.","core_discovery":"On the paper's own terms, the discovery is a three-part gap between what artists want and what the current technical stack delivers. Awareness: 59% of the surveyed professional artists had never heard of robots.txt, and average self-rated familiarity with it was 1.99 on a 5-point scale, just above the fake control term. Agency: among over 1,100 artist websites, over 78% are hosted on eight platforms; four of the top eight provide no way for users to modify robots.txt, and even where a control exists (paid Wix, Squarespace toggle) uptake is near zero or 17%, far below the 75% expressed intention. Efficacy: in six months of passive and active measurement on honey-pot websites, all major AI data crawlers except ByteDance's Bytespider respected robots.txt, but 20 of 23 third-party AI assistant crawlers never fetched robots.txt at all, so they cannot be bound by it. Active blocking via Cloudflare blocks more aggressively (17 AI user agents) but was enabled on only about 5.7% of Cloudflare-hosted top-10k sites and does not cover every AI crawler; the paper concludes that robots.txt and active blocking are complementary, not interchangeable.","pith_inferences":["The results imply that robots.txt is being asked to do two incompatible jobs at once: expressing legal preference and enforcing technical access; regulations like the EU AI Act that condition copyright carve-outs on respecting robots.txt inherit all of the protocol's gaps unless they define what counts as respect.","A natural extension would separate deliberate policy from implementation bugs by checking whether the 20 non-fetching crawlers also ignore HTTP caching directives or fetch robots.txt on retry.","If platform-level toggles became the norm, the measured 59% unawareness would likely drop quickly; a testable prediction is that making a Squarespace-style switch default-on would raise effective protection, though it might also affect artists' search discoverability."],"forward_implications":["If 20 of 23 third-party AI assistant crawlers ignore robots.txt, creators cannot rely on the protocol to stop real-time retrieval of their work by AI assistants; the gap is not the big training crawlers but user-triggered fetching.","If most hosting platforms do not expose robots.txt, then protective intent must be implemented at the platform level (for example, a default-on switch), not by individual artists.","If only 17% of Squarespace artists enable an AI-blocking toggle despite 75% stated intent, discoverability and clear communication of the control matter as much as the control itself.","If active blocking cannot be tuned for dual-purpose crawlers like Googlebot, robots.txt remains necessary even where firewall-level blocking is in place.","If licensing deals make publishers remove robots.txt restrictions, expressed opt-out intent is reversible and economic, not a fixed preference."],"supporting_citations":[{"why":"Supplies the historic robots.txt corpus from Common Crawl snapshots used for the longitudinal adoption analysis.","marker":"[24]"},{"why":"Provides the Dark Visitors list of AI user agents that defines which crawlers count as AI throughout the study.","marker":"[113]"},{"why":"Prior large-scale robots.txt study whose parser errors and predictions this work corrects and extends.","marker":"[70]"},{"why":"RFC 9309 defines the Robots Exclusion Protocol, the central object under test.","marker":"[61]"},{"why":"Method for detecting user-agent-based blocking that the active-blocking measurement adapts.","marker":"[88]"},{"why":"Cloudflare's announcement of the Block AI Bots feature that the case study evaluates.","marker":"[13]"},{"why":"Thematic analysis approach used to code the open-ended survey responses.","marker":"[15]"},{"why":"Cloudflare's Verified Bots list used to infer the operation of the Definitely Automated managed rule.","marker":"[21]"}],"fun_headline_variants":["Artists want to block AI crawlers, but most can't","20 of 23 AI assistants never fetch robots.txt","59% of artists never heard of robots.txt","Even where robots.txt works, it's locked away","Cloudflare blocks more AI, but few sites enable it"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central statistics assume that the 203 artists recruited through the authors' professional networks and Discord channels represent professional visual artists broadly, and that the favorable description of robots.txt shown before the adoption questions did not inflate stated willingness to use it.","fun_headline_variants_meta":{"raw":{"variants":["Artists want to block AI crawlers, but most can't","20 of 23 AI assistants never fetch robots.txt","59% of artists never heard of robots.txt","Even where robots.txt works, it's locked away","Cloudflare blocks more AI, but few sites enable it"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000377,"raw_usage":{"total_tokens":2042,"prompt_tokens":1014,"completion_tokens":1028,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":949}},"tokens_in":630,"tokens_out":1028,"duration_ms":10020,"temperature":1.0,"reasoning_tokens":949,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:30:39.613196+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the artist survey with a probability-sampled panel of working artists and a neutral description of robots.txt that does not call it an 'easy way' or 'quick win'; if awareness exceeds, say, 80% or a majority of artists report being able to edit robots.txt through their host, the paper's headline gap collapses. Independently, instrument a fresh set of 100 AI-assistant endpoints over three months and count how many fetch robots.txt before requesting content; if most fetch and obey, the 20-of-23 non-compliance result is a snapshot, not a stable property.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Dark Visitors list of AI user agents that defines which crawlers count as AI throughout the study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Method for detecting user-agent-based blocking that the active-blocking measurement adapts."}],"review_version":1}