Pith. sign in

REVIEW 3 major objections 7 minor 1 cited by

Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers

T0 review · 3 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper claims that the existing voluntary web controls for AI crawling fail individual artists in three ways: most surveyed artists have never heard of robots.txt, most hosting platforms give them no way to edit it, and most…

desk verdict A careful measurement study of AI crawler controls whose artist survey is the soft spot; the network results are solid and worth citing. read the letter →

arxiv 2411.15091 v2 pith:IBO5YPYA submitted 2024-11-22 cs.HC

classification cs.HC
keywords robots.txtAIcrawlerscontentcreatorsvisualartistswebcontrolactiveblockinggenerativetrainingdatauserstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the web's principal tool for refusing AI crawling—robots.txt—does not currently work for the individual creators who need it most. On the evidence presented, most professional visual artists have never heard of robots.txt (59% of 203 surveyed), most of the hosting platforms artists actually use give them no way to edit it, and a majority of the third-party AI assistant crawlers tested (20 of 23) do not even fetch the file. The consequence would be that the current voluntary, honor-based system leaves artists with little real protection, and that stronger or default-on protections need to be built into platforms. This matters because the same measurement shows strong demand: once told about robots.txt, 75% of previously unaware artists said they would adopt it.

What carries the argument

The load-bearing object is robots.txt, the Robots Exclusion Protocol (RFC 9309): a plain-text file placed at a site's root that names user agents (e.g., GPTBot, ChatGPT-User) and disallows or allows paths. It is an honor-based signal, not an enforceable access control, so its value depends on each crawler's willingness to fetch and obey it. The paper's machinery for testing that willingness is a pair of researcher-controlled honey-pot websites (one disallowing all crawlers, one disallowing each AI user agent individually) whose server logs reveal which crawlers fetch robots.txt and then still request content, plus active triggering of ChatGPT-store GPT apps to force third-party assistant crawlers to visit. For the artist-side claims, the machinery is a 203-respondent survey with an embedded attention-check term and a DNS-level census of 1,182 artist sites to see which hosting providers expose robots.txt controls.

What would settle it

Re-run the artist survey with a probability-sampled panel of working artists and a neutral description of robots.txt that does not call it an 'easy way' or 'quick win'; if awareness exceeds, say, 80% or a majority of artists report being able to edit robots.txt through their host, the paper's headline gap collapses. Independently, instrument a fresh set of 100 AI-assistant endpoints over three months and count how many fetch robots.txt before requesting content; if most fetch and obey, the 20-of-23 non-compliance result is a snapshot, not a stable property.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a three-part gap between what artists want and what the current technical stack delivers. Awareness: 59% of the surveyed professional artists had never heard of robots.txt, and average self-rated familiarity with it was 1.99 on a 5-point scale, just above the fake control term. Agency: among over 1,100 artist websites, over 78% are hosted on eight platforms; four of the top eight provide no way for users to modify robots.txt, and even where a control exists (paid Wix, Squarespace toggle) uptake is near zero or 17%, far below the 75% expressed intention. Efficacy: in six months of passive and active measurement on honey-pot websites, all major AI data crawlers except ByteDance's Bytespider respected robots.txt, but 20 of 23 third-party AI assistant crawlers never fetched robots.txt at all, so they cannot be bound by it. Active blocking via Cloudflare blocks more aggressively (17 AI user agents) but was enabled on only about 5.7% of Cloudflare-hosted top-10k sites and does not cover every AI crawler; the paper concludes that robots.txt and active blocking are complementary, not interchangeable.

Load-bearing premise

The central statistics assume that the 203 artists recruited through the authors' professional networks and Discord channels represent professional visual artists broadly, and that the favorable description of robots.txt shown before the adoption questions did not inflate stated willingness to use it.

Editorial extensions

If this is right

  • If 20 of 23 third-party AI assistant crawlers ignore robots.txt, creators cannot rely on the protocol to stop real-time retrieval of their work by AI assistants; the gap is not the big training crawlers but user-triggered fetching.
  • If most hosting platforms do not expose robots.txt, then protective intent must be implemented at the platform level (for example, a default-on switch), not by individual artists.
  • If only 17% of Squarespace artists enable an AI-blocking toggle despite 75% stated intent, discoverability and clear communication of the control matter as much as the control itself.
  • If active blocking cannot be tuned for dual-purpose crawlers like Googlebot, robots.txt remains necessary even where firewall-level blocking is in place.
  • If licensing deals make publishers remove robots.txt restrictions, expressed opt-out intent is reversible and economic, not a fixed preference.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The results imply that robots.txt is being asked to do two incompatible jobs at once: expressing legal preference and enforcing technical access; regulations like the EU AI Act that condition copyright carve-outs on respecting robots.txt inherit all of the protocol's gaps unless they define what counts as respect.
  • A natural extension would separate deliberate policy from implementation bugs by checking whether the 20 non-fetching crawlers also ignore HTTP caching directives or fetch robots.txt on retry.
  • If platform-level toggles became the norm, the measured 59% unawareness would likely drop quickly; a testable prediction is that making a Squarespace-style switch default-on would raise effective protection, though it might also affect artists' search discoverability.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This paper studies whether the web-control mechanisms available to individual content creators — robots.txt, noai meta tags, and active blocking via reverse proxies — are known to creators, available to them through their hosting platforms, and effective against AI crawlers. The paper combines four complementary measurements: (1) a longitudinal analysis of robots.txt adoption across 40,455 domains that appear in the Tranco top-100k in every month from October 2022 to October 2024, using Common Crawl snapshots cross-validated against the Internet Archive; (2) a user study of 203 artists, plus a measurement of 1,182 artist websites and the control that their eight most common hosting providers expose; (3) passive and active tests of whether 24 AI-related user agents respect robots.txt on two sites under the authors' control; and (4) an evaluation of active blocking, including an inferred behavior model of Cloudflare's 'Block AI Bots' feature and an estimate of its deployment across the Tranco top-10k. The headline findings are that 59% of surveyed artists had never heard of robots.txt, that most hosting platforms do not let artists edit it, that most large AI companies' data crawlers respect robots.txt while 20 of 23 triggered third-party AI assistant crawlers do not fetch it, and that active blocking is more enforceable but sparsely deployed and incomplete in coverage.

Significance. If the measurements hold, this is the most complete empirical map to date of the gap between what individual content creators want from anti-AI-crawling tools and what the current Web actually provides, and it is likely to become the reference point for the state of AI-crawler blocking circa 2024-2025. The network-measurement components are genuinely strong: the robots.txt parser is validated against RFC 9309 edge cases, the Common Crawl longitudinal data is cross-checked against the Internet Archive and an independent crawl, the crawler-compliance tests are direct and falsifiable (20 of 23 triggered third-party assistant crawlers never fetched robots.txt), and the Cloudflare setting is inferred through paired grey-box tests with the authors' own account as ground truth. The paper also ships public code and data via a GitHub repository. There are no fitted parameters and the central claims do not reduce to an input, so circularity is not a concern.

major comments (3)
  1. [§4.2, Appendix D.1 (and Abstract)] The 75% intention-to-adopt figure is load-bearing for the abstract's 'strong demand for tools like robots.txt,' but it is measured by Q26, which is asked immediately after a description that asserts 'over 90% of artists don't realize they can use a simple tool called robots.txt,' calls the tool 'an easy way for artists to protect their work,' and calls adding the file 'a quick win' (Appendix D.1). This is a leading prompt rather than neutral information, so 75% plausibly overstates the adoption intent that a neutral framing would elicit; Q27's preamble ('most companies respect it') raises the same concern for the 77% distrust figure. Please (i) report the un-primed Q22/Q23 results (the 97% desire-for-a-button result) side by side with the primed Q26 result, (ii) add a sensitivity analysis, e.g., re-scoring only the 'very likely' responses, and (iii) add a priming caveat to the Limitations section and hedge the abstract accordingly.
  2. [§4.1, §7, Appendix D.2 (and Abstract)] The 59% awareness figure, which anchors the title and abstract, is estimated from a snowball convenience sample: 203 artists recruited through the authors' social circles, professional Discord channels, and participant referrals (Section 4.1), with 109 of 203 respondents based in North America (89 in the US) and heavy concentration in illustration (163), digital 2D (143), and character/creature design (99) (Appendix D.2). The awareness question (Q24) is asked before the robots.txt description, so it is not affected by the priming in the previous comment, but the sample's representativeness is still unestablished, and recruiting through AI-activism-adjacent professional networks could plausibly bias awareness and sentiment in unquantified ways. I credit the bogus-item attention check ('nearest diffusion tree') and the geographic caveat in Section 7, but the paper should add a confidence interval for the 59% estimate, an explicit discussion of the direction and plausible magnitude of selection bias, and should scope the abstract's generalization to the surveyed population.
  3. [§5.2.2 and §4.2] The '20 of 23' third-party crawler result is an existence proof for a large class of non-compliant assistant crawlers, and the measurement design (triggered fetches to sites under the authors' control) is sound. However, Section 4.2 generalizes this result to 'the majority of AI assistant crawlers do not respect robots.txt,' while the 23 crawlers were obtained by prompting the top-5k GPT-store apps listed on a third-party directory with two specific prompts ('Start action, fetch page: [url]' and 'Get web page content: [url]'). That sampling frame is a self-selected slice of the assistant-crawler ecosystem, likely dominated by small gateway services, and it is not demonstrably representative of all AI assistant crawlers. Please state this sampling frame explicitly in Section 5.2.2, and soften the class-level generalization in Section 4.2 and the abstract to 'the 23 third-party assistant crawlers we were able to trigger.'
minor comments (7)
  1. [Abstract and §4.1] The abstract refers to '203 professional artists,' but only 136 of 203 respondents (67%) self-identify as professional in Q1; the remaining third earn income from art but declined the label, so please make the wording precise, e.g., '203 artists, of whom 67% self-identify as professional.'
  2. [Table 2] The 100% '% Disallow AI' value for Carbonmade reflects the provider's default robots.txt file (which disallows GPTBot and CCBot) rather than artist behavior; the text explains this, but the table should carry an explicit footnote so the column is not read as a measure of artist agency.
  3. [§4.2] 'A significant majority (185, 93%)' appears inconsistent as written: 185/203 is 91%, so the 93% is presumably computed over the 'over 97%' subset who expressed a desire to block; please state the denominator explicitly.
  4. [§2.2] The noai/noimageai check (17 and 16 sites among the Tranco top-10k of October 2024) is reported without methodological detail; a sentence describing the fetch procedure and the measurement date would make this reproducible.
  5. [§6.3] The criterion for identifying the 2,018 (20%) Tranco top-10k sites hosted on Cloudflare is not described; please state the detection method (e.g., DNS, cdn-cgi endpoints, or response headers).
  6. [Figure 4] The caption mixes per-period semantics ('removed restrictions... in each time period') with a cumulative curve ('explicitly allowed'); please clarify the semantics in the caption or legend so the two curves are not misread as comparable.
  7. [§5.1] The two test sites share one IP address, but the paper does not state which robots.txt configuration (the wildcard site or the per-agent site) the active GPT-app fetches targeted; since the per-agent file disallows only the 24 listed agents, this distinction matters for interpreting the 20/23 result, so please state it explicitly.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: every load-bearing claim rests on direct external measurements or survey data, with no fitted parameter renamed as a prediction.

full rationale

This is an observational measurement paper, not a derivation. The abstract's claim that strong demand for robots.txt-like tools is constrained by awareness, agency, and efficacy rests on three independent empirical legs: (1) a survey of 203 artists reporting that 59% had never heard of robots.txt, (2) a measurement of hosting providers showing most do not expose robots.txt controls, and (3) controlled-website experiments showing that 20 of 23 third-party AI assistant crawlers do not fetch robots.txt. Each of these is measured directly against external benchmarks (Common Crawl, Tranco, Dark Visitors, controlled websites, Cloudflare dashboards) rather than derived from an input assumption. The only self-citations, Glaze and Nightshade, appear as survey response options and background context; they are not used to justify the paper's central measurements or conclusions. The survey's robots.txt description is a potential validity concern because it may prime adoption answers, but priming is a bias in measurement, not circularity: the paper does not define its conclusion in terms of the prompt, and the 59% awareness figure is collected before the prompt is shown. No equation reduces to an input, no fitted parameter is relabeled as a prediction, and no uniqueness theorem is imported from the authors' prior work. The paper is self-contained against external evidence, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

No free parameters are fitted in this paper. The measurements rest on several data-source and inference assumptions, the most important being the Dark Visitors user-agent list, Google's parser, the controlled-site observation window, and the representativeness of the snowball survey.

assumptions (6)
  • domain assumption The Dark Visitors list of AI user agents is a comprehensive and correct enumeration of AI crawlers.
    Section 3.1 uses this blog-maintained list (cross-validated with Longpre et al. [70]) to define the 24 user agents in Table 1; all longitudinal, compliance, and active-blocking percentages are computed with respect to this list.
  • domain assumption Google's robots.txt parser correctly interprets RFC 9309 semantics for the studied files.
    Section 3.1 relies on Google's parser, validated by manual review of 100 files and edge-case tests in Appendix B.2; any parser errors would shift the measured restriction percentages.
  • domain assumption Crawler compliance can be judged from observed fetches on two simple controlled websites over six months.
    Section 5.1/5.2 infers respect or non-respect from whether a crawler fetched robots.txt and then content; crawlers that never visited the test sites are marked '-' in Table 1, so the compliance rates describe only observed crawlers.
  • domain assumption User-agent-based differences in HTTP responses correctly identify active blocking of AI crawlers.
    Section 6.1 detects active blocking by comparing responses with a standard user agent versus Anthropic user agents; behavior-based or IP-based blocking is explicitly unmeasured, and the authors treat the result as a lower bound.
  • domain assumption Cloudflare's Block AI Bots setting can be inferred from HTTP responses to a small set of user agents on a headless browser.
    Section 6.3 uses gray-box testing on the authors' own Cloudflare account to map response signatures to settings, then applies that mapping to 2,018 third-party sites; custom firewall rules could misclassify some sites.
  • domain assumption Survey participants' self-reported familiarity reflects actual knowledge.
    Section 4.1 and Appendix D.2 use self-reported familiarity and low ratings on a bogus item as evidence of non-random responding; common-method bias and social desirability remain.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers." pith.science (2026). https://pith.science/paper/IBO5YPYA

@misc{pith2026241115091,
  author       = {Pith},
  title        = {Pith review of: Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IBO5YPYA}},
  note         = {Machine review of arXiv:2411.15091}
}
read the original abstract

The success of generative AI relies heavily on training on data scraped through extensive crawling of the Internet, a practice that has raised significant copyright, privacy, and ethical concerns. While few measures are designed to resist a resource-rich adversary determined to scrape a site, crawlers can be impacted by a range of existing tools such as robots.txt, NoAI meta tags, and active crawler blocking by reverse proxies. In this work, we seek to understand the ability and efficacy of today's networking tools to protect content creators against AI-related crawling. For targeted populations like human artists, do they have the technical knowledge and agency to utilize crawler-blocking tools such as robots.txt, and can such tools be effective? Using large scale measurements and a targeted user study of 203 professional artists, we find strong demand for tools like robots.txt, but significantly constrained by critical hurdles in technical awareness, agency in deploying them, and limited efficacy against unresponsive crawlers. We further test and evaluate network-level crawler blockers provided by reverse proxies. Despite relatively limited deployment today, they offer stronger protections against AI crawlers, but still come with their own set of limitations.

Figures

Figures reproduced from arXiv: 2411.15091 by the authors.

Figure 1
Figure 1. In this example robots.txt file, Googlebot is al [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Percent of sites that fully disallow at least one AI [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 4
Figure 4. Number of sites that explicitly allow at least one [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: Squarespace provides a user-friendly option for [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 7
Figure 7. Figure 7: Flowchart for inferring the Block AI Bots setting [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Agentic Web Requires New Normative Infrastructure

    cs.CY 2026-06 unverdicted novelty 6.0 of 10

    The web's anti-bot regime should be replaced by a framework that presumptively lets user-authorized AI agents act for their principals, requires platforms to disclose access policies, and permits agent blocking only w...

Reference graph

Works this paper leans on

125 extracted references · 70 canonical work pages · cited by 1 Pith paper

  1. [1]

    Adobe. 2024. Adobe General Terms of Use. https://www.adobe.com/legal/terms. html

  2. [2]

    AIOSEO. 2025. The Best WordPress SEO Plugin and Toolkit. https://aioseo.com/

  3. [3]

    Safinah Ali and Cynthia Breazeal. 2023. Studying Artist Sentiments around AI-generated Artwork. (2023), 13 pages. arXiv:2311.13725 [cs.HC] https: //arxiv.org/abs/2311.13725

  4. [4]

    Weigle, and Michael L

    Yasmin AlNoamany, Michele C. Weigle, and Michael L. Nelson. 2013. Access Patterns for Robots and Humans in Web Archives. InProc. of the 13th ACM/IEEE- CS Joint Conference on Digital Libraries . 339–348

  5. [5]

    Babak Amin Azad, Oleksii Starov, Pierre Laperdrix, and Nick Nikiforakis. 2020. Web runner 2049: Evaluating third-party anti-bot services. In Proc. of 17th Detection of Intrusions and Malware, and Vulnerability Assessment . Springer, 135–159

  6. [6]

    Anthropic. 2024. Does Anthropic crawl data from the web, and how can site owners block the crawler? https://support.anthropic.com/en/articles/8896518- does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block- the-crawler

  7. [7]

    ArtStation. 2025. ArtStation Terms of Service. https://www.artstation.com/tos

  8. [8]

    Stefan Baak. 2024. Training Data for the Price of a Sandwich. https://foundation. mozilla.org/en/research/library/generative-ai-training-data/common-crawl/

Show all 125 references
  1. [9]

    Quan Bai, Gang Xiong, Yong Zhao, and Longtao He. 2014. Analysis and De- tection of Bogus Behavior in Web Crawler Measurement. Procedia Computer Science 31 (2014), 1084–1091

  2. [10]

    Kayleigh Barber. 2024. Future’s Jon Steinberg shares his philosophy on AI content licensing deals — Digiday. https://digiday.com/podcasts/futures-jon- steinberg-shares-his-philosophy-on-ai-content-licensing-deals/

  3. [11]

    Muhammad Ahmad Bashir, Sajjad Arshad, Engin Kirda, William Robertson, and Christo Wilson. 2019. A Longitudinal Analysis of the ads.txt Standard. In Proc. ACM Internet Measurement Conference 2019 . 294–307

  4. [12]

    Lucas Bellaiche, Rohin Shahi, Martin Harry Turpin, Anya Ragnhildstveit, Shawn Sprockett, Nathaniel Barr, Alexander Christensen, and Paul Seli. 2023. Humans versus AI: whether and why we prefer human-created compared to AI-created artwork. Cognitive Research: Principles and Imp...

  5. [13]

    Alex Bocharov, Santiago Varagas, Adam Martinetti, Reid Tatoris, and Carlos Azevedo. 2024. Declare your AIndependence: block AI bots, scrapers and crawlers with a single click — Cloudflare. https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots- scrapers-and-cra...

  6. [14]

    Josiah D Boucher, Gillian Smith, and Yunus Doğan Telliel. 2024. Is Resistance Futile?: Early Career Game Developers, Generative AI, and Ethical Skepticism. In Proc. of CHI Conference on Human Factors in Computing Systems 2024 . 1–13

  7. [15]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative Research in Psychology 3, 2 (2006), 77–101

  8. [16]

    Jack Brewster, Zach Fishman, and Isaiah Glick. 2024. AI Chatbots Are Blocked by 67% of Top News Sites, Relying Instead on Low-Quality Sources — News Guard. https://www.newsguardtech.com/special-reports/67-percent-of-top- news-sites-block-ai-chatbots/

  9. [17]

    Gordon Burtch, Dokyun Lee, and Zhichen Chen. 2023. The consequences of generative AI for UGC and online community engagement. SSRN (2023), 26 pages

  10. [18]

    Carbonmade. 2024. Terms and Conditions of Use. https://carbonmade.com/ terms

  11. [19]

    Eva Cetinic and James She. 2022. Understanding and Creating Art with AI: Review and Outlook. ACM Transactions on Multimedia Computing, Communi- cations, and Applications 18, 2 (2022), 1–22

  12. [20]

    Zi Chu, Steven Gianvecchio, and Haining Wang. 2018. Bot or Human? A Behavior-Based Online Bot Detection System. From Database to Cyber Security: Essays Dedicated to Sushil Jajodia on the Occasion of His 70th Birthday (2018), 432–449

  13. [21]

    Cloudflare. 2024. Verified Bots. https://radar.cloudflare.com/traffic/verified- bots

  14. [22]

    European Commission. 2024. First Draft of the General-Purpose AI Code of Practice published, written by independent experts. https://digital-strategy.ec.europa.eu/en/library/first-draft-general-purpose- ai-code-practice-published-written-independent-experts

  15. [23]

    European Commission. 2024. The AI Act Explorer. https: //artificialintelligenceact.eu/ai-act-explorer/

  16. [24]

    Common Crawl. 2025. Common Crawl — Open Repository of Web Crawl Data. https://commoncrawl.org/

  17. [25]

    davepattern. 2024. DDoS from Anthropic AI. https://www.linode.com/ community/questions/24842/ddos-from-anthropic-ai

  18. [26]

    deninho32. 2024. Claudebot attack. https://www.phpbb.com/community/ viewtopic.php?t=2652265

  19. [27]

    Design and Artists Copyright Society (DACS). 2024. Artificial Intelligence and Artists’ Work: A survey of artists on AI. https://cdn.dacs.org.uk/uploads/ documents/News/Artificial-Intelligence-and-Artists-Work-DACS.pdf. (2024), 29 pages

  20. [28]

    Michael Dinzinger and Michael Granitzer. 2024. A Longitudinal Study of Content Control Mechanisms. In Proc. of the ACM Web Conference . 1382–1387

  21. [29]

    Michael Dinzinger, Florian Heß, and Michael Granitzer. 2024. A Survey of Web Content Control for Generative AI. (2024), 12 pages. arXiv:2404.02309 [cs.IR] https://arxiv.org/abs/2404.02309

  22. [30]

    Frank, Matthew Groh, Laura Herman, Neil Leach, Robert Mahari, Alex “Sandy” Pentland, Olga Russakovsky, Hope Schroeder, and Amy Smith

    Ziv Epstein, Aaron Hertzmann, Memo Akten, Hany Farid, Jessica Fjeld, Mor- gan R. Frank, Matthew Groh, Laura Herman, Neil Leach, Robert Mahari, Alex “Sandy” Pentland, Olga Russakovsky, Hope Schroeder, and Amy Smith

  23. [31]

    Fairly Trained. 2024. Statement on AI training. https://www.aitrainingstatement. org/

  24. [32]

    Richard Fletcher. 2024. How many news websites block AI crawlers? Reuters Institute Factsheets (2024), 7 pages

  25. [33]

    noai" and

    Foundation Web Design & Development. 2022. What is DeviantArt’s new "noai" and "noimageai" meta tag and how to install it. https://www.foundationwebdev. com/2022/11/noai-noimageai-meta-tag-how-to-install/

  26. [34]

    Thompson

    Sheera Frenkel and Stuart A. Thompson. 2023. Not for Machines to Harvest: Data Revolts Break Out Against A.I. — The New York Times. https://www.nytimes. com/2023/07/15/technology/artificial-intelligence-models-chat-data.html

  27. [35]

    Manuel B. Garcia. 2024. The Paradox of Artificial Creativity: Challenges and Opportunities of Generative AI Artistry. Creativity Research Journal (2024), 1–14

  28. [36]

    generosus. 2023. PSA | Bytedance and Bytespider Bots | Recommend Block- ing. https://wordpress.org/support/topic/psa-bytedance-and-bytespider-bots- recommend-blocking/

  29. [37]

    Jonathan Gillham. 2024. Block AI Bots from Crawling Websites Using Robots.txt — Originality.ai. https://originality.ai/ai-bot-blocking

  30. [38]

    Google. 2024. Google Robots.txt Parser and Matcher Library. https://github. com/google/robotstxt

  31. [39]

    Paolo Grigis and Antonella De Angeli. 2024. Playwriting with Large Language Models: Perceived Features, Interaction Strategies and Outcomes. In Proc. the International Conference on Advanced Visual Interfaces 2024 . 1–9

  32. [40]

    Alicia Guo, Shreya Sathyanarayanan, Leijie Wang, Jeffrey Heer, and Amy Zhang

  33. [41]

    Eszter Hargittai. 2009. An Update on Survey Measures of Web-Oriented Digital Literacy. Social Science Computer Review 27, 1 (2009), 130–137

  34. [42]

    Joo-Wha Hong and Nathaniel Ming Curran. 2019. Artificial Intelligence, Artists, and Art: Attitudes Toward Artwork Produced by Humans vs. Artificial Intel- ligence. ACM Transactions on Multimedia Computing, Communications, and Applications 15, 2s (2019), 1–16

  35. [43]

    Xinyi Hou, Yanjie Zhao, and Haoyu Wang. 2024. On the (In)Security of LLM App Stores. (2024), 17 pages. arXiv:2407.08422 [cs.CR] https://arxiv.org/abs/ 2407.08422

  36. [44]

    Hongxian Huang, Runshan Fu, and Anindya Ghose. 2023. Generative AI and Content Creators: Evidence from Digital Art Platforms. SSRN (2023), 41 pages

  37. [45]

    Christos Iliou, Theodoros Kostoulas, Theodora Tsikrika, Vasilis Katos, Stefanos Vrochidis, and Ioannis Kompatsiaris. 2021. Detection of Advanced Web Bots by Combining Web Logs with Mouse Behavioural Biometrics. Digital Threats: Research and Practice 2, 3 (2021), 1–26

  38. [46]

    Christos Iliou, Theodoros Kostoulas, Theodora Tsikrika, Vasilis Katos, Stefanos Vrochidis, and Yiannis Kompatsiaris. 2019. Towards a framework for detecting advanced web bots. In Proc. of 14th International Conference on A vailability, Reliability and Security. 1–10

  39. [47]

    Glassman

    Katy Ilonka Gero, Meera Desai, Carly Schnitzler, Nayun Eom, Jack Cushman, and Elena L. Glassman. 2024. Creative Writers’ Attitudes on Writing as Training Data for Large Language Models. In Proc. of CHI Conference on Human Factors in Computing Systems 2025 . 1–16

  40. [48]

    Imperva. 2024. 2024 Bad Bot Report. https://www.imperva.com/resources/ resource-library/reports/2024-bad-bot-report/

  41. [49]

    Gregoire Jacob, Engin Kirda, Christopher Kruegel, and Giovanni Vigna. 2012. PUBCRAWL: Protecting Users and Businesses from CRAWLers. InProc. of 21st USENIX Security Symposium. 507–522

  42. [50]

    Steve TK Jan, Qingying Hao, Tianrui Hu, Jiameng Pu, Sonal Oswal, Gang Wang, and Bimal Viswanath. 2020. Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data Augmentation. In Proc. of 2020 IEEE Symposium on Security and Privacy . IEEE, 1190–1206

  43. [51]

    Harry H Jiang, Lauren Brown, Jessica Cheng, Mehtab Khan, Abhishek Gupta, Deja Workman, Alex Hanna, Johnathan Flowers, and Timnit Gebru. 2023. AI Art and its Impact on Artists. In Proc. of AAAI ACM Conference on AI, Ethics, and Society 2023. 363–374

  44. [52]

    Hannah Johnston and David Thue. 2024. Understanding Visual Artists’ Values and Attitudes towards Collaboration, Technology, and AI. InProc. 50th Graphics Interface Conference. 1–9

  45. [53]

    Ben Jones, Tzu-Wen Lee, Nick Feamster, and Phillipa Gill. 2014. Automated Detection and Fingerprinting of Censorship Block Pages. InProc. of ACM Internet Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers IMC ’25, October 2...

  46. [54]

    Hugo Jonker, Benjamin Krumnow, and Gabry Vlot. 2019. Fingerprint Surface- Based Detection of Web Bot Detectors. In Proc. of 24th European Symposium on Research in Computer Security . Springer, 586–605

  47. [55]

    Reishiro Kawakami and Sukrit Venkatagiri. 2024. The Impact of Generative AI on Artists. In Proc. of 16th Conference on Creativity & Cognition . 79–82

  48. [56]

    Paul Keller. 2023. Protecting Creatives or Impeding Progress? Ma- chine learning and the EU copyright framework — Open Future. https://openfuture.eu/blog/protecting-creatives-or-impeding-progress/

  49. [57]

    Kate Knibbs. 2024. Condé Nast Signs Deal With OpenAI — Wired. https: //www.wired.com/story/conde-nast-openai-deal/

  50. [58]

    Kate Knibbs. 2024. The Race to Block OpenAI’s Scraping Bots Is Slowing Down — Wired. https://www.wired.com/story/open-ai-publisher-deals-scraping-bots/

  51. [59]

    Robb Knight. 2024. Perplexity AI Is Lying about Their User Agent. https: //rknight.me/blog/perplexity-ai-is-lying-about-its-user-agent/

  52. [60]

    Santanu Kolay, Paolo D’Alberto, Ali Dasdan, and Arnab Bhattacharjee. 2008. A Larger Scale Study of Robots.txt. In Proc. of 17th World Wide Web Conference . 1171–1172

  53. [61]

    Martijn Koster, Gary Illyes, Henner Zeller, and Lizzi Sassman. 2022. Robots Exclusion Protocol. RFC 9309. https://doi.org/10.17487/RFC9309

  54. [62]

    Shinil Kwon, Young-Gab Kim, and Sungdeok Cha. 2012. Web robot detection based on pattern-matching technique. Journal of Information Science 38, 2 (2012), 118–126

  55. [63]

    IAB Tech Lab. 2025. Global Privacy Protocol. https://iabtechlab.com/gpp/

  56. [64]

    Rita Latikka, Jenna Bergdahl, Nina Savela, and Atte Oksanen. 2023. AI as an Artist? A Two-Wave Survey Study on Attitudes Toward Using Artificial Intelligence in Art. Poetics 101 (2023), 11 pages

  57. [65]

    Junsup Lee, Sungdeok Cha, Dongkun Lee, and Hyungkyu Lee. 2009. Classi- fication of web robots: an empirical study based on over one billion requests. Computers & Security 28, 8 (2009), 795–802

  58. [66]

    Jie Li, Hancheng Cao, Laura Lin, Youyang Hou, Ruihao Zhu, and Abdallah El Ali. 2024. User Experience Design Professionals’ Perceptions of Generative Artificial Intelligence. In Proc. of CHI Conference on Human Factors in Computing Systems 2024. 1–18

  59. [67]

    Xigao Li, Babak Amin Azad, Amir Rahmati, and Nick Nikiforakis. 2021. Good Bot, Bad Bot: Characterizing Automated Browsing Activity. InProc. of 2021 IEEE Symposium on Security and Privacy . IEEE, 1589–1605

  60. [68]

    Baker & Hostetler LLP. 2024. Case Tracker: Artificial Intelligence, Copyrights and Class Actions. https://www.bakerlaw.com/services/artificial-intelligence- ai/case-tracker-artificial-intelligence-copyrights-and-class-actions/

  61. [69]

    Joseph Saveri Law Firm LLP. 2023. Class Action Filed Against Stability AI, Midjourney, and DeviantArt for DMCA Violations, Right of Publicity Violations, Unlawful Competition, Breach of TOS. https://www.prnewswire.com/news-releases/class-action-filed-against- stability-ai-midj...

  62. [70]

    Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderin- wale, William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh, Manuel Cherep, Ahmad Anis, An Dinh, Caroline Chitongo, D...

  63. [71]

    Lourenço and Orlando O

    Anália G. Lourenço and Orlando O. Belo. 2006. Catching Web Crawlers in the Act. In Proc. of 6th International Conference on Web Engineering . 265–272

  64. [72]

    Juniper Lovato, Julia Witte Zimmerman, Isabelle Smith, Peter Dodds, and Jen- nifer L. Karson. 2024. Foregrounding Artist Opinions: A Survey Study on Transparency, Ownership, and Fairness in AI Generative Art. In Proc. of AAAI ACM Conference on AI, Ethics, and Society 2024 . 905–916

  65. [73]

    Abhijeeth Madhu. 2023. Survey Reveals 9 out of 10 Artists Believe Current Copyright Laws are Outdated in the Age of Generative AI Technology — Book An Artist Blog. https://bookanartist.co/blog/2023-artists-survey-on-ai- technology/

  66. [74]

    Bron Maher. 2024. Revealed: Which of the top 100 UK and US news websites are blocking AI crawlers — PressGazette. https://pressgazette.co.uk/platforms/ news-sites-block-ai-web-crawlers-chatgpt-google/

  67. [75]

    Meta. 2024. Meta Web Crawlers. https://developers.facebook.com/docs/sharing/ webmasters/web-crawlers/

  68. [76]

    Elz˙e Sigut˙e Mikalonyt˙e and Markus Kneer. 2022. Can Artificial Intelligence Make Art?: Folk Intuitions as to whether AI-driven Robots Can Be Viewed as Artists and Produce Art. ACM Transactions on Human-Robot Interaction 11, 4 (2022), 1–19

  69. [77]

    Cullen Miller. 2023. ai.txt: A new way for websites to set permissions for AI — Spawning. https://spawning.substack.com/p/aitxt-a-new-way-for-websites-to- set

  70. [78]

    Piotr Mirowski, Juliette Love, Kory Mathewson, and Shakir Mohamed. 2024. A Robot Walks into a Bar: Can Language Models Serve as Creativity Support Tools for Comedy? An Evaluation of LLMs’ Humour Alignment with Comedians. In Proc. of the ACM Conference on Fairness, Accountabili...

  71. [79]

    Martin Monperrus. 2024. crawler-user-agents. https://github.com/monperrus/ crawler-user-agents

  72. [80]

    Alexis Newton and Kaustubh Dhole. 2023. Is AI Art Another Industrial Rev- olution in the Making? (2023), 6 pages. arXiv:2301.05133 [cs.AI] https: //arxiv.org/abs/2301.05133

  73. [81]

    Republic of Singapore. 2021. Singapore Copyright Act of 2021: Section 244 (English Translation). https://sso.agc.gov.sg/Acts-Supp/22-2021/Published/ ?ProvIds=pr244-

  74. [82]

    OpenAI. 2023. Our approach to AI safety. https://openai.com/index/our- approach-to-ai-safety/

  75. [83]

    Karla Ortiz. 2024. Why AI Models are not inspired like humans. https: //www.kortizblog.com/blog/why-ai-models-are-not-inspired-like-humans

  76. [84]

    Stack Overflow. 2024. Stack Overflow and OpenAI Partner to Strengthen the World’s Most Popular Large Language Models. https://stackoverflow.co/ company/press/archive/openai-partnership

  77. [85]

    palewire. 2025. Who blocks OpenAI, Google AI and Common Crawl? https: //palewi.re/docs/news-homepages/openai-gptbot-robotstxt.html

  78. [86]

    Sungjin Park. 2024. The work of art in the age of generative AI: aura, liberation, and democratization. AI & Society (2024), 1–10

  79. [87]

    Perplexity. 2025. Perplexity Crawlers. https://docs.perplexity.ai/guides/bots

  80. [88]

    Kien Pham, Aécio Santos, and Juliana Freire. 2016. Understanding Website Behavior based on User Agent. InProc. of 39th ACM SIGIR Conference on Research and Development in Information Retrieval . 1053–1056

  81. [89]

    Tara Poteat and Frank Li. 2021. Who You Gonna Call? An Empirical Evaluation of Website security.txt Deployment. In Proc. of ACM Internet Measurement Conference 2021. 526–532

  82. [90]

    Heila Precel, Allison McDonald, Brent Hecht, and Nicholas Vincent. 2024. A Canary in the AI Coal Mine: American Jews May Be Disproportionately Harmed by Intellectual Property Dispossession in Large Language Model Training. In Proc. of CHI Conference on Human Factors in Computi...

  83. [91]

    PRNewswire. 2024. Dotdash Meredith Announces Strategic Partnership with OpenAI, Bringing Iconic Brands and Trusted Content to ChatGPT. https://dotdashmeredith.mediaroom.com/2024-05-07-Dotdash-Meredith- Announces-Strategic-Partnership-with-OpenAI,-Bringing-Iconic-Brands- and-Tr...

  84. [92]

    Gayatri Raman and Erin Brady. 2024. Exploring Use and Perceptions of Genera- tive AI Art Tools by Blind Artists. (2024), 4 pages. arXiv:2409.08226 [cs.HC] https://arxiv.org/abs/2409.08226

  85. [93]

    rejeptai. 2024. Why doesn’t ClaudeBot/Anthropic obey robots.txt? https://www.reddit.com/r/Anthropic/comments/1c8tu5u/why_doesnt_ claudebot_anthropic_obey_robotstxt/

  86. [94]

    Copyright Research and Information Center. 2019. Copyright Law of Japan: Chapter II Rights of Authors (English Translation). https://www.cric.or.jp/ english/clj/cl2.html

  87. [95]

    Stefano Rovetta, Alberto Cabri, Francesco Masulli, and Grażyna Suchacka. 2019. Bot or Not? A Case Study on Bot Recognition from Web Session Logs. Quanti- fying and Processing Biomedical and Behavioral Signals (2019), 197–206

  88. [96]

    Strowes, and Narseo Vallina-Rodriguez

    Quirin Scheitle, Oliver Hohlfeld, Julien Gamba, Jonas Jelten, Torsten Zimmer- mann, Stephen D. Strowes, and Narseo Vallina-Rodriguez. 2018. A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists. InProc. of ACM Internet Measurement Conference 2018 ...

  89. [97]

    M. H. M. Schellekens. 2013. Robot.txt: balancing interests of content producers and content users. Bridging Distances in Technology and Regulation (2013), 173–187

  90. [98]

    Barry Schwartz. 2023. Google-Extended does not stop Google Search Generative Experience from using your site’s content. https://searchengineland.com/google- extended-does-not-stop-google-search-generative-experience-from-using- your-sites-content-433058

  91. [99]

    Yoast SEO. 2025. SEO starts with Yoast. https://yoast.com/

  92. [100]

    Shawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng, Rana Hanocka, and Ben Y Zhao. 2023. Glaze: Protecting Artists from Style Mimicry by Text-to-Image Models. In Proc. of 32nd USENIX Security Symposium . 2187–2204

  93. [101]

    Shawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu, Haitao Zheng, and Ben Y. Zhao. 2024. Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models. In Proc. of IEEE Symposium on Security and Privacy 2024. 807–825

  94. [102]

    Jingyu Shi, Rahul Jain, Runlin Duan, and Karthik Ramani. 2023. Understanding Generative AI in Art: An Interview Study with Artists on G-AI from an HCI Perspective. (2023), 15 pages. arXiv:2310.13149 [cs.HC] https://arxiv.org/abs/ 2310.13149 IMC ’25, October 28–31, 2025, Madiso...

  95. [103]

    Dusan Stevanovic, Aijun An, and Natalija Vlajic. 2012. Feature evaluation for web crawler detection with data mining techniques. Expert Systems with Applications 39, 10 (2012), 8707–8717

  96. [104]

    Dongxun Su, Yanjie Zhao, Xinyi Hou, Shenao Wang, and Haoyu Wang. 2024. GPT Store Mining and Analysis. (2024), 16 pages. arXiv:2405.10210 [cs.LG] https://arxiv.org/abs/2405.10210

  97. [105]

    Grażyna Suchacka, Alberto Cabri, Stefano Rovetta, and Francesco Masulli. 2021. Efficient on-the-fly Web bot detection. Knowledge-Based Systems 223 (2021), 16 pages

  98. [106]

    Mark Sullivan. 2024. AI Companies Ignoring Robots.txt. https://mjtsai.com/ blog/2024/06/24/ai-companies-ignoring-robots-txt/

  99. [107]

    Councill, and C

    Yang Sun, Ziming Zhuang, Isaac G. Councill, and C. Lee Giles. 2007. Deter- mining Bias to Search Engines from Robots.txt. In Proc. of IEEE WIC ACM Web Intelligence and Intelligent Agent Technology 2007 . 149–155

  100. [108]

    Lee Giles

    Yang Sun, Ziming Zhuang, and C. Lee Giles. 2007. A Large-Scale Study of Robots.txt. In Proc. of 16th World Wide Web Conference . 1123–1124

  101. [109]

    David Sénécal. 2024. The Web Scraping Problem: Part 1 — Akamai. https: //www.akamai.com/blog/security/the-web-scraping-problem-part-1

  102. [110]

    Reid Tatoris, Harsh Saxena, and Luis Miglietti. 2025. Trapping misbehaving bots in an AI Labyrinth — Cloudflare. https://blog.cloudflare.com/ai-labyrinth/

  103. [111]

    Antoine Vastel, Walter Rudametkin, Romain Rouvoy, and Xavier Blanc. 2020. FP- Crawlers: Studying the Resilience of Browser Fingerprinting to Block Crawlers. In Proc. of NDSS Workshop on Measurements, Attacks, and Defenses for the Web . 13 pages

  104. [112]

    James Vincent. 2023. Getty Images is suing the creators of AI art tool Stable Diffusion for scraping its content — The Verge. https://www.theverge.com/ 2023/1/17/23558516/ai-art-copyright-stable-diffusion-getty-images-lawsuit

  105. [113]

    Dark Visitors. 2024. Agents. https://darkvisitors.com/agents

  106. [114]

    Dark Visitors. 2025. Track the AI Agents and Bots Crawling Your Website. https://darkvisitors.com/

  107. [115]

    W3Techs. 2024. Usage statistics and market shares of reverse proxy services. https://w3techs.com/technologies/overview/proxy

  108. [116]

    Wikipedia contributors. 2020. Global Privacy Control — Wikipedia. https: //en.wikipedia.org/wiki/Global_Privacy_Control

  109. [117]

    Wix. 2025. Wix.com Terms of Use. https://www.wix.com/about/terms-of-use

  110. [118]

    Chloe Xiang. 2022. Artists Are Revolting Against AI Art on ArtStation — Vice. https://www.vice.com/en/article/ake9me/artists-are-revolt-against-ai- art-on-artstation

  111. [119]

    Chyan Yang and Hsien-Jyh Liao. 2010. Using the Robots. txt and Robots Meta tags to implement online copyright and a related amendment. Library Hi Tech 28, 1 (2010), 94–106

  112. [120]

    Zejun Zhang, Li Zhang, Xin Yuan, Anlan Zhang, Mengwei Xu, and Feng Qian

  113. [121]

    Eric Zhou and Dokyun Lee. 2024. Generative artificial intelligence, human creativity, and art. PNAS Nexus 3, 3 (2024), 8 pages

  114. [122]

    User-agent

    Viola Zhou. 2023. AI is already taking video game illustrators’ jobs in China — Rest of World. https://restofworld.org/2023/ai-china-video-game-layoffs- illustrators/ A Ethics We believe our work has very low ethical risk. Our user study is approved by the IRB at our instituti...

  115. [2023]

    Science 380, 6650 (2023), 1110–1111

    Art and the science of generative AI. Science 380, 6650 (2023), 1110–1111

  116. [2024]

    (2024), 11 pages

    A First Look at GPT Apps: Landscape and Vulnerability. (2024), 11 pages. arXiv:2402.15105 [cs.CR] https://arxiv.org/abs/2402.15105

  117. [2025]

    (2025), 23 pages

    From Pen to Prompt: How Creative Writers Integrate AI into their Writing Practice. (2025), 23 pages. arXiv:2411.03137 [cs.HC] https://arxiv.org/abs/2411. 03137

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.