REVIEW 3 major objections 7 minor 1 cited by
Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers
T0 review · 3 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that the existing voluntary web controls for AI crawling fail individual artists in three ways: most surveyed artists have never heard of robots.txt, most hosting platforms give them no way to edit it, and most…
desk verdict A careful measurement study of AI crawler controls whose artist survey is the soft spot; the network results are solid and worth citing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is robots.txt, the Robots Exclusion Protocol (RFC 9309): a plain-text file placed at a site's root that names user agents (e.g., GPTBot, ChatGPT-User) and disallows or allows paths. It is an honor-based signal, not an enforceable access control, so its value depends on each crawler's willingness to fetch and obey it. The paper's machinery for testing that willingness is a pair of researcher-controlled honey-pot websites (one disallowing all crawlers, one disallowing each AI user agent individually) whose server logs reveal which crawlers fetch robots.txt and then still request content, plus active triggering of ChatGPT-store GPT apps to force third-party assistant crawlers to visit. For the artist-side claims, the machinery is a 203-respondent survey with an embedded attention-check term and a DNS-level census of 1,182 artist sites to see which hosting providers expose robots.txt controls.
What would settle it
Re-run the artist survey with a probability-sampled panel of working artists and a neutral description of robots.txt that does not call it an 'easy way' or 'quick win'; if awareness exceeds, say, 80% or a majority of artists report being able to edit robots.txt through their host, the paper's headline gap collapses. Independently, instrument a fresh set of 100 AI-assistant endpoints over three months and count how many fetch robots.txt before requesting content; if most fetch and obey, the 20-of-23 non-compliance result is a snapshot, not a stable property.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a three-part gap between what artists want and what the current technical stack delivers. Awareness: 59% of the surveyed professional artists had never heard of robots.txt, and average self-rated familiarity with it was 1.99 on a 5-point scale, just above the fake control term. Agency: among over 1,100 artist websites, over 78% are hosted on eight platforms; four of the top eight provide no way for users to modify robots.txt, and even where a control exists (paid Wix, Squarespace toggle) uptake is near zero or 17%, far below the 75% expressed intention. Efficacy: in six months of passive and active measurement on honey-pot websites, all major AI data crawlers except ByteDance's Bytespider respected robots.txt, but 20 of 23 third-party AI assistant crawlers never fetched robots.txt at all, so they cannot be bound by it. Active blocking via Cloudflare blocks more aggressively (17 AI user agents) but was enabled on only about 5.7% of Cloudflare-hosted top-10k sites and does not cover every AI crawler; the paper concludes that robots.txt and active blocking are complementary, not interchangeable.
Load-bearing premise
The central statistics assume that the 203 artists recruited through the authors' professional networks and Discord channels represent professional visual artists broadly, and that the favorable description of robots.txt shown before the adoption questions did not inflate stated willingness to use it.
Editorial extensions
If this is right
- If 20 of 23 third-party AI assistant crawlers ignore robots.txt, creators cannot rely on the protocol to stop real-time retrieval of their work by AI assistants; the gap is not the big training crawlers but user-triggered fetching.
- If most hosting platforms do not expose robots.txt, then protective intent must be implemented at the platform level (for example, a default-on switch), not by individual artists.
- If only 17% of Squarespace artists enable an AI-blocking toggle despite 75% stated intent, discoverability and clear communication of the control matter as much as the control itself.
- If active blocking cannot be tuned for dual-purpose crawlers like Googlebot, robots.txt remains necessary even where firewall-level blocking is in place.
- If licensing deals make publishers remove robots.txt restrictions, expressed opt-out intent is reversible and economic, not a fixed preference.
Reading between the lines
- The results imply that robots.txt is being asked to do two incompatible jobs at once: expressing legal preference and enforcing technical access; regulations like the EU AI Act that condition copyright carve-outs on respecting robots.txt inherit all of the protocol's gaps unless they define what counts as respect.
- A natural extension would separate deliberate policy from implementation bugs by checking whether the 20 non-fetching crawlers also ignore HTTP caching directives or fetch robots.txt on retry.
- If platform-level toggles became the norm, the measured 59% unawareness would likely drop quickly; a testable prediction is that making a Squarespace-style switch default-on would raise effective protection, though it might also affect artists' search discoverability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies whether the web-control mechanisms available to individual content creators — robots.txt, noai meta tags, and active blocking via reverse proxies — are known to creators, available to them through their hosting platforms, and effective against AI crawlers. The paper combines four complementary measurements: (1) a longitudinal analysis of robots.txt adoption across 40,455 domains that appear in the Tranco top-100k in every month from October 2022 to October 2024, using Common Crawl snapshots cross-validated against the Internet Archive; (2) a user study of 203 artists, plus a measurement of 1,182 artist websites and the control that their eight most common hosting providers expose; (3) passive and active tests of whether 24 AI-related user agents respect robots.txt on two sites under the authors' control; and (4) an evaluation of active blocking, including an inferred behavior model of Cloudflare's 'Block AI Bots' feature and an estimate of its deployment across the Tranco top-10k. The headline findings are that 59% of surveyed artists had never heard of robots.txt, that most hosting platforms do not let artists edit it, that most large AI companies' data crawlers respect robots.txt while 20 of 23 triggered third-party AI assistant crawlers do not fetch it, and that active blocking is more enforceable but sparsely deployed and incomplete in coverage.
Significance. If the measurements hold, this is the most complete empirical map to date of the gap between what individual content creators want from anti-AI-crawling tools and what the current Web actually provides, and it is likely to become the reference point for the state of AI-crawler blocking circa 2024-2025. The network-measurement components are genuinely strong: the robots.txt parser is validated against RFC 9309 edge cases, the Common Crawl longitudinal data is cross-checked against the Internet Archive and an independent crawl, the crawler-compliance tests are direct and falsifiable (20 of 23 triggered third-party assistant crawlers never fetched robots.txt), and the Cloudflare setting is inferred through paired grey-box tests with the authors' own account as ground truth. The paper also ships public code and data via a GitHub repository. There are no fitted parameters and the central claims do not reduce to an input, so circularity is not a concern.
major comments (3)
- [§4.2, Appendix D.1 (and Abstract)] The 75% intention-to-adopt figure is load-bearing for the abstract's 'strong demand for tools like robots.txt,' but it is measured by Q26, which is asked immediately after a description that asserts 'over 90% of artists don't realize they can use a simple tool called robots.txt,' calls the tool 'an easy way for artists to protect their work,' and calls adding the file 'a quick win' (Appendix D.1). This is a leading prompt rather than neutral information, so 75% plausibly overstates the adoption intent that a neutral framing would elicit; Q27's preamble ('most companies respect it') raises the same concern for the 77% distrust figure. Please (i) report the un-primed Q22/Q23 results (the 97% desire-for-a-button result) side by side with the primed Q26 result, (ii) add a sensitivity analysis, e.g., re-scoring only the 'very likely' responses, and (iii) add a priming caveat to the Limitations section and hedge the abstract accordingly.
- [§4.1, §7, Appendix D.2 (and Abstract)] The 59% awareness figure, which anchors the title and abstract, is estimated from a snowball convenience sample: 203 artists recruited through the authors' social circles, professional Discord channels, and participant referrals (Section 4.1), with 109 of 203 respondents based in North America (89 in the US) and heavy concentration in illustration (163), digital 2D (143), and character/creature design (99) (Appendix D.2). The awareness question (Q24) is asked before the robots.txt description, so it is not affected by the priming in the previous comment, but the sample's representativeness is still unestablished, and recruiting through AI-activism-adjacent professional networks could plausibly bias awareness and sentiment in unquantified ways. I credit the bogus-item attention check ('nearest diffusion tree') and the geographic caveat in Section 7, but the paper should add a confidence interval for the 59% estimate, an explicit discussion of the direction and plausible magnitude of selection bias, and should scope the abstract's generalization to the surveyed population.
- [§5.2.2 and §4.2] The '20 of 23' third-party crawler result is an existence proof for a large class of non-compliant assistant crawlers, and the measurement design (triggered fetches to sites under the authors' control) is sound. However, Section 4.2 generalizes this result to 'the majority of AI assistant crawlers do not respect robots.txt,' while the 23 crawlers were obtained by prompting the top-5k GPT-store apps listed on a third-party directory with two specific prompts ('Start action, fetch page: [url]' and 'Get web page content: [url]'). That sampling frame is a self-selected slice of the assistant-crawler ecosystem, likely dominated by small gateway services, and it is not demonstrably representative of all AI assistant crawlers. Please state this sampling frame explicitly in Section 5.2.2, and soften the class-level generalization in Section 4.2 and the abstract to 'the 23 third-party assistant crawlers we were able to trigger.'
minor comments (7)
- [Abstract and §4.1] The abstract refers to '203 professional artists,' but only 136 of 203 respondents (67%) self-identify as professional in Q1; the remaining third earn income from art but declined the label, so please make the wording precise, e.g., '203 artists, of whom 67% self-identify as professional.'
- [Table 2] The 100% '% Disallow AI' value for Carbonmade reflects the provider's default robots.txt file (which disallows GPTBot and CCBot) rather than artist behavior; the text explains this, but the table should carry an explicit footnote so the column is not read as a measure of artist agency.
- [§4.2] 'A significant majority (185, 93%)' appears inconsistent as written: 185/203 is 91%, so the 93% is presumably computed over the 'over 97%' subset who expressed a desire to block; please state the denominator explicitly.
- [§2.2] The noai/noimageai check (17 and 16 sites among the Tranco top-10k of October 2024) is reported without methodological detail; a sentence describing the fetch procedure and the measurement date would make this reproducible.
- [§6.3] The criterion for identifying the 2,018 (20%) Tranco top-10k sites hosted on Cloudflare is not described; please state the detection method (e.g., DNS, cdn-cgi endpoints, or response headers).
- [Figure 4] The caption mixes per-period semantics ('removed restrictions... in each time period') with a cumulative curve ('explicitly allowed'); please clarify the semantics in the caption or legend so the two curves are not misread as comparable.
- [§5.1] The two test sites share one IP address, but the paper does not state which robots.txt configuration (the wildcard site or the per-agent site) the active GPT-app fetches targeted; since the per-agent file disallows only the 24 listed agents, this distinction matters for interpreting the 20/23 result, so please state it explicitly.
Circularity Check
No circularity: every load-bearing claim rests on direct external measurements or survey data, with no fitted parameter renamed as a prediction.
full rationale
This is an observational measurement paper, not a derivation. The abstract's claim that strong demand for robots.txt-like tools is constrained by awareness, agency, and efficacy rests on three independent empirical legs: (1) a survey of 203 artists reporting that 59% had never heard of robots.txt, (2) a measurement of hosting providers showing most do not expose robots.txt controls, and (3) controlled-website experiments showing that 20 of 23 third-party AI assistant crawlers do not fetch robots.txt. Each of these is measured directly against external benchmarks (Common Crawl, Tranco, Dark Visitors, controlled websites, Cloudflare dashboards) rather than derived from an input assumption. The only self-citations, Glaze and Nightshade, appear as survey response options and background context; they are not used to justify the paper's central measurements or conclusions. The survey's robots.txt description is a potential validity concern because it may prime adoption answers, but priming is a bias in measurement, not circularity: the paper does not define its conclusion in terms of the prompt, and the 59% awareness figure is collected before the prompt is shown. No equation reduces to an input, no fitted parameter is relabeled as a prediction, and no uniqueness theorem is imported from the authors' prior work. The paper is self-contained against external evidence, so the circularity score is 0.
Assumptions & free parameters
assumptions (6)
- domain assumption The Dark Visitors list of AI user agents is a comprehensive and correct enumeration of AI crawlers.
- domain assumption Google's robots.txt parser correctly interprets RFC 9309 semantics for the studied files.
- domain assumption Crawler compliance can be judged from observed fetches on two simple controlled websites over six months.
- domain assumption User-agent-based differences in HTTP responses correctly identify active blocking of AI crawlers.
- domain assumption Cloudflare's Block AI Bots setting can be inferred from HTTP responses to a small set of user agents on a headless browser.
- domain assumption Survey participants' self-reported familiarity reflects actual knowledge.
Cite this review
Pith. "Pith review of Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers." pith.science (2026). https://pith.science/paper/IBO5YPYA
@misc{pith2026241115091,
author = {Pith},
title = {Pith review of: Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers},
year = {2026},
howpublished = {\url{https://pith.science/paper/IBO5YPYA}},
note = {Machine review of arXiv:2411.15091}
}
read the original abstract
The success of generative AI relies heavily on training on data scraped through extensive crawling of the Internet, a practice that has raised significant copyright, privacy, and ethical concerns. While few measures are designed to resist a resource-rich adversary determined to scrape a site, crawlers can be impacted by a range of existing tools such as robots.txt, NoAI meta tags, and active crawler blocking by reverse proxies. In this work, we seek to understand the ability and efficacy of today's networking tools to protect content creators against AI-related crawling. For targeted populations like human artists, do they have the technical knowledge and agency to utilize crawler-blocking tools such as robots.txt, and can such tools be effective? Using large scale measurements and a targeted user study of 203 professional artists, we find strong demand for tools like robots.txt, but significantly constrained by critical hurdles in technical awareness, agency in deploying them, and limited efficacy against unresponsive crawlers. We further test and evaluate network-level crawler blockers provided by reverse proxies. Despite relatively limited deployment today, they offer stronger protections against AI crawlers, but still come with their own set of limitations.
Figures
Forward citations
Cited by 1 Pith paper
-
The Agentic Web Requires New Normative Infrastructure
The web's anti-bot regime should be replaced by a framework that presumptively lets user-authorized AI agents act for their principals, requires platforms to disclose access policies, and permits agent blocking only w...
Reference graph
Works this paper leans on
-
[1]
Adobe. 2024. Adobe General Terms of Use. https://www.adobe.com/legal/terms. html
2024
-
[2]
AIOSEO. 2025. The Best WordPress SEO Plugin and Toolkit. https://aioseo.com/
2025
-
[3]
Safinah Ali and Cynthia Breazeal. 2023. Studying Artist Sentiments around AI-generated Artwork. (2023), 13 pages. arXiv:2311.13725 [cs.HC] https: //arxiv.org/abs/2311.13725
work page Pith review arXiv 2023
-
[4]
Weigle, and Michael L
Yasmin AlNoamany, Michele C. Weigle, and Michael L. Nelson. 2013. Access Patterns for Robots and Humans in Web Archives. InProc. of the 13th ACM/IEEE- CS Joint Conference on Digital Libraries . 339–348
2013
-
[5]
Babak Amin Azad, Oleksii Starov, Pierre Laperdrix, and Nick Nikiforakis. 2020. Web runner 2049: Evaluating third-party anti-bot services. In Proc. of 17th Detection of Intrusions and Malware, and Vulnerability Assessment . Springer, 135–159
2020
-
[6]
Anthropic. 2024. Does Anthropic crawl data from the web, and how can site owners block the crawler? https://support.anthropic.com/en/articles/8896518- does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block- the-crawler
arXiv 2024
-
[7]
ArtStation. 2025. ArtStation Terms of Service. https://www.artstation.com/tos
2025
-
[8]
Stefan Baak. 2024. Training Data for the Price of a Sandwich. https://foundation. mozilla.org/en/research/library/generative-ai-training-data/common-crawl/
2024
Show all 125 references
-
[9]
Quan Bai, Gang Xiong, Yong Zhao, and Longtao He. 2014. Analysis and De- tection of Bogus Behavior in Web Crawler Measurement. Procedia Computer Science 31 (2014), 1084–1091
2014
-
[10]
Kayleigh Barber. 2024. Future’s Jon Steinberg shares his philosophy on AI content licensing deals — Digiday. https://digiday.com/podcasts/futures-jon- steinberg-shares-his-philosophy-on-ai-content-licensing-deals/
2024
-
[11]
Muhammad Ahmad Bashir, Sajjad Arshad, Engin Kirda, William Robertson, and Christo Wilson. 2019. A Longitudinal Analysis of the ads.txt Standard. In Proc. ACM Internet Measurement Conference 2019 . 294–307
2019
-
[12]
Lucas Bellaiche, Rohin Shahi, Martin Harry Turpin, Anya Ragnhildstveit, Shawn Sprockett, Nathaniel Barr, Alexander Christensen, and Paul Seli. 2023. Humans versus AI: whether and why we prefer human-created compared to AI-created artwork. Cognitive Research: Principles and Imp...
2023
-
[13]
Alex Bocharov, Santiago Varagas, Adam Martinetti, Reid Tatoris, and Carlos Azevedo. 2024. Declare your AIndependence: block AI bots, scrapers and crawlers with a single click — Cloudflare. https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots- scrapers-and-cra...
2024
-
[14]
Josiah D Boucher, Gillian Smith, and Yunus Doğan Telliel. 2024. Is Resistance Futile?: Early Career Game Developers, Generative AI, and Ethical Skepticism. In Proc. of CHI Conference on Human Factors in Computing Systems 2024 . 1–13
2024
-
[15]
Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative Research in Psychology 3, 2 (2006), 77–101
2006
-
[16]
Jack Brewster, Zach Fishman, and Isaiah Glick. 2024. AI Chatbots Are Blocked by 67% of Top News Sites, Relying Instead on Low-Quality Sources — News Guard. https://www.newsguardtech.com/special-reports/67-percent-of-top- news-sites-block-ai-chatbots/
2024
-
[17]
Gordon Burtch, Dokyun Lee, and Zhichen Chen. 2023. The consequences of generative AI for UGC and online community engagement. SSRN (2023), 26 pages
2023
-
[18]
Carbonmade. 2024. Terms and Conditions of Use. https://carbonmade.com/ terms
2024
-
[19]
Eva Cetinic and James She. 2022. Understanding and Creating Art with AI: Review and Outlook. ACM Transactions on Multimedia Computing, Communi- cations, and Applications 18, 2 (2022), 1–22
2022
-
[20]
Zi Chu, Steven Gianvecchio, and Haining Wang. 2018. Bot or Human? A Behavior-Based Online Bot Detection System. From Database to Cyber Security: Essays Dedicated to Sushil Jajodia on the Occasion of His 70th Birthday (2018), 432–449
2018
-
[21]
Cloudflare. 2024. Verified Bots. https://radar.cloudflare.com/traffic/verified- bots
2024
-
[22]
European Commission. 2024. First Draft of the General-Purpose AI Code of Practice published, written by independent experts. https://digital-strategy.ec.europa.eu/en/library/first-draft-general-purpose- ai-code-practice-published-written-independent-experts
2024
-
[23]
European Commission. 2024. The AI Act Explorer. https: //artificialintelligenceact.eu/ai-act-explorer/
2024
-
[24]
Common Crawl. 2025. Common Crawl — Open Repository of Web Crawl Data. https://commoncrawl.org/
2025
-
[25]
davepattern. 2024. DDoS from Anthropic AI. https://www.linode.com/ community/questions/24842/ddos-from-anthropic-ai
2024
-
[26]
deninho32. 2024. Claudebot attack. https://www.phpbb.com/community/ viewtopic.php?t=2652265
2024
-
[27]
Design and Artists Copyright Society (DACS). 2024. Artificial Intelligence and Artists’ Work: A survey of artists on AI. https://cdn.dacs.org.uk/uploads/ documents/News/Artificial-Intelligence-and-Artists-Work-DACS.pdf. (2024), 29 pages
2024
-
[28]
Michael Dinzinger and Michael Granitzer. 2024. A Longitudinal Study of Content Control Mechanisms. In Proc. of the ACM Web Conference . 1382–1387
2024
-
[29]
Michael Dinzinger, Florian Heß, and Michael Granitzer. 2024. A Survey of Web Content Control for Generative AI. (2024), 12 pages. arXiv:2404.02309 [cs.IR] https://arxiv.org/abs/2404.02309
2024 arXiv
-
[30]
Frank, Matthew Groh, Laura Herman, Neil Leach, Robert Mahari, Alex “Sandy” Pentland, Olga Russakovsky, Hope Schroeder, and Amy Smith
Ziv Epstein, Aaron Hertzmann, Memo Akten, Hany Farid, Jessica Fjeld, Mor- gan R. Frank, Matthew Groh, Laura Herman, Neil Leach, Robert Mahari, Alex “Sandy” Pentland, Olga Russakovsky, Hope Schroeder, and Amy Smith
-
[31]
Fairly Trained. 2024. Statement on AI training. https://www.aitrainingstatement. org/
2024
-
[32]
Richard Fletcher. 2024. How many news websites block AI crawlers? Reuters Institute Factsheets (2024), 7 pages
2024
-
[33]
noai" and
Foundation Web Design & Development. 2022. What is DeviantArt’s new "noai" and "noimageai" meta tag and how to install it. https://www.foundationwebdev. com/2022/11/noai-noimageai-meta-tag-how-to-install/
2022
-
[34]
Thompson
Sheera Frenkel and Stuart A. Thompson. 2023. Not for Machines to Harvest: Data Revolts Break Out Against A.I. — The New York Times. https://www.nytimes. com/2023/07/15/technology/artificial-intelligence-models-chat-data.html
2023
-
[35]
Manuel B. Garcia. 2024. The Paradox of Artificial Creativity: Challenges and Opportunities of Generative AI Artistry. Creativity Research Journal (2024), 1–14
2024
-
[36]
generosus. 2023. PSA | Bytedance and Bytespider Bots | Recommend Block- ing. https://wordpress.org/support/topic/psa-bytedance-and-bytespider-bots- recommend-blocking/
2023
-
[37]
Jonathan Gillham. 2024. Block AI Bots from Crawling Websites Using Robots.txt — Originality.ai. https://originality.ai/ai-bot-blocking
2024
-
[38]
Google. 2024. Google Robots.txt Parser and Matcher Library. https://github. com/google/robotstxt
2024
-
[39]
Paolo Grigis and Antonella De Angeli. 2024. Playwriting with Large Language Models: Perceived Features, Interaction Strategies and Outcomes. In Proc. the International Conference on Advanced Visual Interfaces 2024 . 1–9
2024
-
[40]
Alicia Guo, Shreya Sathyanarayanan, Leijie Wang, Jeffrey Heer, and Amy Zhang
-
[41]
Eszter Hargittai. 2009. An Update on Survey Measures of Web-Oriented Digital Literacy. Social Science Computer Review 27, 1 (2009), 130–137
2009
-
[42]
Joo-Wha Hong and Nathaniel Ming Curran. 2019. Artificial Intelligence, Artists, and Art: Attitudes Toward Artwork Produced by Humans vs. Artificial Intel- ligence. ACM Transactions on Multimedia Computing, Communications, and Applications 15, 2s (2019), 1–16
2019
-
[43]
Xinyi Hou, Yanjie Zhao, and Haoyu Wang. 2024. On the (In)Security of LLM App Stores. (2024), 17 pages. arXiv:2407.08422 [cs.CR] https://arxiv.org/abs/ 2407.08422
2024 arXiv
-
[44]
Hongxian Huang, Runshan Fu, and Anindya Ghose. 2023. Generative AI and Content Creators: Evidence from Digital Art Platforms. SSRN (2023), 41 pages
2023
-
[45]
Christos Iliou, Theodoros Kostoulas, Theodora Tsikrika, Vasilis Katos, Stefanos Vrochidis, and Ioannis Kompatsiaris. 2021. Detection of Advanced Web Bots by Combining Web Logs with Mouse Behavioural Biometrics. Digital Threats: Research and Practice 2, 3 (2021), 1–26
2021
-
[46]
Christos Iliou, Theodoros Kostoulas, Theodora Tsikrika, Vasilis Katos, Stefanos Vrochidis, and Yiannis Kompatsiaris. 2019. Towards a framework for detecting advanced web bots. In Proc. of 14th International Conference on A vailability, Reliability and Security. 1–10
2019
-
[47]
Glassman
Katy Ilonka Gero, Meera Desai, Carly Schnitzler, Nayun Eom, Jack Cushman, and Elena L. Glassman. 2024. Creative Writers’ Attitudes on Writing as Training Data for Large Language Models. In Proc. of CHI Conference on Human Factors in Computing Systems 2025 . 1–16
2024
-
[48]
Imperva. 2024. 2024 Bad Bot Report. https://www.imperva.com/resources/ resource-library/reports/2024-bad-bot-report/
2024
-
[49]
Gregoire Jacob, Engin Kirda, Christopher Kruegel, and Giovanni Vigna. 2012. PUBCRAWL: Protecting Users and Businesses from CRAWLers. InProc. of 21st USENIX Security Symposium. 507–522
2012
-
[50]
Steve TK Jan, Qingying Hao, Tianrui Hu, Jiameng Pu, Sonal Oswal, Gang Wang, and Bimal Viswanath. 2020. Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data Augmentation. In Proc. of 2020 IEEE Symposium on Security and Privacy . IEEE, 1190–1206
2020
-
[51]
Harry H Jiang, Lauren Brown, Jessica Cheng, Mehtab Khan, Abhishek Gupta, Deja Workman, Alex Hanna, Johnathan Flowers, and Timnit Gebru. 2023. AI Art and its Impact on Artists. In Proc. of AAAI ACM Conference on AI, Ethics, and Society 2023. 363–374
2023
-
[52]
Hannah Johnston and David Thue. 2024. Understanding Visual Artists’ Values and Attitudes towards Collaboration, Technology, and AI. InProc. 50th Graphics Interface Conference. 1–9
2024
-
[53]
Ben Jones, Tzu-Wen Lee, Nick Feamster, and Phillipa Gill. 2014. Automated Detection and Fingerprinting of Censorship Block Pages. InProc. of ACM Internet Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers IMC ’25, October 2...
2014
-
[54]
Hugo Jonker, Benjamin Krumnow, and Gabry Vlot. 2019. Fingerprint Surface- Based Detection of Web Bot Detectors. In Proc. of 24th European Symposium on Research in Computer Security . Springer, 586–605
2019
-
[55]
Reishiro Kawakami and Sukrit Venkatagiri. 2024. The Impact of Generative AI on Artists. In Proc. of 16th Conference on Creativity & Cognition . 79–82
2024
-
[56]
Paul Keller. 2023. Protecting Creatives or Impeding Progress? Ma- chine learning and the EU copyright framework — Open Future. https://openfuture.eu/blog/protecting-creatives-or-impeding-progress/
2023
-
[57]
Kate Knibbs. 2024. Condé Nast Signs Deal With OpenAI — Wired. https: //www.wired.com/story/conde-nast-openai-deal/
2024
-
[58]
Kate Knibbs. 2024. The Race to Block OpenAI’s Scraping Bots Is Slowing Down — Wired. https://www.wired.com/story/open-ai-publisher-deals-scraping-bots/
2024
-
[59]
Robb Knight. 2024. Perplexity AI Is Lying about Their User Agent. https: //rknight.me/blog/perplexity-ai-is-lying-about-its-user-agent/
2024
-
[60]
Santanu Kolay, Paolo D’Alberto, Ali Dasdan, and Arnab Bhattacharjee. 2008. A Larger Scale Study of Robots.txt. In Proc. of 17th World Wide Web Conference . 1171–1172
2008
-
[61]
Martijn Koster, Gary Illyes, Henner Zeller, and Lizzi Sassman. 2022. Robots Exclusion Protocol. RFC 9309. https://doi.org/10.17487/RFC9309
2022 doi
-
[62]
Shinil Kwon, Young-Gab Kim, and Sungdeok Cha. 2012. Web robot detection based on pattern-matching technique. Journal of Information Science 38, 2 (2012), 118–126
2012
-
[63]
IAB Tech Lab. 2025. Global Privacy Protocol. https://iabtechlab.com/gpp/
2025
-
[64]
Rita Latikka, Jenna Bergdahl, Nina Savela, and Atte Oksanen. 2023. AI as an Artist? A Two-Wave Survey Study on Attitudes Toward Using Artificial Intelligence in Art. Poetics 101 (2023), 11 pages
2023
-
[65]
Junsup Lee, Sungdeok Cha, Dongkun Lee, and Hyungkyu Lee. 2009. Classi- fication of web robots: an empirical study based on over one billion requests. Computers & Security 28, 8 (2009), 795–802
2009
-
[66]
Jie Li, Hancheng Cao, Laura Lin, Youyang Hou, Ruihao Zhu, and Abdallah El Ali. 2024. User Experience Design Professionals’ Perceptions of Generative Artificial Intelligence. In Proc. of CHI Conference on Human Factors in Computing Systems 2024. 1–18
2024
-
[67]
Xigao Li, Babak Amin Azad, Amir Rahmati, and Nick Nikiforakis. 2021. Good Bot, Bad Bot: Characterizing Automated Browsing Activity. InProc. of 2021 IEEE Symposium on Security and Privacy . IEEE, 1589–1605
2021
-
[68]
Baker & Hostetler LLP. 2024. Case Tracker: Artificial Intelligence, Copyrights and Class Actions. https://www.bakerlaw.com/services/artificial-intelligence- ai/case-tracker-artificial-intelligence-copyrights-and-class-actions/
2024
-
[69]
Joseph Saveri Law Firm LLP. 2023. Class Action Filed Against Stability AI, Midjourney, and DeviantArt for DMCA Violations, Right of Publicity Violations, Unlawful Competition, Breach of TOS. https://www.prnewswire.com/news-releases/class-action-filed-against- stability-ai-midj...
2023
-
[70]
Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderin- wale, William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh, Manuel Cherep, Ahmad Anis, An Dinh, Caroline Chitongo, D...
2024 arXiv
-
[71]
Lourenço and Orlando O
Anália G. Lourenço and Orlando O. Belo. 2006. Catching Web Crawlers in the Act. In Proc. of 6th International Conference on Web Engineering . 265–272
2006
-
[72]
Juniper Lovato, Julia Witte Zimmerman, Isabelle Smith, Peter Dodds, and Jen- nifer L. Karson. 2024. Foregrounding Artist Opinions: A Survey Study on Transparency, Ownership, and Fairness in AI Generative Art. In Proc. of AAAI ACM Conference on AI, Ethics, and Society 2024 . 905–916
2024
-
[73]
Abhijeeth Madhu. 2023. Survey Reveals 9 out of 10 Artists Believe Current Copyright Laws are Outdated in the Age of Generative AI Technology — Book An Artist Blog. https://bookanartist.co/blog/2023-artists-survey-on-ai- technology/
2023
-
[74]
Bron Maher. 2024. Revealed: Which of the top 100 UK and US news websites are blocking AI crawlers — PressGazette. https://pressgazette.co.uk/platforms/ news-sites-block-ai-web-crawlers-chatgpt-google/
2024
-
[75]
Meta. 2024. Meta Web Crawlers. https://developers.facebook.com/docs/sharing/ webmasters/web-crawlers/
2024
-
[76]
Elz˙e Sigut˙e Mikalonyt˙e and Markus Kneer. 2022. Can Artificial Intelligence Make Art?: Folk Intuitions as to whether AI-driven Robots Can Be Viewed as Artists and Produce Art. ACM Transactions on Human-Robot Interaction 11, 4 (2022), 1–19
2022
-
[77]
Cullen Miller. 2023. ai.txt: A new way for websites to set permissions for AI — Spawning. https://spawning.substack.com/p/aitxt-a-new-way-for-websites-to- set
2023
-
[78]
Piotr Mirowski, Juliette Love, Kory Mathewson, and Shakir Mohamed. 2024. A Robot Walks into a Bar: Can Language Models Serve as Creativity Support Tools for Comedy? An Evaluation of LLMs’ Humour Alignment with Comedians. In Proc. of the ACM Conference on Fairness, Accountabili...
2024
-
[79]
Martin Monperrus. 2024. crawler-user-agents. https://github.com/monperrus/ crawler-user-agents
2024
-
[80]
Alexis Newton and Kaustubh Dhole. 2023. Is AI Art Another Industrial Rev- olution in the Making? (2023), 6 pages. arXiv:2301.05133 [cs.AI] https: //arxiv.org/abs/2301.05133
2023 arXiv
-
[81]
Republic of Singapore. 2021. Singapore Copyright Act of 2021: Section 244 (English Translation). https://sso.agc.gov.sg/Acts-Supp/22-2021/Published/ ?ProvIds=pr244-
2021
-
[82]
OpenAI. 2023. Our approach to AI safety. https://openai.com/index/our- approach-to-ai-safety/
2023
-
[83]
Karla Ortiz. 2024. Why AI Models are not inspired like humans. https: //www.kortizblog.com/blog/why-ai-models-are-not-inspired-like-humans
2024
-
[84]
Stack Overflow. 2024. Stack Overflow and OpenAI Partner to Strengthen the World’s Most Popular Large Language Models. https://stackoverflow.co/ company/press/archive/openai-partnership
2024
-
[85]
palewire. 2025. Who blocks OpenAI, Google AI and Common Crawl? https: //palewi.re/docs/news-homepages/openai-gptbot-robotstxt.html
2025
-
[86]
Sungjin Park. 2024. The work of art in the age of generative AI: aura, liberation, and democratization. AI & Society (2024), 1–10
2024
-
[87]
Perplexity. 2025. Perplexity Crawlers. https://docs.perplexity.ai/guides/bots
2025
-
[88]
Kien Pham, Aécio Santos, and Juliana Freire. 2016. Understanding Website Behavior based on User Agent. InProc. of 39th ACM SIGIR Conference on Research and Development in Information Retrieval . 1053–1056
2016
-
[89]
Tara Poteat and Frank Li. 2021. Who You Gonna Call? An Empirical Evaluation of Website security.txt Deployment. In Proc. of ACM Internet Measurement Conference 2021. 526–532
2021
-
[90]
Heila Precel, Allison McDonald, Brent Hecht, and Nicholas Vincent. 2024. A Canary in the AI Coal Mine: American Jews May Be Disproportionately Harmed by Intellectual Property Dispossession in Large Language Model Training. In Proc. of CHI Conference on Human Factors in Computi...
2024
-
[91]
PRNewswire. 2024. Dotdash Meredith Announces Strategic Partnership with OpenAI, Bringing Iconic Brands and Trusted Content to ChatGPT. https://dotdashmeredith.mediaroom.com/2024-05-07-Dotdash-Meredith- Announces-Strategic-Partnership-with-OpenAI,-Bringing-Iconic-Brands- and-Tr...
2024
-
[92]
Gayatri Raman and Erin Brady. 2024. Exploring Use and Perceptions of Genera- tive AI Art Tools by Blind Artists. (2024), 4 pages. arXiv:2409.08226 [cs.HC] https://arxiv.org/abs/2409.08226
2024 arXiv
-
[93]
rejeptai. 2024. Why doesn’t ClaudeBot/Anthropic obey robots.txt? https://www.reddit.com/r/Anthropic/comments/1c8tu5u/why_doesnt_ claudebot_anthropic_obey_robotstxt/
2024
-
[94]
Copyright Research and Information Center. 2019. Copyright Law of Japan: Chapter II Rights of Authors (English Translation). https://www.cric.or.jp/ english/clj/cl2.html
2019
-
[95]
Stefano Rovetta, Alberto Cabri, Francesco Masulli, and Grażyna Suchacka. 2019. Bot or Not? A Case Study on Bot Recognition from Web Session Logs. Quanti- fying and Processing Biomedical and Behavioral Signals (2019), 197–206
2019
-
[96]
Strowes, and Narseo Vallina-Rodriguez
Quirin Scheitle, Oliver Hohlfeld, Julien Gamba, Jonas Jelten, Torsten Zimmer- mann, Stephen D. Strowes, and Narseo Vallina-Rodriguez. 2018. A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists. InProc. of ACM Internet Measurement Conference 2018 ...
2018
-
[97]
M. H. M. Schellekens. 2013. Robot.txt: balancing interests of content producers and content users. Bridging Distances in Technology and Regulation (2013), 173–187
2013
-
[98]
Barry Schwartz. 2023. Google-Extended does not stop Google Search Generative Experience from using your site’s content. https://searchengineland.com/google- extended-does-not-stop-google-search-generative-experience-from-using- your-sites-content-433058
2023
-
[99]
Yoast SEO. 2025. SEO starts with Yoast. https://yoast.com/
2025
-
[100]
Shawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng, Rana Hanocka, and Ben Y Zhao. 2023. Glaze: Protecting Artists from Style Mimicry by Text-to-Image Models. In Proc. of 32nd USENIX Security Symposium . 2187–2204
2023
-
[101]
Shawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu, Haitao Zheng, and Ben Y. Zhao. 2024. Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models. In Proc. of IEEE Symposium on Security and Privacy 2024. 807–825
2024
-
[102]
Jingyu Shi, Rahul Jain, Runlin Duan, and Karthik Ramani. 2023. Understanding Generative AI in Art: An Interview Study with Artists on G-AI from an HCI Perspective. (2023), 15 pages. arXiv:2310.13149 [cs.HC] https://arxiv.org/abs/ 2310.13149 IMC ’25, October 28–31, 2025, Madiso...
2023 arXiv
-
[103]
Dusan Stevanovic, Aijun An, and Natalija Vlajic. 2012. Feature evaluation for web crawler detection with data mining techniques. Expert Systems with Applications 39, 10 (2012), 8707–8717
2012
-
[104]
Dongxun Su, Yanjie Zhao, Xinyi Hou, Shenao Wang, and Haoyu Wang. 2024. GPT Store Mining and Analysis. (2024), 16 pages. arXiv:2405.10210 [cs.LG] https://arxiv.org/abs/2405.10210
2024 arXiv
-
[105]
Grażyna Suchacka, Alberto Cabri, Stefano Rovetta, and Francesco Masulli. 2021. Efficient on-the-fly Web bot detection. Knowledge-Based Systems 223 (2021), 16 pages
2021
-
[106]
Mark Sullivan. 2024. AI Companies Ignoring Robots.txt. https://mjtsai.com/ blog/2024/06/24/ai-companies-ignoring-robots-txt/
2024
-
[107]
Councill, and C
Yang Sun, Ziming Zhuang, Isaac G. Councill, and C. Lee Giles. 2007. Deter- mining Bias to Search Engines from Robots.txt. In Proc. of IEEE WIC ACM Web Intelligence and Intelligent Agent Technology 2007 . 149–155
2007
-
[108]
Lee Giles
Yang Sun, Ziming Zhuang, and C. Lee Giles. 2007. A Large-Scale Study of Robots.txt. In Proc. of 16th World Wide Web Conference . 1123–1124
2007
-
[109]
David Sénécal. 2024. The Web Scraping Problem: Part 1 — Akamai. https: //www.akamai.com/blog/security/the-web-scraping-problem-part-1
2024
-
[110]
Reid Tatoris, Harsh Saxena, and Luis Miglietti. 2025. Trapping misbehaving bots in an AI Labyrinth — Cloudflare. https://blog.cloudflare.com/ai-labyrinth/
2025
-
[111]
Antoine Vastel, Walter Rudametkin, Romain Rouvoy, and Xavier Blanc. 2020. FP- Crawlers: Studying the Resilience of Browser Fingerprinting to Block Crawlers. In Proc. of NDSS Workshop on Measurements, Attacks, and Defenses for the Web . 13 pages
2020
-
[112]
James Vincent. 2023. Getty Images is suing the creators of AI art tool Stable Diffusion for scraping its content — The Verge. https://www.theverge.com/ 2023/1/17/23558516/ai-art-copyright-stable-diffusion-getty-images-lawsuit
2023
-
[113]
Dark Visitors. 2024. Agents. https://darkvisitors.com/agents
2024
-
[114]
Dark Visitors. 2025. Track the AI Agents and Bots Crawling Your Website. https://darkvisitors.com/
2025
-
[115]
W3Techs. 2024. Usage statistics and market shares of reverse proxy services. https://w3techs.com/technologies/overview/proxy
2024
-
[116]
Wikipedia contributors. 2020. Global Privacy Control — Wikipedia. https: //en.wikipedia.org/wiki/Global_Privacy_Control
2020
-
[117]
Wix. 2025. Wix.com Terms of Use. https://www.wix.com/about/terms-of-use
2025
-
[118]
Chloe Xiang. 2022. Artists Are Revolting Against AI Art on ArtStation — Vice. https://www.vice.com/en/article/ake9me/artists-are-revolt-against-ai- art-on-artstation
2022
-
[119]
Chyan Yang and Hsien-Jyh Liao. 2010. Using the Robots. txt and Robots Meta tags to implement online copyright and a related amendment. Library Hi Tech 28, 1 (2010), 94–106
2010
-
[120]
Zejun Zhang, Li Zhang, Xin Yuan, Anlan Zhang, Mengwei Xu, and Feng Qian
-
[121]
Eric Zhou and Dokyun Lee. 2024. Generative artificial intelligence, human creativity, and art. PNAS Nexus 3, 3 (2024), 8 pages
2024
-
[122]
User-agent
Viola Zhou. 2023. AI is already taking video game illustrators’ jobs in China — Rest of World. https://restofworld.org/2023/ai-china-video-game-layoffs- illustrators/ A Ethics We believe our work has very low ethical risk. Our user study is approved by the IRB at our instituti...
2023
-
[2023]
Science 380, 6650 (2023), 1110–1111
Art and the science of generative AI. Science 380, 6650 (2023), 1110–1111
2023
-
[2024]
(2024), 11 pages
A First Look at GPT Apps: Landscape and Vulnerability. (2024), 11 pages. arXiv:2402.15105 [cs.CR] https://arxiv.org/abs/2402.15105
2024 arXiv
-
[2025]
(2025), 23 pages
From Pen to Prompt: How Creative Writers Integrate AI into their Writing Practice. (2025), 23 pages. arXiv:2411.03137 [cs.HC] https://arxiv.org/abs/2411. 03137
2025 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.