{"id":"e5c7ae47-7f41-4673-bb45-6badecc99547","arxiv_id":"2510.10315","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Reputable news sites block AI crawlers at roughly six times the rate of misinformation sites, and this gap is widening over time.","lead":"Many established news sites now block AI web crawlers in their robots.txt files, while sites known for misinformation mostly do not. The gap has grown since 2023 and could skew the content that future AI models learn from.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gap may reflect popularity/technical maturity rather than misinformation status; Fig. 3 does not control for Tranco rank.","rationale":"The reader's weakest assumption was that MBFC labels may be noisy or track confounding attributes. I agree, but I focus more narrowly and concretely on the confounding of misinformation status with popularity and technical maturity, which the paper's own Figure 3 fails to control. This is load-bearing because even with perfectly valid labels, the central claim about the accessibility skew requires the gap to be an attribute of misinformation, not of being a smaller or less professional site. The paper's limitation section (§5.2) admits technical capacity is a plausible alternative, but does not test it. Hence the verdict remains CONDITIONAL: the descriptive result is plausible, but the interpretation needs this additional control. This is a partial agreement with the reader because they emphasized label validity; I emphasize the confound problem. The proposed matched-pair/regression test would settle whether the gap survives adjustment, directly addressing the weakest link in the causal interpretation.","tokens_in":15242,"tokens_out":8142,"duration_ms":74229,"concrete_test":"Perform a matched-pair analysis on the 493 misinformation sites with valid robots.txt in §3.2. For each, match one reputable site on log(Tranco rank) within ±0.5 and on platform (presence of '/wp-content/' or '/sites/' in the homepage URL as a CMS indicator). Then compute the proportion with DisallowAll for ≥1 AI agent in both groups and the rank-stratified gap (Tranco <10K, 10K–100K, 100K–1M, >1M). If the matched/stratified gap drops below 15 points (vs. 51 points unadjusted), the 'misinformation effect' is largely a size/technical-maturity artifact. A secondary check: add log-rank and CMS as covariates in a logistic regression for AI-blocking; the coefficient on the misinformation indicator should remain significant and large.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline 60.0% vs 9.1% (Table 2) is internally consistent, but its interpretation as 'misinformation is more open' is not identified. The reputable and misinformation sets differ enormously in prevalence, resources, and likely technical sophistication: 96.4% of reputable sites serve a robots.txt vs 73.8% of misinformation sites, and the misinformation set is skewed to lower Tranco ranks (Fig. 8). The only control for popularity, Fig. 3, plots the Tranco ECDF of sites that already contain a DisallowAll rule; it does not estimate the gap conditional on rank. An unadjusted ECDF comparison cannot rule out that the gap is driven by site popularity, platform defaults, or management capacity rather than by the misinformation/reputable distinction. The paper explicitly acknowledges in §5.2 that the absence of directives 'may reflect limited awareness, technical capacity, or willingness,' but it never attempts to separate these. Since the practical conclusion (LLM training data may skew toward misinformation) depends on the gap being a property of misinformation status, not of correlated site characteristics, this missing control is the load-bearing concern.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares robots.txt files of 3,369 reputable news websites and 710 misinformation websites, as classified by Media Bias/Fact Check, to measure how often each group restricts AI crawlers. It reports that 60.0% of reputable sites with a robots.txt disallow at least one AI agent, versus 9.1% of misinformation sites; the average number of disallowed agents is 15.5 vs 0.77. It also measures active blocking via HTTP requests with AI User-Agent strings, finding both groups engage in it, and uses Internet Archive snapshots to show the gap grew from 23% to ~60% between September 2023 and May 2025. The authors conclude that the accessible web is increasingly skewed toward lower-credibility content.","tokens_in":15500,"tokens_out":5420,"duration_ms":48785,"significance":"This is the first study, to my knowledge, that compares robots.txt AI-exclusion practices across credibility categories. The dataset is large, the crawling is multi-vantage, the AI agent list is carefully curated, and the longitudinal analysis using six IA snapshots is a valuable contribution. The paper is transparent about its methodology and limitations. If the descriptive gap is robust, it raises important and timely questions about AI training data composition and the incentives of low-credibility sites to remain permissive. The paper does not overclaim causality regarding training data, explicitly stating that it does not have visibility into proprietary datasets. Its main weakness is the limited control for popularity and other correlated site characteristics when interpreting the gap as a property of 'misinformation'.","major_comments":[{"comment":"The claim that 'popularity alone does not determine the robots.txt practices' is not supported by the evidence presented. The ECDF in Fig. 3 plots the Tranco rank distribution only among sites that already have a DisallowAll rule; it does not estimate the probability of having such a rule conditional on rank. Because the reputable set has vastly higher robots.txt adoption (96.4% vs 73.8%) and is skewed to higher popularity (Fig. 8), the raw gap in Table 2 could be driven entirely by differences in site scale or management capacity. A stratified comparison within Tranco bands or a logistic regression with rank as a covariate is needed to support the interpretation that misinformation status, rather than correlated attributes, explains the gap. As written, the paper's own limitation in §5.2 is more accurate than the §4.1 claim.","section":"§4.1, Figure 3"},{"comment":"The definition of active blocking is underspecified. The text states that the crawler recorded 'the returned status code and response size in bytes' and filters out non-200 responses, but it never states the exact decision rule for classifying a site as actively blocking. Was a site classified as blocking if a 200 response in the control crawl became a 4xx/5xx in the AI-agent crawl? If response-size shrinkage alone was used, what threshold? Without this criterion, the headline figures (e.g., 16.9% of reputable sites blocking both agents) cannot be reproduced or interpreted. The use of only two AI UAs and a single vantage point is acknowledged, but the binary classification rule is essential and should be stated explicitly.","section":"§4.3"}],"minor_comments":[{"comment":"'DisallowAll' is not a standard REP term; define it (e.g., 'disallow:/') or use 'fully disallowed'.","section":"Throughout"},{"comment":"The two numbers '70,410 robots.txt Retrieved' and '3,672 robots.txt Retrieved' are inconsistent with Table 2 (3,179 + 493 = 3,672). If 70,410 is the total across all vantage points/snapshots, clarify.","section":"Figure 2"},{"comment":"Typo: 'we conduct alongi-tudinal analysis' should be 'conduct a longitudinal analysis'.","section":"§3.2"},{"comment":"The control crawl excludes 18 sites that 'returned 200 responses but were no longer operational'; how was non-operational status determined? Unclear.","section":"§4.3"},{"comment":"The limitations section does not mention the active-blocking methodological choices (only two AI UAs, single vantage point, single request per UA); consider adding.","section":"§5.2"},{"comment":"For the longitudinal analysis, the paper selects the most recent IA snapshot before each date but does not report per-snapshot coverage or whether sites with missing snapshots differ systematically. A brief coverage table would strengthen confidence.","section":"§3.2/§4.4"}],"recommendation":"major_revision","confidential_remarks":"This is a solid measurement study with a timely and important comparison. The main gap is the popularity/management confound: the ECDF in Fig. 3 does not control for rank, yet §4.1 uses it to dismiss the confound. The active-blocking section also needs a precise decision rule. Both are fixable with a re-analysis of existing data, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something genuinely new: it segments AI-related robots.txt restrictions by credibility class and tracks the gap over time. Prior work measured the overall rise of AI blocking (Longpre, Dinzinger, Liu); this is the first to compare reputable news sites against MBFC-labeled misinformation sites longitudinally. The core descriptive finding—60.0% of reputable sites disallow at least one AI agent versus 9.1% of misinformation sites, with the gap widening from 23% to ~60% between Sep 2023 and May 2025—is internally consistent and survives the paper's stated method. The active-blocking measurement is a nice addition, and the authors are careful not to overclaim causality to training data. Credit where due: the crawler methodology is reasonable, the longitudinal use of Internet Archive snapshots is appropriate, and the paper is clearly written.\n\nThe soft spots are real but not fatal. The big one is the popularity/technical-maturity confound. Reputable and misinformation sites differ enormously in size, resources, and likely webmaster sophistication, and Fig. 3 does not actually control for Tranco rank—it just plots ECDFs of sites that already have a DisallowAll rule. The stress-test note is right that this is the load-bearing alternative explanation. The paper acknowledges in §5.2 that missing directives 'may reflect limited awareness, technical capacity, or willingness,' but it never attempts to separate these. That said, the manual inspection of top-10k misinformation sites (10 of 14 impose no AI restrictions) is a partial and credible response; it's not a full control but it shows the effect is not purely a Top-10k popularity story. The MBFC labels are subjective and English-centric, but the authors acknowledge this too, and prior work commonly uses MBFC. The active-blocking probes use only two UAs from one location, so that section is suggestive rather than definitive. No released data or code, which hampers replication but is not a methodological flaw.\n\nThe paper is a solid descriptive contribution, not a paradigm shift. The interpretation as 'misinformation is more open' is plausible and important, but it should be framed as a hypothesis about what drives the gap, not a demonstrated fact. The central descriptive result holds; what needs work is the identification.\n\nWho should read this: anyone working on AI training data sourcing, web measurement, or content moderation. It deserves a serious referee—an editor should send it out, with a request for a rank-matched or regression-controlled analysis and data release. I'd bring it to a reading group and would cite it if I were writing about AI data access asymmetries.","headline":"Solid descriptive measurement of a real and growing asymmetry, but the headline 'misinformation is more open' is one plausible interpretation among several; worth refereeing with a request for popularity-controlled analyses and released artifacts.","tokens_in":15931,"tokens_out":901,"would_cite":true,"duration_ms":10394,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper establishes that 60.0% of reputable news sites disallow at least one AI crawler in robots.txt, versus 9.1% of misinformation sites, a gap that has widened from 23% since September 2023 and may be skewing the content available to A","keywords":["robots.txt","AI crawlers","misinformation","LLM training data","web content access control","gatekeeping","longitudinal analysis","credibility classification"],"falsifier":"Take a matched sample of reputable and misinformation sites with similar popularity rankings, content platforms, and languages; if, after adjustment, the difference in DisallowAll-for-AI rates between the two groups is no longer statistically significant, the claim that misinformation per se drives openness is falsified. Alternatively, comparing the same sites' robots.txt files over the claimed timeline and finding that the reputable sites' disallow rates did not rise from roughly 23% in September 2023 to roughly 60% by May 2025 would falsify the longitudinal trend.","tokens_in":15170,"feed_emoji":"🤖","tokens_out":10033,"duration_ms":71781,"temperature":0.7,"pith_summary":"The paper asks whether the robots.txt files that govern AI web crawlers treat reputable news and misinformation sites differently. It finds a stark asymmetry: as of May 2025, 60.0% of reputable news sites disallow at least one AI agent, versus 9.1% of misinformation sites, and the gap has grown from 23% to nearly 60% since September 2023. Reputable sites block an average of 15.5 AI user agents; misinformation sites block fewer than one. The authors also measure active blocking via HTTP User-agent headers, finding both types use it, but reputable sites' behavior matches their declared robots.txt rules far more closely. If this pattern holds, the content most available to AI training crawlers is increasingly low-credibility.","feed_headline":"60% of news sites block AI crawlers, vs 9% of misinformation sites","feed_subtitle":"The content AI crawlers can freely access is increasingly low-credibility, shaping what language models are trained on.","key_machinery":"The central mechanism is the Robots Exclusion Protocol (robots.txt): a voluntary file at a domain root listing User-agent directives and Disallow rules. The paper operationalizes 'AI gatekeeping' as the presence of a DisallowAll rule for any of a curated list of 63 AI user agents (such as GPTBot, CCBot, and Google-Extended), and complements this with active-blocking measurements that vary only the HTTP User-Agent header to see if the server returns non-200 statuses. The longitudinal component uses six archived snapshots at four-month intervals to trace changes in these directives over time.","core_discovery":"The paper's central discovery is that content-access signaling on the web is credibility-stratified: sites rated reputable by the credibility classification the authors adopt are far more likely to declare blanket bans on AI crawlers in their robots.txt files, and to list far more AI agents, than sites rated as misinformation. The largest specific number: 60.0% vs 9.1% for DisallowAll of at least one AI agent; average blocked agents 15.5 vs 0.77. This gap is not explained by popularity alone, and it has widened sharply over the observation period (September 2023 to May 2025). The authors further show that active blocking—refusing HTTP requests when the User-Agent header names an AI crawler—i","pith_inferences":["A controlled comparison matching reputable and misinformation sites on popularity rank, content management system, and language would likely attenuate the 60%-vs-9% gap, because the paper's own popularity check is an unadjusted cumulative distribution; the misinformation effect may partly be an under-resourced-site effect.","The credibility labels may be confounded with political orientation and English-language coverage; the robots.txt differences could track editorial professionalism or legal exposure rather than factual reliability per se.","A testable extension would measure whether per-site inclusion rates in public web crawl corpora correlate with robots.txt AI-blocking, directly linking declared gatekeeping to actual training-data exposure.","Because the paper measures declarations and one pair of active-blocking agents, not the behavior of the full crawler ecosystem, the true training-data skew could be smaller (if crawlers ignore robots.txt) or larger (if crawlers over-comply) than the declared asymmetry suggests."],"forward_implications":["LLM training corpora assembled from publicly accessible web content will contain a disproportionately large share of misinformation-site content, because those sites rarely forbid AI crawlers.","A self-reinforcing loop can emerge: permissive misinformation sites are scraped repeatedly, and their content reappears in AI-generated text, amplifying its reach.","Reputable news sites' growing blocklists (median over 25 agents by early 2025) mean the common pool of freely crawlable text is increasingly low-credibility or non-news material.","Active blocking is an imperfect substitute for robots.txt because it is all-or-nothing per agent and can interfere with legitimate indexing; the divergence between declared and actual blocking on misinformation sites suggests a patchwork rather than coherent policy.","The widening gap indicates that awareness of AI crawling is itself a differentiator: misinformation operators appear largely unresponsive to AI crawler announcements."],"fun_headline_variants":["News sites block AI crawlers 6x more than misinformation sites","60% of news sites ban AI crawlers; only 9% of misinformation sites do","AI crawlers find misinformation sites more open than reputable news","News sites' AI-blocking jumps from 23% to 60% while misinformation stays at 9%","Misinformation sites rarely block AI crawlers; reputable news sites usually do"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire reputable-versus-misinformation comparison rests on the credibility ratings the authors adopt; if those labels are noisy or track political alignment, technical sophistication, or language rather than factual reliability, the robots.txt gap may be about those correlated attributes instead of misinformation itself.","fun_headline_variants_meta":{"raw":{"variants":["News sites block AI crawlers 6x more than misinformation sites","60% of news sites ban AI crawlers; only 9% of misinformation sites do","AI crawlers find misinformation sites more open than reputable news","News sites' AI-blocking jumps from 23% to 60% while misinformation stays at 9%","Misinformation sites rarely block AI crawlers; reputable news sites usually do"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001438,"raw_usage":{"total_tokens":5662,"prompt_tokens":803,"completion_tokens":4859,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":4769}},"tokens_in":547,"tokens_out":4859,"duration_ms":27904,"temperature":1.0,"reasoning_tokens":4769,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T10:17:39.728345+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a matched sample of reputable and misinformation sites with similar popularity rankings, content platforms, and languages; if, after adjustment, the difference in DisallowAll-for-AI rates between the two groups is no longer statistically significant, the claim that misinformation per se drives openness is falsified. Alternatively, comparing the same sites' robots.txt files over the claimed timeline and finding that the reputable sites' disallow rates did not rise from roughly 23% in September 2023 to roughly 60% by May 2025 would falsify the longitudinal trend.","supporting_citations":[],"review_version":1}