{"id":"11795f4a-bebc-4f5b-ad7d-4a090f66c9f1","arxiv_id":"2505.15837","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across the February 2021 web crawl, over 90 million links point to Wikipedia, mostly from news and science sites, placed in main content, and used as explanations rather than as evidence.","lead":"This paper measures where and why external websites link to Wikipedia, using billions of pages from the February 2021 Common Crawl archive. It counts over 90 million Wikipedia links across 1.68% of web domains, mostly inside article text, and mostly used for explanation rather than evidence.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 95% delegation figure is the linchpin of the 'background knowledge' claim and rests on 500 manually annotated links with no reported inter-annotator reliability or stratified validation.","rationale":"The paper has real strengths: it releases a large, reproducible extraction from Common Crawl, reports precision/recall for the segmentation rules (Table 4), and its descriptive results on domains and segments are plausible. My stress-test centers on Section 6 because the abstract's third finding and the 'knowledge infrastructure' interpretation depend on the 95%/5% split. The reader's CONDITIONAL verdict is appropriate. Sampling 500 units from a population of 48.8 million links can support a rough proportion if the sample is uniform and coding is reliable, but neither condition is currently verifiable: the text alternates between '500 webpages' and '500 Web links' (Section 6 vs Appendix C), no kappa is reported, and the coding categories admit a grey zone. A replication/stratified annotation study is the cheapest way to settle this. I therefore keep the verdict unchanged rather than moving to reject, because the qualitative pattern (delegation is common) is not contradicted by any internal evidence and is likely robust even if the exact 95% is not.","tokens_in":16019,"tokens_out":5415,"duration_ms":53376,"concrete_test":"Run an independent annotation study: draw 2,000 English Wikipedia links via stratified random sampling (strata: main content/boilerplate/user-generated, plus domain-rank quartiles), have two annotators blind to the paper's hypothesis apply the Section 6 definitions to rendered screenshots, and report Cohen's kappa and the stratified delegation proportion with 95% confidence intervals. If kappa < 0.6 or the CI lower bound drops below 0.90, the 95% delegation claim should be revised; if kappa is high and the CI contains 0.95, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that most Wikipedia links are explanatory delegations rather than evidence is generalized in Section 6 from a manual annotation of '500 webpages' (Section 6) / '500 Web links' (Appendix C) to all 48.8 million English links. Three specific issues make this estimate load-bearing: (1) The sample unit is ambiguous—if 500 webpages were sampled and all links on those pages annotated, the effective sample size is not 500 independent links; if 500 links were sampled, the text in Section 6 is wrong. (2) No inter-annotator agreement, codebook excerpt, or adjudication procedure is reported, so the binary delegation/evidence split may not be reliably reproducible; Appendix C itself concedes that intent is difficult to infer without surveys and that the authors reframed the taxonomy as a 'relationship' classification. (3) No confidence intervals or stratified analysis by segment or domain tail are given, yet the proportions are multiplied by the full link count to produce 46.4 million vs 2.4 million. The paper's key qualitative conclusion may survive a revised estimate, but the exact 95% figure and the quantitative generalization are not supported as reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Web2Wiki, a dataset of Wikipedia links extracted from the February 2021 Common Crawl dump, containing 90,805,367 links to Wikipedia articles across 1.68% of crawled domains, with a detailed analysis of the 48.8 million English Wikipedia links. The authors address three questions: which articles are linked (compared with Wikipedia internal in-degree and Reddit citations via ORES topics), where links appear (domain topics via Homepage2Vec and in-page segmentation into boilerplate, main content, and user-generated sections), and why links are made (manual annotation into delegation versus evidence). The headline findings are that Wikipedia links are concentrated in news and science sites, that 91% of links appear in main content, and that 95% of links serve explanatory delegation rather than evidence, supporting the interpretation of Wikipedia as background knowledge infrastructure. The dataset and code are promised for public release.","tokens_in":16237,"tokens_out":4479,"duration_ms":43677,"significance":"If the findings hold, this is the first web-scale characterization of how Wikipedia is referenced across the general Web, going beyond the existing platform-specific studies on Reddit and search engines. The paper's strengths include the large-scale extraction from Common Crawl, the use of external baselines (Wikipedia in-degree, ORES, Homepage2Vec, Reddit) that avoid circularity, the bootstrapped confidence intervals in the topic comparisons, and the planned public release of a reusable dataset and code. The qualitative conclusion that Wikipedia predominantly functions as a delegation target rather than an evidence source is plausibly correct and would be a meaningful contribution to the field. However, the quantitative support for the two headline percentages (91% main-content and 95% delegation) is weaker than the paper presents: the main-content figure depends on a rule-based segmenter with low recall on user-generated content, and the delegation figure comes from a 500-item manual annotation with no reported reliability or sampling detail. These issues are fixable but require additional evidence or a recalibration of the claims.","major_comments":[{"comment":"The 95% delegation versus 5% evidence estimate is load-bearing for the abstract, introduction, and conclusion, but the generalization from the manual annotation to all 48.8 million English links is not supported as reported. Section 6 says the authors 'randomly sampled 500 webpages that link to Wikipedia,' while Appendix C says they sampled '500 Web links'; these are different sampling units. Moreover, the text generalizes the resulting proportions specifically to 'the 48.8 million Wikipedia links in the main content,' even though the annotation sample is not described as restricted to main content. The paper reports no confidence intervals, no inter-annotator agreement, no codebook excerpt, and no stratification by segment or domain tail, yet it multiplies the 95/5 split to obtain 46.4 million delegation and 2.4 million evidence links. Appendix C itself concedes that intent is difficult to infer from links and that the taxonomy was reframed as a relationship classification. The authors should clarify the sampling unit, report reliability and confidence intervals, and either restrict the generalization claim to the population actually sampled or temper the headline claim.","section":"Section 6 / Appendix C"},{"comment":"The claim that 91% of Wikipedia links appear in main content relies on the rule-based segmentation whose evaluation in Appendix B reports recall of only 0.57 for user-generated content and precision of 0.65 for boilerplate. With user-generated content recall at 0.57, the measured 2% share for this segment is likely an underestimate, and the main-content share of 91% may be correspondingly overstated. The paper does not report a sensitivity analysis or discuss how misclassification would affect Table 2. Additionally, Table 2's column header 'Webpages (M)' is inconsistent with the values: 43.7 million main-content webpages exceeds the 14.46 million total linking webpages reported in Table 1, so the column must refer to links rather than webpages. The authors should correct the table label and provide a robustness check on the segment percentages given the measured precision/recall.","section":"Section 5 / Appendix B / Table 4"},{"comment":"There is a mismatch between the paper's framing of RQ3 ('Why do people reference Wikipedia articles?') and the operational definition in Appendix C, which states that the authors moved away from intent inference and instead classified the 'relationship between Wikipedia and the site that invokes it.' The main text nonetheless presents delegation and evidence as motivations and uses this to support the conclusion that Wikipedia serves as background knowledge infrastructure. The authors should consistently describe the taxonomy as a functional or relational classification of link usage, not as a direct measurement of author intent, and should state the associated limitations explicitly in the main text rather than only in an appendix.","section":"Introduction / Section 6 / Appendix C"}],"minor_comments":[{"comment":"The dataset name is inconsistently capitalized as 'Web2Wiki' in the abstract and 'WEB2WIKI' in the body; please standardize.","section":"Abstract / throughout"},{"comment":"The text 'we use Wikipedia in degrees' should read 'Wikipedia in-degrees.'","section":"Section 3"},{"comment":"The phrase 'bucket almost all the Wikipedia sharing links (98%)' is awkward; consider 'classify 98% of the links.'","section":"Section 6"},{"comment":"The x-axis label 'Difference in probabilities' is vague; it should state that the difference is Web (or Reddit) probability minus Wikipedia-internal probability.","section":"Section 4 / Figure 1"},{"comment":"The Pearson correlation r=0.5 is quoted without clarifying whether the correlation is computed on raw counts, log-counts, or probability vectors; please specify the exact variables.","section":"Section 4"},{"comment":"In the phrase 'Raphael’s School of Athens frequently appear,' the singular/plural agreement should be fixed ('frequently appears').","section":"Section 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within ICWSM scope and the dataset release is a valuable contribution. The main concerns are the unsupported generalization of the 95% delegation figure and the imperfect segmentation evaluation behind the 91% main-content figure; both are fixable with additional analysis and careful reframing. I would also encourage the authors to provide at least a sample of the dataset or a preprint release before acceptance to allow independent checking of the reported counts, since the current 'upon acceptance' release delays reproducibility verification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this paper gives you the first large-scale map of where and how Wikipedia gets linked from the wider web, and it releases the dataset. That's real value. The 90M-link extraction from Common Crawl, the 1.68% of domains, the breakdown by ORES topic, and the within-page segmentation are all new measurements, and the figures come with bootstrapped intervals. The segmentation rules are simple but they report precision/recall (main content 0.91/0.96), so the 91% main-content result is on reasonable footing.\n\nThe soft spot is the 'why' section. The 95% delegation vs 5% evidence split is the headline claim in the abstract, but it rests on 500 manually annotated links. No inter-annotator agreement, no confidence intervals, and the sample unit is ambiguous: Section 6 says '500 webpages' and Appendix C says '500 Web links.' On top of that, Appendix C admits the authors moved away from intent to a relationship taxonomy after finding intent hard to infer. That is honest, but it means the 95% number is really an estimate of a relationship category, not a motivation, and generalizing it to 48.8M links is a stretch as reported. The qualitative conclusion – most external Wikipedia links are explanatory references rather than evidence – is plausible and probably survives, but the exact figure needs either a stratified sample, an inter-annotator reliability check, or a much more careful statement of what was measured.\n\nThe 'first dataset' framing is also a bit strong given Alshomary et al. 2019 extracted Wikipedia references from Common Crawl, though their purpose was plagiarism detection, so the characterization is still new.\n\nNo circularity: the baselines (in-degree, ORES, Homepage2Vec) are external, and the manual labels are not fitted to reproduce the conclusions. That's clean.\n\nWho is this for? Anyone working on Wikipedia's external influence, web citation practices, or link prediction. It deserves a serious referee; the dataset alone makes it citable. I'd send it to review but ask for a revision that either tightens the generalization from the 500-link sample or reframes the claim as a qualitative finding. I would not desk-reject this.\n\nBest.","headline":"A genuinely useful web-scale measurement paper whose headline 95% figure is softer than it looks, but the dataset and the topic/placement analyses carry it.","tokens_in":16738,"tokens_out":2000,"would_cite":true,"duration_ms":18730,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 90.8 million web links, the paper finds English Wikipedia is used mainly to delegate explanations—background context—rather than to cite evidence, making Wikipedia web-scale knowledge infrastructure.","keywords":["Wikipedia","hyperlink analysis","web-scale measurement","delegation vs evidence","knowledge infrastructure","webpage segmentation","reference linking","Common Crawl"],"falsifier":"Take a fresh, domain-stratified sample of several thousand Wikipedia links from the released Web2Wiki dataset, have independent annotators classify each as delegation or evidence, and measure inter-annotator agreement. If the delegation share falls well below 95%, or if evidence links dominate on high-link domains such as news and science, the global generalization fails.","tokens_in":15851,"feed_emoji":"🌐","tokens_out":8940,"duration_ms":85157,"temperature":0.7,"pith_summary":"Using a public archive of 2.7 billion web pages, the paper builds the Web2Wiki dataset of 90.8 million links to Wikipedia, reaching 1.68% of web domains. For English Wikipedia it claims the dominant linking pattern is delegation: 95% of sampled links point readers to Wikipedia for background explanation rather than to support a claimed fact. It also finds that 91% of links appear in main page content, and that news, science, and society sites link Wikipedia far more often than business and shopping sites, which tend to relegate links to boilerplate. If these figures hold, Wikipedia functions on the web mainly as background knowledge infrastructure—a shared explanatory layer—rather than as a source of proof. The released dataset lets others test and extend that picture.","feed_headline":"Wikipedia's web role is mostly background explanation, not evidence","feed_subtitle":"A 90.8-million-link study finds 95% of sampled references explain rather than cite.","key_machinery":"The load-bearing mechanism is the Web2Wiki dataset: a regex-based pass over all <a> tags in a web archive, producing link–domain–webpage records for 90.8 million Wikipedia links. Three instruments carry the analysis: Wikipedia's internal in-degree, computed from an XML dump, serves as the importance baseline; a rule-based page segmenter, iteratively developed on 500 inspected pages, labels each link's location as boilerplate, main content, or user contributions (main content precision 0.91, recall 0.96); and a two-category manual taxonomy distinguishes 'delegation' (link offers background explanation) from 'evidence' (link supports a fact or supplies sourced media). The hand annotation of 500 sampled webpages is what produces the 95/5 split that the paper generalizes to all English main-content links.","core_discovery":"On its own terms, the paper's discovery is that Wikipedia's external web footprint is enormous but overwhelmingly explanatory rather than evidential. Extracting every hyperlink to Wikipedia from a 2.7-billion-page public web archive yields 90,805,367 links, of which 48,829,702 point to English Wikipedia from 14,462,267 distinct webpages across 940,239 domains; 80% of those links target only 9% of English articles, and 58% of English articles receive at least one external link. External linking correlates with Wikipedia's own in-degree ($r = 0.5$, $p < 0.001$) but differs by topic from both Wikipedia-internal linking and Reddit sharing. Within pages, 91% of links sit in main content, 7% in boilerplate, and 2% in user-generated sections. Manual annotation of 500 randomly sampled links places 95% in a 'delegation' category—sending readers elsewhere for context—and 5% in an 'evidence' category, from which the paper estimates 46.4 million delegation links and 2.4 million evidence links among English main-content references.","pith_inferences":["Beyond the paper: the delegation/evidence split is likely not uniform across languages; the released multilingual data could be tested for higher evidence shares in legal, medical, or other contexts with stronger attribution norms.","Beyond the paper: if delegation dominates, Wikipedia's influence on public understanding flows through unseen background links, implying that pageview metrics substantially understate Wikipedia's reach.","Beyond the paper: the same link-extraction and two-category coding approach could be applied to other reference sites, such as dictionaries or specialized encyclopedias, to test whether explanatory delegation is a general web pattern or specific to Wikipedia.","Beyond the paper: pairing this single web-archive snapshot with later crawls could measure how quickly Wikipedia links propagate through news cycles and whether the evidence share rises around controversial or rapidly developing topics."],"forward_implications":["If the 95/5 split holds, the practical value of Wikipedia to the web is mostly as a shared explanatory layer: sites offload definitions and background to it, so its quality affects comprehension across millions of pages.","News, science, and society sites are the main carriers of this infrastructure, while business and shopping sites concentrate their links in boilerplate, so Wikipedia's reach is not uniform across web genres.","External web linking tracks Wikipedia's internal importance ($r = 0.5$) more closely than Reddit's sharing patterns, so social media should not be treated as a proxy for how the wider web uses Wikipedia.","The Web2Wiki dataset makes it possible to quantify Wikipedia's external reach article by article, enabling gap and bias audits that go beyond Wikipedia's own internal link network."],"supporting_citations":[{"why":"Frames Wikipedia's relationships with other large online communities and the 'paradox of reuse' that this paper extends to the whole web.","marker":"Vincent, Johnson, and Hecht 2018"},{"why":"Establishes Wikipedia's interdependence with search engines, motivating a web-scale measurement of external references.","marker":"McMahon, Johnson, and Hecht 2017"},{"why":"Supplies the Reddit dataset used to compare social-web linking patterns against the broader web.","marker":"Baumgartner et al. 2020"},{"why":"Provides the ORES topic model used to categorize linked Wikipedia articles into 64 topics.","marker":"Halfaker and Geiger 2020"},{"why":"Provides the webpage topic classifier used to compare Wikipedia-linking domains with a random sample of the web.","marker":"Lugeon, Piccardi, and West 2022"},{"why":"Supplies prior shallow-feature boilerplate detection that motivates the segmentation approach used here.","marker":"Kohlschütter, Fankhauser, and Nejdl 2010"},{"why":"Justifies using Wikipedia-internal in-degree as an approximation of article importance.","marker":"Fortunato et al. 2006"},{"why":"Provides the Wikipedia revision-history data from which internal in-degrees are computed.","marker":"Mitrevski, Piccardi, and West 2020"},{"why":"Prior work using Wikipedia in-degree and PageRank to score entity importance, supporting the baseline choice.","marker":"Thalhammer and Rettinger 2016"}],"fun_headline_variants":["95% of Wikipedia links just explain, not verify","Wikipedia's web role: 90M links, mostly background","News and science lean on Wikipedia—for explanation, not proof","Study: Wikipedia cited as context, rarely as evidence","Most Wikipedia links delegate readers, don't prove facts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 95% delegation figure rests on 500 manually annotated webpages being representative of all 48.8 million English Wikipedia links; if that sample skews toward explanation-heavy pages, the headline split does not generalize.","fun_headline_variants_meta":{"raw":{"variants":["95% of Wikipedia links just explain, not verify","Wikipedia's web role: 90M links, mostly background","News and science lean on Wikipedia—for explanation, not proof","Study: Wikipedia cited as context, rarely as evidence","Most Wikipedia links delegate readers, don't prove facts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000952,"raw_usage":{"total_tokens":4066,"prompt_tokens":958,"completion_tokens":3108,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":3028}},"tokens_in":574,"tokens_out":3108,"duration_ms":23852,"temperature":1.0,"reasoning_tokens":3028,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:47:27.934331+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fresh, domain-stratified sample of several thousand Wikipedia links from the released Web2Wiki dataset, have independent annotators classify each as delegation or evidence, and measure inter-annotator agreement. If the delegation share falls well below 95%, or if evidence links dominate on high-link domains such as news and science, the global generalization fails.","supporting_citations":[],"review_version":1}