{"id":"89592a4d-5c2a-4072-9143-3283660701ee","arxiv_id":"2508.03828","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"MegaWika 2 expands the MegaWika dataset to six times more articles and twice as many scraped citations, with source texts stored inline at precise character offsets.","lead":"MegaWika 2 is a large multilingual dataset that pairs Wikipedia articles with the web sources they cite. It is meant to help fact-checking and research on how claims and citations differ across languages and over time.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim of precise character offsets is unvalidated; the dataset's fact-checking utility depends on a reported offset-accuracy evaluation.","rationale":"The reader's weakest assumption correctly identifies the unvalidated precision of character offsets as the load-bearing condition. I agree with the UNVERDICTED verdict: with only the abstract available, there is insufficient evidence to accept the dataset's central claim. The concern is not that the offsets are wrong, but that the paper does not report any measurement of their correctness. This is a testable empirical claim, not an internal inconsistency. The proposed concrete check would settle the concern: if a stratified sample shows exact or near-exact offsets across languages, the claim is credible; if errors are frequent or language-dependent, the claim must be weakened. Because no full-text evidence was provided for this review, keeping the verdict UNCHANGED is the honest outcome, and the concrete check should be requested from the authors or performed on the released dataset.","tokens_in":571,"tokens_out":2444,"duration_ms":27839,"concrete_test":"Sample at least 100 articles per language from a stratified set of languages included in the dataset. For each sampled citation, take the stored character offset, extract the surrounding article span and the corresponding stored source span. Independently retrieve an archived snapshot of the cited URL from the Wayback Machine or CDX API dated near the original scrape date, and verify: (1) the stored source span occurs verbatim in the archived page, (2) the article span matches the stated Wikipedia revision, and (3) the stored offset equals the character index of the citation in the article text. Report exact-match accuracy and a histogram of offset errors (e.g., 0, ±1, ±10, ±100 characters) per language. If exact-match accuracy falls below a pre-specified threshold or errors concentrate in particular languages, the 'precise character offsets' claim as stated is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the assertion that MegaWika 2 stores scraped source texts inline with 'precise character offsets' of citations in article text. This precision is load-bearing for fact-checking: any downstream use assumes the extracted span reliably identifies the exact sentence or phrase that a Wikipedia citation supports. The abstract provides no validation methodology, no error analysis, and no comparison against ground truth. The offset may be exact in principle, but the paper gives no evidence that the scraping and alignment pipeline achieves exactness in practice. Because the dataset is multilingual, offset errors may also vary systematically by script, tokenization, or URL availability, so a single global accuracy number would not suffice. Without a reported evaluation, the dataset's headline property is an unverified design goal rather than a demonstrated result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces MegaWika 2, a multilingual dataset of Wikipedia articles paired with their citations and scraped web sources. The abstract states that scraped source texts are stored inline with 'precise character offsets' of citations, that the dataset contains six times as many articles and twice as many fully scraped citations as the original MegaWika, and that it is designed for fact-checking and cross-lingual temporal analysis. No other content is provided; the submitted text consists of the abstract alone.","tokens_in":682,"tokens_out":2729,"duration_ms":27088,"significance":"If the offset-precision claim holds, MegaWika 2 would be a substantial resource for fact-checking research, combining large-scale multilingual Wikipedia data with direct alignment of claims to source spans. The reported scale increase over MegaWika is noteworthy, and the provenance of scraped content is a valuable feature. The significance is currently conditional, however: the central utility of the dataset depends on the accuracy of the character offsets and the fidelity of the scraped source texts, and the manuscript provides no evidence for either.","major_comments":[{"comment":"The central claim that source texts are stored with 'precise character offsets' is load-bearing for the fact-checking use case, yet the manuscript reports no validation of offset accuracy, no comparison to ground truth, and no error analysis by language or script. Without such an evaluation, readers cannot assess whether the alignment is reliable enough for downstream fact-checking.","section":"Abstract"},{"comment":"The scale comparisons ('six times as many articles' and 'twice as many fully scraped citations') are not accompanied by any dataset statistics, counts per language, or a definition of what counts as a 'fully scraped' citation. The claims are therefore unverifiable and should be supported by a detailed table in the full paper.","section":"Abstract"},{"comment":"The submission contains only an abstract; there is no description of the scraping, parsing, or alignment pipeline. For a dataset-release paper, this is a critical omission because the reproducibility and quality assessment of the dataset require the pipeline details, including handling of redirected URLs, encoding normalization, and text normalization across languages.","section":"Full text (missing)"}],"minor_comments":[{"comment":"The phrase 'support report generation research ; whereas' contains a stray space before the semicolon; please correct the punctuation.","section":"Abstract"},{"comment":"The terms 'fully scraped citations' and 'precise character offsets' are not defined; please provide precise definitions in the body of the paper.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript as submitted is an abstract only, which is unusual for a dataset-release paper. I recommend requiring a full submission with the pipeline description and an offset-accuracy evaluation before making a decision. The journal may also want to verify that the dataset will be made publicly available with a proper license."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou asked about MegaWika 2. The thing to know: this is an abstract, not a citable result yet. The only claims you can check from the abstract are scale and design intent, and those look reasonable. The load-bearing claim—precise character offsets of citations—is asserted, not demonstrated.\n\nWhat's new: MegaWika 2 is a legitimate extension of the original MegaWika. Six times as many articles and twice as many scraped citations with inline offsets is a real jump for a multilingual fact-checking resource. The design goal of supporting fact checking and cross-language/cross-time analysis is sensible and fills a gap that the original, aimed at QA and retrieval, didn't serve as well.\n\nThe soft spot is exactly the one flagged in the stress-test: offset precision. For fact-checking, the offset has to reliably identify the span a citation supports. The abstract doesn't say how offsets were computed, whether they were checked against ground truth, or whether accuracy holds across scripts and tokenizers. That matters: errors in Latin-script English may not generalize to Arabic or CJK. So the stress-test concern is valid, but it is about missing evidence, not a known flaw. If the full paper reports an offset-accuracy evaluation with per-language breakdowns, the concern dissolves. I cannot judge that from the abstract.\n\nCitation pattern: nothing to flag; self-citation to MegaWika is appropriate.\n\nBottom line: the paper deserves a serious referee if the full version includes the pipeline and validation. I would not cite the dataset in my own work until I've seen that evaluation, but I would bring the full paper to reading group when it's available.","headline":"Scale-up is real; offset precision is asserted but unvalidated in the abstract—hold the citation until the full paper shows per-language accuracy.","tokens_in":1150,"tokens_out":2249,"would_cite":false,"duration_ms":25444,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MegaWika 2 is a multilingual Wikipedia dataset that bundles each article with scraped source texts stored at precise citation offsets, covering six times as many articles and twice as many fully scraped citations as MegaWika.","keywords":["MegaWika 2","Wikipedia","multilingual dataset","citation alignment","fact checking","scraped web sources","character offsets","report generation"],"falsifier":"Take a random sample of MegaWika 2 records, re-fetch the cited URLs or consult archived snapshots from the scraping date, and check whether each stored source text appears at the recorded character offset in the corresponding article; a nontrivial failure rate would show the inline alignment is not as precise as claimed.","tokens_in":405,"feed_emoji":"📚","tokens_out":4388,"duration_ms":47867,"temperature":0.7,"pith_summary":"MegaWika 2 is a multilingual dataset built for fact checking: it pairs Wikipedia articles with the web pages their citations point to, storing the article text, the citation locations, and the scraped source text in one structure. The central claim is that each citation is aligned to a precise character offset in the article and backed by the full text of the cited page, so a fact-checking system can go straight from a statement to the passage that supposedly supports it. The authors report that MegaWika 2 covers six times as many articles and includes twice as many fully scraped citations as the original MegaWika, and that it is designed for fact checking and analyses across time and language rather than only retrieval and report generation.","feed_headline":"MegaWika 2 grows sixfold and doubles fully scraped citations","feed_subtitle":"New corpus pairs Wikipedia citations with scraped source text at exact offset for fact checking.","key_machinery":"The load-bearing object is the aligned citation record: for each citation in a Wikipedia article, the dataset stores a character offset range within the article and the scraped full text of the cited web page inline. This representation removes the need to fetch and parse a URL before checking a claim; the evidence text is already present, positioned against the exact spot in the article where the citation occurs.","core_discovery":"MegaWika 2 is presented as a major upgrade to MegaWika. The dataset represents each Wikipedia article in a rich data structure, and for each citation it stores the scraped text of the cited source inline, along with the exact character offsets telling where the citation appears in the article. On the paper's terms, this turns the corpus into a directly checkable record of what source material stood behind each statement. The reported scale is six times as many articles as MegaWika and twice as many fully scraped citations, with the stated purpose of supporting fact checking and analyses of Wikipedia's sources across time and language.","pith_inferences":["A natural next step not taken up in the paper is a validation study of the offsets; if the alignment is accurate even most of the time, the dataset becomes a ready-made benchmark for citation-precision tasks.","Because scraped source text is stored inline, the dataset could serve as a historical snapshot: diffing the stored text against today's live pages would reveal how cited sources change or vanish, enabling a source-decay analysis beyond what the paper describes.","The same aligned structure could train a model to write claim sentences whose evidence is already attached, since every article provides many statement–source pairs."],"forward_implications":["A fact-checking system can move directly from a Wikipedia sentence to the stored source passage at its cited offset, skipping web retrieval during inference.","The corpus supports large-scale, cross-lingual studies of how Wikipedia citations and their underlying web sources shift over time.","Report generation can use the inline source texts as the ground material for producing summaries or claims with visible provenance.","Researchers can use the character offsets to extract exact claim–source pairs for training or evaluation of grounded text generation."],"supporting_citations":[],"fun_headline_variants":["MegaWika 2 maps each citation to exact source text offsets","Sixfold more articles, double scraped sources for fact-checking","Fact-check Wikipedia's sources with MegaWika 2's exact offsets","MegaWika 2: six times the articles, twice the source scrapes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each stored character offset really points to the cited passage and each scraped source text truly matches the cited page as it then existed, and the paper reports no measurement of how often that alignment fails.","fun_headline_variants_meta":{"raw":{"variants":["MegaWika 2 maps each citation to exact source text offsets","Sixfold more articles, double scraped sources for fact-checking","Fact-check Wikipedia's sources with MegaWika 2's exact offsets","MegaWika 2: six times the articles, twice the source scrapes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000832,"raw_usage":{"total_tokens":3546,"prompt_tokens":772,"completion_tokens":2774,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":388,"completion_tokens_details":{"reasoning_tokens":2693}},"tokens_in":388,"tokens_out":2774,"duration_ms":25208,"temperature":1.0,"reasoning_tokens":2693,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:11:32.918386+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of MegaWika 2 records, re-fetch the cited URLs or consult archived snapshots from the scraping date, and check whether each stored source text appears at the recorded character offset in the corresponding article; a nontrivial failure rate would show the inline alignment is not as precise as claimed.","supporting_citations":[],"review_version":1}