{"id":"d67d877e-7ccd-47ff-9d6d-ea876cc67eac","arxiv_id":"2412.09632","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Using unlearning-based ablation and information-leakage tests, the paper finds UK government websites matter for LLM performance on welfare queries while data.gov.uk datasets are not recalled.","lead":"This paper tests whether UK government websites and the data.gov.uk open data portal actually make it into the training corpora of large language models. It finds that government websites contribute to model performance on welfare-related citizen queries, while data.gov.uk datasets are almost completely unrecalled, and it proposes two assessment methods.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The negative data.gov.uk conclusion (KR4) is unsupported: the leakage test's sensitivity is uncalibrated, and the paper's own controls (BOE, POP) fail for Gemma and Qwen.","rationale":"Reading the paper in good faith, its aim is to assess two UK government data sources as inputs to LLM training. The ablation study (KR1-KR3) is internally plausible: the unlearning method appears minimally intrusive, the control query is unaffected, and the release of code and data is a real positive. However, the central claim in the abstract has two halves, and the second half is the more categorical: 'data.gov.uk is not' a data provider. That conclusion is derived from a leakage experiment in which 190 of 195 tests failed to recall target statistics. The logic requires absence of recall to imply absence from training, but the paper's own controls are guaranteed positives and two of the three tested model families fail them. The authors acknowledge this in Section 3.3 and the limitations, but still state KR4 as a key result. This is not a disagreement with external consensus; it is an internal inference problem: the test has not been shown to be sensitive enough to detect known training data, so a null result cannot support a universal negative. The reader's weakest_assumption identifies exactly this issue, and I agree with the REJECT verdict. A calibration experiment—fine-tuning on the selected CSVs and re-running the identical prompts—would settle whether the test can detect data that is definitely in the corpus; until then, the headline negative finding remains unsupported.","tokens_in":13411,"tokens_out":5922,"duration_ms":59427,"concrete_test":"Fine-tune the same Gemma-2-2B and Qwen2.5-3B models (and Llama-3.1-8B as a reference) on the actual CSV rows of the five selected data.gov.uk datasets using a standard next-token objective, then run the exact prompts from Tables 7-9 and Table 11 (templates a-d, 0/1/5-shot, plus the instruct evaluation). Because the data is now guaranteed to be in the training corpus, the exact-recall rate under this protocol measures the method's sensitivity. If recall stays near zero for Gemma/Qwen, the original zero-recall result is uninformative and KR4 must be withdrawn or reframed as 'not extractable by these methods.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"KR4 ('data.gov.uk is not a data provider for AI') is the negative half of the abstract's headline and rests entirely on the information-leakage study in Section 3. The inference from 'LLM did not output the exact statistic' to 'the statistic was not in the training corpus' requires the prompting procedure to be a sensitive detector of memorized numbers. That sensitivity is not established and is contradicted by the paper's own controls: in Table 11, Gemma 2 2B and Qwen 2.5 3B fail to recall the Bank of England base rate and UK population estimates, which are certainly in their pretraining data. The paper concedes this in Section 3.3 and in the limitations section ('poor performance in controls makes the results ... potentially less robust'). When recall fails on guaranteed positives, a null result on data.gov.uk statistics cannot distinguish 'not in training data' from 'present but not extractable by these prompts at this model scale.' The conclusion also generalizes from five datasets to all of data.gov.uk. The ablation-based first finding (websites matter) is less affected by this particular flaw, but the headline cannot stand without the leakage half.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces two methods for assessing whether UK government data sources contribute to LLM training: (1) an ablation study that applies LLM 'unlearning' to remove a small set of government welfare-related websites from five small open-weight models and measures changes in knowledge errors on 18 citizen queries and one control query, and (2) an information leakage study that prompts the same models to recall statistics from five data.gov.uk datasets, with two external controls (Bank of England base rate and UK population estimate). The authors report that government websites are important data sources for LLMs, with heterogeneity across subject matters, and that data.gov.uk is not a data provider for AI. The paper is framed as a technical report, with code and data released and a non-technical ODI companion report. It also claims broader applicability of the two methods for probing opaque training corpora.","tokens_in":13666,"tokens_out":6601,"duration_ms":56047,"significance":"The paper addresses a genuinely important and timely question—whether and how government data enters proprietary and opaque training corpora—and the two proposed methods are creative. The release of code and data, the use of external controls in the leakage study, and the inclusion of a Google-based prevalence check for the ablation analysis are all positive elements. If the methods were rigorously validated, they could provide organisations with a practical framework for evaluating their data contributions to AI. However, the headline negative conclusion about data.gov.uk is not supported by the evidence as presented: the leakage test is uncalibrated and fails on known-positive controls for two of the three model families, and the paper's own limitations section acknowledges this. The ablation findings, while more plausible, also lack basic statistical grounding. The paper's contribution therefore currently rests on a partially unsupported central claim, though the methods themselves are potentially salvageable.","major_comments":[{"comment":"The conclusion that data.gov.uk is not a data provider for AI is not supported by the leakage experiment, because the method's sensitivity is uncalibrated: in Table 11, Gemma 2 2B and Qwen 2.5 3B also fail to recall the Bank of England base rate and the UK population estimate, which are certainly in their pretraining data. The paper itself acknowledges in the limitations (§4) that 'poor performance in controls makes the results of the information leakage study potentially less robust'. A null result in this setting cannot distinguish 'not in the training corpus' from 'present but not extractable by these prompts at this model scale'. The claim also generalizes from five datasets to the whole of data.gov.uk. Please either add positive controls of similar format and obscurity to the data.gov.uk statistics, report sensitivity per model and per prompting condition, and/or weaken KR4 to a claim about the five tested datasets.","section":"§3.3, Table 11; KR4"},{"comment":"The claim that the ablation causes a clear increase in knowledge errors (KR2 and KR3) rests on a single annotator's coding of 19 hand-picked queries with no inter-rater reliability, no error bars, and no statistical test; the reported 42.6% average increase is not accompanied by variance or any significance measure. The assertion that 'all LLMs were fairly homogeneously affected' is based on visual inspection of Figure 2. Please report per-model and per-query counts with appropriate uncertainty and statistical comparisons, and either provide a second annotator or a clearly defined coding protocol to establish reliability.","section":"§2.4.2–2.4.3, Figures 2 and 3"},{"comment":"The claim of a 'significant negative correlation' between the effect of ablation and the Google prevalence measure is unsupported: no correlation coefficient, p-value, or test is reported, and the prevalence measure is not precisely specified (search engine, query string, date, region, and the rule for the 'first 10 non-government websites'). Since KR3* is one of the paper's key findings, this analysis needs to be made reproducible and statistically quantified, or the claim should be softened to a descriptive observation.","section":"§2.5, Figure 4"},{"comment":"The paper reports '5 out of 195 tests' in §3.3, but Table 11 shows 7 datasets × 3 models × 4 prompting conditions = 84 cells; please clarify how the 195 count is obtained and whether multiple prompt templates are being aggregated. In addition, the selection of the five data.gov.uk datasets is described as a 'random sample' without giving the sampling frame or procedure, which is needed to support any generalization to the entire portal.","section":"§3.2 and §3.3, Table 10"}],"minor_comments":[{"comment":"There are several typographical errors: 'analagous' should be 'analogous', 'espsecially' should be 'especially', and reference [19] lists 'Stablility.ai' which should be 'Stability AI'.","section":"§2.2.1 and footnotes"},{"comment":"The symbols used in Table 11 (!, %, 5, etc.) are visually ambiguous; please use explicit check/cross markers or a clear legend with textual labels, and ensure the table is readable without reference to the surrounding text.","section":"Table 11"},{"comment":"The sentence 'in 5 out of 195 tests, tested LLMs simply did not recall data points in data.gov.uk' appears to state the opposite of what the results show; presumably 'did' was intended instead of 'did not'.","section":"§3.3"},{"comment":"The abstract's unqualified statement that 'data.gov.uk is not' a data provider should be tempered to reflect the limitations acknowledged in §4, since the leakage study's sensitivity is not established.","section":"Abstract and §4"},{"comment":"The single control query in the ablation study concerns US welfare; consider adding a UK-related control query outside the target welfare topics to better test whether the unlearning procedure affects general UK knowledge beyond the ablated websites.","section":"§2.4.1, control query"}],"recommendation":"major_revision","confidential_remarks":"The paper is a technical report with a policy-facing companion, and its empirical standards are somewhat informal for a journal publication. The most serious issue is the unsupported negative conclusion about data.gov.uk, which is load-bearing for the abstract. That said, the ablation method and the overall framing are potentially valuable; with a recalibrated leakage test, a more careful statistical treatment, and appropriately qualified claims, the paper could become a useful contribution. The authors should also be asked to reconcile the reported test count with Table 11 and to provide the sampling details for dataset selection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful technical report, but the abstract overstates what the evidence supports. The ablation half largely works; the leakage half does not.\n\nWhat's new: applying Yao et al.'s unlearning to ablate government websites, plus a qualitative coding framework and the empirical result that small LLMs answer welfare queries worse after ablation, with variation by topic and a plausible negative correlation with online prevalence. The release of code and data is real and helpful. The leakage study adapts Wang et al.'s prompting templates to government statistics, but that adaptation is where the trouble starts.\n\nWhere it goes soft. KR4—'data.gov.uk is not a data provider for AI'—rests on interpreting failed recall as absence from training data. The paper's own controls undercut that: Gemma-2-2B and Qwen-2.5-3B fail to recall the Bank of England base rate and UK population, facts certainly in their pretraining data. The authors admit as much in the limitations section: 'poor performance in controls makes the results ... potentially less robust.' When your detector misses known positives, a null result on five datasets cannot separate 'not in training data' from 'present but not extractable by these prompts at this scale.' Generalizing from five datasets to the whole portal is another leap. The ablation study is less affected, but it still has no error bars, no inter-rater reliability, and a hand-picked set of 18 queries plus one control; the reported '42.6% increase' comes without variance. None of these are fatal to the methods as a pilot, but they mean the paper should be framed as an exploratory methods report, not a settled evaluation.\n\nI do not see a circularity problem: the targets are external, the Google-prevalence check is independent, and the controls are sensible in design even if they fail in execution.\n\nWho gets value: people working on open-government-data strategy and researchers building methods to probe training corpora. The careful method description makes it a worthwhile read despite the unsupported headline. It deserves a serious referee, but with the expectation that the authors either calibrate the leakage method or drop/soften KR4.","headline":"A clearly written technical report whose ablation half mostly works, but whose headline claim about data.gov.uk is undercut by the paper's own failed controls.","tokens_in":14157,"tokens_out":2181,"would_cite":false,"duration_ms":20219,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that UK government websites measurably improve LLM answers to citizen queries while data.gov.uk datasets are almost never recalled by the tested models.","keywords":["generative AI","training data transparency","open government data","data.gov.uk","LLM unlearning","information leakage","ablation study","UK government data policy"],"falsifier":"Run the same leakage prompts against statistics that model documentation explicitly lists as training data; if recall is near zero for those known statistics, the method cannot separate 'not trained on' from 'not prompted successfully.' A simpler version is to re-ask the control statistics (central bank base rate, national population) in several phrasings with models that failed the controls: persistent failure would show the leakage test is too insensitive to support the paper's negative conclusion about the open data portal.","tokens_in":13241,"feed_emoji":"📊","tokens_out":10186,"duration_ms":83086,"temperature":0.7,"pith_summary":"Generative AI training corpora are kept secret, so governments cannot tell whether the data they publish is actually shaping AI models. This paper builds two indirect probes for that question and applies them to the UK government: a 'forgetting' test that removes UK government web pages from a small LLM, and a recall test that asks models to reproduce statistics from the data.gov.uk portal. Its central finding is that UK government websites are a real training-data asset for welfare-related citizen queries, with the strongest effects on topics that are thinly covered by non-government web sources. Its second finding is that the structured datasets on data.gov.uk are almost never recalled by the tested models, and the paper concludes that the portal is not currently a data provider for AI. If these results hold, governments that want to feed AI should pay attention to plain-text web publication as much as to open data portals.","feed_headline":"UK government websites feed AI models; data.gov.uk does not","feed_subtitle":"Ablation tests show welfare pages improve LLM answers, while recall probes find open-data portal stats barely surface.","key_machinery":"The load-bearing object is a two-part audit method. The ablation probe uses LLM unlearning: a model is reverse-fine-tuned so that its loss on a target corpus of UK government welfare pages increases, while a Kullback-Leibler divergence term keeps its loss flat on a safe corpus of general encyclopedic text, so the model forgets the government pages without losing language ability; the comparison of error counts before and after quantifies how much those pages contributed. The leakage probe adapts a known prompting framework: the model is asked to complete a statistic in four templates, under zero-, one-, and five-shot prompting, and an instruct-tuned variant is asked the statistic directly, with correct recall treated as evidence that the statistic was in training data. The experiments use small open-weight models and manually coded evaluation of structural and knowledge errors.","core_discovery":"The paper's central discovery is a contrast between two UK government data channels. After an unlearning-based ablation makes a model forget roughly twenty government welfare pages, knowledge errors on an 18-question citizen-query test rise by an average of 42.6% while fluency and formatting errors stay flat; the paper reads this as evidence that those web pages were doing real work in the model's answers. The effect is uneven across questions, and the paper finds a negative correlation between how much a query degrades after ablation and how often non-government websites can answer the same query, so government web data matters most where alternative online coverage is scarce. The second result comes from an information-leakage test: across five data.gov.uk datasets and three small models, almost none of the statistics are recalled, despite prompts modelled on known leakage methods, and the paper concludes that the portal's datasets are not part of the tested models' training corpora.","pith_inferences":["Editorial inference: the negative result for data.gov.uk may be a property of the small models and prompt formats tested, not of all LLMs or of retrieval-augmented systems, which could still use open data even if it is not memorised.","Editorial inference: the leakage method's failure on controls suggests that 'not recalled' should not be read as 'absent from training' without a sensitivity calibration on statistics known to be in the training corpus.","Editorial inference: a testable extension would be to convert selected data.gov.uk statistics into prose articles and check whether leakage scores rise, isolating format rather than content as the reason the portal is not feeding models."],"forward_implications":["If UK government websites are in LLM training corpora, then the government already influences AI behaviour through its ordinary web content, and changes to that content such as removals, rewrites, or paywalls will shift how models answer citizen queries.","Because the ablation effect is concentrated on topics with little non-government coverage, improving government prose on benefits interactions and eligibility details is the most direct way to improve LLM performance where it currently fails.","If data.gov.uk datasets are not being recalled, publishing statistics as spreadsheets alone is not an effective way to get numbers into AI training; the portal's current format and discoverability would need to change.","The two methods together give any data-holding organisation a reusable way to audit whether its published data and prose are part of AI training mixtures."],"supporting_citations":[{"why":"supplies the LLM unlearning method (gradient ascent on target data with a KL safeguard) used to ablate government websites.","marker":"[9]"},{"why":"establishes that the ablated websites were present in web-crawled training corpora before the tested models' training windows.","marker":"[14]"},{"why":"supplies the 0-/1-/5-shot prompting framework and four templates that the information-leakage study adapts.","marker":"[21]"},{"why":"provides the training-data extraction prompting technique that motivates treating successful recall as evidence of training presence.","marker":"[20]"},{"why":"frames the motivating problem: AI training corpora are opaque, so indirect assessment methods are needed.","marker":"[1]"},{"why":"supplies the Kullback-Leibler divergence used in the unlearning objective to keep non-target knowledge intact.","marker":"[12]"},{"why":"provides the primary tested model family used in the ablation experiment.","marker":"[15]"}],"fun_headline_variants":["LLMs lean on gov websites, not data.gov.uk stats","Forgetting 20 gov pages boosts LLM errors 42.6%; data.gov.uk ignored","Welfare pages matter for AI; data.gov.uk datasets don't","Unlearning welfare pages hurts LLM queries; data.gov.uk never recalled"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole negative claim about data.gov.uk depends on treating a model's failure to volunteer a statistic as proof that the statistic was not in its training data, and the control questions show that some of the same models could not recall even widely known figures such as the central bank base rate.","fun_headline_variants_meta":{"raw":{"variants":["LLMs lean on gov websites, not data.gov.uk stats","Forgetting 20 gov pages boosts LLM errors 42.6%; data.gov.uk ignored","Welfare pages matter for AI; data.gov.uk datasets don't","Unlearning welfare pages hurts LLM queries; data.gov.uk never recalled"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000788,"raw_usage":{"total_tokens":3530,"prompt_tokens":1053,"completion_tokens":2477,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":2396}},"tokens_in":669,"tokens_out":2477,"duration_ms":102819,"temperature":1.0,"reasoning_tokens":2396,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:56:17.422087+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same leakage prompts against statistics that model documentation explicitly lists as training data; if recall is near zero for those known statistics, the method cannot separate 'not trained on' from 'not prompted successfully.' A simpler version is to re-ask the control statistics (central bank base rate, national population) in several phrasings with models that failed the controls: persistent failure would show the leakage test is too insensitive to support the paper's negative conclusion about the open data portal.","supporting_citations":[{"cited_title":"We Must Fix the Lack of Transparency Around the Data Used to Train Foundation Models","cited_arxiv_id":null,"evidence_quote":"frames the motivating problem: AI training corpora are opaque, so indirect assessment methods are needed."}],"review_version":1}