{"id":"b5e20b7e-0145-4d2e-8b4e-3794fc3d9d9e","arxiv_id":"2607.04543","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"By 2026 four of ten U.S. and PRC government document streams show statistically significant AI-writing signals after near-zero 2021 baselines, with country-specific concentration patterns.","lead":"A pilot study finds rising traces of AI-assisted writing in U.S. and PRC government documents by 2026, after near-zero 2021 baselines. The work proposes this as a cheap external monitoring signal for government AI adoption.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Detector validity for Chinese bureaucratic genres is the load-bearing soft spot for the four-source significance claim.","rationale":"The reader correctly identified the central soft spot: Pangram’s accuracy and cross-lingual calibration for these government genres. That assumption is load-bearing for both the statistical claim (four sources significant) and the interpretive claim (U.S. downstream vs. PRC closer-to-policy). The pilot’s design is otherwise careful—pre-LLM baselines, bootstrap CIs, transparent source descriptions, honest limitations—so the concern does not justify rejection; it justifies the existing CONDITIONAL verdict and the call for independent validation and data release. No stronger internal inconsistency or mathematical flaw is present. The concrete multi-detector + human check would settle whether the concern lands without requiring new data collection beyond the already-scraped corpus.","tokens_in":36386,"tokens_out":630,"duration_ms":7164,"concrete_test":"Independently re-score the 2021 and 2026 documents from CAC, MOST, PLA Daily, and Military Review with at least one open detector (e.g., Binoculars or DetectGPT-style) plus a small human expert panel (native Chinese + English government-writing annotators) on a stratified sample of high- and low-fraction_ai windows. If the rank-order of 2026 means or the set of sources whose CIs exclude zero changes under the alternative detectors/human labels, the four-source claim and the U.S./PRC concentration pattern weaken.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that four of ten sources show statistically significant AI-assisted writing by 2026 (with 2021 near-zero baselines), and that the U.S. signal sits downstream of policy while the PRC signal sits closer to it. That claim rests almost entirely on Pangram’s per-document fraction_ai being a valid, cross-lingually calibrated measure of LLM assistance in these specific genres. The paper itself flags that independent Chinese validation is limited (§2 Limitations; Method §1). The four sources that drive the claim are CAC (0.13 [0.07,0.19]), MOST (0.11 [0.02,0.21]), PLA Daily (0.14 [0.09,0.21]), and Military Review (0.10 [0.04,0.17]) (Table 2 / Figures 3–4). Three of the four are Chinese. If Pangram systematically over-flags formal Chinese bureaucratic or military-commentary style (or under-flags English policy prose), the country-pattern contrast and the count of “significant” sources both become artifacts rather than revealed behavior. The 2021 near-zero baselines reduce but do not eliminate this risk: a detector that is well-calibrated on pre-LLM text can still mis-fire once post-2023 stylistic conventions or translation-like fluency appear. No public corpus or detector outputs are released, so external re-scoring is currently impossible.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes measuring traces of language-model assistance in public government documents as a lightweight, externally reproducible monitoring primitive for government AI adoption. In a pilot of ten U.S. and PRC government-related document streams (n≈3,068), it reports near-zero Pangram AI fractions in 2021 and rising scores from 2024, with four sources showing statistically significant 2026 elevations (CAC, MOST, PLA Daily, Military Review). It further claims that, in this sample, the U.S. signal concentrates in outlets downstream of policy work while the PRC signal concentrates closer to policy formulation, and discusses how the signal could complement procurement and disclosure instruments while noting detector brittleness and limited Chinese validation.","tokens_in":36695,"tokens_out":1548,"duration_ms":22534,"significance":"If the detector scores are valid measures of LLM assistance in these genres, the paper supplies a useful, low-cost ecosystem-monitoring instrument for technical AI governance: revealed-behavior evidence that is cheap to recompute, comparable across jurisdictions, and less dependent on government self-report than use-case inventories or procurement records. Strengths include a clean pre-LLM 2021 baseline near zero across all sources, transparent per-source counts and multi-level bootstrap CIs (Tables 1–2, Figures 1–4), carefully scoped pilot claims rather than causal claims about high-level policymakers, and an explicit limitations discussion. The framing as a composable monitoring primitive is a genuine contribution even if the numerical country pattern remains provisional.","major_comments":[{"comment":"§1 Method and §2 Limitations, together with Table 2: the headline claim that four of ten sources show significant 2026 AI-assisted writing, and the U.S.-downstream vs PRC-policy-proximate contrast, rest primarily on Pangram fraction_ai for three Chinese sources (CAC 0.13 [0.07,0.19], MOST 0.11 [0.02,0.21], PLA Daily 0.14 [0.09,0.21]). The manuscript itself states that independent Chinese-language validation is limited. A detector that is well-calibrated on 2021 pre-LLM text can still systematically over-flag formal Chinese bureaucratic or military-commentary style (or under-flag English policy prose) once post-2023 fluency conventions appear. Without a targeted validation—e.g., human expert annotation of a stratified Chinese subsample, dual-detector comparison, or a held-out Chinese bureaucratic calibration set—the four-source count and the country-pattern interpretation remain vulnerabl","section":"§1 Method; §2 Limitations; Table 2"},{"comment":"Table 1 and the 2026 columns of Table 2 / Figures 3–4: several 2026 cells are very small or partial (OSTP n=5, DARPA n=13, State Council n=19, MOST n=42; cutoff 2026-04-25). Wide CIs for DARPA (0.08 [0.00,0.20]) and MOST (0.11 [0.02,0.21]) mean that “statistically significant signs” and the ranking that underpins the downstream-vs-proximate narrative are sensitive to a handful of documents. The paper should either restrict the 2026 significance claim to sources with adequate n, report a pre-registered minimum cell size, or show leave-one-out / influence diagnostics so readers can see whether a few high-scoring items drive the means.","section":"Table 1; Table 2; Figures 3–4"},{"comment":"Abstract and §1 (“externally reproducible”) vs. data availability: the primitive is marketed as lightweight and externally reproducible from public outputs, yet the manuscript does not release document IDs, scraped text, or Pangram window-level scores. Without that release, independent re-scoring with alternative detectors (or future Pangram versions) is impossible, which undercuts both the reproducibility claim and the ability of civil-society actors—the intended users—to verify or extend the panel. Releasing at least URLs, hashes, and per-document fraction_ai (even if full text is restricted) would make the contribution match its framing.","section":"Abstract; §1; Data availability"},{"comment":"§1 results paragraph on U.S. “downstream of high-level policy generation” vs PRC “closer to” policy: this contrast is interpretive and post-hoc. Military Review and DARPA news are labeled downstream; OSTP and AI-keyworded Federal Register are labeled policy-proximate and show no signal; CAC and MOST are labeled policy-proximate and do show signal. The operational definition of “downstream” vs “closer to policy work” is not pre-specified, and alternative groupings (e.g., military vs civilian, commentary vs normative instruments) could reorganize the pattern. Either pre-register a proximity coding, report a sensitivity table under alternative codings, or demote the contrast from a main empirical claim to a suggestive observation.","section":"§1 Results; Abstract"}],"minor_comments":[{"comment":"Figure 1 pools countries with equal source weights; a sentence in the caption or §1 stating how sensitive the pooled means are to dropping PLA Daily or Military Review would help readers assess robustness of the country-level rise.","section":"Figure 1"},{"comment":"Section C source descriptions are thorough, but a short justification for excluding other natural candidates (e.g., DoD press, State Department, NPC documents) would clarify selection bias risk for the pilot panel.","section":"Section C"},{"comment":"Section B quality check is valuable; reporting the post-remediation exclusion count in the main text (not only appendix) would reassure readers that placeholder/stub remediation does not drive the 2026 elevations.","section":"Section B"},{"comment":"In Table 2, bolding intervals that exclude zero is helpful; consider also marking which of the four “significant” sources survive a multiple-comparison correction if the paper continues to count them as a set.","section":"Table 2"},{"comment":"Related work §A.3 could briefly note how the government panel differs from Liang et al. peer-review/scientific-paper estimates in genre formality, which may affect detector false-positive rates.","section":"Section A.3"},{"comment":"The AI Usage Statement is appropriate and clear; keep it.","section":"AI Usage Statement"}],"recommendation":"major_revision","confidential_remarks":"Fit for a technical AI governance workshop is strong; as a journal article the empirical claims need the Chinese-detector validation and data release before the four-source / country-pattern headlines are safe. Novelty is real as a monitoring primitive applied to government text, not as a detector paper. No integrity red flags; the authors are unusually transparent about limitations and pipeline spot-checks."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a solid workshop pilot that does something useful: it treats AI-text detection on government document streams as a cheap, externally runnable monitoring primitive for AI governance, and it actually ships the first panel numbers for ten U.S./PRC sources with a 2021 pre-LLM baseline.\n\nWhat is new is the framing plus the measurements. Detectors have already been run on peer reviews, papers, news, and court filings; applying the same family of tools to CAC, MOST, PLA Daily, Military Review, DARPA news, Federal Register, OSTP, and the gov.cn buckets, and reporting per-source means with bootstrap CIs, is new empirical work. The design is clean for a pilot: near-zero 2021 baselines across every source, equal source weighting, transparent Table 2 and Figures 3–4, and carefully scoped claims (no causal leap to high-level policy use). The country pattern they highlight—U.S. signal in downstream outlets (Military Review 0.10, DARPA 0.08), PRC signal closer to the policy apparatus (CAC 0.13, MOST 0.11, PLA Daily 0.14)—is interesting and honestly presented as sample-specific. Limitations section is candid about detector brittleness, genre confounds, and the gap between drafting assistance and the uses one most wants to measure.\n\nThe soft spot is real and load-bearing for the strongest claim. Three of the four “significant by 2026” sources are Chinese, and the paper itself notes that independent validation of Pangram on Chinese is limited. 2021 near-zero baselines help, but they do not fully rule out post-2023 domain shift or style bias. No public corpus or detector outputs means outsiders cannot re-score. Partial 2026 coverage and proprietary detector are secondary but real. None of this is circular or mathematically broken; it is just the usual pilot constraint that the measurement instrument is not fully stress-tested for the hardest language/genre cells.\n\nThis is for people who care about ecosystem monitoring in AI governance and comparative public administration. It is not a methods breakthrough on detection itself. I would bring it to reading group, cite the framing and the panel numbers, and send it to peer review for a workshop or short empirical track. The central argument holds as a carefully scoped pilot; the Chinese-calibration caveat should be kept front and center in any revision or citation.","headline":"Clean pilot that turns public-document AI detection into a usable governance monitoring signal; the four-source 2026 result is real on the detector they used, but Chinese calibration remains the soft underbelly.","tokens_in":37301,"tokens_out":586,"would_cite":true,"duration_ms":8501,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"By 2026, four of ten U.S. and Chinese government document streams that scored near zero in 2021 show statistically significant traces of AI-assisted writing, with the U.S. signal downstream of policy work and the PRC signal closer to it.","keywords":["government AI adoption","AI text detection","ecosystem monitoring","public documents","revealed behavior","U.S.–China","frontier AI governance","monitoring primitive"],"falsifier":"Have independent expert human annotators score a large stratified sample of the same 2021 and 2024–2026 documents for AI assistance; if the human-judged rise is absent or far smaller than the detector rise, or if the detector still flags high fractions on matched known-human bureaucratic text, the claim fails.","tokens_in":37257,"feed_emoji":"🏛️","tokens_out":1060,"duration_ms":17894,"temperature":0.7,"pith_summary":"Governments matter for frontier AI governance, yet their real day-to-day use of AI is hard to see: procurement and official statements lag, select, and track formal adoption more than practice. This paper proposes a cheap complementary monitor—measuring language-model writing traces in the public documents governments already publish—and treats the resulting score as a standardized monitoring primitive based on revealed behavior. In a pilot of more than 3,000 documents from ten U.S. and PRC streams, 2021 baselines sit near zero; by 2026 four sources rise with confidence intervals excluding zero. In the sample the U.S. signal concentrates in Military Review and DARPA news (downstream of high-level policy), while the PRC signal concentrates at CAC and MOST (closer to rule-writing and science-policy administration). A sympathetic reader cares because the method is externally reproducible by journalists, academics, or other states, can flag activity missing from incomplete self-reported inventories, and can be composed with other instruments without requiring insider access.","feed_headline":"AI writing traces rise in U.S. and Chinese government docs","feed_subtitle":"Four of ten streams near zero in 2021 show clear AI assistance by 2026","key_machinery":"The monitoring primitive is the document-level AI fraction ai in [0,1] returned by a commercial AI-text classifier (Pangram) that segments each public document into windows, classifies each window, and averages them; the same score is applied to any fixed panel of public streams so that changes over time and across sources can be tracked and bootstrapped.","core_discovery":"Across ten public U.S. and PRC government-related document streams, mean per-document AI-writing fractions are near zero in 2021. By 2026 four sources show statistically significant elevation: U.S. Military Review (0.10) and DARPA news (0.08) sit downstream of policy generation, while PRC CAC (0.13) and MOST (0.11) sit closer to it; country-pooled means reach roughly 0.05–0.07 with bootstrap confidence intervals excluding zero. The paper presents this as a pilot demonstration that public-document AI detection can function as a lightweight monitoring primitive for government AI use.","pith_inferences":["If policy-near PRC sources continue to lead, export controls may constrain commercial more than state adoption, especially where governments retain privileged model access.","Absence of signal in high-level U.S. policy outlets may understate elite exposure if AI is used for analysis but not for the final public text.","Once the primitive is public, deliberate detector evasion by agencies becomes itself a measurable governance response.","Expanding the panel across more languages, agencies, and non-AI-keyworded streams would test whether the U.S.–PRC reverse pattern generalizes."],"forward_implications":["Journalists, academics, NGOs, or other states can re-run the same panel monthly on public outputs without insider access or government self-report.","Revealed-behavior scores can surface AI activity omitted from incomplete agency use-case inventories and from DoD/intelligence exemptions.","Where the signal sits (downstream of policy vs. near policy formulation) can indicate which parts of the state adopt first.","Score changes can prioritize sources for qualitative review, procurement searches, interviews, or comparison with disclosed AI policies.","As model-family attribution matures, the same streams may indicate domestic versus foreign or closed versus open model use."],"fun_headline_variants":["AI writing traces climb in U.S. and PRC government documents","Four of ten gov streams show clear AI assistance by 2026","Public docs flag rising AI use in U.S. and Chinese agencies","From near-zero in 2021 to AI signals in U.S. and PRC papers","Pilot tracks AI language traces in government public documents"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The central claim depends on the detector’s per-document AI fraction being a sufficiently accurate measure of real language-model assistance in these English and Chinese government genres, rather than an artifact of formal bureaucratic style, domain shift, or language-specific detector bias.","fun_headline_variants_meta":{"raw":{"variants":["AI writing traces climb in U.S. and PRC government documents","Four of ten gov streams show clear AI assistance by 2026","Public docs flag rising AI use in U.S. and Chinese agencies","From near-zero in 2021 to AI signals in U.S. and PRC papers","Pilot tracks AI language traces in government public documents"]},"model":"grok-4.5","effort":"low","cost_usd":0.008314,"raw_usage":{"total_tokens":1939,"prompt_tokens":782,"num_sources_used":0,"completion_tokens":96,"cost_in_usd_ticks":83140000,"prompt_tokens_details":{"text_tokens":782,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1061,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":782,"tokens_out":96,"duration_ms":10804,"temperature":1.0,"reasoning_tokens":1061,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T17:37:49.962000+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Have independent expert human annotators score a large stratified sample of the same 2021 and 2024–2026 documents for AI assistance; if the human-judged rise is absent or far smaller than the detector rise, or if the detector still flags high fractions on matched known-human bureaucratic text, the claim fails.","supporting_citations":[],"review_version":1}