REVIEW 4 major objections 6 minor 44 references
By 2026, four of ten U.S. and Chinese government document streams that scored near zero in 2021 show statistically significant traces of AI-assisted writing, with the U.S. signal downstream of policy work and the PRC signal closer to it.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 17:37 UTC pith:V67OOCMP
load-bearing objection Clean pilot that turns public-document AI detection into a usable governance monitoring signal; the four-source 2026 result is real on the detector they used, but Chinese calibration remains the soft underbelly. the 4 major comments →
Government AI Use as a Monitoring Primitive: A Public Document Pilot Study
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Across ten public U.S. and PRC government-related document streams, mean per-document AI-writing fractions are near zero in 2021. By 2026 four sources show statistically significant elevation: U.S. Military Review (0.10) and DARPA news (0.08) sit downstream of policy generation, while PRC CAC (0.13) and MOST (0.11) sit closer to it; country-pooled means reach roughly 0.05–0.07 with bootstrap confidence intervals excluding zero. The paper presents this as a pilot demonstration that public-document AI detection can function as a lightweight monitoring primitive for government AI use.
What carries the argument
The monitoring primitive is the document-level AI fraction ai in [0,1] returned by a commercial AI-text classifier (Pangram) that segments each public document into windows, classifies each window, and averages them; the same score is applied to any fixed panel of public streams so that changes over time and across sources can be tracked and bootstrapped.
Load-bearing premise
The central claim depends on the detector’s per-document AI fraction being a sufficiently accurate measure of real language-model assistance in these English and Chinese government genres, rather than an artifact of formal bureaucratic style, domain shift, or language-specific detector bias.
What would settle it
Have independent expert human annotators score a large stratified sample of the same 2021 and 2024–2026 documents for AI assistance; if the human-judged rise is absent or far smaller than the detector rise, or if the detector still flags high fractions on matched known-human bureaucratic text, the claim fails.
If this is right
- Journalists, academics, NGOs, or other states can re-run the same panel monthly on public outputs without insider access or government self-report.
- Revealed-behavior scores can surface AI activity omitted from incomplete agency use-case inventories and from DoD/intelligence exemptions.
- Where the signal sits (downstream of policy vs. near policy formulation) can indicate which parts of the state adopt first.
- Score changes can prioritize sources for qualitative review, procurement searches, interviews, or comparison with disclosed AI policies.
- As model-family attribution matures, the same streams may indicate domestic versus foreign or closed versus open model use.
Where Pith is reading between the lines
- If policy-near PRC sources continue to lead, export controls may constrain commercial more than state adoption, especially where governments retain privileged model access.
- Absence of signal in high-level U.S. policy outlets may understate elite exposure if AI is used for analysis but not for the final public text.
- Once the primitive is public, deliberate detector evasion by agencies becomes itself a measurable governance response.
- Expanding the panel across more languages, agencies, and non-AI-keyworded streams would test whether the U.S.–PRC reverse pattern generalizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes measuring traces of language-model assistance in public government documents as a lightweight, externally reproducible monitoring primitive for government AI adoption. In a pilot of ten U.S. and PRC government-related document streams (n≈3,068), it reports near-zero Pangram AI fractions in 2021 and rising scores from 2024, with four sources showing statistically significant 2026 elevations (CAC, MOST, PLA Daily, Military Review). It further claims that, in this sample, the U.S. signal concentrates in outlets downstream of policy work while the PRC signal concentrates closer to policy formulation, and discusses how the signal could complement procurement and disclosure instruments while noting detector brittleness and limited Chinese validation.
Significance. If the detector scores are valid measures of LLM assistance in these genres, the paper supplies a useful, low-cost ecosystem-monitoring instrument for technical AI governance: revealed-behavior evidence that is cheap to recompute, comparable across jurisdictions, and less dependent on government self-report than use-case inventories or procurement records. Strengths include a clean pre-LLM 2021 baseline near zero across all sources, transparent per-source counts and multi-level bootstrap CIs (Tables 1–2, Figures 1–4), carefully scoped pilot claims rather than causal claims about high-level policymakers, and an explicit limitations discussion. The framing as a composable monitoring primitive is a genuine contribution even if the numerical country pattern remains provisional.
major comments (4)
- [§1 Method; §2 Limitations; Table 2] §1 Method and §2 Limitations, together with Table 2: the headline claim that four of ten sources show significant 2026 AI-assisted writing, and the U.S.-downstream vs PRC-policy-proximate contrast, rest primarily on Pangram fraction_ai for three Chinese sources (CAC 0.13 [0.07,0.19], MOST 0.11 [0.02,0.21], PLA Daily 0.14 [0.09,0.21]). The manuscript itself states that independent Chinese-language validation is limited. A detector that is well-calibrated on 2021 pre-LLM text can still systematically over-flag formal Chinese bureaucratic or military-commentary style (or under-flag English policy prose) once post-2023 fluency conventions appear. Without a targeted validation—e.g., human expert annotation of a stratified Chinese subsample, dual-detector comparison, or a held-out Chinese bureaucratic calibration set—the four-source count and the country-pattern interpretation remain vulnerabl
- [Table 1; Table 2; Figures 3–4] Table 1 and the 2026 columns of Table 2 / Figures 3–4: several 2026 cells are very small or partial (OSTP n=5, DARPA n=13, State Council n=19, MOST n=42; cutoff 2026-04-25). Wide CIs for DARPA (0.08 [0.00,0.20]) and MOST (0.11 [0.02,0.21]) mean that “statistically significant signs” and the ranking that underpins the downstream-vs-proximate narrative are sensitive to a handful of documents. The paper should either restrict the 2026 significance claim to sources with adequate n, report a pre-registered minimum cell size, or show leave-one-out / influence diagnostics so readers can see whether a few high-scoring items drive the means.
- [Abstract; §1; Data availability] Abstract and §1 (“externally reproducible”) vs. data availability: the primitive is marketed as lightweight and externally reproducible from public outputs, yet the manuscript does not release document IDs, scraped text, or Pangram window-level scores. Without that release, independent re-scoring with alternative detectors (or future Pangram versions) is impossible, which undercuts both the reproducibility claim and the ability of civil-society actors—the intended users—to verify or extend the panel. Releasing at least URLs, hashes, and per-document fraction_ai (even if full text is restricted) would make the contribution match its framing.
- [§1 Results; Abstract] §1 results paragraph on U.S. “downstream of high-level policy generation” vs PRC “closer to” policy: this contrast is interpretive and post-hoc. Military Review and DARPA news are labeled downstream; OSTP and AI-keyworded Federal Register are labeled policy-proximate and show no signal; CAC and MOST are labeled policy-proximate and do show signal. The operational definition of “downstream” vs “closer to policy work” is not pre-specified, and alternative groupings (e.g., military vs civilian, commentary vs normative instruments) could reorganize the pattern. Either pre-register a proximity coding, report a sensitivity table under alternative codings, or demote the contrast from a main empirical claim to a suggestive observation.
minor comments (6)
- [Figure 1] Figure 1 pools countries with equal source weights; a sentence in the caption or §1 stating how sensitive the pooled means are to dropping PLA Daily or Military Review would help readers assess robustness of the country-level rise.
- [Section C] Section C source descriptions are thorough, but a short justification for excluding other natural candidates (e.g., DoD press, State Department, NPC documents) would clarify selection bias risk for the pilot panel.
- [Section B] Section B quality check is valuable; reporting the post-remediation exclusion count in the main text (not only appendix) would reassure readers that placeholder/stub remediation does not drive the 2026 elevations.
- [Table 2] In Table 2, bolding intervals that exclude zero is helpful; consider also marking which of the four “significant” sources survive a multiple-comparison correction if the paper continues to count them as a set.
- [Section A.3] Related work §A.3 could briefly note how the government panel differs from Liang et al. peer-review/scientific-paper estimates in genre formality, which may affect detector false-positive rates.
- [AI Usage Statement] The AI Usage Statement is appropriate and clear; keep it.
Circularity Check
Empirical measurement of detector scores on public corpora; no circular derivation or fitted-as-prediction steps.
full rationale
The paper's central claims are observational: collect public documents from ten sources, run a commercial AI-text classifier (Pangram) to obtain per-document fraction_ai in [0,1], then report means and bootstrap CIs by source and year, with 2021 as a pre-LLM baseline. The four-source significance finding and the U.S./PRC location contrast are direct numerical outputs of that pipeline (Table 1 counts, Table 2 means/CIs, Figures 1–4). There is no parameter fitted to a subset of the same data and then re-labeled a prediction; no equation that defines the target quantity in terms of itself; no uniqueness theorem or ansatz imported from the authors' prior work that forces the result; and no renaming of a known pattern as a new derivation. Related-work citations (including one co-authored paper on algorithmic progress) are background only and do not enter the measurement chain. Detector validity is a correctness/assumption risk, not circularity. The derivation is therefore self-contained against external benchmarks and exhibits no circular reduction.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Pangram’s windowed AI-fraction scores are a sufficiently accurate proxy for language-model assistance on both English and Chinese government documents of the genres studied.
- domain assumption 2021 public documents constitute a valid near-zero pre-mass-market-LLM baseline for these sources.
- domain assumption Elevated AI fractions in published government text are informative about government AI adoption/use relevant to AI governance, even if they primarily capture drafting or public-affairs writing.
- ad hoc to paper Equal weighting of sources and equal weighting of documents within source is an appropriate aggregation for the pooled country means.
invented entities (1)
-
public-document AI-detection monitoring primitive
no independent evidence
read the original abstract
Governments are important actors in frontier AI governance, but many facts about their adoption and use of AI systems are difficult to observe directly. Procurement disclosures and official statements are useful, but can also be delayed, selective, and better suited to measuring formal adoption than actual day-to-day use. We propose a complementary monitoring primitive: measuring traces of language-model assistance in public government documents. The approach is lightweight, externally reproducible, and based on revealed behavior rather than stated intent. In a pilot study of ten public document streams from U.S. and PRC government-related sources, we find that, while 2021 baselines are consistently near zero, by 2026, four of our ten sources show statistically significant signs of AI-assisted writing. In our sample, the U.S. signal concentrates in publications downstream of policy work; the PRC signal concentrates closer to it. We close by discussing how this signal could complement existing instruments for monitoring government AI adoption, and where it falls short.
Figures
Reference graph
Works this paper leans on
-
[1]
C., Tripto, N
Ansari, A., Zhang, D. C., Tripto, N. I., and Lee, D. Echoes of Automation : The Increasing Use of LLMs in Newsmaking , April 2026
2026
-
[2]
Claude Gov models for U.S
Anthropic . Claude Gov models for U.S. national security customers. Anthropic News, June 2025. URL https://www.anthropic.com/news/claude-gov-models-for-u-s-national-security-customers
2025
-
[3]
From text to source: Results in detecting large language model-generated content
Antoun, W., Sagot, B., and Seddah, D. From text to source: Results in detecting large language model-generated content. In Calzolari, N., Kan, M.-Y., Hoste, V., Lenci, A., Sakti, S., and Xue, N. (eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pp.\ 7531--7543...
2024
-
[4]
Anthropic Economic Index report: Uneven geographic and enterprise AI adoption, November 2025
Appel, R., McCrory, P., Tamkin, A., McCain, M., Neylon, T., and Stern, M. Anthropic Economic Index report: Uneven geographic and enterprise AI adoption, November 2025
2025
-
[5]
Military review: The professional journal of the U.S
Army University Press . Military review: The professional journal of the U.S. army. Army University Press, U.S. Army Combined Arms Center, Fort Leavenworth, KS, n.d. URL https://www.armyupress.army.mil/Military-Review/
-
[6]
Blomquist, K. Racing for recognition? Theorizing emerging status hierarchies and prestige competition in the AI era. International Affairs, 102 0 (3): 0 949--970, 2026. doi:10.1093/ia/iiag028. URL https://doi.org/10.1093/ia/iiag028
-
[7]
Chen, C. and Jia, X. When researchers use AI : Public trust, ethical judgments, and the perceived value of academic research. AI and Ethics, 6 0 (2): 0 223, March 2026. ISSN 2730-5961. doi:10.1007/s43681-026-01039-w
-
[8]
and Lau, K.-s
Cheung, S. and Lau, K.-s. DeepSeek Use in PRC Military and Public Security Systems . China Brief 25(20), The Jamestown Foundation, October 2025. URL https://jamestown.org/deepseek-use-in-prc-military-and-public-security-systems/
2025
-
[9]
Communications and public affairs
Defense Advanced Research Projects Agency . Communications and public affairs. https://www.darpa.mil/news/public-affairs, n.d
-
[10]
Denain, J.-S. Adding Clinical trial registry entries, Earnings calls prepared remarks and Federal court opinions from CourtListener . X (formerly Twitter), Epoch AI, March 2026. URL https://x.com/i/status/2037646276251279674
arXiv 2026
-
[11]
and Spero, M
Emi, B. and Spero, M. Technical Report on the Pangram AI-Generated Text Classifier , July 2024
2024
-
[12]
Promoting the Use of Trustworthy Artificial Intelligence in the Federal Government
Executive Office of the President . Promoting the Use of Trustworthy Artificial Intelligence in the Federal Government . Executive Order 13960, Federal Register, document 2020-27065, December 2020. URL https://www.federalregister.gov/documents/2020/12/08/2020-27065/promoting-the-use-of-trustworthy-artificial-intelligence-in-the-federal-government. Publish...
2020
-
[13]
The state of AI competition in advanced economies
Haag, A. The state of AI competition in advanced economies. Feds notes, Board of Governors of the Federal Reserve System, Washington, D.C., October 2025. URL https://www.federalreserve.gov/econres/notes/feds-notes/the-state-of-ai-competition-in-advanced-economies-20251006.html
2025
-
[14]
Spotting LLMs With Binoculars : Zero-Shot Detection of Machine-Generated Text , October 2024
Hans, A., Schwarzschild, A., Cherepanova, V., Kazemi, H., Saha, A., Goldblum, M., Geiping, J., and Goldstein, T. Spotting LLMs With Binoculars : Zero-Shot Detection of Machine-Generated Text , October 2024
2024
-
[15]
Rage against the machine? Generative AI exposure, subjective risk, and policy preferences
Haslberger, M., Gingrich, J., and Bhatia, J. Rage against the machine? Generative AI exposure, subjective risk, and policy preferences. Journal of European Public Policy, September 2025. doi:10.1080/13501763.2025.2554903. Advance online publication, September 10, 2025
-
[16]
C., Atkinson, D., Thompson, N., and Sevilla, J
Ho, A., Besiroglu, T., Erdil, E., Owen, D., Rahman, R., Guo, Z. C., Atkinson, D., Thompson, N., and Sevilla, J. Algorithmic progress in language models, March 2024
2024
-
[17]
and Imas, A
Jabarian, B. and Imas, A. Artificial writing and automated detection. Working Paper 34223, National Bureau of Economic Research, 2025. URL https://www.nber.org/papers/w34223
2025
-
[18]
Legacy procurement practices shape how u.s
Johnson, N., Silva, E., Leon, H., Eslami, M., Schwanke, B., Dotan, R., and Heidari, H. Legacy procurement practices shape how u.s. cities govern ai: Understanding government employees’ practices, challenges, and needs. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, pp.\ 772–789. ACM, 2025. doi:10.1145/3...
-
[19]
Anthropic disables fable and mythos ai models after u.s
Kahn, J. Anthropic disables fable and mythos ai models after u.s. government bars it from giving foreigners access. Fortune, June 2026. URL https://fortune.com/2026/06/13/anthropic-disables-fable-mythos-export-controls-national-security-threat/
2026
-
[20]
Agencies report over 3,000 AI use cases in 2025
Kelley, A. Agencies report over 3,000 AI use cases in 2025. Nextgov/FCW, April 2026. URL https://www.nextgov.com/artificial-intelligence/2026/04/agencies-report-over-3000-ai-use-cases-2025/412898/. Published April 16, 2026. Reports that the Intelligence Community and Department of Defense are exempt from inventory reporting
2025
-
[21]
A Watermark for Large Language Models , May 2024
Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T. A Watermark for Large Language Models , May 2024
2024
-
[22]
GPT detectors are biased against non-native English writers, July 2023
Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., and Zou, J. GPT detectors are biased against non-native English writers, July 2023
2023
-
[23]
A., and Zou, J
Liang, W., Izzo, Z., Zhang, Y., Lepp, H., Cao, H., Zhao, X., Chen, L., Ye, H., Liu, S., Huang, Z., McFarland, D. A., and Zou, J. Y. Monitoring AI-Modified Content at Scale : A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews , March 2024
2024
-
[24]
Liang, W., Zhang, Y., Wu, Z., Lepp, H., Ji, W., Zhao, X., Cao, H., Liu, S., He, S., Huang, Z., Yang, D., Potts, C., Manning, C. D., and Zou, J. Quantifying large language model usage in scientific papers. Nature Human Behaviour, 9 0 (12): 0 2599--2609, December 2025. ISSN 2397-3374. doi:10.1038/s41562-025-02273-8
-
[25]
Feature-augmented transformers for robust ai-text detection across domains and generators, 2026
Mady, M., Reschke, J., and Schuller, B. Feature-augmented transformers for robust ai-text detection across domains and generators, 2026. URL https://arxiv.org/abs/2605.03969
Pith/arXiv arXiv 2026
-
[26]
Misra, A., Wang, J., McCullers, S., White, K., and Ferres, J. L. Measuring AI Diffusion : A Population-Normalized Metric for Tracking Global AI Usage , November 2025
2025
-
[27]
D., and Finn, C
Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., and Finn, C. DetectGPT : Zero-Shot Machine-Generated Text Detection using Probability Curvature , July 2023
2023
-
[28]
M., Xiao, S., Xu, Y., and Yang, Y
Naughton, B., Cheung, T. M., Xiao, S., Xu, Y., and Yang, Y. Reorganization of China 's science and technology system. IGCC Working Paper, UC Institute on Global Conflict and Cooperation, 2023. URL https://ucigcc.org/publication/reorganization-of-chinas-science-and-technology-system/
2023
-
[29]
Newhard, J. M. The stock market speaks: How Dr . Alchian learned to build the bomb. Journal of Corporate Finance, 27: 0 116--132, August 2014. ISSN 0929-1199. doi:10.1016/j.jcorpfin.2014.05.002
-
[30]
2025 federal agency artificial intelligence use case inventory
Office of Management and Budget . 2025 federal agency artificial intelligence use case inventory. https://github.com/ombegov/2025-Federal-Agency-AI-Use-Case-Inventory, 2025. Compiled pursuant to E.O.\ 13960, the Advancing American AI Act, and OMB Memorandum M-25-21
2025
-
[31]
S., Rajkumar, N., Mo \"e s, N., Ladish, J., Guha, N., Newman, J., Bengio, Y., South, T., Pentland, A., Koyejo, S., Kochenderfer, M
Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., Hammond, L., Ibrahim, L., Chan, A., Wills, P., Anderljung, M., Garfinkel, B., Heim, L., Trask, A., Mukobi, G., Schaeffer, R., Baker, M., Hooker, S., Solaiman, I., Luccioni, A. S., Rajkumar, N., Mo \"e s, N., Ladish, J., Guha, N., Newman, J., Bengio, Y., South, T., Pentland, A., Koyejo, S...
2024
-
[32]
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI -generated text
Russell, J., Karpinska, M., and Iyyer, M. People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI -generated text. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 5342--5373, Vienna, Austria, 2025. Association for Computational Linguistics. URL htt...
2025
-
[33]
S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S
Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S. Can ai-generated text be reliably detected?, 2025. URL https://arxiv.org/abs/2303.11156
Pith/arXiv arXiv 2025
-
[34]
Compute Trends Across Three Eras of Machine Learning
Sevilla, J., Heim, L., Ho, A., Besiroglu, T., Hobbhahn, M., and Villalobos, P. Compute Trends Across Three Eras of Machine Learning . In 2022 International Joint Conference on Neural Networks ( IJCNN ) , pp.\ 1--8, July 2022. doi:10.1109/IJCNN55064.2022.9891914
-
[35]
and Levy, J
Shah, A. and Levy, J. Access to Justice in the Age of AI : Evidence from U . S . Federal Courts . SSRN working paper, March 2026. URL https://ssrn.com/abstract=6766859
2026
-
[36]
China's AI regulations and how they get made
Sheehan, M. China's AI regulations and how they get made. Technical report, Carnegie Endowment for International Peace, July 2023. URL https://carnegieendowment.org/research/2023/07/chinas-ai-regulations-and-how-they-get-made
2023
-
[37]
Sutter, K. M. U.s. export controls and china: Advanced semiconductors. CRS Report R48642, Congressional Research Service, September 2025. URL https://www.congress.gov/crs-product/R48642. Updated September 19, 2025
2025
-
[38]
Federal Agencies Largely Miss the Mark on Documenting AI Compliance Plans as Required by AI Executive Order
Tobin-Miyaji , M. Federal Agencies Largely Miss the Mark on Documenting AI Compliance Plans as Required by AI Executive Order . EPIC, November 2024. URL https://epic.org/federal-agencies-largely-miss-the-mark-on-documenting-ai-compliance-plans-as-required-by-ai-executive-order/. Published November 21, 2024
2024
-
[39]
Department of Defense
U.S. Department of Defense . Military and Security Developments Involving the People 's Republic of China . Technical report, December 2025
2025
-
[40]
Government Accountability Office
U.S. Government Accountability Office . Artificial Intelligence: Agencies Have Begun Implementation but Need to Complete Key Requirements . Technical Report GAO-24-105980, U.S. Government Accountability Office , December 2023. URL https://www.gao.gov/products/gao-24-105980
2023
-
[41]
Government Accountability Office
U.S. Government Accountability Office . Artificial intelligence: Generative ai use and management at federal agencies. Report to Congressional Requesters GAO-25-107653, U.S. Government Accountability Office, July 2025. URL https://www.gao.gov/assets/gao-25-107653.pdf
2025
-
[42]
Congress on Superintelligence
Wildeford, P. Congress on Superintelligence . The AI Policy Network (AIPN), June 2026. URL https://theaipn.org/issue/quotes/
2026
-
[43]
Wu, J., Yang, S., Zhan, R., Yuan, Y., Chao, L. S., and Wong, D. F. A Survey on LLM-Generated Text Detection : Necessity , Methods , and Future Directions . Computational Linguistics, 51 0 (1): 0 275--338, March 2025. ISSN 0891-2017. doi:10.1162/coli_a_00549
-
[44]
and Mathur, V
Zimmerman, M. and Mathur, V. Ombegov/2024- Federal-AI-Use-Case-Inventory . Office of the Federal Chief Information Officer, April 2026
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.