{"id":"97f6f006-d4c8-4e55-9503-f244d10af3b8","arxiv_id":"2502.09747","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LLM-assisted writing rose sharply after ChatGPT's launch and plateaued by 2024, reaching estimated shares of roughly 18% in consumer complaints, 24% in corporate press releases, 14% in UN releases, and up to 15% in small-firm job postings.","lead":"Using a statistical text-analysis detector, this paper estimates how much real-world writing is now shaped by large language models. It finds that by 2024 roughly 18% of US consumer complaint text, 24% of corporate press releases, and 14% of UN press releases showed detectable signs of LLM assistance, with adoption plateauing after an initial post-ChatGPT surge.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central percentages rest on an unvalidated equivalence between real-world LLM-assisted writing and the paper's synthetic two-prompt GPT-generation pipeline; calibration is shown only on synthetic mixtures, so the population estimates are not yet proven unbiased for actual human-in-the-loop LLM…","rationale":"The paper is a well-executed large-scale application of a previously published detector, and the qualitative post-ChatGPT surge-then-plateau pattern is consistent across all four domains and survives multiple robustness checks, including training data generated by different GPT models (Supp. Fig. 4). The synthetic validation demonstrates genuine internal calibration: the estimator recovers known mixture proportions with small error, which is real evidence for the stability of the estimation procedure. My concern is external validity rather than internal consistency. The positive class used to train and validate the detector is generated by one specific two-prompt GPT pipeline, and no evidence is provided that this synthetic class matches the distribution of real LLM-assisted writing in any of the four domains. Because every headline percentage is an output of this estimator, the central claim is not yet established as an unbiased measure of actual LLM use; what is established is that a measurable fraction of text resembles the training pipeline. The paper's limitations section partially mitigates this concern by acknowledging that heavily edited or very human-like LLM output will be missed, which supports a lower-bound interpretation. However, the validation tables also show a consistent additive positive bias of roughly 2–3 percentage points, including at zero true alpha, and there is no test of whether the false-positive rate remains stable on post-2023 human writing. These issues are addressable with a focused human-in-the-loop calibration study, and the qualitative findings are unlikely to collapse. The conditional verdict is therefore appropriate, and my stress-test does not move it.","tokens_in":15727,"tokens_out":6171,"duration_ms":69357,"concrete_test":"Build a held-out benchmark of genuinely LLM-assisted documents in each of the four domains: for each domain, have a set of writers complete the same real-world task (e.g., filing a consumer complaint, drafting a press release, writing a job posting, drafting a UN-style release) with ChatGPT assistance and then edit the draft as they normally would, and have a matched control group write without any LLM; establish ground truth by the experimental assignment plus sentence-level human annotation. Run the published detector on this benchmark and compare its estimated alpha to the true fraction of LLM-assisted sentences.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline estimates—roughly 18% of consumer complaints, up to 24% of press releases, just below 10% of small-firm job postings, and nearly 14% of UN press releases—are all outputs of a mixture estimator whose \"LLM-assisted\" class is defined by a specific synthetic generation protocol: GPT-3.5-turbo compresses a pre-ChatGPT human text into a bullet skeleton and then expands that skeleton into full text (Supp. Figs. 5–6). The validation in Supp. Tables 1–5 mixes only this synthetic positive class with pre-ChatGPT human text at known proportions; there is no independently labeled set of real documents produced by humans using LLMs in the wild. The reported calibration error of less than 3.3 percentage points therefore measures how well the detector recovers the fraction of text that resembles this particular pipeline, not the fraction of text actually written or substantially modified by LLMs. Real LLM-assisted writing—direct prompting, human editing of model drafts, use of other model families—can have a different lexical signature. The validation tables also show a systematic positive bias of roughly 2–3 percentage points at every ground-truth level, including +1.8 to +2.9 percentage points at alpha = 0, and this offset is not used to debias the headline numbers. The paper's \"lower bound\" framing covers heavily edited or very human-like LLM output being missed, but it does not cover the opposite risk: human text whose style drifts toward GPT-like patterns, whether from templates or from human writers imitating AI, will be counted as adoption. Thus the central claim is not yet externally calibrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper applies a word-frequency mixture estimator, originally developed by the authors to measure AI-modified text in peer reviews, to four large English-language text corpora: 687,241 consumer complaints to the CFPB, 537,413 corporate press releases from Newswire/PRNewswire/PRWeb, 304.3 million LinkedIn job postings, and 15,919 UN press releases. The central empirical claim is that LLM-assisted writing rose sharply starting 3–5 months after ChatGPT's November 2022 release and then plateaued, reaching roughly 18% of consumer complaint text, up to 24% of corporate press release text, just below 10% of small-firm job postings, and nearly 14% of UN press releases by late 2024. The paper also reports heterogeneity in adoption by geography, urbanization, education, and firm age/size. The detection method's ground truth is built entirely from a synthetic pipeline in which GPT-3.5-turbo compresses pre-ChatGPT human text into bullet skeletons and re-expands it, and the validation sets mix only this synthetic positive class with pre-ChatGPT human text.","tokens_in":16058,"tokens_out":9633,"duration_ms":80205,"significance":"If the headline percentages are unbiased, the paper provides the first population-level, cross-domain measurement of LLM-assisted writing, with immediate relevance to policy discussions about AI adoption, labor markets, and institutional communication. The study's strengths are its scale, the consistency of the temporal pattern across four independent domains, the use of an open and transparent estimator rather than commercial black-box detectors, and the robustness check across GPT model versions in Supplementary Figure 4. The authors also honestly acknowledge that heavily edited or human-like LLM text escapes detection and frame their numbers as lower bounds. However, the central estimates inherit a load-bearing external-validity assumption: that synthetic GPT-generated outlines-and-expansions are representative of real-world human-in-the-loop LLM-assisted writing. That assumption is not tested against any independently labeled real documents, and the validation tables in the supplement show a systematic positive bias at every ground-truth level, including at zero.","major_comments":[{"comment":"The ground truth for the 'LLM-assisted' class is entirely synthetic: pre-ChatGPT human text is compressed into bullet-point skeletons and then re-expanded by GPT-3.5-turbo (Supp. Figs. 5–6), and the validation corpora in Supp. Tables 1–5 mix only this synthetic positive class with pre-ChatGPT human text. No independently labeled set of real documents produced by humans using LLMs in the wild is used for validation. The reported prediction error of less than 3.3 percentage points therefore measures how well the estimator recovers the fraction of text that resembles this specific two-prompt pipeline, not the fraction of text actually written or substantially modified by LLMs. Direct prompting, human editing of model drafts, and other model families can produce different lexical signatures. The paper's 'lower bound' framing covers heavily edited or very human-like LLM output being missed, but it does not cover the opposite risk: human text whose style drifts toward GPT-like patterns, from either temporal style drift or humans imitating LLM output. This assumption is load-bearing for every headline percentage.","section":"Supplementary Information, Model Fitting; Supp. Figs. 5–6; Supp. Tables 1–5"},{"comment":"The validation tables show a systematic positive bias of roughly 2–3 percentage points at every ground-truth level, including at α=0. For example, Supp. Table 1 reports an estimate of 1.8% at ground-truth 0.0%; Supp. Table 3 reports 2.9%, 2.1%, and 2.3% for PRNewswire, PRWeb, and Newswire at α=0; and Supp. Table 5 reports 2.0% for Scientist at α=0. This offset is not subtracted or otherwise corrected in the headline estimates. For the smallest headline claim ('just below 10%' in small-firm job postings), this is a substantial relative correction, and even for the largest estimates (18–24%) it is a non-negligible absolute correction. The manuscript should either explicitly debias the estimates using the measured false-positive rates at α=0 and at the relevant mixing levels, or demonstrate that the qualitative conclusions are unchanged after applying the corresponding correction.","section":"Supp. Tables 1–5"},{"comment":"The stabilization/plateau pattern is presented as a main finding, but the interpretation is confounded with possible changes in detector sensitivity over time. Because the synthetic positive class is generated by specific GPT models (GPT-3.5-turbo and GPT-4 variants), improvements in LLM indistinguishability—or in human writers' tendency to imitate LLM style—could cause the estimated fraction to flatten or decline even if true adoption continues to rise. The manuscript acknowledges this in the Discussion and in footnote 2, but the abstract and results still present stabilization as an empirical regularity. The validation in Supp. Tables 1–5 only assesses calibration on pre-ChatGPT human text mixed with synthetic LLM text; it does not test whether the detector's sensitivity is stationary across 2023–2024. A sensitivity analysis that varies the assumed sensitivity trajectory over time (e.g., by re-generating the synthetic positive class with successive model versions and showing the time-series conclusions are robust) is needed to support the plateau claim.","section":"Fig. 1; Results; Discussion"}],"minor_comments":[{"comment":"The sentence 'Using the sample of small companies based on the number of vacancies posted, our findings reveal... (Fig. 1d, Fig. 4)' refers to Figure 1d, which displays UN press releases; the correct panel for job postings appears to be Fig. 1c.","section":"Results, LLM Adoption in LinkedIn Job Postings; Fig. 1"},{"comment":"The caption states that GPT-3.5-turbo, 'used in main analysis,' was 'released January 25, 2024'; GPT-3.5-turbo was released earlier, so this date likely refers to a specific model snapshot rather than the model family, and the caption should be corrected or clarified.","section":"Supplementary Figure 4"},{"comment":"The Introduction says adoption surged '3-4 months' after ChatGPT's release, while this section says 'about 2 quarters post rollout'; please reconcile the timing statements.","section":"Results, LLM Adoption in Corporate Press Releases"},{"comment":"The p-values ('less than 0.001') and 'highly statistically significant' claims are reported without specifying the statistical test, the unit of analysis, or whether any multiple-comparison correction was applied; please add these details.","section":"Results, Geographic and Demographic Disparities"},{"comment":"The consumer-complaint data end in August 2024, so 'By late 2024' is slightly imprecise; please adjust to 'by August 2024' or 'by late summer 2024.'","section":"Abstract"},{"comment":"The definition of small firms as 'companies with either 10 or fewer registered employees in 2021 or companies posting less than or equal to about 2 postings per year' is ambiguous in light of the later statement that the median number of postings is 3; please clarify whether the threshold is 2 or 3.","section":"Supplementary Information, LinkedIn Job Posting Data"}],"recommendation":"major_revision","confidential_remarks":"The paper tackles an important and timely question and the scale of the data is impressive. My main concern is methodological: the estimator is validated only on synthetic data, and the validation shows a systematic positive bias that is not corrected. The authors' own framing as a 'lower bound' does not address the false-positive risk. I would encourage the editor to seek a reviewer with expertise in text-detection methodology and calibration, and to ask the authors for either an external validation set with real human-in-the-loop LLM documents or a substantial tempering of the quantitative claims. The paper is not fatally flawed, but the headline numbers need to be presented with their model-dependence made explicit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know up front. The cross-domain design is genuinely new—no one has put consumer complaints, corporate press releases, job postings, and UN press releases on the same measurement footing—and the core temporal pattern, a sharp rise a few months after ChatGPT then a plateau, is consistent across all four corpora and survives several robustness checks. That pattern is the paper's real contribution. But the headline percentages (18% complaints, up to 24% press releases, ~10% small-firm postings, 14% UN) are outputs of a detector whose 'LLM-assisted' class is defined by a synthetic outline-then-expand pipeline using GPT-3.5-turbo on pre-ChatGPT human text. The validation in the supplementary tables mixes only that synthetic positive class with pre-ChatGPT text; there is no independently labeled sample of real documents produced by humans using LLMs in the wild. So the calibration error of <3.3 pp tells you how well the detector recovers the fraction of text resembling that specific pipeline, not the true fraction of LLM-assisted writing.\n\nMore concretely, the validation tables show a systematic positive bias of roughly 1.8–2.9 percentage points at every ground-truth level, including at zero, in the consumer, press-release, and UN validations. The authors call the pre-ChatGPT estimate a 'false positive rate' but then don't subtract it from the post-ChatGPT numbers. The result is that the quoted adoption figures likely overstate true adoption by about 2–3 pp. The 'lower bound' framing covers the risk of missing heavily edited or very human-like LLM output, but it doesn't cover the opposite risk: human text that drifts toward GPT-like style—templates, or people imitating AI—will be counted as adoption.\n\nNone of this sinks the qualitative story. The surge, the plateau, the lower-education and younger-firm patterns are all well supported in shape if not in precise level. The paper is readable, the data work is large, and the authors are transparent about several limitations. The job-posting data end in October 2023 and are licensed, which is minor but worth noting for reproducibility.\n\nI'd send this to a serious referee. The central methodological gap is addressable: debias the estimates using the validation offsets, validate against a small set of independently labeled real LLM-assisted documents, and soften the exact percentages. If that happens, this becomes a durable measurement paper. I'd cite it for the trajectory and cross-domain comparison, not for the absolute numbers.","headline":"Solid cross-domain measurement of the post-ChatGPT surge, but the absolute adoption percentages need debiasing and external validation before I'd quote them.","tokens_in":16582,"tokens_out":4553,"would_cite":true,"duration_ms":43695,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper estimates that LLMs assisted with roughly 18% of consumer complaint text, 24% of corporate press releases, 10% of small-firm job postings, and 14% of UN press releases by late 2024.","keywords":["large language models","LLM adoption","AI-assisted writing","text attribution","consumer complaints","job postings","press releases","ChatGPT"],"falsifier":"Take a corpus of documents whose true AI-assistance status is known independently—for example, job postings or press releases drafted with an LLM and then edited by humans, with editing logs intact—and run this estimator on it. If the estimated $\\alpha$ falls well below the true fraction of LLM-influenced text, the reported population figures are best read as lower bounds rather than point estimates.","tokens_in":15537,"feed_emoji":"✍️","tokens_out":5332,"duration_ms":47231,"temperature":0.7,"pith_summary":"This paper sets out to measure how much of everyday writing is now done with help from large language models, using four large text corpora: consumer financial complaints, corporate press releases, online job postings from small firms, and United Nations press releases. Tracking these from before ChatGPT's launch through late 2024, it estimates that roughly 18% of consumer complaint text, up to 24% of corporate press release text, just under 10% of small-firm job posting text, and nearly 14% of UN press release text is LLM-assisted. It also finds that adoption rose sharply about three to five months after ChatGPT's release and then flattened out by 2023–2024. A sympathetic reader would take these as the first population-level, cross-domain estimates of LLM adoption in writing, with implications for how we read corporate, institutional, and consumer-facing text.","feed_headline":"LLMs now write ~18% of consumer complaints, 24% of press releases","feed_subtitle":"Adoption surged months after ChatGPT's debut, then plateaued across four domains by 2024.","key_machinery":"The engine of the analysis is a domain-adapted text-mixture estimator. For each corpus, word-frequency distributions are built from two reference pools: human writing collected before ChatGPT's release and synthetic text produced by prompting current GPT models. Fitting a mixture model to observed monthly text yields an estimate of $\\alpha$, the proportion of sentences substantially modified by LLMs. The same estimator is fitted separately for each press-release platform, job category, and complaint/UN corpus, and its bias is checked by mixing pre-ChatGPT human text with GPT output at known concentrations. That calibration, with prediction error below 3.3 percentage points in the validation tables, is what turns raw detection scores into population-level adoption numbers.","core_discovery":"The central claim is that LLM-assisted writing has become a measurable, large-scale feature of public and commercial text, not an edge case. Using a statistical estimator of the fraction $\\alpha$ of sentences that were generated or substantially modified by an LLM, the paper reports adoption levels of about 18% for financial consumer complaints, 23–24% for at least one corporate press-release platform, up to 15% for job postings from young small firms, and 14% for UN press releases by the end of the study period. The paper further claims that these adoption curves share a common shape: a lag of several months after ChatGPT's debut, a steep rise through 2023, and a plateau by 2024, which it reads as either saturation or the growing indistinguishability of AI output.","pith_inferences":["Because the method misses heavily edited or highly human-like LLM text, the paper's own numbers are lower bounds; the true prevalence in late 2024 could be materially higher than 18–24%.","If the plateau reflects detector blindness rather than saturation, apparent stabilization may mask continued growth—a testable prediction when future detectors calibrated on newer models are applied to the same 2024 corpora.","A direct extension would compare these population estimates with self-reported usage from surveys or platform telemetry, which could cross-validate the framework without relying on synthetic ground truth.","The finding that consumer complaints in lower-education areas show higher LLM adoption suggests these tools may be functioning as an equalizer in consumer advocacy, but the paper does not test whether AI-assisted complaints are more likely to receive redress."],"forward_implications":["Adoption of LLM-assisted writing is now a majority-adjacent phenomenon in corporate press releases, with the top platform reaching about 24% of text by late 2024.","The consistent plateau across domains suggests the first wave of adoption had largely run its course by 2024, whether through saturation or through models becoming harder to detect.","Smaller and younger firms lead in job-posting adoption, with post-2015 firms reaching 10–15% in some roles, pointing to an organizational-age gradient in AI uptake.","Geographic and demographic heterogeneity is modest but real: more urbanized areas show higher complaint adoption, while lower-education areas show slightly higher rates.","International organizations, exemplified by UN press releases, reach roughly 14% LLM-modified content, indicating institutional adoption in high-stakes communication."],"supporting_citations":[{"why":"supplies the core statistical framework for estimating the fraction of LLM-modified text at population scale.","marker":"[12]"},{"why":"extends the same detection approach to scientific papers, providing the methodological precedent for cross-domain adoption tracking.","marker":"[13]"},{"why":"documents that GPT detectors can misclassify non-native English writing, which the paper cites when discussing false-positive risks.","marker":"[14]"},{"why":"provides survey-based evidence of individual ChatGPT adoption that the paper contrasts with its population-level text estimates.","marker":"[8]"},{"why":"offers complementary survey evidence on rapid generative AI adoption, used as context for the observed surge.","marker":"[9]"},{"why":"supplies evidence on how LLM-written job posts affect applications and hiring, which the paper invokes when discussing implications of job-posting adoption.","marker":"[27]"},{"why":"is a related single-domain study of LLM adoption in consumer complaints, used as a comparison point.","marker":"[11]"}],"fun_headline_variants":["LLMs now craft 18% of complaints, 24% of press releases","ChatGPT-effect: 1 in 5 complaints, 1 in 4 releases written by AI","AI writing surges then plateaus: 18-24% of public text","From complaints to UN: LLMs write up to 24% of communications"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that text generated by current GPT models in the validation setup is a faithful stand-in for all real-world LLM-assisted writing, and that the pre-ChatGPT false-positive rate stays constant through 2024.","fun_headline_variants_meta":{"raw":{"variants":["LLMs now craft 18% of complaints, 24% of press releases","ChatGPT-effect: 1 in 5 complaints, 1 in 4 releases written by AI","AI writing surges then plateaus: 18-24% of public text","From complaints to UN: LLMs write up to 24% of communications"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000264,"raw_usage":{"total_tokens":1619,"prompt_tokens":977,"completion_tokens":642,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":553}},"tokens_in":593,"tokens_out":642,"duration_ms":6082,"temperature":1.0,"reasoning_tokens":553,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T20:36:59.487605+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a corpus of documents whose true AI-assistance status is known independently—for example, job postings or press releases drafted with an LLM and then edited by humans, with editing logs intact—and run this estimator on it. If the estimated $\\alpha$ falls well below the true fraction of LLM-influenced text, the reported population figures are best read as lower bounds rather than point estimates.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"extends the same detection approach to scientific papers, providing the methodological precedent for cross-domain adoption tracking."},{"cited_title":"& Horton, J","cited_arxiv_id":null,"evidence_quote":"supplies evidence on how LLM-written job posts affect applications and hiring, which the paper invokes when discussing implications of job-posting adoption."},{"cited_title":"& Shin, J","cited_arxiv_id":null,"evidence_quote":"is a related single-domain study of LLM adoption in consumer complaints, used as a comparison point."}],"review_version":1}