{"id":"26588f96-e0ca-45d3-afb4-2b81cba5625d","arxiv_id":"2506.17467","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A compilation of three research programs showing that GPT detectors are biased against non-native writers, that population-level estimates place AI-modified text at up to 16.9% of AI-conference reviews and up to 24% in some domains, and that GPT-4 feedback overlaps with human peer review at…","lead":"This dissertation measures how much of modern writing may be AI-assisted, from conference reviews to job postings, and shows that AI text detectors unfairly flag writing by non-native English speakers. It also shows that GPT-4 feedback overlaps substantially with human peer review, pointing toward a role for AI in widening access to research critique.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-launch human style drift, not just LLM use, could produce the estimated alpha; the semi-synthetic validations do not exercise this counterfactual.","rationale":"The reader's weakest assumption correctly identifies the stationarity of the human reference distribution P as the load-bearing premise for all population-level adoption estimates. My stress-test confirms this is the single most consequential threat: the method's MLE in Eq. (3.2) with P fixed by Eq. (3.3) cannot distinguish 'human text that shifted toward AI-flavored style' from 'text actually modified or generated by an LLM.' The paper's own validation strategy is strong in several respects: it checks robustness across venues, prompts, parts of speech, and LLM families, and Ch. 4 adds a temporal-split validation that covers pre-ChatGPT drift. These checks genuinely support the internal consistency of the estimator under the stated generative model. However, none of them provides a post-ChatGPT human-only corpus, so the key counterfactual remains untested. The concern is not that the method is circular or internally inconsistent; it is an identification problem that a skeptical reader can probe, exactly as the CONDITIONAL verdict implies. My proposed synthetic style-drift test would quantify how much human-only lexical drift toward AI-favored adjectives is needed to produce the observed alpha increases, settling whether the headline range is robust or an artifact of the stationarity assumption. No evidence of misconduct or data fabrication appears; the limitation is at the level of modeling assumptions, matching the reader's assessment.","tokens_in":59318,"tokens_out":4066,"duration_ms":46245,"concrete_test":"Construct a synthetic 'style-drift' corpus by taking pre-ChatGPT ICLR reviews and probabilistically replacing a small fraction of adjectives with the top AI-favored adjectives from Table 3.2, at a rate calibrated to observed pre-ChatGPT year-over-year word-frequency drift (e.g., extrapolating the 2018-2022 trend in Figure 3.1 by one year). Run the MLE estimator on this purely human, no-LLM corpus. If estimated alpha rises above the paper's 5% significance threshold (or above the ~3.5% error band reported for the temporal-split validation in Fig. 4.3), then plausible human style drift can fully explain the post-ChatGPT alpha increase, and the headline overstates LLM adoption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (6.5-16.9% of peer-review text substantially modified by LLMs) is identified by assuming that the human reference distribution P estimated in Eq. (3.3) from pre-ChatGPT reviews remains the correct human distribution after Nov 30, 2022. The MLE then attributes any post-launch shift toward the AI adjective profile (e.g., 'commendable', 'meticulous', 'intricate') entirely to alpha. But the counterfactual human distribution is unobserved: humans may adopt AI-flavored words through exposure to LLM output, changes in reviewing norms, or reviewer-pool demographic shifts, all of which would inflate alpha. The semi-synthetic validation (Sec. 3.3.6) mixes pre-ChatGPT human reviews with ChatGPT-generated reviews, so it only verifies the estimator under the model's own generative assumptions. The temporal-split validation in Ch. 4 (Fig. 4.3) is stronger: it separates training data (up to 2020) from validation data (Jan-Nov 2022), but it stops before ChatGPT's release and therefore does not test the regime where style drift from LLM exposure is plausible. The paper's own Limitations (Sec. 3.5) concede that 'temporal distribution shift in token frequencies due to, e.g., changes in topics, reviewers, etc.' introduces error, and that shifts in the non-native speaker population could impact accuracy, but no bound is given. Without a post-ChatGPT human-only reference, the gap between 'measured alpha' and 'true LLM modification' is an untested identification assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The dissertation develops computational methods for measuring and characterizing the impact of large language models on writing and information ecosystems. Chapter 2 shows that widely used GPT detectors systematically misclassify non-native English writing as AI-generated and that simple prompting can bypass detectors. Chapter 3 introduces a maximum-likelihood 'distributional GPT quantification' framework: using pre-ChatGPT human-written reviews and ChatGPT-generated reference reviews to estimate the token-occurrence distributions P and Q, the method estimates α, the fraction of sentences in a target corpus substantially modified by an LLM. Applying this to ML conference peer reviews, the dissertation reports α estimates of 6.5%–16.9% after ChatGPT's release, with higher estimates near deadlines, for low-confidence reviewers, and for reviews without citations. Chapter 4 extends the framework to scientific abstracts and introductions across arXiv, bioRxiv, and Nature portfolio journals, showing rising LLM-modified content after 2022. Chapter 5 applies the same method to consumer complaints, corporate press releases, job postings, and UN press releases, reporting double-digit percentages of LLM-modified sentences. Chapter 6 evaluates GPT-4's ability to provide scientific feedback, finding substantial overlap with human reviewer comments and generally positive assessments in a prospective user study.","tokens_in":59581,"tokens_out":2539,"duration_ms":27890,"significance":"If the identification assumptions hold, this is an important and timely contribution. The population-level estimation framework is genuinely novel: it avoids instance-level detection, is many orders of magnitude cheaper than classifier-based detectors, and is validated extensively on semi-synthetic mixtures, under prompt shift, across parts of speech, for proofreading and outline-expansion use cases, and with alternative LLMs. The temporal patterns across multiple independent corpora—peer reviews, scientific abstracts, press releases, and complaints—are internally consistent and make a strong circumstantial case that LLM-assisted writing increased sharply after late 2022. The dissertation also provides a valuable cautionary result on detector bias against non-native writers. However, the central quantitative claims rest on an untested identification assumption: that the pre-ChatGPT human distribution P remains the correct counterfactual for post-ChatGPT human writing. The paper's own limitations section acknowledges temporal distribution shift and changing non-native-speaker populations as potential error sources without bounding them.","major_comments":[{"comment":"The central identification assumption is that the human token-occurrence distribution P, estimated from pre-ChatGPT reviews, remains the correct distribution for human-written text after November 30, 2022. The post-ChatGPT increase in α is then attributed to LLM adoption, but the counterfactual human distribution is unobserved. If human style drifted toward AI-flavored adjectives (e.g., 'commendable', 'meticulous', 'intricate') through exposure to LLM output, changing review norms, or reviewer-pool shifts, the estimated α would overstate true LLM modification. The semi-synthetic validation in Section 3.3.6 mixes pre-ChatGPT human reviews with the same ChatGPT-generated Q used for estimation, so it verifies the estimator only under the model's own generative assumptions. This is the load-bearing point for the headline 6.5–16.9% claim, and it needs a direct sensitivity analysis or an external post-ChatGPT human-only reference corpus.","section":"§3.3.1–3.3.5, Eq. (3.3), §3.4.4"},{"comment":"The temporal-split validation in Section 4.2.2 is the strongest validation in the dissertation because it separates training data (up to 2020) from validation data (January–November 2022) by more than a year and still achieves estimation error below 3.5%. However, this split stops before ChatGPT's launch, so it does not test the regime where human style drift from LLM exposure is most plausible. To support the population-level adoption estimates in Chapters 4 and 5, the author needs either a post-ChatGPT human-only validation corpus (e.g., writing from communities or venues with verified zero LLM use, or pre-registered human-written text) or a formal bound on the bias in α under a plausible range of human style drift.","section":"Chapter 4, Fig. 4.3"},{"comment":"The Limitations section concedes that 'temporal distribution shift in token frequencies due to, e.g., changes in topics, reviewers, etc.' introduces error and that substantial shifts in the non-native speaker population could affect accuracy, but it provides no quantification or bound. Since the magnitude of the headline estimates (6.5–16.9%) is comparable to the estimated baseline false-positive rate of 1.6–2.4% plus the post-ChatGPT increase of roughly 5–15 percentage points, an unquantified drift of even a few percentage points could materially change the conclusions. The manuscript should report a sensitivity analysis that perturbs P in the direction of observed post-ChatGPT token-frequency shifts and shows how α changes.","section":"§3.5 Limitations"}],"minor_comments":[{"comment":"There is a typo: 'non-native English witters' should be 'non-native English writers', and the sentence 'A lot of Room of improvement, it is crucial to develop more robust detection methods' is grammatically incomplete and should be revised.","section":"§2.2.3"},{"comment":"The word 'discrepency' should be 'discrepancy'.","section":"§2.3"},{"comment":"The caption states 'EMNLP '23' as an orange dot, but the figure legend is small and the point is easy to miss; consider adding a label directly or a table reference in the caption.","section":"Chapter 3, Figure 3.4"},{"comment":"The sentence 'Among the conferences with pre- and post-ChatGPT data, ICLR experienced the most significant increase' is slightly confusing because NeurIPS also has pre- and post-ChaGPT data; clarifying that ICLR showed the largest absolute increase would improve readability.","section":"§3.4.4"},{"comment":"The figure reports adoption 'plateauing at 17.7% through August 2024' for consumer complaints, but the discussion in §5.2.2 describes geographic and demographic disparities; a brief note on how the national estimate relates to the state-level estimates in Fig. 5.3 would help the reader connect the two analyses.","section":"Chapter 5, Fig. 5.1"}],"recommendation":"major_revision","confidential_remarks":"The dissertation is a compilation of previously published, highly visible papers, and the chapters are largely self-contained. For a journal submission, the main added value is the synthesis across domains and the unified methodological framing. The core methodological concern is the untested post-ChatGPT human-counterfactual assumption; this is not an instance of 'outside current consensus' but a genuine identification gap. If the author can add a bounded sensitivity analysis or a post-ChatGPT human-only validation, the central claims would be substantially strengthened. I would not reject on the current evidence, but I would not accept without addressing the identification concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This dissertation is a well-organized compilation of three research lines you likely already know: detector bias, distributional GPT quantification, and LLM feedback. If you have read the ICML paper and the Patterns paper, there is nothing new here empirically; the contribution is the unified framing and the careful presentation of validation. What is genuinely good: the population-level MLE method is clever, cheap, and more stable than instance-level detectors, and the validation effort is thorough—semi-synthetic mixtures, temporal splits, robustness to POS choice and prompt shifts. The bias chapter remains important and well-executed. The feedback chapter has a large prospective user study.\n\nThe soft spots are real but not fatal. The central estimates of alpha (6.5–16.9% of review sentences modified) rest on the assumption that the pre-ChatGPT human reference distribution P remains the correct counterfactual after ChatGPT's release. The stress-test note is right that human style drift from exposure to LLM output, shifting reviewer pools, or topic changes would inflate alpha. The semi-synthetic validation does not exercise that counterfactual because it mixes the same ChatGPT-generated Q used for estimation. The paper itself acknowledges temporal shift and non-native population shifts in Section 3.5, but gives no bound. The Chapter 4 temporal split stops in November 2022, so it does not test the post-launch regime. I think the honest reading is that the numbers are plausible upper bounds, and the paper mostly hedges appropriately—it says 'could have been substantially modified,' not 'were definitely.'\n\nA second, minor soft spot is Chapter 6's self-referential evaluation: GPT-4 both generates comments and matches them via semantic similarity. The human verification of the pipeline helps, but the matching model is the same family as the generator.\n\nOverall: this is a solid dissertation and a useful reference. The math behind the estimator is sound; the citation pattern is honest; the authors flag their own limitations. For a journal or conference paper, novelty is low because of the compilation, but the underlying methods deserve engagement. I would bring it to a reading group to discuss identification assumptions in population-level detection, and I would cite the method if I write on LLM adoption. I would send this to peer review if it were a new submission, but the appropriate venue is probably a methods-journal format rather than a results-claim format.","headline":"A well-structured compilation of published work; the method is sound and the validation is extensive, but the headline LLM-adoption numbers rest on an untested stationarity assumption and should be read as upper bounds.","tokens_in":60149,"tokens_out":2327,"would_cite":true,"duration_ms":24227,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This dissertation claims that the fraction of text substantially modified by large language models can be estimated at the population level by fitting a two-component word-frequency mixture model, and reports that 6.5% to 16.9% of…","keywords":["large language models","AI-generated text detection","population-level estimation","peer review","GPT detector bias","LLM scientific feedback","distributional GPT quantification","writing ecosystems"],"falsifier":"Collect a matched corpus of post-ChatGPT, human-only peer reviews from a community that demonstrably avoided LLM assistance, with topics and author demographics matched to the machine-learning conferences studied; if the estimation method still reports a share of LLM-modified text well above its measured error when the true share is zero, the pre-ChatGPT human baseline is shifting for reasons other than LLM modification.","tokens_in":59044,"feed_emoji":"📊","tokens_out":11829,"duration_ms":113063,"temperature":0.7,"pith_summary":"This dissertation argues that the population-level impact of large language models on writing can be measured even when individual AI sentences cannot be reliably identified. Its central contribution is a word-frequency estimation method, called distributional GPT quantification, that fits a mixture of human and AI text distributions to a corpus and reports the fraction of sentences substantially modified or generated by an LLM. Applying it to peer reviews at major machine-learning conferences, the dissertation estimates that between 6.5% and 16.9% of review sentences after ChatGPT's launch were substantially modified by LLMs, beyond proofreading or minor edits. The same method is extended to scientific abstracts, consumer complaints, corporate press releases, job postings, and UN press releases, where double-digit percentages of LLM-modified content appear after ChatGPT's release. A sympathetic reader would care because the work promises a scalable, low-cost way to monitor AI adoption in high-stakes writing ecosystems, and it pairs that measurement with findings that GPT detectors are biased against non-native English writers and that LLM feedback can overlap substantially with human peer review.","feed_headline":"Word counting puts 6.5-16.9% of peer-review sentences as LLM-edited","feed_subtitle":"A population-level word-frequency test tracks LLM adoption where single-sentence detectors fail.","key_machinery":"The load-bearing mechanism is the two-component mixture model with maximum-likelihood estimation over word-occurrence statistics. The paper defines a vocabulary of adjectives, chosen for stability over adverbs, verbs, and nouns, and models each sentence's probability as the product of per-token occurrence probabilities under the human distribution $P$ and under the LLM distribution $Q$; the corpus log-likelihood is $\\sum_i \\log((1-\\alpha)P(x_i)+\\alpha Q(x_i))$, maximized to estimate $\\alpha$. What makes it work is the empirical observation that words such as 'commendable', 'meticulous', and 'intricate' rose sharply in frequency in post-ChatGPT review corpora while staying flat for years beforehand, so the estimator detects a distributional shift toward LLM-flavored adjectives rather than attempting to label any one sentence.","core_discovery":"On its own terms, the central discovery is that the fraction of a corpus that has been substantially modified by an LLM can be recovered from the relative frequencies of a small vocabulary of words without classifying a single document. The model assumes every sentence is drawn from $(1-\\alpha)P + \\alpha Q$, where $P$ is the human-written reference distribution and $Q$ the LLM distribution, both estimated from labeled reference corpora; maximum likelihood over the occurrence probabilities yields $\\alpha$. On semi-synthetic blends of official human reviews and LLM-generated reviews, the estimator recovers the true $\\alpha$ within 0 to 2.4 percentage points across in-distribution and out-of-distribution venues. On real post-ChatGPT corpora it finds a sharp rise in $\\alpha$ for machine-learning conference reviews (ICLR 2024 at 10.6%, EMNLP 2023 at 16.9%, NeurIPS 2023 at 9.1%, CoRL 2023 at 6.5%) but no significant rise in Nature-portfolio reviews, and it shows that proofreading alone cannot explain the increase while expanding a bulleted outline into full review text can. The dissertation also reports that reviews with higher estimated LLM modification are more likely to be submitted near the deadline, to carry low self-rated confidence, to omit citations, and to sit close to the centroid of all reviews of the same paper.","pith_inferences":["A consequence the dissertation leaves implicit: the telltale adjective set it identifies could be used as a diachronic tracer, allowing later corpora to be dated or audited for LLM influence even if the originating model changes.","A testable extension is to apply the estimator to any domain that has a stable pre-2022 human-written archive, such as news op-eds, student essays, policy documents, or clinical notes, and compare adoption rates across them; the method's requirement is only a matched human baseline and a plausible LLM reference corpus.","Because the estimator reads word frequencies, prompt-engineering to avoid AI-flavored adjectives would push the measured share down; if that became common practice, the current figures would be a lower bound on true LLM modification rather than an upper bound.","The homogenization correlation suggests a direct test: within venues where the estimated LLM share rises, the average pairwise similarity of reviews of the same paper should increase over time; if it does not, the convergence signal may reflect reviewer demographics or topic shift rather than LLM use."],"forward_implications":["A substantial minority of peer-review sentences at top machine-learning venues, between 6.5% and 16.9%, were substantially modified by LLMs in the first post-ChatGPT cycle, with the share varying by venue and review behavior.","LLM-assisted text is not evenly distributed: it concentrates near deadlines, in low-confidence reviews, in reviews without citations, and in reviews that converge toward the average review, implying measurable homogenization of feedback.","The same estimator finds double-digit shares of LLM-modified sentences in consumer complaints, corporate press releases, job postings, and UN press releases after ChatGPT's launch, so AI-assisted writing is not confined to academia.","Individual-level GPT detectors flag non-native English writing as AI-generated at high rates and can be bypassed with a single self-edit prompt, so institutions that rely on them will both penalize non-native writers and miss AI text.","LLM-generated feedback on scientific manuscripts overlaps with human reviewer comments at rates comparable to reviewer-reviewer agreement, about 30-39% versus 28-35%, and is paper-specific, suggesting a viable complement to human review."],"supporting_citations":[{"why":"Supplies the distributional GPT quantification method the dissertation uses and extends to scientific publishing and society-wide domains.","marker":"[120]"},{"why":"Shows that instance-level GPT detectors misclassify non-native English writing, motivating the population-level estimation approach.","marker":"[126]"},{"why":"Provides DetectGPT, the zero-shot instance-level baseline the proposed estimator is compared against in validation.","marker":"[145]"},{"why":"Identifies ChatGPT, whose release defines the pre/post split and whose output generates the LLM reference corpus.","marker":"[157]"},{"why":"Supplies evidence that LLM feedback overlaps with human reviewer comments and emphasizes recurring topics, used in the homogenization discussion.","marker":"[129]"},{"why":"Provides the US 8th-grade essay corpus used alongside TOEFL essays to measure detector bias on native versus non-native writing.","marker":"[94]"},{"why":"Supports the choice to focus on ChatGPT by quantifying its dominant share of generative-AI traffic.","marker":"[205]"}],"fun_headline_variants":["LLM-edited peer-review sentences: 6.5-16.9% and rising","Word-frequency hack reveals hidden LLM edits in reviews","AI detector bias hits non-native writers; new estimator sidesteps","Population-level LLM fingerprint found in peer reviews"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the human writing patterns estimated from pre-ChatGPT reviews still describe human writing after ChatGPT's launch; if reviewers shifted toward AI-flavored adjectives because of exposure, style change, or topic drift rather than direct LLM use, the estimated share of LLM-modified text would overstate true LLM modification.","fun_headline_variants_meta":{"raw":{"variants":["LLM-edited peer-review sentences: 6.5-16.9% and rising","Word-frequency hack reveals hidden LLM edits in reviews","AI detector bias hits non-native writers; new estimator sidesteps","Population-level LLM fingerprint found in peer reviews"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000348,"raw_usage":{"total_tokens":1930,"prompt_tokens":995,"completion_tokens":935,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":861}},"tokens_in":611,"tokens_out":935,"duration_ms":7922,"temperature":1.0,"reasoning_tokens":861,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:08:29.146121+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a matched corpus of post-ChatGPT, human-only peer reviews from a community that demonstrably avoided LLM assistance, with topics and author demographics matched to the machine-learning conferences studied; if the estimation method still reports a share of LLM-modified text well above its measured error when the true share is zero, the pre-ChatGPT human baseline is shifting for reasons other than LLM modification.","supporting_citations":[{"cited_title":"Van Rossum","cited_arxiv_id":null,"evidence_quote":"Supports the choice to focus on ChatGPT by quantifying its dominant share of generative-AI traffic."}],"review_version":2}