{"id":"28d19bca-ca9d-4bc7-b475-8e077ddc3b93","arxiv_id":"2412.11385","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A formal three-step corpus method identifies 21 words whose recent spike in scientific abstracts tracks ChatGPT overuse, and tests of training data, architecture, and human feedback leave RLHF as a plausible but unproven cause.","lead":"Scientists are using words like 'delve' and 'intricate' much more often in abstracts, and this paper builds a repeatable method to identify 21 such words and asks why ChatGPT favors them. Tests point toward human-feedback training as a possible cause, but the evidence is mixed and the study is honest about being exploratory.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No pre-LLM negative control: the three-step method could identify 'focal words' in an earlier period, so the 21-word list's causal attribution to LLM usage is not yet calibrated.","rationale":"The reader correctly sees the causal link as conditional. My concern sharpens this: besides manual annotation and prompt realism, the method has no false-positive calibration. The three-step selection is essentially a high-recall filter; by construction it will find some words in any two time slices. Without a pre-LLM control, the 21-word list could reflect ordinary low-frequency drift plus ChatGPT's stylistic quirks rather than LLM-driven change. The proposed test directly addresses this and is feasible with the authors' existing code and API access. I do not see this as fatal; the paper is transparent and the method is a useful transferable artifact, but the causal claim in the abstract should be conditional on passing the negative control. This is consistent with the reader's conditional verdict, so no change to the verdict itself is needed.","tokens_in":15004,"tokens_out":5415,"duration_ms":51088,"concrete_test":"Apply the exact three-step method to an earlier PubMed pair such as 2016 vs 2018 (or 2018 vs 2020), using ChatGPT-3.5 with the same prompts to generate abstracts from the earlier year's papers. Count the resulting 'focal words' and compare their semantic profile and count to the 21 in Table 2. If a comparable number of focal words emerges, or if words like 'delve' are overused in the generated abstracts even though they did not spike in the pre-LLM period, the 2024 list is not specific to LLM-era language change; if the pre-LLM control yields few or no focal words, the method's calibration is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('21 focal words whose increased occurrence... likely the result of LLM usage') depends on the pipeline in §2: a 2020→2024 chi-square spike, a manual 'unexplained' filter, and overuse by ChatGPT-3.5 in a single two-stage summarize-then-expand task. The weakest point is that no negative control calibrates this procedure. The same pipeline applied to a pre-LLM pair of years (e.g., 2016 vs 2018) would presumably also produce some unexplained spikes; because ChatGPT-3.5's overuse list is generated independently of the real-world cause of those spikes, any pre-LLM 'focal words' would be false positives by the paper's own causal logic. The paper offers no estimate of this false-positive rate. The problem is compounded by §2's explicit statement that the two-stage prompting procedure is only a suspicion about how scientists used LLMs in 2022–2024; if real usage differed, the overuse list is an artifact of the prompt pipeline. The RLHF entropy comparison in §5 is acknowledged as indirect, but the focal-word identification itself needs the missing control before the causal claim can be accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a formal, transferable method for identifying lexical items whose increased occurrence in scientific abstracts is plausibly attributable to LLM usage. The method combines (a) a chi-square-based spike in PubMed abstracts between 2020 and 2024, (b) a manual 'unexplained spike' annotation, and (c) significant overuse by ChatGPT-3.5 in a two-stage summarize-then-expand abstract-generation task, yielding 21 focal words. The paper then explores why these words are overused, finding no supporting evidence for architecture, algorithms, or training data, and mixed evidence regarding RLHF: a Llama 2-Base vs. Llama 2-Chat entropy comparison is consistent with an RLHF contribution, while an exploratory Prolific study shows a significant aversion to 'delve' when it appears in the first sentence but no overall preference for focal-word abstracts.","tokens_in":15231,"tokens_out":4194,"duration_ms":38325,"significance":"If the central claim holds, the paper provides a reproducible and transferable method for detecting LLM-driven lexical change in scientific writing, and it offers a clearly posed 'puzzle of lexical overrepresentation' with initial probes into causal factors. The manuscript is unusually transparent: the full pipeline is described, code and stimuli are on GitHub, the authors explicitly acknowledge the limitations of the Llama entropy comparison and the exploratory nature of the human study, and the absence of evidence for architecture- or data-based explanations is honestly reported. The empirical finding that human participants react differently to 'delve' than to other focal words is a useful, falsifiable observation for future work. The main weakness is that the causal attribution of the 21 focal words to LLM usage is not calibrated against a pre-LLM negative control, and the manual annotation and prompt-design choices are load-bearing but unvalidated.","major_comments":[{"comment":"The three-step method has no negative control. Applying the same pipeline to a pre-LLM pair of years (e.g., 2016 vs. 2018) would presumably produce a set of words that show an unexplained spike and are overused by ChatGPT-3.5, yet those words' increased occurrence would not be caused by LLM usage. Because the central claim is that the 21 focal words' increased occurrence is 'likely the result of LLM usage,' the paper needs to report such a control and estimate the false-positive rate of the pipeline; without it, the causal attribution is not yet supported.","section":"Section 2"},{"comment":"The manual annotation of 'unexplained' spikes is load-bearing but lacks reliability evidence. The authors state that they 'independently reviewed' the list and 'in cases of disagreement, we included the word on our list,' but no inter-annotator agreement statistic is reported and the criteria for excluding words with 'an obvious explanation' are not operationalized. The maximally lenient disagreement rule is likely to bias the list toward false positives; please report agreement rates and provide a transparent, reproducible criterion for what counts as an unexplained spike.","section":"Section 2"},{"comment":"The two-stage prompt (summarize then expand) is justified only by the sentence, 'We suspect that the most common way of using an LLM to generate an abstract... involved providing important fragments of a paper.' This suspicion is load-bearing because the ChatGPT-overuse list in step 3 is entirely a function of the specific prompt design. If real LLM-assisted writing was done differently (e.g., direct generation, editing human prose, or other prompt styles), the 21 focal words could be artifacts of the pipeline. Please validate the assumption with evidence about actual usage, or at least show that the focal-word list is stable across plausible alternative prompt designs.","section":"Section 2"},{"comment":"The Llama 2-Base vs. Llama 2-Chat entropy comparison does not isolate focal-word overrepresentation. The large drop in per-word entropy for AI abstracts in the chat model could be driven by any number of other stylistic properties of ChatGPT-generated text (e.g., formulaic sentence frames, reduced syntactic variety, or repetition of non-focal function words). The paper acknowledges this limitation in principle, but the conclusion that 'fine-tuning and RLHF... might be important contributors' to lexical overrepresentation specifically requires a test that controls for focal-word density, for example by comparing entropy for human abstracts with and without artificially inserted focal words, or by regressing entropy differences on focal-word frequency.","section":"Section 5, Table 1"}],"minor_comments":[{"comment":"The corpus description says 'more than 5.2 billion tokens (inflected forms)' but does not specify the tokenization or normalization procedure used on PubMed abstracts; please provide these details for reproducibility.","section":"Section 2"},{"comment":"The focal-word list appears both in Appendix A and in Table 2 with inconsistent column naming; please harmonize the two presentations or make one a cross-reference.","section":"Appendix A and Table 2"},{"comment":"The statement that 'considerably more than half' of the focal-word abstracts were delve-initial should be replaced with the exact proportion, and the criterion for 'first sentence' should be defined (e.g., the first sentence of the generated abstract).","section":"Section 6"},{"comment":"In the per-word entropy formula, the variables L, n, and the conditioning context are not defined; please clarify whether this is the average per-token entropy over a sequence and how the probability p(x_i) is computed.","section":"Equation (1)"},{"comment":"The table rows labeled '8b Llama 3-Base' and '8b Llama 3.1-Base' are hard to parse; please format model names consistently (e.g., Llama 3 8B).","section":"Appendix D"}],"recommendation":"major_revision","confidential_remarks":"The missing negative control is the key barrier to the paper's central causal claim, but it is a fixable one within the paper's scope (e.g., apply the pipeline to a pre-LLM year pair and report the resulting false-positive rate). The manual-annotation and prompt-validity issues also require substantial revision, but the authors have already shown a willingness to report limitations transparently, so I expect they can address these points in a revision. The paper is a good fit for an NLP or computational-linguistics venue; the topic has broad interest and the method is reproducible. No concerns about citation or novelty disclosure beyond what is stated in the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. The core contribution is the three-step pipeline: significant 2020-2024 spike, manual \"unexplained\" filter, then ChatGPT overuse. That is a real improvement on the informal lists floating around, and because the code and data are on GitHub, it's reproducible. The 21 focal words are concrete and the method is transferable to other corpora and models, as the paper says.\n\nThe causal part is where I'd be careful. The authors are appropriately cautious about RLHF - they call the experiment exploratory and the model testing indirect. That's honest. But the stress-test note lands: there is no pre-LLM negative control. If you ran the same pipeline on 2016 vs 2018, you would get some unexplained spikes, and since ChatGPT overuse is generated independently of the real cause, some of those would be false positives by the paper's own logic. Without an estimate of that false-positive rate, the abstract's \"likely the result of LLM usage\" is stronger than the evidence supports. The manual annotation also needs a reliability measure, and the two-stage prompt is a suspicion, not a demonstrated match to real usage. The paper flags the prompt issue in Section 2, but the focal-word list depends on it.\n\nThe Llama entropy comparison is a nice idea but indirect: base vs chat differ in fine-tuning and RLHF, but also in other ways, and the paper admits it can't pin the difference on focal words. The human study is underpowered and the authors split it post hoc. They acknowledge this. So the RLHF conclusion is plausible but far from established.\n\nOverall, the method is the contribution. The causal probes are preliminary. That's a reasonable exploratory paper, not a settled answer. It will be useful to anyone tracking LLM influence on academic writing, and the focal-word list is a concrete starting point.\n\nRecommendation: send it to review. A serious referee will ask for a negative control and inter-annotator agreement, but the method is worth publishing and the review process can tighten the causal claims.","headline":"A genuinely transferable method for spotting LLM-typical words; the causal story is plausible but the missing pre-LLM control keeps it from being settled.","tokens_in":15746,"tokens_out":1716,"would_cite":true,"duration_ms":15984,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Recent spikes in words like 'delve' in scientific abstracts are likely LLM-driven, and a new method isolates 21 such words.","keywords":["lexical overrepresentation","large language models","scientific English","corpus analysis","reinforcement learning from human feedback","linguistic change","ChatGPT","delve"],"falsifier":"The cleanest test: rerun the three-step pipeline on the same PubMed years but with a different way of generating AI abstracts, for example feeding ChatGPT-3.5 the full abstract text instead of a summary, or prompting it to paraphrase rather than expand notes. If the 21 focal words no longer emerge as overused, the focal list is an artifact of the specific two-stage prompt rather than a property of LLM-assisted scientific writing. A second decisive observation would be a transparency audit of an RLHF training run showing whether responses containing 'delve' actually received higher human feedback scores than matched responses without it; if they did not, RLHF cannot be the cause.","tokens_in":14787,"feed_emoji":"📈","tokens_out":7951,"duration_ms":63850,"temperature":0.7,"pith_summary":"This paper tries to establish that the recent surge of words like 'delve,' 'intricate,' and 'underscore' in scientific abstracts is a real, measurable, LLM-driven change in Scientific English, and not just anecdote. It contributes a formal three-step method: find words whose frequency per million tokens spiked between 2020 and 2024, keep only spikes with no obvious scientific or worldly explanation, then ask whether ChatGPT-3.5 overuses those same words when writing abstracts. Applied to more than 5.2 billion tokens of PubMed, the method yields 21 'focal words.' The paper then asks why these words are overused and reports negative evidence against training data, architecture, and algorithm choices, plus mixed evidence that reinforcement learning from human feedback plays a role. The experimental part is exploratory and suggests human readers may actually be put off by 'delve' in the first sentence of an abstract.","feed_headline":"Method pins 'delve' and 20 other word spikes on ChatGPT","feed_subtitle":"A three-step PubMed screening identifies the words LLMs likely injected into scientific writing.","key_machinery":"The load-bearing mechanism is a three-step screening pipeline. Step 1 counts occurrences per million tokens in PubMed abstracts for 1975 through May 2024 and keeps roughly 7,300 tokens with a significant $\\chi^2$ increase between 2020 and 2024. Step 2 filters this list by hand to 50 tokens whose spike lacks an obvious explanation in science or world events. Step 3 generates 9,953 AI abstracts from 10,000 real 2020 PubMed abstracts using a two-stage ChatGPT-3.5 prompt (summarize, then write an abstract from the summary), tests each candidate token for significant overuse with a $\\chi^2$ comparison against the human originals, and keeps the 21 tokens that pass all three gates. A secondary instrument is per-word entropy, computed for Llama 2-Base and Llama 2-Chat on the same human and AI abstracts; the entropy gap between the two models is used to isolate fine-tuning and RLHF as the factor that differs.","core_discovery":"On the paper's own terms, the central discovery is that a substantial share of the recent lexical shift in biomedical abstracts can be attributed to LLM assistance, embodied in 21 focal words whose increase is significant, unexplained by external events, and mirrored by statistically significant overuse in ChatGPT-3.5-generated abstracts. The authors do not claim to have fully solved 'the puzzle of lexical overrepresentation': they find no evidence that architecture, tokenization or other algorithms, or training and fine-tuning data explain the overuse. Comparison of Llama 2-Base with Llama 2-Chat, which differ mainly in fine-tuning and RLHF, shows that the chat model is considerably less 'surprised' (lower per-word entropy) by AI-generated abstracts containing focal words, which is consistent with RLHF contributing to the overuse. Their online preference study failed to show an overall preference for abstracts containing focal words and found that when 'delve' opened the abstract, participants significantly preferred the version without it, which the authors interpret as public wariness toward that particular word.","pith_inferences":["The method could be run on non-English scientific corpora to see whether LLM-driven lexical overrepresentation is a universal phenomenon of current models or an artifact of English-language training.","If 'delve' is becoming socially marked, one testable prediction is that its frequency in LLM outputs should decline over time as RLHF raters begin to penalize it; monitoring deployed model versions would settle this.","The paper's 'decoupling of form and content' hypothesis implies that other stylistic tics of LLMs, such as bullet-point structure or hedging phrases, might be detectable by the same spike-and-overuse screening, extending the tool beyond single words.","A stronger experiment would compare preference ratings for focal words embedded in otherwise identical abstracts varying only the word, avoiding the forced-insertion artifacts the authors identify; that design would directly estimate the RLHF reward signal."],"forward_implications":["If the method is sound, the 21 focal words give a concrete, measurable fingerprint of LLM-assisted writing in biomedical abstracts; the same fingerprint can be computed for other corpora and years.","Because almost all focal words were already rising before ChatGPT, the paper implies LLMs are accelerating an existing lexical drift rather than inventing it from scratch.","If RLHF- or fine-tuning-driven overuse is real, then the preference of human raters for certain words has directly shaped machine vocabulary, and future rounds of feedback will reshape it again.","The 'delve' effect suggests public discourse about AI-typical words can feed back into human preferences, potentially changing the next round of RLHF data.","The method can be transferred to LLMs other than ChatGPT-3.5, and the paper's appendix shows GPT-4o-mini behaves similarly for most focal words."],"supporting_citations":[{"why":"Quantifies ChatGPT's footprint in scientific writing through excess vocabulary; the baseline this paper's spike-and-overuse approach builds on.","marker":"Kobak et al. 2024"},{"why":"Maps the increasing use of LLMs in scientific papers, providing evidence that AI-generated text has entered the literature.","marker":"Liang et al. 2024b"},{"why":"Shows AI-modified content in conference peer reviews, supporting the premise that LLM output is widespread in academic writing.","marker":"Liang et al. 2024a"},{"why":"Releases Llama 2 Base and Llama 2 Chat, the model pair the entropy comparison uses to isolate fine-tuning and RLHF.","marker":"Touvron et al. 2023"},{"why":"Describes the InstructGPT RLHF pipeline that the paper hypothesizes as a source of lexical overrepresentation.","marker":"Ouyang et al., 2022"},{"why":"Introduces deep reinforcement learning from human preferences, the training mechanism behind the RLHF hypothesis.","marker":"Christiano et al., 2017"},{"why":"Shows how to fine-tune language models from human preferences, the direct mechanism tested by the Llama comparison.","marker":"Ziegler et al., 2019"},{"why":"Provides the Leipzig Corpora Collection, one of the comparison corpora used to argue focal words are not overrepresented in training-like data.","marker":"Goldhahn et al. 2012"},{"why":"Provides the International Corpus of English data used to test whether focal-word overuse comes from a particular English variety.","marker":"Kirk and Nelson 2018"},{"why":"Supplies the PubMed abstracts that form the 5.2-billion-token corpus from which focal words are extracted.","marker":"National Library of Medicine, 2023"}],"fun_headline_variants":["21 words, including 'delve', show ChatGPT's mark on science","Why LLMs overuse 'delve'? Study tracks 21 words to AI","ChatGPT's 'delve' spike: 21 words tied to LLMs in abstracts","Study: 21 word surges in science papers likely from LLMs","Unraveling 'delve': 21 words expose ChatGPT's influence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the authors' hand-screening of 'unexplained' spikes is reliable, and that their two-stage ChatGPT-3.5 summarize-then-expand task approximates how scientists actually used LLMs in 2022-2024; the authors state this only as a suspicion, and if real usage differed, the 21 focal words could be an artifact of their pipeline rather than a genuine LLM-driven shift.","fun_headline_variants_meta":{"raw":{"variants":["21 words, including 'delve', show ChatGPT's mark on science","Why LLMs overuse 'delve'? Study tracks 21 words to AI","ChatGPT's 'delve' spike: 21 words tied to LLMs in abstracts","Study: 21 word surges in science papers likely from LLMs","Unraveling 'delve': 21 words expose ChatGPT's influence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000645,"raw_usage":{"total_tokens":2988,"prompt_tokens":991,"completion_tokens":1997,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1895}},"tokens_in":607,"tokens_out":1997,"duration_ms":12738,"temperature":1.0,"reasoning_tokens":1895,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:58:55.419427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The cleanest test: rerun the three-step pipeline on the same PubMed years but with a different way of generating AI abstracts, for example feeding ChatGPT-3.5 the full abstract text instead of a summary, or prompting it to paraphrase rather than expand notes. If the 21 focal words no longer emerge as overused, the focal list is an artifact of the specific two-stage prompt rather than a property of LLM-assisted scientific writing. A second decisive observation would be a transparency audit of an RLHF training run showing whether responses containing 'delve' actually received higher human feedback scores than matched responses without it; if they did not, RLHF cannot be the cause.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the International Corpus of English data used to test whether focal-word overuse comes from a particular English variety."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the PubMed abstracts that form the 5.2-billion-token corpus from which focal words are extracted."}],"review_version":1}