Pith. sign in

REVIEW 4 major objections 5 minor 22 references

Most biomedical publications show signs of LLM-assisted writing

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read By the end of 2025, 89% of open-access biomedical papers carried excess LLM-associated vocabulary, according to a new word-frequency analysis.

desk verdict Useful empirical breakdowns and a clean lower-bound estimator, but the '89%' point estimate overreaches: tightness is asserted, not shown. read the letter →

arxiv 2608.10715 v1 pith:HSUCUD5T submitted 2026-08-11 cs.CL cs.AIcs.CYcs.DLcs.SI

classification cs.CLcs.AIcs.CYcs.DLcs.SI
keywords LLM-assistedwritingbiomedicalpublicationswordfrequencyanalysismarkerwordsprevalenceestimationPubMedCentralacademicintegrityChatGPT
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that by December 2025, 89% of open-access biomedical papers in PubMed Central showed signs of LLM-assisted writing or editing, up from about 19% in 2023. The authors argue that earlier frequency-gap methods systematically underestimated the true prevalence, and they offer a way to convert a lower bound into a point estimate by assuming, with simulation support, that the best lower bound is tight. The result matters because it suggests that LLM assistance has become the default in biomedical publishing, with implications for research integrity, peer review, and language equity. The method tracks the share of papers containing any of a fixed set of 'marker words' whose usage rose sharply after ChatGPT's release.

What carries the argument

The central identity is the two-component mixture q = (1−β)p_human + β p_LLM, which yields the lower bound β ≥ (q − p_human)/(1 − p_human). The paper's estimate β̂ is the maximum over word-set thresholds of this lower bound, computed with the counterfactual p̂_human from linear extrapolation of 2018–2022 frequencies, justified by simulation. The 379 marker words and the threshold selection for word-set size carry the analysis; the critical step is the 'max lower bound' assumption that the optimal word set has p_LLM close to 1, which converts a bound into a point estimate.

What would settle it

Take a set of biomedical papers published in 2025 whose authors declare no LLM use and whose text passes manual inspection, and compare their marker-word frequency to the linear extrapolation from 2018–2022; if the frequency runs significantly above the trend, the counterfactual is biased and the 89% figure is inflated.

Watch

Extended reading notes

Core claim

Using full texts of 1,194,287 open-access biomedical papers from PubMed Central, the authors estimate that by December 2025, 89% of papers contained excess LLM-associated vocabulary. The estimate comes from a set of 379 non-content marker words (e.g., 'these', 'potential', 'delves') identified in prior work. For each candidate word set, they compare the observed share of papers containing any marker word to a counterfactual share extrapolated from the 2018–2022 linear trend. The excess is converted into an LLM-usage prevalence using the mixture identity q = (1−β)p_human + β p_LLM with p_LLM bounded above by 1, and the maximum lower bound over word-set sizes is taken as the estimate. The paper argues this is a tighter and more accurate estimate than earlier frequency-gap or mixture-model approaches, supported by a simulation that recovers the true β.

Load-bearing premise

The whole estimate rests on believing that the pre-2023 linear trend in how often humans used these marker words would have continued unchanged through 2025 if ChatGPT had never existed, and that among the candidate marker-word sets one achieves near-certain use in LLM-edited text.

Editorial extensions

If this is right

  • If correct, the 89% figure means LLM assistance is now the norm in biomedical publishing, not a minority practice.
  • Discussion sections are roughly twice as LLM-influenced as Methods (68% vs 32% in length-controlled 255-word crops), suggesting authors use LLMs most for framing and interpretation.
  • The gap between native-English-majority countries (37%) and others (72%) indicates LLM writing tools are closing language gaps but also creating large cross-country differences in dependence.
  • Because frequency-based methods estimate direct LLM editing and surveys report similar or higher usage, policies that presume disclosure may need to assume near-universal exposure to LLM-assisted text.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the marker-word trend continues, the extrapolated counterfactual will become increasingly unreliable, so the 89% figure is a moving target that will need re-estimation with fresh human-written baselines, such as preprints that predate ChatGPT or non-academic writing.
  • The reported country differences could reflect differences in English proficiency and reliance on LLM translation rather than direct LLM drafting; a follow-up could separate translation-assisted from drafting-assisted writing.
  • The method could be ported to other corpora, such as grant applications, clinical notes, or policy documents, wherever a pre-LLM baseline exists.
  • The tightness assumption is testable: analyzing version-control histories of manuscripts would give a direct measure of p_LLM and validate or correct the estimation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a frequency-based estimator for the prevalence of LLM-assisted writing in a corpus. The authors apply it to open-access biomedical papers from PubMed Central, estimating β̂=0.89 for full papers in December 2025, with Discussion sections at 0.68 and Methods at 0.32 when length is controlled via 255-word crops. The method builds on a 379-marker-word list from the authors' prior work, models the counterfactual human usage p̂_human by linear extrapolation of 2018–2022 frequencies, and uses β̂ = max_T (q−p̂_human)/(1−p̂_human) after discarding thresholds with large standard errors. A simulation experiment is reported to support the estimator's accuracy. The paper also compares results across sections and affiliation countries, and frames the result as an unbiased estimate rather than a lower bound.

Significance. If the tightness of the lower bound and the faithfulness of the counterfactual extrapolation were established, this would be a valuable contribution: it would provide the first corpus-level estimate of LLM usage that goes beyond a lower bound, with a transparent public dataset and code (GitHub). The section-level and country-level comparisons are policy-relevant and extend prior work. However, the load-bearing 'unbiased' claim is not currently supported: the estimator remains a lower bound whose gap from the true β is not quantified in real data, and the simulation validates the procedure under the exact assumption (p_LLM close to 1) that is at issue. The core scientific claim is therefore conditional on additional sensitivity analyses and reporting of the unobserved quantities (q and p̂_human for selected thresholds).

major comments (4)
  1. [§2, Eq. (4) and the sentence 'As this maximum was typically achieved at q close to 1...'] The tightness assumption is not logically justified. From q = (1−β)p_h + β p_LLM one can only conclude p_LLM ≥ q, not p_LLM ≈ 1; for example, with q = 0.98 and p_h = 0.80, the estimator returns β̂_LB = 0.90 while the true β could be 1.00, a bias of 0.10 even though q is close to 1. The residual bias is (q−p_h)(1−p_LLM)/((p_LLM−p_h)(1−p_h)), which is not bounded by the reported data. The authors should report q and p̂_human for every selected threshold (Table 1 rows and Figure 1f/g), and provide a sensitivity analysis of β̂ as a function of an assumed p_LLM (e.g., 0.90, 0.95, 0.99) or a formal upper bound on the bias. Until then, the abstract's '89%' should be stated as 'at least 89%'.
  2. [§4, 'Simulation experiment'] The simulation draws p_LLM = (1+δ) p_human with δ ~ U(0.5, 5), which by construction drives the union probability over a broad threshold set close to 1, so the simulation validates the estimator only in the regime where p_LLM ≈ 1, i.e., the very condition the paper asserts without real-data support. The simulation does not test the estimator under moderate p_LLM (e.g., 0.8–0.95), where the lower-bound gap is material, nor under a misspecified p_human trend (e.g., quadratic drift or a level shift). Please rerun the simulation with fixed p_LLM values and with a nonlinear counterfactual, and report bias and coverage for β in [0,1]. The current Figure 1h therefore does not substantiate the claim of 'high accuracy' for the real-data scenario.
  3. [§3, 'Limitations' paragraph beginning 'Our estimate hinges...'] The paper correctly identifies that β̂ is only valid if p̂_human faithfully estimates p_human in 2025. This is a load-bearing assumption for the 0.89 point estimate, yet no placebo test is provided: for example, fitting the linear extrapolation on 2014–2017 data to predict 2018–2022, or applying the same estimator to a control set of words that are not associated with LLM use, would calibrate the extrapolation error. Without such a demonstration, the excess (q − p̂_human) could reflect non-LLM vocabulary drift or human stylistic adaptation to LLM output, as the authors themselves acknowledge. The authors should either add such a calibration or explicitly present the headline number as a lower bound subject to this assumption.
  4. [§2, 'To find an optimal value of T' and Fig. 1g] Selecting the threshold T that maximizes β̂_LB over a grid introduces an upward finite-sample selection bias: each individual β̂_LB is a valid lower bound under the model, but the maximum of many noisy lower-bound estimates will tend to exceed the true maximum lower-bound curve. The simulation incorporates the same selection and reports negligible bias, but that simulation again relies on near-saturated p_LLM. Please report the number of thresholds within one standard error of the maximum, provide a selection-adjusted standard error (e.g., bootstrap or max-bias correction), or show that the maximum is attained over a plateau rather than at a single noisy point.
minor comments (5)
  1. [Abstract] The phrase '89% of papers show excess of LLM-associated vocabulary' should read 'at least 89%' given that the estimator is a lower bound; the same wording appears in the Discussion and should be made consistent.
  2. [§2, displayed equations] The notation alternates between 'p human' and 'p_h' (or 'p_human') in Eqs. (2)–(4). Please define one symbol (e.g., p_h for human usage frequency) and use it consistently; also distinguish the estimated p̂_human from the true p_human in the derivation.
  3. [§4, 'Standard errors'] The variance formula for Var[β̂] appears to treat q and p̂_human as independent; the methods text should note when this independence assumption is used and whether the covariance term is negligible in practice.
  4. [§4, 'PubMed Central'] The text uses 'Pubmed' in the first line of the Methods section; it should be 'PubMed' for consistency with the rest of the manuscript.
  5. [Supplementary Tables] Table S1 and S2 contain references without full citation details (e.g., author-year placeholders such as '2024' and '2025' are used in the table cells). Please include the full author list or a footnote mapping the table entries to the reference list.

Circularity Check

3 steps flagged · score 6.0 of 10

The 89% headline is the maximum lower bound over thresholds fitted to the same 2025 data, renamed as an unbiased estimate; the marker list comes from the authors' own prior work, and the validating simulation builds in the pLLM≈1 condition.

  1. self definitional [Section 2, first paragraph]
    "In our prior work (Kobak et al., 2025), we identified 379 non-content words that showed markedly increased usage in PubMed abstracts in 2024 compared to earlier years, and attributed this excess to LLM-assisted writing or editing. These LLM markers are words that can be used in any research context, such as these, potential, or delves. We confirmed that in our PMC dataset almost all of these words saw increased usage in abstracts in 2025 (Figure 1a), so we based our subsequent analysis on this list."

    The marker vocabulary is the dependent variable of the authors' own prior study: words selected because their post-ChatGPT frequency increased, with the increase attributed to LLMs. Using those same words to measure post-ChatGPT LLM usage makes the 2025 estimate a re-measurement of the selection criterion. The only external justification is a self-citation by overlapping authors. If the 2024 excess had any non-LLM cause (e.g., stylistic drift), that cause is inherited by the 2025 estimate. This is not a minor self-citation: the entire estimator is defined on this list.

  2. fitted input called prediction [Section 2, threshold optimization paragraph]
    "To find an optimal value ofT, we used a grid ofT values. For each T value, we found q and ˆphuman for 2025 using regression on yearly 2018–2022 frequencies (Figure 1f), and selected T with the highest ˆβLB (Figure 1g), after discarding all T values yielding standard error above 0.025 (see Methods). As this maximum was typically achieved at q close to 1 (and hence pLLM close to 1, which is the approximation used in computing the lower bound), we assume that our lower bound is acceptably tight and use it as the estimate of LLM usage frequency β̂=max_T(β̂_LB)."

    The reported estimate is by construction the maximum lower bound over thresholds, with T chosen on the same 2025 data that is then summarized. The headline 89% is therefore not an independently derived point estimate; it is the largest of the lower bounds after in-sample optimization. Converting that lower bound into an unbiased estimate rests entirely on the assertion that q close to 1 implies pLLM close to 1, which is not a consequence of the mixture equation q=(1−β)p_h+βp_LLM. This assertion is exactly the condition needed for unbiasedness, so the central prediction reduces to its own lower bound plus an unverified assumption.

1 more flagged steps
  1. other [Methods, Simulation experiment]
    "For each word, its phuman was sampled from a gamma distribution with shape parameter 2 and scale parameter 0.02. Its pLLM was set (1 +δ) times larger, with δ∼U( 0.5, 5). ... We obtained | ˆβ−β|< 0.02 for all simulated values β∈[ 0, 1] (Figure 1h), confirming the reliability of our estimation procedure."

    The simulation is the only evidence that β̂ is unbiased, but its generative model guarantees the very condition the estimator needs: pLLM is at least 1.5×phuman for every word, and the procedure takes unions over hundreds of words, so the union pLLM for the selected threshold sets is near 1. Under that condition, β̂_LB≈β by construction. The simulation therefore confirms unbiasedness only under the tightness assumption it was supposed to test; it provides no information about the real pLLM. This is a circular validation: the conclusion (estimator is unbiased) is encoded in the input distribution.

full rationale

The paper's algebraic lower bound β̂_LB=(q−p̂_human)/(1−p̂_human) is mathematically valid, and the paper is candid that it is a lower bound. However, the central reported quantity is not this bound but the claim that 89% of papers were LLM-assisted. That step rests on three linked reductions. First, the marker vocabulary is taken from the authors' own prior work, where the same words were selected because they showed post-ChatGPT excess and attributed to LLMs; using them to 'estimate' LLM usage is partially a re-measurement of the selection criterion, and the only support is a self-citation. Second, the threshold T is fitted to the same 2025 data by maximizing β̂_LB, and the final value β̂=max_T β̂_LB is then called 'the estimate of LLM usage frequency'; the headline is thus the maximum lower bound, not an independently derived point estimate. The leap from lower bound to estimate is justified only by the assertion that q close to 1 implies pLLM close to 1, which is not a consequence of the mixture equation and is exactly the condition required for unbiasedness. Third, the simulation validation does not test that condition: it generates pLLM = (1+δ)p_human with δ≥0.5 over hundreds of words, guaranteeing the union pLLM is near 1 for the threshold sets the procedure selects, so the estimator recovers β by construction. The paper's own Limitations section honestly flags the p̂_human extrapolation as an assumption, which is a correctness risk rather than circularity; but the tightness of the lower bound is not similarly supported. Because the 89% figure reduces to a lower bound plus an untested saturation assumption, and the only validation encodes that assumption, the central claim is partially circular. Score 6.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper's estimate rests on a marker-word list inherited from the authors' prior work, a threshold fitted to the 2025 data, and two unverified assumptions: the linear extrapolation of the human baseline and the near-saturation of p_LLM for the chosen word set. The simulation builds in the saturation condition, so it does not independently validate the point estimate.

free parameters (3)
  • Marker-word set for each section (via threshold T) = Not reported; chosen on a 19-point grid to maximize β̂_LB for 2025
    The set G(T) determines the measured frequency q and therefore the estimate. Optimizing T on the 2025 data directly tunes the headline number.
  • Threshold grid and exclusion cutoffs (SE ≤ 0.025, p̂_human < 0.999) = Grid of 19 thresholds; cutoffs 0.025 and 0.999
    These hand-chosen rules filter which lower bounds are admitted and affect the maximum and hence the final estimate.
  • Inherited 379-marker-word list = 379 words from Kobak et al., 2025, selected for 2024 frequency excess in PubMed abstracts
    The entire estimate is conditional on this list. It was produced by the same research group and fitted to the same class of phenomenon, so it is a carried-over fitted quantity.
assumptions (5)
  • domain assumption Mixture model q = (1−β) p_human + β p_LLM for document-level marker-word presence
    Assumes no other process changes marker-word frequency besides the human-LLM mixture; the paper's equations rely on this decomposition.
  • domain assumption Linear extrapolation of 2018-2022 frequencies is a faithful counterfactual for human writing in 2023-2025
    Stated in Limitations as the central assumption. If human style drifts for non-LLM reasons, or adapts due to LLM exposure, the excess is biased.
  • ad hoc to paper For the selected threshold, p_LLM ≈ 1, i.e., almost every LLM-assisted text contains at least one marker word from G(T)
    Needed to convert the lower bound into a point estimate. Justified only by the simulation, not by real-data evidence.
  • domain assumption Marker words are stable non-content indicators of LLM usage rather than proxies for topic or editorial shifts
    Required for the causal attribution from word-frequency excess to LLM assistance; no independent check against topical drift is provided.
  • ad hoc to paper In the simulation, 500 words are independent with p_LLM = (1+δ) p_human, δ ~ U(0.5,5)
    This setup guarantees a large word-set with p_LLM near 1, which is exactly the condition the method needs; it may not transfer to real corpora with dependencies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Most biomedical publications show signs of LLM-assisted writing." pith.science (2026). https://pith.science/paper/HSUCUD5T

@misc{pith2026260810715,
  author       = {Pith},
  title        = {Pith review of: Most biomedical publications show signs of LLM-assisted writing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HSUCUD5T}},
  note         = {Machine review of arXiv:2608.10715}
}
read the original abstract

Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%. We believe that our estimates are crucial to shape future guidelines and policies.

Figures

Figures reproduced from arXiv: 2608.10715 by the authors.

Figure 1
Figure 1. Overview of the estimation and validation process. (a) [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Estimated LLM usage for individual sections [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

22 extracted references · 12 canonical work pages

  1. [1]

    Quantitative analysis of AI-generated texts in academic research: A study of AI presence in arxiv submissions using AI detection tool.arXiv preprint arXiv:2403.13812,

    Arslan Akram. Quantitative analysis of AI-generated texts in academic research: A study of AI presence in arxiv submissions using AI detection tool.arXiv preprint arXiv:2403.13812,

  2. [3]

    Insights 2024 — attitudes toward AI — Elsevier,

    Elsevier. Insights 2024 — attitudes toward AI — Elsevier,

  3. [4]

    archive.org/web/20260410222804/https://www

    URL https://web. archive.org/web/20260410222804/https://www. elsevier.com/insights/attitudes-toward-ai/ the-current-ai-landscape. Michael Eppler, Conner Ganjavi, Lorenzo Storino Ramac- ciotti, Pietro Piazza, Severin Rodler, Enrico Checcucci, Juan Gomez Rivas, Karl F. Kowalewski, Ines Rivero Be- lench´on, Stefano Puliatti, Mark Taratkin, Alessandro Veccia,...

  4. [5]

    Can GenAI improve academic performance? evidence from the social and behavioral sciences.arXiv preprint arXiv:2510.02408,

    Dragan Filimonovic, Christian Rutzer, and Conny Wunsch. Can GenAI improve academic performance? evidence from the social and behavioral sciences.arXiv preprint arXiv:2510.02408,

  5. [7]

    Human-LLM coevo- lution: Evidence from academic writing

    Mingmeng Geng and Roberto Trotta. Human-LLM coevo- lution: Evidence from academic writing. InFindings of the Association for Computational Linguistics: ACL 2025, pages 12689–12696, 2

  6. [8]

    contamination

    Andrew Gray. ChatGPT “contamination”: estimating the prevalence of LLMs in the scholarly literature.arXiv preprint arXiv:2403.16887,

  7. [9]

    Estimating the prevalence of LLM-assisted text in scholarly writing.arXiv preprint arXiv:2512.01560,

    7 Holzwarth et al.LLM-associated vocabulary in biomedical publications Andrew Gray. Estimating the prevalence of LLM-assisted text in scholarly writing.arXiv preprint arXiv:2512.01560,

  8. [10]

    How much are LLMs changing the language of academic papers after Chat- GPT? a multi-database and full text analysis.Scientomet- rics 2026, pages 1–21, 9

    Kayvan Kousha and Mike Thelwall. How much are LLMs changing the language of academic papers after Chat- GPT? a multi-database and full text analysis.Scientomet- rics 2026, pages 1–21, 9

Show all 22 references
  1. [11]

    Monitoring AI- modified content at scale: a case study on the impact of ChatGPT on AI conference peer reviews

    Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Hao- tian Ye, Sheng Liu, Zhi Huang, et al. Monitoring AI- modified content at scale: a case study on the impact of ChatGPT on AI conference peer reviews. InProceedings of the 41st...

  2. [12]

    Llms as research tools: A large scale survey of researchers’ usage and perceptions.arXiv preprint arXiv:2411.05025,

    Zhehui Liao, Maria Antoniak, Inyoung Cheong, Evie Yu- Yen Cheng, Ai-Heng Lee, Kyle Lo, Joseph Chee Chang, and Amy X Zhang. Llms as research tools: A large scale survey of researchers’ usage and perceptions.arXiv preprint arXiv:2411.05025,

  3. [13]

    ChatGPT as linguistic equalizer? quantifying LLM- driven lexical shifts in academic writing.arXiv preprint arXiv:2504.12317,

    Dingkang Lin, Naixuan Zhao, Dan Tian, and Jiang Li. ChatGPT as linguistic equalizer? quantifying LLM- driven lexical shifts in academic writing.arXiv preprint arXiv:2504.12317,

  4. [14]

    Towards the relationship be- tween AIGC in manuscript writing and author pro- files: evidence from preprints in LLMs.arXiv preprint arXiv:2404.15799,

    Jialin Liu and Yi Bu. Towards the relationship be- tween AIGC in manuscript writing and author pro- files: evidence from preprints in LLMs.arXiv preprint arXiv:2404.15799,

  5. [15]

    Delving into PubMed records: Some terms in medical writing have drastically changed after the arrival of ChatGPT.MedRxiv, pages 2024–05,

    Kentaro Matsui. Delving into PubMed records: Some terms in medical writing have drastically changed after the arrival of ChatGPT.MedRxiv, pages 2024–05,

  6. [16]

    Writing without borders: AI and cross-cultural convergence in academic writing quality.Humanities and Social Sciences Communications 2025 12:1, 12:1058–, 7

    Arjun Prakash, Shruti Aggarwal, Jeevan John Varghese, and Joel John Varghese. Writing without borders: AI and cross-cultural convergence in academic writing quality.Humanities and Social Sciences Communications 2025 12:1, 12:1058–, 7

  7. [17]

    Does genai rewrite how we write? an empirical study on two-million preprints.arXiv preprint arXiv:2510.17882,

    Minfeng Qi, Zhongmin Cao, Qin Wang, Ningran Li, and Tianqing Zhu. Does genai rewrite how we write? an empirical study on two-million preprints.arXiv preprint arXiv:2510.17882,

  8. [19]

    URL https://www.wiley.com/en-de/about-us/ ai-resources/ai-study/. Wiley. ExplanAItions 2025-2026: The evolution of AI in re- search., 2

  9. [20]

    Hiromu Yakura, Ezequiel Lopez-Lopez, Levin Brinkmann, Ignacio Serna, Prateek Gupta, Ivan Soraperra, and Iyad Rahwan

    URL https://www.wiley.com/en-de/ about-us/ai-resources/ai-study/. Hiromu Yakura, Ezequiel Lopez-Lopez, Levin Brinkmann, Ignacio Serna, Prateek Gupta, Ivan Soraperra, and Iyad Rahwan. Empirical evidence of large language model’s influence on human spoken communication.arXiv pre...

  10. [21]

    LLM hallucinations in the wild: Large-scale evidence from non-existent citations

    Zhenyue Zhao, Yihe Wang, Toby Stuart, Mathijs De Vaan, Paul Ginsparg, and Yian Yin. LLM hallucinations in the wild: Large-scale evidence from non-existent citations. arXiv preprint arXiv:2605.07723,

  11. [22]

    9 Holzwarth et al.LLM-associated vocabulary in biomedical publications Supplementary T ables Period Section Dataset Estimate Method Ref. 12/2022–02/2023 abstract Various journals 0.10 Various LLM detectors 2024 08/2023 abstract arXiv, bioRxiv 0.13 LLM detector (custom) 2025 11...

  12. [2024]

    Delving into the utilisation of ChatGPT in scientific publications in astronomy.arXiv preprint arXiv:2406.17324,

    Simone Astarita, Sandor Kruk, Jan Reerink, and Pablo G´omez. Delving into the utilisation of ChatGPT in scientific publications in astronomy.arXiv preprint arXiv:2406.17324,

  13. [2025]

    Is ChatGPT trans- forming academics’ writing style?arXiv preprint arXiv:2404.08627,

    Mingmeng Geng and Roberto Trotta. Is ChatGPT trans- forming academics’ writing style?arXiv preprint arXiv:2404.08627,

  14. [2026]

    The epistemic downside of using LLM- based generative AI in academic writing.Publications 2025, 13:63, 12

    Bor Luen Tang. The epistemic downside of using LLM- based generative AI in academic writing.Publications 2025, 13:63, 12

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.