Pith. sign in

REVIEW 8 cited by

Delving into LLM-assisted writing in biomedical publications through excess vocabulary

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07016 v5 pith:4MJ2PG6U submitted 2024-06-11 cs.CL cs.AIcs.CYcs.DLcs.SI

classification cs.CLcs.AIcs.CYcs.DLcs.SI
keywords biomedicalllmswritingabstractsexcessmodelsresearchvocabulary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) like ChatGPT can generate and revise text with human-level performance. These models come with clear limitations: they can produce inaccurate information, reinforce existing biases, and be easily misused. Yet, many scientists use them for their scholarly writing. But how wide-spread is such LLM usage in the academic literature? To answer this question for the field of biomedical research, we present an unbiased, large-scale approach: we study vocabulary changes in over 15 million biomedical abstracts from 2010--2024 indexed by PubMed, and show how the appearance of LLMs led to an abrupt increase in the frequency of certain style words. This excess word analysis suggests that at least 13.5% of 2024 abstracts were processed with LLMs. This lower bound differed across disciplines, countries, and journals, reaching 40% for some subcorpora. We show that LLMs have had an unprecedented impact on scientific writing in biomedical research, surpassing the effect of major world events such as the Covid pandemic.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 35 citations worldwide. Full citation record

  1. Using Large Language Models for Idea Generation in Innovation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    GPT-4-generated product ideas had higher average purchase intent than student ideas and made up 35 of the top 40 ideas, while being rated less novel and more similar to each other.

  2. Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it

    cs.CL 2026-07 conditional novelty 6.0 of 10

    LLMs overuse the 'not X, but Y' self-correction pattern in persuasive registers and underuse it in informal Q&A; a prompt or a detachable LoRA dial adjusts it to human levels.

  3. Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback

    cs.CL 2025-08 conditional novelty 6.0 of 10

    People prefer text containing the words that an instruction-tuned model uses far more than its base version, linking human feedback training to LLM word overuse.

  4. Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English

    cs.CL 2025-08 conditional novelty 6.0 of 10

    After ChatGPT's release, science and tech podcast speakers used AI-associated words like 'surpass' and 'align' more often, while control synonyms showed no average shift.

  5. Exploring the change in scientific readability following the release of ChatGPT

    cs.CY 2025-06 conditional novelty 6.0 of 10

    arXiv abstracts became less readable between 2010 and 2024, with a statistically significant jump in complexity in 2023 and 2024 following the release of ChatGPT.

  6. Exploring the Structure of AI-Induced Language Change in Scientific English

    cs.CL 2025-06 conditional novelty 6.0 of 10

    In PubMed abstracts, AI-associated 'spiking' words rise together with their synonyms rather than replacing them, and declining words show less systematic, more organic patterns.

  7. GPT Editors, Not Authors: The Stylistic Footprint of LLMs in Academic Preprints

    cs.CL 2025-05 reject novelty 5.0 of 10

    Across 2,408 arXiv preprints, LLM-typical word usage does not cluster in any section, indicating that AI assistance, when used, is uniform rather than limited to specific parts of a paper.

  8. What Shapes Writers' Decisions to Disclose AI Use?

    cs.HC 2025-05 conditional novelty 3.0 of 10

    A literature synthesis identifies 12 procedural, social, and personal factors that may shape writers' decisions to disclose AI use.

Pith tools