REVIEW 4 cited by
ChatGPT "contamination": estimating the prevalence of LLMs in the scholarly literature
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The use of ChatGPT and similar Large Language Model (LLM) tools in scholarly communication and academic publishing has been widely discussed since they became easily accessible to a general audience in late 2022. This study uses keywords known to be disproportionately present in LLM-generated text to provide an overall estimate for the prevalence of LLM-assisted writing in the scholarly literature. For the publishing year 2023, it is found that several of those keywords show a distinctive and disproportionate increase in their prevalence, individually and in combination. It is estimated that at least 60,000 papers (slightly over 1% of all articles) were LLM-assisted, though this number could be extended and refined by analysis of other characteristics of the papers or by identification of further indicative keywords.
Forward citations
Cited by 4 Pith papers
-
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
People prefer text containing the words that an instruction-tuned model uses far more than its base version, linking human feedback training to LLM word overuse.
-
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
After ChatGPT's release, science and tech podcast speakers used AI-associated words like 'surpass' and 'align' more often, while control synonyms showed no average shift.
-
Exploring the Structure of AI-Induced Language Change in Scientific English
In PubMed abstracts, AI-associated 'spiking' words rise together with their synonyms rather than replacing them, and declining words show less systematic, more organic patterns.
-
GPT Editors, Not Authors: The Stylistic Footprint of LLMs in Academic Preprints
Across 2,408 arXiv preprints, LLM-typical word usage does not cluster in any section, indicating that AI assistance, when used, is uniform rather than limited to specific parts of a paper.
Discussion (0). Sign in to comment.