REVIEW 15 cited by
Mapping the Increasing Use of LLMs in Scientific Papers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Scientific publishing lays the foundation of science by disseminating research findings, fostering collaboration, encouraging reproducibility, and ensuring that scientific knowledge is accessible, verifiable, and built upon over time. Recently, there has been immense speculation about how many people are using large language models (LLMs) like ChatGPT in their academic writing, and to what extent this tool might have an effect on global scientific practices. However, we lack a precise measure of the proportion of academic writing substantially modified or produced by LLMs. To address this gap, we conduct the first systematic, large-scale analysis across 950,965 papers published between January 2020 and February 2024 on the arXiv, bioRxiv, and Nature portfolio journals, using a population-level statistical framework to measure the prevalence of LLM-modified content over time. Our statistical estimation operates on the corpus level and is more robust than inference on individual instances. Our findings reveal a steady increase in LLM usage, with the largest and fastest growth observed in Computer Science papers (up to 17.5%). In comparison, Mathematics papers and the Nature portfolio showed the least LLM modification (up to 6.3%). Moreover, at an aggregate level, our analysis reveals that higher levels of LLM-modification are associated with papers whose first authors post preprints more frequently, papers in more crowded research areas, and papers of shorter lengths. Our findings suggests that LLMs are being broadly used in scientific writings.
Forward citations
Cited by 15 Pith papers
-
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
Under a strict identity-level definition, roughly one in twenty 2025 NeurIPS and USENIX Security papers carries at least two likely hallucinated academic citations that survived peer review.
-
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
People prefer text containing the words that an instruction-tuned model uses far more than its base version, linking human feedback training to LLM word overuse.
-
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
After ChatGPT's release, science and tech podcast speakers used AI-associated words like 'surpass' and 'align' more often, while control synonyms showed no average shift.
-
Exploring the Structure of AI-Induced Language Change in Scientific English
In PubMed abstracts, AI-associated 'spiking' words rise together with their synonyms rather than replacing them, and declining words show less systematic, more organic patterns.
-
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
ScienceMeter evaluates language model knowledge updates across three axes, preservation of old scientific claims, acquisition of new claims, and projection to future findings, and finds all current methods fall short.
-
The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text
Arabic text written by LLMs carries detectable stylometric signatures, and fine-tuned XLM-RoBERTa detectors reach near-perfect F1 on academic abstracts but degrade on social media.
-
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Reinforcement learning on questions extracted from CRISPR expert forums improves LLM accuracy on a new benchmark (Genome-Bench) by over 15 percentage points.
-
The Influence of Fraudulent AI-Generated Responses on Software Engineering Surveys
Suspicious or AI-assisted survey answers leave most SE quantitative distributions stable but materially change qualitative framing, code prominence, and interpretive evidence.
-
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
BRoverbs lets researchers test whether language models understand Portuguese proverbs; commercial models nearly master it, small models often guess randomly.
-
Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes
Analyzing 98,000 Scopus computer science papers, the paper finds a global rise in AI-like writing after ChatGPT and reports regional differences in citation returns, but the key differential-gain result is statistical...
-
Low-Perplexity LLM-Generated Sequences and Where To Find Them
Only about 40% of low-perplexity 6-token spans generated by Pythia-6.9B can be exactly matched to The Pile, and the authors categorize matched and unmatched spans into four classes.
-
AI for Auto-Research: Roadmap & User Guide
The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.
-
From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines
LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.
-
Emphasizing Deliberation and Critical Thinking in an AI Hype World
An HCI researcher argues that resisting AI solutionism requires slow, deliberate use and critical thinking rather than outright bans.
-
What Shapes Writers' Decisions to Disclose AI Use?
A literature synthesis identifies 12 procedural, social, and personal factors that may shape writers' decisions to disclose AI use.
Discussion (0). Sign in to comment.