Pith. sign in

REVIEW 15 cited by

Mapping the Increasing Use of LLMs in Scientific Papers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01268 v1 pith:G5ZALOF7 submitted 2024-04-01 cs.CL cs.AIcs.DLcs.LGcs.SI

classification cs.CLcs.AIcs.DLcs.LGcs.SI
keywords scientificllmsfindingsacademicanalysisfirstlevelmeasure
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scientific publishing lays the foundation of science by disseminating research findings, fostering collaboration, encouraging reproducibility, and ensuring that scientific knowledge is accessible, verifiable, and built upon over time. Recently, there has been immense speculation about how many people are using large language models (LLMs) like ChatGPT in their academic writing, and to what extent this tool might have an effect on global scientific practices. However, we lack a precise measure of the proportion of academic writing substantially modified or produced by LLMs. To address this gap, we conduct the first systematic, large-scale analysis across 950,965 papers published between January 2020 and February 2024 on the arXiv, bioRxiv, and Nature portfolio journals, using a population-level statistical framework to measure the prevalence of LLM-modified content over time. Our statistical estimation operates on the corpus level and is more robust than inference on individual instances. Our findings reveal a steady increase in LLM usage, with the largest and fastest growth observed in Computer Science papers (up to 17.5%). In comparison, Mathematics papers and the Nature portfolio showed the least LLM modification (up to 6.3%). Moreover, at an aggregate level, our analysis reveals that higher levels of LLM-modification are associated with papers whose first authors post preprints more frequently, papers in more crowded research areas, and papers of shorter lengths. Our findings suggests that LLMs are being broadly used in scientific writings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 39 citations worldwide. Full citation record

  1. Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences

    cs.DL 2026-07 conditional novelty 7.0 of 10

    Under a strict identity-level definition, roughly one in twenty 2025 NeurIPS and USENIX Security papers carries at least two likely hallucinated academic citations that survived peer review.

  2. Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback

    cs.CL 2025-08 conditional novelty 6.0 of 10

    People prefer text containing the words that an instruction-tuned model uses far more than its base version, linking human feedback training to LLM word overuse.

  3. Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English

    cs.CL 2025-08 conditional novelty 6.0 of 10

    After ChatGPT's release, science and tech podcast speakers used AI-associated words like 'surpass' and 'align' more often, while control synonyms showed no average shift.

  4. Exploring the Structure of AI-Induced Language Change in Scientific English

    cs.CL 2025-06 conditional novelty 6.0 of 10

    In PubMed abstracts, AI-associated 'spiking' words rise together with their synonyms rather than replacing them, and declining words show less systematic, more organic patterns.

  5. ScienceMeter: Tracking Scientific Knowledge Updates in Language Models

    cs.CL 2025-05 reject novelty 6.0 of 10

    ScienceMeter evaluates language model knowledge updates across three axes, preservation of old scientific claims, acquisition of new claims, and projection to future findings, and finds all current methods fall short.

  6. The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Arabic text written by LLMs carries detectable stylometric signatures, and fine-tuned XLM-RoBERTa detectors reach near-perfect F1 on academic abstracts but degrade on social media.

  7. Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Reinforcement learning on questions extracted from CRISPR expert forums improves LLM accuracy on a new benchmark (Genome-Bench) by over 15 percentage points.

  8. The Influence of Fraudulent AI-Generated Responses on Software Engineering Surveys

    cs.SE 2026-07 conditional novelty 5.0 of 10

    Suspicious or AI-assisted survey answers leave most SE quantitative distributions stable but materially change qualitative framing, code prominence, and interpretive evidence.

  9. BRoverbs -- Measuring how much LLMs understand Portuguese proverbs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    BRoverbs lets researchers test whether language models understand Portuguese proverbs; commercial models nearly master it, small models often guess randomly.

  10. Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes

    cs.CY 2025-09 reject novelty 5.0 of 10

    Analyzing 98,000 Scopus computer science papers, the paper finds a global rise in AI-like writing after ChatGPT and reports regional differences in citation returns, but the key differential-gain result is statistical...

  11. Low-Perplexity LLM-Generated Sequences and Where To Find Them

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Only about 40% of low-perplexity 6-token spans generated by Pythia-6.9B can be exactly matched to The Pile, and the authors categorize matched and unmatched spans into four classes.

  12. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  13. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 unverdicted novelty 3.0 of 10

    LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.

  14. Emphasizing Deliberation and Critical Thinking in an AI Hype World

    cs.HC 2025-07 unverdicted novelty 3.0 of 10

    An HCI researcher argues that resisting AI solutionism requires slow, deliberate use and critical thinking rather than outright bans.

  15. What Shapes Writers' Decisions to Disclose AI Use?

    cs.HC 2025-05 conditional novelty 3.0 of 10

    A literature synthesis identifies 12 procedural, social, and personal factors that may shape writers' decisions to disclose AI use.

Pith tools