REVIEW 4 major objections 6 minor 2 references
Exploring the change in scientific readability following the release of ChatGPT
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that arXiv abstracts became steadily harder to read from 2010 to 2024, with a statistically significant extra decline after ChatGPT's release in November 2022.
desk verdict Useful descriptive measurement of arXiv readability decline and a post-2022 acceleration, but the causal framing toward AI outruns the identification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the battery of four readability formulas, each a fixed arithmetic function of word count, sentence count, characters, or syllables, applied to every abstract in the arXiv dataset. The analysis then aggregates these per-paper scores by year and by arXiv's eight primary categories, and tests for shifts using consecutive-year differences, percentage changes, rolling three-year standard deviations, word-presence comparisons, and paired t-tests on version pairs.
What would settle it
A controlled reanalysis that restricts to authors who posted abstracts both before and after November 2022, and that adjusts for abstract length and field mix, would settle the matter: if the 2023–2024 readability jump disappears under those controls, the aggregate effect is compositional rather than a change in writing style.
Extended reading notes
Core claim
Using four readability formulas—the Automated Readability Index, Coleman–Liau Index, Flesch Reading Ease, and Flesch–Kincaid Grade Level—the study reports that average readability scores of arXiv abstracts rose (became more complex) every year from 2010 to 2024, with high Pearson correlations (r near 0.98) and p-values below 0.001 for all four metrics. The largest consecutive-year jumps in scores occurred between 2022 and 2023 and again between 2023 and 2024. This pattern also appears when comparing older and updated versions of the same papers: abstracts revised after ChatGPT's release scored measurably harder to read than their pre-release versions, and abstracts containing ChatGPT-associated words such as 'pivotal,' 'intricate,' 'realm,' and 'showcasing' had more difficult readability scores than those without.
Load-bearing premise
The load-bearing premise is that year-to-year changes in average readability scores reflect changes in how authors write, rather than changes in who submits papers, which fields grow, or how long abstracts get.
Editorial extensions
If this is right
- Readability metrics can serve as a scalable, low-cost indicator of LLM influence on scientific text, complementing word-frequency analyses.
- The post-2022 readability jump appears across most arXiv categories, so the trend is not confined to a single field.
- Abstracts updated after ChatGPT's release tend to score as harder to read than their own earlier versions, pointing to AI-assisted revision as a plausible contributor.
- The steady 2010–2024 decline in readability is statistically significant under all four formulas, extending earlier findings of decreasing scientific-text readability.
Reading between the lines
- If the aggregate shift is driven by LLM-assisted writing, similar readability jumps should appear in peer-reviewed journals—possibly with a delay due to the slower publication cycle—and in other preprint repositories.
- The author leaves implicit that readability formulas respond to length and jargon; AI tools that produce longer, more formal abstracts could lower readability scores even without authors deliberately making text more difficult.
- The paper's four-word probe could be expanded into a larger lexicon of LLM-associated tokens, and combined with author metadata, to test whether the effect is concentrated among particular author groups or early-adopting disciplines.
- A direct comparison with non-English or non-STEM preprint servers would clarify whether the readability shift is specific to arXiv's community or a broader phenomenon of AI-assisted academic writing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes the readability of all arXiv abstracts posted between 2010 and June 7, 2024, using four standard formulas (ARI, CLI, FRE, FKRGL). It reports a steady annual decline in readability over the full period and identifies notable shifts in 2023 and 2024, which it interprets as likely influenced by ChatGPT. The analysis aggregates yearly averages, examines consecutive-year changes, uses a word-list comparison, and performs a version-update study on samples of papers with revisions. It also breaks down results by arXiv's eight primary categories for 2022–2024.
Significance. The study provides a large-scale, descriptive account of readability trends across a major preprint platform, with four independent metrics and a long observation window. The long-term decline in readability is strongly supported by the high Pearson correlations, and the category-level analyses show consistent directions of change. The paper is transparent about several limitations and uses external readability formulas rather than fitting parameters to its own data, which is a strength. Its main contribution is the descriptive trend; the causal or quasi-causal interpretation regarding ChatGPT is the weakest part of the manuscript.
major comments (4)
- [§4.2 and Appendix A] The central claim that readability 'likely' changed because of ChatGPT (Abstract) and that there were 'notable shifts occurring after November 2022' (Conclusion) is not identified against compositional changes. Appendix A reports a 19.96% submission surge in 2023, and the dataset shows a growing share of computer science papers, which tend to have higher readability scores. The category-level t-tests in Table 6 reduce but do not eliminate this concern because papers are double-counted across categories, within-category composition is not examined, and the unweighted category means do not reflect the population-weighted overall mean. Please add analyses that control for abstract length, primary subcategory, and submission volume, or explicitly limit the conclusions to a descriptive association.
- [§4.2, first two approaches] The claim that readability changed 'significantly' in 2023 and 2024 is not supported by any formal statistical test. Figures 2 and 3 show descriptive differences and percentage changes, but no confidence intervals or hypothesis tests compare the 2022–2023 and 2023–2024 changes against the distribution of earlier year-pair changes. Without such a test, the word 'significant' in the abstract and RQ2 is not justified. Please add a formal test (e.g., interrupted time series, permutation test, or a comparison of the observed jump against the historical distribution).
- [§4.2, fifth approach (Tables 3–4)] The version-update analysis, which is the strongest design because it holds authorship fixed, yields heterogeneous results: only one of the three Post_2022 samples shows significant changes for all four metrics, and the other two show significance only for CLI. The text's conclusion that 'all four measures were higher on average for the subset of papers updated after 2022' (Section 4.2) is not supported by the t-tests, which show frequent non-significance. Please report effect sizes with confidence intervals, pool the samples appropriately, and temper the wording to match the mixed evidence.
- [§3.2 and Figure 3] The publication-year assignment uses the latest version date, and the V1 robustness check only reassigns the year while still using the latest-version abstract. Consequently, papers originally posted in 2022 but updated in 2023 or 2024 are counted in the later years with their updated text, which could inflate the post-2022 jump even if the writing of new submissions did not change. Please separately report the version-update analysis for papers with and without revisions, or demonstrate that using the original-version abstracts with the first-version date yields the same trend.
minor comments (6)
- [Table 1] The Flesch Reading Ease formula is stated with a coefficient of 85.6 for syllables per word, but the standard Flesch formula uses 84.6. Please verify and correct the formula; this may also affect the reported FRE values if Textstat follows the standard.
- [§3.1] The ARI formula is attributed to Kincaid et al. (1975), but the Automated Readability Index was originally developed by Senter and Smith (1967). Please correct the citation for the ARI metric.
- [§3.1 and Appendix A] The text says the 2024 data include 'the first few days of June' in Section 3.1 and 'the first seven days of June' in Appendix A. Please use a consistent description of the cutoff date.
- [§4.2, third approach] The rolling standard deviation analysis is described only in prose; providing a table or figure of the rolling standard deviation values would make the claim about 2021–2023 and 2022–2024 being the highest more transparent.
- [§4.2, word-list approach] The word-list comparison is a useful auxiliary analysis, but the presence of 'pivotal,' 'intricate,' 'realm,' and 'showcasing' may be correlated with topic rather than with AI editing. Please acknowledge this limitation more explicitly when interpreting the larger readability differences for abstracts containing these words.
- [Table 4] The t-test results are reported with raw p-values; given that multiple tests are performed, please state whether any multiple-comparison correction was considered, or justify the unadjusted p-values.
Circularity Check
No circularity: readability trends are computed directly from external formulas and compared over time; no fitted parameter or self-cited uniqueness claim is load-bearing.
full rationale
The paper's central result is a direct aggregation of four standard readability formulas (ARI, CLI, FRE, FKRGL) applied to arXiv abstracts, with no parameters fitted to the dataset. The post-2022 shift is identified by comparing consecutive-year mean scores, percentage changes, rolling standard deviations, an external word list from prior work (Geng & Trotta, 2024), and a within-paper version-update comparison. None of these steps defines the outcome in terms of the hypothesis: the formulas are fixed and external, the word list comes from independent prior research, and the version-update analysis holds authorship fixed by comparing older and updated abstracts of the same papers. The two citations to the author's own work (Alsudais, 2021a, 2021b) appear only in a general related-work sentence and carry no argumentative weight. The abstract's phrase 'point to the likely influence of AI' is an interpretation, and the paper itself acknowledges other factors and confounding; that is a causal-inference limitation, not a circularity. There is no equation in which a prediction equals an input by construction, no fitted parameter renamed as a prediction, and no uniqueness theorem imported from the authors' prior work. Therefore the derivation chain is self-contained with respect to circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The four readability formulas (ARI, CLI, FRE, FKRGL) as implemented in Textstat produce valid, comparable scores for scientific abstracts.
- domain assumption Abstracts with fewer than 100 words can be excluded without biasing the yearly comparisons.
- domain assumption The upload date of the latest version approximates the date the abstract was written.
- domain assumption The words 'pivotal', 'intricate', 'realm', and 'showcasing' are valid markers of ChatGPT-influenced writing.
Cite this review
Pith. "Pith review of Exploring the change in scientific readability following the release of ChatGPT." pith.science (2026). https://pith.science/paper/3XHXH2HP
@misc{pith2026250621825,
author = {Pith},
title = {Pith review of: Exploring the change in scientific readability following the release of ChatGPT},
year = {2026},
howpublished = {\url{https://pith.science/paper/3XHXH2HP}},
note = {Machine review of arXiv:2506.21825}
}
read the original abstract
The rise and growing popularity of accessible large language models have raised questions about their impact on various aspects of life, including how scientists write and publish their research. The primary objective of this paper is to analyze a dataset consisting of all abstracts posted on arXiv.org between 2010 and June 7th, 2024, to assess the evolution of their readability and determine whether significant shifts occurred following the release of ChatGPT in November 2022. Four standard readability formulas are used to calculate individual readability scores for each paper, classifying their level of readability. These scores are then aggregated by year and across the eight primary categories covered by the platform. The results show a steady annual decrease in readability, suggesting that abstracts are likely becoming increasingly complex. Additionally, following the release of ChatGPT, a significant change in readability is observed for 2023 and the analyzed months of 2024. Similar trends are found across categories, with most experiencing a notable change in readability during 2023 and 2024. These findings offer insights into the broader changes in readability and point to the likely influence of AI on scientific writing.
Reference graph
Works this paper leans on
-
[79]
https://doi.org/10.1016/j.econlet.2019.02.017 Plavén-Sigray, P., Matheson, G. J., Schiffler, B. C., & Thompson, W. H. (2017). The readability of scientific texts is decreasing over time. Elife, 6, e27725. Porwal, P., & Devare, M. H. (2024). Scientific impact analysis: Unraveling the link between linguistic properties and citations. Journal of Informetrics...
arXiv 2017
-
[2020]
PACIS 2021 Proceedings, 200. Hamed, A. A., & Wu, X. (2024). Detection of ChatGPT fake science with the xFakeSci learning algorithm. Scientific Reports, 14(1), 16231. Horbach, S. P., Schneider, J. W., & Sainte -Marie, M. (2022). Ungendered writing: Writing styles are unlikely to account for gender differences in funding rates in the natural and technical s...
arXiv 2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.