REVIEW 4 major objections 5 minor 18 references
Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read AI-assisted writing raises citation counts, but the gains are uneven: Global East authors adopted AI tools more aggressively, while Western authors captured more benefit per unit of adoption, leaving recognition gaps intact.
desk verdict Large corpus and a few genuinely new analyses, but the headline claims are not supported by the paper's own statistics. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is Binoculars, a zero-shot detector that scores texts by comparing perplexity between a strong 'performer' LLM and a weaker 'observer' LLM; lower scores mean more AI-like writing, and the paper treats this continuous score as its proxy for LLM adoption. Around that score, the analysis builds three pillars: zero-inflated negative binomial regressions for total and cross-hemispheric citations with GPT-era, hemisphere, and humanness interactions; an ordered-logit model for journal quartile; and a country-level citation network measured by PageRank, betweenness, conductance, K-core, and assortativity to test structural integration.
What would settle it
A ground-truth test: take matched sets of human-written and LLM-assisted abstracts from Eastern and Western authors, run Binoculars on them, and compare scores. If human-written Eastern abstracts score as AI-like as LLM-assisted Western ones, the regional results are detector artifacts. Alternatively, a leads-and-lags regression around ChatGPT's release would test whether citation behavior on humanness was actually stable across time.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that AI-assisted writing is a real but uneven currency in academic recognition. Using Binoculars scores as a continuous measure of human-likeness (lower scores mean more AI assistance), the authors find a global post-ChatGPT decline in human-likeness, steeper for the Global East; a citation model in which greater humanness predicts fewer citations; a Western advantage in citation returns per unit of AI-likeness; an ordered-logit result suggesting human-sounding texts are more likely to land in higher journal quartiles; and citation-network changes—more than doubled regional assortativity, modest Eastern PageRank gains, and only 3 of 25 perip
Load-bearing premise
The Binoculars score is a regionally unbiased measure of true LLM use in writing: if the detector flags non-Western or non-native English writing as 'AI-like' regardless of actual tool use, the regional adoption gap and the unequal citation returns both collapse.
Editorial extensions
If this is right
- Articles that read as more AI-like get more citations on average, so scholars face a structural incentive to use LLMs for visibility.
- Because Global East authors adopted AI writing more aggressively but gained less per unit, LLM assistance reduces surface language barriers without closing the recognition gap.
- Prestigious journals' preference for human-sounding text penalizes the authors who have leaned most into AI assistance, creating a trade-off between citation visibility and formal prestige.
- Post-ChatGPT citation networks become more regionally clustered rather than more integrated, so AI-style convergence does not translate into structural center-periphery change.
- Successive GPT releases—especially GPT-4o-mini—push detectable AI-likeness further, suggesting the stylistic shift is still accelerating.
Reading between the lines
- I infer that if detector bias against Global East writing is real, the paper's adoption gap is overstated and its per-unit return gap may be understated; the two regional findings would not both survive in their current magnitudes.
- The 'humanlike penalty' suggests an arms-race dynamic: as AI-fluent prose becomes the baseline, human-sounding writing may become a costly signal of either prestige or non-adoption, a mechanism the paper does not model.
- A testable extension: compare citation returns for AI-assisted writing in fields with explicit AI-use disclosure or bans, where the signal value of AI-likeness should shift.
- The network's eastward tilt is consistent with a long pre-ChatGPT trend in computer science, so attributing the modest rebalancing to LLM adoption is not supported by the paper's design.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes 238,218 Scopus-indexed computer science articles (2021–2025) using the Binoculars zero-shot detector to measure the AI-likeness of abstracts, then relates this measure to citation counts, cross-regional citation flows, journal tier, and citation-network structure before and after ChatGPT's release. The central claim is that authors in the Global East adopt AI tools more aggressively, yet Western authors gain more per unit of adoption because of pre-existing penalties for humanlike writing, and that prestigious journals continue to privilege human-sounding texts. The paper also presents a causal argument from ChatGPT release to citation outcomes through reduced human-likeness, and a descriptive network analysis showing modest Eastern visibility gains but little structural integration.
Significance. If the headline findings were empirically supported, the paper would make a useful contribution to the literature on AI, academic writing, and global inequalities in scholarly recognition. The paper has notable strengths: a large corpus, a zero-shot detector applied consistently, explicit regression tables in the appendix, and a descriptive network analysis whose scope is stated carefully. The authors also appropriately acknowledge some limitations, including detector bias and the observational nature of the network results. However, the central claims as stated in the abstract and introduction are not supported by the statistical evidence presented in the paper. The key three-way interaction is non-significant, the journal-prestige model is never estimated, and the causal premise of stable citation behavior is contradicted by the paper's own results. These are load-bearing issues for the paper's main message.
major comments (4)
- [§4.2, Table 7] The abstract's central assertion that 'Western authors gain more per unit of adoption' is directly tested by the three-way interaction Post-GPT×West×Binoculars in Equation 3. That coefficient is 1.150 with p=0.134, which is not statistically significant at conventional levels. The text in §4.2 correctly calls this 'suggestive rather than definitive,' but the abstract and introduction present it as an established finding. A non-significant coefficient cannot support the paper's headline claim.
- [§3.4, Eq. (4)] The paper claims in the abstract and discussion that 'prestigious journals continue to privilege more human-sounding texts,' but no results from the ordered-logit model specified in Equation 4 are reported anywhere in the text, tables, or appendix. The model is described but never estimated or displayed. This claim is therefore unverified, and the assertion should either be removed or supported with the actual model output.
- [§3.4, Figure 1; Table 7] The causal argument from ChatGPT release to citation outcomes rests on Premise 4, that 'citing behavior with respect to human-likeness remains approximately stable before and after ChatGPT's release.' However, Table 7 reports a statistically significant Post-GPT×Binoculars interaction (-0.721, p<0.01), and the text in §4.2 reports a similar significant effect (β=-0.660, p=0.031). This indicates that the relationship between human-likeness and citations changed after ChatGPT, directly contradicting the stability premise. The limitation section acknowledges this concern, but the paper still concludes that 'causality holds,' which is not warranted.
- [§3.2, §5.2] The measure of AI adoption—the Binoculars score—is the foundation of the regional comparison. The limitations section states that 'the majority of detection tools place a baseline penalty on Global East authorship,' but the paper provides no analysis showing whether, or to what extent, this penalty applies to Binoculars in this corpus. Without such validation, the finding that Global East authors adopt AI more aggressively could be an artifact of detector bias rather than actual usage. This is especially important because the regional adoption gap drives the subsequent citation-return analysis.
minor comments (5)
- [§3.4, Eq. (3)] Equation 3 is labeled a 'zero inflated binomial regression,' but the model described and estimated is a zero-inflated negative binomial. Please correct the terminology.
- [§4.1] Nonstandard regional labels appear, such as 'Asiatic Region' and 'Northern America.' Use 'Asia' and 'North America' for consistency with the SJR mapping.
- [Table 5] The table uses 'GPT4o+' as a variable name, while the text refers to 'GPT-4o-mini.' These labels should be unified to avoid confusion.
- [Figure 3] The caption refers to 'predicted binoculars centrality,' but the analysis is of the Binoculars score, not a centrality measure. This wording is misleading and should be clarified.
- [§3.1 / Abstract] The abstract says 'over 230,000' articles; the introduction says 'over 150,000.' These numbers are not inconsistent if they refer to different subsets, but the manuscript should state clearly which subset is used for which analysis.
Circularity Check
No significant circularity: the paper's analyses use external detectors and external citation-aging models; the sole self-citation (AlShebli et al. 2022) is peripheral context, not load-bearing.
full rationale
Walking the claimed derivation chain: (1) AI-likeness is measured with Binoculars (Hans et al. 2024), an externally developed zero-shot detector; the paper does not define AI adoption in terms of its own citation outcomes. (2) Adoption trends are estimated by regressing Binoculars score on time/region dummies (Eq. 1), and citation outcomes are modeled separately with a zero-inflated negative binomial using Binoculars as a moderator (Eq. 3). Using the same proxy as outcome and mediator is a standard observational design, not a reduction to inputs; the causal claim is explicitly framed via a quasi-instrument (ChatGPT release) with a stated stability assumption. (3) The exposure control is taken from Wang et al. (2013), an external bibliometric model, not fitted to the citation outcome in this paper. (4) The only reference to the authors' own prior work, AlShebli et al. (2022), is used to contextualize the eastward shift in AI research networks and is not load-bearing for the citation-return or journal-prestige claims. The acknowledged limitations (detector bias against Global East texts, unstable instrument, unsupported ordered-logit results) are validity and overclaiming concerns, not circularity. Therefore no step reduces, by construction or by self-citation, to its own inputs. Score reflects the peripheral self-citation and the paper's partial reliance on a single self-cited contextual result; no circular reasoning was found.
Assumptions & free parameters
free parameters (1)
- Exposure score log-normal parameters (mu, sigma) =
not reported; taken from Wang et al. 2013 model
assumptions (3)
- domain assumption Binoculars score is a valid proxy for actual LLM adoption in academic writing
- ad hoc to paper Citation behavior with respect to humanness is stable before and after ChatGPT (Premise 4)
- domain assumption Scopus relevance-based sampling and SJR matching are representative of global CS output
Cite this review
Pith. "Pith review of Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes." pith.science (2026). https://pith.science/paper/AO6SH36J
@misc{pith2026250908306,
author = {Pith},
title = {Pith review of: Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes},
year = {2026},
howpublished = {\url{https://pith.science/paper/AO6SH36J}},
note = {Machine review of arXiv:2509.08306}
}
read the original abstract
The rapid adoption of generative AI tools is reshaping how scholars produce and communicate knowledge, raising questions about who benefits and who is left behind. We analyze over 230,000 Scopus-indexed computer science articles between 2021 and 2025 to examine how AI-assisted writing alters scholarly visibility across regions. Using zero-shot detection of AI-likeness, we track stylistic changes in writing and link them to citation counts, journal placement, and global citation flows before and after ChatGPT. Our findings reveal uneven outcomes: authors in the Global East adopt AI tools more aggressively, yet Western authors gain more per unit of adoption due to pre-existing penalties for "humanlike" writing. Prestigious journals continue to privilege more human-sounding texts, creating tensions between visibility and gatekeeping. Network analyses show modest increases in Eastern visibility and tighter intra-regional clustering, but little structural integration overall. These results highlight how AI adoption reconfigures the labor of academic writing and reshapes opportunities for recognition.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Bedoor AlShebli, Enshu Cheng, Marcin Waniek, Ramesh Jagannathan, Pablo Hernández-Lagos, and Talal Rahwan. 2022. Beijing’s central role in global artificial intelligence research.Scientific Reports12, 1 (Dec. 2022), 21461. doi:10.1038/s41598-022-25714-0
-
[2]
Abdulrahman M. Al-Zahrani and. 2024. The impact of generative AI tools on researchers and research: Implications for academia in higher education. Innovations in Education and Teaching International61, 5 (2024), 1029–1043. https: //doi.org/10.1080/14703297.2023.2271445
arXiv 2024
-
[3]
S. A. Athaluri, S. V. Manthena, V. S. R. K. M. Kesapragada, V. Yarlagadda, T. Dave, and R. T. S. Duddumpudi. 2023. Exploring the Boundaries of Reality: Investigating the Phenomenon of Artificial Intelligence Hallucination in Scientific Writing Through ChatGPT References.Cureus15, 4 (2023), e37432. doi:10.7759/cureus. 37432
doi:10.7759/cureus 2023
-
[4]
Atanasov, Nan Liu, Yue Qiu, Tien Yin Wong, Yih-Chung Tham, and Yingfeng Zheng
Huzi Cheng, Bin Sheng, Aaron Lee, Varun Chaudary, Atanas G. Atanasov, Nan Liu, Yue Qiu, Tien Yin Wong, Yih-Chung Tham, and Yingfeng Zheng. 2024. Have AI-Generated Texts from LLM Infiltrated the Realm of Scientific Writing? A Large-Scale Analysis of Preprint Platforms.bioRxiv(2024). doi:10.1101/2024.03. 25.586710
-
[5]
Mingmeng Geng and Roberto Trotta. 2024. Is ChatGPT Transforming Academics’ Writing Style? arXiv:2404.08627 [cs.CL] https://arxiv.org/abs/2404.08627
arXiv 2024
-
[6]
Mingmeng Geng and Roberto Trotta. 2025. Human-LLM Coevolution: Evidence from Academic Writing. arXiv:2502.09606 [cs.CL] https://arxiv.org/abs/2502. 09606
arXiv 2025
-
[7]
David I. Hanauer, Cheryl L. Sheridan, and Karen Englander. 2019. Linguistic Injustice in the Writing of Research Articles in English as a Second Language: Data From Taiwanese and Mexican Researchers.Written Communication36, 1 (2019), 136–154. doi:10.1177/0741088318804821
-
[8]
Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2024. Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text. InInternational Conference on Machine Learning. PMLR, 17519–17537
work page 2024
Show all 18 references
-
[9]
Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D Manning, and James Y. Zou. 2024. Mapping the Increasing Use of LLMs in Scientific Papers. arXiv:2404.01268 ...
2024 arXiv
-
[10]
Dingkang Lin, Naixuan Zhao, Dan Tian, and Jiang Li. 2025. ChatGPT as Lin- guistic Equalizer? Quantifying LLM-Driven Lexical Shifts in Academic Writing. arXiv:2504.12317 [cs.CL] https://arxiv.org/abs/2504.12317
2025 arXiv
-
[11]
Chengzhi Mao, Carl Vondrick, Hao Wang, and Junfeng Yang. 2024. Raidar: geneRative AI Detection viA Rewriting. arXiv:2401.12970 [cs.CL] https://arxiv. Conference’17, July 2017, Washington, DC, USA Khan et al. org/abs/2401.12970
2024 arXiv
-
[12]
Shakked Noy and Whitney Zhang. 2023. Experimental evidence on the produc- tivity effects of generative artificial intelligence.Science381, 6654 (2023), 187–192. doi:10.1126/science.adh2586
2023 doi
-
[13]
Jovan Shopovski. 2024. Generative Artificial Intelligence, AI for Scientific Writ- ing: A Literature Review.Preprints(June 2024). https://doi.org/10.20944/ preprints202406.0011.v1
2024
-
[14]
Cuiping Song and Yanping Song. 2023. Enhancing academic writing skills and motivation: assessing the efficacy of ChatGPT in AI-assisted language learning for EFL students.Frontiers in PsychologyVolume 14 - 2023 (2023). doi:10.3389/ fpsyg.2023.1260843
2023
-
[15]
Vivek Verma, Eve Fleisig, Nicholas Tomlin, and Dan Klein. 2024. Ghostbuster: Detecting Text Ghostwritten by Large Language Models. arXiv:2305.15047 [cs.CL] https://arxiv.org/abs/2305.15047
2024 arXiv
-
[16]
Dashun Wang, Chaoming Song, and Albert-László Barabási. 2013. Quantifying Long-Term Scientific Impact.Science342, 6154 (2013), 127–132. doi:10.1126/ science.1237825
2013
-
[17]
Richard Watermeyer, Lawrie Phipps, Donna Lanclos, and Cathryn Knight. 2024. Generative AI and the Automating of Academia.Postdigital Science and Education 6, 2 (2024), 446–466. doi:10.1007/s42438-023-00440-6
2024 doi
-
[18]
Tatyana Yakhontova. 2020. English Writing of Non-Anglophone Researchers. Journal of Korean Medical Science35, 26 (June 2020). doi:10.3346/jkms.2020.35.e216 Variable Estimate (SE) Signif. Post-GPT−0.0038(0.0004) ∗∗∗ Hemisphere: West0.0017(0.0005) ∗∗∗ Post-GPT×West0.0014(0.0007)...
2020 doi
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.