Pith. sign in

REVIEW 4 major objections 5 minor 18 references

Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read AI-assisted writing raises citation counts, but the gains are uneven: Global East authors adopted AI tools more aggressively, while Western authors captured more benefit per unit of adoption, leaving recognition gaps intact.

desk verdict Large corpus and a few genuinely new analyses, but the headline claims are not supported by the paper's own statistics. read the letter →

arxiv 2509.08306 v1 pith:AO6SH36J submitted 2025-09-10 cs.CY

classification cs.CY
keywords LLMadoptionacademicwritingcitationoutcomesAIdetectionBinocularsGlobalEastscholarlyvisibilitycomputerscience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the arrival of ChatGPT has not made scholarly visibility fairer; it has changed the terms of visibility without dismantling old asymmetries. Analyzing more than 230,000 Scopus-indexed computer science articles from 2021 to 2025, the authors use a zero-shot AI-likeness detector to show that authors in the Global East shifted their writing toward AI norms more strongly after ChatGPT, while authors in the West retained higher baseline 'humanness.' Yet the citation system rewards AI-assisted style more generously for Western authors than for Eastern authors, because humanlike writing was already penalized unevenly, and prestigious journals continue to favor human-sounding prose. If correct, the paper shows that LLMs act less like an equalizer and more like a new layer of gatekeeping: they raise citation counts on average, tighten intra-regional citation clusters, and leave the core-periphery structure of global computer science largely intact.

What carries the argument

The load-bearing instrument is Binoculars, a zero-shot detector that scores texts by comparing perplexity between a strong 'performer' LLM and a weaker 'observer' LLM; lower scores mean more AI-like writing, and the paper treats this continuous score as its proxy for LLM adoption. Around that score, the analysis builds three pillars: zero-inflated negative binomial regressions for total and cross-hemispheric citations with GPT-era, hemisphere, and humanness interactions; an ordered-logit model for journal quartile; and a country-level citation network measured by PageRank, betweenness, conductance, K-core, and assortativity to test structural integration.

What would settle it

A ground-truth test: take matched sets of human-written and LLM-assisted abstracts from Eastern and Western authors, run Binoculars on them, and compare scores. If human-written Eastern abstracts score as AI-like as LLM-assisted Western ones, the regional results are detector artifacts. Alternatively, a leads-and-lags regression around ChatGPT's release would test whether citation behavior on humanness was actually stable across time.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that AI-assisted writing is a real but uneven currency in academic recognition. Using Binoculars scores as a continuous measure of human-likeness (lower scores mean more AI assistance), the authors find a global post-ChatGPT decline in human-likeness, steeper for the Global East; a citation model in which greater humanness predicts fewer citations; a Western advantage in citation returns per unit of AI-likeness; an ordered-logit result suggesting human-sounding texts are more likely to land in higher journal quartiles; and citation-network changes—more than doubled regional assortativity, modest Eastern PageRank gains, and only 3 of 25 perip

Load-bearing premise

The Binoculars score is a regionally unbiased measure of true LLM use in writing: if the detector flags non-Western or non-native English writing as 'AI-like' regardless of actual tool use, the regional adoption gap and the unequal citation returns both collapse.

Editorial extensions

If this is right

  • Articles that read as more AI-like get more citations on average, so scholars face a structural incentive to use LLMs for visibility.
  • Because Global East authors adopted AI writing more aggressively but gained less per unit, LLM assistance reduces surface language barriers without closing the recognition gap.
  • Prestigious journals' preference for human-sounding text penalizes the authors who have leaned most into AI assistance, creating a trade-off between citation visibility and formal prestige.
  • Post-ChatGPT citation networks become more regionally clustered rather than more integrated, so AI-style convergence does not translate into structural center-periphery change.
  • Successive GPT releases—especially GPT-4o-mini—push detectable AI-likeness further, suggesting the stylistic shift is still accelerating.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I infer that if detector bias against Global East writing is real, the paper's adoption gap is overstated and its per-unit return gap may be understated; the two regional findings would not both survive in their current magnitudes.
  • The 'humanlike penalty' suggests an arms-race dynamic: as AI-fluent prose becomes the baseline, human-sounding writing may become a costly signal of either prestige or non-adoption, a mechanism the paper does not model.
  • A testable extension: compare citation returns for AI-assisted writing in fields with explicit AI-use disclosure or bans, where the signal value of AI-likeness should shift.
  • The network's eastward tilt is consistent with a long pre-ChatGPT trend in computer science, so attributing the modest rebalancing to LLM adoption is not supported by the paper's design.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper analyzes 238,218 Scopus-indexed computer science articles (2021–2025) using the Binoculars zero-shot detector to measure the AI-likeness of abstracts, then relates this measure to citation counts, cross-regional citation flows, journal tier, and citation-network structure before and after ChatGPT's release. The central claim is that authors in the Global East adopt AI tools more aggressively, yet Western authors gain more per unit of adoption because of pre-existing penalties for humanlike writing, and that prestigious journals continue to privilege human-sounding texts. The paper also presents a causal argument from ChatGPT release to citation outcomes through reduced human-likeness, and a descriptive network analysis showing modest Eastern visibility gains but little structural integration.

Significance. If the headline findings were empirically supported, the paper would make a useful contribution to the literature on AI, academic writing, and global inequalities in scholarly recognition. The paper has notable strengths: a large corpus, a zero-shot detector applied consistently, explicit regression tables in the appendix, and a descriptive network analysis whose scope is stated carefully. The authors also appropriately acknowledge some limitations, including detector bias and the observational nature of the network results. However, the central claims as stated in the abstract and introduction are not supported by the statistical evidence presented in the paper. The key three-way interaction is non-significant, the journal-prestige model is never estimated, and the causal premise of stable citation behavior is contradicted by the paper's own results. These are load-bearing issues for the paper's main message.

major comments (4)
  1. [§4.2, Table 7] The abstract's central assertion that 'Western authors gain more per unit of adoption' is directly tested by the three-way interaction Post-GPT×West×Binoculars in Equation 3. That coefficient is 1.150 with p=0.134, which is not statistically significant at conventional levels. The text in §4.2 correctly calls this 'suggestive rather than definitive,' but the abstract and introduction present it as an established finding. A non-significant coefficient cannot support the paper's headline claim.
  2. [§3.4, Eq. (4)] The paper claims in the abstract and discussion that 'prestigious journals continue to privilege more human-sounding texts,' but no results from the ordered-logit model specified in Equation 4 are reported anywhere in the text, tables, or appendix. The model is described but never estimated or displayed. This claim is therefore unverified, and the assertion should either be removed or supported with the actual model output.
  3. [§3.4, Figure 1; Table 7] The causal argument from ChatGPT release to citation outcomes rests on Premise 4, that 'citing behavior with respect to human-likeness remains approximately stable before and after ChatGPT's release.' However, Table 7 reports a statistically significant Post-GPT×Binoculars interaction (-0.721, p<0.01), and the text in §4.2 reports a similar significant effect (β=-0.660, p=0.031). This indicates that the relationship between human-likeness and citations changed after ChatGPT, directly contradicting the stability premise. The limitation section acknowledges this concern, but the paper still concludes that 'causality holds,' which is not warranted.
  4. [§3.2, §5.2] The measure of AI adoption—the Binoculars score—is the foundation of the regional comparison. The limitations section states that 'the majority of detection tools place a baseline penalty on Global East authorship,' but the paper provides no analysis showing whether, or to what extent, this penalty applies to Binoculars in this corpus. Without such validation, the finding that Global East authors adopt AI more aggressively could be an artifact of detector bias rather than actual usage. This is especially important because the regional adoption gap drives the subsequent citation-return analysis.
minor comments (5)
  1. [§3.4, Eq. (3)] Equation 3 is labeled a 'zero inflated binomial regression,' but the model described and estimated is a zero-inflated negative binomial. Please correct the terminology.
  2. [§4.1] Nonstandard regional labels appear, such as 'Asiatic Region' and 'Northern America.' Use 'Asia' and 'North America' for consistency with the SJR mapping.
  3. [Table 5] The table uses 'GPT4o+' as a variable name, while the text refers to 'GPT-4o-mini.' These labels should be unified to avoid confusion.
  4. [Figure 3] The caption refers to 'predicted binoculars centrality,' but the analysis is of the Binoculars score, not a centrality measure. This wording is misleading and should be clarified.
  5. [§3.1 / Abstract] The abstract says 'over 230,000' articles; the introduction says 'over 150,000.' These numbers are not inconsistent if they refer to different subsets, but the manuscript should state clearly which subset is used for which analysis.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the paper's analyses use external detectors and external citation-aging models; the sole self-citation (AlShebli et al. 2022) is peripheral context, not load-bearing.

full rationale

Walking the claimed derivation chain: (1) AI-likeness is measured with Binoculars (Hans et al. 2024), an externally developed zero-shot detector; the paper does not define AI adoption in terms of its own citation outcomes. (2) Adoption trends are estimated by regressing Binoculars score on time/region dummies (Eq. 1), and citation outcomes are modeled separately with a zero-inflated negative binomial using Binoculars as a moderator (Eq. 3). Using the same proxy as outcome and mediator is a standard observational design, not a reduction to inputs; the causal claim is explicitly framed via a quasi-instrument (ChatGPT release) with a stated stability assumption. (3) The exposure control is taken from Wang et al. (2013), an external bibliometric model, not fitted to the citation outcome in this paper. (4) The only reference to the authors' own prior work, AlShebli et al. (2022), is used to contextualize the eastward shift in AI research networks and is not load-bearing for the citation-return or journal-prestige claims. The acknowledged limitations (detector bias against Global East texts, unstable instrument, unsupported ordered-logit results) are validity and overclaiming concerns, not circularity. Therefore no step reduces, by construction or by self-citation, to its own inputs. Score reflects the peripheral self-citation and the paper's partial reliance on a single self-cited contextual result; no circular reasoning was found.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central quantitative claims rest on three unverified premises: that Binoculars measures AI adoption (rather than stylistic otherness), that citation behavior toward humanness was stable across ChatGPT's release, and that the Scopus sample is representative. The exposure-score parameters are taken from Wang et al. without reporting fitted values. No new entities are introduced.

free parameters (1)
  • Exposure score log-normal parameters (mu, sigma) = not reported; taken from Wang et al. 2013 model
    The exposure control in Equation 3 requires mu and sigma for the aging function, but the values are not stated in the paper. The citation models cannot be fully reconstructed without them, and any arbitrary choice affects the exposure term.
assumptions (3)
  • domain assumption Binoculars score is a valid proxy for actual LLM adoption in academic writing
    Section 3.2 and throughout; the entire adoption analysis interprets lower Binoculars scores as higher AI assistance, but no ground-truth validation is provided and the authors acknowledge detector bias against Global East authorship in the Limitations.
  • ad hoc to paper Citation behavior with respect to humanness is stable before and after ChatGPT (Premise 4)
    Figure 1 and Section 4.2; this assumption is required for the causal chain, but the paper's own ZINB model finds a significant PostGPT x Binoculars interaction for raw citations (p=0.031), contradicting the premise.
  • domain assumption Scopus relevance-based sampling and SJR matching are representative of global CS output
    Section 3.1 and Limitations; the authors admit the relevance metric may introduce systemic biases and the pagination limit restricts monthly samples to 5000, which affects the representativeness of the regression sample.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes." pith.science (2026). https://pith.science/paper/AO6SH36J

@misc{pith2026250908306,
  author       = {Pith},
  title        = {Pith review of: Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AO6SH36J}},
  note         = {Machine review of arXiv:2509.08306}
}
read the original abstract

The rapid adoption of generative AI tools is reshaping how scholars produce and communicate knowledge, raising questions about who benefits and who is left behind. We analyze over 230,000 Scopus-indexed computer science articles between 2021 and 2025 to examine how AI-assisted writing alters scholarly visibility across regions. Using zero-shot detection of AI-likeness, we track stylistic changes in writing and link them to citation counts, journal placement, and global citation flows before and after ChatGPT. Our findings reveal uneven outcomes: authors in the Global East adopt AI tools more aggressively, yet Western authors gain more per unit of adoption due to pre-existing penalties for "humanlike" writing. Prestigious journals continue to privilege more human-sounding texts, creating tensions between visibility and gatekeeping. Network analyses show modest increases in Eastern visibility and tighter intra-regional clustering, but little structural integration overall. These results highlight how AI adoption reconfigures the labor of academic writing and reshapes opportunities for recognition.

Figures

Figures reproduced from arXiv: 2509.08306 by the authors.

Figure 1
Figure 1. A Diagrammatic Depiction of the Path to Causality [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Evolution of Binoculars score over time on an inverted y-axis, separated by region. The trends suggest a global uptick [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Interaction Plot: ChatGPT Effects by Region. Predicted “binoculars” centrality (left) and probability of a correct [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: GPT-Wave Effects Relative to Pre-GPT Baseline. Point estimates and 95% confidence intervals for the four GPT [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Chord diagrams depicting citation exchanges among the 20 most prolific countries in computer science research [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

18 extracted references · 9 canonical work pages

  1. [1]

    Bedoor AlShebli, Enshu Cheng, Marcin Waniek, Ramesh Jagannathan, Pablo Hernández-Lagos, and Talal Rahwan. 2022. Beijing’s central role in global artificial intelligence research.Scientific Reports12, 1 (Dec. 2022), 21461. doi:10.1038/s41598-022-25714-0

  2. [2]

    Al-Zahrani and

    Abdulrahman M. Al-Zahrani and. 2024. The impact of generative AI tools on researchers and research: Implications for academia in higher education. Innovations in Education and Teaching International61, 5 (2024), 1029–1043. https: //doi.org/10.1080/14703297.2023.2271445

  3. [3]

    S. A. Athaluri, S. V. Manthena, V. S. R. K. M. Kesapragada, V. Yarlagadda, T. Dave, and R. T. S. Duddumpudi. 2023. Exploring the Boundaries of Reality: Investigating the Phenomenon of Artificial Intelligence Hallucination in Scientific Writing Through ChatGPT References.Cureus15, 4 (2023), e37432. doi:10.7759/cureus. 37432

  4. [4]

    Atanasov, Nan Liu, Yue Qiu, Tien Yin Wong, Yih-Chung Tham, and Yingfeng Zheng

    Huzi Cheng, Bin Sheng, Aaron Lee, Varun Chaudary, Atanas G. Atanasov, Nan Liu, Yue Qiu, Tien Yin Wong, Yih-Chung Tham, and Yingfeng Zheng. 2024. Have AI-Generated Texts from LLM Infiltrated the Realm of Scientific Writing? A Large-Scale Analysis of Preprint Platforms.bioRxiv(2024). doi:10.1101/2024.03. 25.586710

  5. [5]

    Mingmeng Geng and Roberto Trotta. 2024. Is ChatGPT Transforming Academics’ Writing Style? arXiv:2404.08627 [cs.CL] https://arxiv.org/abs/2404.08627

  6. [6]

    Mingmeng Geng and Roberto Trotta. 2025. Human-LLM Coevolution: Evidence from Academic Writing. arXiv:2502.09606 [cs.CL] https://arxiv.org/abs/2502. 09606

  7. [7]

    Hanauer, Cheryl L

    David I. Hanauer, Cheryl L. Sheridan, and Karen Englander. 2019. Linguistic Injustice in the Writing of Research Articles in English as a Second Language: Data From Taiwanese and Mexican Researchers.Written Communication36, 1 (2019), 136–154. doi:10.1177/0741088318804821

  8. [8]

    Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2024. Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text. InInternational Conference on Machine Learning. PMLR, 17519–17537

Show all 18 references
  1. [9]

    Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D Manning, and James Y. Zou. 2024. Mapping the Increasing Use of LLMs in Scientific Papers. arXiv:2404.01268 ...

  2. [10]

    Dingkang Lin, Naixuan Zhao, Dan Tian, and Jiang Li. 2025. ChatGPT as Lin- guistic Equalizer? Quantifying LLM-Driven Lexical Shifts in Academic Writing. arXiv:2504.12317 [cs.CL] https://arxiv.org/abs/2504.12317

  3. [11]

    Chengzhi Mao, Carl Vondrick, Hao Wang, and Junfeng Yang. 2024. Raidar: geneRative AI Detection viA Rewriting. arXiv:2401.12970 [cs.CL] https://arxiv. Conference’17, July 2017, Washington, DC, USA Khan et al. org/abs/2401.12970

  4. [12]

    Shakked Noy and Whitney Zhang. 2023. Experimental evidence on the produc- tivity effects of generative artificial intelligence.Science381, 6654 (2023), 187–192. doi:10.1126/science.adh2586

  5. [13]

    Jovan Shopovski. 2024. Generative Artificial Intelligence, AI for Scientific Writ- ing: A Literature Review.Preprints(June 2024). https://doi.org/10.20944/ preprints202406.0011.v1

  6. [14]

    Cuiping Song and Yanping Song. 2023. Enhancing academic writing skills and motivation: assessing the efficacy of ChatGPT in AI-assisted language learning for EFL students.Frontiers in PsychologyVolume 14 - 2023 (2023). doi:10.3389/ fpsyg.2023.1260843

  7. [15]

    Vivek Verma, Eve Fleisig, Nicholas Tomlin, and Dan Klein. 2024. Ghostbuster: Detecting Text Ghostwritten by Large Language Models. arXiv:2305.15047 [cs.CL] https://arxiv.org/abs/2305.15047

  8. [16]

    Dashun Wang, Chaoming Song, and Albert-László Barabási. 2013. Quantifying Long-Term Scientific Impact.Science342, 6154 (2013), 127–132. doi:10.1126/ science.1237825

  9. [17]

    Richard Watermeyer, Lawrie Phipps, Donna Lanclos, and Cathryn Knight. 2024. Generative AI and the Automating of Academia.Postdigital Science and Education 6, 2 (2024), 446–466. doi:10.1007/s42438-023-00440-6

  10. [18]

    Tatyana Yakhontova. 2020. English Writing of Non-Anglophone Researchers. Journal of Korean Medical Science35, 26 (June 2020). doi:10.3346/jkms.2020.35.e216 Variable Estimate (SE) Signif. Post-GPT−0.0038(0.0004) ∗∗∗ Hemisphere: West0.0017(0.0005) ∗∗∗ Post-GPT×West0.0014(0.0007)...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.