Pith. sign in

REVIEW 4 major objections 5 minor 62 references

"You Cannot Sound Like GPT": Signs of language discrimination and resistance in computer science publishing

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Peer reviewers penalize authors from non-English-speaking countries, and ChatGPT only changes the tell.

desk verdict The interviews are the real contribution; the regression is honest but the abstract oversells a causal reading of 'bias.' read the letter →

arxiv 2505.08127 v1 pith:NPS3ED66 submitted 2025-05-12 cs.CY

classification cs.CY
keywords languageideologiespeerreviewChatGPTdiscriminationsemioticsindexicalityICLRlinguisticbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that ChatGPT does not solve linguistic bias in scientific peer review; it changes the surface on which the bias is read. Analyzing 76,453 reviews of 20,827 submissions to the ICLR conference (2018–2024), the authors find that manuscripts with more authors from countries where English is less widely spoken receive significantly more critiques of writing clarity, fewer praises of clarity, and lower overall scores, even after other epistemic critiques are statistically held constant. After ChatGPT's release in November 2022, the pattern only weakened modestly. Interviews with 14 multilingual ICLR participants explain why: reviewers shifted from grammar-based cues to 'ChatGPT style' and non-linguistic cues—long author lists, acknowledgment names, perceived experimental volume—as signs that the author is not a 'native' English speaker, and they still tie that perceived identity to the quality of the science. The paper's wager is that language ideologies are durable: when one sign disappears, readers recruit another.

What carries the argument

The load-bearing mechanism is the indexical sign: a concrete textual feature—a missing plural, a long sentence, an LLM-typical phrase, a long author list—stands to a reviewer for an imagined type of person, and that imagined person stands in turn for an evaluation of the science. This two-step chain explains both the quantitative results and their persistence: removing the first sign with ChatGPT does not break the chain, because reviewers recruit new signs, while authors engage in 'language labor' to strip any feature that could expose their language background. Methodologically, the chain is operationalized with sentence-level classifiers that count clarity critiques and praises from review text, and with panel regressions comparing pre- and post-ChatGPT evaluation patterns.

What would settle it

A controlled experiment would settle it: send the same manuscript text to a large pool of reviewers with only authorship cues varied (affiliation country, author names, or acknowledgment names) and measure clarity critiques and scores; if the cues do not move evaluations, the central bias claim would be refuted. Alternatively, if an independent, origin-blind linguistic rating of manuscript clarity fully accounted for the regional gap in clarity critiques, the discrimination interpretation would collapse.

Watch

Extended reading notes

Core claim

The central claim is that clarity judgment in peer review is not an objective measure of readability but a socially loaded index of the imagined author. Across the full ICLR corpus, a higher percentage of authors affiliated with Asian, Chinese, or English-proficiency-tested countries is associated with more clarity critiques, less clarity praise, and lower ratings; these associations are significant before and, in attenuated form, after ChatGPT, and for most groups the pre/post shift is not statistically significant. The qualitative arm supplies the mechanism: grammatical idiosyncrasies (tense and plural errors, long clunky sentences) index a 'non-native speaker,' and that imagined speaker indexes lesser science; when ChatGPT removes the grammatical signs, reviewers report reading AI style, word choice, jargon density, author count, and acknowledgment names as new signs of origin. The title quote—'you cannot sound like GPT'—captures the trap: authors who polish with LLMs are heard as non-native in a new way, because AI style itself is now read as the accent of the language-marginalized. The paper presents this indexical chain as the reason availability of LLMs produces only a muted statistical shift, and as evidence that linguistic exclusion is reproduced rather than dissolved by the technology.

Load-bearing premise

The inference that regional differences in clarity critiques reflect reviewer discrimination rather than real differences in writing quality assumes that, after controlling for the other epistemic content of reviews, no unmeasured differences in manuscript clarity or quality remain—yet the dataset contains no independent measure of either author first language or objective writing clarity.

Editorial extensions

If this is right

  • AI-assisted polishing can narrow but not eliminate the regional gap in clarity critiques, because new non-grammatical signs of author origin take the place of grammatical ones.
  • Language-marginalized authors face a shifting 'language labor' tax: in addition to English, they must learn the current conventions for not sounding like AI, since AI style is itself read as non-native.
  • Reviewers' equation of 'good English' with 'good science' means clarity critiques double as status judgments, producing a durable inequality that is only partially masked by anonymous review.
  • The paper's own recommendations follow: more multilingual publishing venues, conferences run with local languages, and language coursework for graduate programs, rather than reliance on author-side LLM polishing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if the indexical-shift logic is right, any widely adopted writing or translation tool will eventually acquire its own stylistic signature, and that signature will become a demographic marker wherever readers judge clarity while caring about author origin.
  • Editorial inference: a pre-registered experiment—identical text presented with different author-origin cues (affiliation, names, acknowledgments)—would directly test the mechanism; the paper's framework predicts clarity ratings and trust in the science should move with the cues.
  • Editorial inference: the findings imply that evaluation design could weaken the chain more effectively than author-side tools: separating grammatical correctness from argumentative clarity in rubrics, or having trained editors rather than reviewers judge prose, would give the bias fewer signs to attach to.
  • Editorial inference: because even interviewees described feeling 'damned if you do, damned if you don't,' the study suggests that uniform AI adoption may create a new hierarchy between writers who can mimic a specific elite register and those who cannot, rather than flattening linguistic difference.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper studies linguistic bias in peer review at ICLR. It combines a statistical analysis of roughly 76,000 reviews from 2018–2024 with interviews of 14 multilingual scholars and one area chair. The quantitative part regresses clarity praise/critique and review ratings on the percentage of authors affiliated with institutions in Asian, Chinese, or TOEFL-required countries, controlling for reviewer-labeled epistemic content, review length, and manuscript random effects, and compares pre- and post-ChatGPT periods. The qualitative part documents indexical reasoning by which reviewers associate writing features with author demographics and with science quality, and describes the shift toward detecting 'ChatGPT style' as a new index of non-nativeness. The paper claims significant bias against authors from countries where English is less widely spoken, a muted effect of ChatGPT availability, and argues that ChatGPT has not broken the link between writing features, author origin, and perceived science quality.

Significance. If the central claim holds, the paper makes a valuable contribution to scholarship on linguistic disadvantage in scientific publishing by offering large-scale quantitative evidence and a theoretically grounded interview study of indexicality in the GPT era. The strengths are real: the code is publicly available, the manuscript honestly reports overlapping confidence intervals and the non-causal nature of the regression design, the qualitative methods are clearly described, and the interviews provide rich, credible evidence that reviewers and authors actively engage in indexical reasoning about language background. The theoretical framing through Peircean semiotics and raciolinguistic ideologies is novel for this empirical setting. The main weakness is that the quantitative analysis cannot adjudicate between reviewer bias and actual regional differences in writing quality, so the causal framing in the abstract and conclusion is not supported by the regression evidence alone.

major comments (4)
  1. [Abstract; Section 4.2; Table 4] The central claim, stated in the abstract and conclusion, that 'reviewers critique paper clarity significantly more' for authors from TOEFL-required countries is a causal inference, but the regression in Section 4.2 includes no independent measure of manuscript writing quality and no author-level first-language data. Section 4.1.1 itself calls the country-of-institution proxy 'a deeply inexact match' for language background, and Section 5.1 explicitly states 'our evaluation is not causal.' The coefficients in Table 4 would arise equally if papers from those regions are, on average, less clear in ways that reviewers correctly perceive. This is the load-bearing inference of the paper, and the abstract and conclusion must be revised to present the quantitative results as descriptive associations, with the causal interpretation reserved for the qualitative findings.
  2. [Section 4.2; Section 5.1] The controls for 'other epistemic value judgments' are derived from the same review text that produces the dependent variable. If bias affects all dimensions of evaluation, these are 'bad controls' that may absorb part of the very bias under study; moreover, they do not measure manuscript clarity independently. Therefore the assertion in Section 5.1 that regional differences persist 'even after statistically accounting for the other substantive content of the reviews' overstates the degree of confounding control. The paper should explicitly discuss this limitation and soften the corresponding claim.
  3. [Section 4.3; Sections 5.3–5.4] The interview evidence convincingly demonstrates that language ideologies and indexical reasoning exist among reviewers and authors, and that ChatGPT-style text is now read as a sign of demographic background. However, interviews describe perceptions and self-reported behaviors, not measured reviewer behavior on actual submissions. They therefore do not establish that the regional gaps in clarity critiques observed in Table 4 are caused by this bias. The conclusion currently conflates the descriptive quantitative association with the qualitative mechanism, and the paper should keep these two forms of evidence clearly separated in its final claims.
  4. [Section 4.1.3; Section 5.1; Figure 4] The post-GPT indicator is a conference-year split rather than a measure of which papers used LLMs, and the 'muted shift' is inferred from point estimates whose confidence intervals overlap across periods. The abstract's 'only a muted shift' claim is therefore fragile: the data are equally consistent with no change in the bias, with a compositional change in the author pool, or with ChatGPT simply improving the average clarity of submissions. The authors should present the shift as suggestive and give explicit attention to these alternative explanations, not just the indexicality account.
minor comments (5)
  1. [Section 1, paragraph 2] The phrase 'to assist users in to acquiring cultural capital' contains a typo: 'in to' should be 'in'.
  2. [Figure 4 caption] The caption reads 'Clarity ( )' for the third panel; the minus sign appears to be missing, and the error bars are not defined in the caption.
  3. [Appendix A.1.1] The sentence 'Every submission receives a textual review and numerical score score by each reviewer' duplicates 'score'; the duplicate should be removed.
  4. [Section 5.1] The phrase 'surprised by the muted improvements to equity of clarity critique and praise' is awkward and overstates the evidence; it would be more precise to say 'surprised by the muted improvement in equity of clarity critique and praise.'
  5. [References] Reference [38] is listed as 'Under review (2024)' without a venue; if the paper has been published, the citation should be updated.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the quantitative associations are estimated, not derived by construction, and the self-citations are contextual rather than load-bearing.

full rationale

The paper's central quantitative claim is an empirical association estimated from ICLR reviews: panel OLS regressions relate reviewer-labeled clarity critiques, clarity praise, and ratings to the percentage of authors from Asia, China, or TOEFL-required countries, with controls for other epistemic categories, review length, and manuscript random effects (Section 4.2; Tables 3-5). The dependent variables are sentence-level labels produced by RoBERTa classifiers trained on an external annotation schema (Hua et al., Ref. [25]) and are not defined in terms of the regional predictors, so the regional coefficients are not equalities by construction. The pre-/post-ChatGPT split is a calendar-time indicator (November 2022), not a fitted function of the outcome; the cited Liang et al. estimates of ChatGPT sentence prevalence (Refs. [37,38]) include a co-author of this paper, but they are used only as motivation and for a contextual 20% assumption in Appendix A.1.3, and the main coefficients do not depend on those estimates. The interviews are independent qualitative evidence, and the paper explicitly disclaims causality in Section 5.1 ('our evaluation is not causal'), which keeps the 'muted shift' and 'significant bias' language as interpretive framing rather than as a derivation from the model. No equation reduces a predicted quantity to an input, and no fitted parameter is renamed as a prediction. The causal interpretation of the regression coefficients is a validity and confounding concern, not a circularity concern; the only noteworthy self-citation issue is minor and does not carry the argument.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central quantitative claim rests on measurement and interpretive assumptions more than on fitted parameters. No ad hoc constants are introduced; instead, the load-bearing premises are the country proxy, the validity of the automatic review labelers, and the interpretive move from regional differences to bias.

assumptions (6)
  • domain assumption Reviewer language judgments are expressions of language ideologies, not objective measurements of writing quality.
    Invoked in the Introduction and Section 2 to justify treating clarity critiques as evidence about bias rather than as measures of actual writing skill. If false, the quantitative result would be reinterpreted as a quality signal rather than discrimination.
  • domain assumption Country of the earliest listed institutional affiliation is a usable proxy for author language background.
    Section 4.1.1 acknowledges the country label is a "deeply inexact" match for language background, yet every regional regression coefficient depends on this proxy.
  • domain assumption Stanford's TOEFL exemption policy is a meaningful operationalization of whether a country's institutions use English as a language of instruction.
    Section 4.1.1 uses this binary to define the TOEFL variable for all countries, collapsing heterogeneous language situations into a single institutional category.
  • domain assumption The post-GPT binary variable captures both direct and indirect effects of ChatGPT availability on the review ecosystem.
    Appendix A.1.3 states the variable is not about individual papers known to contain AI text, but about year-level exposure to AI-modified content, which includes indirect behavioral change.
  • domain assumption RoBERTa sentence classifiers trained on external peer-review data accurately identify clarity critique, clarity praise, and other epistemic categories in ICLR reviews.
    Section 4.1.1 describes fine-tuning two models on training data from prior work, but no accuracy, inter-annotator agreement, or domain validation on the ICLR corpus is reported.
  • domain assumption Residual regional differences after controlling for labeled review content can be interpreted as bias rather than unmeasured differences in manuscript quality.
    This assumption is required to move from the correlational finding to the word "bias" in the abstract and conclusion. Section 5.1 acknowledges non-causality but retains the bias framing.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "You Cannot Sound Like GPT": Signs of language discrimination and resistance in computer science publishing." pith.science (2026). https://pith.science/paper/NPS3ED66

@misc{pith2026250508127,
  author       = {Pith},
  title        = {Pith review of: "You Cannot Sound Like GPT": Signs of language discrimination and resistance in computer science publishing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NPS3ED66}},
  note         = {Machine review of arXiv:2505.08127}
}
read the original abstract

LLMs have been celebrated for their potential to help multilingual scientists publish their research. Rather than interpret LLMs as a solution, we hypothesize their adoption can be an indicator of existing linguistic exclusion in scientific writing. Using the case study of ICLR, an influential, international computer science conference, we examine how peer reviewers critique writing clarity. Analyzing almost 80,000 peer reviews, we find significant bias against authors associated with institutions in countries where English is less widely spoken. We see only a muted shift in the expression of this bias after the introduction of ChatGPT in late 2022. To investigate this unexpectedly minor change, we conduct interviews with 14 conference participants from across five continents. Peer reviewers describe associating certain features of writing with people of certain language backgrounds, and such groups in turn with the quality of scientific work. While ChatGPT masks some signs of language background, reviewers explain that they now use ChatGPT "style" and non-linguistic features as indicators of author demographics. Authors, aware of this development, described the ongoing need to remove features which could expose their "non-native" status to reviewers. Our findings offer insight into the role of ChatGPT in the reproduction of scholarly language ideologies which conflate producers of "good English" with producers of "good science."

Figures

Figures reproduced from arXiv: 2505.08127 by the authors.

Figure 1
Figure 1. ICLR has seen a rapid increase in submissions and [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The number of authors from countries which would [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Sentences from reviews which our method infers are about clarity also make inferences about the identity of authors, [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Shifts in reviewer scores on overall paper ratings, praise of writing clarity, and critiques of writing clarity before and [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Correlation heatmap of evaluative categories [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

62 extracted references · 56 canonical work pages

  1. [1]

    Alex A Ahmed, Levin Kim, and Anna L Hoffmann. 2022. ‘This app can help you change your voice’: Authenticity and authority in mobile applications for transgender voice training. Convergence 28, 5 (2022), 1283–1302

  2. [2]

    Nur Ahmed, Muntasir Wahed, and Neil C Thompson. 2023. The growing influence of industry in AI research. Science 379, 6635 (2023), 884–886

  3. [3]

    Sanna J Ali, Angèle Christin, Andrew Smart, and Riitta Katila. 2023. Walking the walk of AI ethics: Organizational challenges and the individualization of risk among ethics entrepreneurs. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 217–226

  4. [4]

    Tatsuya Amano, Valeria Ramírez-Castañeda, Violeta Berdejo-Espinola, Israel Borokini, Shawan Chowdhury, Marina Golivets, Juan David González-Trujillo, Flavia Montaño-Centellas, Kumar Paudel, Rachel Louise White, et al. 2023. The manifold costs of being a non-native English speaker in science. PLoS Biology 21, 7 (2023), e3002184

  5. [5]

    Tatsuya Amano, Valeria Ramírez-Castañeda, Violeta Berdejo-Espinola, Israel Borokini, Shawan Chowdhury, Marina Golivets, Juan David González-Trujillo, Flavia Montaño-Centellas, Kumar Paudel, Rachel Louise White, and Diogo Verís- simo. 2023. The manifold costs of being a non-native English speaker in science. You Cannot Sound Like GPT FAccT ’25, June 23–26,...

  6. [6]

    Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. 2022. The values encoded in machine learning research. In Pro- ceedings of the 2022 ACM conference on fairness, accountability, and transparency . 173–184

  7. [7]

    Irene Bloemraad and Cecilia Menjívar. 2022. Precarious times, professional ten- sions: The ethics of migration research and the drive for scientific accountability. International Migration Review 56, 1 (2022), 4–32

  8. [8]

    Pierre Bourdieu. 2018. Distinction a social critique of the judgement of taste. In Inequality. Routledge, 287–318

Show all 62 references
  1. [9]

    2002.A geopolitics of academic writing

    A Suresh Canagarajah. 2002.A geopolitics of academic writing. Vol. 163. University of Pittsburgh Press

  2. [10]

    Mengjie Cheng, Daniel Scott Smith, Xiang Ren, Hancheng Cao, Sanne Smith, and Daniel A McFarland. 2023. How new ideas diffuse in science. American sociological review 88, 3 (2023), 522–561

  3. [11]

    Theodore Eugene Day. 2015. The big consequences of small biases: A simulation of peer review. Research Policy 44, 6 (2015), 1266–1270

  4. [12]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding.arXiv preprint arXiv:1810.04805 (2018)

  5. [13]

    Mario S Di Bitetti and Julián A Ferreras. 2017. Publish (in English) or perish: The effect on citation rate of using languages other than English in scientific publications. Ambio 46 (2017), 121–127

  6. [14]

    Norbert Elias. 1939. The civilizing process. (1939)

  7. [15]

    Nelson Flores and Jonathan Rosa. 2015. Undoing appropriateness: Raciolinguistic ideologies and language diversity in education. Harvard educational review 85, 2 (2015), 149–171

  8. [16]

    Susan Gal and Judith T Irvine. 2019. Signs of difference: Language and ideology in social life. Cambridge University Press

  9. [17]

    Charles J Gomez, Andrew C Herman, and Paolo Parigi. 2022. Leading countries in global science increasingly receive more citations than other countries doing similar research. Nature Human Behaviour 6, 7 (2022), 919–929

  10. [18]

    Michael D Gordin. 2015. Scientific Babel: How science was done before and after global English. University of Chicago Press

  11. [19]

    Anthony Ha. 2024. NeurIPS keynote speaker apologizes for reference to Chi- nese student. https://techcrunch.com/2024/12/15/neurips-keynote-speaker- apologizes-for-reference-to-chinese-student/

  12. [20]

    Rainer Enrique Hamel. 2007. The Dominance of English in the International Scientific Periodical Literature and the Future of Language Use in Science. AILA Review 20 (2007), 53–71. https://doi.org/10.1075/aila.20.06ham Publisher: John Benjamins Publishing Company ERIC Number: EJ1069221

  13. [21]

    Gordon H Hanson and Matthew J Slaughter. 2016. High-skilled immigration and the rise of STEM occupations in US employment . Technical Report. National Bureau of Economic Research Cambridge, MA

  14. [22]

    Larry Hardesty. 2020. ICLR: The AI conference that helped redefine the field. Amazon Science (April 2020). https://www.amazon.science/blog/iclr-the-ai- conference-that-helped-redefine-the-field

  15. [23]

    Markus Helmer, Manuel Schottdorf, Andreas Neef, and Demian Battaglia. 2017. Gender bias in scholarly peer review. Elife 6 (2017), e21718

  16. [24]

    Antonio J Herrera. 1999. Language bias discredits the peer-review system.Nature 397, 6719 (1999), 467–467

  17. [25]

    Xinyu Hua, Mitko Nikolov, Nikhil Badugu, and Lu Wang. 2019. Argument mining for understanding peer reviews. arXiv preprint arXiv:1903.10104 (2019)

  18. [26]

    Lisa I Iezzoni. 2018. Explicit disability bias in peer review. Medical Care 56, 4 (2018), 277–278

  19. [27]

    Miyako Inoue. 2004. What does language remember?: Indexical inversion and the naturalized history of Japanese women. Journal of Linguistic Anthropology 14, 1 (2004), 39–56

  20. [28]

    Webb Keane. 2003. Semiotics and the social analysis of material things.Language & communication 23, 3-4 (2003), 409–425

  21. [29]

    Neha Kennard, Tim O’Gorman, Rajarshi Das, Akshay Sharma, Chhandak Bagchi, Matthew Clinton, Pranay Kumar Yelugam, Hamed Zamani, and Andrew McCal- lum. 2021. DISAPERE: A dataset for discourse structure in peer review discussions. arXiv preprint arXiv:2110.08520 (2021)

  22. [30]

    Saurabh Khanna, Jon Ball, Juan Pablo Alperin, and John Willinsky. 2022. Recali- brating the scope of scholarly publishing: A modest step in a vast decolonization process. Quantitative Science Studies 3, 4 (2022), 912–930

  23. [31]

    Thomas S. Kuhn. 1977. Objectivity, value judgement, and theory choice. In The essential tension: selected studies in scientific tradition and change . University of Chicago Press, Chicago, 320–339

  24. [32]

    Michèle Lamont. 2009. How professors think: inside the curious world of academic judgment. Harvard University Press, Cambridge, Mass. OCLC: ocn237048345

  25. [33]

    Bruno Latour. 1987. Science in action: How to follow scientists and engineers through society. Harvard university press

  26. [34]

    You are Apple, why are you speaking to me in Turkish?

    Didem Leblebici. 2024. “You are Apple, why are you speaking to me in Turkish?”: the role of English in voice assistant interactions. Multilingua 43, 4 (2024), 455– 485

  27. [35]

    Haley Lepp and Parth Sarin. 2024. A global AI community requires language diverse publishing.. In Proceedings of the Workshop on Global AI Cultures at the International Conference on Learning Representations . https://arxiv.org/pdf/2408. 14772

  28. [36]

    Karen Levy and Solon Barocas. 2018. Privacy at the Margins| refractive surveil- lance: Monitoring customers to manage workers. International Journal of Com- munication 12 (2018), 23

  29. [37]

    Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuan- dong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, et al. 2024. Moni- toring ai-modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews. arXiv preprint arX...

  30. [38]

    Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, et al. 2024. Mapping the Increasing Use of LLMs in Scientific Papers. Under review (2024)

  31. [39]

    Yi Lu. 2025. Trump’s China Policy Is a Disaster for Higher Education . https://www. chronicle.com/article/trumps-china-policy-is-a-disaster-for-higher-education

  32. [40]

    Yingyi Ma. 2024. US security and immigration policies threaten its AI leadership. Brookings Institute (2024). https://www.brookings.edu/articles/us-security-and- immigration-policies-threaten-its-ai-leadership/

  33. [41]

    Floyd Merrell. 2005. Charles Sanders Peirce’s concept of the sign. InThe Routledge companion to semiotics and linguistics . Routledge, 44–55

  34. [42]

    Leon Moosavi. 2022. The myth of academic tolerance: the stigmatisation of East Asian students in Western higher education. Asian Ethnicity 23, 3 (2022), 484–503

  35. [43]

    Dakota Murray, Kyle Siler, Vincent Lariviére, Wei Mun Chan, Andrew M Collings, Jennifer Raymond, and Cassidy R Sugimoto. 2018. Gender and international diversity improves equity in peer review. BioRxiv (2018), 400515

  36. [44]

    Mathias Wullum Nielsen, Christine Friis Baker, Emer Brady, Michael Bang Pe- tersen, and Jens Peter Andersen. 2021. Weak evidence of country-and institution- related status bias in the peer review of abstracts. Elife 10 (2021), e64561

  37. [45]

    Charles Sanders Peirce. 1974. Collected papers of charles sanders peirce . Vol. 5. Harvard University Press

  38. [46]

    Surangika Ranathunga and Nisansa De Silva. 2022. Some languages are more equal than others: Probing deeper into the linguistic disparity in the nlp world. arXiv preprint arXiv:2210.08523 (2022)

  39. [47]

    Cecilia L Ridgeway. 2014. Why status matters for inequality.American sociological review 79, 1 (2014), 1–16

  40. [48]

    Scimago. 2023. Journal Rankings in Computer Science. https://www.scimagojr. com/journalrank.php?area=1700

  41. [49]

    Ali Akbar Septiandri, Marios Constantinides, Mohammad Tahaei, and Daniele Quercia. 2023. WEIRD FAccTs: How Western, Educated, Industrialized, Rich, and Democratic is FAccT?. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 160–171

  42. [50]

    Olivia M Smith, Kayla L Davis, Riley B Pizza, Robin Waterman, Kara C Dobson, Brianna Foster, Julie C Jarvey, Leonard N Jones, Wendy Leuenberger, Nan Nourn, et al. 2023. Peer review perpetuates barriers for historically excluded groups. Nature Ecology & Evolution 7, 4 (2023), 512–523

  43. [51]

    Mengyi Sun, Jainabou Barry Danfa, and Misha Teplitskiy. 2022. Does double- blind peer review reduce bias? Evidence from a top computer science conference. Journal of the Association for Information Science and Technology 73, 6 (2022), 811–819

  44. [52]

    Misha Teplitskiy, Daniel Acuna, Aïda Elamrani-Raoult, Konrad Körding, and James Evans. 2018. The sociology of scientific validity: How professional net- works shape judgement in peer review. Research Policy 47, 9 (Nov. 2018), 1825–

  45. [53]

    The United States House of Representatives Select Committee on the CCP

  46. [54]

    Charles Tilly et al. 1998. Durable inequality. Vol. 685. University of California Press Berkeley

  47. [55]

    Stefan Timmermans and Iddo Tavory. 2012. Theory construction in qualitative research: From grounded theory to abductive analysis. Sociological theory 30, 3 (2012), 167–186

  48. [56]

    Andrew Tomkins, Min Zhang, and William D Heavlin. 2017. Reviewer bias in single-versus double-blind peer review. Proceedings of the National Academy of Sciences 114, 48 (2017), 12708–12713

  49. [57]

    Mark Warschauer, Waverly Tseng, Soobin Yim, Thomas Webster, Sharin Jacob, Qian Du, and Tamara Tate. 2023. The affordances and contradictions of AI- generated text for writers of english as a second or foreign language. Journal of Second Language Writing 62 (2023)

  50. [58]

    Woolard and Bambi B

    Kathryn A. Woolard and Bambi B. Schieffelin. 1994. Language Ideology. Annual Review of Anthropology 23 (1994), 55–82. http://www.jstor.org/stable/2156006

  51. [59]

    Sylvia Wynter. 2003. Unsettling the coloniality of being/power/truth/freedom: Towards the human, after man, its overrepresentation—An argument. CR: The new centennial review 3, 3 (2003), 257–337. FAccT ’25, June 23–26, 2025, Athens, Greece Haley Lepp and Daniel Scott Smith

  52. [60]

    none" category. To do this, annotators, “graduate students in computer science who have undergone training and calibration

    Shira Zilberstein. 2021. National Academies of Science and Medicine: Diversity, Equity and Inclusion in Peer Review. (2021). A REGRESSION V ARIABLES In this section, we discuss the methods for producing our regression variables. A.1 Variables of interest We label the unique re...

  53. [1841]

    https://doi.org/10.1016/j.respol.2018.06.014

  54. [2025]

    Jonathan Levin (President of Stanford) On Transparency from Universities on National Security Risks Posed by Chinese Nationals in STEM Programs

    Letter to Dr. Jonathan Levin (President of Stanford) On Transparency from Universities on National Security Risks Posed by Chinese Nationals in STEM Programs. https://selectcommitteeontheccp.house.gov/media/letters/letter-dr- jonathan-levin-president-stanford-transparency-univ...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.