{"id":"45775a95-13f5-44eb-a1f0-1b5fe7435b5d","arxiv_id":"2505.08127","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Reviewers at ICLR critique writing clarity more for authors from non-English-dominant countries, and after ChatGPT they use AI style as a new cue to infer language background, linking it to perceived science quality.","lead":"This paper analyzes nearly 80,000 ICLR peer reviews and interviews 14 multilingual computer scientists to study how writing quality is judged. It finds that reviewers critique writing more when authors are from countries where English is less widely spoken, and that ChatGPT only partly masks the signs reviewers use to guess an author's language background.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'significant bias' claim in the abstract is causal, but Section 5.1 disclaims causality and the regression lacks any independent measure of manuscript clarity or author language background.","rationale":"The reader's weakest_assumption correctly identifies the core problem: the quantitative claim of bias is an observational association between author region and reviewer clarity critiques, with no baseline for actual writing quality or author first language. I agree that this is the load-bearing point. For the abstract's 'significant bias' claim to hold, the residual regional differences must be due to reviewers' discriminatory responses rather than to unmeasured differences in the manuscripts' clarity. The paper explicitly says in Section 5.1 that 'our evaluation is not causal,' yet the abstract and conclusion use causal and discrimination vocabulary; this mismatch is not merely presentational because the paper's novelty and policy recommendations depend on bias, not just association. The existing controls do not resolve the issue: the 'epistemic value judgments' are derived from the same review sentences via a RoBERTa classifier (Appendix A.1.1), so they are not an independent measure of paper quality, and clarity itself is never externally benchmarked. Author region is imputed from the earliest institutional domain, a coarse proxy for language background openly acknowledged in Section 4.1.1. Measurement error in the key independent variable and the outcome could bias coefficients in either direction. The qualitative interview evidence is a genuine strength: it documents explicit indexical reasoning, including a reviewer who reports grading more harshly when they believe an author is Chinese, and authors working to remove 'non-native' cues. That supports the existence of discriminatory ideologies and the plausibility that some of the quantitative gap reflects bias, so I would not reject the paper; the central claim is conditionally plausible. But the interviews are perceptions and recollections from fourteen informants, not a controlled comparison, and they cannot establish the magnitude or whether the aggregate regional gap reflects bias rather than real clarity differences. A secondary concern is the 'muted shift' framing: the paper's own pre/post comparisons show overlapping confidence intervals for most coefficients (Section 5.1), so 'muted shift' is arguably an overinterpretation of non-significant differences unless a formal interaction model is reported. I leave this as secondary because even if the shift claim were corrected, the main bias claim still requires an external clarity baseline. The proposed concrete check—adding an independent, blinded clarity rating to the regression—would test the load-bearing assumption directly. If the coefficients attenuate, the paper should downgrade 'bias' to 'association'; if they persist, the central claim is materially strengthened. The verdict should remain CONDITIONAL.","tokens_in":23160,"tokens_out":8382,"duration_ms":89203,"concrete_test":"Take a random subsample of approximately 500 ICLR submissions stratified by the share of TOEFL-required authors. Remove author names, affiliations, and acknowledgments; have two or more trained copyeditors (or a validated automated clarity metric, ideally both) independently rate each manuscript's writing clarity. Add this external clarity score as a control in the Section 4.2/Table 4 model and compare the Asia/China/TOEFL coefficients with the current estimates. If the coefficients remain within a pre-registered tolerance (say, at least 80% of the original effect) and significant, the bias interpretation is supported. If they attenuate materially, the observed association is confounded by actual clarity differences, and the abstract's 'significant bias' claim should be softened pending a controlled experiment.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing issue is the inference from the Table 4 Clarity(–) coefficients to 'significant bias' (Abstract; Conclusion). Section 5.1 itself states 'our evaluation is not causal,' yet the abstract and conclusion use a causal, discrimination-oriented framing. The models in Section 4.2 control only for reviewer-labeled epistemic categories, review length, diversity, and manuscript random effects. They contain no independent measure of actual writing quality and no author-level first-language data; country of earliest institution is a very coarse proxy (Section 4.1.1 admits it is 'a deeply inexact match'). If papers from TOEFL-required, Asian, or Chinese-affiliated teams are on average less clear in ways reviewers can perceive, the observed coefficients would be expected even with zero discrimination. The epistemic controls do not close this gap because they come from the same reviews and can themselves be shaped by the same bias, and they do not measure clarity. The qualitative interviews make a plausible case that discriminatory indexical reasoning exists and that ChatGPT-style text is read socially, but they do not establish that the quantitative regional gaps are caused by bias rather than confounding. The central load-bearing assumption is that residual regional differences in clarity critiques should be interpreted as reviewer bias, not as unmeasured differences in writing quality. That assumption is currently untested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies linguistic bias in peer review at ICLR. It combines a statistical analysis of roughly 76,000 reviews from 2018–2024 with interviews of 14 multilingual scholars and one area chair. The quantitative part regresses clarity praise/critique and review ratings on the percentage of authors affiliated with institutions in Asian, Chinese, or TOEFL-required countries, controlling for reviewer-labeled epistemic content, review length, and manuscript random effects, and compares pre- and post-ChatGPT periods. The qualitative part documents indexical reasoning by which reviewers associate writing features with author demographics and with science quality, and describes the shift toward detecting 'ChatGPT style' as a new index of non-nativeness. The paper claims significant bias against authors from countries where English is less widely spoken, a muted effect of ChatGPT availability, and argues that ChatGPT has not broken the link between writing features, author origin, and perceived science quality.","tokens_in":23384,"tokens_out":3920,"duration_ms":40233,"significance":"If the central claim holds, the paper makes a valuable contribution to scholarship on linguistic disadvantage in scientific publishing by offering large-scale quantitative evidence and a theoretically grounded interview study of indexicality in the GPT era. The strengths are real: the code is publicly available, the manuscript honestly reports overlapping confidence intervals and the non-causal nature of the regression design, the qualitative methods are clearly described, and the interviews provide rich, credible evidence that reviewers and authors actively engage in indexical reasoning about language background. The theoretical framing through Peircean semiotics and raciolinguistic ideologies is novel for this empirical setting. The main weakness is that the quantitative analysis cannot adjudicate between reviewer bias and actual regional differences in writing quality, so the causal framing in the abstract and conclusion is not supported by the regression evidence alone.","major_comments":[{"comment":"The central claim, stated in the abstract and conclusion, that 'reviewers critique paper clarity significantly more' for authors from TOEFL-required countries is a causal inference, but the regression in Section 4.2 includes no independent measure of manuscript writing quality and no author-level first-language data. Section 4.1.1 itself calls the country-of-institution proxy 'a deeply inexact match' for language background, and Section 5.1 explicitly states 'our evaluation is not causal.' The coefficients in Table 4 would arise equally if papers from those regions are, on average, less clear in ways that reviewers correctly perceive. This is the load-bearing inference of the paper, and the abstract and conclusion must be revised to present the quantitative results as descriptive associations, with the causal interpretation reserved for the qualitative findings.","section":"Abstract; Section 4.2; Table 4"},{"comment":"The controls for 'other epistemic value judgments' are derived from the same review text that produces the dependent variable. If bias affects all dimensions of evaluation, these are 'bad controls' that may absorb part of the very bias under study; moreover, they do not measure manuscript clarity independently. Therefore the assertion in Section 5.1 that regional differences persist 'even after statistically accounting for the other substantive content of the reviews' overstates the degree of confounding control. The paper should explicitly discuss this limitation and soften the corresponding claim.","section":"Section 4.2; Section 5.1"},{"comment":"The interview evidence convincingly demonstrates that language ideologies and indexical reasoning exist among reviewers and authors, and that ChatGPT-style text is now read as a sign of demographic background. However, interviews describe perceptions and self-reported behaviors, not measured reviewer behavior on actual submissions. They therefore do not establish that the regional gaps in clarity critiques observed in Table 4 are caused by this bias. The conclusion currently conflates the descriptive quantitative association with the qualitative mechanism, and the paper should keep these two forms of evidence clearly separated in its final claims.","section":"Section 4.3; Sections 5.3–5.4"},{"comment":"The post-GPT indicator is a conference-year split rather than a measure of which papers used LLMs, and the 'muted shift' is inferred from point estimates whose confidence intervals overlap across periods. The abstract's 'only a muted shift' claim is therefore fragile: the data are equally consistent with no change in the bias, with a compositional change in the author pool, or with ChatGPT simply improving the average clarity of submissions. The authors should present the shift as suggestive and give explicit attention to these alternative explanations, not just the indexicality account.","section":"Section 4.1.3; Section 5.1; Figure 4"}],"minor_comments":[{"comment":"The phrase 'to assist users in to acquiring cultural capital' contains a typo: 'in to' should be 'in'.","section":"Section 1, paragraph 2"},{"comment":"The caption reads 'Clarity ( )' for the third panel; the minus sign appears to be missing, and the error bars are not defined in the caption.","section":"Figure 4 caption"},{"comment":"The sentence 'Every submission receives a textual review and numerical score score by each reviewer' duplicates 'score'; the duplicate should be removed.","section":"Appendix A.1.1"},{"comment":"The phrase 'surprised by the muted improvements to equity of clarity critique and praise' is awkward and overstates the evidence; it would be more precise to say 'surprised by the muted improvement in equity of clarity critique and praise.'","section":"Section 5.1"},{"comment":"Reference [38] is listed as 'Under review (2024)' without a venue; if the paper has been published, the citation should be updated.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper fits FAccT well and the qualitative component is strong and original. My main concern is the gap between the causal language in the abstract and conclusion and the correlational design; this is fixable by reframing the quantitative claims and clearly delimiting what each method can establish. The self-citation to Liang et al. is used only to justify the post-GPT split and is not a serious issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper deserves a serious referee, but the abstract and conclusion should be dialed back. What's genuinely new is the interview evidence: reviewers now treat 'ChatGPT style' as an index of non-native background, and authors actively work to avoid sounding like GPT. That is a concrete observation about how language bias mutates rather than disappears, and the qualitative work is careful and convincing. The quantitative piece is also useful—roughly 80k ICLR reviews, pre/post ChatGPT, showing clarity critiques correlate with the share of authors from TOEFL-required countries, and that the association only partially shrinks after 2022. The regressions are simple, internally consistent, and Section 5.1 explicitly says the analysis is not causal. Credit where due: the paper ships code, uses public data, and the interviews are described with proper attention to confidentiality and recruitment bias.\\n\\nThe soft spot is the gap between the abstract and the evidence. The abstract says 'significant bias against authors...' and the conclusion repeats it. But the regression has no independent measure of manuscript clarity or author first language. Country of earliest institution is a coarse proxy, and the controls are reviewer-labeled epistemic categories drawn from the same reviews—those controls can be shaped by the same bias they are meant to absorb. If papers from these regions are on average less clear in ways reviewers perceive, you would see the same coefficients with zero discrimination. The interviews make the bias interpretation plausible, but they don't rescue the regression from that confounding. The authors seem aware of this; the limitations section acknowledges confounders and the disclaimers are there. The fix is mostly framing: call the quantitative result a descriptive association consistent with bias, report classifier validation for the RoBERTa models, and lean on the interviews for the discrimination claim.\\n\\nWho is this for? Anyone working on peer review, fairness in AI venues, language ideology, or the social impact of LLMs. It deserves peer review rather than desk rejection. With revisions to soften the causal language and a bit more transparency on the text classifiers, it would be a solid contribution. I'd bring it to our reading group, and I'd probably cite the interview finding on indexical shift in my own work.","headline":"The interviews are the real contribution; the regression is honest but the abstract oversells a causal reading of 'bias.'","tokens_in":23876,"tokens_out":1714,"would_cite":true,"duration_ms":19793,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Peer reviewers penalize authors from non-English-speaking countries, and ChatGPT only changes the tell.","keywords":["language ideologies","peer review","ChatGPT","discrimination","semiotics","indexicality","ICLR","linguistic bias"],"falsifier":"A controlled experiment would settle it: send the same manuscript text to a large pool of reviewers with only authorship cues varied (affiliation country, author names, or acknowledgment names) and measure clarity critiques and scores; if the cues do not move evaluations, the central bias claim would be refuted. Alternatively, if an independent, origin-blind linguistic rating of manuscript clarity fully accounted for the regional gap in clarity critiques, the discrimination interpretation would collapse.","tokens_in":22967,"feed_emoji":"🌐","tokens_out":10168,"duration_ms":99694,"temperature":0.7,"pith_summary":"This paper argues that ChatGPT does not solve linguistic bias in scientific peer review; it changes the surface on which the bias is read. Analyzing 76,453 reviews of 20,827 submissions to the ICLR conference (2018–2024), the authors find that manuscripts with more authors from countries where English is less widely spoken receive significantly more critiques of writing clarity, fewer praises of clarity, and lower overall scores, even after other epistemic critiques are statistically held constant. After ChatGPT's release in November 2022, the pattern only weakened modestly. Interviews with 14 multilingual ICLR participants explain why: reviewers shifted from grammar-based cues to 'ChatGPT style' and non-linguistic cues—long author lists, acknowledgment names, perceived experimental volume—as signs that the author is not a 'native' English speaker, and they still tie that perceived identity to the quality of the science. The paper's wager is that language ideologies are durable: when one sign disappears, readers recruit another.","feed_headline":"ChatGPT shifts bias from grammar to 'GPT style' in peer review","feed_subtitle":"Reviewers still tie perceived author origin to science quality, so masking one tell just creates another.","key_machinery":"The load-bearing mechanism is the indexical sign: a concrete textual feature—a missing plural, a long sentence, an LLM-typical phrase, a long author list—stands to a reviewer for an imagined type of person, and that imagined person stands in turn for an evaluation of the science. This two-step chain explains both the quantitative results and their persistence: removing the first sign with ChatGPT does not break the chain, because reviewers recruit new signs, while authors engage in 'language labor' to strip any feature that could expose their language background. Methodologically, the chain is operationalized with sentence-level classifiers that count clarity critiques and praises from review text, and with panel regressions comparing pre- and post-ChatGPT evaluation patterns.","core_discovery":"The central claim is that clarity judgment in peer review is not an objective measure of readability but a socially loaded index of the imagined author. Across the full ICLR corpus, a higher percentage of authors affiliated with Asian, Chinese, or English-proficiency-tested countries is associated with more clarity critiques, less clarity praise, and lower ratings; these associations are significant before and, in attenuated form, after ChatGPT, and for most groups the pre/post shift is not statistically significant. The qualitative arm supplies the mechanism: grammatical idiosyncrasies (tense and plural errors, long clunky sentences) index a 'non-native speaker,' and that imagined speaker indexes lesser science; when ChatGPT removes the grammatical signs, reviewers report reading AI style, word choice, jargon density, author count, and acknowledgment names as new signs of origin. The title quote—'you cannot sound like GPT'—captures the trap: authors who polish with LLMs are heard as non-native in a new way, because AI style itself is now read as the accent of the language-marginalized. The paper presents this indexical chain as the reason availability of LLMs produces only a muted statistical shift, and as evidence that linguistic exclusion is reproduced rather than dissolved by the technology.","pith_inferences":["Editorial inference: if the indexical-shift logic is right, any widely adopted writing or translation tool will eventually acquire its own stylistic signature, and that signature will become a demographic marker wherever readers judge clarity while caring about author origin.","Editorial inference: a pre-registered experiment—identical text presented with different author-origin cues (affiliation, names, acknowledgments)—would directly test the mechanism; the paper's framework predicts clarity ratings and trust in the science should move with the cues.","Editorial inference: the findings imply that evaluation design could weaken the chain more effectively than author-side tools: separating grammatical correctness from argumentative clarity in rubrics, or having trained editors rather than reviewers judge prose, would give the bias fewer signs to attach to.","Editorial inference: because even interviewees described feeling 'damned if you do, damned if you don't,' the study suggests that uniform AI adoption may create a new hierarchy between writers who can mimic a specific elite register and those who cannot, rather than flattening linguistic difference."],"forward_implications":["AI-assisted polishing can narrow but not eliminate the regional gap in clarity critiques, because new non-grammatical signs of author origin take the place of grammatical ones.","Language-marginalized authors face a shifting 'language labor' tax: in addition to English, they must learn the current conventions for not sounding like AI, since AI style is itself read as non-native.","Reviewers' equation of 'good English' with 'good science' means clarity critiques double as status judgments, producing a durable inequality that is only partially masked by anonymous review.","The paper's own recommendations follow: more multilingual publishing venues, conferences run with local languages, and language coursework for graduate programs, rather than reliance on author-side LLM polishing."],"supporting_citations":[{"why":"Supplies the language-ideology framework: statements about language are never merely statements, they entail ideological positions.","marker":"[16]"},{"why":"Provides the semiotic theory of signs and indexicality used to model how writing features evoke an author's imagined identity.","marker":"[45]"},{"why":"Provides the manually annotated training data and argument-mining approach used to label review sentences into evaluative aspects and polarity.","marker":"[25]"},{"why":"Contributes the peer-review discourse dataset whose six evaluative dimensions structure the paper's dependent and control variables.","marker":"[29]"},{"why":"Documents the rise of ChatGPT-modified sentences in scientific preprints, motivating the pre/post-GPT comparison and the conference choice.","marker":"[38]"},{"why":"Provides the estimate that roughly 20% of post-2023 submissions contain substantially generated text, grounding the Post-GPT variable.","marker":"[37]"},{"why":"Shows that linguistically adapted speakers remain marked by the listening subject, the parallel used to explain why ChatGPT-style writing still indexes non-native status.","marker":"[15]"},{"why":"Establishes the manifold costs of being a non-native English speaker in science, framing the 'tax' the paper measures in clarity critiques.","marker":"[5]"}],"fun_headline_variants":["Peer review bias survives ChatGPT's grammar mask","GPT style becomes new accent in peer review bias","LLMs change the tell: from grammar to AI style","AI style outs non-native authors despite ChatGPT","Linguistic bias persists as reviewers read GPT style"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The inference that regional differences in clarity critiques reflect reviewer discrimination rather than real differences in writing quality assumes that, after controlling for the other epistemic content of reviews, no unmeasured differences in manuscript clarity or quality remain—yet the dataset contains no independent measure of either author first language or objective writing clarity.","fun_headline_variants_meta":{"raw":{"variants":["Peer review bias survives ChatGPT's grammar mask","GPT style becomes new accent in peer review bias","LLMs change the tell: from grammar to AI style","AI style outs non-native authors despite ChatGPT","Linguistic bias persists as reviewers read GPT style"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1314,"prompt_tokens":1018,"completion_tokens":296,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":634,"completion_tokens_details":{"reasoning_tokens":224}},"tokens_in":634,"tokens_out":296,"duration_ms":3668,"temperature":1.0,"reasoning_tokens":224,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:02:26.733559+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment would settle it: send the same manuscript text to a large pool of reviewers with only authorship cues varied (affiliation country, author names, or acknowledgment names) and measure clarity critiques and scores; if the cues do not move evaluations, the central bias claim would be refuted. Alternatively, if an independent, origin-blind linguistic rating of manuscript clarity fully accounted for the regional gap in clarity critiques, the discrimination interpretation would collapse.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the language-ideology framework: statements about language are never merely statements, they entail ideological positions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the semiotic theory of signs and indexicality used to model how writing features evoke an author's imagined identity."},{"cited_title":"Argument Mining for Understanding Peer Reviews","cited_arxiv_id":"1903.10104","evidence_quote":"Provides the manually annotated training data and argument-mining approach used to label review sentences into evaluative aspects and polarity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the rise of ChatGPT-modified sentences in scientific preprints, motivating the pre/post-GPT comparison and the conference choice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that linguistically adapted speakers remain marked by the listening subject, the parallel used to explain why ChatGPT-style writing still indexes non-native status."}],"review_version":1}