Pith. sign in

REVIEW 4 major objections 6 minor 79 references

Disparities in Peer Review Tone and the Role of Reviewer Anonymity

T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that reviewer tone in first-round reviews of accepted papers varies with corresponding authors' gender, race, region, and institutional rank, and that signing reviews narrows appreciative and constructive gaps but not…

desk verdict A large, well-validated descriptive study of tone disparities in peer review, but the headline claim that identity disclosure reduces those disparities is not supported by the paper's own statistics. read the letter →

arxiv 2507.14741 v1 pith:CQEP7NQM submitted 2025-07-19 cs.CL

classification cs.CL
keywords peerreviewrevieweranonymityopenlinguisticbiassentimentanalysistoneclassificationn-gramlargelanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper examines more than 80,000 first-round peer reviews of accepted papers from two open-access journals and asks whether the language of evaluation differs by author demographics and by whether reviewers sign their names. It claims that female, white, and western-affiliated corresponding authors receive more appreciative and positive language, especially under anonymous review, while non-white and eastern-affiliated authors receive more critical and negative language and more correction-oriented feedback. Comparing anonymous and signed reviews, the study reports that disclosure is associated with a marked reduction in appreciative and constructive tone gaps, while critical-evaluative gaps persist. A sympathetic reader would care because these findings bear directly on policy debates about anonymous versus open peer review: they suggest that signing reviews may dampen some identity-driven tone differences without eliminating them.

What carries the argument

The analysis is carried by a three-tier linguistic pipeline and a regression framework. First, a SciBERT classifier fine-tuned on peer review sentences assigns positive, neutral, or negative sentiment. Second, a curated lexicon of 269 bigrams and trigrams, grouped into manuscript appreciation, requests for clarification, constructive suggestions, and strong methodological critique, tracks recurring evaluative phrasing normalized per 1,000 words. Third, GPT-4o classifies each sentence into appreciative, constructive-analytical, questioning, critical-evaluative, or contextual-summarizing tones, validated against authors (91.1% agreement) and independent researchers (95.1% agreement); a weighted tone score combines sentence frequency and sentence length. Ordinary least squares regressions with robust standard errors, field-of-science fixed effects, and separate models for anonymous and disclosed reviews then estimate how each tone varies with gender, race, region, institutional rank, academic age, and reviewer disclosure.

What would settle it

A decisive test would be a randomized or policy-forced comparison in which the same reviewers evaluate comparable manuscripts under anonymous and signed conditions, for example by tracking a journal's switch to mandatory open review; if tone and demographic gaps are identical across conditions, the paper's claim that disclosure reduces disparities would be refuted, while persistence of critical gaps under mandatory signing would support it.

Watch

Extended reading notes

Core claim

The central claim is that reviewer tone in first-round reviews of ultimately accepted papers is measurably associated with the demographics and institutional standing of the corresponding author, and that reviewer identity disclosure changes the pattern. In the paper's own terms, female, white, and western-affiliated authors were more likely to receive appreciative and positive sentiment, especially under anonymous review, whereas non-white and eastern-affiliated authors were more frequently met with critical and negative language and with constructive or clarifying suggestions. When reviewers disclosed their identities, the elevated appreciation for white, female, and western authors largely disappeared and several n-gram level gaps in critique, constructiveness, and clarification also lost significance; however, critical-evaluative tone remained higher for eastern-affiliated and non-top-100 authors, and constructive-analytical tone toward eastern-affiliated authors persisted. The paper reads this as evidence that disclosure moderates extremes in tone and encourages more balanced feedback, without evidence that disclosure makes reviews harsher.

Load-bearing premise

The load-bearing premise is that the difference in tone between signed and unsigned reviews reflects the effect of losing anonymity, but reviewers choose whether to sign, so reviewers who disclose may simply be more generous or constructive people; if so, the observed 'disclosure' effect is not caused by disclosure.

Editorial extensions

If this is right

  • If the central claim holds, anonymous single-blind review of accepted papers already contains measurable identity-linked differences in evaluative language, even in the subset of reviews where bias should be lowest.
  • Signing reviews would plausibly narrow appreciative and constructive tone gaps, partly because disclosed reviews use more appreciative, constructive, and questioning language overall.
  • Critical-evaluative disparities for eastern-affiliated and non-top-100 authors would persist under disclosure, so identity transparency alone would not equalize evaluation.
  • The finding that disclosed reviews show stronger growth in appreciative and constructive tone from 2019 to 2024 and no increase in critical tone would counter the worry that open review makes reviewers harsher.
  • Because only first-round reviews of accepted papers are observed, the same disparities would be expected to be at least as large in reviews of rejected manuscripts if the mechanism is bias in tone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The anonymous-versus-signed comparison is observational, so the paper cannot rule out that reviewers who sign are self-selected and systematically different in generosity or writing style; the directional conclusion about disclosure should be read with that caveat.
  • If the critical-tone gap is driven by cues other than identity, such as topic, methods, or writing style correlated with region and institution, then mandatory disclosure would not remove it; a testable extension is to randomize reviewers' signing status or exploit a journal's policy change.
  • The paper's focus on accepted papers means its estimates may understate total bias, since rejected manuscripts receive no public reviews; including rejected reviews from journals that publish them would test whether tone disparities intensify.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper examines peer review language in more than 80,000 first-round reviews of accepted manuscripts published in Nature Communications and PLOS One (2019–2024). It combines a fine-tuned SciBERT sentiment classifier, an n-gram analysis of evaluative phrases, and GPT-4o sentence-level tone classification to measure appreciative, constructive-analytical, questioning, and critical-evaluative tone. Tone scores are regressed on corresponding-author demographics (gender, race, region, institutional rank, academic age, productivity, prior impact), paper features, review features, and reviewer identity disclosure. The paper reports widespread demographic disparities in sentiment and tone, and argues that disclosed reviewer identity is associated with more appreciative and constructive language and with the disappearance of several demographic tone disparities, mainly in appreciative tone, while critical tone disparities persist. The Discussion draws policy implications in favor of reviewer disclosure.

Significance. If the descriptive associations are robust, the paper is a useful empirical contribution: it assembles a large, carefully documented corpus, validates GPT-4o tone labels against independent researchers (95.1%) and against the reviewed authors themselves (91.1%), publishes the full n-gram lexicon and survey instruments in the supplement, and uses Huber-White robust standard errors in the regressions. The quantitative linguistic description of demographic disparities in review tone is valuable in its own right. However, the central policy conclusion about reviewer disclosure is not supported by the statistical evidence as reported, because the anonymous-versus-disclosed comparison is confounded by self-selection and is based on informal contrasts across separate subsample regressions rather than formal interaction tests. The paper's significance therefore hinges on a revision that either supplies interaction tests and a defensible identification strategy or recalibrates the claims to descriptive associations.

major comments (4)
  1. [Results, "Tone Variation across Author Groups"; Supplementary Table 4] The claim that demographic disparities in tone disappear under identity disclosure is inferred from comparing coefficients across two separate regressions, one for anonymous and one for disclosed reviews, in Supplementary Table 4. Non-significance in the smaller disclosed subsample does not imply that the coefficient is smaller than in the larger anonymous subsample. For example, for gender on appreciative tone, the anonymous coefficient is -0.283 (SE 0.107) and the disclosed coefficient is 0.033 (SE 0.224); the difference is about 0.316, giving an approximate z of 1.27, which is not statistically significant. The same problem affects the race comparison on appreciation (0.360 vs 0.277, SEs 0.113 and 0.253) and the region comparison (0.809 vs 0.366, SEs 0.120 and 0.267). The paper should estimate models with interaction terms between reviewer identity disclosure and the relevant author demographics, and report those interaction coefficients and their standard errors. Without such tests, the headline statement that "disclosure appears to foster a more balanced and equitable evaluative environment" is unsupported by the paper's own statistics.
  2. [Materials and Methods, Data ("Reviewer anonymity"); Discussion] The anonymous-versus-disclosed comparison is used to draw causal or quasi-causal conclusions about the effect of identity disclosure, but disclosure status is chosen by reviewers, not assigned. The paper has no reviewer-level covariates, no reviewer fixed effects, no instrument, and no random assignment; Table 1 reports only the counts of anonymous and signed reviews. If reviewers who sign are systematically more generous, more constructive, or more relationally connected to the authors, the observed tone differences may reflect selection rather than the consequences of disclosure. The sentence in the Discussion that "reviewer identity disclosure appears to foster a more balanced and equitable evaluative environment" overstates what an observational contrast can establish. The manuscript should either add a design that can address selection (for example, within-reviewer comparisons when the same reviewer reviews under both conditions, or reviewer-level propensity adjustments conditional on observable reviewer characteristics) or explicitly reframe the disclosure results as descriptive associations and soften the policy recommendations accordingly.
  3. [Materials and Methods, "Using SciBERT to classify sentiment"; Results, "Sentiment Differences Across Author Groups"] The sentiment analysis rests on a fine-tuned SciBERT classifier whose validation accuracy is 0.677 and weighted F1 is 0.671. The paper does not report class-wise precision, recall, or a confusion matrix, nor does it report how the two-sample proportion Z-tests in Figure 2b behave under plausible classification error rates or different decision thresholds. Since the sentiment findings are presented as evidence of demographic disparities in positive and negative review tone, the paper should either provide per-class diagnostics plus a sensitivity analysis demonstrating that the reported disparities are not artifacts of measurement error, or explicitly relegate the sentiment results to an exploratory supporting analysis. The strong claims in the Results about negative sentiment being "more commonly directed toward" specific demographic groups need stronger measurement support.
  4. [Results, "Tone Variation across Author Groups"; Supplementary Table 4] There is an internal inconsistency in the interpretation of the regional coefficient. The text says that "the tendency for reviewers to use more appreciative language when evaluating submissions from female, white, or eastern-affiliated corresponding authors disappears entirely when reviewer identity is disclosed," but Supplementary Table 4 and Figure 4 define the regional variable as "Western group," with a positive coefficient for appreciative tone in the anonymous model (0.809, SE 0.120). This means Western-affiliated authors, not eastern-affiliated authors, receive more appreciative language under anonymous review. The Discussion paragraph states the direction correctly ("Female, white, and western-affiliated authors were more likely to receive appreciative and positive sentiment"), so the Results paragraph appears to contain a substantive reversal rather than a mere wording issue. This should be corrected and the surrounding interpretation checked for consistency.
minor comments (6)
  1. [Supplementary Note 2, Survey content] The definition of the "Questioning" tone in the consent-and-instructions text is identical to the definition of "Constructive-Analytical" ("Provides a detailed examination of the study's methodology, data, or findings, offering suggestions for improvement"). This likely affected how survey participants interpreted the tone labels and should be fixed in the supplement and in any replication materials.
  2. [Materials and Methods, "Weighted Tone Scoring Method"] The weighted tone score uses alpha = 0.2, and the text states that any value between 0.1 and 0.4 would have been reasonable, but no sensitivity analysis is reported. A short robustness table showing how the main regression coefficients and their standard errors vary across alpha values in this range would address concerns about a hand-set free parameter.
  3. [Results, "Evaluative Language Patterns across Author Groups"; Materials and Methods, "Identifying and Mapping N-grams"] The n-gram and sentiment analyses run many two-group comparisons across demographic axes, age bins, and n-gram categories, each with its own significance asterisks. No multiple-comparison correction is applied. Reporting false-discovery-rate-adjusted p-values, or at least stating how many comparisons were made, would make the pattern of significant results easier to evaluate.
  4. [Figure 2 caption and labels] The figure caption says the y-axis shows "subgroup contribution within sentiment," while the main text describes the y-axis as the proportion of reviews within each subgroup falling into a given sentiment category. These are different quantities; the caption and axis labels should be aligned with the actual computation.
  5. [Table 1 and regression sample sizes] Table 1 reports 80,187 reviews, but the full-data regressions in Supplementary Table 3 use 76,644 observations. The manuscript should state how many reviews were dropped due to missing tone scores or other features, so the discrepancy is transparent.
  6. [General] The article does not include a data availability or code availability statement. Given the policy-oriented nature of the claims and the unusual data assembly process, a clear statement about whether the review texts, derived features, and analysis code will be shared would strengthen reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: tone outcomes are independently constructed from review text, and demographic/disclosure predictors come from external data, so no claimed result reduces by construction to its inputs.

full rationale

The paper's outcomes (sentiment, n-gram frequencies, and GPT-4o tone scores) are generated from peer-review text, while the predictors (author gender, race, region, institutional rank, and reviewer disclosure status) come from independent sources such as Genderize.io, NamSor, OpenAlex, and journal metadata. The weighted tone score in Eq. (1) uses a hand-set alpha of 0.2, but this parameter does not encode any demographic or disclosure effect, so it cannot force the reported associations. Validation of the tone classifier against impartial researchers and corresponding authors does not make the outcome a function of the predictors. The only self-citations are AlShebli et al. [59,60], cited for a citation-window convention in measuring prior impact; this is not load-bearing for the paper's central claims. The paper's conclusion that disclosure reduces demographic disparities rests on comparing coefficients across stratified subsamples rather than estimating disclosure-by-demographic interactions; that is a statistical inference concern, not circularity, because the compared coefficients are not fitted inputs renamed as predictions. No equation, definition, or self-citation chain reduces the reported findings to the paper's own inputs.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central claims rest on several measurement choices: the GPT-4o tone labels, the hand-built n-gram lexicon, the weighted tone score with alpha=0.2, and the name-based demographic inference. Each is validated to a degree, but none has shipped artifacts. The anonymity conclusion additionally rests on the untested assumption that self-selection into signing does not drive the tone differences.

free parameters (6)
  • alpha in weighted tone score = 0.2
    Hand-chosen weight balancing word and sentence proportions; any value between 0.1 and 0.4 is said to be reasonable, but no sensitivity analysis is shown. This parameter directly defines the dependent variable in all tone regressions.
  • N-gram frequency thresholds = trigrams > 100 occurrences, bigrams > 1000 occurrences
    Hand-chosen cutoffs that produce the final lexicon of 269 n-grams; the resulting lexicon underpins all n-gram analyses.
  • Gender and race inference confidence threshold = 70%
    Papers are excluded if the Genderize or NamSor confidence is below 70%; this affects sample composition and could bias demographic comparisons.
  • Top 100 institutional rank definition = Any of THE, QS, or Nature Index top 100
    Institutional prestige is reduced to a binary variable based on inclusion in any of three rankings; no continuous measure or robustness check is provided.
  • Two-year citation window for prior impact = citations within first 2 years after publication
    Adopted from AlShebli et al.; this choice defines the 'average prior impact' control variable but is not varied or justified beyond citation to prior work.
  • Academic age bins = 0-10, 11-25, 26+ years
    Chosen cutoffs for early, mid, and senior career; used in sentiment subgroup analysis and as a control, without justification or robustness checks.
assumptions (6)
  • domain assumption Tone labels assigned by GPT-4o accurately reflect the four predefined categories across the corpus.
    Validated on 100 reviews (95.1% external agreement) and 21,150 sentences judged by 531 authors (91.1% agreement), but validation samples are small or self-selected and the API model version is unspecified.
  • domain assumption Inferred gender and race from author names are accurate enough for group comparisons.
    Genderize and NamSor accuracies of 94.4% and 83.2% are estimated on a self-selected survey subsample (8% response rate); misclassification may be correlated with region and affect group comparisons.
  • ad hoc to paper First-round reviews of accepted manuscripts are an unbiased or conservative setting for studying review bias.
    Authors state that if bias appears in accepted reviews it likely persists in rejected ones, but this extrapolation is unverified; rejected reviews are unavailable.
  • domain assumption OLS regressions with Huber-White robust standard errors give valid statistical inference despite clustering of reviews within papers and authors.
    No clustering by manuscript or corresponding author is performed; residual correlations can understate standard errors.
  • domain assumption The four tone categories and the four n-gram categories capture the evaluatively relevant language of peer review.
    Categories are based on feedback theory but operationalized by the authors; sentences that do not fit are labeled contextual-summarizing.
  • ad hoc to paper The weighted tone score with alpha=0.2 represents tone prevalence faithfully.
    The mixing weight is chosen by judgment; robustness to alpha in [0.1, 0.4] is asserted but not demonstrated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Disparities in Peer Review Tone and the Role of Reviewer Anonymity." pith.science (2026). https://pith.science/paper/CQEP7NQM

@misc{pith2026250714741,
  author       = {Pith},
  title        = {Pith review of: Disparities in Peer Review Tone and the Role of Reviewer Anonymity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CQEP7NQM}},
  note         = {Machine review of arXiv:2507.14741}
}
read the original abstract

The peer review process is often regarded as the gatekeeper of scientific integrity, yet increasing evidence suggests that it is not immune to bias. Although structural inequities in peer review have been widely debated, much less attention has been paid to the subtle ways in which language itself may reinforce disparities. This study undertakes one of the most comprehensive linguistic analyses of peer review to date, examining more than 80,000 reviews in two major journals. Using natural language processing and large-scale statistical modeling, it uncovers how review tone, sentiment, and supportive language vary across author demographics, including gender, race, and institutional affiliation. Using a data set that includes both anonymous and signed reviews, this research also reveals how the disclosure of reviewer identity shapes the language of evaluation. The findings not only expose hidden biases in peer feedback, but also challenge conventional assumptions about anonymity's role in fairness. As academic publishing grapples with reform, these insights raise critical questions about how review policies shape career trajectories and scientific progress.

Figures

Figures reproduced from arXiv: 2507.14741 by the authors.

Figure 1
Figure 1. • Publication year: We recorded the publication year for each paper to analyze trends in peer review over time and observe shifts in reviewer tone influenced by evolving editorial practices, standards, and global events. Our dataset focuses on papers published from 2019 to 2024, aligning with the policies of Nature Communications and PLOS ONE, which introduced trans￾parent peer review options in 2016 and 2019, respe… view at source ↗
Figure 1
Figure 1. Overview of the dataset. The figure presents key characteristics of our dataset, including: (a) the distribution of corresponding author attributes such as institutional affiliation, gender, race, and academic age; (b) paper-specific features like field of science, team size, and submission timelines (for a more detailed breakdown of the fields of science, refer to Supplementary [PITH_FULL_IMAGE:figures/full_fig_p0… view at source ↗
Figure 2
Figure 2. Sentiment analysis of peer reviews across author subgroups and academic age bins. (a) Distribution of subgroup contributions to each sentiment category (Positive, Neutral, Negative) across academic age bins: [0,10], [11,25], and [26+]. Marker color indicates gender (Male vs. Female), border color reflects race (White vs. Non-white), marker shape encodes region of affiliation (Western vs. Eastern), and fill opacity r… view at source ↗
Figures from the paper (2 more)
Figure 3
Figure 3. Figure 3: Group-level differences in linguistic pattern frequencies across author demographics. Bars show the average frequency (per 1,000 words) of four grouped bigram/trigram categories (“manuscript appreciation”, “constructive suggestions”, “requests for clarification”, and “…
Figure 4
Figure 4. Figure 4: Regression coefficients and 95% confidence intervals estimated for each tone category using the full dataset (gray) and reviewer-stratified models (blue = anonymous, orange = disclosed). The first row displays the effect of reviewer identity disclosure, included as a b…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

79 extracted references · 76 canonical work pages

  1. [1]

    Tennant, J. P. & Ross-Hellauer, T. The limitations to our understanding of peer review. Research integrity and peer review 5, 6 (2020)

  2. [2]

    J., Sugimoto, C

    Lee, C. J., Sugimoto, C. R., Zhang, G. & Cronin, B. Bias in peer review. Journal of the American Society for information Science and Technology 64, 2–17 (2013)

  3. [3]

    Aczel, B. et al. The present and future of peer review: Ideas, interventions, and evidence.Proceedings of the National Academy of Sciences 122, e2401232121 (2025)

  4. [4]

    Schwartz, S. J. & Zamboanga, B. L. The peer-review and editorial system: Ways to fix something that might be broken. Perspectives on Psychological Science 4, 54–61 (2009)

  5. [5]

    Who’s afraid of peer review? (2013)

    Bohannon, J. Who’s afraid of peer review? (2013)

  6. [6]

    & Heavlin, W

    Tomkins, A., Zhang, M. & Heavlin, W. D. Reviewer bias in single-versus double-blind peer review. Proceedings of the National Academy of Sciences 114, 12708–12713 (2017)

  7. [7]

    Peer review: a flawed process at the heart of science and journals

    Smith, R. Peer review: a flawed process at the heart of science and journals. Journal of the royal society of medicine 99, 178–182 (2006)

  8. [8]

    & Raphael, E

    Mulligan, A., Hall, L. & Raphael, E. Peer review in a changing world: An international study measuring the attitudes of researchers. Journal of the American Society for Information Science and Technology 64, 132–161 (2013)

Show all 79 references
  1. [9]

    Jefferson, T., Rudin, M., Folse, S. B. & Davidoff, F. Editorial peer review for improving the quality of reports of biomedical studies. Cochrane Database of Systematic Reviews 1 (2006)

  2. [10]

    Silbiger, N. J. & Stubler, A. D. Unprofessional peer reviews disproportionately harm underrepre- sented groups in stem. PeerJ 7, e8247 (2019)

  3. [11]

    Beaumont, L. J. Peer reviewers need a code of conduct too. Nature 572, 439–440 (2019)

  4. [12]

    this work is antithetical to the spirit of research

    Hyland, K. & Jiang, F. K. “this work is antithetical to the spirit of research”: An anatomy of harsh peer reviews. Journal of English for Academic Purposes 46, 100867 (2020)

  5. [13]

    Verharen, J. P. Chatgpt identifies gender disparities in scientific peer review. Elife 12, RP90230 (2023)

  6. [14]

    Budden, A. E. et al. Double-blind review favours increased representation of female authors. Trends in ecology & evolution 23, 4–6 (2008)

  7. [15]

    & Maruˇsi´c, A

    Buljan, I., Garcia-Costa, D., Grimaldo, F., Squazzoni, F. & Maruˇsi´c, A. Large-scale language analysis of peer review reports. Elife 9, e53249 (2020)

  8. [16]

    Squazzoni, F. et al. Peer review and gender bias: A study on 145 scholarly journals.Science advances 7, eabd0299 (2021)

  9. [17]

    J., Kahn, S

    Ceci, S. J., Kahn, S. & Williams, W. M. Exploring gender bias in six key domains of academic science: An adversarial collaboration. Psychological Science in the Public Interest24, 15–73 (2023). 17

  10. [18]

    Smith, O. M. et al. Peer review perpetuates barriers for historically excluded groups. Nature Ecology & Evolution 7, 512–523 (2023)

  11. [19]

    & Fang, F

    Casadevall, A. & Fang, F. C. Causes for the persistence of impact factor mania. MBio 5, 10–1128 (2014)

  12. [20]

    von Wedel, D. et al. Affiliation bias in peer review of abstracts by a large language model. JAMA 331, 252–253 (2024)

  13. [21]

    & ˙Zela´zniewicz, A

    Kowal, M., Sorokowski, P., Kulczycki, E. & ˙Zela´zniewicz, A. The impact of geographical bias when judging scientific studies. Scientometrics 127, 265–273 (2022)

  14. [22]

    G., Politzer-Ahles, S

    Cuskley, C., Roberts, S. G., Politzer-Ahles, S. & Verhoef, T. Double-blind reviewing and gender biases at evolang conferences: An update. Journal of Language Evolution 5, 92–99 (2020)

  15. [23]

    & Teplitskiy, M

    Sun, M., Barry Danfa, J. & Teplitskiy, M. Does double-blind peer review reduce bias? evidence from a top computer science conference. Journal of the Association for Information Science and Technology 73, 811–819 (2022)

  16. [24]

    Strauss, D., Gran-Ruaz, S., Osman, M., Williams, M. T. & Faber, S. C. Racism and censorship in the editorial and peer review process. Frontiers in Psychology14, 1120938 (2023)

  17. [25]

    & Dinesh, S

    Kulal, A., N, A., Shareena, P. & Dinesh, S. Unmasking favoritism and bias in academic publishing: An empirical study on editorial practices. Public Integrity 1–22 (2025)

  18. [26]

    & Roth, D

    Zhang, J., Zhang, H., Deng, Z. & Roth, D. Investigating fairness disparities in peer review: A language model enhanced approach. arXiv preprint arXiv:2211.06398 (2022)

  19. [27]

    & Cohan, A

    Beltagy, I., Lo, K. & Cohan, A. Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676 (2019)

  20. [28]

    K., Navlakha, M., Agarwal, M

    Bharti, P. K., Navlakha, M., Agarwal, M. & Ekbal, A. Politepeer: does peer review hurt? a dataset to gauge politeness intensity in the peer reviews. Language Resources and Evaluation 1–23 (2023)

  21. [29]

    Ghosal, T., Kumar, S., Bharti, P. K. & Ekbal, A. Peer review analyze: A novel benchmark resource for computational analysis of peer reviews. Plos one 17, e0259238 (2022)

  22. [30]

    & Falk Delgado, A

    Falk Delgado, A., Garretson, G. & Falk Delgado, A. The language of peer review reports on articles published in the bmj, 2014–2017: an observational study. Scientometrics 120, 1225–1235 (2019)

  23. [31]

    Analysis of the language used in the reports of peer-review journals

    Al-Khasawneh, F. Analysis of the language used in the reports of peer-review journals. Applied Research on English Language 11, 79–94 (2022)

  24. [32]

    Textual features of peer review predict top-cited papers: An interpretable machine learning perspective

    Sun, Z. Textual features of peer review predict top-cited papers: An interpretable machine learning perspective. Journal of Informetrics 18, 101501 (2024)

  25. [33]

    Mehrabian, A. et al. Silent messages, vol. 8 (Wadsworth Belmont, CA, 1971)

  26. [34]

    W.,(1997)

    Picard, R. W.,(1997). affective computing (1997). 18

  27. [35]

    N., Langerhuizen, D

    Steffens, A. N., Langerhuizen, D. W., Doornberg, J. N., Ring, D. & Janssen, S. J. Emotional tones in scientific writing: comparison of commercially funded studies and non-commercially funded ortho- pedic studies. Acta Orthopaedica 92, 240–243 (2021)

  28. [36]

    Ramachandran, L., Gehringer, E. F. & Yadav, R. K. Automated assessment of the quality of peer reviews using natural language processing techniques. International Journal of Artificial Intelligence in Education 27, 534–581 (2017)

  29. [37]

    A., Jawaid, M

    Jawaid, S. A., Jawaid, M. & Jafary, M. H. Characteristics of reviewers and quality of reviews: a retrospective study of reviewers at pakistan journal of medical sciences. Pakistan J Med Sci 22, 101–6 (2006)

  30. [38]

    Interactive functions of language in peer reviews of medical papers written by non- native users of english

    Kourilov ´a, M. Interactive functions of language in peer reviews of medical papers written by non- native users of english. Unesco ALSED-LSP Newsletter 19, 4–21 (1996)

  31. [39]

    & Diani, G

    Hyland, K. & Diani, G. Introduction: Academic evaluation and review genres. In Academic evalua- tion: Review genres in university settings, 1–14 (Springer, 2009)

  32. [40]

    & Mckeachie, W

    Kleinsasser, R. & Mckeachie, W. Teaching tips: Strategies, research, and theory for college and university teachers. The Modern Language Journal 78, 545 (2011)

  33. [41]

    Nicol, D. J. & Macfarlane-Dick, D. Formative assessment and self-regulated learning: A model and seven principles of good feedback practice. Studies in higher education 31, 199–218 (2006)

  34. [42]

    & Timperley, H

    Hattie, J. & Timperley, H. The power of feedback.Review of educational research77, 81–112 (2007)

  35. [43]

    & Molloy, E

    Boud, D. & Molloy, E. Feedback in higher and professional education. Understanding It and Doing It Well 2013 (2013)

  36. [44]

    Differing perceptions in the feedback process

    Carless, D. Differing perceptions in the feedback process. Studies in higher education 31, 219–233 (2006)

  37. [45]

    R., James, R., Berghella, V

    Kern-Goldberger, A. R., James, R., Berghella, V . & Miller, E. S. The impact of double-blind peer review on gender bias in scientific publishing: a systematic review. American journal of obstetrics and gynecology 227, 43–50 (2022)

  38. [46]

    S., Knutsen, J

    McDowell, G. S., Knutsen, J. D., Graham, J. M., Oelker, S. K. & Lijek, R. S. Co-reviewing and ghostwriting by early-career researchers in the peer review of manuscripts. Elife 8, e48425 (2019)

  39. [47]

    The peer-review crisis (2022)

    Flaherty, C. The peer-review crisis (2022). URL https://www.insidehighered.com/ news/2022/06/13/peer-review-crisis-creates-problems-journals-and-scholars . Accessed: 2025-07-15

  40. [48]

    & Smits, J

    Huisman, J. & Smits, J. Duration and quality of the peer review process: the author’s perspective. Scientometrics 113, 633–650 (2017)

  41. [49]

    Drozdz, J. A. & Ladomery, M. R. The peer review process: past, present, and future. British Journal of Biomedical Science 81, 12054 (2024)

  42. [50]

    McIntosh, L. D. & Hudson Vitale, C. Safeguarding scientific integrity: A case study in examining manipulation in the peer review process. Accountability in Research 32, 195–213 (2025). 19

  43. [51]

    Self-archiving and license to publish

    Nature Communications. Self-archiving and license to publish. https://www.nature.com/ ncomms/editorial-policies/self-archiving-and-license-to-publish . Ac- cessed: 2023-10-15

  44. [52]

    Plos text and data mining

    PLOS. Plos text and data mining. https://api.plos.org/text-and-data-mining. html. Accessed: 2023-10-15

  45. [53]

    Genderize.io (2024)

    ApS, D. Genderize.io (2024). URL https://genderize.io/. Accessed: 2024-09-29

  46. [54]

    Onomastics, N. A. Namsor sas (2024). URL https://namsor.app/. Accessed: 2024-09-29

  47. [55]

    Education, T. T. H. World university rankings — times higher education (the) (2024). URL https://www.timeshighereducation.com/world-university-rankings. Ac- cessed: 2024-09-29

  48. [56]

    Limited, Q. Q. S. Qs universities rankings - top global universities and colleges (2024). URLhttps: //www.topuniversities.com/university-rankings. Accessed: 2024-09-29

  49. [57]

    Nature index (2024)

    Nature, S. Nature index (2024). URL https://www.nature.com/nature-index/. Ac- cessed: 2024-09-29

  50. [58]

    & Piwowar, H

    Priem, J. & Piwowar, H. Openalex: A fully-open index of scholarly works, authors, institutions, and more (2022). URL https://openalex.org. Accessed via API

  51. [59]

    AlShebli, B. et al. Beijing’s central role in global artificial intelligence research.Scientific reports 12, 21461 (2022)

  52. [60]

    A., Evans, J

    AlShebli, B., Memon, S. A., Evans, J. A. & Rahwan, T. China and the us produce more impactful ai research when collaborating together. Scientific Reports 14, 28576 (2024)

  53. [61]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Devlin, J. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)

  54. [62]

    Nature communications: Writing your report (2024)

    Communications, N. Nature communications: Writing your report (2024). URL https://www. nature.com/ncomms/for-reviewers/writing-your-report. Accessed: 2024-09- 29

  55. [63]

    manuscript appreciation

    PlosOne. Plosone: Guidelines for reviewers (2024). URL https://journals.plos.org/ plosone/s/reviewer-guidelines. Accessed: 2024-09-29. 20 Main Figures 29% 32% 26% 7% 3% 3% Race 69%, Western affiliated 31%, Eastern affiliated Gender Institution Ranking 59%, Non-White 41%, White...

  56. [64]

    Read the Review: Read the entire review carefully to fully understand its content and context

  57. [65]

    Contextual-Summarizing

    Identify Contextual-Summarizing Sentences: If a sentence provides context, summarizes, or reiterates the research without expressing any evaluative tone, label it as “Contextual-Summarizing”

  58. [66]

    Handle Sarcasm Carefully: If a sentence contains sarcasm, do not classify it based on its literal meaning

    Identify Evaluative Tones:For all other sentences, classify them according to the relevant tone category. Handle Sarcasm Carefully: If a sentence contains sarcasm, do not classify it based on its literal meaning. Instead, determine the true intent

  59. [67]

    Preserve Sentence Order: The review text must be returned exactly as it appears, with each sentence tagged individually

  60. [68]

    If two tones are equally present, list both, separated by a semicolon

    Determine Dominant Tone: If a sentence exhibits multiple tones, assign the most dominant one. If two tones are equally present, list both, separated by a semicolon

  61. [69]

    Sincerely, Dr. Smith

    Detect and Label Signatures: • If the last sentence of the review contains only a name, title, or polite closing phrase (e.g., “Sincerely, Dr. Smith”), assign the tag “Name”. • This includes sentences that contain titles such as “Dr.”, “Professor”, or affiliations like “NYU, H...

  62. [70]

    Contextual-Summarizing

    Return Output as JSON: • Each sentence is a key. • The assigned tone is the value. • If a sentence is Contextual-Summarizing, explicitly label it as “Contextual-Summarizing”, • If a sentence does not match any tone, label it as “None”, • If the last sentence is a signature, la...

  63. [71]

    Constructive-Analytical: Provides a detailed examination of the study’s methodology, data, and results, offering targeted suggestions for improvement

  64. [72]

    Appreciative: This tone expresses genuine gratitude or approval, recognizing the strengths or contributions of the study

  65. [73]

    Critical-Evaluative: This tone highlights significant flaws, weaknesses, or limitations in the study, focusing on areas where the research does not meet standards or expectations

  66. [74]

    sentence 1

    Questioning: This tone seeks further information or clarification about specific aspects of the study, indicating a need for more detail. Output Format: JSON { tagged sentences:{ “sentence 1”: “Tone 1”, “sentence 2”: “Tone 2”, “sentence 3”: “Tone 3” } } 32 Peer Review Example ...

  67. [75]

    (1) and constant in eq

    Where is the multiplicative factor alpha in eq. (1) and constant in eq. (2) ?

  68. [76]

    In the result section p5, line 10. Did you multiply the normalized spectrum of the light source with the intensity ratio between the measured spectral profiles of the light source and the sample or with the intensity ratio between the normalized spectral profiles of light sour...

  69. [77]

    the question is, if the absorption of a certain tissue is not that small at 800nm, does the method still work? if not, how to solve this problem?

    The wavelength 800 nm was selected as the band with low absorption to verify the feasibility of the proposed method. the question is, if the absorption of a certain tissue is not that small at 800nm, does the method still work? if not, how to solve this problem?

  70. [78]

    In figure 5a, there is no much difference in darkness between rgb images of gt and sb? in contrast, gt and sb looks similar

  71. [79]

    In figure 5a, six colored rectangular boxes are obscure. 39 This hybrid weighting ensures that both short, high-impact tone cues and longer, detailed feedback are fairly represented, offering a more accurate and nuanced measure of the tone conveyed in peer review texts. 40 Sup...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.