REVIEW 2 major objections 3 minor 43 references
Gender and Discipline Shape Length, Content and Tone of Grant Peer Review Reports
T0 review · 2 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A text analysis of 39,280 grant reviews shows that report length, content focus, and tone vary systematically with research discipline and with reviewer and applicant gender.
desk verdict Solid large-scale descriptive study of grant review text; the classifier measurement-error concern is real but fixable, and the paper deserves peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument runs on a measurement pipeline: 39,280 English-language review reports from the Swiss National Science Foundation (2016–2023), split into 1,304,621 sentences; six fine-tuned transformer classifiers (built from SPECTER2, a transformer model pre-trained on scientific texts) that assign each sentence to evaluation-criteria and sentiment categories, trained on 3,000 human-annotated sentences; aggregation of sentence labels to review-level prevalence; and mixed-effects linear regressions with proposal-level random intercepts, with double machine learning used to relax linearity assumptions.
What would settle it
Re-annotate a stratified random sample of review sentences balanced by discipline, reviewer gender, and applicant gender, and compare human labels with the classifier labels; if classification accuracy varies across these strata in the directions of the reported effects, the disciplinary and gender differences may be artifacts. Alternatively, a randomized experiment that blinds reviewers to applicant gender would test whether the female-applicant tone difference survives.
Extended reading notes
Core claim
Gender and disciplinary culture shape the length, content, and tone of grant peer review reports. Compared with Life Sciences reviews, Social Sciences and Humanities reviews are longer and more critical, with less attention to the applicant's track record, while Mathematics, Informatics, Natural Sciences and Technology reviews are more concise, more positive, and put more weight on the track record and feasibility. Female reviewers write longer reviews than male reviewers, cover the funder's evaluation criteria more closely, and express more positive sentiment; female applicants receive slightly more positive and slightly less negative sentiment than male applicants. The differences survive adjustment for a wide set of applicant, proposal, call, and reviewer covariates, including in partially linear double-machine-learning robustness checks.
Load-bearing premise
The sentence-level machine learning classifiers measure content and sentiment equally well across disciplines and across male and female reviewers and applicants; if their error rates differ by group, the reported gaps could be measurement artifacts rather than real differences.
Editorial extensions
If this is right
- Review reports are not directly comparable across disciplines, so multidisciplinary panels comparing reports on different proposals face systematic differences in length, emphasis, and tone.
- Funding agencies could use these patterns to calibrate reviewers or design structured report formats that encourage more comparable coverage of evaluation criteria.
- The finding that female reviewers align more closely with formal criteria and express more positive sentiment suggests that the gender composition of reviewer pools can shape how a proposal is characterized.
- The slightly more positive reviews received by female applicants do not by themselves show bias in funding outcomes, because the study does not link report text to panel decisions or success rates.
- Disciplinary conventions operationalize merit differently: MINT reviews emphasize track record and feasibility, while LS and SSH reviews emphasize methods and relevance or originality.
Reading between the lines
- If the sentiment and content classifiers have unequal accuracy across disciplines or genders, part of the reported gaps could be measurement rather than true difference; stratifying classifier evaluation by field and reviewer gender would test this directly.
- The patterns imply that review reports are shaped by conventions that predate the individual reviewer, so reviewer training or structured checklists could plausibly reduce between-discipline and between-gender variation.
- A natural extension would connect report text to panel outcomes, testing whether the longer, more critical SSH style or the shorter, positive MINT style translates into different funding rates at equal grades.
- Because the data cover only English-language reports and binary gender, extending the analysis to other languages and to non-binary gender would show whether the patterns hold outside this setting.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper analyzes 39,280 English-language peer review reports submitted to the Swiss National Science Foundation between 2016 and 2023, covering 11,385 proposals across Social Sciences and Humanities (SSH), Life Sciences (LS), and Mathematics, Informatics, Natural Sciences and Technology (MINT). Using six fine-tuned transformer classifiers applied to over 1.3 million sentences, the authors measure the prevalence of four evaluation criteria (Track Record, Relevance/Originality/Topicality, Suitability of Methods, Feasibility) and two sentiment categories (Positive, Negative). They then estimate mixed-effects regressions relating these review-level outcomes to reviewer gender, applicant gender, and research domain, with a set of controls and robustness checks including interaction models, grade-adjusted models, and partially linear double-machine-learning models. The central claims are that SSH reviews are longer and more critical with less focus on track record, MINT reviews are more concise, positive, and track-record-focused, female reviewers write longer, more criteria-aligned, and more positive reviews, and female applicants receive slightly more positive sentiment and more methodological comments.
Significance. If the results hold, the paper makes a valuable contribution to the under-studied area of grant peer review by moving beyond numerical scores to the actual textual content and tone of reports. Its strengths include a large, near-complete dataset from a national funding agency; transparent, locally run machine-learning methods with publicly shared models and code; human annotation with reported inter-coder agreement; and robustness analyses via double machine learning with clustering and Bonferroni correction. The descriptive findings across 21 disciplines are internally consistent and plausible, and the review-length results, which are measured directly rather than through classifiers, are on solid ground. However, the novel content and tone findings depend on classifier outputs whose measurement error properties are not fully established, and at least one headline applicant-gender effect is fragile under a stated robustness specification. These issues are addressable and do not undermine the overall direction of the results, but they need to be resolved before the content/tone claims can be taken as fully supported.
major comments (2)
- [Materials and Methods: Text classification; Statistical analysis] The content and sentiment outcomes are sentence-level classifier predictions aggregated to the review level and then used as outcomes in mixed-effects models. The reported performance metrics are aggregate (Table S6), with class-specific F1 values as low as 0.62 for Suitability of Methods and 0.71 for Negative; no performance is reported separately by research domain or by reviewer/applicant gender. If classification error is differential across these groups, the estimated coefficients in Table 3—particularly the MINT-vs-LS difference in Positive (4.65 percentage points) and the marginal applicant-female Positive effect (0.50, CI [0.02, 0.97])—could be inflated or created. The paper's own limitation statement in the Discussion acknowledges that machine learning models 'may struggle to capture subtle or discipline-specific expressions of critique or praise.' Because this is the central evidence for the content/tone claims, the authors should report stratified held-out classification performance by domain and gender, and/or conduct sensitivity analyses that assume plausible differential misclassification rates, and propagate classifier uncertainty into the regression estimates.
- [Supplementary Materials, Table S5] The abstract's claim that female applicants receive reviews with slightly more positive sentiment is not robust under the partially linear double-machine-learning specification: the applicant-female coefficient for Positive is 0.49 with a Bonferroni-corrected 95% CI of [-0.09, 1.08], which includes zero. The main mixed-effects estimate in Table 3 is also close to zero (0.50, CI [0.02, 0.97]). The manuscript should either temper this part of the central claim or explicitly explain why the linear model is preferred and why the DML interval should not alter the conclusion.
minor comments (3)
- [Fig. 1 and Fig. S3] The label 'Ethnology and and Social Geography' contains a duplicated 'and' in both the main figure and the supplementary figure; please correct this typo.
- [Materials and Methods: Robustness tests] The linear rescaling of the nine-point grade scale to a six-point scale should be justified or subjected to a sensitivity analysis, since the two scales may not be interval-equivalent; Fig. S2 shows the raw scales but not how the rescaling affects the grade coefficient in Table S4.
- [Discussion] The statement that macro F1 scores 'ranged from 0.79 to 0.91, demonstrating good reliability' should be qualified in the main text, because the class-specific F1 for Suitability of Methods (0.62) and Negative (0.71) in Table S6 is considerably lower and may affect the interpretation of the corresponding regression coefficients.
Circularity Check
No circularity: outcomes are measured text properties, classifiers are validated on held-out annotations, and regression coefficients are estimated from data rather than imposed.
full rationale
The paper's derivation chain is self-contained. Review length is a direct word count. Content and sentiment outcomes are sentence-level classifier labels, aggregated to reviews, then regressed on reviewer/applicant gender and domain with proposal-level random intercepts. The classifiers were trained on 3,000 human-annotated sentences drawn from the same SNSF corpus; this shared corpus is a generalizability limitation but not circularity, because the training labels come from human annotation and held-out test performance (macro F1 0.79-0.91) is reported, with models publicly released. The regression is a conventional mixed-effects model; no coefficient is fitted to reproduce the paper's conclusions. Self-citations to Okasa et al. [21] are not load-bearing because that companion paper supplies the independently benchmarked classifier, which is externally checkable via the public HuggingFace models, and the external SNSF-score study [25] is prior independent work. The paper itself flags the residual risk that 'even well-performing machine learning models may struggle to capture subtle or discipline-specific expressions of critique or praise' (Discussion), which is a measurement-error caveat rather than a circular step. No equation defines an input in terms of an output, and no fitted parameter is renamed as a prediction. Aggregate classifier F1 scores do not prove error is equal across disciplines and genders, but that is a potential bias in the measurement, not a reduction of the claimed results to the model inputs.
Assumptions & free parameters
assumptions (5)
- standard math The mixed-effects model linearity and normality assumptions hold; the random intercept captures proposal-level correlation.
- domain assumption The binary self-reported gender categories (male/female) accurately represent the relevant gender identity for the analysis.
- domain assumption The machine learning classifiers provide unbiased prevalence estimates across disciplines and gender.
- domain assumption English-language reviews are representative of all reviews.
- ad hoc to paper The linear rescaling of the nine-point grade scale to a six-point scale preserves comparability.
Cite this review
Pith. "Pith review of Gender and Discipline Shape Length, Content and Tone of Grant Peer Review Reports." pith.science (2026). https://pith.science/paper/2KEWXOGG
@misc{pith2026250700103,
author = {Pith},
title = {Pith review of: Gender and Discipline Shape Length, Content and Tone of Grant Peer Review Reports},
year = {2026},
howpublished = {\url{https://pith.science/paper/2KEWXOGG}},
note = {Machine review of arXiv:2507.00103}
}
read the original abstract
Peer review by experts is central to the evaluation of grant proposals, but little is known about how gender and disciplinary differences shape the content and tone of grant peer review reports. We analyzed 39,280 review reports submitted to the Swiss National Science Foundation between 2016 and 2023, covering 11,385 proposals for project funding across 21 disciplines from the Social Sciences and Humanities (SSH), Life Sciences (LS), and Mathematics, Informatics, Natural Sciences, and Technology (MINT). Using supervised machine learning, we classified over 1.3 million sentences by evaluation criteria and sentiment. Reviews in SSH were significantly longer and more critical, with less focus on the applicant's track record, while those in MINT were more concise and positive, with a higher focus on the track record, as compared to those in LS. Compared to male reviewers, female reviewers write longer reviews that more closely align with the evaluation criteria and express more positive sentiments. Female applicants tend to receive reviews with slightly more positive sentiment than male applicants. Gender and disciplinary culture influence how grant proposals are reviewed - shaping the tone, length, and focus of peer review reports. These differences have important implications for fairness and consistency in research funding.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Bornmann, L. Scientific Peer Review. Annual Review of Information Science and Technology 45, 197–245 (2011)
work page 2011
-
[2]
Jerrim, J. & Vries, R. Are Peer Reviews of Grant Proposals Reliable? An Analysis of Economic and Social Research Council (ESRC) Funding Applications. The Social Science Journal 60, 91–109 (2023)
work page 2023
-
[3]
in Challenges in Research Policy (eds Sivertsen, G
Langfeldt, L. in Challenges in Research Policy (eds Sivertsen, G. & Langfeldt, L.) 29–36 (Springer Nature Switzerland, Cham, 2025)
work page 2025
-
[4]
Nicholls, R. D. Peer Review Under Review. Science 286, 1853–1853 (1999)
work page 1999
-
[5]
Langfeldt, L., Reymert, I. & Svartefoss, S. M. Distrust in Grant Peer Review— Reasons and Remedies. Science and Public Policy 51, 28–41 (2024)
work page 2024
-
[6]
Lee, C. J., Sugimoto, C. R., Zhang, G. & Cronin, B. Bias in Peer Review. Journal of the American Society for Information Science and Technology 64, 2–17 (2013)
work page 2013
-
[7]
Tamblyn, R., Girard, N., Qian, C. J. & Hanley, J. Assessment of Potential Bias in Research Grant Peer Review in Canada. CMAJ 190, E489–E499 (2018)
work page 2018
-
[8]
Huber, J., Inoua, S., Kerschbamer, R., Konig-Kersting, C. & Palan, S. Nobel and Novice: Author Prominence Affects Peer Review.Proceedings of the National Academy of Sciences of the United States of America 119, e2205779119 (2022)
work page 2022
Show all 43 references
-
[9]
W., Marsh, H
Jayasinghe, U. W., Marsh, H. W. & Bond, N. Peer Review in the Funding of Research in Higher Education: The Australian Experience. Educational Evaluation and Policy Analysis 23, 343–364 (2001)
2001
-
[10]
Cicchetti, D. V. The Reliability of Peer Review for Manuscript and Grant Submis- sions: A Cross-Disciplinary Investigation. Behavioral and Brain Sciences 14, 119– 135 (1991)
1991
-
[11]
Ghosal, T., Kumar, S., Bharti, P. K. & Ekbal, A. Peer Review Analyze: A Novel Benchmark Resource for Computational Analysis of Peer Reviews. PLoS One 17, e0259238 (2022)
2022
-
[12]
Verharen, J. P. H. ChatGPT Identifies Gender Disparities in Scientific Peer Review. eLife 12, RP90230 (2023)
2023
-
[13]
Severin, A., Strinzel, M., Egger, M., Barros, T., Sokolov, A., Mouatt, J. V. & M¨ uller, S. Relationship between Journal Impact Factor and the Thoroughness and Helpful- ness of Peer Reviews. PLoS Biology 21, e3002238 (2023)
2023
-
[14]
& Shankar, K
Luo, J., Feliciani, T., Reinhart, M., Hartstein, J., Das, V., Alabi, O. & Shankar, K. Analyzing Sentiments in Peer Review Reports: Evidence from Two Science Funding Agencies. Quantitative Science Studies 2, 1271–1295 (2021). 25
2021
-
[15]
A., Grant, S., Chen, M.-C., Lindner, M
Erosheva, E. A., Grant, S., Chen, M.-C., Lindner, M. D., Nakamura, R. K. & Lee, C. J. NIH Peer Review: Criterion Scores Completely Account for Racial Disparities in Overall Impact Scores. Science Advances 6, eaaz4868 (2020)
2020
-
[16]
Peer Review of Grant Applications: What Do We Know? The Lancet 352, 301–305 (1998)
Wessely, S. Peer Review of Grant Applications: What Do We Know? The Lancet 352, 301–305 (1998)
1998
-
[17]
& Van De Rijt, A
Bol, T., De Vaan, M. & Van De Rijt, A. The Matthew Effect in Science Funding. Proceedings of the National Academy of Sciences 115, 4887–4890 (2018)
2018
-
[18]
von, Wolf, T
Tunstall, L., Werra, L. von, Wolf, T. & G´ eron, A. Natural Language Processing with Transformers: Building Language Applications with HuggingFace (O’Reilly, Beijing, Boston, Farnham, Sebastopol, Tokyo, 2022)
2022
-
[19]
M., Dercksen, K., Dycke, N., Goldberg, A., Hope, T., Hovy, D., Kummerfeld, J
Kuznetsov, I., Afzal, O. M., Dercksen, K., Dycke, N., Goldberg, A., Hope, T., Hovy, D., Kummerfeld, J. K., Lauscher, A., Leyton-Brown, K., Lu, S., Mausam, Mieskes, M., N´ ev´ eol, A., Pruthi, D., Qu, L., Schwartz, R., Smith, N. A., Solorio, T., Wang, J., Zhu, X., Rogers, A., S...
2024 arXiv
-
[20]
Sentiment Analysis: Mining Opinions, Sentiments, and Emotions Second edition
Liu, B. Sentiment Analysis: Mining Opinions, Sentiments, and Emotions Second edition. (Cambridge university press, Cambridge New York, 2020)
2020
-
[21]
& M¨ uller, S
Okasa, G., de Le´ on, A., Strinzel, M., Jorstad, A., Milzow, K., Egger, M. & M¨ uller, S. A Supervised Machine Learning Approach for Assessing Grant Peer Review Reports 2024
2024
-
[22]
Cleavage Identities in Voters’ Own Words: Harnessing Open-ended Survey Responses
Zollinger, D. Cleavage Identities in Voters’ Own Words: Harnessing Open-ended Survey Responses. American Journal of Political Science 68, 139–159 (2024)
2024
-
[23]
& Robins, J
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W. & Robins, J. Double/Debiased Machine Learning for Treatment and Structural Parameters. The Econometrics Journal 21, C1–C68 (2018)
2018
-
[24]
& Wold, A
Wenneras, C. & Wold, A. Nepotism and Sexism in Peer-Review. Nature 387, 341– 343 (1997)
1997
-
[25]
& Egger, M
Severin, A., Martins, J., Heyard, R., Delavy, F., Jorstad, A. & Egger, M. Gender and Other Potential Biases in Peer Review: Cross-Sectional Analysis of 38 250 External Peer Review Reports. BMJ Open 10, e035058 (2020)
2020
-
[26]
Making Reviewers Visible: Openness, Accountability, and Credit
Godlee, F. Making Reviewers Visible: Openness, Accountability, and Credit. JAMA 287, 2762 (2002)
2002
-
[27]
What Is Open Peer Review? A Systematic Review
Ross-Hellauer, T. What Is Open Peer Review? A Systematic Review. F1000Research 6, 588 (2017). 26
2017
-
[28]
Squazzoni, F., Ahrweiler, P., Barros, T., Bianchi, F., Birukou, A., Blom, H. J. J., Bravo, G., Cowley, S., Dignum, V., Dondio, P., Grimaldo, F., Haire, L., Hoyt, J., Hurst, P., Lammey, R., MacCallum, C., Maruˇ si´ c, A., Mehmani, B., Murray, H., Nicholas, D., Pedrazzi, G., Pue...
2020
-
[29]
A., Bakker, M
Buljan, I., Garcia-Costa, D., Grimaldo, F., Klein, R. A., Bakker, M. & Maruˇ si´ c, A. Development and Application of a Comprehensive Glossary for the Identifica- tion of Statistical and Methodological Concepts in Peer Review Reports. Journal of Informetrics 18, 101555 (2024)
2024
-
[30]
Palmer, A., Smith, N. A. & Spirling, A. Using Proprietary Language Models in Academic Research Requires Explicit Justification. Nature Computational Science 4, 2–3 (2023)
2023
-
[31]
G., Norman, C
Hren, D., Pina, D. G., Norman, C. R. & Maruˇ si´ c, A. What Makes or Breaks Compet- itive Research Proposals? A Mixed-Methods Analysis of Research Grant Evaluation Reports. Journal of Informetrics 16, 101289 (2022)
2022
-
[32]
Enhancing Peer Review Skills in Higher Education: A Mixed-Methods Study on Challenges and Training Needs
Stupacher, J. Enhancing Peer Review Skills in Higher Education: A Mixed-Methods Study on Challenges and Training Needs. Open Science Framework. https://doi. org/10.31219/osf.io/89xju_v2 (2025)
2025 doi
-
[33]
& Bornmann, L
Olbrecht, M. & Bornmann, L. Panel Peer Review of Grant Applications: What Do We Know from Research in Social Psychology on Judgment and Decision-Making in Groups? Research Evaluation 19, 293–304 (2010)
2010
-
[34]
Cld2: Google’s Compact Language Detector 2 2025
Ooms, J. Cld2: Google’s Compact Language Detector 2 2025. https://doi.org/ 10.32614/CRAN.package.cld2
2025 doi
-
[35]
& Feldman, S
Singh, A., D’Arcy, M., Cohan, A., Downey, D. & Feldman, S. SciRepEval: A Multi- Format Benchmark for Scientific Document Representations in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (Association for Computational Linguistics, Singap...
2023
-
[36]
& Kl´ en, R
Rainio, O., Teuho, J. & Kl´ en, R. Evaluation Metrics and Statistical Tests for Machine Learning. Scientific Reports 14, 6086 (2024)
2024
-
[37]
Llama 3 Model Card 2024
AI@Meta. Llama 3 Model Card 2024. https : / / huggingface . co / meta - llama / Meta-Llama-3-8B-Instruct
2024
-
[38]
& Milzow, K
Strinzel, M., Okasa, G., Jorstad, A., M¨ uller, S., de Le´ on, A., Egger, M. & Milzow, K. Data Management Plan (DMP): A Supervised Machine Learning Approach for Assessing Grant Peer Review Reports 2024. https://doi.org/10.46446/DMP- peer-review-assessment-ML . 27
2024 doi
-
[39]
& Hill, J
Gelman, A. & Hill, J. Data Analysis Using Regression and Multilevel/Hierarchical Models (Cambridge University Press, Cambridge; New York, 2007)
2007
-
[40]
Robinson, P. M. Root-N-consistent Semiparametric Regression. Econometrica 56, 931 (1988)
1988
-
[41]
S., Chernozhukov, V., Spindler, M
Bach, P., Kurz, M. S., Chernozhukov, V., Spindler, M. & Klaassen, S. DoubleML: An Object-Oriented Implementation of Double Machine Learning in R. Journal of Statistical Software 108 (2024)
2024
-
[42]
Random Forests
Breiman, L. Random Forests. Machine Learning 45, 5–32 (2001)
2001
-
[43]
D., Kato, K., Ma, Y
Chiang, H. D., Kato, K., Ma, Y. & Sasaki, Y. Multiway Cluster Robust Dou- ble/Debiased Machine Learning. Journal of Business & Economic Statistics 40, 1046–1056 (2022). 28 Supplementary Materials Gender and Discipline Shape Length, Content and Tone of Grant Peer Review Reports...
2022
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.