{"id":"76b28bb2-5152-4249-ba26-c00b50301847","arxiv_id":"2507.00103","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Grant peer review report length, content focus, and sentiment vary systematically by discipline and by the gender of reviewers and applicants.","lead":"Analyzing nearly 40,000 Swiss grant review reports, the study finds that discipline and gender measurably shape the length, content, and tone of reviews. It shows SSH reviews are longer and more critical, MINT reviews more concise and positive, and female reviewers write longer, more positive reports.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Differential classifier measurement error across disciplines and genders is the key threat: aggregate F1 scores do not establish that the MINT/SSH and female-reviewer differences in content and sentiment are real rather than artifacts of label bias.","rationale":"The paper has genuine strengths: a large corpus, locally fine-tuned transformer models, openly shared code and models, explicit discussion of limitations, and robustness checks including double machine learning and grade-adjusted models. The review-length findings are direct and credible. However, the content and tone findings, which carry the title's central claim, depend on the classifiers measuring the same construct across research domains and reviewer/applicant groups. The aggregate F1 values do not rule out differential measurement error, and the low class-specific F1 for rarer categories such as Suitability of Methods and Negative sentiment makes the risk concrete. I considered the causal language in the abstract as an alternative concern, but the authors explicitly state the analysis is observational and correlational, so the more specific and testable threat is measurement bias in the text-derived outcomes. The reader's conditional verdict is appropriate: the concern is substantial enough to require additional analysis but not, on current evidence, sufficient to reject the paper. If the proposed stratified performance check and measurement-error correction show the coefficients survive, the conditional verdict can stand or be upgraded; if they do not, the content/tone claims would need to be substantially weakened. For now, the reader's conditional verdict remains the right recommendation.","tokens_in":21636,"tokens_out":6096,"duration_ms":79222,"concrete_test":"Re-annotate or use the existing 3,000-sentence gold-standard annotations to compute per-domain (LS/MINT/SSH) and per-reviewer-gender/per-applicant-gender precision, recall, and F1 for each of the six classifiers, using the same five-fold cross-validation protocol. Then re-estimate the Table 3 regressions after applying a measurement-error correction, such as stratum-specific misclassification matrix inversion or multiple imputation of true labels from predicted probabilities. Check whether the key coefficients—MINT vs LS Positive (4.65 pp), SSH Negative (1.59 pp), and Female Applicant Positive (0.50 pp)—remain materially unchanged in sign and significance. If stratified F1 differences are small and the corrected estimates match Table 3, the concern is resolved; if not, the central claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's novel content and tone outcomes are sentence-level classifier labels aggregated to review level (Methods: Text classification; Statistical analysis). The central comparisons—MINT reviews more positive and more track-record focused, SSH reviews more negative, female reviewers more positive—require that classification error be non-differential across research domains and reviewer/applicant gender. The reported macro F1 values (0.79–0.91) and Table S6 are aggregate; class-specific F1 is low for the rarer classes (Suitability of Methods F1=0.62, Negative F1=0.71), and no performance is reported separately by domain or gender. If, for example, the Positive classifier labels MINT sentences as positive more often because of field vocabulary while missing hedged praise in SSH, the 4.65 percentage-point MINT-LS positive difference and the smaller 0.50 percentage-point female-applicant positive effect could be inflated or even created. The mixed-effects models treat predicted labels as observed outcomes, so classifier uncertainty is not propagated into coefficients or confidence intervals; the applicant-female positive confidence interval [0.02, 0.97] is close to zero and would not survive even modest differential error. This is the load-bearing assumption for the content/tone half of the central claim, which is the paper's novel contribution; review length alone is measured directly and is not at issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper analyzes 39,280 English-language peer review reports submitted to the Swiss National Science Foundation between 2016 and 2023, covering 11,385 proposals across Social Sciences and Humanities (SSH), Life Sciences (LS), and Mathematics, Informatics, Natural Sciences and Technology (MINT). Using six fine-tuned transformer classifiers applied to over 1.3 million sentences, the authors measure the prevalence of four evaluation criteria (Track Record, Relevance/Originality/Topicality, Suitability of Methods, Feasibility) and two sentiment categories (Positive, Negative). They then estimate mixed-effects regressions relating these review-level outcomes to reviewer gender, applicant gender, and research domain, with a set of controls and robustness checks including interaction models, grade-adjusted models, and partially linear double-machine-learning models. The central claims are that SSH reviews are longer and more critical with less focus on track record, MINT reviews are more concise, positive, and track-record-focused, female reviewers write longer, more criteria-aligned, and more positive reviews, and female applicants receive slightly more positive sentiment and more methodological comments.","tokens_in":21901,"tokens_out":4248,"duration_ms":48397,"significance":"If the results hold, the paper makes a valuable contribution to the under-studied area of grant peer review by moving beyond numerical scores to the actual textual content and tone of reports. Its strengths include a large, near-complete dataset from a national funding agency; transparent, locally run machine-learning methods with publicly shared models and code; human annotation with reported inter-coder agreement; and robustness analyses via double machine learning with clustering and Bonferroni correction. The descriptive findings across 21 disciplines are internally consistent and plausible, and the review-length results, which are measured directly rather than through classifiers, are on solid ground. However, the novel content and tone findings depend on classifier outputs whose measurement error properties are not fully established, and at least one headline applicant-gender effect is fragile under a stated robustness specification. These issues are addressable and do not undermine the overall direction of the results, but they need to be resolved before the content/tone claims can be taken as fully supported.","major_comments":[{"comment":"The content and sentiment outcomes are sentence-level classifier predictions aggregated to the review level and then used as outcomes in mixed-effects models. The reported performance metrics are aggregate (Table S6), with class-specific F1 values as low as 0.62 for Suitability of Methods and 0.71 for Negative; no performance is reported separately by research domain or by reviewer/applicant gender. If classification error is differential across these groups, the estimated coefficients in Table 3—particularly the MINT-vs-LS difference in Positive (4.65 percentage points) and the marginal applicant-female Positive effect (0.50, CI [0.02, 0.97])—could be inflated or created. The paper's own limitation statement in the Discussion acknowledges that machine learning models 'may struggle to capture subtle or discipline-specific expressions of critique or praise.' Because this is the central evidence for the content/tone claims, the authors should report stratified held-out classification performance by domain and gender, and/or conduct sensitivity analyses that assume plausible differential misclassification rates, and propagate classifier uncertainty into the regression estimates.","section":"Materials and Methods: Text classification; Statistical analysis"},{"comment":"The abstract's claim that female applicants receive reviews with slightly more positive sentiment is not robust under the partially linear double-machine-learning specification: the applicant-female coefficient for Positive is 0.49 with a Bonferroni-corrected 95% CI of [-0.09, 1.08], which includes zero. The main mixed-effects estimate in Table 3 is also close to zero (0.50, CI [0.02, 0.97]). The manuscript should either temper this part of the central claim or explicitly explain why the linear model is preferred and why the DML interval should not alter the conclusion.","section":"Supplementary Materials, Table S5"}],"minor_comments":[{"comment":"The label 'Ethnology and and Social Geography' contains a duplicated 'and' in both the main figure and the supplementary figure; please correct this typo.","section":"Fig. 1 and Fig. S3"},{"comment":"The linear rescaling of the nine-point grade scale to a six-point scale should be justified or subjected to a sensitivity analysis, since the two scales may not be interval-equivalent; Fig. S2 shows the raw scales but not how the rescaling affects the grade coefficient in Table S4.","section":"Materials and Methods: Robustness tests"},{"comment":"The statement that macro F1 scores 'ranged from 0.79 to 0.91, demonstrating good reliability' should be qualified in the main text, because the class-specific F1 for Suitability of Methods (0.62) and Negative (0.71) in Table S6 is considerably lower and may affect the interpretation of the corresponding regression coefficients.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"This is a transparent and well-executed observational study, but the central content and tone findings rest on classifier outputs whose differential measurement error has not been assessed. I recommend asking the authors to supply stratified classifier performance by domain and gender, or equivalent sensitivity analyses, and to address the fragility of the applicant-female positive-sentiment effect under the DML specification. The companion methodology paper [21] should be made available to reviewers with the relevant performance details. The data-sharing restrictions are understandable, but the public availability of models and code is a positive factor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a genuinely useful contribution. It takes a large, real corpus of grant peer review reports and does something nobody has done at this scale: it measures not just scores but what reviewers actually write about—track record, methods, feasibility, and positive/negative tone. The length results are rock solid and interesting on their own. The content and sentiment results are plausible and the analysis is generally careful: mixed-effects models, a full set of controls, double-machine-learning robustness checks, and the code and models are public.\n\nThe soft spot is the one the stress test flags, and I think it's the right one. The classifiers' labels are treated as if they were observed outcomes, but the classifiers have real error, especially for rare classes like Suitability of Methods (F1=0.62) and Negative (F1=0.71). The paper reports only aggregate performance, so we don't know whether error is differential across disciplines or reviewer/applicant gender. If, say, MINT reviews use vocabulary that the sentiment classifier overscores, the 4.65 pp MINT-LS positive gap could be partly a measurement artifact. That is a real threat to the content/tone half of the central claim.\n\nBut I would not call it fatal. The gaps are large, the directions line up with prior work on numeric scores, and the authors already acknowledge that discipline-specific expression is hard to capture. The applicant-female positive effect is the exception: 0.50 pp with a CI that just barely excludes zero. That particular finding should be described as weak evidence, and the DML version actually loses significance. The authors should be asked to stratify classifier performance by domain and gender and to propagate classification uncertainty in the regressions. That is a fixable problem, not a reason to desk-reject.\n\nThe paper is honest about its limitations—English-only, binary gender, observational, no panel outcome. That helps. I would send it to referees. The field needs this kind of descriptive work on grant review text; it will be cited.","headline":"Solid large-scale descriptive study of grant review text; the classifier measurement-error concern is real but fixable, and the paper deserves peer review.","tokens_in":22394,"tokens_out":1861,"would_cite":true,"duration_ms":21544,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A text analysis of 39,280 grant reviews shows that report length, content focus, and tone vary systematically with research discipline and with reviewer and applicant gender.","keywords":["grant peer review","gender differences","disciplinary culture","sentiment analysis","supervised machine learning","text classification","Swiss National Science Foundation","peer review bias"],"falsifier":"Re-annotate a stratified random sample of review sentences balanced by discipline, reviewer gender, and applicant gender, and compare human labels with the classifier labels; if classification accuracy varies across these strata in the directions of the reported effects, the disciplinary and gender differences may be artifacts. Alternatively, a randomized experiment that blinds reviewers to applicant gender would test whether the female-applicant tone difference survives.","tokens_in":21459,"feed_emoji":"📝","tokens_out":5398,"duration_ms":56865,"temperature":0.7,"pith_summary":"The paper tries to establish that how a grant proposal is reviewed in writing—how long the report is, which criteria it emphasizes, and how positive or critical its tone is—depends systematically on the research discipline and on the genders of the reviewer and applicant. It does this by classifying more than 1.3 million sentences from 39,280 English-language review reports submitted to the Swiss National Science Foundation, then modeling the resulting content and sentiment measures with mixed-effects regressions. If the claim is right, review reports are not a neutral yardstick across fields: the same proposal could be described differently depending on disciplinary conventions and reviewer demographics, which matters for fairness in funding decisions.","feed_headline":"39,000 grant reviews show how field and gender shape reviews","feed_subtitle":"Discipline and reviewer gender visibly change how grant reports are written.","key_machinery":"The argument runs on a measurement pipeline: 39,280 English-language review reports from the Swiss National Science Foundation (2016–2023), split into 1,304,621 sentences; six fine-tuned transformer classifiers (built from SPECTER2, a transformer model pre-trained on scientific texts) that assign each sentence to evaluation-criteria and sentiment categories, trained on 3,000 human-annotated sentences; aggregation of sentence labels to review-level prevalence; and mixed-effects linear regressions with proposal-level random intercepts, with double machine learning used to relax linearity assumptions.","core_discovery":"Gender and disciplinary culture shape the length, content, and tone of grant peer review reports. Compared with Life Sciences reviews, Social Sciences and Humanities reviews are longer and more critical, with less attention to the applicant's track record, while Mathematics, Informatics, Natural Sciences and Technology reviews are more concise, more positive, and put more weight on the track record and feasibility. Female reviewers write longer reviews than male reviewers, cover the funder's evaluation criteria more closely, and express more positive sentiment; female applicants receive slightly more positive and slightly less negative sentiment than male applicants. The differences survive adjustment for a wide set of applicant, proposal, call, and reviewer covariates, including in partially linear double-machine-learning robustness checks.","pith_inferences":["If the sentiment and content classifiers have unequal accuracy across disciplines or genders, part of the reported gaps could be measurement rather than true difference; stratifying classifier evaluation by field and reviewer gender would test this directly.","The patterns imply that review reports are shaped by conventions that predate the individual reviewer, so reviewer training or structured checklists could plausibly reduce between-discipline and between-gender variation.","A natural extension would connect report text to panel outcomes, testing whether the longer, more critical SSH style or the shorter, positive MINT style translates into different funding rates at equal grades.","Because the data cover only English-language reports and binary gender, extending the analysis to other languages and to non-binary gender would show whether the patterns hold outside this setting."],"forward_implications":["Review reports are not directly comparable across disciplines, so multidisciplinary panels comparing reports on different proposals face systematic differences in length, emphasis, and tone.","Funding agencies could use these patterns to calibrate reviewers or design structured report formats that encourage more comparable coverage of evaluation criteria.","The finding that female reviewers align more closely with formal criteria and express more positive sentiment suggests that the gender composition of reviewer pools can shape how a proposal is characterized.","The slightly more positive reviews received by female applicants do not by themselves show bias in funding outcomes, because the study does not link report text to panel decisions or success rates.","Disciplinary conventions operationalize merit differently: MINT reviews emphasize track record and feasibility, while LS and SSH reviews emphasize methods and relevance or originality."],"supporting_citations":[{"why":"Supplies the annotation codebook, fine-tuned transformer classifiers, and classification performance metrics that the entire analysis rests on.","marker":"[21]"},{"why":"Provides the prior SNSF score-based analysis whose disciplinary patterns the sentiment results align with.","marker":"[25]"},{"why":"Motivates sentiment analysis of grant review reports and offers a comparison for how sentiment relates to rejection.","marker":"[14]"},{"why":"Establishes the sentence-level content classification and keyness analysis approach reused here.","marker":"[13]"},{"why":"Supplies the double machine learning estimator used in the robustness checks that relax linearity assumptions.","marker":"[23]"},{"why":"Provides the SPECTER2 transformer model pre-trained on scientific texts that the six classifiers were fine-tuned from.","marker":"[35]"},{"why":"Informs the comparison of methodological emphasis across fields in journal peer review, which the paper contrasts with its grant review findings.","marker":"[29]"}],"fun_headline_variants":["Grant reviews: discipline and gender alter length, tone, and focus","Female reviewers write longer, more positive grant reports","Grant reviews: SSH longer and critical, MINT concise and positive","Field and gender reshape grant peer review reports","Discipline and reviewer gender shape grant report tone"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The sentence-level machine learning classifiers measure content and sentiment equally well across disciplines and across male and female reviewers and applicants; if their error rates differ by group, the reported gaps could be measurement artifacts rather than real differences.","fun_headline_variants_meta":{"raw":{"variants":["Grant reviews: discipline and gender alter length, tone, and focus","Female reviewers write longer, more positive grant reports","Grant reviews: SSH longer and critical, MINT concise and positive","Field and gender reshape grant peer review reports","Discipline and reviewer gender shape grant report tone"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001117,"raw_usage":{"total_tokens":4620,"prompt_tokens":888,"completion_tokens":3732,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":3654}},"tokens_in":504,"tokens_out":3732,"duration_ms":29147,"temperature":1.0,"reasoning_tokens":3654,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:26:46.783421+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a stratified random sample of review sentences balanced by discipline, reviewer gender, and applicant gender, and compare human labels with the classifier labels; if classification accuracy varies across these strata in the directions of the reported effects, the disciplinary and gender differences may be artifacts. Alternatively, a randomized experiment that blinds reviewers to applicant gender would test whether the female-applicant tone difference survives.","supporting_citations":[{"cited_title":"& M¨ uller, S","cited_arxiv_id":null,"evidence_quote":"Supplies the annotation codebook, fine-tuned transformer classifiers, and classification performance metrics that the entire analysis rests on."},{"cited_title":"& Egger, M","cited_arxiv_id":null,"evidence_quote":"Provides the prior SNSF score-based analysis whose disciplinary patterns the sentiment results align with."},{"cited_title":"& Shankar, K","cited_arxiv_id":null,"evidence_quote":"Motivates sentiment analysis of grant review reports and offers a comparison for how sentiment relates to rejection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the sentence-level content classification and keyness analysis approach reused here."},{"cited_title":"& Robins, J","cited_arxiv_id":null,"evidence_quote":"Supplies the double machine learning estimator used in the robustness checks that relax linearity assumptions."},{"cited_title":"& Feldman, S","cited_arxiv_id":null,"evidence_quote":"Provides the SPECTER2 transformer model pre-trained on scientific texts that the six classifiers were fine-tuned from."},{"cited_title":"A., Bakker, M","cited_arxiv_id":null,"evidence_quote":"Informs the comparison of methodological emphasis across fields in journal peer review, which the paper contrasts with its grant review findings."}],"review_version":1}