{"id":"dde18a02-f216-4a6d-a1b5-6ad4d9169b05","arxiv_id":"2507.14741","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Using NLP on 80,000 peer reviews, the study shows tone disparities across author demographics and finds reviewer identity disclosure is associated with more balanced, appreciative language, though critique gaps remain.","lead":"This study analyzes more than 80,000 peer reviews from two journals and finds that review tone varies with author gender, race, region, and institution. It reports that when reviewers sign their names, some tone gaps shrink, but critical tone gaps persist.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that disclosure eliminates demographic tone disparities is not supported by the paper's own statistics: it relies on contrasting significance across separate subsample regressions rather than testing disclosure-by-demographic interactions.","rationale":"The reader's weakest assumption concerns self-selection into reviewer disclosure, which is a real threat to causal interpretation. My focus is different and more internal: even as an association, the paper's evidence for 'disappearing' disparities is methodologically insecure because it compares significance across separate subsample models rather than testing whether coefficients differ. This is the most load-bearing concern because the central claim about disclosure is built on exactly those contrasts. The paper has genuine strengths: a large corpus, explicit exclusions, a transparent prompt for GPT-4o, and two validation exercises for tone classification with reasonable agreement rates. However, the regression analysis does not support the strong 'disparities disappear under disclosure' interpretation without formal interaction tests. The descriptive demographic associations and the pooled disclosure effects may still be informative, so conditional acceptance with a required re-analysis is the appropriate stance; my concern does not change the reader's conditional verdict but sharpens the specific condition that must be met.","tokens_in":26811,"tokens_out":3656,"duration_ms":48631,"concrete_test":"Re-estimate the four pooled OLS tone models in Supplementary Table 3 with the addition of Reviewer identity disclosure × demographic interaction terms for gender, race, institutional rank, and region. For each tone category, report the four interaction coefficients with confidence intervals and a joint Wald test. If the joint p-value exceeds 0.05, the claim that disclosure significantly reduces demographic tone disparities is unsupported; if it is below 0.05 with coefficients in the expected direction, the claim gains support. This directly tests whether the anonymous versus disclosed coefficients differ, rather than relying on their separate significance levels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion that identity disclosure reduces demographic tone disparities is inferred from comparing coefficients across two separate regressions: one for anonymous reviews and one for disclosed reviews (Supplementary Table 4). In the full-data models, reviewer disclosure is a significant predictor, but no interaction terms between disclosure and author demographics are estimated or reported. Coefficients that are significant in the anonymous subsample and non-significant in the disclosed subsample are interpreted as 'disappearing' disparities. This is not a valid statistical comparison: non-significance in a smaller subsample does not imply that the effect is smaller. For example, for gender on appreciative tone, the anonymous coefficient is -0.283 (SE 0.107) and the disclosed coefficient is 0.033 (SE 0.224); the difference is about 0.316 with an approximate z of 1.27, so the two coefficients are not significantly different. Similar patterns hold for race on appreciation (0.360 vs 0.277, SEs 0.113 and 0.253) and region on appreciation (0.809 vs 0.366, SEs 0.120 and 0.267). The paper therefore does not currently provide statistical evidence that disparities shrink or disappear under disclosure; it only shows that some demographic coefficients are significant in the large anonymous sample and some are not in the smaller disclosed sample. This is a load-bearing gap because the headline policy claim about disclosure rests on those contrasts.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper examines peer review language in more than 80,000 first-round reviews of accepted manuscripts published in Nature Communications and PLOS One (2019–2024). It combines a fine-tuned SciBERT sentiment classifier, an n-gram analysis of evaluative phrases, and GPT-4o sentence-level tone classification to measure appreciative, constructive-analytical, questioning, and critical-evaluative tone. Tone scores are regressed on corresponding-author demographics (gender, race, region, institutional rank, academic age, productivity, prior impact), paper features, review features, and reviewer identity disclosure. The paper reports widespread demographic disparities in sentiment and tone, and argues that disclosed reviewer identity is associated with more appreciative and constructive language and with the disappearance of several demographic tone disparities, mainly in appreciative tone, while critical tone disparities persist. The Discussion draws policy implications in favor of reviewer disclosure.","tokens_in":27123,"tokens_out":5789,"duration_ms":73291,"significance":"If the descriptive associations are robust, the paper is a useful empirical contribution: it assembles a large, carefully documented corpus, validates GPT-4o tone labels against independent researchers (95.1%) and against the reviewed authors themselves (91.1%), publishes the full n-gram lexicon and survey instruments in the supplement, and uses Huber-White robust standard errors in the regressions. The quantitative linguistic description of demographic disparities in review tone is valuable in its own right. However, the central policy conclusion about reviewer disclosure is not supported by the statistical evidence as reported, because the anonymous-versus-disclosed comparison is confounded by self-selection and is based on informal contrasts across separate subsample regressions rather than formal interaction tests. The paper's significance therefore hinges on a revision that either supplies interaction tests and a defensible identification strategy or recalibrates the claims to descriptive associations.","major_comments":[{"comment":"The claim that demographic disparities in tone disappear under identity disclosure is inferred from comparing coefficients across two separate regressions, one for anonymous and one for disclosed reviews, in Supplementary Table 4. Non-significance in the smaller disclosed subsample does not imply that the coefficient is smaller than in the larger anonymous subsample. For example, for gender on appreciative tone, the anonymous coefficient is -0.283 (SE 0.107) and the disclosed coefficient is 0.033 (SE 0.224); the difference is about 0.316, giving an approximate z of 1.27, which is not statistically significant. The same problem affects the race comparison on appreciation (0.360 vs 0.277, SEs 0.113 and 0.253) and the region comparison (0.809 vs 0.366, SEs 0.120 and 0.267). The paper should estimate models with interaction terms between reviewer identity disclosure and the relevant author demographics, and report those interaction coefficients and their standard errors. Without such tests, the headline statement that \"disclosure appears to foster a more balanced and equitable evaluative environment\" is unsupported by the paper's own statistics.","section":"Results, \"Tone Variation across Author Groups\"; Supplementary Table 4"},{"comment":"The anonymous-versus-disclosed comparison is used to draw causal or quasi-causal conclusions about the effect of identity disclosure, but disclosure status is chosen by reviewers, not assigned. The paper has no reviewer-level covariates, no reviewer fixed effects, no instrument, and no random assignment; Table 1 reports only the counts of anonymous and signed reviews. If reviewers who sign are systematically more generous, more constructive, or more relationally connected to the authors, the observed tone differences may reflect selection rather than the consequences of disclosure. The sentence in the Discussion that \"reviewer identity disclosure appears to foster a more balanced and equitable evaluative environment\" overstates what an observational contrast can establish. The manuscript should either add a design that can address selection (for example, within-reviewer comparisons when the same reviewer reviews under both conditions, or reviewer-level propensity adjustments conditional on observable reviewer characteristics) or explicitly reframe the disclosure results as descriptive associations and soften the policy recommendations accordingly.","section":"Materials and Methods, Data (\"Reviewer anonymity\"); Discussion"},{"comment":"The sentiment analysis rests on a fine-tuned SciBERT classifier whose validation accuracy is 0.677 and weighted F1 is 0.671. The paper does not report class-wise precision, recall, or a confusion matrix, nor does it report how the two-sample proportion Z-tests in Figure 2b behave under plausible classification error rates or different decision thresholds. Since the sentiment findings are presented as evidence of demographic disparities in positive and negative review tone, the paper should either provide per-class diagnostics plus a sensitivity analysis demonstrating that the reported disparities are not artifacts of measurement error, or explicitly relegate the sentiment results to an exploratory supporting analysis. The strong claims in the Results about negative sentiment being \"more commonly directed toward\" specific demographic groups need stronger measurement support.","section":"Materials and Methods, \"Using SciBERT to classify sentiment\"; Results, \"Sentiment Differences Across Author Groups\""},{"comment":"There is an internal inconsistency in the interpretation of the regional coefficient. The text says that \"the tendency for reviewers to use more appreciative language when evaluating submissions from female, white, or eastern-affiliated corresponding authors disappears entirely when reviewer identity is disclosed,\" but Supplementary Table 4 and Figure 4 define the regional variable as \"Western group,\" with a positive coefficient for appreciative tone in the anonymous model (0.809, SE 0.120). This means Western-affiliated authors, not eastern-affiliated authors, receive more appreciative language under anonymous review. The Discussion paragraph states the direction correctly (\"Female, white, and western-affiliated authors were more likely to receive appreciative and positive sentiment\"), so the Results paragraph appears to contain a substantive reversal rather than a mere wording issue. This should be corrected and the surrounding interpretation checked for consistency.","section":"Results, \"Tone Variation across Author Groups\"; Supplementary Table 4"}],"minor_comments":[{"comment":"The definition of the \"Questioning\" tone in the consent-and-instructions text is identical to the definition of \"Constructive-Analytical\" (\"Provides a detailed examination of the study's methodology, data, or findings, offering suggestions for improvement\"). This likely affected how survey participants interpreted the tone labels and should be fixed in the supplement and in any replication materials.","section":"Supplementary Note 2, Survey content"},{"comment":"The weighted tone score uses alpha = 0.2, and the text states that any value between 0.1 and 0.4 would have been reasonable, but no sensitivity analysis is reported. A short robustness table showing how the main regression coefficients and their standard errors vary across alpha values in this range would address concerns about a hand-set free parameter.","section":"Materials and Methods, \"Weighted Tone Scoring Method\""},{"comment":"The n-gram and sentiment analyses run many two-group comparisons across demographic axes, age bins, and n-gram categories, each with its own significance asterisks. No multiple-comparison correction is applied. Reporting false-discovery-rate-adjusted p-values, or at least stating how many comparisons were made, would make the pattern of significant results easier to evaluate.","section":"Results, \"Evaluative Language Patterns across Author Groups\"; Materials and Methods, \"Identifying and Mapping N-grams\""},{"comment":"The figure caption says the y-axis shows \"subgroup contribution within sentiment,\" while the main text describes the y-axis as the proportion of reviews within each subgroup falling into a given sentiment category. These are different quantities; the caption and axis labels should be aligned with the actual computation.","section":"Figure 2 caption and labels"},{"comment":"Table 1 reports 80,187 reviews, but the full-data regressions in Supplementary Table 3 use 76,644 observations. The manuscript should state how many reviews were dropped due to missing tone scores or other features, so the discrepancy is transparent.","section":"Table 1 and regression sample sizes"},{"comment":"The article does not include a data availability or code availability statement. Given the policy-oriented nature of the claims and the unusual data assembly process, a clear statement about whether the review texts, derived features, and analysis code will be shared would strengthen reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful descriptive core and unusually thorough validation of the LLM tone classifications, but the central causal claim about reviewer disclosure is not currently supported by the reported statistics. The most important fix is the formal interaction test between disclosure and author demographics; the self-selection problem is also fundamental and should be addressed head-on or the claims should be explicitly downgraded to associational. I do not see evidence of fabrication or misconduct, and the missing analyses appear feasible with the existing data, so major revision rather than rejection seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a large and genuinely useful descriptive study, but its central policy claim—that reviewer identity disclosure shrinks demographic tone disparities—is not supported by the paper's own regressions. The authors compare coefficients that are significant in the anonymous subsample and non-significant in the disclosed subsample, then call that \"disappearing.\" That is not a valid test of a difference. For gender on appreciative tone, the anonymous coefficient is -0.283 (SE 0.107) and the disclosed is 0.033 (SE 0.224); the difference is about 0.3 with z around 1.3. Race and region contrasts are similarly non-significant. So the headline result about disclosure reducing disparities is, statistically, not there.\n\nWhat the paper does well: the scale (80,187 reviews from two journals), the three-tier NLP pipeline, and especially the tone validation. They had corresponding authors rate over 21,000 sentences (91.1% agreement) and impartial researchers rate a smaller set (95.1%). They also validated the gender and race classifiers against self-reports. Those are real strengths. The descriptive patterns—non-white and eastern-affiliated authors receive more critical and constructive language, white and western authors more appreciative language—are internally consistent and worth reporting.\n\nThe self-selection into signing is a second major problem. Reviewers choose whether to sign; the paper has no instrument, no reviewer fixed effects, no reviewer-level covariates. Kind or constructive reviewers may simply be more likely to sign. The full-data association between disclosure and more appreciative tone is fine as an association, but the Discussion's causal language (\"disclosure appears to foster a more balanced and equitable evaluative environment\") is not earned. Smaller issues: sentiment classifier accuracy is modest (0.677 validation accuracy, F1 0.671), no clustering by manuscript, many Z-tests without multiple-comparison correction, and no data or code release.\n\nWho this is for: editors and scientometrics researchers, plus NLP people working on review text. It deserves a serious referee because the data and descriptive core are valuable and the validation approach is exemplary. But the central claim needs reanalysis with actual interaction tests (disclosure × demographics) and ideally some attempt to address selection—reviewer fixed effects or matched comparisons would help. The causal language should be softened to associations. I would send it to peer review with the expectation of major revision, not desk-reject it.","headline":"A large, well-validated descriptive study of tone disparities in peer review, but the headline claim that identity disclosure reduces those disparities is not supported by the paper's own statistics.","tokens_in":27619,"tokens_out":2380,"would_cite":true,"duration_ms":31718,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that reviewer tone in first-round reviews of accepted papers varies with corresponding authors' gender, race, region, and institutional rank, and that signing reviews narrows appreciative and constructive gaps but not…","keywords":["peer review","reviewer anonymity","open peer review","linguistic bias","sentiment analysis","tone classification","n-gram analysis","large language models"],"falsifier":"A decisive test would be a randomized or policy-forced comparison in which the same reviewers evaluate comparable manuscripts under anonymous and signed conditions, for example by tracking a journal's switch to mandatory open review; if tone and demographic gaps are identical across conditions, the paper's claim that disclosure reduces disparities would be refuted, while persistence of critical gaps under mandatory signing would support it.","tokens_in":26595,"feed_emoji":"📝","tokens_out":6264,"duration_ms":72744,"temperature":0.7,"pith_summary":"The paper examines more than 80,000 first-round peer reviews of accepted papers from two open-access journals and asks whether the language of evaluation differs by author demographics and by whether reviewers sign their names. It claims that female, white, and western-affiliated corresponding authors receive more appreciative and positive language, especially under anonymous review, while non-white and eastern-affiliated authors receive more critical and negative language and more correction-oriented feedback. Comparing anonymous and signed reviews, the study reports that disclosure is associated with a marked reduction in appreciative and constructive tone gaps, while critical-evaluative gaps persist. A sympathetic reader would care because these findings bear directly on policy debates about anonymous versus open peer review: they suggest that signing reviews may dampen some identity-driven tone differences without eliminating them.","feed_headline":"Signed reviews shrink author-identity tone gaps, study says","feed_subtitle":"An 80,000-review analysis ties anonymous review to uneven appreciation and criticism across author demographics.","key_machinery":"The analysis is carried by a three-tier linguistic pipeline and a regression framework. First, a SciBERT classifier fine-tuned on peer review sentences assigns positive, neutral, or negative sentiment. Second, a curated lexicon of 269 bigrams and trigrams, grouped into manuscript appreciation, requests for clarification, constructive suggestions, and strong methodological critique, tracks recurring evaluative phrasing normalized per 1,000 words. Third, GPT-4o classifies each sentence into appreciative, constructive-analytical, questioning, critical-evaluative, or contextual-summarizing tones, validated against authors (91.1% agreement) and independent researchers (95.1% agreement); a weighted tone score combines sentence frequency and sentence length. Ordinary least squares regressions with robust standard errors, field-of-science fixed effects, and separate models for anonymous and disclosed reviews then estimate how each tone varies with gender, race, region, institutional rank, academic age, and reviewer disclosure.","core_discovery":"The central claim is that reviewer tone in first-round reviews of ultimately accepted papers is measurably associated with the demographics and institutional standing of the corresponding author, and that reviewer identity disclosure changes the pattern. In the paper's own terms, female, white, and western-affiliated authors were more likely to receive appreciative and positive sentiment, especially under anonymous review, whereas non-white and eastern-affiliated authors were more frequently met with critical and negative language and with constructive or clarifying suggestions. When reviewers disclosed their identities, the elevated appreciation for white, female, and western authors largely disappeared and several n-gram level gaps in critique, constructiveness, and clarification also lost significance; however, critical-evaluative tone remained higher for eastern-affiliated and non-top-100 authors, and constructive-analytical tone toward eastern-affiliated authors persisted. The paper reads this as evidence that disclosure moderates extremes in tone and encourages more balanced feedback, without evidence that disclosure makes reviews harsher.","pith_inferences":["The anonymous-versus-signed comparison is observational, so the paper cannot rule out that reviewers who sign are self-selected and systematically different in generosity or writing style; the directional conclusion about disclosure should be read with that caveat.","If the critical-tone gap is driven by cues other than identity, such as topic, methods, or writing style correlated with region and institution, then mandatory disclosure would not remove it; a testable extension is to randomize reviewers' signing status or exploit a journal's policy change.","The paper's focus on accepted papers means its estimates may understate total bias, since rejected manuscripts receive no public reviews; including rejected reviews from journals that publish them would test whether tone disparities intensify."],"forward_implications":["If the central claim holds, anonymous single-blind review of accepted papers already contains measurable identity-linked differences in evaluative language, even in the subset of reviews where bias should be lowest.","Signing reviews would plausibly narrow appreciative and constructive tone gaps, partly because disclosed reviews use more appreciative, constructive, and questioning language overall.","Critical-evaluative disparities for eastern-affiliated and non-top-100 authors would persist under disclosure, so identity transparency alone would not equalize evaluation.","The finding that disclosed reviews show stronger growth in appreciative and constructive tone from 2019 to 2024 and no increase in critical tone would counter the worry that open review makes reviewers harsher.","Because only first-round reviews of accepted papers are observed, the same disparities would be expected to be at least as large in reviews of rejected manuscripts if the mechanism is bias in tone."],"supporting_citations":[{"why":"Supplies the SciBERT model that is fine-tuned for sentiment classification.","marker":"[27]"},{"why":"Supplies the labeled peer review sentences used to fine-tune the sentiment classifier.","marker":"[29]"},{"why":"Provides prior evidence that anonymous reviews tend to be more candid and critical, the premise the disclosure comparison builds on.","marker":"[38]"},{"why":"Documents reviewer bias in single- versus double-blind review, motivating the question of whether reviewer identity matters.","marker":"[6]"},{"why":"A large-scale language analysis of peer review reports that found limited linguistic effects of reviewer gender and review model, which this study extends in scale and scope.","marker":"[15]"},{"why":"A large-scale study of peer review and gender bias across many journals, providing the mixed evidence landscape the paper responds to.","marker":"[16]"},{"why":"Evidence that double-blind review reduces gender bias at conferences, used to argue that blinding is an incomplete solution.","marker":"[22]"},{"why":"Evidence on whether double-blind peer review reduces bias, which the discussion uses to frame disclosure as one possible reform.","marker":"[23]"}],"fun_headline_variants":["Anonymous reviews skew tone by author demographics","Signed reviews ease but don't erase peer review tone gaps","Reviewer anonymity linked to uneven peer feedback","Demographic tone gaps fade when reviewers sign off","Anonymity inflates appraisal gaps in peer review"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the difference in tone between signed and unsigned reviews reflects the effect of losing anonymity, but reviewers choose whether to sign, so reviewers who disclose may simply be more generous or constructive people; if so, the observed 'disclosure' effect is not caused by disclosure.","fun_headline_variants_meta":{"raw":{"variants":["Anonymous reviews skew tone by author demographics","Signed reviews ease but don't erase peer review tone gaps","Reviewer anonymity linked to uneven peer feedback","Demographic tone gaps fade when reviewers sign off","Anonymity inflates appraisal gaps in peer review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1333,"prompt_tokens":902,"completion_tokens":431,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":360}},"tokens_in":518,"tokens_out":431,"duration_ms":5313,"temperature":1.0,"reasoning_tokens":360,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:49:23.269054+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would be a randomized or policy-forced comparison in which the same reviewers evaluate comparable manuscripts under anonymous and signed conditions, for example by tracking a journal's switch to mandatory open review; if tone and demographic gaps are identical across conditions, the paper's claim that disclosure reduces disparities would be refuted, while persistence of critical gaps under mandatory signing would support it.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the labeled peer review sentences used to fine-tune the sentiment classifier."},{"cited_title":"Interactive functions of language in peer reviews of medical papers written by non- native users of english","cited_arxiv_id":null,"evidence_quote":"Provides prior evidence that anonymous reviews tend to be more candid and critical, the premise the disclosure comparison builds on."},{"cited_title":"& Heavlin, W","cited_arxiv_id":null,"evidence_quote":"Documents reviewer bias in single- versus double-blind review, motivating the question of whether reviewer identity matters."},{"cited_title":"& Maruˇsi´c, A","cited_arxiv_id":null,"evidence_quote":"A large-scale language analysis of peer review reports that found limited linguistic effects of reviewer gender and review model, which this study extends in scale and scope."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A large-scale study of peer review and gender bias across many journals, providing the mixed evidence landscape the paper responds to."},{"cited_title":"G., Politzer-Ahles, S","cited_arxiv_id":null,"evidence_quote":"Evidence that double-blind review reduces gender bias at conferences, used to argue that blinding is an incomplete solution."},{"cited_title":"& Teplitskiy, M","cited_arxiv_id":null,"evidence_quote":"Evidence on whether double-blind peer review reduces bias, which the discussion uses to frame disclosure as one possible reform."}],"review_version":1}