REVIEW 2 major objections 1 minor 4 references
More positive sentiment in peer reviews links to shorter durations, especially for evaluation and impact aspects.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-26 12:22 UTC pith:XI6JAFQK
load-bearing objection The paper reports weak negative correlations between positive aspect-level sentiments and shorter review durations in Nature Communications data, with stronger signals for Evaluation/Results and Impact/Value aspects that vary by round, but the unvalidated NLP pipeline makes the specific claims hard to trust. the 2 major comments →
Which Review Aspect Has a Greater Impact on the Duration of Open Peer Review in Multiple Rounds? -- Evidence from Nature Communications
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Review sentiment has a weak but statistically significant negative correlation with peer review duration, indicating that more positive reviews tend to be associated with shorter review periods. Aspects concerning Evaluation and Results and Impact and Research Value show relatively stronger correlations with review duration. The relationships between aspect-level sentiment and review duration also differ significantly across review rounds.
What carries the argument
Aspect-level sentiment scores extracted from peer review reports and their correlation with review duration.
Load-bearing premise
The fine-grained aspect extraction and sentiment classification model accurately identifies aspects and their polarity from raw review text without substantial mislabeling or overlap.
What would settle it
Reanalyzing the same review reports with manual checks or a different extraction model and finding no significant sentiment-duration correlation would challenge the central claim.
If this is right
- Authors may prioritize revisions on evaluation, results, and research value aspects to potentially shorten review times.
- Reviewers and editors could use aspect sentiment to better anticipate and manage review timelines.
- Approaches to improve efficiency need to differ between initial and later review rounds.
- Targeted revisions based on these aspects may help accelerate scholarly communication.
Where Pith is reading between the lines
- The same aspect-sentiment approach could be applied to reviews from journals other than Nature Communications to test generalizability.
- Improved sentiment models might enable tools that flag aspects likely to extend review duration.
- Early-round sentiment patterns might allow predictions of total review length in future applications.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper examines correlations between aspect-level sentiment extracted from peer review reports in Nature Communications and the duration of the multi-round open peer review process. It employs a two-stage NLP pipeline (fine-grained aspect extraction followed by sentiment classification) and reports a weak but statistically significant negative correlation between overall review sentiment and duration, with relatively stronger associations for the Evaluation and Results and Impact and Research Value aspects; these relationships also vary across review rounds. The goal is to inform targeted revisions and improve review efficiency.
Significance. If the aspect extraction and sentiment labels prove reliable, the work could provide actionable evidence linking specific review content to process timing, potentially aiding authors, reviewers, and editors. The round-wise differences and aspect-specific findings add granularity beyond overall sentiment. However, the lack of any reported model validation, performance metrics, or sample details substantially limits the current significance, as the reported correlations rest entirely on unverified labels.
major comments (2)
- [Design/methodology/approach] Design/methodology/approach: The two-stage NLP pipeline (aspect extraction then sentiment classification) is described at a high level but provides no training details, architecture, performance metrics (precision/recall/F1), or human validation results on Nature Communications reviews. All correlations in the Findings section depend directly on the quality of these labels; systematic errors in aspect boundaries or polarity would render the reported negative correlation and round-specific differences uninterpretable.
- [Findings] Findings: The abstract states statistically significant correlations but supplies no information on sample size, exact measurement of duration, effect sizes, multiple-testing correction across aspects/rounds, or model accuracy. These omissions are load-bearing because they prevent assessment of whether the 'weak' negative correlation and stronger effects for Evaluation/Results and Impact/Value aspects are robust or artifacts.
minor comments (1)
- The abstract and methodology description would benefit from explicit definitions of 'review duration' and 'sentiment score' to allow replication.
Simulated Author's Rebuttal
We thank the referee for these constructive comments, which identify key areas where additional methodological transparency and statistical detail will strengthen the manuscript. We will revise accordingly to provide the requested information on the NLP pipeline and reporting of findings.
read point-by-point responses
-
Referee: [Design/methodology/approach] Design/methodology/approach: The two-stage NLP pipeline (aspect extraction then sentiment classification) is described at a high level but provides no training details, architecture, performance metrics (precision/recall/F1), or human validation results on Nature Communications reviews. All correlations in the Findings section depend directly on the quality of these labels; systematic errors in aspect boundaries or polarity would render the reported negative correlation and round-specific differences uninterpretable.
Authors: We agree that the current description of the two-stage NLP pipeline is insufficiently detailed. In the revised manuscript we will expand the Design/methodology/approach section with: (i) the training corpora and fine-tuning procedures used for aspect extraction and sentiment classification, (ii) the specific model architectures, (iii) quantitative performance metrics (precision, recall, F1) obtained on held-out validation data, and (iv) results of a human validation study performed on a sample of Nature Communications reviews. These additions will allow readers to assess label quality and the robustness of the reported correlations. revision: yes
-
Referee: [Findings] Findings: The abstract states statistically significant correlations but supplies no information on sample size, exact measurement of duration, effect sizes, multiple-testing correction across aspects/rounds, or model accuracy. These omissions are load-bearing because they prevent assessment of whether the 'weak' negative correlation and stronger effects for Evaluation/Results and Impact/Value aspects are robust or artifacts.
Authors: We accept that the abstract and Findings section require fuller statistical reporting. We will update both to state the total sample size, the precise operationalization of review duration, the observed correlation coefficients (effect sizes), whether and how multiple-testing correction was applied across aspects and rounds, and the model accuracy figures. These revisions will make it possible to evaluate the strength and reliability of the weak negative correlation and the aspect- and round-specific patterns. revision: yes
Circularity Check
No circularity: empirical correlations computed from independent duration data after NLP labeling
full rationale
The paper describes a two-stage pipeline that first applies aspect extraction and sentiment classification to peer-review text, then computes Pearson or similar correlations between the resulting sentiment scores and the externally measured review duration. No equations, fitted parameters, or predictions are defined in terms of the target correlations themselves. No self-citations are invoked to justify uniqueness or load-bearing premises. The reported negative correlations, aspect-specific strengths, and round-wise differences are therefore direct empirical measurements rather than reductions to the paper's own inputs or prior self-referential results.
Axiom & Free-Parameter Ledger
read the original abstract
Purpose: Peer review is essential to scientific publishing, but increasing submission volumes have placed growing pressure on reviewers and editors. This study examines the relationship between sentiment toward specific review aspects and peer review duration. It also investigates how this relationship varies across disciplines and review rounds, with the aim of supporting targeted manuscript revision and improving review efficiency. Design/methodology/approach: We adopt a two-stage approach. First, fine-grained aspects are extracted from peer review reports, and a sentiment classification model is used to determine the sentiment associated with each aspect. Second, correlations between aspect-level sentiment and peer review duration are analyzed. Sentiment scores are also calculated for different review rounds to determine whether these relationships change over successive rounds. Findings: Review sentiment has a weak but statistically significant negative correlation with peer review duration, indicating that more positive reviews tend to be associated with shorter review periods. Aspects concerning Evaluation and Results and Impact and Research Value show relatively stronger correlations with review duration. The relationships between aspect-level sentiment and review duration also differ significantly across review rounds. Originality/value: This study connects the textual content of peer review reports with the temporal characteristics of the review process. By identifying review aspects that are more closely associated with review duration, it provides evidence that may help authors prioritize revisions and assist reviewers and editors in improving review efficiency. The findings contribute to reducing the burden of peer review and accelerating scholarly communication and knowledge dissemination.
Reference graph
Works this paper leans on
-
[1]
Benos, D. J., Bashari, E., Chaves, J. M., Gaggar, A., Kapoor, N., LaFrance, M., ... & Zotov, A. (2007). The ups and downs of peer review. Advances in physiology education, 31(2), 145-152. Bharti, P. K., Ghosal, T., Agarwal, M., & Ekbal, A. (2023). PEERRec: An AI -based approach to automatically generate recommendations and predict decisions in peer review...
-
[2]
Luo, J., Feliciani, T., Reinhart, M., Hartste in, J., Das, V ., Alabi, O., & Shankar, K. (2021). Analysing sentiments in peer review reports: evidence from two science funding agencies. Quantitative Science Studies, 2(4), 1271-1295. Lyman R.L. (2013). A three -decade history of the duration of peer revi ew. Journal of Scholarly Publishing, 44(3), 211-220....
work page internal anchor Pith review Pith/arXiv arXiv 2021
-
[3]
Aspect Level Sentiment Classification with Deep Memory Network
Qin, C. and Zhang, C. (2023), "Which structure of academic articles do referees pay more attention to?: perspective of peer review and full -text of academic articles", Aslib Journal of Information Management, 75(5), 884-916. Qiu, G., Liu, B., Bu, J., & Chen, C. (2011). Opinion word expansion and target extraction through double propagation. Computational...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.1007/s11192-024-05070-8 2023
-
[4]
“This article is interesting, however
Zhang, G., Wang, L., Xie, W., Shang, F., Xia, X., Jiang, C. and Wang, X. (2022), "“This article is interesting, however”: exploring the language use in the peer review comment of articles published in the BMJ", Aslib Journal of Information Management, 74(3), 399-416
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.