REVIEW 3 major objections 4 minor 2 references
Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Averaging 30 ChatGPT predictions yields weak-to-moderate correlation with peer review outcomes at ICLR and SciPost, and none at F1000Research.
desk verdict An honest, useful dataset extension whose positive correlations are undermined by unresolved training-data contamination. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the averaging ensemble: each submission is scored 30 separate times by ChatGPT-4o-mini through the API, and the mean of the 30 scores for a paper is compared with the human reviewer aggregate by Spearman rank correlation. Averaging smooths the model's stochastic output and is the step the paper credits for stronger results than single-shot predictions. The reviewer instructions for each venue, lightly reformatted as system prompts, define the scoring scale; a chain-of-thought variant that postpones the score to the end of the report is tested as an alternative prompt structure. Confidence intervals for the population correlations are estimated by bootstrapping.
What would settle it
Run the same 30-iteration protocol on submissions whose reviewer scores were not public before the model's training cutoff, and check whether the Spearman correlations stay above zero; if they fall to zero, the reported predictions were contamination from training data.
Extended reading notes
Core claim
The central claim is that an ensemble of repeated ChatGPT scorings, prompted with the actual reviewer guidelines, can partially reproduce the rank order of human peer review at some venues. On ICLR 2017, the averaged predictions correlated with reviewer scores at Spearman rho = 0.38 with title/abstract input and rho = 0.46 with full text. On SciPost Physics, the correlations were rho = 0.25 for validity, rho = 0.25 for originality, rho = 0.20 for significance, and rho = 0.08 for clarity. On F1000Research, the correlation was rho = 0.00 from title and abstract, rising only to rho = 0.09 with full text and rho = 0.10 with chain-of-thought prompting. The paper argues that this demonstrates weak but real predictive capacity in some contexts, not a universal ability, and that the optimal input format and prompt style vary by platform.
Load-bearing premise
The load-bearing assumption is that ChatGPT has not memorized the public review scores it is tested against, so its positive correlations reflect judgement rather than recall.
Editorial extensions
If this is right
- Editors at venues where correlations are positive could use averaged ChatGPT scores as a rough triage signal to flag submissions for desk rejection, provided authors consent and the system does not train on submitted text.
- Venue-specific pilot testing is mandatory: the same protocol that fails at F1000Research works at ICLR and SciPost, so no general assumption about LLM review prediction should be made.
- Full text is not always better: adding full text raised the ICLR correlation from 0.38 to 0.46 but barely moved F1000Research, so the best input must be determined empirically for each venue.
- Chain-of-thought prompts did not improve and sometimes degraded results, suggesting that standard reviewer instructions, with only light adaptation, are the safer prompt design.
- Because all tested scores are public, the positive correlations may not transfer to genuinely private review processes until the memorization question is resolved.
Reading between the lines
- A prospective test on live submissions whose outcomes are not yet public would separate genuine predictive signal from memorized training data; if correlations survive, the method could move from a research finding to an operational tool.
- The F1000Research failure may come from the coarse three-level decision scale combined with ChatGPT's tendency to pile up on 'Approved with Reservations'; using a wider score scale or forcing a distribution might recover signal, a hypothesis the paper does not test.
- The title-and-abstract correlations may reflect how convincingly a paper is written rather than its underlying truth; if so, LLM triage is better described as a proxy for presentation quality than for scientific validity.
- Running the same protocol with a local, non-API model on the same public datasets would show whether the effects are specific to ChatGPT or general to instruction-tuned LLMs.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper examines whether ChatGPT (specifically GPT-4o-mini) can predict pre-publication peer review outcomes by averaging 30 model responses for each submission. It uses three platforms with public review scores: F1000Research, ICLR 2017, and SciPost Physics, correlating averaged ChatGPT scores with human reviewer scores via Spearman's rho. Results: F1000Research shows essentially zero correlation (rho=0.00 for title/abstract), ICLR shows moderate correlation (rho=0.38 for title/abstract, rho=0.46 for full text), and SciPost Physics shows weak positive correlations for validity, originality, and significance (rho 0.20-0.25). The paper concludes that ChatGPT can provide weak pre-publication quality assessments in some contexts, but performance varies by platform and input type.
Significance. If the positive correlations reflect genuine predictive capacity rather than memorization of public review scores, the paper offers a practical, low-cost triage tool for editors and an important methodological lesson about averaging LLM judgments. The study covers three distinct venues, uses public data, provides confusion matrices, and explicitly reports limitations. The paper is transparent about the training-data contamination threat, which it labels a 'major conceptual limitation.' However, because the F1000Research null result is uninformative (the model assigns 0.5 to 244 of 250 papers), the paper's positive ICLR and SciPost findings currently lack a decisive defense against the leakage alternative. The central claim is therefore plausible but not yet established.
major comments (3)
- [Discussion, first paragraph; Table 1] The training-data contamination issue is not adequately mitigated. The authors argue that the F1000Research null result, combined with divergence between ChatGPT and human averages, provides 'some reassurance' that prior knowledge of scores was unlikely to be the main reason behind the positive correlations. This reasoning does not hold: Table 1 shows that ChatGPT's average score was 0.5 for 244 of 250 F1000Research papers, so the Spearman correlation is constrained to be near zero regardless of whether the model had memorized any scores. A model that has memorized ICLR or SciPost scores could easily still produce a near-constant F1000Research score if it failed to map its memory to the article-specific scale or chose a cautious middle option. To support the central claim that ChatGPT has predictive capacity, the authors should provide a direct leakage test, for example by comparing performance on submissions published after the model's training cutoff, or by probing whether ChatGPT can reproduce specific public scores given only titles/abstracts, or by validating on a private dataset with a similar format. Without such a test, the positive correlations for ICLR and SciPost remain consistent with recall rather than prediction.
- [Results: SciPost Physics; Discussion, final paragraph] The dimension-specific claims for SciPost Physics are not supported as independent assessments. The paper acknowledges that reviewer scores across the four dimensions are highly correlated, so a single underlying quality signal—whether inferred or memorized—could produce all four positive rhos. The authors should report the human-score inter-dimension correlation matrix and, if feasible, provide partial correlations or a multivariate analysis to show that ChatGPT's dimension scores add information beyond a common quality factor. As it stands, the reader cannot determine whether ChatGPT is assessing originality, validity, and significance separately or simply responding to an overall quality impression.
- [Results: F1000Research; Table 1] The conclusion that ChatGPT 'failed to predict' F1000Research outcomes overstates what the data show. Because the model assigned an average score of 0.5 to 244 of 250 papers, the near-zero Spearman correlation is a floor-variance artifact: the independent variable has almost no variance, so the correlation cannot be large even if a latent predictive signal exists. The appropriate characterization is that ChatGPT's outputs are insufficiently differentiated with the given prompt and scoring scheme, not that the model has no predictive capacity for this platform. This distinction matters for the paper's comparative claims about platform-specific performance.
minor comments (4)
- [Introduction / Methods, research questions] The research questions are mislabeled: RQ3 appears twice. The second occurrence (about system prompts) should be RQ2.
- [Table 3] The rows labeled 'Standard', 'Chain-of-thought', and 'Full text LaTeX' should be more explicit—'Standard (title/abstract)', 'Chain-of-thought (title/abstract)', and 'Full text LaTeX (standard prompt)'—to avoid ambiguity.
- [Results: SciPost Physics, Figure 4] Figure 4 reports correlations based on 104 papers with available LaTeX source, but the figure caption and text should state this more prominently, since all other analyses use 250 papers.
- [Discussion, third paragraph] The citation 'Saad et al., 2024' is described as a private dataset study, but the cited work is an observational study of ChatGPT in peer review; please clarify whether it truly used non-public review scores.
Circularity Check
No significant derivation circularity; the paper's self-cited averaging method is a hypothesis tested on new data, and the training-data leakage limitation is an acknowledged validity threat rather than a by-construction circle.
full rationale
This is an empirical evaluation, not a derivation, and no step in its inference chain reduces to its own inputs. The pipeline—download public manuscripts and reviews, prompt ChatGPT with reviewer guidelines, average 30 responses, correlate with human scores—does not fit any parameter to the target outcomes within the paper, and no equation is defined in terms of the criterion. The authors cite their own prior work (Thelwall 2024a, 2024b) to motivate title/abstract inputs and response averaging, but those are methodological precedents tested afresh here, and the ICLR result also replicates an external study (Zhou et al. 2024); self-citation is therefore not load-bearing. The strongest challenge is the paper's own 'major conceptual limitation': ICLR2017 and SciPost scores are public and predate GPT-4o-mini, so positive correlations (rho=0.38/0.46; rho=0.20-0.25) could reflect training-data recall rather than predictive capacity. This is a genuine external-validity threat—especially since the F1000 null result is partly a floor artifact (244/250 articles at a ChatGPT average of 0.5)—but the paper does not construct the model, fit it to these scores, or define its metric in terms of them, so the concern is contamination/leakage, not derivation-circularity. A conservative reading would therefore lower confidence in the 'prediction' label, but it does not make the analysis equivalent to its inputs by construction.
Assumptions & free parameters
assumptions (3)
- domain assumption Human reviewer scores on the three platforms are a valid ground truth for research quality and peer review outcomes.
- domain assumption The ChatGPT-4o-mini model's outputs are stable enough across 30 iterations to be averaged into a meaningful point estimate.
- domain assumption The extracted plain text and LaTeX files preserve enough content for the model to assess the papers.
Cite this review
Pith. "Pith review of Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms." pith.science (2026). https://pith.science/paper/7H4QT2DS
@misc{pith2026241109763,
author = {Pith},
title = {Pith review of: Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms},
year = {2026},
howpublished = {\url{https://pith.science/paper/7H4QT2DS}},
note = {Machine review of arXiv:2411.09763}
}
read the original abstract
While previous studies have demonstrated that Large Language Models (LLMs) can predict peer review outcomes to some extent, this paper builds on that by introducing two new contexts and employing a more robust method - averaging multiple ChatGPT scores. The findings that averaging 30 ChatGPT predictions, based on reviewer guidelines and using only the submitted titles and abstracts, failed to predict peer review outcomes for F1000Research (Spearman's rho=0.00). However, it produced mostly weak positive correlations with the quality dimensions of SciPost Physics (rho=0.25 for validity, rho=0.25 for originality, rho=0.20 for significance, and rho = 0.08 for clarity) and a moderate positive correlation for papers from the International Conference on Learning Representations (ICLR) (rho=0.38). Including the full text of articles significantly increased the correlation for ICLR (rho=0.46) and slightly improved it for F1000Research (rho=0.09), while it had variable effects on the four quality dimension correlations for SciPost LaTeX files. The use of chain-of-thought system prompts slightly increased the correlation for F1000Research (rho=0.10), marginally reduced it for ICLR (rho=0.37), and further decreased it for SciPost Physics (rho=0.16 for validity, rho=0.18 for originality, rho=0.18 for significance, and rho=0.05 for clarity). Overall, the results suggest that in some contexts, ChatGPT can produce weak pre-publication quality assessments. However, the effectiveness of these assessments and the optimal strategies for employing them vary considerably across different platforms, journals, and conferences. Additionally, the most suitable inputs for ChatGPT appear to differ depending on the platform.
Figures
Reference graph
Works this paper leans on
-
[1]
Aczel, B., Szaszi, B., & Holcombe, A. O. (2021). A billion-dollar donation: estimating the cost of researchers’ time spent on peer review. Research Integrity and Peer Review, 6, 1-8. Bar-Ilan, J., & Halevi, G. (2018). Temporal characteristics of retracted articles. Scientometrics, 116(3), 1771-1783. Bornmann, L., Haunschild, R., & Mutz, R. (2021). Growth ...
work page 2021
-
[25]
https://doi.org/10.1057/s41599-020-00703-8 Du, J., Wang, Y., Zhao, W., Deng, Z., Liu, S., Lou, R., & Yin, W. (2024). L LMs assist NLP researchers: Critique paper (meta-) reviewing. arXiv preprint arXiv:2406.16253. Franceschet, A., Lucas, J., O’Neill, B., Pando, E., & Thomas, M. (2022). Editor fatigue: can political science journals increase review invitat...
arXiv 2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.