Pith. sign in

REVIEW 3 major objections 4 minor 2 references

Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Averaging 30 ChatGPT predictions yields weak-to-moderate correlation with peer review outcomes at ICLR and SciPost, and none at F1000Research.

desk verdict An honest, useful dataset extension whose positive correlations are undermined by unresolved training-data contamination. read the letter →

arxiv 2411.09763 v1 pith:7H4QT2DS submitted 2024-11-14 cs.DL cs.CL

classification cs.DLcs.CL
keywords ChatGPTpeerreviewpredictionresearchevaluationlargelanguagemodelspre-publicationtriageSpearmancorrelationaveragingensemblereviewerguidelines
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Peer review is slow and expensive, so a cheap machine pre-screening would help editors triage submissions. This paper tests whether ChatGPT-4o-mini, given the same instructions as human reviewers and asked to score each submission 30 times, can predict actual pre-publication review outcomes. Averaging the 30 scores produced statistically significant, weak-to-moderate Spearman correlations with reviewer scores for ICLR 2017 papers (rho = 0.38 from titles and abstracts, rho = 0.46 from full text) and with three of the four SciPost Physics quality dimensions (rho = 0.20 to 0.25), but no correlation at all for F1000Research (rho = 0.00). The paper concludes that LLM-based triage is possible in some venues, but only after per-venue pilot testing and calibration, and never as a replacement for human judgement. It also finds that averaging multiple iterations is more reliable than single predictions and that chain-of-thought prompting does not help.

What carries the argument

The load-bearing mechanism is the averaging ensemble: each submission is scored 30 separate times by ChatGPT-4o-mini through the API, and the mean of the 30 scores for a paper is compared with the human reviewer aggregate by Spearman rank correlation. Averaging smooths the model's stochastic output and is the step the paper credits for stronger results than single-shot predictions. The reviewer instructions for each venue, lightly reformatted as system prompts, define the scoring scale; a chain-of-thought variant that postpones the score to the end of the report is tested as an alternative prompt structure. Confidence intervals for the population correlations are estimated by bootstrapping.

What would settle it

Run the same 30-iteration protocol on submissions whose reviewer scores were not public before the model's training cutoff, and check whether the Spearman correlations stay above zero; if they fall to zero, the reported predictions were contamination from training data.

Watch

Extended reading notes

Core claim

The central claim is that an ensemble of repeated ChatGPT scorings, prompted with the actual reviewer guidelines, can partially reproduce the rank order of human peer review at some venues. On ICLR 2017, the averaged predictions correlated with reviewer scores at Spearman rho = 0.38 with title/abstract input and rho = 0.46 with full text. On SciPost Physics, the correlations were rho = 0.25 for validity, rho = 0.25 for originality, rho = 0.20 for significance, and rho = 0.08 for clarity. On F1000Research, the correlation was rho = 0.00 from title and abstract, rising only to rho = 0.09 with full text and rho = 0.10 with chain-of-thought prompting. The paper argues that this demonstrates weak but real predictive capacity in some contexts, not a universal ability, and that the optimal input format and prompt style vary by platform.

Load-bearing premise

The load-bearing assumption is that ChatGPT has not memorized the public review scores it is tested against, so its positive correlations reflect judgement rather than recall.

Editorial extensions

If this is right

  • Editors at venues where correlations are positive could use averaged ChatGPT scores as a rough triage signal to flag submissions for desk rejection, provided authors consent and the system does not train on submitted text.
  • Venue-specific pilot testing is mandatory: the same protocol that fails at F1000Research works at ICLR and SciPost, so no general assumption about LLM review prediction should be made.
  • Full text is not always better: adding full text raised the ICLR correlation from 0.38 to 0.46 but barely moved F1000Research, so the best input must be determined empirically for each venue.
  • Chain-of-thought prompts did not improve and sometimes degraded results, suggesting that standard reviewer instructions, with only light adaptation, are the safer prompt design.
  • Because all tested scores are public, the positive correlations may not transfer to genuinely private review processes until the memorization question is resolved.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A prospective test on live submissions whose outcomes are not yet public would separate genuine predictive signal from memorized training data; if correlations survive, the method could move from a research finding to an operational tool.
  • The F1000Research failure may come from the coarse three-level decision scale combined with ChatGPT's tendency to pile up on 'Approved with Reservations'; using a wider score scale or forcing a distribution might recover signal, a hypothesis the paper does not test.
  • The title-and-abstract correlations may reflect how convincingly a paper is written rather than its underlying truth; if so, LLM triage is better described as a proxy for presentation quality than for scientific validity.
  • Running the same protocol with a local, non-API model on the same public datasets would show whether the effects are specific to ChatGPT or general to instruction-tuned LLMs.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper examines whether ChatGPT (specifically GPT-4o-mini) can predict pre-publication peer review outcomes by averaging 30 model responses for each submission. It uses three platforms with public review scores: F1000Research, ICLR 2017, and SciPost Physics, correlating averaged ChatGPT scores with human reviewer scores via Spearman's rho. Results: F1000Research shows essentially zero correlation (rho=0.00 for title/abstract), ICLR shows moderate correlation (rho=0.38 for title/abstract, rho=0.46 for full text), and SciPost Physics shows weak positive correlations for validity, originality, and significance (rho 0.20-0.25). The paper concludes that ChatGPT can provide weak pre-publication quality assessments in some contexts, but performance varies by platform and input type.

Significance. If the positive correlations reflect genuine predictive capacity rather than memorization of public review scores, the paper offers a practical, low-cost triage tool for editors and an important methodological lesson about averaging LLM judgments. The study covers three distinct venues, uses public data, provides confusion matrices, and explicitly reports limitations. The paper is transparent about the training-data contamination threat, which it labels a 'major conceptual limitation.' However, because the F1000Research null result is uninformative (the model assigns 0.5 to 244 of 250 papers), the paper's positive ICLR and SciPost findings currently lack a decisive defense against the leakage alternative. The central claim is therefore plausible but not yet established.

major comments (3)
  1. [Discussion, first paragraph; Table 1] The training-data contamination issue is not adequately mitigated. The authors argue that the F1000Research null result, combined with divergence between ChatGPT and human averages, provides 'some reassurance' that prior knowledge of scores was unlikely to be the main reason behind the positive correlations. This reasoning does not hold: Table 1 shows that ChatGPT's average score was 0.5 for 244 of 250 F1000Research papers, so the Spearman correlation is constrained to be near zero regardless of whether the model had memorized any scores. A model that has memorized ICLR or SciPost scores could easily still produce a near-constant F1000Research score if it failed to map its memory to the article-specific scale or chose a cautious middle option. To support the central claim that ChatGPT has predictive capacity, the authors should provide a direct leakage test, for example by comparing performance on submissions published after the model's training cutoff, or by probing whether ChatGPT can reproduce specific public scores given only titles/abstracts, or by validating on a private dataset with a similar format. Without such a test, the positive correlations for ICLR and SciPost remain consistent with recall rather than prediction.
  2. [Results: SciPost Physics; Discussion, final paragraph] The dimension-specific claims for SciPost Physics are not supported as independent assessments. The paper acknowledges that reviewer scores across the four dimensions are highly correlated, so a single underlying quality signal—whether inferred or memorized—could produce all four positive rhos. The authors should report the human-score inter-dimension correlation matrix and, if feasible, provide partial correlations or a multivariate analysis to show that ChatGPT's dimension scores add information beyond a common quality factor. As it stands, the reader cannot determine whether ChatGPT is assessing originality, validity, and significance separately or simply responding to an overall quality impression.
  3. [Results: F1000Research; Table 1] The conclusion that ChatGPT 'failed to predict' F1000Research outcomes overstates what the data show. Because the model assigned an average score of 0.5 to 244 of 250 papers, the near-zero Spearman correlation is a floor-variance artifact: the independent variable has almost no variance, so the correlation cannot be large even if a latent predictive signal exists. The appropriate characterization is that ChatGPT's outputs are insufficiently differentiated with the given prompt and scoring scheme, not that the model has no predictive capacity for this platform. This distinction matters for the paper's comparative claims about platform-specific performance.
minor comments (4)
  1. [Introduction / Methods, research questions] The research questions are mislabeled: RQ3 appears twice. The second occurrence (about system prompts) should be RQ2.
  2. [Table 3] The rows labeled 'Standard', 'Chain-of-thought', and 'Full text LaTeX' should be more explicit—'Standard (title/abstract)', 'Chain-of-thought (title/abstract)', and 'Full text LaTeX (standard prompt)'—to avoid ambiguity.
  3. [Results: SciPost Physics, Figure 4] Figure 4 reports correlations based on 104 papers with available LaTeX source, but the figure caption and text should state this more prominently, since all other analyses use 250 papers.
  4. [Discussion, third paragraph] The citation 'Saad et al., 2024' is described as a private dataset study, but the cited work is an observational study of ChatGPT in peer review; please clarify whether it truly used non-public review scores.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant derivation circularity; the paper's self-cited averaging method is a hypothesis tested on new data, and the training-data leakage limitation is an acknowledged validity threat rather than a by-construction circle.

full rationale

This is an empirical evaluation, not a derivation, and no step in its inference chain reduces to its own inputs. The pipeline—download public manuscripts and reviews, prompt ChatGPT with reviewer guidelines, average 30 responses, correlate with human scores—does not fit any parameter to the target outcomes within the paper, and no equation is defined in terms of the criterion. The authors cite their own prior work (Thelwall 2024a, 2024b) to motivate title/abstract inputs and response averaging, but those are methodological precedents tested afresh here, and the ICLR result also replicates an external study (Zhou et al. 2024); self-citation is therefore not load-bearing. The strongest challenge is the paper's own 'major conceptual limitation': ICLR2017 and SciPost scores are public and predate GPT-4o-mini, so positive correlations (rho=0.38/0.46; rho=0.20-0.25) could reflect training-data recall rather than predictive capacity. This is a genuine external-validity threat—especially since the F1000 null result is partly a floor artifact (244/250 articles at a ChatGPT average of 0.5)—but the paper does not construct the model, fit it to these scores, or define its metric in terms of them, so the concern is contamination/leakage, not derivation-circularity. A conservative reading would therefore lower confidence in the 'prediction' label, but it does not make the analysis equivalent to its inputs by construction.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper does not introduce free parameters or invented entities. It relies on three domain assumptions: the validity of human review scores as a benchmark, the stability of the LLM across iterations, and the fidelity of the text extraction. The training-data contamination question is treated as an acknowledged limitation rather than an axiom.

assumptions (3)
  • domain assumption Human reviewer scores on the three platforms are a valid ground truth for research quality and peer review outcomes.
    The paper treats averaged human scores as the criterion against which ChatGPT is judged, without problematizing whether reviewer scores themselves are reliable. This is standard for this line of research but is an assumption about the validity of the gold standard.
  • domain assumption The ChatGPT-4o-mini model's outputs are stable enough across 30 iterations to be averaged into a meaningful point estimate.
    The analysis assumes that the variance across iterations is noise and that the mean captures a stable model tendency, which is supported by the convergence shown in Figures 1-4 but is not formally proven.
  • domain assumption The extracted plain text and LaTeX files preserve enough content for the model to assess the papers.
    The authors note that equations were stripped to symbolic representations and that SciPost LaTeX files were 'largely meaningless' in some cases. This undermines the assumption that full-text inputs represent the actual papers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms." pith.science (2026). https://pith.science/paper/7H4QT2DS

@misc{pith2026241109763,
  author       = {Pith},
  title        = {Pith review of: Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7H4QT2DS}},
  note         = {Machine review of arXiv:2411.09763}
}
read the original abstract

While previous studies have demonstrated that Large Language Models (LLMs) can predict peer review outcomes to some extent, this paper builds on that by introducing two new contexts and employing a more robust method - averaging multiple ChatGPT scores. The findings that averaging 30 ChatGPT predictions, based on reviewer guidelines and using only the submitted titles and abstracts, failed to predict peer review outcomes for F1000Research (Spearman's rho=0.00). However, it produced mostly weak positive correlations with the quality dimensions of SciPost Physics (rho=0.25 for validity, rho=0.25 for originality, rho=0.20 for significance, and rho = 0.08 for clarity) and a moderate positive correlation for papers from the International Conference on Learning Representations (ICLR) (rho=0.38). Including the full text of articles significantly increased the correlation for ICLR (rho=0.46) and slightly improved it for F1000Research (rho=0.09), while it had variable effects on the four quality dimension correlations for SciPost LaTeX files. The use of chain-of-thought system prompts slightly increased the correlation for F1000Research (rho=0.10), marginally reduced it for ICLR (rho=0.37), and further decreased it for SciPost Physics (rho=0.16 for validity, rho=0.18 for originality, rho=0.18 for significance, and rho=0.05 for clarity). Overall, the results suggest that in some contexts, ChatGPT can produce weak pre-publication quality assessments. However, the effectiveness of these assessments and the optimal strategies for employing them vary considerably across different platforms, journals, and conferences. Additionally, the most suitable inputs for ChatGPT appear to differ depending on the platform.

Figures

Figures reproduced from arXiv: 2411.09763 by the authors.

Figure 1
Figure 1. Spearman correlations between reviewer recommendations and average ChatGPT [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Spearman correlations between reviewer recommendations and average ChatGPT [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 4
Figure 4. Spearman correlations between reviewer recommendations and average ChatGPT [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: Spearman correlations between reviewer recommendations and average ChatGPT [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

2 extracted references · 1 canonical work pages

  1. [1]

    la Caixa

    Aczel, B., Szaszi, B., & Holcombe, A. O. (2021). A billion-dollar donation: estimating the cost of researchers’ time spent on peer review. Research Integrity and Peer Review, 6, 1-8. Bar-Ilan, J., & Halevi, G. (2018). Temporal characteristics of retracted articles. Scientometrics, 116(3), 1771-1783. Bornmann, L., Haunschild, R., & Mutz, R. (2021). Growth ...

  2. [25]

    https://doi.org/10.1057/s41599-020-00703-8 Du, J., Wang, Y., Zhao, W., Deng, Z., Liu, S., Lou, R., & Yin, W. (2024). L LMs assist NLP researchers: Critique paper (meta-) reviewing. arXiv preprint arXiv:2406.16253. Franceschet, A., Lucas, J., O’Neill, B., Pando, E., & Thomas, M. (2022). Editor fatigue: can political science journals increase review invitat...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.