{"id":"3022cdd4-6281-4b4d-9100-159f29af76e3","arxiv_id":"2411.09763","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Averaging 30 ChatGPT scores yields weak to moderate correlations with peer review scores on ICLR and SciPost Physics, but no correlation on F1000Research.","lead":"This study tested whether averaging 30 ChatGPT-4o-mini scores can predict real peer review results on three publication platforms. It found weak positive correlations for ICLR and SciPost Physics, and no correlation for F1000Research, suggesting AI pre-screening is highly context-dependent.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing issue is training-data leakage: ICLR 2017 and SciPost Physics review scores are public and predate GPT-4o-mini, so the positive correlations may reflect memorized outcomes rather than predictive capacity; the paper's own Discussion concedes this is unresolved.","rationale":"Good-faith reading: the paper is honest, well-structured, and appropriately cautious about limitations. It reports a straightforward empirical correlation study, and its headline finding is not implausible. The most load-bearing condition for that finding is that the correlations measure prediction rather than memory. The paper cannot establish that condition, because all positive-correlation datasets (ICLR 2017 and SciPost Physics) are public and pre-date the model's training. The authors acknowledge this in the Discussion, but they do not resolve it. Their two supporting arguments are weak: (1) the F1000 null is uninformative because the ChatGPT scores are degenerate (Table 1), and (2) consistency with earlier private-data studies is indirect and not a test of these datasets. I therefore agree with the reader's weakest_assumption. I considered other possible objections—small variance in ChatGPT scores, correlational rather than decision-level evaluation, and the absence of code/data—but none is as decisive: even a perfect methodology cannot distinguish prediction from leakage without a temporal or private-data holdout. The proposed post-cutoff test is the single check that would settle the concern. If it reproduces the correlations, the paper's conclusion is materially supported; if not, the central claim should be weakened to 'ChatGPT can sometimes match public review outcomes already in its training data,' and the practical triage recommendation would require stronger evidence.","tokens_in":12438,"tokens_out":4782,"duration_ms":46727,"concrete_test":"Run the same 30-iteration averaged protocol on submissions with review outcomes that postdate GPT-4o-mini's knowledge cutoff, e.g., ICLR 2024 submissions (reviews released in January 2024, after the model's training window) or an independent private dataset from a journal that does not publish scores. Compare title/abstract and full-text Spearman rhos to the paper's ICLR rho=0.38/0.46. If the post-cutoff correlations reproduce the reported range (≈0.3–0.5), the predictive claim survives; if they collapse to ≈0, the original positive correlations are best explained by training-data contamination rather than genuine predictive capacity. A secondary check: repeat the F1000 experiment on a non-public set of submissions with non-degenerate reviewer scores to confirm the null is real and not a floor effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To sustain the central claim—that ChatGPT can provide weak pre-publication quality assessments—the positive correlations for ICLR (rho=0.38–0.46) and SciPost (rho=0.20–0.25) must be attributable to prediction, not recall. This is the least secure link. Both venues' review scores are public and predate GPT-4o-mini's training data, so the model could plausibly have memorized them; the authors explicitly say it is 'unclear whether ChatGPT had encountered the scores' (Discussion). The F1000 null result does not rescue this: Table 1 shows ChatGPT assigned 0.5 to 244 of 250 articles, so rho≈0 there is a floor/variance artifact and carries no information about leakage elsewhere. The appeal to two private-data studies is indirect and does not test these exact public scores. Notably, the SciPost dimension correlations are not independent: reviewer scores across dimensions are highly correlated, so one memorized overall-quality signal could produce all four positive rhos. If leakage drives the positive results, the central claim collapses into 'the model recalls public reviews' and the triage application is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper examines whether ChatGPT (specifically GPT-4o-mini) can predict pre-publication peer review outcomes by averaging 30 model responses for each submission. It uses three platforms with public review scores: F1000Research, ICLR 2017, and SciPost Physics, correlating averaged ChatGPT scores with human reviewer scores via Spearman's rho. Results: F1000Research shows essentially zero correlation (rho=0.00 for title/abstract), ICLR shows moderate correlation (rho=0.38 for title/abstract, rho=0.46 for full text), and SciPost Physics shows weak positive correlations for validity, originality, and significance (rho 0.20-0.25). The paper concludes that ChatGPT can provide weak pre-publication quality assessments in some contexts, but performance varies by platform and input type.","tokens_in":12598,"tokens_out":3890,"duration_ms":43651,"significance":"If the positive correlations reflect genuine predictive capacity rather than memorization of public review scores, the paper offers a practical, low-cost triage tool for editors and an important methodological lesson about averaging LLM judgments. The study covers three distinct venues, uses public data, provides confusion matrices, and explicitly reports limitations. The paper is transparent about the training-data contamination threat, which it labels a 'major conceptual limitation.' However, because the F1000Research null result is uninformative (the model assigns 0.5 to 244 of 250 papers), the paper's positive ICLR and SciPost findings currently lack a decisive defense against the leakage alternative. The central claim is therefore plausible but not yet established.","major_comments":[{"comment":"The training-data contamination issue is not adequately mitigated. The authors argue that the F1000Research null result, combined with divergence between ChatGPT and human averages, provides 'some reassurance' that prior knowledge of scores was unlikely to be the main reason behind the positive correlations. This reasoning does not hold: Table 1 shows that ChatGPT's average score was 0.5 for 244 of 250 F1000Research papers, so the Spearman correlation is constrained to be near zero regardless of whether the model had memorized any scores. A model that has memorized ICLR or SciPost scores could easily still produce a near-constant F1000Research score if it failed to map its memory to the article-specific scale or chose a cautious middle option. To support the central claim that ChatGPT has predictive capacity, the authors should provide a direct leakage test, for example by comparing performance on submissions published after the model's training cutoff, or by probing whether ChatGPT can reproduce specific public scores given only titles/abstracts, or by validating on a private dataset with a similar format. Without such a test, the positive correlations for ICLR and SciPost remain consistent with recall rather than prediction.","section":"Discussion, first paragraph; Table 1"},{"comment":"The dimension-specific claims for SciPost Physics are not supported as independent assessments. The paper acknowledges that reviewer scores across the four dimensions are highly correlated, so a single underlying quality signal—whether inferred or memorized—could produce all four positive rhos. The authors should report the human-score inter-dimension correlation matrix and, if feasible, provide partial correlations or a multivariate analysis to show that ChatGPT's dimension scores add information beyond a common quality factor. As it stands, the reader cannot determine whether ChatGPT is assessing originality, validity, and significance separately or simply responding to an overall quality impression.","section":"Results: SciPost Physics; Discussion, final paragraph"},{"comment":"The conclusion that ChatGPT 'failed to predict' F1000Research outcomes overstates what the data show. Because the model assigned an average score of 0.5 to 244 of 250 papers, the near-zero Spearman correlation is a floor-variance artifact: the independent variable has almost no variance, so the correlation cannot be large even if a latent predictive signal exists. The appropriate characterization is that ChatGPT's outputs are insufficiently differentiated with the given prompt and scoring scheme, not that the model has no predictive capacity for this platform. This distinction matters for the paper's comparative claims about platform-specific performance.","section":"Results: F1000Research; Table 1"}],"minor_comments":[{"comment":"The research questions are mislabeled: RQ3 appears twice. The second occurrence (about system prompts) should be RQ2.","section":"Introduction / Methods, research questions"},{"comment":"The rows labeled 'Standard', 'Chain-of-thought', and 'Full text LaTeX' should be more explicit—'Standard (title/abstract)', 'Chain-of-thought (title/abstract)', and 'Full text LaTeX (standard prompt)'—to avoid ambiguity.","section":"Table 3"},{"comment":"Figure 4 reports correlations based on 104 papers with available LaTeX source, but the figure caption and text should state this more prominently, since all other analyses use 250 papers.","section":"Results: SciPost Physics, Figure 4"},{"comment":"The citation 'Saad et al., 2024' is described as a private dataset study, but the cited work is an observational study of ChatGPT in peer review; please clarify whether it truly used non-public review scores.","section":"Discussion, third paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and practically relevant question, and the authors are appropriately transparent about their data and limitations. The main barrier to acceptance is the unresolved training-data leakage threat, which undercuts the central claim of predictive capacity. I would encourage the editor to invite a revision that either (a) provides a direct contamination test, or (b) substantially tempers the conclusions to describe correlations with public review outcomes rather than predictive capacity. If the authors can supply a robust leakage check, the paper could become a solid contribution. If not, the claims should be reframed as descriptive associations, which would lower the paper's novelty but make it defensible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read if you care about whether LLMs can triage submissions. The paper gives us new numbers for two venues nobody had tested (F1000Research, SciPost Physics) and confirms ICLR with a stronger design than Zhou et al. But the central result is fragile: the positive correlations for ICLR and SciPost could be memory of public scores, and the paper's own control argument doesn't hold.\n\nWhat's actually new and good: This is the first test of averaging 30 ChatGPT scores on pre-publication datasets — the method is borrowed from the author's earlier work, but the application to F1000 and SciPost is new. The full-text versus title/abstract comparison produces a surprising reversal from Zhou et al. (full text helps ICLR here, hurts there). The paper is transparent: it reports null results, gives confidence intervals, and puts the full prompts in the appendix. The F1000 finding (rho=0.00) is honest and useful, even if it's partly a floor effect.\n\nSoft spots: Training-data leakage is the load-bearing issue. ICLR 2017 and SciPost review scores are public and predate GPT-4o-mini, so the model could have memorized them. The paper says this is 'a major conceptual limitation' and then tries to argue it away with the F1000 null — but Table 1 shows ChatGPT assigned 0.5 to 244 of 250 F1000 articles, so rho=0 there is a variance artifact, not evidence against memorization. The appeal to two private-data studies is indirect; it doesn't test these specific public scores. Also, SciPost's four dimensions are highly correlated, so one memorized quality signal could produce all positive rhos. Those are real problems, and they're unresolved. Minor issues: the F1000 0.5 midpoint for 'Approved with Reservations' is arbitrary, and there's no data/code availability statement, which weakens reproducibility.\n\nWho is this for? Researchers studying LLM-assisted peer review and research evaluation. It won't change practice, but it's a solid empirical stepping stone if you keep the contamination caveat in mind. I'd send it to a serious referee — the question is important, the design is mostly careful, and the leakage issue can be addressed in revision (e.g., test papers after the training cutoff or use private data). A desk reject would be a mistake.","headline":"An honest, useful dataset extension whose positive correlations are undermined by unresolved training-data contamination.","tokens_in":13188,"tokens_out":2841,"would_cite":true,"duration_ms":27045,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Averaging 30 ChatGPT predictions yields weak-to-moderate correlation with peer review outcomes at ICLR and SciPost, and none at F1000Research.","keywords":["ChatGPT","peer review prediction","research evaluation","large language models","pre-publication triage","Spearman correlation","averaging ensemble","reviewer guidelines"],"falsifier":"Run the same 30-iteration protocol on submissions whose reviewer scores were not public before the model's training cutoff, and check whether the Spearman correlations stay above zero; if they fall to zero, the reported predictions were contamination from training data.","tokens_in":12170,"feed_emoji":"🤖","tokens_out":6584,"duration_ms":58458,"temperature":0.7,"pith_summary":"Peer review is slow and expensive, so a cheap machine pre-screening would help editors triage submissions. This paper tests whether ChatGPT-4o-mini, given the same instructions as human reviewers and asked to score each submission 30 times, can predict actual pre-publication review outcomes. Averaging the 30 scores produced statistically significant, weak-to-moderate Spearman correlations with reviewer scores for ICLR 2017 papers (rho = 0.38 from titles and abstracts, rho = 0.46 from full text) and with three of the four SciPost Physics quality dimensions (rho = 0.20 to 0.25), but no correlation at all for F1000Research (rho = 0.00). The paper concludes that LLM-based triage is possible in some venues, but only after per-venue pilot testing and calibration, and never as a replacement for human judgement. It also finds that averaging multiple iterations is more reliable than single predictions and that chain-of-thought prompting does not help.","feed_headline":"Averaged ChatGPT scores predict peer review at some venues, not others","feed_subtitle":"Weak positive correlations at ICLR and SciPost, zero at F1000Research: LLM triage needs venue-level testing.","key_machinery":"The load-bearing mechanism is the averaging ensemble: each submission is scored 30 separate times by ChatGPT-4o-mini through the API, and the mean of the 30 scores for a paper is compared with the human reviewer aggregate by Spearman rank correlation. Averaging smooths the model's stochastic output and is the step the paper credits for stronger results than single-shot predictions. The reviewer instructions for each venue, lightly reformatted as system prompts, define the scoring scale; a chain-of-thought variant that postpones the score to the end of the report is tested as an alternative prompt structure. Confidence intervals for the population correlations are estimated by bootstrapping.","core_discovery":"The central claim is that an ensemble of repeated ChatGPT scorings, prompted with the actual reviewer guidelines, can partially reproduce the rank order of human peer review at some venues. On ICLR 2017, the averaged predictions correlated with reviewer scores at Spearman rho = 0.38 with title/abstract input and rho = 0.46 with full text. On SciPost Physics, the correlations were rho = 0.25 for validity, rho = 0.25 for originality, rho = 0.20 for significance, and rho = 0.08 for clarity. On F1000Research, the correlation was rho = 0.00 from title and abstract, rising only to rho = 0.09 with full text and rho = 0.10 with chain-of-thought prompting. The paper argues that this demonstrates weak but real predictive capacity in some contexts, not a universal ability, and that the optimal input format and prompt style vary by platform.","pith_inferences":["A prospective test on live submissions whose outcomes are not yet public would separate genuine predictive signal from memorized training data; if correlations survive, the method could move from a research finding to an operational tool.","The F1000Research failure may come from the coarse three-level decision scale combined with ChatGPT's tendency to pile up on 'Approved with Reservations'; using a wider score scale or forcing a distribution might recover signal, a hypothesis the paper does not test.","The title-and-abstract correlations may reflect how convincingly a paper is written rather than its underlying truth; if so, LLM triage is better described as a proxy for presentation quality than for scientific validity.","Running the same protocol with a local, non-API model on the same public datasets would show whether the effects are specific to ChatGPT or general to instruction-tuned LLMs."],"forward_implications":["Editors at venues where correlations are positive could use averaged ChatGPT scores as a rough triage signal to flag submissions for desk rejection, provided authors consent and the system does not train on submitted text.","Venue-specific pilot testing is mandatory: the same protocol that fails at F1000Research works at ICLR and SciPost, so no general assumption about LLM review prediction should be made.","Full text is not always better: adding full text raised the ICLR correlation from 0.38 to 0.46 but barely moved F1000Research, so the best input must be determined empirically for each venue.","Chain-of-thought prompts did not improve and sometimes degraded results, suggesting that standard reviewer instructions, with only light adaptation, are the safer prompt design.","Because all tested scores are public, the positive correlations may not transfer to genuinely private review processes until the memorization question is resolved."],"supporting_citations":[{"why":"Supplies the prior ICLR result this paper extends and the title/abstract scoring baseline it improves on.","marker":"Zhou et al., 2024"},{"why":"Previous experiments on ChatGPT quality scores that motivate averaging 30 predictions; the paper builds directly on this method.","marker":"Thelwall, 2024ab"},{"why":"Created the PeerRead dataset from which the ICLR 2017 submissions and reviewer scores were extracted.","marker":"Kang et al., 2018"},{"why":"One of the two private-dataset studies cited as evidence that positive correlations are not merely an artefact of public scores.","marker":"Saad et al., 2024"},{"why":"Introduces chain-of-thought prompting, the alternative prompt strategy the paper evaluates and finds not reliably helpful.","marker":"Zhang et al., 2022"}],"fun_headline_variants":["Averaged ChatGPT scores predict peer review only at some venues","Peer review triage: ChatGPT averages succeed at ICLR, not F1000","ChatGPT score averaging yields weak venue-specific review predictions","LLM review forecasts: strong ICLR correlation, zero F1000","Repeated ChatGPT scoring mimics human review at selective venues"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that ChatGPT has not memorized the public review scores it is tested against, so its positive correlations reflect judgement rather than recall.","fun_headline_variants_meta":{"raw":{"variants":["Averaged ChatGPT scores predict peer review only at some venues","Peer review triage: ChatGPT averages succeed at ICLR, not F1000","ChatGPT score averaging yields weak venue-specific review predictions","LLM review forecasts: strong ICLR correlation, zero F1000","Repeated ChatGPT scoring mimics human review at selective venues"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1482,"prompt_tokens":1094,"completion_tokens":388,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":710,"completion_tokens_details":{"reasoning_tokens":298}},"tokens_in":710,"tokens_out":388,"duration_ms":3759,"temperature":1.0,"reasoning_tokens":298,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:19:04.792609+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 30-iteration protocol on submissions whose reviewer scores were not public before the model's training cutoff, and check whether the Spearman correlations stay above zero; if they fall to zero, the reported predictions were contamination from training data.","supporting_citations":[],"review_version":1}