Pith. sign in

REVIEW 3 major objections 5 minor 8 references

Is ChatGPT as reliable as individual reviewers assessing the quality of published journal articles from PDFs or titles and abstracts?

T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read ChatGPT-5.4 ranks journal articles as reliably as individual expert reviewers

desk verdict The averaging of 30 ChatGPT scores explains the apparent 'more reliable than reviewers' result; the paper is transparent and the null results are still useful. read the letter →

arxiv 2607.25965 v1 pith:QVQGO4FV submitted 2026-07-28 cs.DL

classification cs.DL
keywords largelanguagemodelsresearchqualityassessmentpeerreviewreliabilityChatGPTscoringtitle/abstractvsPDFevaluationSpearmanrankcorrelationREF-styleinter-rateragreement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Using confidential internal review scores for 200 published journal articles across three fields, the paper asks whether ChatGPT-5.4's quality rankings are as accurate as those of individual expert reviewers. For one field with per-reviewer data, the rank correlation between two human reviewers was about half the correlation between ChatGPT's averaged scores and each reviewer, suggesting the model may rank closer to the 'true' quality order—though the difference is not statistically significant. The paper also shows that feeding ChatGPT the full PDF produces more detailed, specific evaluative comments than feeding only a title and abstract, yet the PDF-based scores are not better at predicting expert quality ratings. The practical upshot: ChatGPT scores can help rank documents, but its detailed PDF critiques should not be mistaken for genuine deep evaluation.

What carries the argument

The core mechanism is the averaged-score ranking pipeline: each article is scored 30 times (5 prompt variants × 6 runs), the scores averaged, and the averages ranked. The rank-ratio argument then compares human-human Spearman correlation to ChatGPT-human correlation: if the LLM were purely noisy, its correlation with humans would be zero; if human-like, equal to human-human; if better, greater. The ratio is treated as the LLM's 'fraction' of human-level expertise. Statistical significance is assessed via bootstrap confidence intervals.

What would settle it

Take the same UoA3-style dataset but compute a Spearman correlation using just one of the 30 ChatGPT runs (not the averaged score) against each reviewer; if the single-run correlation drops to about 0.13 (the human-human level) or below, the claim that ChatGPT is as reliable as individual reviewers fails. Alternatively, recruit a large set of reviewers and show human-human correlation rises above ChatGPT-human correlation when both are computed on the same number of independent judgments.

Watch

Extended reading notes

Core claim

The paper's central claim is that, for the dataset examined, ChatGPT-5.4's score ranks agree with individual human expert reviewers at least as well as those reviewers agree with each other. The evidence is a single field where individual reviewer scores were available: the Spearman correlation between two sets of reviewer scores was about 0.13, while ChatGPT-human correlations ranged from about 0.25 to 0.32 depending on input (full text, PDF, or title/abstract). Because the comparison assumes human ranks are noisy approximations of a true quality order, the higher ChatGPT-human correlation is taken to mean the model's rankings are closer to the truth. The paper is careful to note the differ

Load-bearing premise

The argument assumes that the average of many ChatGPT responses can be fairly compared against a single human score, and that human ranks are noisy approximations of a true quality order—if either assumption fails, the conclusion that ChatGPT is as reliable as individual reviewers collapses.

Editorial extensions

If this is right

  • ChatGPT-5.4's averaged scores rank journal articles in rough agreement with expert quality scores across three fields.
  • The model's rankings may be more reliable than individual reviewers in the one field tested, though not significantly so.
  • Feeding full PDFs does not improve score prediction over titles and abstracts, despite producing more detailed evaluative reports.
  • PDF input's detailed critiques likely stem from reading full text, but they don't translate into better scores.
  • Using ChatGPT for low-stakes quality evaluation (formative feedback, calibration, journal ranking) may be defensible; high-stakes decisions (promotion, funding, REF selection) should not rely on it.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The fairness of the headline comparison hinges on averaging 30 ChatGPT scores versus a single human score; a fairer test might match one ChatGPT response against one human score, which would likely lower the LLM's apparent reliability.
  • The near-identical rankings from full-text and PDF inputs suggest the model mostly ignores figures and tables, so image-rich PDFs add little signal—likely true for other LLMs, but untested here.
  • A testable extension: use a single raw ChatGPT response (not averaged) and a larger set of reviewers to see whether the reliability advantage survives; the paper's own confidence intervals suggest it may not.
  • If ChatGPT's detailed critiques don't improve scores, then research managers should treat LLM-written evaluation reports as summaries, not as evidence of deep understanding.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper uses confidential internal REF-style expert scores from three UK Units of Assessment (UoA3: 98 articles with individual reviewer scores; UoA13: 44; UoA34: 58) to ask whether ChatGPT-5.4 scores can match individual expert reviewers (RQ1, UoA3 only), whether PDF input improves score prediction over titles/abstracts (RQ2), and whether the qualitative content of PDF-based reports differs from that of title/abstract-based reports (RQ3). The authors report positive rank correlations with expert scores, no clear improvement from PDF input, and more detailed evaluative comments in PDF-based reports. They interpret the UoA3 results as weakly suggesting ChatGPT-5.4 may be more reliable than individual reviewers, while acknowledging that the difference is not statistically significant.

Significance. If the RQ1 claim were supported, this would be a notable benchmark for LLM-assisted research evaluation: private expert scores not in training data, multiple prompt variants, three model sizes, and three fields. The study is transparent about data collection and limitations, and the WATA analysis is a statistically grounded qualitative complement. However, the central RQ1 comparison is not like-for-like because the ChatGPT side is an average of 30 responses whereas each human contributes one score. Under a standard measurement-error model, averaging alone can produce exactly the observed ChatGPT–human correlation even if a single ChatGPT response has no more reliability than a single human reviewer. This undermines the headline inference. The RQ2/RQ3 results are still useful, but the paper's strongest claim needs substantially more support.

major comments (3)
  1. [§3.1, Table 2] The RQ1 comparison is not like-for-like. Table 2 reports human–human Spearman ρ=0.126 for single reviewer score sets, but the ChatGPT–human correlations use the average of 30 LLM responses. Averaging 30 independent conditionally exchangeable responses inflates the expected correlation with any single human score. If a single ChatGPT response had exactly human reliability, the expected correlation between the 30-response mean and one human reviewer would be approximately 0.126/√(0.126+0.874/30) ≈ 0.32 — essentially the observed 0.319 for title/abstract input. Thus the data are fully consistent with ChatGPT having no per-response advantage; the apparent advantage is an artefact of the averaging protocol. The paper acknowledges this only in the final conclusion ('not individual scores but the average of 30 scores with structured prompts') and does not report single-response correlations. Pl
  2. [§2.4] The 'fraction of human level expertise' is defined as the ratio of the LLM–human correlation to the human–human correlation. This ratio is not a valid effect size: Spearman correlations are nonlinear, and the denominator is a single-pair correlation, not the reliability of the composite used on the LLM side. Moreover, the ratio is directly inflated by the averaging protocol. The argument in §2.4 that 'if the LLM is more powerful than a single human, then it can correlate more strongly with individual human scores' is only valid when both sides are single scores. The analysis should define the estimand explicitly (e.g., single-response Spearman vs. single-reviewer Spearman) and avoid interpreting the ratio as a percentage of human expertise.
  3. [§3.2, Figures 2–4] The practical conclusion that PDF input is unnecessary for obtaining the best scores rests on non-significant differences between overlapping confidence intervals. Because the same articles are scored under multiple input conditions, the correlations are dependent; a paired bootstrap or an appropriate test for dependent correlations (e.g., Steiger's Z) should be reported. As written, the evidence supports 'no significant improvement from PDF input' but not the stronger practical advice that PDFs are unnecessary, which is highlighted in the title and abstract. This is especially relevant for UoA13, where most correlations are not significantly different from zero.
minor comments (5)
  1. [Abstract] The abstract states that 'the rank correlations with expert scores are almost all statistically significantly positive', but the body reports that for UoA13 most correlations are not significantly different from zero. Please qualify the abstract to match the UoA-specific results.
  2. [Throughout] The model is referred to as 'ChatGPT 5.4', 'ChatGPT-5.4', and 'ChatGPT-5' in different places (e.g., RQ1 wording, Section 2.3, Discussion). Standardize the naming and specify whether '5.4' is a version of the API or a model release.
  3. [Figure 1 / Table 2] The text says 'the confidence intervals mostly overlap', but Table 2 shows the title/abstract ChatGPT Spearman CI (0.197–0.456) does not overlap with the reviewer–reviewer CI (0.112–0.156). Please clarify that the overlap applies to some input types only, and explain the implications for interpretation.
  4. [§2.1] The description of the UoA3 reviewer sets could be clearer: were 'Reviewer 1' and 'Reviewer 2' fixed individual reviewers across all articles, or are they role labels used by different pairs? The random-shuffle analysis in Table 2 suggests the latter, but the text should state this explicitly.
  5. [§3.3] The WATA manual theme grouping is described as repeated until stable, but no inter-coder reliability index is reported. Given that the qualitative results support the RQ3 interpretation, an independent second coder or a reliability check would strengthen the analysis.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: the core comparisons rest on private expert scores and untrained model outputs; the RQ1 caveat is an averaging asymmetry, not a definitional shortcut.

full rationale

The paper's derivation chain is not circular. The expert scores are private departmental mini-REF ratings (section 2.1), explicitly chosen because they cannot be in LLM training data, and the ChatGPT scores are generated from the paper's own prompts (section 2.3) rather than fitted to the target scores. No parameter is estimated from the outcome data, and no prediction is defined in terms of the quantity it is claimed to predict. The main quantitative concern is RQ1 (section 3.1, Table 2): ChatGPT scores are the mean of 30 responses ('The final ChatGPT score for each article and input type ... was the arithmetic mean of the 30 scores', section 2.3) and are compared with a single human reviewer score, so the higher ChatGPT-human correlation is expected even if a single ChatGPT response had human-level reliability. Indeed, using the paper's own human-human Spearman value (about 0.126), the Spearman-Brown expectation for a 30-response average is about 0.32, close to the observed 0.319 for title/abstract input. That shows the apparent advantage is explainable by measurement design, but it does not make the design circular: the aggregate is not defined in terms of the human scores, and no fitted parameter is renamed as a prediction. The paper also partially discloses the scope in the conclusion: 'not individual scores but the average of 30 scores with structured prompts'. This is a validity/comparability limitation rather than a circular derivation. The reliance on prior work by the same authors for prompt design and model choice (Thelwall, 2026b; Thelwall & Mohammadi, 2026) is self-citation, but it is not load-bearing proof of the empirical claim; the present results are evaluated against external, non-public human scores and are not forced by that citation. The shared REF rubric between humans and ChatGPT is a task-description match, not a derivation of the score from the answer. The RQ3 conclusion that detailed PDF comments are 'not genuine evaluations' because they do not improve scores is a stipulative interpretive claim, not a circular reduction. Overall, no circular step rises above minor, non-load-bearing self-citational practice; the skeptical case concerns statistical interpretation rather than circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

No fitted constants or new entities are introduced. The central claims rest on assumptions about reviewer noise, the ground-truth proxy, and the fairness of comparing an averaged LLM score to a single human score. The only notable hand-chosen number is the 30-response averaging, which is load-bearing for RQ1.

free parameters (1)
  • LLM score aggregation (mean of 30) = 30 responses (5 prompt variants × 6 runs)
    Chosen from prior work to reduce individual response noise. Directly affects RQ1 by comparing a de-noised ensemble to single human scores, inflating apparent reliability.
assumptions (4)
  • domain assumption Human reviewer scores are independent random noisy measurements of a single true article quality
    Section 2.4 uses this to interpret correlation ratios as the LLM's 'fraction of human expertise'.
  • domain assumption The departmental agreed/reviewer-average score is the best available proxy for true article quality
    Section 2.4; used as the ground truth for RQ2 comparisons across all three UoAs.
  • domain assumption The first two UoA3 reviewers are a representative sample of reviewer behavior; additional reviewers excluded
    Section 2.1 and 3.1; exclusion avoids bias from disagreement-triggered additions, but the deliberate pairing of one specialist and one non-specialist may lower human-human correlation.
  • domain assumption Word Association Thematic Analysis (chi-squared with Benjamini-Hochberg correction) identifies meaningful differences between report sets
    Section 2.4 RQ3; relies on statistical word frequency comparison plus manual theme clustering, introducing subjectivity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Is ChatGPT as reliable as individual reviewers assessing the quality of published journal articles from PDFs or titles and abstracts?." pith.science (2026). https://pith.science/paper/QVQGO4FV

@misc{pith2026260725965,
  author       = {Pith},
  title        = {Pith review of: Is ChatGPT as reliable as individual reviewers assessing the quality of published journal articles from PDFs or titles and abstracts?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QVQGO4FV}},
  note         = {Machine review of arXiv:2607.25965}
}
read the original abstract

Whilst Large Language Models (LLMs) have a weak to moderate ability to score published journal articles for research quality, they have not been compared with individual expert reviewers. It is also unknown whether quality scores from ChatGPT based on PDFs can improve on those from titles and abstracts through a deeper evaluation. To address both issues, this article uses expert scores (98 internal departmental ratings for UK Unit of Assessment [UoA] 3 Allied Health Professions, 44 for UoA13 Architecture, and 58 library and information science articles from UoA34), comparing them against ChatGPT-5.4 scores from both title/abstract and PDF inputs. For UoA3, individual reviewer scores were also compared against each other and ChatGPT-5.4. The rank correlations with expert scores are almost all statistically significantly positive, but differences between the correlations are mostly not, despite weakly suggesting that ChatGPT-5.4 can be more reliable than individual reviewers for UoA3. Moreover, whilst ChatGPT-5.4 provides more detailed evaluations of PDFs than of titles/abstracts, its score predictions do not seem to improve. Thus, whilst the results broadly confirm the value of ChatGPT scores for ranking academic documents, its apparently deeper evaluative comments on PDFs are misleading in the sense of not translating to improved score predictions.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

8 extracted references · 2 linked inside Pith

  1. [1]

    All reports were decomposed into their component words and all 30 reports for the input type were concatenated into a single mega -report for each article . Although the multiple reports on the same article are always different, they can focus on a specific aspect of that article, so combining reports avoids article - specific terms appearing as statistic...

  2. [2]

    The proportion of mega-reports in both sets (pdf, and title/abstract) containing each word was calculated (e.g., 3% of pdf mega-reports and 1% of title/abstract mega-reports contain “error”)

  3. [3]

    A chi-squared test with a Benjamini Hochberg correction was used to list all words in descending order of statistically significant difference between mega-report sets

  4. [4]

    conclude

    All statistically significant (p<0.001) terms were manually checked in the mega- reports to identify their typical contexts . For example, the context of “ conclude” might be “ produce summary ” but after reading all the mega-reports its (wider) context was assessed to be “insufficient information to make a strong conclusion”. Terms with multiple differen...

  5. [5]

    conclude

    Words were manually grouped into themes of related contexts. For example, the terms “conclude” and “abstract” were clustered into a “ Insufficient information in the abstract to make a strong conclusion” theme

  6. [6]

    evaluation

    Stages 4 and 5 were repeated until the results stabilised: contexts were revisited for large themes to give finer grained contexts if possible and isolated contexts were re -checked. For example , an initial “evaluation” theme was split into positive and negative evaluation themes, whereas the “Insufficient information in the abstract to make a strong con...

  7. [2000]

    mistake”, “omitted

    procedure ( Thelwall, 2021), and then clusters the words into themes . Although there are many different methods to compare different sets of texts, this has the advantage of automatically identifying differences with the backing of statistical evidence. In this sense it is the opposite of reflexive thematic analysis, which relies fully on the subjective ...

  8. [2025]

    Thelwall, M

    arXiv preprint arXiv:2504.09737. Thelwall, M. (2021). Word association thematic analysis: A social media text exploration strategy. Morgan & Claypool Publishers. Thelwall, M. (2025a). Is Google Gemini better than ChatGPT at evaluating research quality? Journal of Data and Information Science , 10(2), 1 –5. https://doi.org/10.2478/jdis-2025-0014 with exten...

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.