REVIEW 3 major objections 5 minor 6 references
Research quality evaluation by AI in the era of Large Language Models: Advantages, disadvantages, and systemic effects
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This review argues that LLM-generated quality scores already rival or beat citation-based indicators as research quality signals, and that the main open questions are the AI's biases and its vulnerability to strategic abstract-writing.
desk verdict A clearly written review of LLM-based research quality evaluation, but the 'already superior to bibliometrics' claim outruns the evidence because no study in the paper compares LLM scores and citation indicators against the same human quality scores on the same outputs. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the 'indicator' concept from evaluative bibliometrics: a quantity that associates with research quality without purporting to measure it. For LLMs, the specific mechanism is prompt-based scoring—supplying a quality definition, usually the REF2021 four-point scale, with an article's title and abstract, asking the model for a score, and then averaging scores across many repetitions to reduce random variation. This turns the LLM into a text-based proxy for expert review, in contrast to citation counts, which are influence-based proxies.
What would settle it
A controlled experiment could settle it: take a set of articles with known expert quality scores, submit each to an LLM with its departmental affiliation and author names removed, and compare averaged scores to the expert scores. If the correlation collapses when identifying information is stripped, the public-information leakage explanation wins; if it holds, genuine text-based quality judgement is supported. A second check would modify abstracts (inflate or deflate quality claims) while keeping the underlying research identical and see whether LLM scores move accordingly.
Extended reading notes
Core claim
The paper's central claim is that 'LLM-based quality evaluations seem to be already superior to bibliometrics as research quality indicators, although with clear and not yet well understood biases.' The evidence comes from experiments feeding ChatGPT the titles and abstracts of published articles along with the UK REF quality definitions (rigour, originality, significance; 1* to 4*), then averaging repeated scores. Averaged ChatGPT-4o scores correlate with expert quality scores better than citation-based indicators do in most fields, the review reports, and the approach works in arts and humanities where citations are nearly useless, while failing mainly in clinical medicine. The author stresses that neither LLM scores nor citations measure quality; they are indicators, and any responsible use must weigh the known tradeoffs.
Load-bearing premise
The entire superiority claim rests on assuming that correlations between ChatGPT scores and departmental or expert REF scores reflect the AI's genuine ability to judge quality, rather than its exploitation of publicly available information about the universities' REF profiles or stylistic patterns in abstracts.
Editorial extensions
If this is right
- National research evaluation exercises like the REF could begin offering LLM scores alongside citation data as supporting evidence for panels.
- Quality assessment would extend to very recent papers and to arts, humanities, and social science fields where citation indicators are weak or useless.
- Researchers would face a new incentive: writing abstracts that optimise LLM-assessed originality, rigour, and significance claims, potentially encouraging overselling.
- Journal editors would have an incentive to permit exaggerated abstracts if LLM-based journal indicators replace impact factors, threatening the integrity of the public record.
- Any rollout would need bias audits for age, field, abstract length, gender, institution, and methodology preferences before scores are used.
Reading between the lines
- A direct test of the leakage hypothesis is feasible using preprints or anonymised abstracts, and the author's own work has not yet run that test; this is the fastest way to decide whether the claimed superiority is real.
- If LLM scores do track quality through abstract claims, then gaming countermeasures could include adversarial style normalisation or requiring LLMs to justify scores from specific sentences, which would be harder to fake than overall impressions.
- The same scoring mechanism could be turned into an audit instrument for human peer review, giving a stable, reproducible baseline against which reviewer bias and noise could be measured.
- Because LLMs can score any output type (books, proposals, datasets), the indicator role they may take over from bibliometrics is broader than citations, and so are the systemic effects on what kinds of work researchers choose to pursue.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reviews the potential of large language models (LLMs) as research quality indicators in comparison with bibliometrics, drawing on a mixture of small-, medium-, and large-scale empirical studies, mostly by the author and co-workers. It argues that LLM-based scores are more accurate than citation-based indicators (higher correlations with human scores in most fields), have broader coverage of fields and recent years, and can in principle assess more quality dimensions. It also discusses similarities (both are indirect indicators, not direct measures) and disadvantages (unknown biases, lower transparency, limited research into use contexts, and greater risk of gaming via abstract manipulation). It concludes that LLM scores have technical potential to complement or surpass bibliometrics but that there are too many unknowns for immediate use in important contexts, recommending further research on biases, limitations, and gaming before minor supporting roles are considered.
Significance. The paper addresses a timely and important question for the research evaluation community. Its principal strength is the systematic mapping of the systemic and normative dimensions of LLM-based indicators, especially the discussion of gaming incentives on abstracts and journal editorial practices, which goes beyond simple accuracy comparisons. The paper is also commendably explicit about the tentativeness of the evidence, using hedged language such as 'not conclusive' and 'suggestive'. The reviewed evidence includes both published and preprint studies, and the author acknowledges the leakage confound in the large-scale REF study. However, the central comparative claim of superiority over bibliometrics rests on indirect comparisons across different samples, levels of analysis, and outcome measures, and much of the supporting evidence is author-affiliated and not independently replicated. If the claim is taken as a hypothesis needing further testing, the paper is a valuable agenda-setting review; as an evidence-based conclusion, it currently overreaches.
major comments (3)
- ['Advantages: More accurate'] The claim that LLM scores are 'more accurate' than bibliometrics is not established by the cited evidence because the comparisons are not on the same footing. The key support (Thelwall & Yaghi, 2024b) correlates ChatGPT scores with departmental average REF2021 scores, whereas the bibliometric correlations cited for comparison (Thelwall et al., 2023b, 2023c) are at the level of individual articles or journals against expert scores. Correlations with departmental averages benefit from aggregation that removes individual-rater noise, so higher correlations on that basis do not demonstrate that an article-level indicator is more accurate. The paper should either report a same-sample, same-level comparison or explicitly reframe the conclusion as a hypothesis pending such a test.
- ['LLM-generated research quality indicators'] The paper acknowledges a plausible leakage confound in the main large-scale study: 'ChatGPT might have leveraged public information about departmental REF quality profiles when scoring individual articles.' This confound is not merely a peripheral weakness; it threatens the central evidence for the 'more accurate' and 'greater coverage' claims, including the notable claim that ChatGPT is 'useless only for clinical medicine.' Because the articles were selected from high- and low-scoring departments and the scores are averaged over 30 iterations, the model could plausibly exploit the public departmental averages. The manuscript should either present a direct test that rules out this mechanism (for example, comparing scores for articles from departments with similar profiles, or using a blinded protocol) or substantially soften the superiority conclusion in light of the unresolved confound.
- ['Advantages: Greater coverage of science'] The comparison that citations are 'useless' for arts and humanities while ChatGPT is 'only useless for clinical medicine' is based on different studies with different operationalizations of 'useless' (correlation thresholds, levels of aggregation, and output sets). For instance, Thelwall et al. (2023b) examine article-level citation correlations with expert scores, while Thelwall & Yaghi (2024b) examine departmental-average correlations. The claim may be true, but the evidence as presented does not support the comparative assertion. The authors should either provide a comparable analysis across fields for both indicators or present this as a preliminary observation requiring direct comparison.
minor comments (5)
- [Throughout] There are inconsistent renderings of the model name: 'ChatGPT 40-mini' appears in the medium-scale study description and 'ChatGPT 4o-mini' elsewhere; please use a consistent notation.
- [LLM-generated research quality indicators] The reference 'Thelwall & Yaghi, 2024a' is cited as evidence for the claim that LLMs can assess multiple quality dimensions, but this reference is listed as 'Submitted' rather than published; the dependence on a non-peer-reviewed manuscript should be made explicit, or the claim should be attributed to a published source.
- [Conclusion] The sentence 'LLM-based quality evaluations being already superior to bibliometrics' is stronger than the hedged language elsewhere in the paper ('not conclusive', 'suggestive'); rephrasing to 'potentially superior' or 'may already be superior' would better align the conclusion with the stated limitations.
- [Disadvantages] The phrase 'as of September 2024' in the opening of the Disadvantages section is inconsistent with later statements referencing February 2025; please update the temporal reference for consistency.
- [References] Several key empirical references are non-peer-reviewed preprints (e.g., Thelwall & Yaghi, 2024b; Thelwall & Jiang, 2025; Thelwall & Kurt, 2024). While preprints are acceptable in a fast-moving field, the manuscript should note the status of these sources in the text or reference list so readers can gauge the evidence base.
Circularity Check
No circularity; the review's claims are empirical summaries with acknowledged limitations, not derivations that reduce to their inputs.
full rationale
This is a narrative review rather than a formal derivation, so most circularity patterns do not apply. The central claim that LLM scores are 'already superior to bibliometrics' is supported by empirical studies, including several by the author, but those studies are external, falsifiable investigations using REF human scores as ground truth; they are not fitted parameters renamed as predictions, nor does the paper define LLM quality in terms of the claimed conclusion. The paper explicitly concedes the main confound ('ChatGPT might have leveraged public information about departmental REF quality profiles when scoring individual articles'), which is an evidential limitation, not a circular reduction. Heavy self-citation is present but constitutes normal review practice and is supported by independent empirical checks and external studies; it does not create a load-bearing self-referential argument. No equation or definition makes the conclusion equivalent to its inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption Human expert scores (e.g., REF quality ratings) are a valid ground truth for research quality.
- domain assumption Correlation with human scores is an appropriate measure of indicator quality.
- domain assumption LLM scores based on titles and abstracts reflect the content's quality rather than stylistic artifacts or leaked metadata.
Cite this review
Pith. "Pith review of Research quality evaluation by AI in the era of Large Language Models: Advantages, disadvantages, and systemic effects." pith.science (2026). https://pith.science/paper/5QMI6A56
@misc{pith2026250607748,
author = {Pith},
title = {Pith review of: Research quality evaluation by AI in the era of Large Language Models: Advantages, disadvantages, and systemic effects},
year = {2026},
howpublished = {\url{https://pith.science/paper/5QMI6A56}},
note = {Machine review of arXiv:2506.07748}
}
read the original abstract
Artificial Intelligence (AI) technologies like ChatGPT now threaten bibliometrics as the primary generators of research quality indicators. They are already used in at least one research quality evaluation system and evidence suggests that they are used informally by many peer reviewers. Since using bibliometrics to support research evaluation continues to be controversial, this article reviews the corresponding advantages and disadvantages of AI-generated quality scores. From a technical perspective, generative AI based on Large Language Models (LLMs) equals or surpasses bibliometrics in most important dimensions, including accuracy (mostly higher correlations with human scores), and coverage (more fields, more recent years) and may reflect more research quality dimensions. Like bibliometrics, current LLMs do not "measure" research quality, however. On the clearly negative side, LLM biases are currently unknown for research evaluation, and LLM scores are less transparent than citation counts. From a systemic perspective, the key issue is how introducing LLM-based indicators into research evaluation will change the behaviour of researchers. Whilst bibliometrics encourage some authors to target journals with high impact factors or to try to write highly cited work, LLM-based indicators may push them towards writing misleading abstracts and overselling their work in the hope of impressing the AI. Moreover, if AI-generated journal indicators replace impact factors, then this would encourage journals to allow authors to oversell their work in abstracts, threatening the integrity of the academic record.
Reference graph
Works this paper leans on
-
[1]
Baccini, A., De Nicolao, G., & Petrovich, E. (2019). Citation gaming induced by bibliometric evaluation: A country-level comparative analysis. PLoS One, 14(9), e0221212. Barnett, A., Allen, L., Aldcroft, A., Lash, T. L., & McCreanor, V. (2024). Examining uncertainty in journal peer reviewers’ recommendations: a cross -sectional study. Royal Society Open S...
work page 2019
-
[6]
Detecting expressions with multimodal transformers
https://doi.org/10.3389/fncom.2012.00063. Kelly, C. D., & Jennions, M. D. (2006). The h index and career assessment by numbers. Trends in Ecology & Evolution, 21(4), 167-170. Kordzadeh, N., & Ghasemaghaei, M. (2022). Algorithmic bias: review, synthesis, and future research directions. European Journal of Information Systems, 31(3), 388-409. Kousha, K., & ...
work page Pith review arXiv 2006
-
[9]
Wilsdon, J., Allen, L., Belfiore, E., Campbell, P., Curry, S., Hill, S., & Johnson, B. (2015). The metric tide: Independent review of the role of metrics in research assessment and management. https://www.ukri.org/publications/review-of-metrics-in-research- assessment-and-management/ Winker, M. (2015). The promise of post ‐publication peer review: how do ...
arXiv 2015
-
[40]
https://doi.org/10.1080/08989621.2014.899909. Devlin, J. (2018). B ERT: Pre -training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. de Winter, J. (2024). Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstracts. Scientometrics,...
-
[244]
Langfeldt, L., Nedeva, M., Sörlin, S., & Thomas, D. A. (2020). Co -existing notions of research quality: A framework to study context -specific understandings of good research. Minerva, 58(1), 115-137. Liang, W., Izzo, Z., Zhang, Y., Lepp, H., Cao, H., Zhao, X., & Zou, J. Y. (2024). Monitoring ai - modified content at scale: A case study on the impact of ...
arXiv 2020
-
[2024]
(pp. 9340-9351). Zhuang, Z., Chen, J., Xu, H., Jiang, Y., & Lin, J. (2025). Large language models for automated scholarly paper review: A survey. arXiv preprint arXiv:2501.10326
arXiv 2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.