REVIEW 4 major objections 4 minor 4 references
Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis
T0 review · 4 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Across 65 studies, LLM essay scoring shows moderate to good agreement with human raters, but the evidence is too inconsistent to support a single summary estimate.
desk verdict Useful, well-run synthesis of 65 studies, but the '0.3-0.80 moderate-to-good' headline rests on post-hoc trimming and needs a transparent re-analysis before the claim is trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The synthesis is carried by the agreement indices — Quadratic Weighted Kappa, Pearson correlation, and Spearman's rho — which quantify how closely LLM scores track human ratings. The authors use the reported distribution of these indices, after removing extreme values below 0.10 and above 0.90, to establish a typical 0.30–0.80 range. The other load-bearing element is the coding framework that records study characteristics (model, rubric type, task, education level, language) to explain variation across studies.
What would settle it
A re-analysis of the included studies' raw or fully reported data, using only average agreement values and only studies with adequate human-rater reliability, would settle the claim: if the typical range fell below 0.30, the paper's 'moderate to good' conclusion would not hold.
Extended reading notes
Core claim
The central finding is that LLM-human agreement in automated essay scoring is generally moderate to good, with most reported agreement values between 0.30 and 0.80 after excluding extreme outliers. Because studies use incomparable metrics and reporting practices, the authors do not pool a single estimate; instead they establish that agreement is highly context-dependent, varying with model capability, prompting strategies, and whether essays come from standardized datasets or classroom settings. They also flag that many studies do not report human-rater reliability, so low agreement cannot always be attributed to the LLM rather than to inconsistent human scoring.
Load-bearing premise
The conclusion that agreement is typically moderate to good rests on treating the 65 studies' reported agreement numbers as comparable, even though some studies report only their best values and many do not verify how reliable the human raters were.
Editorial extensions
If this is right
- If the moderate-to-good range holds, LLM-based scoring can serve as a complement or screening tool in essay assessment, especially with advanced models and detailed rubrics.
- A meta-analysis that models study-level factors simultaneously would be the next step, because the observed patterns (model version, prompting, dataset) have not been tested together.
- Researchers should report human-rater reliability alongside LLM-human agreement, otherwise low agreement could be misread as a model failure.
- More studies are needed for K-12 writers and for essays in languages other than English before broad classroom use.
- Standardized reporting practices are needed to make future studies directly comparable.
Reading between the lines
- If studies tend to report their best kappa rather than the average, the true typical agreement is likely lower than 0.30–0.80; a re-analysis using raw data from the included studies could test this.
- The 'context-dependent' conclusion implies that LLM scoring should be validated locally for each rubric and population, rather than expected to generalize from one study.
- A practical extension would be a shared reporting standard requiring all agreement indices, both average and best, so future syntheses can pool estimates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a PRISMA 2020-guided research synthesis of 65 published and unpublished studies (2022–2025) that report quantitative agreement between LLM-generated and human essay scores. It describes the search, screening, coding, and distribution of study characteristics, and then summarizes reported agreement. The abstract and Conclusion claim that LLM-human agreement is 'generally moderate to good,' with indices 'mostly ranging between 0.30 and 0.80,' while emphasizing high context-dependence and the field's lack of standardized reporting. It does not compute pooled estimates, and it identifies missing human-rater reliability evidence as an important caveat.
Significance. If properly supported, the synthesis would be a useful contribution because it covers a fast-moving empirical literature more broadly than earlier reviews (e.g., Huang et al., 2025), includes preprints and non-English essays, and follows systematic-review procedures: PRISMA 2020, five databases plus hand-searching, dual screening with consensus, and pilot-tested dual coding. The authors also deserve credit for explicitly not computing pooled mean agreement and for acknowledging best-vs-average reporting and metric heterogeneity. The main conclusion, however, rests on a non-transparent post-hoc exclusion of extreme agreement values and on combining non-equivalent metrics; these issues must be repaired before the 'moderate to good' claim can be considered established.
major comments (4)
- [Results, 'Agreement levels'] The central descriptive claim is not supported by the procedure as reported. The text gives the full QWK range as -0.10 to 0.97 and the Pearson/Spearman range as 0.17 to 0.91, then says: 'After removing these extreme values (such as lower than 0.10 and higher than 0.90), we examined the remaining QWK values and found their typical range fell between the 0.3-0.80 range.' The same is applied to correlation coefficients. This is close to circular: excluding values below 0.10 and above 0.90 before summarizing will make any centrally located distribution look 'moderate to good.' The manuscript does not state how many estimates/studies were removed, whether removal was at the estimate or study level, what 'typical range' means (IQR, central 80%, informal inspection), or what the untrimmed distribution looks like. Please report the full distribution (e.g., histogram or sorted values), the numbe
- [Results and Discussion, 'Agreement levels' / 'Observed agreement on LLM-human agreement patterns'] The summary combines QWK, Pearson, and Spearman values into a single 0.30-0.80 band. These indices quantify different properties (weighted agreement vs. linear association), and their numerical values are not interchangeable, especially for ordinal scores with different numbers of categories. The paper correctly declines to pool estimates because of metric heterogeneity, but then the abstract and Discussion state that 'agreement indices ... mostly ranging between 0.30 and 0.80' and 'Most reported agreement values (QWK, Pearson correlation, or Spearman's rho) ranged from 0.3 to 0.8.' The aggregate band should be replaced by metric-specific distributions, or the claim should be explicitly restricted to conclusions that do not require comparability across metrics.
- [Results, 'Agreement levels' and Discussion, 'Observed agreement on LLM-human agreement patterns'] Unit-of-analysis and reporting-selection problems are acknowledged but not resolved. Some studies contribute many agreement estimates (e.g., per prompt, per trait, per model), and some report only the best value, while others report averages or ranges. The review does not define a study-level summary rule, nor does it report the number of estimates versus studies underlying the stated range. Because best-value reporting creates upward selection, the 'moderate to good' summary may overstate typical performance. Please state the analytical unit, report how many estimates each study contributes or at least the distribution of counts, and discuss the likely direction of bias in the summary.
- [Methods / Results / Appendix A] The synthesis does not provide a study-level data extraction table or supplement. Appendix A lists the 65 references, but the reader cannot see the extracted models, metrics, and numeric agreement values, nor check which values were considered extreme and removed. For a research synthesis whose main output is a descriptive distribution of agreement values, the underlying data should be made available, at minimum as a supplemental table. This is necessary for verification of the trimming, the heterogeneity claims, and any future re-analysis.
minor comments (4)
- [Introduction/Abstract] 'Pearson correlation' is listed as an 'agreement index'; correlation is an association/consistency measure, not agreement in the same sense as QWK. Consider more careful terminology.
- [Results, 'Agreement levels'] The phrase 'such as lower than 0.10 and higher than 0.90' is vague; exact thresholds and rationale are needed. Also define 'typical range' formally.
- [Appendix A] The 'About the Author' bios appended after the reference list appear to be a submission/formatting artifact rather than part of a scientific paper; remove them.
- [Results, 'Characteristics of Studies'] The text says the Gjorevski et al. (2025) study is counted for both US and Germany because of dual affiliations; this should be footnoted in Table 1 as well.
Circularity Check
No circularity: descriptive synthesis of external empirical studies, not a derivation from its own inputs.
full rationale
This paper is a PRISMA-guided systematic review and descriptive synthesis of 65 independently conducted empirical studies of LLM–human agreement in essay scoring. It does not derive a mathematical prediction from fitted parameters, and it does not invoke a self-citation chain to establish its central claim. The abstract's statement that agreement indices 'mostly ranged between 0.30 and 0.80' is a post-hoc descriptive summary of the reported values; the authors explicitly refrain from pooling estimates ('we did not compute pooled mean or median agreement estimates') and they candidly acknowledge heterogeneity and reporting inconsistencies. The most debatable step is the removal of 'extreme values (such as lower than 0.10 and higher than 0.90)' before stating the typical range; this is a transparency/robustness limitation, not a circular reduction, because the reported range is not defined as the trimming rule and is not equivalent to the input by construction. The paper also discloses that 'many studies did not provide sufficient evidence regarding human rating reliability' and that some studies report only their best agreement values, which are validity concerns rather than instances of the conclusion being forced by its own premises. There are no load-bearing self-citations: the included Li et al. (2024) is a different author group from the present first author, and no uniqueness theorem or prior work by the same authors is used to justify the findings. Overall, no step in the synthesis reduces to its own inputs by definition or by fitted-parameter recycling.
Assumptions & free parameters
free parameters (1)
- Extreme-value exclusion cutoff =
QWK below 0.10 and above 0.90 removed before stating the typical range
assumptions (3)
- domain assumption Agreement indices (QWK, Pearson, Spearman) reported across studies can be meaningfully compared and summarized descriptively despite metric differences
- standard math Landis & Koch (1977)/Altman (1991) thresholds define 'moderate to good' agreement for ordinal essay scores in this context
- domain assumption Search and screening identified all relevant published and unpublished English-language studies (PRISMA 2020)
Cite this review
Pith. "Pith review of Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis." pith.science (2026). https://pith.science/paper/FTQ6BEYB
@misc{pith2026251214561,
author = {Pith},
title = {Pith review of: Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis},
year = {2026},
howpublished = {\url{https://pith.science/paper/FTQ6BEYB}},
note = {Machine review of arXiv:2512.14561}
}
read the original abstract
Despite the growing promise of large language models (LLMs) in automated essay scoring (AES), empirical findings regarding their reliability compared to human raters remain mixed. Following the PRISMA 2020 guidelines, we synthesized 65 published and unpublished studies from January 2022 to August 2025 that examined agreement between LLM-generated scores and human ratings. Agreement levels varied substantially both across and within studies, with reported values spanning a wide range. Overall, the findings suggest that LLM-human agreement is highly context-dependent. Implications, challenges, and directions for future research are discussed.
Figures
Reference graph
Works this paper leans on
-
[243]
https://www.jstor.org/stable/20371545 Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 stateme...
arXiv 2021
-
[1513]
A., Albatarni, S., Eltanbouly, S., & Elsayed, T
https://doi.org/10.1080/14703297.2025.2469089 Mansour, W. A., Albatarni, S., Eltanbouly, S., & Elsayed, T. (2024). Can large language models automatically score proficiency of written essays? arXiv preprint arXiv:2403.06149. https://doi.org/10.48550/arXiv.2403.06149 Masikisiki, B., Marivate, V., & Hlope, Y. (2023). Investigating the efficacy of large lang...
arXiv 2025
-
[2024]
(Lecture Notes in Computer Science, Vol. 14830, pp. 276–283). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-64299-9_22 Chu, S. Y., Kim, J. W., Wong, B., & Yi, M. Y. (2024). Rationale behind essay scores: Enhancing S- LLM’s multi-trait essay scoring with rationale generated by LLMs. arXiv preprint arXiv:2410.14202. https://doi.org/10.48550...
work page Pith review arXiv doi:10.48550/arxiv.2410.14202 2024
-
[2058]
https://doi.org/10.1007/s10639-024-12891-w Cai, Y., Liang, K., Lee, S., Wang, Q., & Wu, Y. (2025). Rank-then-score: Enhancing large language models for automated essay scoring. arXiv preprint arXiv:2504.05736. https://doi.org/10.48550/arXiv.2504.05736 Chen, S., Lan, Y., & Yuan, Z. (2024). A multi-task automated assessment system for essay scoring. In Proc...
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.