{"id":"5ef64bb3-00c4-42f8-8711-15abac57e6fa","arxiv_id":"2508.02442","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Five LLMs showed low human-LLM agreement and weak within-model stability when scoring 67 Italian psychology essays on a four-criterion rubric.","lead":"This study tested five large language models as graders of 67 Italian university psychology essays against human scores. The models showed weak agreement with human raters and unstable scores across repeated runs, suggesting they are not yet reliable for open-ended academic assessment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Without human inter-rater reliability, low human-LLM agreement cannot be attributed to LLM deficiency; the central validity claim rests on an unverified benchmark.","rationale":"The reader correctly identified the most load-bearing assumption: the human scores are treated as the validity benchmark for interpreting human-LLM agreement, yet the abstract provides no evidence on human rater reliability. This matters because the central conclusion is specifically about failing to replicate human judgment, not merely about scoring instability. Without knowing how much agreement exists between human raters, a low human-LLM kappa is ambiguous: it could reflect LLM error, human noise, or both. The paper does provide an independent reliability finding (low within-model Kendall's W across prompt replications), which partially supports the overall caution about LLM use in this setting. But that finding supports a reliability concern, not specifically the validity claim about disciplinary insight and contextual sensitivity. The small sample and single-course context are acknowledged by the authors through 'limited in scope,' so they do not constitute a hidden flaw. Given that the full text is garbled and only the abstract is available, the UNVERDICTED verdict is appropriate; my stress-test identifies no reason to move that verdict, and the reader's weakest-assumption point is the right one to press first.","tokens_in":7371,"tokens_out":4026,"duration_ms":47576,"concrete_test":"Compute inter-rater reliability (e.g., quadratic weighted kappa or ICC) among the human raters on each of the four rubric criteria from the 67 essays. If human-human agreement is comparably low to the reported human-LLM values, then the low QWK does not distinguish LLM performance from human-rater noise; if human-human QWK is high (e.g., >0.7), the low human-LLM QWK can be attributed to the models. This single check determines whether the reference standard is sufficient to support the conclusion that LLMs fail to replicate human judgment.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that current LLMs 'may struggle to replicate human judgment'—depends on human scores being a reliable and valid reference standard. The abstract reports consistently low, non-significant Quadratic Weighted Kappa values for human-LLM agreement but nowhere reports inter-rater reliability among the human raters for the four rubric criteria (Pertinence, Coherence, Originality, Feasibility). If human raters disagree with one another to a similar degree, low human-LLM agreement would be expected even for a model that captures the shared component of human judgment; the gap cannot then be credited to LLM-specific unreliability. The separate finding of weak within-model replication reliability (median Kendall's W < 0.30) does indicate LLM instability, so the overall conclusion is not fully unmoored. However, the interpretive leap to 'disciplinary insight and contextual sensitivity' is unsupported without data on the human benchmark's own reliability and without external outcome validation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an empirical study of five large language models (Claude 3.5, DeepSeek v2, Gemini 2.5, GPT-4, and Mistral 24B) used to score 67 Italian-language psychology essays on a four-criterion rubric (Pertinence, Coherence, Originality, Feasibility). Each model scored all essays across three prompt replications to measure intra-model stability. The abstract reports consistently low and non-significant human-LLM agreement (Quadratic Weighted Kappa), weak within-model reliability across replications (median Kendall's W < 0.30), and mixed inter-model agreement, with moderate convergence for Coherence and Originality but negligible concordance for Pertinence and Feasibility. The authors conclude that current LLMs may struggle to replicate human judgment in tasks requiring disciplinary insight and contextual sensitivity, and that human oversight remains critical in interpretive domains.","tokens_in":7545,"tokens_out":4408,"duration_ms":49800,"significance":"If the finding is robust, the paper provides a useful, falsifiable negative result in the ongoing evaluation of LLMs for automated essay scoring. Its strengths include the real-world higher-education context, the use of multiple LLMs, and the use of established reliability metrics (Quadratic Weighted Kappa and Kendall's W). The study also has a clear and honest limitation statement ('Although limited in scope'), which is appropriate. However, the significance is undercut by the absence of reported human inter-rater reliability, which is necessary to interpret low human-LLM agreement as an LLM-specific deficiency. The small sample (N=67) and a single discipline/course also limit the generalizability, though these limits are acknowledged.","major_comments":[{"comment":"The abstract reports low and non-significant human-LLM agreement as the central validity evidence, but it does not report human inter-rater reliability for the four rubric criteria. Without evidence that human raters agree with one another, low human-LLM agreement could reflect an unreliable reference standard rather than an LLM deficiency. The paper must report human inter-rater reliability (e.g., multi-rater kappa or intraclass correlation) and, if it is poor, temper the conclusion that LLMs 'struggle to replicate human judgment.'","section":"Abstract"},{"comment":"The abstract reports 'non-significant' Quadratic Weighted Kappa values without giving point estimates, confidence intervals, or a discussion of statistical power. With only 67 essays, non-significance may result from low power rather than the absence of agreement. The paper should report the exact QWK values, 95% confidence intervals or bootstrap intervals, and the smallest effect size detectable with this sample size.","section":"Abstract"},{"comment":"The abstract does not specify the exact prompts used for the three replications, the precise model versions (e.g., GPT-4-turbo vs. GPT-4-base), or the generation parameters (temperature, max tokens). Since the study's reliability claim depends on the sensitivity of responses to prompt replication, full transparency about the prompt construction and the scoring procedure is essential for evaluating whether the low Kendall's W is a property of the models or of the prompt engineering.","section":"Abstract / Methods"}],"minor_comments":[{"comment":"The terms 'reliability' and 'validity' are used in a general sense; the paper should state the specific definitions (e.g., reliability as consistency across replications, validity as agreement with human judgment) to avoid conflation.","section":"Abstract"},{"comment":"The four rubric criteria (Pertinence, Coherence, Originality, Feasibility) are likely ordinal, but the paper should explicitly justify the use of Quadratic Weighted Kappa and specify the weighting scheme; if the criteria are treated as nominal, a different agreement coefficient would be appropriate.","section":"Methods"},{"comment":"The abstract reports 'median Kendall's W < 0.30' without clarifying whether the median is computed across the five models or across the four criteria; the paper should present the full distribution, including per-model and per-criterion values, and discuss possible range restriction in the score distributions that may deflate W.","section":"Results"},{"comment":"The claim that 'current LLMs may struggle to replicate human judgment in tasks requiring disciplinary insight and contextual sensitivity' goes beyond the direct evidence of low agreement; the authors should restrict the conclusion to what the data show (low agreement and low replication reliability) and label the 'disciplinary insight' interpretation as one plausible explanation among others.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The full-text file supplied to me was corrupted in transmission and not legible; this report is based on the abstract and the reader's summary. The reader's stress-test concern about the missing human inter-rater reliability is valid and load-bearing. If the full paper contains human reliability data, the authors should highlight it in the abstract; if not, the validity interpretation is underdetermined. I recommend major revision to address the benchmark-reliability issue and the statistical reporting gaps."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a modest but genuine empirical contribution, not a field-changing result. The abstract describes a clean comparison of five named LLMs on 67 Italian psychology essays, with three prompt replications and a four-criterion rubric. Standard agreement measures (Quadratic Weighted Kappa, Kendall's W) are used, and the report of inter-model convergence for some criteria adds depth. The finding that within-model reliability across replications was weak (median Kendall's W < 0.30) is a useful caution for anyone thinking about using these models as standalone graders in interpretive domains.\n\nThe main soft spot is the one the stress-test note flags: the validity benchmark is human judgment, but the abstract never reports human inter-rater reliability. Without that, low human-LLM QWK values are hard to attribute to the LLMs. If human raters disagree with each other to a similar degree, low agreement between the model and any single rater is expected even for a model that captures the shared component of human judgment. The weak within-model reliability does support the \"don't use these alone\" conclusion, but the stronger gloss about lacking \"disciplinary insight and contextual sensitivity\" goes beyond what is reported. Also, 67 essays from one course is a small evidence base; the authors do say the scope is limited, which is honest.\n\nI could not assess the full text because the supplied version was garbled, so my read is abstract-level. If the full paper includes human rater reliability statistics, the main concern shrinks to a sample-size caveat. Nothing in the abstract suggests citation problems or slipshod method design.\n\nWho this is for: people working on automated essay scoring, assessment validity, or LLM evaluation in education. It is a serious empirical study that deserves referee time, but the referee should ask for the human rater data and a clearer separation between the replication-instability finding and the human-validity claim. Yes, send it to peer review.\n\nI would bring this to a reading group focused on LLM evaluation but not to a general one. I would likely cite it if writing about automated grading in interpretive domains.","headline":"Useful empirical caution about LLM essay scoring, but the central validity claim leans on an unverified human benchmark.","tokens_in":8030,"tokens_out":1622,"would_cite":true,"duration_ms":20252,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that five current LLMs, scoring Italian psychology essays with a four-criterion rubric, agree with human raters no better than chance on any criterion, and that their own scores are not reproducible across repeated runs.","keywords":["large language models","automated essay scoring","higher education","reliability","validity","quadratic weighted kappa","Kendall's W","Italian-language essays"],"falsifier":"Have at least two independent trained human raters score the same 67 essays on the same four criteria and compute pairwise human-human Quadratic Weighted Kappa. If human-human agreement turns out to be as low as human-LLM agreement, the paper's attribution of the gap to LLM unreliability would not hold; if human agreement is high, the negative result against the LLMs would be confirmed.","tokens_in":7215,"feed_emoji":"🎓","tokens_out":3546,"duration_ms":38731,"temperature":0.7,"pith_summary":"The paper tests whether five large language models can take over the human grading of open-ended university essays. It uses 67 essays from an Italian psychology course, scored on four rubric criteria, with each model grading the full set three times. The central finding is negative: human-LLM agreement is consistently low and non-significant, and each model's own scores vary too much across repetitions to be trusted. The authors conclude that current LLMs are not reliable standalone graders in tasks that need disciplinary insight and contextual sensitivity, and that human oversight must remain in the loop.","feed_headline":"Five LLMs fail to replicate human essay grading","feed_subtitle":"Human-LLM agreement stayed near zero on all four rubric criteria and scores varied across repeated runs.","key_machinery":"The argument is carried by a four-criterion analytic rubric—Pertinence, Coherence, Originality, Feasibility—applied to a corpus of 67 Italian-language essays from a university psychology course. Each of the five LLMs scored all essays three times (three prompt replications), and three agreement statistics were used: Quadratic Weighted Kappa ($\\kappa_q$) for human-LLM agreement, Kendall's coefficient of concordance $W$ for within-model stability across replications, and a second agreement analysis for inter-model convergence. These statistics are the load-bearing instrumentation: the conclusion of unsuitability rests entirely on the low $\\kappa_q$ and weak $W$ values.","core_discovery":"The study's claim is that, in a real higher-education setting, current LLMs cannot replicate human assessment of essay quality when the criteria call for discipline-specific understanding. Across 67 essays, five models, four rubric criteria, and three prompt replications, Quadratic Weighted Kappa values were low and non-significant, and within-model stability across replications was weak, with a median Kendall's $W$ below 0.30. The LLMs also showed systematic distortions, notably inflating the Coherence criterion, and they handled context-dependent dimensions inconsistently. Inter-model agreement was moderate only for Coherence and Originality and negligible for the other two criteria, which the authors read as evidence that agreement among models does not imply agreement with human judgment.","pith_inferences":["The paper's negative result would be more decisive if human inter-rater reliability were high; a natural next experiment is to compute human-human agreement on the same 67 essays, because if human raters also disagree substantially, the rubric itself would be the limiting factor rather than the LLMs.","A testable extension is to give the same models a rubric enriched with disciplinary examples and see whether human-LLM agreement rises; if it does, the deficiency may lie in prompt or rubric design rather than in a fixed model ceiling.","Because all five models showed similar non-significance, the bottleneck may be shared across model families rather than specific to any one architecture, suggesting a comparison with instruction-tuned models trained on domain scoring data as a follow-up."],"forward_implications":["None of the five tested models would be a trustworthy standalone grader for open-ended academic essays, because human-LLM agreement is low and non-significant on all four rubric criteria.","A model's score is not stable even for itself: repeated runs of the same essay produce weak within-model agreement, with a median Kendall's $W$ below 0.30.","The tendency to inflate the Coherence criterion means LLM scores would systematically overstate one component of essay quality, distorting any composite grade built from them.","Moderate inter-model agreement on Coherence and Originality does not rescue the approach, since agreement among models does not translate into agreement with human raters."],"supporting_citations":[],"fun_headline_variants":["LLM essay grading fails human agreement test","Five AI models can't replicate human essay scores","Low human-LLM agreement in real essay grading","AI essay scoring unstable across repeated runs","LLMs inflate coherence, miss nuance in essays"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the human scores used for comparison are themselves dependable; the paper does not report how well two human graders agreed on the same essays, so low human-LLM agreement could partly reflect human graders disagreeing with each other rather than the LLMs being wrong.","fun_headline_variants_meta":{"raw":{"variants":["LLM essay grading fails human agreement test","Five AI models can't replicate human essay scores","Low human-LLM agreement in real essay grading","AI essay scoring unstable across repeated runs","LLMs inflate coherence, miss nuance in essays"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1296,"prompt_tokens":913,"completion_tokens":383,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":327}},"tokens_in":529,"tokens_out":383,"duration_ms":5061,"temperature":1.0,"reasoning_tokens":327,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:56:32.004894+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have at least two independent trained human raters score the same 67 essays on the same four criteria and compute pairwise human-human Quadratic Weighted Kappa. If human-human agreement turns out to be as low as human-LLM agreement, the paper's attribution of the gap to LLM unreliability would not hold; if human agreement is high, the negative result against the LLMs would be confirmed.","supporting_citations":[],"review_version":1}