{"id":"573322eb-29ee-4095-9445-65e6a0a8716f","arxiv_id":"2411.16337","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a small German essay study, OpenAI's o1 agreed best with teacher ratings, but all models scored content less reliably and gave higher marks than teachers.","lead":"This study compared five AI language models with 37 human teachers scoring 20 German student essays on ten criteria. The best model, OpenAI's o1, agreed with teachers on language quality but tended to score content more generously.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on an unvalidated teacher ground truth and untested model differences: with only 3–7 teacher ratings per essay and no reported inter-rater reliability, the Spearman coefficients on N=20 are too unstable to establish that o1 outperforms GPT-4.","rationale":"The reader's conditional verdict is appropriate, and the weakest assumption they identified is the same one I find most load-bearing: the teacher ratings used as ground truth have no reported reliability, and the comparative claim is not accompanied by any uncertainty estimate. The paper has real strengths: 37 real teachers, multiple criteria, repeated LLM runs, and a comparison of open- and closed-source models. These do not, however, compensate for the fact that each essay rests on only 3–7 teacher ratings and that the Spearman correlations are computed on 20 essays. An additional internal inconsistency supports the concern: Table 2 gives GPT-3.5 an ICC of .84 versus o1's .80, so the abstract's phrase that o1 'outperforms all other LLMs' is not even supported on the internal-consistency dimension. The paper's own limitation section acknowledges the human-rater variability problem. Therefore, the correct disposition is unchanged: conditional acceptance, with the required revision being an explicit teacher-reliability analysis and a test of whether the o1-versus-GPT-4 difference is statistically reliable, plus release of the raw ratings for reproducibility.","tokens_in":17278,"tokens_out":6153,"duration_ms":61208,"concrete_test":"Using the raw 1,090 teacher ratings (and the essay–teacher assignment), compute per-criterion inter-rater reliability (e.g., ICC(2,1) and ICC(2,k)). Then bootstrap the model ranking: for many resamples, draw 3–7 teacher ratings per essay with replacement, recompute each LLM's Spearman correlation with the teacher mean, and form a bootstrap/permutation confidence interval for r_o1 − r_GPT4. If the teacher ICC is low (say <0.5) or the difference interval includes zero, the abstract's comparative claim is not supported by this dataset. Releasing the raw ratings and analysis code would make this check reproducible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the teacher-average benchmark is a stable, valid ground truth. Section 3.2 reports only 3–7 raters per essay (mean 5.45) and gives no inter-rater reliability for the 37 teachers; Section 5.6 itself concedes 'the absence of a gold standard due to the inherent variability among human raters.' Averaging a handful of teacher ratings does not remove this problem, because the resulting mean is still a noisy estimate of any true essay quality. With N=20, the Spearman coefficients in Table 3 are noisy point estimates, and the paper never tests whether o1's r=.742 is significantly better than GPT-4's r=.575 or GPT-3.5's r=.418. The headline 'o1 outperforms all other LLMs' could therefore reflect which teachers happened to rate each essay. If teacher error is non-negligible, both the model ranking and the workload-reduction conclusion are not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a comparative evaluation of five LLMs (GPT-3.5, GPT-4, o1-preview, LLaMA 3-70B, Mixtral 8x7B) as automated essay scorers for German secondary-school essays. Twenty real student essays were rated by 37 teachers on ten criteria (content- and language-related), and the teacher averages were used as ground truth. Each LLM scored each essay ten times, and the mean of the ten runs was correlated with the teacher means using Spearman's rho; inter-run reliability was measured with ICC, and rating distributions were compared with Mann-Whitney U tests. The central claim is that the closed-source GPT models, especially o1, outperform open-source models, with o1 achieving Spearman's r = .74 with human ratings on the overall score and an ICC of .80, and that LLM-based assessment could support teachers particularly for language-related criteria.","tokens_in":17453,"tokens_out":4213,"duration_ms":40498,"significance":"If the reported effects were robust, the study would be a useful contribution to the emerging literature on LLM-based essay scoring in a non-English context, adding evidence on multidimensional criteria and run-to-run reliability. The authors assemble a real-world corpus of German essays with teacher ratings, make their prompting protocol transparent, and provide per-criterion analyses that go beyond holistic scores. The inclusion of open-source models and the explicit reporting of low ICC values for LLaMA 3 and Mixtral are informative. However, the statistical evidence for the headline superiority claim is thin, and several load-bearing assertions go beyond what the data can support. The credibility of the study would be substantially improved by addressing the reliability of the human benchmark and the significance of model differences.","major_comments":[{"comment":"The teacher-average ground truth is used to compute all Spearman correlations in Table 3, but no inter-rater reliability (e.g., teacher-level ICC) is reported for the 37 teachers, and the number of ratings per essay is only 3–7 (mean 5.45). The paper itself states in §5.6 that 'the absence of a gold standard due to the inherent variability among human raters' is a challenge. Without a variance-components analysis or a teacher ICC, it is not known how much of the variation in the teacher means reflects true essay quality versus sampling noise. With N=20 essays, the Spearman coefficients are highly unstable, and the conclusion that o1 'aligns with human assessments' is therefore not established. Please report teacher-level agreement (e.g., ICC per criterion) or at least the raw distribution of teacher ratings, and temper the wording of the central claim accordingly.","section":"§3.2 and §5.6"},{"comment":"The claim that the novel o1 model 'outperforms all other LLMs' rests on point estimates of Spearman's rho across 50 correlations (5 models × 10 criteria) with no correction for multiple comparisons and no significance test on the differences between models. For instance, the overall-score correlations are r = .742 (o1) and r = .575 (GPT-4); with N=20, this difference is not shown to be significant. A test for correlated correlations (e.g., Steiger's z) or bootstrap confidence intervals should be provided, and a multiple-comparison correction (Benjamini-Hochberg) should be applied to the p-values in Table 3. Without this, the model ranking is not supported beyond the level of descriptive statistics.","section":"Table 3 and §4.2"},{"comment":"The abstract states that 'the novel o1 model outperforms all other LLMs, achieving Spearman's r = .74 with human assessments in the overall score, and an internal consistency of ICC = .80.' This is internally inconsistent with Table 2, which shows ICC = .84 for GPT-3.5 and ICC = .73 for GPT-4, i.e., o1's ICC is not the highest among the closed-source models. The sentence implies o1 is best on both alignment and consistency, but the data only support the former (and even that is contested above). Please revise the abstract and Section 5.1 to state that o1's internal consistency is comparable to, but not higher than, that of GPT-3.5.","section":"Abstract and Table 2"},{"comment":"The paper filters out N=169 missing or out-of-range open-source model outputs before computing correlations and ICCs, but it does not report the distribution of these exclusions across essays or criteria. If LLaMA 3 and Mixtral failed disproportionately on certain essays or criteria, the low ICC values (Table 2) and weak correlations (Table 3) could be partly artifacts of the filtering rather than genuine evaluative behavior. Please report how many outputs were excluded per model, per essay, and per criterion, and run a sensitivity analysis (e.g., re-imputing or analyzing the unfiltered outputs).","section":"§3.3"}],"minor_comments":[{"comment":"The sentence 'Notably, their are differences between model versions' contains a typo: 'their' should be 'there.'","section":"§4.4"},{"comment":"The model naming is inconsistent: Table 2 uses 'GPT-o1' while the text and other tables use 'o1' or 'GPT-o1' inconsistently; please use a single label throughout.","section":"Table 2 and text"},{"comment":"The figure caption refers to 'the red line represents the linear least-squares regression,' but the paper reports Spearman correlations; please clarify whether the regression line is an illustrative fit and not the basis of the reported coefficients.","section":"Figure 3"},{"comment":"The phrase 'without limitating user confidence' should be 'without limiting user confidence.'","section":"§5.5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a reasonable exploratory study, but the abstract and Section 5.1 overclaim in ways that the reported statistics do not support. The most serious issue is the absence of teacher-level inter-rater reliability for a ground truth derived from only 3–7 ratings per essay; given the authors themselves acknowledge the lack of a gold standard in §5.6, the central 'o1 outperforms all other LLMs' claim is not yet established. I would be willing to look at a revised version that adds teacher ICC, significance tests for model differences, and a corrected abstract. The scope fit for LAK is acceptable, but the statistical reporting needs strengthening."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a genuinely useful addition to the educational NLP literature: it compares five LLMs against 37 teachers on 20 real German student essays across ten didactic criteria, including content and language dimensions. That alone fills a real gap, since most AES work is English, holistic, or limited to GPT models. The repeated-runs ICC design is a nice touch, and the qualitative finding that GPT models align best with language-related criteria is plausible and consistent with prior work. The soft spots are mostly statistical, and they are not minor. The teacher average is the ground truth, but the paper never reports inter-rater reliability among the 37 teachers. With only 3–7 ratings per essay, that average is a noisy target, so the Spearman correlations on N=20 are unstable point estimates. The paper also never tests whether o1's r=.742 is actually better than GPT-4's .575 or GPT-3.5's .418; with this sample size, those differences could easily be noise. And the abstract overstates the case by citing o1's ICC=.80 while Table 2 shows GPT-3.5's ICC=.84 — so o1 does not outperform all models on internal consistency. The lack of multiple-comparison correction across ten criteria and five models is a further concern, though it doesn't change the overall pattern. I don't think the central argument collapses. The direction — closed-source models, especially newer ones, track human ratings better on surface features — is credible and matches prior findings. But the workload-reduction conclusion in the abstract and conclusion is too strong for a 20-essay pilot. The paper would benefit from confidence intervals or bootstrap tests on the Spearman differences, a teacher inter-rater reliability figure, and a more careful wording of what can be concluded. This is a solid work-in-progress rather than a definitive benchmark. I'd send it to peer review, but with the expectation of major revision. A reviewer should push on the teacher ground truth and the significance testing. The right audience is researchers in educational NLP and learning analytics; they'll get value from the German real-world dataset and the multidimensional rubric design, even if the o1-specific claim needs to be read cautiously.","headline":"Useful pilot study of LLM essay scoring in German, but the o1 headline is over-stated given N=20 and an unvalidated teacher ground truth.","tokens_in":664,"tokens_out":2651,"would_cite":true,"duration_ms":37457,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims o1 aligns with averaged teacher ratings on German essay scoring (Spearman r = .74), is internally consistent (ICC = .80), and can support teachers on language-related criteria, while open-source models show no reliable…","keywords":["Large language models","Automated essay scoring","Teacher ratings","German student essays","Multidimensional assessment","o1","Reliability","Learning analytics"],"falsifier":"Recompute o1's Spearman correlation against each individual teacher's rating instead of the averaged teacher score for the same 20 essays; if the median such correlation is below, say, 0.3 or its confidence interval includes zero, the reported r = .74 is an artifact of averaging away teacher disagreement rather than true alignment.","tokens_in":17070,"feed_emoji":"📝","tokens_out":9932,"duration_ms":81371,"temperature":0.7,"pith_summary":"This paper tries to establish whether large language models can grade German student essays reliably enough to support teachers, using a ten-criterion rubric that separates content from language. On 20 real essays from Years 7 and 8, rated by 37 teachers, it finds that OpenAI's o1 model correlates with the averaged teacher ratings at Spearman r = .74 for the overall score and repeats its own ratings consistently (ICC = .80), while GPT-4 correlates at r = .58 and the open-source models LLaMA 3-70B and Mixtral 8x7B show no meaningful agreement. The alignment is strongest on language-related criteria such as spelling, literal speech, and verbal imagery, and weakest on content criteria such as introduction and main part. The paper concludes that LLM-based assessment can reduce teacher workload for language-related aspects if the models' tendency to score higher than teachers is calibrated, but that content quality still needs human judgment.","feed_headline":"AI essay rater o1 matches teacher scores at r=.74","feed_subtitle":"Best model agrees with teachers on language, but over-scores content; open-source models lag far behind.","key_machinery":"The argument is carried by a comparative scoring protocol rather than a single mathematical identity. Each essay is rated on ten six-point Likert criteria (five content, four language, and one overall judgment); the teacher ground truth is the average of the 3-7 ratings each essay received (mean 5.45), and each LLM score is the average of ten zero-shot runs of the same prompt at temperature 0.7. Alignment is measured with Spearman's rank correlation between LLM and teacher averages per criterion, and reliability is measured with the intraclass correlation coefficient (ICC) across the ten LLM runs. The paper also uses Mann-Whitney U tests to detect systematic leniency differences and inter-criteria correlation matrices to infer how much each criterion drives the overall score. The load-bearing comparison is the rank alignment between a single averaged LLM score and a single averaged teacher score on only 20 essays.","core_discovery":"The central claim, stated on the paper's own terms, is that o1 outperforms all other tested LLMs in essay scoring: it achieves Spearman's r = .74 with human assessments on the overall score and ICC = .80 internal consistency across ten repeated runs, and it is the only model with significant correlations in nine of ten criteria. The paper further claims that GPT models generally align with teachers on language-related criteria but systematically assign higher scores, so their overall ratings diverge from human strictness, while content-heavy criteria like plot logic and main part are precisely where agreement is weakest. Open-source models, by contrast, have very low reliability (ICC near -0.04 and 0.01) and near-zero correlations with teachers, making them unsuitable for this task in their current form. These findings support the conclusion that o1 is a promising assistive tool for reducing teacher workload in evaluating language-related aspects, provided its leniency bias is addressed.","pith_inferences":["If teacher ratings are as noisy as the small per-essay rater count suggests, the true alignment of o1 with any individual teacher may be much lower than .74; a fair test would compare o1 against each teacher separately rather than against the averaged consensus score.","The high inter-criterion correlations in GPT-3.5 and GPT-4 (above .86 and .75 respectively) indicate these models may be producing a single global quality impression repackaged into ten criterion scores, so criterion-level scores should not be read as independent diagnoses of student strengths.","The findings suggest a division-of-labor deployment: LLMs handle mechanics and style, teachers concentrate on content and feedback quality; this is a testable design for workload reduction studies.","One could extend the same protocol to argumentative essays or other languages to check whether o1's language-criterion advantage generalizes or is specific to German narrative writing."],"forward_implications":["o1 can serve as a reliable second reader for language-related scoring criteria such as spelling, expression, and literal speech, where its correlations with teacher averages are highest.","Teachers should still make the final call on content criteria such as plot logic and main part, where LLM-teacher agreement is weak and models differ significantly from human score distributions.","Any deployed system should aggregate multiple LLM runs rather than trust a single output, since single-run reliability is imperfect (o1 ICC = .80).","The leniency bias toward higher scores means calibration on teacher ratings would be needed before LLM scores are used as grades.","Open-source models LLaMA 3-70B and Mixtral 8x7B, as tested with this zero-shot prompt, are not reliable enough for essay scoring support."],"supporting_citations":[{"why":"Supplies the GPT-4 model used in the comparison and its API behavior.","marker":"[37]"},{"why":"Supplies the o1-preview model whose performance is the paper's main result.","marker":"[38]"},{"why":"Supplies the LLaMA 3-70B open-source model whose near-zero correlation anchors the closed-versus-open contrast.","marker":"[1]"},{"why":"Supplies the Mixtral 8x7B open-source model as the second open-weight comparison point.","marker":"[22]"},{"why":"Provides the ICC thresholds used to interpret internal consistency as moderate-to-good or poor.","marker":"[25]"},{"why":"Prior evidence that GPT-4 ratings are consistent, cited to situate the closed-source reliability finding.","marker":"[17]"},{"why":"Earlier finding that open-source models like LLaMA 2 lack human alignment, used to contextualize the paper's open-source result.","marker":"[48]"},{"why":"German short-answer grading baseline that motivates testing full free-text German essays beyond holistic scores.","marker":"[45]"}],"fun_headline_variants":["o1 matches teacher essay scores at r=.74","GPT o1 leads essay grading, open-source lags far behind","AI essay rater o1 aligns with teachers, over-scores content","Open-source LLMs fail essay scoring, o1 outperforms all","o1 best LLM for essay grading, matches teacher r=.74"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison assumes the average of three to seven teacher ratings per essay is a trustworthy gold standard, but the paper never reports how much the teachers agree with each other.","fun_headline_variants_meta":{"raw":{"variants":["o1 matches teacher essay scores at r=.74","GPT o1 leads essay grading, open-source lags far behind","AI essay rater o1 aligns with teachers, over-scores content","Open-source LLMs fail essay scoring, o1 outperforms all","o1 best LLM for essay grading, matches teacher r=.74"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000207,"raw_usage":{"total_tokens":1422,"prompt_tokens":991,"completion_tokens":431,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":341}},"tokens_in":607,"tokens_out":431,"duration_ms":4857,"temperature":1.0,"reasoning_tokens":341,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:13:20.725508+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute o1's Spearman correlation against each individual teacher's rating instead of the averaged teacher score for the same 20 essays; if the median such correlation is below, say, 0.3 or its confidence interval includes zero, the reported r = .74 is an artifact of averaging away teacher disagreement rather than true alignment.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the GPT-4 model used in the comparison and its API behavior."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the o1-preview model whose performance is the paper's main result."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Mixtral 8x7B open-source model as the second open-weight comparison point."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior evidence that GPT-4 ratings are consistent, cited to situate the closed-source reliability finding."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"German short-answer grading baseline that motivates testing full free-text German essays beyond holistic scores."}],"review_version":1}