{"id":"6319b139-cf93-4151-82d0-c908f94c4bef","arxiv_id":"2412.18719","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GPT-4 grading of short astronomy essays matched instructor grades when given a model answer and rubric, and was closer to the instructor than peer grading.","lead":"This paper tests whether GPT-4 can grade short science essays from three astronomy MOOCs as reliably as a human instructor. Across 120 student answers, GPT-4 given a model answer and a rubric matched instructor grades, and was closer to the instructor than peer grading on average.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim overrelies on null results: 'not statistically different' (p=1.000) is treated as 'approximately matched,' yet per-question data show gaps up to 0.71/4 points in History & Philosophy; no equivalence margin or direct error comparison to peer grading is reported.","rationale":"The reader identified the single-instructor gold-standard as the weakest assumption. That is an important external-validity limitation, but the more load-bearing problem is internal: the study's quantitative evidence for the central claim is a non-significant test treated as a positive result. If the aggregate p-value is driven by low power and heterogeneous question-level effects, then even the narrow claim 'GPT-4 grades like this instructor' is not established in the courses where it matters most. The History & Philosophy course is the most subjective and is also where Table 2 shows the largest gaps; the paper's narrative even flags Q3 and Q4 as poor agreement. A revised analysis with equivalence margins and direct error comparisons would either confirm the claim or show it depends on aggregation. The reader's conditional verdict is appropriate; my concern is a different reason for the condition, so the verdict is unchanged. No data/code are released, so the check requires the authors to provide the paired scores or run it themselves.","tokens_in":22930,"tokens_out":15646,"duration_ms":143100,"concrete_test":"Reanalyze the paired Instructor vs Prompt 2 data (and Instructor vs peer) with a pre-specified equivalence margin, e.g., +-5% of the assignment maximum (0.2 points for 4-point HPA, 0.45 for 9-point ETS, 0.5 for 10-point ABIO). Compute two one-sided tests (TOST) on the mean difference per course and overall, using the same bootstrap resampling, and report 90% confidence intervals. In addition, run a paired Wilcoxon signed-rank test on per-student absolute errors |Instructor - GPT4| vs |Instructor - peer|. If any course's 90% CI exceeds the margin — as the HPA Q2 mean gap of 0.71 points strongly suggests — or the absolute-error test is non-significant, the headline claims of 'approximately matched' and 'performs better than peer grading' are not supported by the current analysis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central assertion is that GPT-4 with an instructor answer plus rubric produces grades 'not statistically different' from the instructor and 'performs better than peer grading.' The support for both halves is a set of null-hypothesis tests (Friedman/Conover, Table 1) and descriptive per-question tables. A null result is not evidence of agreement: no equivalence margin is pre-specified, and 'p=1.000' after Bonferroni simply means the raw pairwise test was not significant. With only 10 answers per question, the per-question tests in Table 3 have very low power; the paper itself notes that per-class tests were too small to detect differences. More importantly, the paper's own descriptive statistics contradict the 'approximately matched for all three online courses' claim. In Table 2, HPA Q2 instructor mean is 2.39 vs Prompt 2 3.10 on a 4-point scale (a 17.8 percentage-point gap); HPA Q3 is 2.70 vs 3.20; Table 4 gives HPA RMS of 0.48 points (12% of max). The text admits agreement is 'poor' on HPA Q3 and Q4. These are educationally meaningful differences: on a 4-point rubric, 0.5-0.7 points is often a full grade category. The aggregate Friedman test can be non-significant even when such gaps exist, because it ranks within each answer and can cancel across heterogeneous questions. Similarly, 'GPT-4 performs better than peer grading' is inferred from peer grades differing significantly from the instructor while GPT-4 does not. This indirect comparison does not test whether GPT-4's absolute disagreement with the instructor is actually smaller than peers'. The Figure 2 dispersion comparison is descriptive and untested. Without an equivalence margin or a direct paired test on absolute errors, the practical claim that GPT-4 can replace peer grading is not established.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports an experiment in which GPT-4 graded short science writing assignments from three astronomy-related MOOCs, using three prompt conditions: instructor model answer only; instructor answer plus instructor-written rubric; and instructor answer plus an LLM-generated rubric. Grades from 120 students across 12 questions were compared with the original instructor grades and with Coursera peer grades, using Friedman/Conover tests, bootstrap confidence intervals, RMS differences, and per-student scatterplots. The paper claims that GPT-4 with an answer plus a rubric produced grades that were not statistically different from the instructor's and that GPT-4 outperformed peer grading in matching the instructor.","tokens_in":23266,"tokens_out":2668,"duration_ms":43371,"significance":"If the claim holds, the result is practically useful for large-scale MOOC assessment and for large introductory science courses, where instructor grading of writing is infeasible. The paper's methodological strengths include the use of non-parametric tests with post-hoc corrections, bootstrap-based standard errors, and a per-student analysis that goes beyond aggregate means. The appendices provide full question texts and rubrics, which helps reproducibility. However, the central claim rests on a single instructor as the gold standard, on null-hypothesis tests rather than equivalence tests, and on a small purposeful sample; these do not refute the claim but substantially limit its current evidentiary strength.","major_comments":[{"comment":"The claim that GPT-4 'approximately matched' instructor grades is based on non-significant Friedman/Conover comparisons (p = 1.000 in Table 1), with no pre-specified equivalence margin. A null result from a rank-based test does not demonstrate agreement, and the descriptive statistics in Table 2 show educationally meaningful gaps, e.g., History and Philosophy Q2 instructor mean 2.39 versus Prompt 2 mean 3.10 on a 4-point scale (0.71 points, about 18% of the maximum), and HPA Q3 2.70 versus 3.20. The text itself notes 'poor' agreement on HPA Q3 and Q4. The authors should report an equivalence test with a justified margin (e.g., ±0.5 rubric points or a Cohen's d bound) and should temper the 'approximately matched for all three online courses' conclusion where the data do not support it.","section":"Results, Tables 1 and 2; Large Language Model vs. Instructor"},{"comment":"The instructor who created the rubrics and model answers is also the sole human grader whose grades serve as the gold standard. As the Author Contributions state, 'Evaluation material for the MOOCs was created by Matthew Wenger, who also acted as the instructor grader for this project.' This creates a favorable alignment: the LLM is prompted with the same instructor's rubric and answer and is then compared with that same instructor's scores. The Discussion concedes 'we have assumed instructors to be perfect, when in fact they are fallible.' The manuscript should explicitly frame the result as reproducing one instructor's grading, not as accurate grading in general, and should discuss how rubric-derived idiosyncrasy could inflate the apparent agreement.","section":"Author Contributions and Discussion (gold-standard instructor)"},{"comment":"Per-question samples of 10 answers (12 questions total) with purposeful sampling to spread peer grades provide very low power for the per-question bootstrap p-values in Table 3, all of which exceed 0.05. The paper acknowledges that per-class tests were underpowered, but then uses the non-significant per-question results to support the claim of no difference. A non-significant p-value with n=10 cannot support 'no statistically significant difference' as evidence of agreement. The authors should either report effect sizes with confidence intervals for each question, or explicitly restrict their generalizability claim to the aggregate course-level comparison and present the per-question results only descriptively.","section":"Research Data and Results, Table 3"},{"comment":"The statement 'GPT-4 performs better than peer grading' is an indirect comparison: peer grades differ significantly from the instructor while GPT-4 grades do not (Table 1), and descriptive dispersions in Figures 2 and 3 are smaller for GPT-4 in some courses. No direct statistical test compares the absolute or squared errors of GPT-4 versus peer graders relative to the instructor. Also, the reported ICC of 0.92 is computed across graders without clarifying which graders are included, and the LLM was run once per prompt, so no estimate of LLM run-to-run variability is given. A direct paired comparison of |GPT-4 - instructor| versus |peer - instructor|, with appropriate clustering, would directly support the claimed superiority over peer grading.","section":"Comparisons for Individual Students and Reliability of Large Language Models and Peer Grading"}],"minor_comments":[{"comment":"The sentence 'The text for all writing assignment questions and grading rubrics used in this research study are provided in Appendix A' is followed by 'Model answers are available upon request'; the model answers should be included in the appendices or supplementary material for full reproducibility.","section":"Methods, Research Data"},{"comment":"Table 4 is introduced as 'Table 3' in the text ('shown in Table 3'), and the caption numbering is inconsistent; this should be corrected.","section":"Results, Table 4 and surrounding text"},{"comment":"Several rubric entries contain typographical errors, e.g., 'The write only includes one wavelength instead of two' and 'The writer correctly answers the question correctly'; these should be proofread.","section":"Appendix A, ETS rubrics"},{"comment":"The caption says 'Dashed lines in the histograms indicate the means of the three classes for the two measures of grade difference,' but it is unclear which classes the dashed lines correspond to; please clarify or label the lines.","section":"Results, Figure 2 caption"},{"comment":"Some references are incomplete (e.g., Bojic, Kovacevic, & Caparkapa, 2023 lacks publication details), and the relationship between this manuscript and the authors' earlier arXiv paper (Golchin et al., 2024) should be stated more explicitly to avoid duplicate-publication concerns.","section":"Introduction, Previous Work"}],"recommendation":"major_revision","confidential_remarks":"The paper overlaps substantially with the authors' earlier arXiv preprint (Golchin et al., 2024), and the current manuscript appears to be a revised journal submission; the editor may want to verify that the submission includes sufficient new material and proper attribution. The single-instructor design and the use of null results as evidence of equivalence are the main scientific risks, but both are fixable through reanalysis and more cautious framing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the headline result is not new. It's a re-analysis and extension of the same group's Golchin et al. 2024, and Table 3 is explicitly adapted from it. What is new is the per-student comparison, the dispersion plots, and the test of LLM-generated rubrics. Those are legitimate additions, but the core claim—GPT-4 with a model answer and rubric 'approximately matches' instructor grades—still leans on a null result.\n\nThe study does several things right. It uses real MOOC data: 120 students, 12 questions across three courses, graded by an instructor, peers, and GPT-4 under three prompt conditions. The statistical approach (Friedman with Conover post-hoc, bootstrap CIs) is appropriate for the skewed distributions. The authors are transparent about limitations: purposeful sampling, small per-question samples, and the fallibility of the instructor as gold standard. The LLM-generated rubric condition is a genuine extension and worth reporting.\n\nThe soft spots are real and central. 'Not statistically different' (p=1.000) is not evidence of agreement. No equivalence margin was specified. With 10 answers per question, the per-question tests have low power, and the paper itself notes this. More importantly, the descriptive statistics contradict the abstract's 'approximately matched for all three' claim. In Table 2, HPA Q2 instructor mean is 2.39 vs Prompt 2's 3.10 on a 4-point scale—a 0.71-point gap. HPA RMS in Table 4 is 0.48, just under half a point, which on a 4-point rubric is a meaningful grade-category shift. The text concedes agreement is 'poor' on HPA Q3 and Q4. The aggregate Friedman test can mask these gaps because it ranks within each answer and pools across heterogeneous questions. Similarly, 'GPT-4 performs better than peer grading' is inferred indirectly: peers differ from the instructor, GPT-4 doesn't. That is not a direct test of whether GPT-4's absolute disagreement is smaller. The per-student scatterplot (Fig. 2) is descriptive, with no test on the error distributions.\n\nThe circularity concern is also fair: the instructor who wrote the rubrics and model answers is the same person whose grades serve as the gold standard. So 'matching the instructor' may just mean GPT-4 learned that instructor's idiosyncrasies. That's acknowledged in the Discussion but not mitigated.\n\nWho is this for? Someone in educational assessment or MOOC research will want to read it as a proof-of-concept. It deserves a serious referee—the data are real, the question matters, and the limitations are addressable. But the authors should be pushed to reframe the claims as 'no detectable difference' with an equivalence margin, or better, to report direct paired comparisons of absolute errors and to release the data. As is, the abstract overstates what the design can support.","headline":"New data on an old claim: the paper repackages the prior GPT-4 grading result with per-student comparisons, but the 'match' still rests on null results and a single gold-standard grader.","tokens_in":23867,"tokens_out":2725,"would_cite":false,"duration_ms":23079,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4 can match an instructor's grades on short science essays — and beat peer grading — when its prompt includes a model answer and a rubric.","keywords":["GPT-4 grading","automated essay scoring","peer grading","massive open online courses","science writing assessment","rubric generation","large language models","astronomy education"],"falsifier":"Take the same 120 essays and have two independent instructors who did not write the rubrics grade them, then run GPT-4 with the same prompt template. If GPT-4 sits no closer to the rubric-writing instructor than the two outside instructors sit to each other, the claim that the LLM 'matches the instructor' collapses into the weaker claim that it matches one instructor's standards. A sharper version: re-run the study with rubrics and model answers written by a different instructor; if the agreement pattern flips, rubric authorship, not grading skill, is what the LLM is reproducing.","tokens_in":22740,"feed_emoji":"📝","tokens_out":13068,"duration_ms":102683,"temperature":0.7,"pith_summary":"This paper tries to establish that a large language model can grade short, content-based science writing as reliably as the instructor who designed the assignment, and more reliably than the peer grading that massive open online courses currently rely on. Across 120 student answers to 12 questions from three astronomy-themed MOOCs, GPT-4 produced grade distributions that were not statistically different from the instructor's whenever the prompt included both a model answer and a grading rubric, whether the rubric was the instructor's own or generated by GPT-4 itself. When the model received only a model answer, its grades differed significantly from the instructor's, and so did peer grades. If the result holds, automated grading could replace or lighten peer grading in low-stakes online courses and large introductory science classes, where a human cannot grade thousands of essays.","feed_headline":"Matches instructor's essay grades when given a rubric","feed_subtitle":"GPT-4's scores were indistinguishable from the instructor's; peer grades were not.","key_machinery":"The mechanism is a prompt template with three conditions that control how much grading guidance GPT-4 receives. Every condition embeds the student response and the assignment's total point value; the first adds only the instructor's model answer, the second adds the model answer plus the instructor's rubric, and the third asks GPT-4 to write its own rubric from the course description, the question, and the model answer before scoring. The rubric is the load-bearing element: leaving it out makes GPT-4's grades differ significantly from the instructor's, while adding either the instructor's rubric or an LLM-generated one brings the grades into statistical agreement ($p = 1.000$). The statistical framework is non-parametric — the Friedman test with Conover post-hoc comparisons and Bonferroni correction — chosen because the score distributions violate normality and homogeneity of variance, with bootstrap resampling used for per-question standard errors and p-values.","core_discovery":"The paper's central claim is that GPT-4, given an instructor's model answer plus a grading rubric, produces grades that are not statistically different from the instructor's on short science writing assignments, and that this beats peer grading. When the prompt supplied the instructor's rubric (the second condition) or a GPT-4-generated rubric (the third condition), the LLM's grades were statistically indistinguishable from the instructor's ($p = 1.000$ after a Bonferroni-corrected Conover post-hoc test on the Friedman analysis), while both peer grades and rubric-free GPT-4 grading differed significantly from the instructor ($p < 0.001$). The same pattern held for individual students: the mean instructor-minus-GPT-4 gap was near zero with smaller dispersion than the instructor-minus-peer gap in all three courses. An intraclass correlation coefficient of 0.92 across graders is offered as evidence of consistency, and the authors conclude that with a model answer and a rubric in hand, an LLM can stand in for the instructor in low-stakes settings and that LLM-generated rubrics match the utility of instructor rubrics.","pith_inferences":["Because the reference instructor wrote the model answers and rubrics, part of what GPT-4 matches may be that instructor's scoring style rather than an objective standard; an immediate test would be having a second independent instructor grade the same 120 essays and comparing how far GPT-4 sits from each of them.","A workflow the paper does not design but its data support is a human-audit loop: let the LLM grade everything with a short written justification, and have the instructor review only borderline cases, since the disagreement histograms imply the audit load would be small.","A testable extension beyond the paper: agreement between LLM and instructor should track rubric granularity, with point-by-point analytic rubrics yielding tighter agreement than holistic scales; the astronomy-versus-history-and-philosophy contrast is consistent with this but too small to prove it.","The paper's own concession that instructors are fallible implies the next benchmark should be agreement with a consensus of multiple expert graders, not any single instructor; if two experts disagree with each other by as much as GPT-4 disagrees with the reference instructor, 'matches the instructor' stops being evidence of accuracy."],"forward_implications":["Grading of writing in low-stakes MOOCs can be automated: an LLM equipped with a model answer and rubric can score thousands of essays in real time, removing the peer-grading burden that currently caps assignment frequency.","Assignments that lack rubrics need not be redesigned, because GPT-4-generated rubrics produced grades statistically indistinguishable from instructor-rubric grades, extending the approach to archival courses.","The method's accuracy degrades on open-ended, speculative questions, as the history and philosophy course showed the weakest agreement, so content-based factual prompts are the near-term application.","The paper states the same pipeline transfers to large university general-education science courses, where it plans to use the approach for formative assessment with grade reasoning attached."],"supporting_citations":[{"why":"Establishes the peer-grading baseline in the same astronomy MOOC, including the moderate instructor–peer correlation the LLM result is measured against.","marker":"Formanek et al., 2017"},{"why":"Supplies the finding that predicting rubric scores is essential to automated essay grading, motivating the rubric-equipped prompt conditions.","marker":"Kumar & Boulanger, 2020"},{"why":"Closest prior effort, using a transformer language model to validate peer-assigned essay scores in a MOOC.","marker":"Morris et al., 2023"},{"why":"Provides the rubric-design principles (analytic criteria, small subjective scales) that the course rubrics were built to follow.","marker":"Jonsson & Svingby, 2007"},{"why":"Bootstrap resampling, used for the per-question standard errors and p-values in Tables 2 and 3.","marker":"Efron, 1979"},{"why":"Defines the intraclass correlation coefficient used to report inter-rater reliability of 0.92 among graders.","marker":"Shrout & Fleiss, 1979"},{"why":"The Friedman test, the non-parametric omnibus test comparing instructor, peer, and LLM grade distributions.","marker":"Friedman, 1937; Friedman, 1940"},{"why":"Prior version of this study ('Large Language Models as MOOCs Graders'), from which the design and the bootstrap p-value table are adapted.","marker":"Golchin et al., 2024"},{"why":"Shows that automated short-answer grading systems can be fooled, the attack surface the paper names as an open limitation.","marker":"Filighera et al., 2020"}],"fun_headline_variants":["GPT-4 grades essays as reliably as instructors","AI grader matches instructor on science essays, beats peers","LLM grading rivals instructor accuracy with a rubric","GPT-4 with rubric matches instructor's essay scores","Automated essay grading: GPT-4 equals instructor reliability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a single instructor's grades are the gold standard, and that same instructor wrote the model answers and rubrics given to GPT-4, so 'matching the instructor' may amount to reproducing one person's scoring style rather than grading accurately in general; the paper itself concedes that instructors are fallible.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 grades essays as reliably as instructors","AI grader matches instructor on science essays, beats peers","LLM grading rivals instructor accuracy with a rubric","GPT-4 with rubric matches instructor's essay scores","Automated essay grading: GPT-4 equals instructor reliability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000653,"raw_usage":{"total_tokens":3034,"prompt_tokens":1029,"completion_tokens":2005,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":1930}},"tokens_in":645,"tokens_out":2005,"duration_ms":14086,"temperature":1.0,"reasoning_tokens":1930,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:32:01.557281+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same 120 essays and have two independent instructors who did not write the rubrics grade them, then run GPT-4 with the same prompt template. If GPT-4 sits no closer to the rubric-writing instructor than the two outside instructors sit to each other, the claim that the LLM 'matches the instructor' collapses into the weaker claim that it matches one instructor's standards. A sharper version: re-run the study with rubrics and model answers written by a different instructor; if the agreement pattern flips, rubric authorship, not grading skill, is what the LLM is reproducing.","supporting_citations":[],"review_version":1}