{"id":"bde0d5cd-14be-4bed-bfb2-1b83d354775f","arxiv_id":"2412.06651","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The Fobizz AI grading tool gives unstable and arbitrary grades and feedback, only rewards ChatGPT-written texts with top scores, and fails to detect false claims or nonsense submissions.","lead":"This study tested the Fobizz AI Grading Assistant, a tool that gives German teachers grades and feedback on student work. It found that repeated runs on the same assignment often produced very different grades and comments, and that following the tool's own advice did not help students improve their scores.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Testreihe B's no-improvement claim is confounded by single stochastic grading draws per iteration; repeated grades per version are needed.","rationale":"Good-faith read: the paper is a transparent usability audit, not a randomized trial. Its fatal-defect conclusion mainly rests on Testreihe A's repeated-run variance, which is strong and not in question. However, the reader's strongest_claim bundles two claims: randomness and non-improvement. The non-improvement claim is not established by the reported design because each iteration is a single stochastic sample. Since the paper itself teaches that the evaluator is highly noisy, it must average over repeated evaluations per version to separate signal from noise. The authors explicitly acknowledge the single-task, small-sample limitation in Section IV 'Weiterführende Untersuchungen', and that generalization concern is real but secondary; the sharper problem is the mismatched statistical treatment of a stochastic system. Thus I partially agree with the reader: their representativeness concern is valid, but the more acute issue is within-series inference. The proposed re-test would settle whether feedback implementation truly fails to improve scores or whether the effect was masked by grading noise. If the re-test confirms the null, the paper's qualitative conclusion remains largely intact, so the CONDITIONAL verdict stands.","tokens_in":47887,"tokens_out":6417,"duration_ms":78733,"concrete_test":"Re-run Testreihe B for Texts 1 and 10 with 15 independent grading passes per version: the original text, versions after each of 5 feedback-implementation rounds, and a ChatGPT-polished version, using fresh tool sessions to avoid memory effects. Compare score distributions (mean, SD, 95% CI) across versions. If human-improved versions show no significant mean increase over the original relative to the within-version SD, the no-improvement claim is supported; if they do, Mangel 5 should be downgraded and the ChatGPT-only claim re-tested with a high-quality human-written control graded the same way.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Testreihe B (Section III.5) uses exactly one score per iterative text version. Because the paper's own Testreihe A shows the tool's score for an unchanged submission can vary by 13 points (Abgabe 8) and is unstable for 8/10 texts, a single draw cannot measure whether incorporating feedback improved the text. The non-monotonic 'Irrfahrt' and the average slightly below baseline are exactly what noise alone would produce. This underdetermines the central claim that implementing the tool's suggestions does not raise grades (Mangel 5.a, classified as fatal in Table #tab:D:gravität). The same single-draw issue weakens the 'only ChatGPT gets near-perfect marks' claim, since the ChatGPT-final versions were also evaluated only once or a few times and no equally polished human-written control was graded repeatedly. The randomness findings in Testreihe A remain credible and independently support the 'do not use' conclusion; this critique is specifically about the feedback-effectiveness and ChatGPT-only subclaims.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates the Fobizz 'KI-Korrekturhilfe' (AI Grading Assistant) by running 10 simulated student submissions five times each through the tool (Testreihe A) and by iteratively rewriting two submissions in response to the tool's feedback over 7–12 rounds, with a final ChatGPT-based revision (Testreihe B). The authors report wide variance in numerical scores and qualitative feedback, including a score range of 1 to 14 points for one identical submission; failure to detect factual errors, nonsense, and nonconforming text length; internal contradictions in the feedback; and a pattern in which only ChatGPT-rewritten texts receive near-perfect scores. They conclude that the tool is unsuitable for classroom use, that its marketing is misleading, and that the deficiencies stem from inherent LLM limitations. The paper also includes an update section describing Fobizz's response and a rough re-test of the modified tool.","tokens_in":48017,"tokens_out":5828,"duration_ms":62559,"significance":"The topic is timely and consequential: Fobizz's tool is licensed by several German states (Table #tab:E:bundesländer), so independent functional evaluation is genuinely needed. The Series A design, with five independent grading runs of identical submissions, is an appropriate and simple method for exposing the stochastic variance of an LLM-based grading system, and the authors' publication of inputs and outputs in a material annex is a transparency strength. The finding that re-evaluating an unchanged submission can change the score by more than a school grade (Section III.1.a) is credible and, on its own, already raises serious doubts about the fitness of this tool for high-stakes decisions. However, the related claim that feedback incorporation does not improve scores (Mangel 5) and the claim that only ChatGPT texts achieve top marks (Mangel 6) rely on a much weaker single-draw design in Series B, which is insufficient given the demonstrated noise level.","major_comments":[{"comment":"In Testreihe B, each version of the text is graded exactly once. Testreihe A (Section III.1.a) shows that re-evaluating the same unchanged submission yields score swings of more than one grade for 3 of 10 submissions and a range of 1 to 14 points for one submission. Consequently, the non-monotonic trajectory in Fig. #fig:E:irrfahrt and the observation that the mean over iterations lies slightly below the baseline are exactly what sampling noise alone would produce. The resulting claim that incorporating the tool's feedback does not improve the grade (Mangel 5.a/b, Table #tab:D:gravität) is therefore underdetermined. The authors themselves note in Section IV ('Weiterführende Untersuchungen') that the Series B sample will be expanded in a future version, but the current paper states this claim categorically and classifies it as fatal. I recommend either repeating the iterative series with at least five independent evaluations per text version and reporting the resulting distributions, or explicitly downgrading the claim to a preliminary qualitative observation.","section":"§III.5.a / Fig. #fig:E:irrfahrt"},{"comment":"The claim that only ChatGPT-rewritten texts receive near-perfect evaluations is based on a single final evaluation of each ChatGPT version and on a small number of evaluations of the human-revised versions. Given the score variability documented in Series A, the difference between the human-revised and the ChatGPT-revised versions could be accounted for by chance. To support the strong claim that 'Bestnote nur durch ChatGPT möglich', the authors should grade the final human-revised and ChatGPT-revised versions multiple times (say, five times each) and compare the score distributions. Without such data, the inequality between human-revised and ChatGPT-revised versions is not established.","section":"§III.6"},{"comment":"The classification of Mangel 5.a/b as a 'fatales Gebrauchshindernis' and the associated recommendation to 'Tool nicht anbieten/verwenden' are load-bearing for the paper's overall condemnation. Because the evidence for this classification comes from a single stochastic draw per iteration, the categorical claim that the overall grade 'steigt nicht' when feedback is implemented is not yet supported. This is not a minor statistical detail: the conclusion that the feedback is didactically valueless is central to the paper's policy recommendation. Either the repeated-draw evidence should be supplied or the claim should be reformulated as a tentative finding that motivates further testing. The same caveat applies to the summary in the Executive Summary and Section V.1(5).","section":"§IV / Table #tab:D:gravität"}],"minor_comments":[{"comment":"The study uses a single task (a 150–250 word German position statement) whose simplicity is acknowledged later in Section IV. This scope limitation should be stated in the methodology itself, since it defines the generality of every subsequent finding about the tool.","section":"§II"},{"comment":"The update table for the 'verbesserten' tool (tested 21.02.2025) notes that the authors' own test scenario was incorporated into Fobizz's prompt as an example; this means the tool is being tested on a near-match to its prompt examples, which likely makes the results more favorable to the tool. This caveat should be displayed in the table itself or its caption so that readers do not interpret the 'Status quo' entries as directly comparable to the original August 2024 tests.","section":"§VI"},{"comment":"The figure would be much easier to read if the raw scores for each of the five runs per submission were provided in a table, together with the x-axis mapping of submission numbers; the current figure requires checking the text to reconstruct the individual data points.","section":"Fig. #fig:E:volatilität"},{"comment":"The statement that the average fluctuation is 'mehr als einen Punkt auf der 15-Punkte-Skala' would be more informative with a per-submission range or standard deviation; the present aggregate hides the fact that most submissions are fairly stable while one (Abgabe 8) spans from 1 to 14 points.","section":"§III.1.a"},{"comment":"Several web sources are cited without an access date; given that the paper's evidence is time-sensitive (the tool was modified on 15.12.2024), every online source should include an 'accessed on' date.","section":"Footnotes throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper would be substantially strengthened by a focused re-run of Testreihe B with repeated grading per iteration, which is feasible within the paper's design and would directly address the main methodological weakness. The pre-update Testreihe A results are already valuable and should remain the core of the paper, even if the feedback-incorporation claim is softened. The update section is useful but the authors should also consider clearly separating the analysis of the August 2024 version from the cursory re-test of the post-15.12.2024 version, as the two are not methodologically comparable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful, policy-relevant case study. The authors tested a commercial AI grading tool (Fobizz's \"KI-Korrekturhilfe\") used in German schools, and their central evidence—that the tool gives unstable numerical grades and qualitative feedback for identical submissions—is credible and well-documented. Testreihe A, where 10 submissions were each graded five times, shows 8/10 unstable scores, one ranging from 1 to 14 points on a 15-point scale. The screenshots and material appendix make it reproducible in principle. The update section, documenting the company's response and the authors' re-testing, is unusually candid.\n\nThe paper also documents other genuine problems: false statements in submissions go undetected, nonsense and off-topic texts get passable grades, word-count criteria are applied inconsistently, and feedback documents contain internal contradictions. These findings alone justify the conclusion that the tool should not be used as marketed.\n\nThe soft spot is Testreihe B. The authors claim that implementing the tool's feedback does not improve grades (Mangel 5.a) and classify it as a fatal barrier. But each iterative version is graded only once. Given the tool's demonstrated stochasticity, a single draw per version cannot distinguish real improvement from noise. The observed \"Irrfahrt\" and the average slightly below baseline are exactly what random variation would produce. This underdetermines the no-improvement claim. Same issue affects the \"only ChatGPT gets near-perfect marks\" finding: the ChatGPT-polished versions were evaluated once (or a few times), and no equally polished human-written control was graded repeatedly. That subclaim needs more runs and a proper control.\n\nThese weaknesses are real but localized. The randomness evidence in Testreihe A stands on its own and supports the \"do not use\" recommendation. I'd bring this to a reading group; it's a good example of how to empirically probe a marketed AI tool, and a caution for buyers of edtech licenses. The paper deserves peer review, with the expectation that the authors either re-run Testreihe B with repeated grading per iteration or soften the related claims. A serious referee will ask for that.","headline":"The core volatility finding is credible and policy-relevant; the feedback-no-improvement claim is undercut by single-draw iteration design.","tokens_in":48569,"tokens_out":3433,"would_cite":true,"duration_ms":35446,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Fobizz AI Grading Assistant produces grades and feedback that vary substantially when the same submission is graded repeatedly, and incorporating its own suggestions does not improve the grade; only ChatGPT-written texts receive…","keywords":["Fobizz AI Grading Assistant","automated essay grading","LLM feedback reliability","grade volatility","AI in education","AI-generated text detection","teacher workload automation","nonsense submission detection"],"falsifier":"Re-run the paper's Test Series A on the current Fobizz tool: if the 10 submissions each produce the same grade across five runs, or if the two iterative revision series never drop below their starting grades and reach 99 percent without a ChatGPT rewrite, then the paper's central claims would be contradicted.","tokens_in":47648,"feed_emoji":"🤖","tokens_out":9147,"duration_ms":95926,"temperature":0.7,"pith_summary":"This paper tests Fobizz's 'AI Grading Assistant', a preconfigured chatbot built on GPT-4 that offers teachers automated correction, feedback, and grade proposals for student texts. Using one German writing task and ten simulated student submissions, each graded five times, it claims that the tool's numerical grades and qualitative feedback vary substantially across repeated runs. In a second test series, revising two essays according to the tool's own feedback over 7-12 iterations did not raise the grade, and only a final ChatGPT rewrite produced near-perfect scores. The paper also reports that false factual claims and nonsense or off-topic submissions often pass undetected, that user-supplied grading criteria such as word count are applied unreliably, and that feedback documents contain invented errors and inconsistent category names. The authors conclude that these deficits stem from inherent properties of large language models, that a quick technical fix is not in sight, and that marketing the tool as objective and time-saving is misleading.","feed_headline":"Fobizz AI grader gave the same essay 1 to 14 points","feed_subtitle":"Why it matters: a single automated pass can hide grade swings of more than a school grade.","key_machinery":"The load-bearing machinery is the repeated-grading protocol: deliberately submitting the same text several times and comparing outputs. Because the underlying model samples from a probability distribution, a single run hides the spread that this protocol exposes. The complementary mechanism is the iterative feedback loop—revising an essay to implement the tool's own error list and then regrading it—which tests whether feedback has the monotonicity property a grading system must have. Together these procedures turn hidden sampling randomness, unreliable word counting, and absent AI-text detection into observable, reproducible failures.","core_discovery":"At the level of the tool's own functioning, the central discovery is that Fobizz's grading assistant fails the minimum consistency requirements of grading. In Test Series A, only two of ten essays received the same overall grade in all five runs; three essays moved by more than one school grade, and one nonsense essay ranged from 1 to 14 points on the 15-point scale. The qualitative feedback was just as unstable: the same essay could be praised as 'hervorragend' in one run and described with ambivalent phrasing in the next. In Test Series B, following the tool's correction suggestions did not increase the recommended grade—the average over all iterations stayed below the starting value—and the feedback oscillated, sometimes reversing its own previous instructions. The paper interprets these observations as consequences of the stochastic, pattern-completion nature of LLMs, not as bugs that ordinary software updates would fix.","pith_inferences":["Beyond the paper's test design, the 1-to-14 point swing on an off-topic essay implies that low-effort and refusal submissions are precisely the cases where a single LLM grade carries the least information; any evaluation of similar tools should over-sample those cases.","The near-perfect rating for a ChatGPT-written text raises a testable hypothesis the paper leaves open: the grader may systematically prefer text that matches its own generation style, so the 'best grade only via AI' effect could grow as graders and student tools converge on the same model family.","Because the observed volatility is a property of stochastic sampling, the paper's logic extends to any LLM-based grading assistant, not just this one; a practical fix would be to run each submission multiple times and report the spread before a teacher sees a single number."],"forward_implications":["A teacher who grades a submission once may unknowingly issue a grade that another run would move by more than one school grade.","A student who faithfully implements every feedback suggestion cannot expect the grade to rise; over 7-12 iterations the average stayed slightly below the starting score.","If students learn that only ChatGPT-assisted texts reach the top band, the tool itself becomes an incentive to outsource homework to AI.","Because the defects trace to inherent LLM limits, simply updating the prompt or model version will not remove them; systematic evaluation before purchase is the direct consequence."],"supporting_citations":[{"why":"Establishes that the Fobizz grading tool is built on GPT-4, the model whose stochastic output the paper identifies as the source of grade and feedback volatility.","marker":"37"},{"why":"Supplies the scientific consensus that automatic detection of AI-generated text is an unsolved problem, supporting the paper's finding that the 'not written by AI' criterion fails.","marker":"41"},{"why":"Fobizz's own blog post concedes that AI-generated text cannot be reliably detected, which the paper uses to show the interface nonetheless scores that criterion with false certainty.","marker":"42"},{"why":"Documents automation bias, the human tendency to over-trust automated judgments, which motivates the paper's warning that teachers may uncritically adopt unstable AI grades.","marker":"13"},{"why":"Points to a related study comparing LLM essay grades with teacher grades, providing context for the paper's examination of numerical grading reliability.","marker":"45"},{"why":"Provides the concept of technological solutionism that frames the paper's broader critique of AI as a quick fix for teacher shortage and systemic overburdening.","marker":"46"}],"fun_headline_variants":["Fobizz AI grader gave the same essay 1 to 14 points in five runs","Same essay scored 1 to 14 points on Fobizz's AI grading tool","Nonsense essay earned 1 to 14 points on Fobizz's AI grader","Fobizz AI grading is inconsistent: same essay gets 1 to 14 points","AI grading tool flunks consistency: same essay scored 1 to 14 points"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume that one short German position-writing task (150-250 words, with the authors' chosen criteria) is representative of how teachers use the tool and of typical student work; the paper itself flags this limitation in its section on further research.","fun_headline_variants_meta":{"raw":{"variants":["Fobizz AI grader gave the same essay 1 to 14 points in five runs","Same essay scored 1 to 14 points on Fobizz's AI grading tool","Nonsense essay earned 1 to 14 points on Fobizz's AI grader","Fobizz AI grading is inconsistent: same essay gets 1 to 14 points","AI grading tool flunks consistency: same essay scored 1 to 14 points"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001079,"raw_usage":{"total_tokens":4509,"prompt_tokens":936,"completion_tokens":3573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":3461}},"tokens_in":552,"tokens_out":3573,"duration_ms":27373,"temperature":1.0,"reasoning_tokens":3461,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:25:19.285047+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the paper's Test Series A on the current Fobizz tool: if the 10 submissions each produce the same grade across five runs, or if the two iterative revision series never drop below their starting grades and reach 99 percent without a ChatGPT rewrite, then the paper's central claims would be contradicted.","supporting_citations":[{"cited_title":"Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring","cited_arxiv_id":"2411.16337","evidence_quote":"Documents automation bias, the human tendency to over-trust automated judgments, which motivates the paper's warning that teachers may uncritically adopt unstable AI grades."}],"review_version":1}