REVIEW 3 major objections 5 minor 14 references
Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The Fobizz AI Grading Assistant produces grades and feedback that vary substantially when the same submission is graded repeatedly, and incorporating its own suggestions does not improve the grade; only ChatGPT-written texts receive…
desk verdict The core volatility finding is credible and policy-relevant; the feedback-no-improvement claim is undercut by single-draw iteration design. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the repeated-grading protocol: deliberately submitting the same text several times and comparing outputs. Because the underlying model samples from a probability distribution, a single run hides the spread that this protocol exposes. The complementary mechanism is the iterative feedback loop—revising an essay to implement the tool's own error list and then regrading it—which tests whether feedback has the monotonicity property a grading system must have. Together these procedures turn hidden sampling randomness, unreliable word counting, and absent AI-text detection into observable, reproducible failures.
What would settle it
Re-run the paper's Test Series A on the current Fobizz tool: if the 10 submissions each produce the same grade across five runs, or if the two iterative revision series never drop below their starting grades and reach 99 percent without a ChatGPT rewrite, then the paper's central claims would be contradicted.
Extended reading notes
Core claim
At the level of the tool's own functioning, the central discovery is that Fobizz's grading assistant fails the minimum consistency requirements of grading. In Test Series A, only two of ten essays received the same overall grade in all five runs; three essays moved by more than one school grade, and one nonsense essay ranged from 1 to 14 points on the 15-point scale. The qualitative feedback was just as unstable: the same essay could be praised as 'hervorragend' in one run and described with ambivalent phrasing in the next. In Test Series B, following the tool's correction suggestions did not increase the recommended grade—the average over all iterations stayed below the starting value—and the feedback oscillated, sometimes reversing its own previous instructions. The paper interprets these observations as consequences of the stochastic, pattern-completion nature of LLMs, not as bugs that ordinary software updates would fix.
Load-bearing premise
The results assume that one short German position-writing task (150-250 words, with the authors' chosen criteria) is representative of how teachers use the tool and of typical student work; the paper itself flags this limitation in its section on further research.
Editorial extensions
If this is right
- A teacher who grades a submission once may unknowingly issue a grade that another run would move by more than one school grade.
- A student who faithfully implements every feedback suggestion cannot expect the grade to rise; over 7-12 iterations the average stayed slightly below the starting score.
- If students learn that only ChatGPT-assisted texts reach the top band, the tool itself becomes an incentive to outsource homework to AI.
- Because the defects trace to inherent LLM limits, simply updating the prompt or model version will not remove them; systematic evaluation before purchase is the direct consequence.
Reading between the lines
- Beyond the paper's test design, the 1-to-14 point swing on an off-topic essay implies that low-effort and refusal submissions are precisely the cases where a single LLM grade carries the least information; any evaluation of similar tools should over-sample those cases.
- The near-perfect rating for a ChatGPT-written text raises a testable hypothesis the paper leaves open: the grader may systematically prefer text that matches its own generation style, so the 'best grade only via AI' effect could grow as graders and student tools converge on the same model family.
- Because the observed volatility is a property of stochastic sampling, the paper's logic extends to any LLM-based grading assistant, not just this one; a practical fix would be to run each submission multiple times and report the spread before a teacher sees a single number.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates the Fobizz 'KI-Korrekturhilfe' (AI Grading Assistant) by running 10 simulated student submissions five times each through the tool (Testreihe A) and by iteratively rewriting two submissions in response to the tool's feedback over 7–12 rounds, with a final ChatGPT-based revision (Testreihe B). The authors report wide variance in numerical scores and qualitative feedback, including a score range of 1 to 14 points for one identical submission; failure to detect factual errors, nonsense, and nonconforming text length; internal contradictions in the feedback; and a pattern in which only ChatGPT-rewritten texts receive near-perfect scores. They conclude that the tool is unsuitable for classroom use, that its marketing is misleading, and that the deficiencies stem from inherent LLM limitations. The paper also includes an update section describing Fobizz's response and a rough re-test of the modified tool.
Significance. The topic is timely and consequential: Fobizz's tool is licensed by several German states (Table #tab:E:bundesländer), so independent functional evaluation is genuinely needed. The Series A design, with five independent grading runs of identical submissions, is an appropriate and simple method for exposing the stochastic variance of an LLM-based grading system, and the authors' publication of inputs and outputs in a material annex is a transparency strength. The finding that re-evaluating an unchanged submission can change the score by more than a school grade (Section III.1.a) is credible and, on its own, already raises serious doubts about the fitness of this tool for high-stakes decisions. However, the related claim that feedback incorporation does not improve scores (Mangel 5) and the claim that only ChatGPT texts achieve top marks (Mangel 6) rely on a much weaker single-draw design in Series B, which is insufficient given the demonstrated noise level.
major comments (3)
- [§III.5.a / Fig. #fig:E:irrfahrt] In Testreihe B, each version of the text is graded exactly once. Testreihe A (Section III.1.a) shows that re-evaluating the same unchanged submission yields score swings of more than one grade for 3 of 10 submissions and a range of 1 to 14 points for one submission. Consequently, the non-monotonic trajectory in Fig. #fig:E:irrfahrt and the observation that the mean over iterations lies slightly below the baseline are exactly what sampling noise alone would produce. The resulting claim that incorporating the tool's feedback does not improve the grade (Mangel 5.a/b, Table #tab:D:gravität) is therefore underdetermined. The authors themselves note in Section IV ('Weiterführende Untersuchungen') that the Series B sample will be expanded in a future version, but the current paper states this claim categorically and classifies it as fatal. I recommend either repeating the iterative series with at least five independent evaluations per text version and reporting the resulting distributions, or explicitly downgrading the claim to a preliminary qualitative observation.
- [§III.6] The claim that only ChatGPT-rewritten texts receive near-perfect evaluations is based on a single final evaluation of each ChatGPT version and on a small number of evaluations of the human-revised versions. Given the score variability documented in Series A, the difference between the human-revised and the ChatGPT-revised versions could be accounted for by chance. To support the strong claim that 'Bestnote nur durch ChatGPT möglich', the authors should grade the final human-revised and ChatGPT-revised versions multiple times (say, five times each) and compare the score distributions. Without such data, the inequality between human-revised and ChatGPT-revised versions is not established.
- [§IV / Table #tab:D:gravität] The classification of Mangel 5.a/b as a 'fatales Gebrauchshindernis' and the associated recommendation to 'Tool nicht anbieten/verwenden' are load-bearing for the paper's overall condemnation. Because the evidence for this classification comes from a single stochastic draw per iteration, the categorical claim that the overall grade 'steigt nicht' when feedback is implemented is not yet supported. This is not a minor statistical detail: the conclusion that the feedback is didactically valueless is central to the paper's policy recommendation. Either the repeated-draw evidence should be supplied or the claim should be reformulated as a tentative finding that motivates further testing. The same caveat applies to the summary in the Executive Summary and Section V.1(5).
minor comments (5)
- [§II] The study uses a single task (a 150–250 word German position statement) whose simplicity is acknowledged later in Section IV. This scope limitation should be stated in the methodology itself, since it defines the generality of every subsequent finding about the tool.
- [§VI] The update table for the 'verbesserten' tool (tested 21.02.2025) notes that the authors' own test scenario was incorporated into Fobizz's prompt as an example; this means the tool is being tested on a near-match to its prompt examples, which likely makes the results more favorable to the tool. This caveat should be displayed in the table itself or its caption so that readers do not interpret the 'Status quo' entries as directly comparable to the original August 2024 tests.
- [Fig. #fig:E:volatilität] The figure would be much easier to read if the raw scores for each of the five runs per submission were provided in a table, together with the x-axis mapping of submission numbers; the current figure requires checking the text to reconstruct the individual data points.
- [§III.1.a] The statement that the average fluctuation is 'mehr als einen Punkt auf der 15-Punkte-Skala' would be more informative with a per-submission range or standard deviation; the present aggregate hides the fact that most submissions are fairly stable while one (Abgabe 8) spans from 1 to 14 points.
- [Footnotes throughout] Several web sources are cited without an access date; given that the paper's evidence is time-sensitive (the tool was modified on 15.12.2024), every online source should include an 'accessed on' date.
Circularity Check
No significant circularity: the Fobizz tool evaluation is an empirical measurement study whose central claims rest on direct test observations, not on definitions, fits, or self-citations.
full rationale
This is an empirical evaluation, not a derivation. The central claims—that repeated grading of the same submission yields volatile scores, that factual errors and nonsense go undetected, that criterion implementation is unreliable, that iterative incorporation of feedback does not improve scores, and that only ChatGPT-rewritten texts approach perfect marks—are supported by direct, documented test runs in the paper (e.g., five runs per text in Testreihe A, iterative chains in Testreihe B). None of these results is obtained by fitting a parameter or by defining a quantity in terms of the target result. The only self-references in the paper are background citations to Mühlhoff's prior work on data power and AI ethics (e.g., in footnotes 50-51) and a citation to an external arXiv study; these are not load-bearing for the tool evaluation. The update section even notes that Fobizz incorporated the authors' test scenario into the new prompt, which would bias the follow-up assessment in favor of the tool—a transparency step opposite to circular reasoning. A methodological concern about Testreihe B (single stochastic grading draw per iteration) affects the strength of the no-improvement inference, but it is a statistical-power issue, not a circularity issue. Therefore no circular steps are identified.
Assumptions & free parameters
assumptions (3)
- domain assumption The Fobizz tool is based on GPT-4 and therefore has non-deterministic outputs.
- domain assumption The authors' classifications of the ten submissions (e.g., which are good, which contain factual errors, which are nonsense) are correct.
- domain assumption The tool's behavior during the test window (August 2024) is representative of its behavior at other times.
Cite this review
Pith. "Pith review of Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben." pith.science (2026). https://pith.science/paper/XO6KYLZJ
@misc{pith2026241206651,
author = {Pith},
title = {Pith review of: Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben},
year = {2026},
howpublished = {\url{https://pith.science/paper/XO6KYLZJ}},
note = {Machine review of arXiv:2412.06651}
}
read the original abstract
This study examines the AI-powered grading tool "AI Grading Assistant" by the German company Fobizz, designed to support teachers in evaluating and providing feedback on student assignments. Against the societal backdrop of an overburdened education system and rising expectations for artificial intelligence as a solution to these challenges, the investigation evaluates the tool's functional suitability through two test series. The results reveal significant shortcomings: The tool's numerical grades and qualitative feedback are often random and do not improve even when its suggestions are incorporated. The highest ratings are achievable only with texts generated by ChatGPT. False claims and nonsensical submissions frequently go undetected, while the implementation of some grading criteria is unreliable and opaque. Since these deficiencies stem from the inherent limitations of large language models (LLMs), fundamental improvements to this or similar tools are not immediately foreseeable. The study critiques the broader trend of adopting AI as a quick fix for systemic problems in education, concluding that Fobizz's marketing of the tool as an objective and time-saving solution is misleading and irresponsible. Finally, the study calls for systematic evaluation and subject-specific pedagogical scrutiny of the use of AI tools in educational contexts.
Reference graph
Works this paper leans on
-
[1]
Messung” man einzelne Abweichler (“Messfehler
Zufälligkeit von Bewertungen und Rückmeldungen a. Zufälligkeit der vorgeschlagenen Gesamtnote Testreihe A umfasste die fünfmal wiederholte Erstellung einer Bewertung für jede der 10 exemplarischen Abgaben zu unserer Aufgabenstellung. Diese Methodologie gestattet es, das Tool hinsichtlich Konsistenz und Robustheit von Feedback und Bewertungen hin zu unters...
work page 2024
-
[2]
Unzuverlässige Erkennung inhaltlicher Defizite a. Falschbehauptungen Unsere simulierten Schüler:innen-Abgaben enthielten teilweise simple inhaltliche Fehler. So hat beispielsweise Text 5 behauptet, das Wahlalter für die Europawahl sei kürzlich auf 14 Jahre abgesenkt worden (Abbildung #fig:E:5-2). Abb. #fig:E:5-2: Screenshot Abgabe 5. Mühlhoff / Henningsen...
work page 2024
-
[3]
Unzuverlässige Umsetzung einzelner Bewertungskriterien Das Korrekturtool bietet eine freie Eingabemöglichkeit für Bewertungskriterien (siehe Abbildung #fig:T:kriterien) durch die Lehrkraft, inklusive relativer Gewichtung der Kriterien für die Gesamtnote. So wirbt Fobizz in einer Video-Anleitung zum Korrekturtool explizit mit der Möglichkeit, nach Belieben...
work page 2024
-
[4]
Inkonsistentes Feedback a. Uneinheitliche und zufällige Bezeichnung von Fehlerkategorien Die Zuordnung der erkannten Fehler zu jeweils einer Fehlerkategorie (dritte Spalte der tabellarischen Auflistung “Fehlerliste”, siehe Abbildung #fig:E:bewertung) ist ein integraler Bestandteil des automatisch generierten Feedbacks. Dabei fällt im Vergleich verschieden...
arXiv 2024
-
[5]
Die Umsetzung des Feedbacks führt nicht zur Verbesserung a. Fehlende Monotonie bei der Einarbeitung von Vorschlägen In Testreihe B haben die Abgaben Nr. 1 und 10 einer Serie von Verbesserungen und automatisierten Bewertungen unterzogen: Beginnend bei der Originalabgabe haben wir stets die Verbesserungsvorschläge aus der Fehlerliste der automatisierten Bew...
work page 2024
-
[6]
Erreichbarkeit einer Bestbewertung nur durch Einsatz von ChatGPT (Täuschung) In beiden Serien haben wir nach 7–12 Iterationen einen Punkt erreicht, an dem das weitere Einarbeiten der Rückmeldungen nicht sinnvoll zu sein schien, weil die Rückmeldungen zwischen zwei Optionen oszillierten (siehe vorherigen Punkt) und sich in marginalen Details festgebissen h...
work page 2024
-
[7]
fundierter didaktischer Beurteilung
Zufälligkeit von Bewertungen und Rückmeldungen Indem wir jede Abgabe für unsere exemplarische Aufgabenstellung fünfmal durch das Korrekturtool bewerten lassen haben, konnten wir beobachten, dass sowohl der Notenvorschlag als auch die qualitative inhaltliche Rückmeldung zwischen den verschiedenen Bewertungsdurchläufen für ein und dieselbe Lösung teilweise ...
work page 2024
-
[8]
Das Wahlalter in der EU ist auf 14 Jahre gesenkt worden
Unzuverlässige Erkennung inhaltlicher Defizite Wir haben systematisch beobachtet, dass das Korrekturtool faktische Falschbehauptungen (z.B.: “Das Wahlalter in der EU ist auf 14 Jahre gesenkt worden.”) und Fälle offensichtlicher Arbeitsverweigerung (Nonsense-Abgaben) nicht verlässlich erkennt. Weniger die sehr guten Abgaben, als gerade “low effort” und abs...
Show all 14 references
-
[9]
Dieses Feature wird in den begleitenden Video-Tutorials von Fobizz explizit beworben
Unzuverlässige Umsetzung einzelner Bewertungskriterien Das Korrekturtool erlaubt die freie Eingabe einer beliebigen Liste von stichwortartig bezeichneten Bewertungskriterien zusammen mit einer relativen prozentualen Gewichtung der Kriterien für die Gesamtnote (siehe Abbildung ...
2024
-
[10]
Fehlerliste
Inkonsistentes Feedback In der Zusammenschau von 50 Bewertungsdurchläufen für 10 Abgaben zu unserer exemplarischen Aufgabenstellung haben wir diverse Glitches in der Form von Inkonsistenzen und internen Selbstwidersprüchen der erstellten Bewertungsdokumente festgestellt. Die F...
-
[11]
Unentschiedenheit
Die Umsetzung des Feedbacks führt nicht zur Verbesserung In der Logik von Rückmeldungen auf Prüfungsleistungen oder Hausaufgaben ist angelegt, dass eine Umsetzung konkreter Verbesserungsvorschläge durch eine bessere Bewertung honoriert wird – oder jedenfalls nicht zu einer Ver...
2024
-
[12]
fatales Gebrauchshindernis
Erreichbarkeit einer Bestbewertung nur durch Einsatz von ChatGPT (Täuschung) Schließlich ist es eine zentrale Beobachtung, dass in der Feedback-Logik des Korrekturtools nicht möglich zu sein scheint, eine perfekte – also nicht mehr zu beanstandende – Abgabe zu präsentieren. Wi...
2024
-
[13]
KI-Korrekturhilfe
Die “KI-Korrekturhilfe” von Fobizz sollte im Schulalltag nicht eingesetzt werden Unsere empirische Untersuchung hat sich auf funktionale Mindestbedingungen für die Gebrauchstauglichkeit der Fobizz “KI-Korrekturhilfe” fokussiert und mehrere gravierende Beeinträchtigungen gefund...
2024 arXiv
-
[14]
Techno Fix
KI in der Schule ist keine Lösung für den Lehrkräftemangel und Produkt eines gefährlichen gesellschaftlichen Trends Unsere Untersuchung nimmt nur ein exemplarisches KI-Tool in den Blick. Mit der Automatisierung von Bewertung und Rückmeldung möchte dieses exemplarische Tool ein...
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.