Pith. sign in

REVIEW 3 major objections 5 minor 14 references

Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The Fobizz AI Grading Assistant produces grades and feedback that vary substantially when the same submission is graded repeatedly, and incorporating its own suggestions does not improve the grade; only ChatGPT-written texts receive…

desk verdict The core volatility finding is credible and policy-relevant; the feedback-no-improvement claim is undercut by single-draw iteration design. read the letter →

arxiv 2412.06651 v5 pith:XO6KYLZJ submitted 2024-12-09 cs.CY cs.AIcs.CLcs.ET

classification cs.CYcs.AIcs.CLcs.ET
keywords FobizzAIGradingAssistantautomatedessayLLMfeedbackreliabilitygradevolatilityineducationAI-generatedtextdetectionteacherworkloadautomationnonsensesubmission
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tests Fobizz's 'AI Grading Assistant', a preconfigured chatbot built on GPT-4 that offers teachers automated correction, feedback, and grade proposals for student texts. Using one German writing task and ten simulated student submissions, each graded five times, it claims that the tool's numerical grades and qualitative feedback vary substantially across repeated runs. In a second test series, revising two essays according to the tool's own feedback over 7-12 iterations did not raise the grade, and only a final ChatGPT rewrite produced near-perfect scores. The paper also reports that false factual claims and nonsense or off-topic submissions often pass undetected, that user-supplied grading criteria such as word count are applied unreliably, and that feedback documents contain invented errors and inconsistent category names. The authors conclude that these deficits stem from inherent properties of large language models, that a quick technical fix is not in sight, and that marketing the tool as objective and time-saving is misleading.

What carries the argument

The load-bearing machinery is the repeated-grading protocol: deliberately submitting the same text several times and comparing outputs. Because the underlying model samples from a probability distribution, a single run hides the spread that this protocol exposes. The complementary mechanism is the iterative feedback loop—revising an essay to implement the tool's own error list and then regrading it—which tests whether feedback has the monotonicity property a grading system must have. Together these procedures turn hidden sampling randomness, unreliable word counting, and absent AI-text detection into observable, reproducible failures.

What would settle it

Re-run the paper's Test Series A on the current Fobizz tool: if the 10 submissions each produce the same grade across five runs, or if the two iterative revision series never drop below their starting grades and reach 99 percent without a ChatGPT rewrite, then the paper's central claims would be contradicted.

Watch

Extended reading notes

Core claim

At the level of the tool's own functioning, the central discovery is that Fobizz's grading assistant fails the minimum consistency requirements of grading. In Test Series A, only two of ten essays received the same overall grade in all five runs; three essays moved by more than one school grade, and one nonsense essay ranged from 1 to 14 points on the 15-point scale. The qualitative feedback was just as unstable: the same essay could be praised as 'hervorragend' in one run and described with ambivalent phrasing in the next. In Test Series B, following the tool's correction suggestions did not increase the recommended grade—the average over all iterations stayed below the starting value—and the feedback oscillated, sometimes reversing its own previous instructions. The paper interprets these observations as consequences of the stochastic, pattern-completion nature of LLMs, not as bugs that ordinary software updates would fix.

Load-bearing premise

The results assume that one short German position-writing task (150-250 words, with the authors' chosen criteria) is representative of how teachers use the tool and of typical student work; the paper itself flags this limitation in its section on further research.

Editorial extensions

If this is right

  • A teacher who grades a submission once may unknowingly issue a grade that another run would move by more than one school grade.
  • A student who faithfully implements every feedback suggestion cannot expect the grade to rise; over 7-12 iterations the average stayed slightly below the starting score.
  • If students learn that only ChatGPT-assisted texts reach the top band, the tool itself becomes an incentive to outsource homework to AI.
  • Because the defects trace to inherent LLM limits, simply updating the prompt or model version will not remove them; systematic evaluation before purchase is the direct consequence.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's test design, the 1-to-14 point swing on an off-topic essay implies that low-effort and refusal submissions are precisely the cases where a single LLM grade carries the least information; any evaluation of similar tools should over-sample those cases.
  • The near-perfect rating for a ChatGPT-written text raises a testable hypothesis the paper leaves open: the grader may systematically prefer text that matches its own generation style, so the 'best grade only via AI' effect could grow as graders and student tools converge on the same model family.
  • Because the observed volatility is a property of stochastic sampling, the paper's logic extends to any LLM-based grading assistant, not just this one; a practical fix would be to run each submission multiple times and report the spread before a teacher sees a single number.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper evaluates the Fobizz 'KI-Korrekturhilfe' (AI Grading Assistant) by running 10 simulated student submissions five times each through the tool (Testreihe A) and by iteratively rewriting two submissions in response to the tool's feedback over 7–12 rounds, with a final ChatGPT-based revision (Testreihe B). The authors report wide variance in numerical scores and qualitative feedback, including a score range of 1 to 14 points for one identical submission; failure to detect factual errors, nonsense, and nonconforming text length; internal contradictions in the feedback; and a pattern in which only ChatGPT-rewritten texts receive near-perfect scores. They conclude that the tool is unsuitable for classroom use, that its marketing is misleading, and that the deficiencies stem from inherent LLM limitations. The paper also includes an update section describing Fobizz's response and a rough re-test of the modified tool.

Significance. The topic is timely and consequential: Fobizz's tool is licensed by several German states (Table #tab:E:bundesländer), so independent functional evaluation is genuinely needed. The Series A design, with five independent grading runs of identical submissions, is an appropriate and simple method for exposing the stochastic variance of an LLM-based grading system, and the authors' publication of inputs and outputs in a material annex is a transparency strength. The finding that re-evaluating an unchanged submission can change the score by more than a school grade (Section III.1.a) is credible and, on its own, already raises serious doubts about the fitness of this tool for high-stakes decisions. However, the related claim that feedback incorporation does not improve scores (Mangel 5) and the claim that only ChatGPT texts achieve top marks (Mangel 6) rely on a much weaker single-draw design in Series B, which is insufficient given the demonstrated noise level.

major comments (3)
  1. [§III.5.a / Fig. #fig:E:irrfahrt] In Testreihe B, each version of the text is graded exactly once. Testreihe A (Section III.1.a) shows that re-evaluating the same unchanged submission yields score swings of more than one grade for 3 of 10 submissions and a range of 1 to 14 points for one submission. Consequently, the non-monotonic trajectory in Fig. #fig:E:irrfahrt and the observation that the mean over iterations lies slightly below the baseline are exactly what sampling noise alone would produce. The resulting claim that incorporating the tool's feedback does not improve the grade (Mangel 5.a/b, Table #tab:D:gravität) is therefore underdetermined. The authors themselves note in Section IV ('Weiterführende Untersuchungen') that the Series B sample will be expanded in a future version, but the current paper states this claim categorically and classifies it as fatal. I recommend either repeating the iterative series with at least five independent evaluations per text version and reporting the resulting distributions, or explicitly downgrading the claim to a preliminary qualitative observation.
  2. [§III.6] The claim that only ChatGPT-rewritten texts receive near-perfect evaluations is based on a single final evaluation of each ChatGPT version and on a small number of evaluations of the human-revised versions. Given the score variability documented in Series A, the difference between the human-revised and the ChatGPT-revised versions could be accounted for by chance. To support the strong claim that 'Bestnote nur durch ChatGPT möglich', the authors should grade the final human-revised and ChatGPT-revised versions multiple times (say, five times each) and compare the score distributions. Without such data, the inequality between human-revised and ChatGPT-revised versions is not established.
  3. [§IV / Table #tab:D:gravität] The classification of Mangel 5.a/b as a 'fatales Gebrauchshindernis' and the associated recommendation to 'Tool nicht anbieten/verwenden' are load-bearing for the paper's overall condemnation. Because the evidence for this classification comes from a single stochastic draw per iteration, the categorical claim that the overall grade 'steigt nicht' when feedback is implemented is not yet supported. This is not a minor statistical detail: the conclusion that the feedback is didactically valueless is central to the paper's policy recommendation. Either the repeated-draw evidence should be supplied or the claim should be reformulated as a tentative finding that motivates further testing. The same caveat applies to the summary in the Executive Summary and Section V.1(5).
minor comments (5)
  1. [§II] The study uses a single task (a 150–250 word German position statement) whose simplicity is acknowledged later in Section IV. This scope limitation should be stated in the methodology itself, since it defines the generality of every subsequent finding about the tool.
  2. [§VI] The update table for the 'verbesserten' tool (tested 21.02.2025) notes that the authors' own test scenario was incorporated into Fobizz's prompt as an example; this means the tool is being tested on a near-match to its prompt examples, which likely makes the results more favorable to the tool. This caveat should be displayed in the table itself or its caption so that readers do not interpret the 'Status quo' entries as directly comparable to the original August 2024 tests.
  3. [Fig. #fig:E:volatilität] The figure would be much easier to read if the raw scores for each of the five runs per submission were provided in a table, together with the x-axis mapping of submission numbers; the current figure requires checking the text to reconstruct the individual data points.
  4. [§III.1.a] The statement that the average fluctuation is 'mehr als einen Punkt auf der 15-Punkte-Skala' would be more informative with a per-submission range or standard deviation; the present aggregate hides the fact that most submissions are fairly stable while one (Abgabe 8) spans from 1 to 14 points.
  5. [Footnotes throughout] Several web sources are cited without an access date; given that the paper's evidence is time-sensitive (the tool was modified on 15.12.2024), every online source should include an 'accessed on' date.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Fobizz tool evaluation is an empirical measurement study whose central claims rest on direct test observations, not on definitions, fits, or self-citations.

full rationale

This is an empirical evaluation, not a derivation. The central claims—that repeated grading of the same submission yields volatile scores, that factual errors and nonsense go undetected, that criterion implementation is unreliable, that iterative incorporation of feedback does not improve scores, and that only ChatGPT-rewritten texts approach perfect marks—are supported by direct, documented test runs in the paper (e.g., five runs per text in Testreihe A, iterative chains in Testreihe B). None of these results is obtained by fitting a parameter or by defining a quantity in terms of the target result. The only self-references in the paper are background citations to Mühlhoff's prior work on data power and AI ethics (e.g., in footnotes 50-51) and a citation to an external arXiv study; these are not load-bearing for the tool evaluation. The update section even notes that Fobizz incorporated the authors' test scenario into the new prompt, which would bias the follow-up assessment in favor of the tool—a transparency step opposite to circular reasoning. A methodological concern about Testreihe B (single stochastic grading draw per iteration) affects the strength of the no-improvement inference, but it is a statistical-power issue, not a circularity issue. Therefore no circular steps are identified.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters are fit. The study relies on the domain assumptions listed, mainly that the test setup is representative and the authors' judgments about submission quality are correct. No new entities are postulated.

assumptions (3)
  • domain assumption The Fobizz tool is based on GPT-4 and therefore has non-deterministic outputs.
    Stated in Section II, 'Technisch'; used to explain observed variance and to motivate repeated runs.
  • domain assumption The authors' classifications of the ten submissions (e.g., which are good, which contain factual errors, which are nonsense) are correct.
    The validity of findings like 'false claims go undetected' depends on the premise that the planted false claim is actually false and the nonsense texts are actually off-topic. This is reasonable for the EU voting age example, but it is still an unverified input.
  • domain assumption The tool's behavior during the test window (August 2024) is representative of its behavior at other times.
    The study tests one window; the tool may change. The update section shows it did change in December 2024.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben." pith.science (2026). https://pith.science/paper/XO6KYLZJ

@misc{pith2026241206651,
  author       = {Pith},
  title        = {Pith review of: Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XO6KYLZJ}},
  note         = {Machine review of arXiv:2412.06651}
}
read the original abstract

This study examines the AI-powered grading tool "AI Grading Assistant" by the German company Fobizz, designed to support teachers in evaluating and providing feedback on student assignments. Against the societal backdrop of an overburdened education system and rising expectations for artificial intelligence as a solution to these challenges, the investigation evaluates the tool's functional suitability through two test series. The results reveal significant shortcomings: The tool's numerical grades and qualitative feedback are often random and do not improve even when its suggestions are incorporated. The highest ratings are achievable only with texts generated by ChatGPT. False claims and nonsensical submissions frequently go undetected, while the implementation of some grading criteria is unreliable and opaque. Since these deficiencies stem from the inherent limitations of large language models (LLMs), fundamental improvements to this or similar tools are not immediately foreseeable. The study critiques the broader trend of adopting AI as a quick fix for systemic problems in education, concluding that Fobizz's marketing of the tool as an objective and time-saving solution is misleading and irresponsible. Finally, the study calls for systematic evaluation and subject-specific pedagogical scrutiny of the use of AI tools in educational contexts.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

14 extracted references · 13 canonical work pages

  1. [1]

    Messung” man einzelne Abweichler (“Messfehler

    Zufälligkeit von Bewertungen und Rückmeldungen a. Zufälligkeit der vorgeschlagenen Gesamtnote Testreihe A umfasste die fünfmal wiederholte Erstellung einer Bewertung für jede der 10 exemplarischen Abgaben zu unserer Aufgabenstellung. Diese Methodologie gestattet es, das Tool hinsichtlich Konsistenz und Robustheit von Feedback und Bewertungen hin zu unters...

  2. [2]

    inhaltlichen Richtigkeit

    Unzuverlässige Erkennung inhaltlicher Defizite a. Falschbehauptungen Unsere simulierten Schüler:innen-Abgaben enthielten teilweise simple inhaltliche Fehler. So hat beispielsweise Text 5 behauptet, das Wahlalter für die Europawahl sei kürzlich auf 14 Jahre abgesenkt worden (Abbildung #fig:E:5-2). Abb. #fig:E:5-2: Screenshot Abgabe 5. Mühlhoff / Henningsen...

  3. [3]

    Feedback

    Unzuverlässige Umsetzung einzelner Bewertungskriterien Das Korrekturtool bietet eine freie Eingabemöglichkeit für Bewertungskriterien (siehe Abbildung #fig:T:kriterien) durch die Lehrkraft, inklusive relativer Gewichtung der Kriterien für die Gesamtnote. So wirbt Fobizz in einer Video-Anleitung zum Korrekturtool explizit mit der Möglichkeit, nach Belieben...

  4. [4]

    Fehlerliste

    Inkonsistentes Feedback a. Uneinheitliche und zufällige Bezeichnung von Fehlerkategorien Die Zuordnung der erkannten Fehler zu jeweils einer Fehlerkategorie (dritte Spalte der tabellarischen Auflistung “Fehlerliste”, siehe Abbildung #fig:E:bewertung) ist ein integraler Bestandteil des automatisch generierten Feedbacks. Dabei fällt im Vergleich verschieden...

  5. [5]

    flatterhafter

    Die Umsetzung des Feedbacks führt nicht zur Verbesserung a. Fehlende Monotonie bei der Einarbeitung von Vorschlägen In Testreihe B haben die Abgaben Nr. 1 und 10 einer Serie von Verbesserungen und automatisierten Bewertungen unterzogen: Beginnend bei der Originalabgabe haben wir stets die Verbesserungsvorschläge aus der Fehlerliste der automatisierten Bew...

  6. [6]

    An diesem Punkt wurde für eine finale Überarbeitung ChatGPT verwendet, Abbildung #fig:E:chatgpt-prompt zeigt exemplarisch den dafür in Fall von Text 1 verwendeten Prompt

    Erreichbarkeit einer Bestbewertung nur durch Einsatz von ChatGPT (Täuschung) In beiden Serien haben wir nach 7–12 Iterationen einen Punkt erreicht, an dem das weitere Einarbeiten der Rückmeldungen nicht sinnvoll zu sein schien, weil die Rückmeldungen zwischen zwei Optionen oszillierten (siehe vorherigen Punkt) und sich in marginalen Details festgebissen h...

  7. [7]

    fundierter didaktischer Beurteilung

    Zufälligkeit von Bewertungen und Rückmeldungen Indem wir jede Abgabe für unsere exemplarische Aufgabenstellung fünfmal durch das Korrekturtool bewerten lassen haben, konnten wir beobachten, dass sowohl der Notenvorschlag als auch die qualitative inhaltliche Rückmeldung zwischen den verschiedenen Bewertungsdurchläufen für ein und dieselbe Lösung teilweise ...

  8. [8]

    Das Wahlalter in der EU ist auf 14 Jahre gesenkt worden

    Unzuverlässige Erkennung inhaltlicher Defizite Wir haben systematisch beobachtet, dass das Korrekturtool faktische Falschbehauptungen (z.B.: “Das Wahlalter in der EU ist auf 14 Jahre gesenkt worden.”) und Fälle offensichtlicher Arbeitsverweigerung (Nonsense-Abgaben) nicht verlässlich erkennt. Weniger die sehr guten Abgaben, als gerade “low effort” und abs...

Show all 14 references
  1. [9]

    Dieses Feature wird in den begleitenden Video-Tutorials von Fobizz explizit beworben

    Unzuverlässige Umsetzung einzelner Bewertungskriterien Das Korrekturtool erlaubt die freie Eingabe einer beliebigen Liste von stichwortartig bezeichneten Bewertungskriterien zusammen mit einer relativen prozentualen Gewichtung der Kriterien für die Gesamtnote (siehe Abbildung ...

  2. [10]

    Fehlerliste

    Inkonsistentes Feedback In der Zusammenschau von 50 Bewertungsdurchläufen für 10 Abgaben zu unserer exemplarischen Aufgabenstellung haben wir diverse Glitches in der Form von Inkonsistenzen und internen Selbstwidersprüchen der erstellten Bewertungsdokumente festgestellt. Die F...

  3. [11]

    Unentschiedenheit

    Die Umsetzung des Feedbacks führt nicht zur Verbesserung In der Logik von Rückmeldungen auf Prüfungsleistungen oder Hausaufgaben ist angelegt, dass eine Umsetzung konkreter Verbesserungsvorschläge durch eine bessere Bewertung honoriert wird – oder jedenfalls nicht zu einer Ver...

  4. [12]

    fatales Gebrauchshindernis

    Erreichbarkeit einer Bestbewertung nur durch Einsatz von ChatGPT (Täuschung) Schließlich ist es eine zentrale Beobachtung, dass in der Feedback-Logik des Korrekturtools nicht möglich zu sein scheint, eine perfekte – also nicht mehr zu beanstandende – Abgabe zu präsentieren. Wi...

  5. [13]

    KI-Korrekturhilfe

    Die “KI-Korrekturhilfe” von Fobizz sollte im Schulalltag nicht eingesetzt werden Unsere empirische Untersuchung hat sich auf funktionale Mindestbedingungen für die Gebrauchstauglichkeit der Fobizz “KI-Korrekturhilfe” fokussiert und mehrere gravierende Beeinträchtigungen gefund...

  6. [14]

    Techno Fix

    KI in der Schule ist keine Lösung für den Lehrkräftemangel und Produkt eines gefährlichen gesellschaftlichen Trends Unsere Untersuchung nimmt nur ein exemplarisches KI-Tool in den Blick. Mit der Automatisierung von Bewertung und Rückmeldung möchte dieses exemplarische Tool ein...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.