The Fobizz AI grading tool gives unstable and arbitrary grades and feedback, only rewards ChatGPT-written texts with top scores, and fails to detect false claims or nonsense submissions.
Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The manual assessment and grading of student writing is a time-consuming yet critical task for teachers. Recent developments in generative AI, such as large language models, offer potential solutions to facilitate essay-scoring tasks for teachers. In our study, we evaluate the performance and reliability of both open-source and closed-source LLMs in assessing German student essays, comparing their evaluations to those of 37 teachers across 10 pre-defined criteria (i.e., plot logic, expression). A corpus of 20 real-world essays from Year 7 and 8 students was analyzed using five LLMs: GPT-3.5, GPT-4, o1, LLaMA 3-70B, and Mixtral 8x7B, aiming to provide in-depth insights into LLMs' scoring capabilities. Closed-source GPT models outperform open-source models in both internal consistency and alignment with human ratings, particularly excelling in language-related criteria. The novel o1 model outperforms all other LLMs, achieving Spearman's $r = .74$ with human assessments in the overall score, and an internal consistency of $ICC=.80$. These findings indicate that LLM-based assessment can be a useful tool to reduce teacher workload by supporting the evaluation of essays, especially with regard to language-related criteria. However, due to their tendency for higher scores, the models require further refinement to better capture aspects of content quality.
fields
cs.CY 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben
The Fobizz AI grading tool gives unstable and arbitrary grades and feedback, only rewards ChatGPT-written texts with top scores, and fails to detect false claims or nonsense submissions.