In a small German essay study, OpenAI's o1 agreed best with teacher ratings, but all models scored content less reliably and gave higher marks than teachers.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
In a small German essay study, OpenAI's o1 agreed best with teacher ratings, but all models scored content less reliably and gave higher marks than teachers.