REVIEW 4 major objections 4 minor 1 cited by
Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims o1 aligns with averaged teacher ratings on German essay scoring (Spearman r = .74), is internally consistent (ICC = .80), and can support teachers on language-related criteria, while open-source models show no reliable…
desk verdict Useful pilot study of LLM essay scoring in German, but the o1 headline is over-stated given N=20 and an unvalidated teacher ground truth. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a comparative scoring protocol rather than a single mathematical identity. Each essay is rated on ten six-point Likert criteria (five content, four language, and one overall judgment); the teacher ground truth is the average of the 3-7 ratings each essay received (mean 5.45), and each LLM score is the average of ten zero-shot runs of the same prompt at temperature 0.7. Alignment is measured with Spearman's rank correlation between LLM and teacher averages per criterion, and reliability is measured with the intraclass correlation coefficient (ICC) across the ten LLM runs. The paper also uses Mann-Whitney U tests to detect systematic leniency differences and inter-criteria correlation matrices to infer how much each criterion drives the overall score. The load-bearing comparison is the rank alignment between a single averaged LLM score and a single averaged teacher score on only 20 essays.
What would settle it
Recompute o1's Spearman correlation against each individual teacher's rating instead of the averaged teacher score for the same 20 essays; if the median such correlation is below, say, 0.3 or its confidence interval includes zero, the reported r = .74 is an artifact of averaging away teacher disagreement rather than true alignment.
Extended reading notes
Core claim
The central claim, stated on the paper's own terms, is that o1 outperforms all other tested LLMs in essay scoring: it achieves Spearman's r = .74 with human assessments on the overall score and ICC = .80 internal consistency across ten repeated runs, and it is the only model with significant correlations in nine of ten criteria. The paper further claims that GPT models generally align with teachers on language-related criteria but systematically assign higher scores, so their overall ratings diverge from human strictness, while content-heavy criteria like plot logic and main part are precisely where agreement is weakest. Open-source models, by contrast, have very low reliability (ICC near -0.04 and 0.01) and near-zero correlations with teachers, making them unsuitable for this task in their current form. These findings support the conclusion that o1 is a promising assistive tool for reducing teacher workload in evaluating language-related aspects, provided its leniency bias is addressed.
Load-bearing premise
The whole comparison assumes the average of three to seven teacher ratings per essay is a trustworthy gold standard, but the paper never reports how much the teachers agree with each other.
Editorial extensions
If this is right
- o1 can serve as a reliable second reader for language-related scoring criteria such as spelling, expression, and literal speech, where its correlations with teacher averages are highest.
- Teachers should still make the final call on content criteria such as plot logic and main part, where LLM-teacher agreement is weak and models differ significantly from human score distributions.
- Any deployed system should aggregate multiple LLM runs rather than trust a single output, since single-run reliability is imperfect (o1 ICC = .80).
- The leniency bias toward higher scores means calibration on teacher ratings would be needed before LLM scores are used as grades.
- Open-source models LLaMA 3-70B and Mixtral 8x7B, as tested with this zero-shot prompt, are not reliable enough for essay scoring support.
Reading between the lines
- If teacher ratings are as noisy as the small per-essay rater count suggests, the true alignment of o1 with any individual teacher may be much lower than .74; a fair test would compare o1 against each teacher separately rather than against the averaged consensus score.
- The high inter-criterion correlations in GPT-3.5 and GPT-4 (above .86 and .75 respectively) indicate these models may be producing a single global quality impression repackaged into ten criterion scores, so criterion-level scores should not be read as independent diagnoses of student strengths.
- The findings suggest a division-of-labor deployment: LLMs handle mechanics and style, teachers concentrate on content and feedback quality; this is a testable design for workload reduction studies.
- One could extend the same protocol to argumentative essays or other languages to check whether o1's language-criterion advantage generalizes or is specific to German narrative writing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a comparative evaluation of five LLMs (GPT-3.5, GPT-4, o1-preview, LLaMA 3-70B, Mixtral 8x7B) as automated essay scorers for German secondary-school essays. Twenty real student essays were rated by 37 teachers on ten criteria (content- and language-related), and the teacher averages were used as ground truth. Each LLM scored each essay ten times, and the mean of the ten runs was correlated with the teacher means using Spearman's rho; inter-run reliability was measured with ICC, and rating distributions were compared with Mann-Whitney U tests. The central claim is that the closed-source GPT models, especially o1, outperform open-source models, with o1 achieving Spearman's r = .74 with human ratings on the overall score and an ICC of .80, and that LLM-based assessment could support teachers particularly for language-related criteria.
Significance. If the reported effects were robust, the study would be a useful contribution to the emerging literature on LLM-based essay scoring in a non-English context, adding evidence on multidimensional criteria and run-to-run reliability. The authors assemble a real-world corpus of German essays with teacher ratings, make their prompting protocol transparent, and provide per-criterion analyses that go beyond holistic scores. The inclusion of open-source models and the explicit reporting of low ICC values for LLaMA 3 and Mixtral are informative. However, the statistical evidence for the headline superiority claim is thin, and several load-bearing assertions go beyond what the data can support. The credibility of the study would be substantially improved by addressing the reliability of the human benchmark and the significance of model differences.
major comments (4)
- [§3.2 and §5.6] The teacher-average ground truth is used to compute all Spearman correlations in Table 3, but no inter-rater reliability (e.g., teacher-level ICC) is reported for the 37 teachers, and the number of ratings per essay is only 3–7 (mean 5.45). The paper itself states in §5.6 that 'the absence of a gold standard due to the inherent variability among human raters' is a challenge. Without a variance-components analysis or a teacher ICC, it is not known how much of the variation in the teacher means reflects true essay quality versus sampling noise. With N=20 essays, the Spearman coefficients are highly unstable, and the conclusion that o1 'aligns with human assessments' is therefore not established. Please report teacher-level agreement (e.g., ICC per criterion) or at least the raw distribution of teacher ratings, and temper the wording of the central claim accordingly.
- [Table 3 and §4.2] The claim that the novel o1 model 'outperforms all other LLMs' rests on point estimates of Spearman's rho across 50 correlations (5 models × 10 criteria) with no correction for multiple comparisons and no significance test on the differences between models. For instance, the overall-score correlations are r = .742 (o1) and r = .575 (GPT-4); with N=20, this difference is not shown to be significant. A test for correlated correlations (e.g., Steiger's z) or bootstrap confidence intervals should be provided, and a multiple-comparison correction (Benjamini-Hochberg) should be applied to the p-values in Table 3. Without this, the model ranking is not supported beyond the level of descriptive statistics.
- [Abstract and Table 2] The abstract states that 'the novel o1 model outperforms all other LLMs, achieving Spearman's r = .74 with human assessments in the overall score, and an internal consistency of ICC = .80.' This is internally inconsistent with Table 2, which shows ICC = .84 for GPT-3.5 and ICC = .73 for GPT-4, i.e., o1's ICC is not the highest among the closed-source models. The sentence implies o1 is best on both alignment and consistency, but the data only support the former (and even that is contested above). Please revise the abstract and Section 5.1 to state that o1's internal consistency is comparable to, but not higher than, that of GPT-3.5.
- [§3.3] The paper filters out N=169 missing or out-of-range open-source model outputs before computing correlations and ICCs, but it does not report the distribution of these exclusions across essays or criteria. If LLaMA 3 and Mixtral failed disproportionately on certain essays or criteria, the low ICC values (Table 2) and weak correlations (Table 3) could be partly artifacts of the filtering rather than genuine evaluative behavior. Please report how many outputs were excluded per model, per essay, and per criterion, and run a sensitivity analysis (e.g., re-imputing or analyzing the unfiltered outputs).
minor comments (4)
- [§4.4] The sentence 'Notably, their are differences between model versions' contains a typo: 'their' should be 'there.'
- [Table 2 and text] The model naming is inconsistent: Table 2 uses 'GPT-o1' while the text and other tables use 'o1' or 'GPT-o1' inconsistently; please use a single label throughout.
- [Figure 3] The figure caption refers to 'the red line represents the linear least-squares regression,' but the paper reports Spearman correlations; please clarify whether the regression line is an illustrative fit and not the basis of the reported coefficients.
- [§5.5] The phrase 'without limitating user confidence' should be 'without limiting user confidence.'
Circularity Check
No significant circularity: the paper's empirical comparison of LLM ratings against averaged teacher ratings is self-contained and does not reduce to its inputs by construction.
full rationale
The paper is an empirical evaluation study. Its central quantitative claims — Spearman correlations between each LLM's averaged ratings and the averaged teacher ratings, ICC values across repeated LLM runs, and Mann-Whitney comparisons of score distributions — are computed from independently collected human ratings and LLM outputs. No parameter is fitted to the target correlations, no model output is defined in terms of the teacher scores, and no 'prediction' is derived from the quantity it claims to predict. The teacher-average ground truth is external to the LLM pipeline, and the LLM ratings are generated from a fixed zero-shot prompt without training on the essays or on the teacher scores. The small sample size (N=20) and unreported teacher inter-rater reliability weaken the robustness of the conclusions, but those are validity concerns, not circularity. The authors do cite their own prior work (e.g., Kasneci et al. 2023; Seßler et al. 2023; Bewersdorff et al. 2023), but these citations are used for background motivation and related-work context, not as the load-bearing justification for the empirical results. The statement in Section 5.6 that o1 'automatically incorporates CoT reasoning, which likely contributes to its superior performance' is an explanatory hypothesis, not a derivation from an assumed premise. Accordingly, no circular step meeting the quoted-evidence standard is present, and the appropriate score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The average of teacher ratings is a valid ground truth for essay quality per essay and criterion.
- domain assumption The ten criteria, derived from teacher feedback and refined by a German didactics expert, cover the relevant dimensions of essay quality.
- ad hoc to paper A single zero-shot prompt without prompt engineering gives a fair comparison of all five LLMs.
- domain assumption Averaging ten runs with temperature 0.7 yields a stable model rating analogous to averaging multiple human ratings.
Cite this review
Pith. "Pith review of Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring." pith.science (2026). https://pith.science/paper/GQN5QCHW
@misc{pith2026241116337,
author = {Pith},
title = {Pith review of: Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring},
year = {2026},
howpublished = {\url{https://pith.science/paper/GQN5QCHW}},
note = {Machine review of arXiv:2411.16337}
}
abstract
The manual assessment and grading of student writing is a time-consuming yet critical task for teachers. Recent developments in generative AI, such as large language models, offer potential solutions to facilitate essay-scoring tasks for teachers. In our study, we evaluate the performance and reliability of both open-source and closed-source LLMs in assessing German student essays, comparing their evaluations to those of 37 teachers across 10 pre-defined criteria (i.e., plot logic, expression). A corpus of 20 real-world essays from Year 7 and 8 students was analyzed using five LLMs: GPT-3.5, GPT-4, o1, LLaMA 3-70B, and Mixtral 8x7B, aiming to provide in-depth insights into LLMs' scoring capabilities. Closed-source GPT models outperform open-source models in both internal consistency and alignment with human ratings, particularly excelling in language-related criteria. The novel o1 model outperforms all other LLMs, achieving Spearman's $r = .74$ with human assessments in the overall score, and an internal consistency of $ICC=.80$. These findings indicate that LLM-based assessment can be a useful tool to reduce teacher workload by supporting the evaluation of essays, especially with regard to language-related criteria. However, due to their tendency for higher scores, the models require further refinement to better capture aspects of content quality.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben
The Fobizz AI grading tool gives unstable and arbitrary grades and feedback, only rewards ChatGPT-written texts with top scores, and fails to detect false claims or nonsense submissions.
Reference graph
Works this paper leans on
-
[1]
AI@Meta. 2024. Llama 3 Model Card
2024
-
[2]
Hani Alers, Aleksandra Malinowska, Gregory Meghoe, and Enso Apfel. 2024. Using ChatGPT-4 to Grade Open Question Exams. In Advances in Information and Communication, Kohei Arai (Ed.). Springer, Springer Nature Switzerland, Cham, 1–9
work page 2024
-
[3]
Anthropic. 2023. Claude. https://www.anthropic.com/
work page 2023
-
[4]
Xiaoyu Bai and Manfred Stede. 2023. A survey of current machine learning approaches to student free-text evaluation for intelligent tutoring. International Journal of Artificial Intelligence in Education 33, 4 (2023), 992–1030
work page 2023
-
[5]
Majdi Beseiso, Omar A Alzubi, and Hasan Rashaideh. 2021. A novel automated essay scoring approach for reliable higher educational assessments. Journal of Computing in Higher Education 33 (2021), 727–746
work page 2021
-
[6]
Arne Bewersdorff, Christian Hartmann, Marie Hornberger, Kathrin Seßler, Maria Bannert, Enkelejda Kasneci, Gjergji Kasneci, Xiaoming Zhai, and Claudia Nerdel. 2024. Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education. arXiv:2401.00832
arXiv 2024
-
[7]
Arne Bewersdorff, Kathrin Seßler, Armin Baur, Enkelejda Kasneci, and Claudia Nerdel. 2023. Assessing student errors in experimentation using artificial in- telligence and large language models: A comparative study with human raters. Computers and Education: Artificial Intelligence 5 (2023), 100177
work page 2023
-
[8]
Shravya Bhat, Huy Anh Nguyen, Steven Moore, John C Stamper, Majd Sakr, and Eric Nyberg. 2022. Towards Automated Generation and Evaluation of Questions in Educational Domains.. In EDM, Antonija Mitrovic and Nigel Bosch (Eds.). International Educational Data Mining Society, Durham, United Kingdom, 701– 704
work page 2022
Show all 56 references
-
[9]
Peter Birkel and Claudia Birkel. 2002. Wie einig sind sich Lehrer bei der Auf- satzbeurteilung? Eine Replikationsstudie zur Untersuchung von Rudolf Weiss. Psychologie in Erziehung und Unterricht, 49, 3 (2002), 219–224
2002
-
[10]
Daniel Blanchard, Joel Tetreault, Derrick Higgins, Aoife Cahill, and Martin Chodorow. 2013. TOEFL11: A Corpus of Non-Native English . Technical Report. Educational Testing Service
2013
-
[11]
Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Al- ternative to Human Evaluations?. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki...
2023
-
[12]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...
2019
-
[13]
Afrizal Doewes and Mykola Pechenizkiy. 2021. On the Limitations of Human- Computer Agreement in Automated Essay Scoring.. In The 14th International Conference on Educational Data Mining (EDM21) . International Educational Data Mining Society, Paris, France, 475–480
2021
-
[14]
Jonas Flodén. 2024. Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT. British Educational Research Journal 00 (2024), 1–24
2024
-
[15]
Google Gemini Team. 2024. Gemini: A Family of Highly Capable Multimodal Models
2024
-
[16]
Jürgen Grzesik and Michael Fischer. 1984. Was leisten Kriterien für die Aufsatzbeurteilung?: Theoretische, empirische und praktische Aspekte des Ge- brauchs von Kriterien und der Mehrfachbeurteilung nach globalem Ersteindruck . Westdeutscher-Verlag, Wiesbaden, Germany
1984
-
[17]
Veronika Hackl, Alexandra Elena Müller, Michael Granitzer, and Maximilian Sailer. 2023. Is GPT-4 a reliable rater? Evaluating consistency in GPT-4’s text ratings. In Frontiers in Education, Vol. 8. Frontiers Media SA, 1272229
2023
-
[18]
Ben Hamner, Jaison Morgan, Iynnvandev, Mark Shermis, and Tom Vander Ark
-
[19]
Hendrik Haverkamp, Malte Hecht, and Kirsten Schindler. 2024. Lernförderliches Feedback KI-basiert vermitteln. Der Deutschunterricht 5 (2024)
2024
-
[20]
Hyangeun Ji, Insook Han, and Yujung Ko. 2023. A systematic review of conver- sational AI in language education: Focusing on the collaboration with human teachers. Journal of Research on Technology in Education 55, 1 (2023), 48–63
2023
-
[21]
AQ Jiang, A Sablayrolles, A Mensch, C Bamford, DS Chaplot, D de las Casas, F Bressand, G Lengyel, G Lample, L Saulnier, et al. 2023. Mistral 7B (2023). Technical Report. Mistral AI
2023
-
[22]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. Technical Report. Mistral AI
2024
-
[23]
Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Can AI grade your essays? LAK ’25, March 03–07, 2025, Dublin, Ireland Hüllermeier, et al. 2023. ChatGPT for good? On opportunit...
2023
-
[24]
Zixuan Ke and Vincent Ng. 2019. Automated Essay Scoring: A Survey of the State of the Art.. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19. International Joint Conferences on Artificial Intelligence Organization, 6300–6308
2019
-
[25]
Terry K Koo and Mae Y Li. 2016. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of chiropractic medicine 15, 2 (2016), 155–163
2016
-
[26]
Kultusministerkonferenz. 2023. Lehrkräfteeinstellungsbedarf und -angebot in der Bundesrepublik Deutschland 2023 – 2035: Zusammengefasste Modellrechnungen der Länder. Dokumentation 238. Sekretariat der Ständigen Konferenz der Kultus- minister der Länder in der Bundesrepublik De...
2023
-
[27]
Gyeong-Geon Lee, Ehsan Latif, Xuansheng Wu, Ninghao Liu, and Xiaoming Zhai. 2024. Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence 6 (2024), 100213
2024
-
[28]
Artificial Intelligence (AI)
Dogan Gursoy Mesut Cicek and Lu Lu. 2024. Adverse impacts of revealing the presence of “Artificial Intelligence (AI)” technology in product and service descriptions on purchase intentions: the mediating role of emotional trust and the moderating role of perceived risk.Journal ...
2024
-
[29]
Atsushi Mizumoto and Masaki Eguchi. 2023. Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics 2, 2 (2023), 100050
2023
-
[30]
Leo Morjaria, Levi Burns, Keyna Bracken, Anthony J Levinson, Quang N Ngo, Mark Lee, and Matthew Sibbald. 2024. Examining the Efficacy of ChatGPT in Marking Short-Answer Assessments in an Undergraduate Medical Program. International Medical Education 3, 1 (2024), 32–43
2024
-
[31]
Nora Müller, Till Utesch, and Vera Busse. 2023. Qualität statt Quantität? Zum Zusammenhang von Schreibförderungs-und Feedbackpraktiken mit Textqualität unter Berücksichtigung von migrationsbedingter Mehrsprachigkeit. Unt.wiss. Zeits. f. Lernforschung 51, 2 (2023), 169–198
2023
-
[32]
Sonia Alejandrina Sotelo Muñoz, Giovanna Gutiérrez Gayoso, Alberto Caceres Huambo, Rogelio Domingo Cahuana Tapia, Jorge Layme Incaluque, Oscar Ed- uardo Pongo Aguila, Juan Cielo Ramírez Cajamarca, Jesus Enrique Reyes Acevedo, Herbert Victor Huaranga Rivera, and José Luis Arias...
2023
-
[33]
Frank Mußmann, Martin Riethmüller, Thomas Hardwig, Stefan Peters, Marcel Parciak, Ilka Charlotte Ohms, and Stefan Klötzer. 2016. Niedersächsische Arbeit- szeitstudie 2015 / 2016: Lehrkräfte an öffentlichen Schulen Ergebnisbericht. Technical Report. Georg-August-Universität Göt...
2016
-
[34]
Ben Naismith, Phoebe Mulcaire, and Jill Burstein. 2023. Automated evaluation of written discourse coherence using GPT-4. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023) , Ekaterina Kochmar, Jill Burstein, Andrea Hor...
2023
-
[35]
Tanya Nazaretsky, Moriah Ariely, Mutlu Cukurova, and Giora Alexandron. 2022. Teachers’ trust in AI-powered educational technology and a professional devel- opment program to improve it. British journal of educational technology 53, 4 (2022), 914–931
2022
-
[36]
Huy A Nguyen, Hayden Stec, Xinying Hou, Sarah Di, and Bruce M McLaren. 2023. Evaluating chatgpt’s decimal skills and feedback generation in a digital learning game. In European Conference on Technology Enhanced Learning , Olga Viberg, Ioana Jivet, Pedro J. Muñoz-Merino, Maria ...
2023
-
[37]
OpenAI. 2023. GPT-4 technical report. Technical Report. OpenAI
2023
-
[38]
OpenAI. 2024. OpenAI o1 System Card . Technical Report. OpenAI
2024
-
[39]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 (2022), 27730–27744
2022
-
[40]
Ulrike Padó, Yunus Eryilmaz, and Larissa Kirschner. 2023. Short-Answer Grad- ing for German: Addressing the Challenges. International Journal of Artificial Intelligence in Education (2023), 1–32
2023
-
[41]
Gustavo Pinto, Isadora Cardoso-Pereira, Danilo Monteiro, Danilo Lucena, Alberto Souza, and Kiev Gama. 2023. Large language models for education: Grading open- ended questions using chatgpt. In Proceedings of the XXXVII Brazilian Symposium on Software Engineering. Association f...
2023
-
[42]
Sanna Pohlmann-Rother, Edgar Schoreit, and Anja Kürzinger. 2016. Schreibkom- petenzen von Erstklässlern quantitativ-empirisch erfassen-Herausforderungen und Zugewinn eines analytisch-kriterialen Vorgehens gegenüber einer holistis- chen Bewertung. Journal for educational resear...
2016
-
[43]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9
2019
-
[44]
Dadi Ramesh and Suresh Kumar Sanampudi. 2022. An automated essay scoring systems: a systematic literature review. Artificial Intelligence Review 55, 3 (2022), 2495–2527
2022
-
[45]
Jörg Sawatzki, Tim Schlippe, and Marian Benner-Wickner. 2021. Deep learning techniques for automatic short answer grading: Predicting scores for English and German answers. In International conference on artificial intelligence in education technology, Eric C. K. Cheng, Rekha ...
2021
-
[46]
Pauline Schröter, Hannelore Söldner, Lars Hoffmann, Anja Riemenschnei- der, Jörg Jost, and Dorothee Wieser. 2022. Wie vergleichbar sind die Bew- ertungen von Abiturarbeiten im Fach Deutsch? Empirische Studien zu ver- schiedenen Bewertungsmodellen. In Das unvergleichliche Abitu...
2022
-
[47]
Kathrin Seßler, Tao Xiang, Lukas Bogenrieder, and Enkelejda Kasneci. 2023. Peer: Empowering writing with large language models. In Responsive and Sustainable Educational Futures, Olga Viberg, Ioana Jivet, Pedro J. Muñoz-Merino, Maria Perifanou, and Tina Papathoma (Eds.). Sprin...
2023
-
[48]
Maja Stahl, Leon Biermann, Andreas Nehring, and Henning Wachsmuth. 2024. Exploring LLM Prompting Strategies for Joint Essay Scoring and Feedback Gen- eration. arXiv:2404.15845 [cs.CL]
2024 arXiv
-
[49]
Chul Sung, Tejas Indulal Dhamecha, and Nirmal Mukhi. 2019. Improving short answer grading using transformer-based pre-training. In Artificial Intelligence in Education, Seiji Isotani, Eva Millán, Amy Ogan, Peter Hastings, Bruce McLaren, and Rose Luckin (Eds.). Springer Interna...
2019
-
[50]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models . Technical Report
2023
-
[51]
Masaki Uto, Yikuan Xie, and Maomi Ueno. 2020. Neural automated essay scor- ing incorporating handcrafted features. In Proceedings of the 28th International Conference on Computational Linguistics , Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). International Committee on C...
2020
-
[52]
Poortinga, and Theo M
Hester van Herk, Ype H. Poortinga, and Theo M. M. Verhallen. 2004. Response Styles in Rating Scales: Evidence of Method Bias in Data From Six EU Countries. Journal of Cross-Cultural Psychology 35, 3 (2004), 346–360
2004
-
[53]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837
2022
-
[54]
Jin Xue, Xiaoyi Tang, and Liyan Zheng. 2021. A hierarchical BERT-based transfer learning approach for multi-dimensional essay scoring. Ieee Access 9 (2021), 125403–125415
2021
-
[55]
Na Zhai and Xiaomei Ma. 2022. Automated writing evaluation (AWE) feedback: a systematic investigation of college students’ acceptance. Computer Assisted Language Learning 35, 9 (2022), 2817–2842
2022
-
[2012]
https://www.kaggle
The hewlett foundation: Automated essay scoring. https://www.kaggle. com/c/asap-aes
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.