Five LLMs showed low human-LLM agreement and weak within-model stability when scoring 67 Italian psychology essays on a four-criterion rubric.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CY 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Assessing the Reliability and Validity of Large Language Models for Automated Assessment of Student Essays in Higher Education
Five LLMs showed low human-LLM agreement and weak within-model stability when scoring 67 Italian psychology essays on a four-criterion rubric.