The paper introduces a 1,030-question benchmark of 4,109 expert-annotated student responses with step-wise scores and error causes, along with two consistency metrics for evaluating LLM short answer scoring.
Automated cross-prompt scoring of essay traits,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
The paper introduces a 1,030-question benchmark of 4,109 expert-annotated student responses with step-wise scores and error causes, along with two consistency metrics for evaluating LLM short answer scoring.