GPT-4o-mini's scoring of music analysis essays agrees only moderately with teacher mean scores, with strategy-specific bias: Fs+CoT under-scores, RAG over-scores, and self-consistency is repeatable but weakly accurate per response.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
GPT-4o-mini's scoring of music analysis essays agrees only moderately with teacher mean scores, with strategy-specific bias: Fs+CoT under-scores, RAG over-scores, and self-consistency is repeatable but weakly accurate per response.