Off-the-shelf LLMs match or exceed inter-examiner agreement on a new 32k-response double-marked GCSE dataset spanning five subjects and handwritten scripts.
arXiv preprint arXiv:2401.06431 , year=
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CL 3years
2026 3verdicts
UNVERDICTED 3representative citing papers
LLM representations encode essay quality in a linearly decodable form that emerges across layers and includes identifiable scoring neurons whose distribution shifts with essay length.
MAPLE uses meta-learning with prototypical networks to learn transferable representations and achieves state-of-the-art cross-prompt essay scoring on ELLIPSE, LAILA, and parts of ASAP datasets.
citing papers explorer
-
LLM Performance on a Real, Double-Marked GCSE Benchmark
Off-the-shelf LLMs match or exceed inter-examiner agreement on a new 32k-response double-marked GCSE dataset spanning five subjects and handwritten scripts.
-
From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models
LLM representations encode essay quality in a linearly decodable form that emerges across layers and includes identifiable scoring neurons whose distribution shifts with essay length.
-
MAPLE: A Meta-learning Framework for Cross-Prompt Essay Scoring
MAPLE uses meta-learning with prototypical networks to learn transferable representations and achieves state-of-the-art cross-prompt essay scoring on ELLIPSE, LAILA, and parts of ASAP datasets.