REVIEW 8 cited by
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In recent years, large language models (LLMs) achieve remarkable success across a variety of tasks. However, their potential in the domain of Automated Essay Scoring (AES) remains largely underexplored. Moreover, compared to English data, the methods for Chinese AES is not well developed. In this paper, we propose Rank-Then-Score (RTS), a fine-tuning framework based on large language models to enhance their essay scoring capabilities. Specifically, we fine-tune the ranking model (Ranker) with feature-enriched data, and then feed the output of the ranking model, in the form of a candidate score set, with the essay content into the scoring model (Scorer) to produce the final score. Experimental results on two benchmark datasets, HSK and ASAP, demonstrate that RTS consistently outperforms the direct prompting (Vanilla) method in terms of average QWK across all LLMs and datasets, and achieves the best performance on Chinese essay scoring using the HSK dataset.
Forward citations
Cited by 8 Pith papers
-
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
A 30-day, knowledge-tracing-grounded simulated learner benchmark for tutoring agents finds that no base model or harness alone determines quality and that almost all tested combinations plateau within days.
-
Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards
RLAES-AGFO trains an LLM with GRPO and an LLM-judged 166-item feedback rubric, reaching QWK 0.803 on ASAP while generating feedback rated as well as GPT-5.5.
-
WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays
WrAFT, a modular LLM system, scores TOEFL essays with QWK 0.84/RMSE 0.44 and generates surface and deep feedback that human raters approve 93-96% of the time.
-
Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis
Across 65 studies, LLM-human essay-score agreement mostly falls between 0.30 and 0.80, but the spread is wide, reporting is heterogeneous, and no pooled estimate is given.
-
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
ZPD-SCA, an expert-annotated Chinese reading benchmark, shows LLMs judge reading difficulty for student age groups poorly in zero-shot settings and improve, but remain biased, with in-context examples.
-
Using Large Language Models to Assess Teachers' Pedagogical Content Knowledge
In video-based teacher knowledge assessments, GPT-4 scoring was more lenient than both human raters and a supervised ML model, while rater-related factors dominated construct-irrelevant variance.
-
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring
A student-teacher multi-agent pipeline with positive-only feedback improves QWK agreement with human essay scores by 21% on a multimodal benchmark, with gains concentrated in traits where baselines were weakest.
-
Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
In a fixed hybrid AES pipeline on ASAP 2.0, GPT-5 mini summarization gives the best human-agreement (QWK 0.8435), ahead of GPT-5 and GPT-5 nano, though the differences are small and no significance tests or baselines ...
Discussion (0). Continue with ORCID to comment.