Pith. sign in

REVIEW 8 cited by

Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.05736 v1 pith:C34E3PMX submitted 2025-04-08 cs.CL cs.AI

classification cs.CLcs.AI
keywords essayscoringlanguagelargemodelmodelsacrossautomated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, large language models (LLMs) achieve remarkable success across a variety of tasks. However, their potential in the domain of Automated Essay Scoring (AES) remains largely underexplored. Moreover, compared to English data, the methods for Chinese AES is not well developed. In this paper, we propose Rank-Then-Score (RTS), a fine-tuning framework based on large language models to enhance their essay scoring capabilities. Specifically, we fine-tune the ranking model (Ranker) with feature-enriched data, and then feed the output of the ranking model, in the form of a candidate score set, with the essay content into the scoring model (Scorer) to produce the final score. Experimental results on two benchmark datasets, HSK and ASAP, demonstrate that RTS consistently outperforms the direct prompting (Vanilla) method in terms of average QWK across all LLMs and datasets, and achieves the best performance on Chinese essay scoring using the HSK dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

    cs.CY 2026-08 conditional novelty 6.0 of 10

    A 30-day, knowledge-tracing-grounded simulated learner benchmark for tutoring agents finds that no base model or harness alone determines quality and that almost all tested combinations plateau within days.

  2. Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards

    cs.CL 2026-07 conditional novelty 6.0 of 10

    RLAES-AGFO trains an LLM with GRPO and an LLM-judged 166-item feedback rubric, reaching QWK 0.803 on ASAP while generating feedback rated as well as GPT-5.5.

  3. WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays

    cs.AI 2026-07 conditional novelty 6.0 of 10

    WrAFT, a modular LLM system, scores TOEFL essays with QWK 0.84/RMSE 0.44 and generates surface and deep feedback that human raters approve 93-96% of the time.

  4. Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis

    cs.CL 2025-12 conditional novelty 5.0 of 10

    Across 65 studies, LLM-human essay-score agreement mostly falls between 0.30 and 0.80, but the spread is wide, reporting is heterogeneous, and no pooled estimate is given.

  5. ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities

    cs.CL 2025-08 conditional novelty 5.0 of 10

    ZPD-SCA, an expert-annotated Chinese reading benchmark, shows LLMs judge reading difficulty for student age groups poorly in zero-shot settings and improve, but remain biased, with in-context examples.

  6. Using Large Language Models to Assess Teachers' Pedagogical Content Knowledge

    cs.AI 2025-05 conditional novelty 5.0 of 10

    In video-based teacher knowledge assessments, GPT-4 scoring was more lenient than both human raters and a supervised ML model, while rater-related factors dominated construct-irrelevant variance.

  7. CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A student-teacher multi-agent pipeline with positive-only feedback improves QWK agreement with human essay scores by 21% on a multimodal benchmark, with gains concentrated in traits where baselines were weakest.

  8. Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment

    cs.CL 2026-07 conditional novelty 4.0 of 10

    In a fixed hybrid AES pipeline on ASAP 2.0, GPT-5 mini summarization gives the best human-agreement (QWK 0.8435), ahead of GPT-5 and GPT-5 nano, though the differences are small and no significance tests or baselines ...

Pith tools