Pith. sign in

REVIEW 4 cited by

A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.02165 v2 pith:KXDXBVGV submitted 2024-10-03 cs.AI cs.CL

classification cs.AIcs.CL
keywords gradinggradeoptasagquestionssagsautomaticchallengescontent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-ended short-answer questions (SAGs) have been widely recognized as a powerful tool for providing deeper insights into learners' responses in the context of learning analytics (LA). However, SAGs often present challenges in practice due to the high grading workload and concerns about inconsistent assessments. With recent advancements in natural language processing (NLP), automatic short-answer grading (ASAG) offers a promising solution to these challenges. Despite this, current ASAG algorithms are often limited in generalizability and tend to be tailored to specific questions. In this paper, we propose a unified multi-agent ASAG framework, GradeOpt, which leverages large language models (LLMs) as graders for SAGs. More importantly, GradeOpt incorporates two additional LLM-based agents - the reflector and the refiner - into the multi-agent system. This enables GradeOpt to automatically optimize the original grading guidelines by performing self-reflection on its errors. Through experiments on a challenging ASAG task, namely the grading of pedagogical content knowledge (PCK) and content knowledge (CK) questions, GradeOpt demonstrates superior performance in grading accuracy and behavior alignment with human graders compared to representative baselines. Finally, comprehensive ablation studies confirm the effectiveness of the individual components designed in GradeOpt.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    An iterative framework lets LLMs learn procedural assessment skills for rubric construction, improving automated scoring on all ten ASAP-SAS items and often exceeding expert rubrics while showing cross-item transfer.

  2. "**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems

    cs.CR 2026-06 unverdicted novelty 5.0 of 10

    LLM-based automatic grading systems are highly vulnerable to prompt injection attacks that force high scores regardless of answer quality, and existing defenses fail to mitigate them.

  3. LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability

    cs.CY 2026-06 unverdicted novelty 5.0 of 10

    Liberal partial-credit prompting reduces question-level grading error for all six tested LLMs, with ChatGPT 5.5 Thinking (LIBERAL) achieving the lowest MAE of 1.87.

  4. Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering

    cs.CY 2026-05 unverdicted novelty 3.0 of 10

    LLM graders achieve substantial human agreement on math and science MCAS items but vary on ELA, performing best as sources of formative narrative feedback rather than summative numerical scores.

Pith tools