Pith. sign in

REVIEW 3 cited by

Towards LLM-based Autograding for Short Textual Answers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.11508 v2 pith:USWMHS32 submitted 2023-09-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords gradingautogradingllmstextualanswersevaluationlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Grading exams is an important, labor-intensive, subjective, repetitive, and frequently challenging task. The feasibility of autograding textual responses has greatly increased thanks to the availability of large language models (LLMs) such as ChatGPT and the substantial influx of data brought about by digitalization. However, entrusting AI models with decision-making roles raises ethical considerations, mainly stemming from potential biases and issues related to generating false information. Thus, in this manuscript, we provide an evaluation of a large language model for the purpose of autograding, while also highlighting how LLMs can support educators in validating their grading procedures. Our evaluation is targeted towards automatic short textual answers grading (ASAG), spanning various languages and examinations from two distinct courses. Our findings suggest that while "out-of-the-box" LLMs provide a valuable tool to provide a complementary perspective, their readiness for independent automated grading remains a work in progress, necessitating human oversight.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. "**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems

    cs.CR 2026-06 unverdicted novelty 5.0 of 10

    LLM-based automatic grading systems are highly vulnerable to prompt injection attacks that force high scores regardless of answer quality, and existing defenses fail to mitigate them.

  2. LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability

    cs.CY 2026-06 unverdicted novelty 5.0 of 10

    Liberal partial-credit prompting reduces question-level grading error for all six tested LLMs, with ChatGPT 5.5 Thinking (LIBERAL) achieving the lowest MAE of 1.87.

  3. Fine-tuning for Better Few Shot Prompting: An Empirical Comparison for Short Answer Grading

    cs.LG 2025-08 conditional novelty 5.0 of 10

    Fine-tuning GPT-4o-mini on about 150 examples raised short-answer grading F1 from 0.68 to 0.73; QLoRA fine-tuning of Llama 3.1 8B only reached 0.65 after adding synthetic data.

Pith tools