Pith. sign in

REVIEW 3 cited by

AI-assisted Automated Short Answer Grading of Handwritten University Level Mathematics Exams

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11728 v1 pith:RFEGH4Q4 submitted 2024-08-21 math.HO

classification math.HO
keywords gradinggpt-4handwrittenautomatedfeedbackmathematicspre-trainedresponse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Effective and timely feedback in educational assessments is essential but labor-intensive, especially for complex tasks. Recent developments in automated feedback systems, ranging from deterministic response grading to the evaluation of semi-open and open-ended essays, have been facilitated by advances in machine learning. The emergence of pre-trained Large Language Models, such as GPT-4, offers promising new opportunities for efficiently processing diverse response types with minimal customization. This study evaluates the effectiveness of a pre-trained GPT-4 model in grading semi-open handwritten responses in a university-level mathematics exam. Our findings indicate that GPT-4 provides surprisingly reliable and cost-effective initial grading, subject to subsequent human verification. Future research should focus on refining grading rules and enhancing the extraction of handwritten responses to further leverage these technologies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 5 citations worldwide. Full citation record

  1. Automated Grading of Students' Handwritten Graphs: A Comparison of Meta-Learning and Vision-Large Language Models

    cs.LG 2025-07 conditional novelty 5.0 of 10

    The best meta-learning models reach 56.9% in 2-way grading and the best vision-language models reach 50.0% in 3-way grading of handwritten economics graphs, both near chance levels.

  2. Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Cropping scanned exam pages to a question's bounding box improved handwritten-answer grading accuracy and cut token/compute cost across eight vision-language models.

  3. Pensieve Grader: An AI-Powered, Ready-to-Use Platform for Effortless Handwritten STEM Grading

    cs.AI 2025-07 reject novelty 4.0 of 10

    Pensieve Grader claims to cut grading time by 65% and match instructors on 95.4% of high-confidence grades, but the evidence is sparse and the time-savings model is self-referential.

Pith tools