Pith. sign in

REVIEW 4 cited by

Automatic Short Math Answer Grading via In-context Meta-learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.15219 v3 pith:CPIWXFWE submitted 2022-05-30 cs.CL cs.LG

classification cs.CLcs.LG
keywords languagemodelgradingquestionsanswerapproachesautomaticframework
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Automatic short answer grading is an important research direction in the exploration of how to use artificial intelligence (AI)-based tools to improve education. Current state-of-the-art approaches use neural language models to create vectorized representations of students responses, followed by classifiers to predict the score. However, these approaches have several key limitations, including i) they use pre-trained language models that are not well-adapted to educational subject domains and/or student-generated text and ii) they almost always train one model per question, ignoring the linkage across a question and result in a significant model storage problem due to the size of advanced language models. In this paper, we study the problem of automatic short answer grading for students' responses to math questions and propose a novel framework for this task. First, we use MathBERT, a variant of the popular language model BERT adapted to mathematical content, as our base model and fine-tune it for the downstream task of student response grading. Second, we use an in-context learning approach that provides scoring examples as input to the language model to provide additional context information and promote generalization to previously unseen questions. We evaluate our framework on a real-world dataset of student responses to open-ended math questions and show that our framework (often significantly) outperforms existing approaches, especially for new questions that are not seen during training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Veln(ia)s is in the Details: Evaluating LLM Judgment on Latvian and Lithuanian Short Answer Matching

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Open-source LLMs mostly detect fine-grained matched versus non-matched short answers in Latvian and Lithuanian, with 70b-class models near-perfect and smaller models showing variable, model-specific weaknesses.

  2. Automated Grading of Students' Handwritten Graphs: A Comparison of Meta-Learning and Vision-Large Language Models

    cs.LG 2025-07 conditional novelty 5.0 of 10

    The best meta-learning models reach 56.9% in 2-way grading and the best vision-language models reach 50.0% in 3-way grading of handwritten economics graphs, both near chance levels.

  3. How trust networks shape students' opinions about the proficiency of artificially intelligent assistants

    physics.soc-ph 2025-06 conditional novelty 5.0 of 10

    In simulated classrooms, trust relationships determine whether students accurately perceive an AI assistant's proficiency: allies-only networks converge correctly unless a single partisan dissenter is present, while o...

  4. Do Tutors Learn from Equity Training and Can Generative AI Assess It?

    cs.HC 2024-12 conditional novelty 5.0 of 10

    A scenario-based equity lesson produced only marginal learning gains for 81 tutors, but GPT-4o with few-shot prompting assessed tutors' equity responses with about 89% accuracy, making it the recommended low-cost option.

Pith tools