Pith. sign in

REVIEW 2 cited by

CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02323 v4 pith:767QURK7 submitted 2025-04-03 cs.CL

classification cs.CL
keywords scoringcotalchain-of-thoughtengineeringfeedbackpromptpromptingacross
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored various prompt engineering approaches in educational contexts, the degree to which these approaches generalize across domains--such as science, computing, and engineering--remains underexplored. In this paper, we introduce Chain-of-Thought Prompting + Active Learning (CoTAL), an LLM-based approach to formative assessment scoring that (1) leverages Evidence-Centered Design (ECD) to align assessments and rubrics with curriculum goals, (2) applies human-in-the-loop prompt engineering to automate response scoring, and (3) incorporates chain-of-thought (CoT) prompting and teacher and student feedback to iteratively refine questions, rubrics, and LLM prompts. Our findings demonstrate that CoTAL improves GPT-4's scoring performance across domains, achieving gains of up to 38.9% over a non-prompt-engineered baseline (i.e., without labeled examples, chain-of-thought prompting, or iterative refinement). Teachers and students judge CoTAL to be effective at scoring and explaining responses, and their feedback produces valuable insights that enhance grading accuracy and explanation quality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering

    cs.SE 2026-08 conditional novelty 6.0 of 10

    AgentForge lets novices practice software repair by playing one of four roles alongside AI agents; in a 37-person study, completion and self-reported learning were high, while code review proved the hardest role.

  2. Evidence-Decision-Feedback: Theory-Driven Adaptive Scaffolding for LLM Agents

    cs.MA 2026-02 conditional novelty 5.0 of 10

    EDF organizes LLM tutoring agents into evidence, decision, and feedback modules, and its Copa instantiation shows within-system correlations between task mastery, scaffold fading, and interpretable feedback in 33 dyads.

Pith tools