REVIEW 2 cited by
CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored various prompt engineering approaches in educational contexts, the degree to which these approaches generalize across domains--such as science, computing, and engineering--remains underexplored. In this paper, we introduce Chain-of-Thought Prompting + Active Learning (CoTAL), an LLM-based approach to formative assessment scoring that (1) leverages Evidence-Centered Design (ECD) to align assessments and rubrics with curriculum goals, (2) applies human-in-the-loop prompt engineering to automate response scoring, and (3) incorporates chain-of-thought (CoT) prompting and teacher and student feedback to iteratively refine questions, rubrics, and LLM prompts. Our findings demonstrate that CoTAL improves GPT-4's scoring performance across domains, achieving gains of up to 38.9% over a non-prompt-engineered baseline (i.e., without labeled examples, chain-of-thought prompting, or iterative refinement). Teachers and students judge CoTAL to be effective at scoring and explaining responses, and their feedback produces valuable insights that enhance grading accuracy and explanation quality.
Forward citations
Cited by 2 Pith papers
-
AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering
AgentForge lets novices practice software repair by playing one of four roles alongside AI agents; in a 37-person study, completion and self-reported learning were high, while code review proved the hardest role.
-
Evidence-Decision-Feedback: Theory-Driven Adaptive Scaffolding for LLM Agents
EDF organizes LLM tutoring agents into evidence, decision, and feedback modules, and its Copa instantiation shows within-system correlations between task mastery, scaffold fading, and interpretable feedback in 33 dyads.
Discussion (0). Continue with ORCID to comment.