{"id":"41ed8b77-e0f6-4e14-a3b0-66a930a80b60","arxiv_id":"2502.10410","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Refining an LLM auto-evaluator with expert-teacher themes and few-shot examples improved agreement with human scores on quiz quality, but only on the same questions used for refinement.","lead":"A UK education body built an AI judge to grade the quality of AI-generated lesson quizzes, and improved its agreement with human teachers by refining the judge's instructions with examples from expert feedback. The case study shows a practical cycle for keeping AI teaching tools accurate and safe before they reach classrooms.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prompt refinement and few-shot exemplars were built from the same 311 MCQs used for the before/after evaluation, so the reported improvement does not establish generalizable alignment.","rationale":"The reader's weakest assumption correctly identified train/eval overlap. My stress-test sharpens it: not only was the thematic analysis derived from the same 311 MCQs, but the improved prompt embeds few-shot examples of specific MCQs from the evaluation set (e.g., 'Tobacco', 'probability 0.5', 'TNC'), so the after-condition has direct access to labels for those items. This makes the before/after QWK/MSE comparison uninterpretable as evidence of better auto-evaluation on new content. The paper is honest about other limitations (single rater, MCQ focus) but not this leakage. The appropriate outcome remains conditional acceptance: the case study illustrates an internal workflow but should not be cited as demonstrating generalizable improvement until a held-out or fresh evaluation is performed. Since the reader already reached CONDITIONAL, my verdict is unchanged.","tokens_in":12594,"tokens_out":2495,"duration_ms":22677,"concrete_test":"Split the 311 MCQs into development and held-out evaluation sets before any thematic analysis. Use only development MCQs to derive themes and few-shot exemplars and to refine the prompt. Then measure MSE and QWK before vs after on the held-out set, or better, on a fresh batch of Aila-generated MCQs unseen at prompt-construction time. If the improvement disappears or reverses, the reported effect is an artifact of in-sample tuning.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central evidence that the auto-evaluation agent improved after refinement is a before/after comparison on the same 311 MCQs used to construct the refinement. The Methods state that thematic analysis of teachers' justifications (for scores 1, 3, and 5, plus discrepancy cases) produced exemplar MCQs that were added to the prompt; those exemplars are visibly drawn from the evaluation pool (Appendix 4, Appendix 5), and the comparison in Table 1/Appendix 3 then scores those same questions. This is not merely a subtle statistical issue: the improved prompt contains few-shot demonstrations of exact items from the test set, so the model is being evaluated on memorized examples. Consequently, the drop in mean-based MSE (3.83 to 2.95, p = 0.00640) and rise in QWK (0.17 to 0.32) only show that the prompt can be tuned to fit these 311 items; they do not establish that alignment generalizes to new MCQs. The paper's limitation section acknowledges single-rater noise and MCQ-specificity, but not this leakage.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper describes Oak National Academy's development of an LLM-based auto-evaluation agent for assessing the quality of AI-generated lesson resources, focusing on a case study of the 'minimally different answers' criterion for multiple-choice quiz distractors. Twenty qualified teachers rated 311 MCQs, and the authors performed a thematic analysis of teacher justifications to refine the auto-evaluation prompt with additional criteria and few-shot examples. They report that after refinement, the agreement between the auto-evaluation agent and human raters improved, as measured by a decrease in mean-based MSE from 3.83 to 2.95 (p=0.00640) and an increase in QWK from 0.17 to 0.32. The paper positions this as evidence that human expert feedback can be incorporated into auto-evaluation prompts to improve alignment, and it makes the prompts and code publicly available.","tokens_in":12781,"tokens_out":3723,"duration_ms":30510,"significance":"The problem the paper addresses—evaluating AI-generated educational content at scale—is timely and important, and the authors bring a real deployment context, a large OER corpus, and open-sourced prompts. The thematic analysis of expert justifications is a sensible approach to making teacher knowledge explicit. However, the central evidence for improvement is compromised by the use of the same 311 questions for both prompt construction and evaluation: the few-shot examples in Appendix 5 are drawn from that exact evaluation set. As a result, the reported gains do not establish that the refined prompt would generalize to new MCQs. The paper's value is therefore primarily as a demonstration of in-sample prompt refinement and as a shared resource, not as a validated measure of alignment. If the authors re-run the evaluation on a held-out set, the contribution could be substantial.","major_comments":[{"comment":"The before/after comparison uses the same 311 MCQs from which the thematic analysis and few-shot examples were derived (Appendix 2 and Appendix 5). The improved prompt contains exact question-and-answer exemplars taken from this evaluation set, meaning the post-refinement scores are measured on data the prompt was explicitly tuned to. Consequently, the decrease in MSE from 3.83 to 2.95 and the QWK increase from 0.17 to 0.32 are in-sample results and do not support the claim that auto-evaluation alignment improves generally. Please add a held-out evaluation on MCQs not used in prompt construction, or substantially weaken the claim to describe fitting to a fixed benchmark.","section":"Analysis and Table 1 / Appendix 3"},{"comment":"The p-value of 0.00640 for the MSE decrease is reported without specifying the statistical test, whether it is paired, or how the 10 auto-evaluation runs were aggregated; the QWK increase is described as 'statistically significant' but no test, confidence interval, or p-value is given. Please provide the full statistical procedures so the reader can assess whether the comparison is valid.","section":"Results, Table 1, and Appendix 3"},{"comment":"The reference standard consists of a single human rater per MCQ, as acknowledged in the Discussion. This limits the maximum achievable agreement and makes the before/after comparison sensitive to rater noise. The paper should state the implications of single-rater labels for the interpretation of QWK and MSE, and ideally provide a reliability analysis (e.g., a small subset double-scored) to quantify the noise floor.","section":"Method"},{"comment":"The paper describes using 'the mean of the 10 scores given by the auto-evaluation per evaluation' but does not explain how the model was prompted to produce multiple scores for the same question, whether the outputs varied, and how the variance affects the reported metrics. Please clarify the aggregation procedure and report variance or a measure of instability across runs.","section":"Method/Results"}],"minor_comments":[{"comment":"The reference to 'D'Sa & Wisbal-Dionaldo, 20217' contains a typo in the year; it should be 2017.","section":"References and throughout"},{"comment":"The formatting of the improved prompt is inconsistent, with several 'Input:/Output:' labels appearing without corresponding examples or with misplaced line breaks, making it hard to follow the exact prompt structure.","section":"Appendix 5"},{"comment":"Figure 3 is described in the text but the figure itself is not visible in the version I reviewed; please ensure the figure is included and clearly labeled with axis definitions.","section":"Figure 3"},{"comment":"The paper states participants assessed 311 questions with an average of 16.4 per participant; after the exclusion of one participant it would be 19 participants, so please make the participant count explicit and consistent.","section":"Method"},{"comment":"The description of QWK 0.32 as a 'moderate to large' improvement is overstated; values in this range are typically considered fair to moderate. Please calibrate the interpretation or provide a reference for the thresholds used.","section":"Results/Appendix 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a practitioner case study with open code and prompts, which is refreshing. However, the lack of a held-out evaluation is a significant threat to the central claim. A revision that adds a held-out set (or re-frames the contribution as a methodology description) would be appropriate for this venue. I would lean toward major revision rather than rejection because the core approach is sound and the paper's limitations do not invalidate the qualitative findings about distractor quality."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline result is an in-sample fit, not a validated improvement. The genuinely useful pieces here are the thematic categories for MCQ distractor quality (plausibility, commonality, structural coherence), the concrete few-shot examples, and the open release of prompts and code on GitHub. If you work on content evaluation for edtech, this is worth reading and potentially building on.\n\nThe weak point is exactly where the stress-test points. The improved prompt was built by analyzing human justifications on the same 311 MCQs and then adding few-shot examples drawn from that pool. The before/after comparison then scores those same questions, and some of the exemplar MCQs in Appendix 5 are visibly exact items from the evaluation set. That means the drop in mean-based MSE (3.83 to 2.95, p = 0.00640) and the QWK rise (0.17 to 0.32) do not establish that the agent generalizes to new lessons. The p-value is not a license to infer an underlying improvement in alignment, because the sample was used to fit the prompt. The paper acknowledges single-rater noise and MCQ-specificity, but it does not acknowledge this leakage, and that omission matters for the empirical claim.\n\nI do not think this kills the paper. The thematic analysis is still informative, the method is clearly described, and the prompt-engineering loop is a sensible internal QA process. But the central claim should be reframed as \"we can fit this auto-evaluator to our existing corpus using expert justifications,\" not \"the evaluator is now more accurate.\" A held-out set of MCQs or a fresh batch of Aila lessons, scored before and after, would turn this into a solid validation.\n\nFor a serious referee: yes. The paper has a real, small contribution, open artifacts, and a correctable flaw. I would send it out with a request for either held-out validation or a clearly restricted claim. I would not cite the quantitative improvement as evidence, but I might cite the codified themes and open prompt as a useful applied example.","headline":"Useful thematic categories and open code, but the reported before/after agreement is in-sample prompt tuning, not a validated improvement.","tokens_in":13300,"tokens_out":3382,"would_cite":true,"duration_ms":30836,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An LLM-based auto-evaluation agent for quiz distractors can be made to agree far more closely with expert teacher ratings by rewriting its prompt around themes and examples drawn from teachers' own justifications.","keywords":["auto-evaluation agent","LLM-as-judge","AI-generated lesson resources","multiple-choice questions","distractor quality","prompt refinement","human-AI alignment","teacher expertise"],"falsifier":"Take a fresh set of quiz questions that were not part of the 311 used here, have qualified teachers score them on the 'minimally different answers' criterion, run both the original and the refined auto-evaluation prompts on the same questions, and compare MSE and QWK against the human scores. If the refined prompt does not beat the original prompt on this held-out set by roughly the same margin (or at all), the reported improvement is prompt tuning to the training questions rather than a general gain in alignment with teachers.","tokens_in":12417,"feed_emoji":"🎓","tokens_out":9324,"duration_ms":68233,"temperature":0.7,"pith_summary":"This paper argues that a scalable quality check for AI-generated lesson content can be made to track expert teacher judgement if the evaluator's prompt is refined using teachers' own written justifications. In the case study, the auto-evaluation agent is an LLM used as a judge, scoring whether the wrong answers in a multiple-choice quiz are 'minimally different' from the correct answer—plausible distractors that still test understanding rather than obvious fillers. Qualified teachers scored 311 quiz questions on that same criterion and explained their scores; a thematic analysis of those explanations was then folded into a rewritten prompt with example inputs and outputs. After the rewrite, the agent's mean squared disagreement with the teachers dropped from 3.83 to 2.95 (p = 0.00640) and its quadratic weighted kappa agreement rose from 0.17 to 0.32. The practical point is that a sufficiently aligned auto-evaluator lets a small team check thousands of generated lessons for quality and safety instead of reading every one.","feed_headline":"AI quiz grader tuned with teacher feedback cuts error score by 23%","feed_subtitle":"Teacher justifications built into the prompt raised agreement with human scorers on 311 quiz questions","key_machinery":"The carrying mechanism is the auto-evaluation prompt itself, used in an LLM-as-judge setup in which the model (gpt-4o-2024-08-06, temperature 0.5) returns a score and a justification for each of 24 quality and safety benchmarks. The case study focuses on one benchmark, 'answers are minimally different,' scored on a 1–5 Likert scale, with ten runs per question averaged to produce the auto-score. The refinement machinery has three parts: a thematic analysis of expert teachers' justifications that codifies weak versus strong distractors; a set of exemplar questions and outputs that turn that codification into few-shot examples; and a rewritten prompt that lists three explicit criteria—plausibility, commonality, and structural coherence—each with sample inputs and outputs. The prompt carries the argument because it is the only component that changes between the before and after comparison; the underlying model, the questions, and the human scores are held fixed.","core_discovery":"The central claim is that an 'LLM-as-judge' auto-evaluation agent can act as a reliable proxy for expert teachers in assessing AI-generated lesson resources, provided its prompts are built from and refined against expert human judgement. The paper's illustrative case study concerns distractor quality in multiple-choice quizzes: twenty qualified teachers evaluated 311 quiz questions using the same 1–5 Likert scale as the automated agent, and their justifications were thematically coded. Weak distractors were characterised by opposite sentiment to the correct answer, different grammatical structure, or the correct answer echoing the question; strong distractors shared a category with the correct answer, related to a common theme, included common misconceptions, and matched the structure of the correct answer. These themes, together with exemplar questions, were added to the evaluation prompt as explicit criteria and few-shot examples. On the same 311 questions, the refined agent's scores moved closer to the teachers' scores: mean-based MSE fell from 3.83 to 2.95 (p = 0.00640), QWK rose from 0.17 to 0.32, exact agreement rose from 19% to 27%, and the proportion of questions on which the agent scored lower than the teacher fell from 75% to 62%.","pith_inferences":["The reported gains are measured on the same 311 questions used to build the refined prompt, so the before/after comparison is likely to overstate how much the agent would improve on unseen lessons; a held-out test set would settle the size of the real effect.","The three codified criteria (plausibility, commonality, structural coherence) are probably portable across subjects and key stages, but their relative weights may need local recalibration where curricula emphasise different misconceptions.","The same rubric could be inverted into a generation-time constraint: the lesson generator could be instructed to avoid opposite-sentiment distractors and structurally mismatched options, making low-quality items rarer rather than merely better flagged.","Using multiple human raters per question, and weighting by teacher experience, would produce a less noisy gold standard; with a cleaner target, the auto-evaluator's observed agreement ceiling (QWK = 0.32) is likely an underestimate of what prompt refinement could achieve."],"forward_implications":["The same human-in-the-loop refinement cycle can be extended to the other quality and safety benchmarks in the evaluation suite, including bias, misconceptions, and quiz progression.","With an auto-evaluator whose scores track expert teachers, generation changes such as swapping the underlying model or altering retrieval settings can be compared on large lesson sets before release.","Because the agent returns a written justification with each score, its output can flag specific parts of a lesson for a teacher to check, turning evaluation into targeted feedback.","Openly releasing the evaluation prompts and code gives other organisations a concrete starting point for building their own teacher-aligned content checks.","High-scoring lessons that receive aligned human and auto scores can be used as training data for fine-tuning the lesson generator, completing the loop from evaluation to generation."],"supporting_citations":[{"why":"Supplies the LLM-as-Judge methodology that the auto-evaluation agent is built on.","marker":"Chiang & Lee, 2023"},{"why":"Defines the evidence-informed curriculum principles from which the 24 evaluation benchmarks, including the 'minimally different answers' criterion, are derived.","marker":"McCrea, 2023"},{"why":"Provides the multiple-choice item-analysis rationale (distractor efficiency) that makes distractor quality a meaningful thing to benchmark.","marker":"D'Sa & Wisbal-Dionaldo, 2017"},{"why":"Supports the authors' proposed next step of fine-tuning the lesson generator on high-quality lessons identified through aligned auto and human evaluation.","marker":"Ouyang et al., 2022"},{"why":"Provides the retrieval-augmented generation evidence (accuracy gains from a high-quality corpus) that motivates Aila's content-anchored design.","marker":"Government Social Research, 2024"}],"fun_headline_variants":["Teacher-driven prompts sharpen AI quiz evaluator","AI judge improved with teacher justifications","Refined auto-evaluator aligns closer with teachers","Teacher refinement cuts AI quiz scoring error by 23%","Prompt tuning with teacher feedback boosts quiz scoring"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the before/after improvement reflects genuine better alignment rather than overfitting, but the comparison is made on the same 311 questions whose teacher justifications were used to build the refined prompt, so the gain on unseen lessons is not yet established.","fun_headline_variants_meta":{"raw":{"variants":["Teacher-driven prompts sharpen AI quiz evaluator","AI judge improved with teacher justifications","Refined auto-evaluator aligns closer with teachers","Teacher refinement cuts AI quiz scoring error by 23%","Prompt tuning with teacher feedback boosts quiz scoring"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000677,"raw_usage":{"total_tokens":3111,"prompt_tokens":1008,"completion_tokens":2103,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":2033}},"tokens_in":624,"tokens_out":2103,"duration_ms":14718,"temperature":1.0,"reasoning_tokens":2033,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:48:04.379537+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fresh set of quiz questions that were not part of the 311 used here, have qualified teachers score them on the 'minimally different answers' criterion, run both the original and the refined auto-evaluation prompts on the same questions, and compare MSE and QWK against the human scores. If the refined prompt does not beat the original prompt on this held-out set by roughly the same margin (or at all), the reported improvement is prompt tuning to the training questions rather than a general gain in alignment with teachers.","supporting_citations":[],"review_version":1}