AI tutoring models systematically fail to detect student misconceptions when flawed reasoning coincidentally produces the correct answer, with 71% of failures concentrated in two predictable question types.
Innovating As- sessment with Conversational Agents: A Technology- Enhanced Approach to Formative Assessments
3 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.
fields
cs.CY 3years
2026 3representative citing papers
Verbalized confidence from small LMs enables cost-effective cascade routing for automated educational scoring, matching large-model accuracy at 76% lower cost when discrimination is strong.
LLM graders achieve substantial human agreement on math and science MCAS items but vary on ELA, performing best as sources of formative narrative feedback rather than summative numerical scores.
citing papers explorer
-
Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning
AI tutoring models systematically fail to detect student misconceptions when flawed reasoning coincidentally produces the correct answer, with 71% of failures concentrated in two predictable question types.
-
Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment
Verbalized confidence from small LMs enables cost-effective cascade routing for automated educational scoring, matching large-model accuracy at 76% lower cost when discrimination is strong.
-
Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering
LLM graders achieve substantial human agreement on math and science MCAS items but vary on ELA, performing best as sources of formative narrative feedback rather than summative numerical scores.