A new benchmark of 55 algebra misconceptions and 220 examples shows GPT-4-turbo diagnoses around 53% of misconceptions overall, 75% when topic-constrained, and 83.9% when educator feedback is included.
Exploring Automated Distractor Generation for Math Multiple-choice Questions via Large Language Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Multiple-choice questions (MCQs) are ubiquitous in almost all levels of education since they are easy to administer, grade, and are a reliable format in assessments and practices. One of the most important aspects of MCQs is the distractors, i.e., incorrect options that are designed to target common errors or misconceptions among real students. To date, the task of crafting high-quality distractors largely remains a labor and time-intensive process for teachers and learning content designers, which has limited scalability. In this work, we study the task of automated distractor generation in the domain of math MCQs and explore a wide variety of large language model (LLM)-based approaches, from in-context learning to fine-tuning. We conduct extensive experiments using a real-world math MCQ dataset and find that although LLMs can generate some mathematically valid distractors, they are less adept at anticipating common errors or misconceptions among real students.
citation-role summary
citation-polarity summary
fields
cs.HC 1years
2024 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
A Benchmark for Math Misconceptions: Bridging Gaps in Middle School Algebra with AI-Supported Instruction
A new benchmark of 55 algebra misconceptions and 220 examples shows GPT-4-turbo diagnoses around 53% of misconceptions overall, 75% when topic-constrained, and 83.9% when educator feedback is included.