Pith. sign in

REVIEW 2 cited by

Exploring Automated Distractor Generation for Math Multiple-choice Questions via Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.02124 v3 pith:SQRGUQCE submitted 2024-04-02 cs.CL

classification cs.CL
keywords distractorsmathmcqsautomatedcommondistractorerrorsgeneration
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multiple-choice questions (MCQs) are ubiquitous in almost all levels of education since they are easy to administer, grade, and are a reliable format in assessments and practices. One of the most important aspects of MCQs is the distractors, i.e., incorrect options that are designed to target common errors or misconceptions among real students. To date, the task of crafting high-quality distractors largely remains a labor and time-intensive process for teachers and learning content designers, which has limited scalability. In this work, we study the task of automated distractor generation in the domain of math MCQs and explore a wide variety of large language model (LLM)-based approaches, from in-context learning to fine-tuning. We conduct extensive experiments using a real-world math MCQ dataset and find that although LLMs can generate some mathematically valid distractors, they are less adept at anticipating common errors or misconceptions among real students.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Role-Aware Multi-Agent Framework for Financial Education Question Answering with LLMs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    A role-aware multi-agent pipeline with retrieval and expert critique raises financial multiple-choice accuracy by 6.6-8.3 percentage points over zero-shot CoT across four LLMs.

  2. A Benchmark for Math Misconceptions: Bridging Gaps in Middle School Algebra with AI-Supported Instruction

    cs.HC 2024-12 conditional novelty 5.0 of 10

    A new benchmark of 55 algebra misconceptions and 220 examples shows GPT-4-turbo diagnoses around 53% of misconceptions overall, 75% when topic-constrained, and 83.9% when educator feedback is included.

Pith tools