Pith. sign in

REVIEW 3 cited by

Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.09136 v1 pith:CH5LZXLO submitted 2024-07-12 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords studenterrorsmodelsreasoningexistinggenerationlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) present an opportunity to scale high-quality personalized education to all. A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving. However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor their feedback to these errors. Inspired by real-world teaching practice where teachers identify student errors and customize their response based on them, we focus on verifying student solutions and show how grounding to such verification improves the overall quality of tutor response generation. We collect a dataset of 1K stepwise math reasoning chains with the first error step annotated by teachers. We show empirically that finding the mistake in a student solution is challenging for current models. We propose and evaluate several verifiers for detecting these errors. Using both automatic and human evaluation we show that the student solution verifiers steer the generation model towards highly targeted responses to student errors which are more often correct with less hallucinations compared to existing baselines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors

    cs.CY 2025-07 accept novelty 6.0 of 10

    Best submitted systems scored 58-72 macro F1 on four three-class pedagogical assessment tracks and 97 macro F1 on nine-class tutor identification, showing automatic evaluation of AI math tutors works but still has roo...

  2. Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A large LLM-to-LLM math tutoring simulation across 11 languages shows English-language hints often yield the largest accuracy gains for student models, but the low-resource-language results lack statistical support.

  3. BD at BEA 2025 Shared Task: MPNet Ensembles for Pedagogical Mistake Identification and Localization in AI Tutor Responses

    cs.CL 2025-06 conditional novelty 4.0 of 10

    An ensemble of MPNet classifiers trained with grouped cross-validation and class-weighted loss reaches competitive macro-F1 on mistake identification and location in AI tutor responses.

Pith tools