Pith. sign in

REVIEW 2 cited by

Can Large Language Models Replicate ITS Feedback on Open-Ended Math Questions?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.06414 v2 pith:P6LFAWV5 submitted 2024-05-10 cs.CL

classification cs.CL
keywords feedbackerrorslargellmsmathmodelsopen-endedquestions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Intelligent Tutoring Systems (ITSs) often contain an automated feedback component, which provides a predefined feedback message to students when they detect a predefined error. To such a feedback component, we often resort to template-based approaches. These approaches require significant effort from human experts to detect a limited number of possible student errors and provide corresponding feedback. This limitation is exemplified in open-ended math questions, where there can be a large number of different incorrect errors. In our work, we examine the capabilities of large language models (LLMs) to generate feedback for open-ended math questions, similar to that of an established ITS that uses a template-based approach. We fine-tune both open-source and proprietary LLMs on real student responses and corresponding ITS-provided feedback. We measure the quality of the generated feedback using text similarity metrics. We find that open-source and proprietary models both show promise in replicating the feedback they see during training, but do not generalize well to previously unseen student errors. These results suggest that despite being able to learn the formatting of feedback, LLMs are not able to fully understand mathematical errors made by students.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A large LLM-to-LLM math tutoring simulation across 11 languages shows English-language hints often yield the largest accuracy gains for student models, but the low-resource-language results lack statistical support.

  2. Exploring LLM-Generated Feedback for Economics Essays: How Teaching Assistants Evaluate and Envision Its Use

    cs.HC 2025-05 conditional novelty 5.0 of 10

    In a think-aloud study, five economics teaching assistants found AI-generated essay feedback useful as suggestions, especially when it included highlighted evidence and intermediate judgments.

Pith tools