Pith. sign in

REVIEW 2 cited by

Learning gain differences between ChatGPT and human tutor generated algebra hints

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.06871 v1 pith:JEJRUEUO submitted 2023-02-14 cs.CY cs.CLcs.HC

classification cs.CYcs.CLcs.HC
keywords chatgpthintsalgebracontenthumanlearninggainswere
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs), such as ChatGPT, are quickly advancing AI to the frontiers of practical consumer use and leading industries to re-evaluate how they allocate resources for content production. Authoring of open educational resources and hint content within adaptive tutoring systems is labor intensive. Should LLMs like ChatGPT produce educational content on par with human-authored content, the implications would be significant for further scaling of computer tutoring system approaches. In this paper, we conduct the first learning gain evaluation of ChatGPT by comparing the efficacy of its hints with hints authored by human tutors with 77 participants across two algebra topic areas, Elementary Algebra and Intermediate Algebra. We find that 70% of hints produced by ChatGPT passed our manual quality checks and that both human and ChatGPT conditions produced positive learning gains. However, gains were only statistically significant for human tutor created hints. Learning gains from human-created hints were substantially and statistically significantly higher than ChatGPT hints in both topic areas, though ChatGPT participants in the Intermediate Algebra experiment were near ceiling and not even with the control at pre-test. We discuss the limitations of our study and suggest several future directions for the field. Problem and hint content used in the experiment is provided for replicability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Benchmark for Math Misconceptions: Bridging Gaps in Middle School Algebra with AI-Supported Instruction

    cs.HC 2024-12 conditional novelty 5.0 of 10

    A new benchmark of 55 algebra misconceptions and 220 examples shows GPT-4-turbo diagnoses around 53% of misconceptions overall, 75% when topic-constrained, and 83.9% when educator feedback is included.

  2. A Novel Approach to Scalable and Automatic Topic-Controlled Question Generation in Education

    cs.CY 2025-01 conditional novelty 4.0 of 10

    A contrastive mixed-context fine-tuning method lets a 60M-parameter T5 model generate topic-controlled educational questions, with best topical alignment from augmented data and a Jaccard Wikipedia metric.

Pith tools