Pith. sign in

REVIEW 1 cited by

Generative AI for Programming Education: Benchmarking ChatGPT, GPT-4, and Human Tutors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.17156 v3 pith:6RO4XJDN submitted 2023-06-29 cs.CY cs.AIcs.CL

classification cs.CYcs.AIcs.CL
keywords modelsprogrammingeducationgpt-4performancescenarioschatgpthuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative AI and large language models hold great promise in enhancing computing education by powering next-generation educational technologies for introductory programming. Recent works have studied these models for different scenarios relevant to programming education; however, these works are limited for several reasons, as they typically consider already outdated models or only specific scenario(s). Consequently, there is a lack of a systematic study that benchmarks state-of-the-art models for a comprehensive set of programming education scenarios. In our work, we systematically evaluate two models, ChatGPT (based on GPT-3.5) and GPT-4, and compare their performance with human tutors for a variety of scenarios. We evaluate using five introductory Python programming problems and real-world buggy programs from an online platform, and assess performance using expert-based annotations. Our results show that GPT-4 drastically outperforms ChatGPT (based on GPT-3.5) and comes close to human tutors' performance for several scenarios. These results also highlight settings where GPT-4 still struggles, providing exciting future directions on developing techniques to improve the performance of these models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. DS@GT at CheckThat! 2025: Ensemble Methods for Detection of Scientific Discourse on Social Media

    cs.CL 2025-07 conditional novelty 4.0 of 10

    The DS@GT system achieved 0.8611 macro-F1 on the CheckThat! 2025 Task 4a development set by combining a fine-tuned DeBERTa model with GPT-4o few-shot prompting, outperforming the DeBERTaV3 baseline of 0.8375.

Pith tools