Pith. sign in

REVIEW 2 cited by

AI-enhanced Auto-correction of Programming Exercises: How Effective is GPT-3.5?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.10737 v1 pith:ERO6F4FA submitted 2023-10-24 cs.CY cs.AIcs.HCcs.SE

classification cs.CYcs.AIcs.HCcs.SE
keywords feedbackcodegpt-3effectiveerrorscorrectnessgeneratedlarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Timely formative feedback is considered as one of the most important drivers for effective learning. Delivering timely and individualized feedback is particularly challenging in large classes in higher education. Recently Large Language Models such as GPT-3 became available to the public that showed promising results on various tasks such as code generation and code explanation. This paper investigates the potential of AI in providing personalized code correction and generating feedback. Based on existing student submissions of two different real-world assignments, the correctness of the AI-aided e-assessment as well as the characteristics such as fault localization, correctness of hints, and code style suggestions of the generated feedback are investigated. The results show that 73 % of the submissions were correctly identified as either correct or incorrect. In 59 % of these cases, GPT-3.5 also successfully generated effective and high-quality feedback. Additionally, GPT-3.5 exhibited weaknesses in its evaluation, including localization of errors that were not the actual errors, or even hallucinated errors. Implications and potential new usage scenarios are discussed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. To Google or To ChatGPT? A Comparison of CS2 Students' Information Gathering Approaches and Outcomes

    cs.HC 2025-01 conditional novelty 6.0 of 10

    In a 32-participant within-subjects study, CS2 students scored significantly better on a conceptual quiz for currying when learning via web search than via ChatGPT, while query behavior differed markedly between the t...

  2. Automating Autograding: Large Language Models as Test Suite Generators for Introductory Programming

    cs.CY 2024-11 conditional novelty 6.0 of 10

    GPT-4-generated test suites for 26 CS1 problems identified 92.8% of valid student solutions and caught invalid solutions more often than instructor suites did.

Pith tools