REVIEW 3 cited by
AI-assisted coding: Experiments with GPT-4
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Artificial intelligence (AI) tools based on large language models have acheived human-level performance on some computer programming tasks. We report several experiments using GPT-4 to generate computer code. These experiments demonstrate that AI code generation using the current generation of tools, while powerful, requires substantial human validation to ensure accurate performance. We also demonstrate that GPT-4 refactoring of existing code can significantly improve that code along several established metrics for code quality, and we show that GPT-4 can generate tests with substantial coverage, but that many of the tests fail when applied to the associated code. These findings suggest that while AI coding tools are very powerful, they still require humans in the loop to ensure validity and accuracy of the results.
Forward citations
Cited by 3 Pith papers
-
Autonomous Code Evolution Meets NP-Completeness
An LLM-based agent framework evolved five 2024 SAT solver codebases over 70 cycles and produced solvers that the authors report outperform the 2025 SAT Competition champions.
-
Is Solving Better Than Evaluating GenAI Solutions?
Evaluating GenAI-generated solutions neither helped nor hurt exam and course performance in an algorithms course, despite raising homework scores.
-
The Impact of AI on Educational Assessment: A Framework for Constructive Alignment
A framework linking Bloom's learning levels and AI-use levels to keep assessment valid, plus a small TU Delft lecturer survey suggesting staff are biased by their own AI use.
Discussion (0). Sign in to comment.