Pith. sign in

REVIEW 3 cited by

AI-assisted coding: Experiments with GPT-4

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.13187 v1 pith:5DIO3SMD submitted 2023-04-25 cs.AI cs.SE

classification cs.AIcs.SE
keywords codegpt-4experimentstoolscodingcomputerdemonstrateensure
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Artificial intelligence (AI) tools based on large language models have acheived human-level performance on some computer programming tasks. We report several experiments using GPT-4 to generate computer code. These experiments demonstrate that AI code generation using the current generation of tools, while powerful, requires substantial human validation to ensure accurate performance. We also demonstrate that GPT-4 refactoring of existing code can significantly improve that code along several established metrics for code quality, and we show that GPT-4 can generate tests with substantial coverage, but that many of the tests fail when applied to the associated code. These findings suggest that while AI coding tools are very powerful, they still require humans in the loop to ensure validity and accuracy of the results.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Autonomous Code Evolution Meets NP-Completeness

    cs.AI 2025-09 conditional novelty 7.0 of 10

    An LLM-based agent framework evolved five 2024 SAT solver codebases over 70 cycles and produced solvers that the authors report outperform the 2025 SAT Competition champions.

  2. Is Solving Better Than Evaluating GenAI Solutions?

    cs.CY 2026-07 conditional novelty 6.0 of 10

    Evaluating GenAI-generated solutions neither helped nor hurt exam and course performance in an algorithms course, despite raising homework scores.

  3. The Impact of AI on Educational Assessment: A Framework for Constructive Alignment

    cs.HC 2025-06 conditional novelty 4.0 of 10

    A framework linking Bloom's learning levels and AI-use levels to keep assessment valid, plus a small TU Delft lecturer survey suggesting staff are biased by their own AI use.

Pith tools