Pith. sign in

REVIEW 4 cited by

Transforming Software Development: Evaluating the Efficiency and Challenges of GitHub Copilot in Real-World Projects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17910 v1 pith:P3BSKATX submitted 2024-06-25 cs.SE cs.AI

classification cs.SEcs.AI
keywords codecopilotdevelopmenttaskschallengessoftwarebenefitscoding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative AI technologies promise to transform the product development lifecycle. This study evaluates the efficiency gains, areas for improvement, and emerging challenges of using GitHub Copilot, an AI-powered coding assistant. We identified 15 software development tasks and assessed Copilot's benefits through real-world projects on large proprietary code bases. Our findings indicate significant reductions in developer toil, with up to 50% time saved in code documentation and autocompletion, and 30-40% in repetitive coding tasks, unit test generation, debugging, and pair programming. However, Copilot struggles with complex tasks, large functions, multiple files, and proprietary contexts, particularly with C/C++ code. We project a 33-36% time reduction for coding-related tasks in a cloud-first software development lifecycle. This study aims to quantify productivity improvements, identify underperforming scenarios, examine practical benefits and challenges, investigate performance variations across programming languages, and discuss emerging issues related to code quality, security, and developer experience.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code

    cs.SE 2026-07 conditional novelty 6.5 of 10

    Post-merge, agentic code needs ~46–51% more corrective/bug-fix maintenance and introduces more security and dependency findings than human code, with higher burden in low-review projects.

  2. Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Persistent SAEs learn per-feature persistence coefficients from reconstruction, splitting features into fast local detectors and slow topic-tracking states that retain prompt-injection signals over long contexts.

  3. Three-Phase Evaluation of AI-Assisted Software Development Life Cycle

    cs.SE 2026-07 conditional novelty 5.5 of 10

    In a sequential three-phase pilot, higher AI autonomy and AWS Kiro versus GitHub Copilot were associated with fewer development hours, higher RITM scores, and lower mental demand, with modestly higher frustration.

  4. The Ground Is Shifting: A Reflection on the Foundations of Software Measurement

    cs.SE 2026-08 conditional novelty 5.0 of 10

    AI-generated commits and squash-merging violate the human-origin assumption underlying software measurement, and the field should run a large-scale AI-assisted replication agenda.

Pith tools