Pith. sign in

REVIEW 2 major objections 1 minor 59 references

AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate

T0 review · 2 major / 1 minor · reviewed 2026-07-03 · grok-4.3

Pith's one-line read An enterprise mandate for AI coding tools doubled per-developer throughput to 2.09 times the pre-mandate level by April 2026.

desk verdict This paper supplies one of the largest real-world panels on an enterprise AI coding mandate, with a documented 2.09x throughput rise tied to adoption, but the staggered DiD leaves room for selection effects. read the letter →

arxiv 2607.01904 v1 pith:XEBVFJTJ submitted 2026-07-02 cs.SE

classification cs.SE
keywords AIcodingtoolsproductivitygainspullrequestsdifference-in-differencescodereviewautomationenterpriseadoptionlongitudinalanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tracks productivity at a company that required its developers to use AI coding assistants in pursuit of doubling output. Using data on 802 developers and 196,212 pull requests from early 2024 to spring 2026, it finds that average output per person rose to more than double the starting point. A staggered difference-in-differences analysis connects the increase within each developer to when they began using the tools and how long they had been using them. The company mandate appears to have accelerated adoption rather than directly causing the gains. Review workloads shifted heavily toward automation while key quality indicators stayed constant.

What carries the argument

Staggered difference-in-differences design comparing each developer's output before and after their personal adoption date to measure the contribution of AI use.

What would settle it

Observing no productivity increase in a randomized controlled trial where some developers receive AI tools and others do not would falsify the link between adoption and the observed gains.

Watch

Extended reading notes

Core claim

In a panel of 802 developers and 196,212 pull requests spanning January 2024 to April 2026, per-capita throughput eventually doubled, reaching 2.09x the pre-mandate baseline in April 2026. A staggered difference-in-differences design links the within-developer share of this gain to AI adoption and to a further gain that grows with accumulated use, with the mandate acting as a catalyst rather than a direct driver. Adoption also restructured code review around automation: per-reviewer load roughly doubled and automated review overtook human review, while merge and revert rates held steady.

Load-bearing premise

The staggered timing of individual developers' AI adoptions creates a valid counterfactual for estimating the effect of AI use on their productivity.

Editorial extensions

If this is right

  • Throughput gains from AI tools reached 2.09 times baseline when adoption was promoted via mandate.
  • Gains increased with the length of accumulated AI use.
  • Code review load per reviewer doubled while automated review became the majority.
  • Quality measures such as merge and revert rates remained unchanged.
  • The productivity increase was shared across different seniority levels.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the adoption channel holds, companies without mandates may see smaller or slower gains from the same tools.
  • The concentration of gains in newer code suggests AI may be more effective for initial development than for maintenance.
  • Review process redesigns may be needed as automated checks scale with higher code volume.
  • Similar studies in other firms could test whether the 2x target is replicable outside this AI-forward setting.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript reports a longitudinal case study of an enterprise '2x' mandate to double merged pull requests per engineer via AI coding tools. In a panel of 802 developers and 196,212 pull requests (Jan 2024–Apr 2026), per-capita throughput reached 2.09× the pre-mandate baseline by April 2026. A staggered difference-in-differences design attributes the within-developer share of the gain to AI adoption and to further increases that grow with accumulated use, with the mandate acting as a catalyst. The authors explicitly note non-random assignment of adoption and usage and frame the evidence as implicating an adoption-and-use channel rather than exact causal attribution. The study also documents restructuring of code review (per-reviewer load doubled, automated review overtaking human review) while merge and revert rates remained stable.

Significance. If the identification holds, the work supplies one of the largest-scale longitudinal field deployments of AI coding tools, documenting substantial throughput gains and workflow shifts in a real enterprise setting. The dataset size, multi-year span, and explicit caveat on non-random assignment provide a rare quantitative window into mandate-driven adoption. Credit is due for the direct reporting of the 2.09× ratio as an observed throughput measure rather than a fitted parameter and for the cautious interpretation of the DiD results.

major comments (2)
  1. [Staggered difference-in-differences design] The staggered DiD design (described in the methods and results sections) is load-bearing for the central claim that within-developer gains are linked to AI adoption and accumulated use. The manuscript does not report event-study pre-trends, tests for selection on gains, or robustness to Callaway-Sant'Anna or Sun-Abraham estimators. Given the abstract's explicit statement that adoption was not randomly assigned, these checks are needed to assess whether timing of adoption supplies a valid counterfactual or whether the post-adoption coefficients partly reflect selection.
  2. [Results on accumulated use] The claim that 'a further gain that grows with accumulated use' is linked to AI (abstract and results) requires a clear specification of the usage-intensity measure and its interaction with time since adoption. Without reported robustness to developer-specific trends or alternative specifications, this component of the within-developer attribution remains vulnerable to the same selection concerns noted above.
minor comments (1)
  1. [Abstract and results] The abstract states the gain is 'broadly shared across seniority yet concentrated in newer code and not separable across model generations.' A table or figure breaking out these heterogeneity results by seniority, code age, and model would improve clarity and allow readers to assess the scope of the findings.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments. We address each major point below, agreeing that the suggested robustness checks would strengthen the identification section and committing to revisions.

read point-by-point responses
  1. Referee: [Staggered difference-in-differences design] The staggered DiD design (described in the methods and results sections) is load-bearing for the central claim that within-developer gains are linked to AI adoption and accumulated use. The manuscript does not report event-study pre-trends, tests for selection on gains, or robustness to Callaway-Sant'Anna or Sun-Abraham estimators. Given the abstract's explicit statement that adoption was not randomly assigned, these checks are needed to assess whether timing of adoption supplies a valid counterfactual or whether the post-adoption coefficients partly reflect selection.

    Authors: We agree that additional checks would strengthen the paper. Although the manuscript already caveats non-random assignment and frames results as implicating an adoption-and-use channel rather than exact causality, we will add event-study pre-trend plots, tests for selection on gains, and robustness using Callaway-Sant'Anna and Sun-Abraham estimators in the revised version. revision: yes

  2. Referee: [Results on accumulated use] The claim that 'a further gain that grows with accumulated use' is linked to AI (abstract and results) requires a clear specification of the usage-intensity measure and its interaction with time since adoption. Without reported robustness to developer-specific trends or alternative specifications, this component of the within-developer attribution remains vulnerable to the same selection concerns noted above.

    Authors: We will expand the methods section to explicitly define the usage-intensity measure (cumulative AI-assisted PRs) and its interaction with time since adoption. We will also add robustness specifications that include developer-specific trends and alternative functional forms for the accumulated-use term. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: direct ratios and standard DiD on observational panel

full rationale

The paper's central results are observed per-capita throughput ratios (2.09x) computed directly from the 196,212 PR panel and a conventional staggered difference-in-differences specification linking within-developer changes to adoption timing. No equations reduce a claimed prediction to a fitted input by construction, no self-citations supply load-bearing uniqueness theorems or ansatzes, and the authors explicitly qualify the design as implicating rather than proving causality. The derivation chain is therefore self-contained against the raw data and standard econometric methods.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The analysis rests on standard econometric assumptions for staggered DiD without introducing new free parameters, invented entities, or ad-hoc constructs beyond the observed data.

assumptions (1)
  • domain assumption Parallel trends assumption holds for the staggered adoption timing in the difference-in-differences design.
    Invoked to attribute within-developer gains to AI adoption and usage intensity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate." pith.science (2026). https://pith.science/paper/XEBVFJTJ

@misc{pith2026260701904,
  author       = {Pith},
  title        = {Pith review of: AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XEBVFJTJ}},
  note         = {Machine review of arXiv:2607.01904}
}
read the original abstract

Enterprises increasingly mandate AI coding tools and report large productivity gains, yet longitudinal evidence on how such a mandate unfolds is scarce. In this paper, we present a quantitative case study of a documented enterprise "2x" mandate at a mid-sized, AI-forward company that has been committed to doubling merged pull requests per engineer since mid-2025. In a panel of 802 developers and 196,212 pull requests (January 2024-April 2026), per-capita throughput eventually doubled, reaching 2.09x the pre-mandate baseline in April 2026, among the largest gains reported from a field deployment of AI coding tools to our knowledge. A staggered difference-in-differences design links the within-developer share of this gain to AI adoption and to a further gain that grows with accumulated use, with the mandate acting as a catalyst rather than a direct driver. Because adoption and usage intensity were not randomly assigned, we read this evidence as strongly implicating an adoption-and-use channel rather than as exact causal attribution. The gain is broadly shared across seniority yet concentrated in newer code and not separable across model generations. Adoption also restructured code review around automation: per-reviewer load roughly doubled and automated review overtook human review, while merge and revert rates held steady.

Figures

Figures reproduced from arXiv: 2607.01904 by the authors.

Figure 1
Figure 1. Monthly AI tool users. Inset: Claude Code token spend. a) Pull-request (PR) and review history: We extract full PR history: title, author (anonymized), state, timestamps, size (lines added/removed, files changed), labels, and the associated commits, comments, and review events. Then, we derive per￾PR throughput, review counts, reverts, and how long each PR takes to travel from first commit to merge (or close)—a metr… view at source ↗
Figure 2
Figure 2. Six simulated PR trajectories consistent with our data. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. The share of AI-authored PRs climbs from near zero [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (10 more)
Figure 5
Figure 5. Figure 5: Output rises more for developers who use AI more [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: The share of PRs receiving any human review (black) falls to 68%, and automated review—AI review bots (purple)—overtakes it shortly after the mandate, reaching 84%. 16 4 0 4 8 12 16 Jan 24 Apr 24 Jul 24 Oct 24 Jan 25 Apr 25 Jul 25 Oct 25 Jan 26 Apr 26 median reviews / …
Figure 7
Figure 7. Figure 7: Review activity rises after adoption, faster in light [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 9
Figure 9. Figure 9: Mean reviews per reviewer, decomposed into com [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: The human-review share decomposed into substantive (navy) and silent approval-only (muted violet) reviews, among non-bot pull requests on the estimation sample. Substantive review erodes while silent approvals hold, so the residual human review is increasingly bare ap…
Figure 11
Figure 11. Figure 11: The per-PR AI premium across the DORA cycle [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 13
Figure 13. Figure 13: Per-capita Claude Code token spend among active [PITH_FULL_IMAGE:figures/full_fig_p014_13.png]
Figure 12
Figure 12. Figure 12: The same bottleneck over time: 90th-percentile end [PITH_FULL_IMAGE:figures/full_fig_p014_12.png]
Figure 14
Figure 14. Figure 14: Pooled AI-adoption event study on monthly pull re [PITH_FULL_IMAGE:figures/full_fig_p015_14.png]
Figure 15
Figure 15. Figure 15: Accumulated use, not the model frontier. (A) Event studies around each release on [PITH_FULL_IMAGE:figures/full_fig_p016_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

59 extracted references · 59 canonical work pages

  1. [1]

    2025 Stack Overflow developer survey: AI,

    Stack Overflow, “2025 Stack Overflow developer survey: AI,” 2025, 84% of respondents use or plan to use AI tools in their development process. Accessed 2026-06-23. [Online]. Available: https://survey.stackoverflow.co/2025/ai

  2. [2]

    2025 DORA state of AI-assisted software development report,

    Google and DORA, “2025 DORA state of AI-assisted software development report,” 2025, reports AI adoption among software development professionals at 90%. Accessed 2026-06-23. [Online]. Available: https://dora.dev/research/2025/ 10

  3. [3]

    Shopify CEO says staffers need to prove jobs can’t be done by AI before asking for more headcount,

    A. Palmer, “Shopify CEO says staffers need to prove jobs can’t be done by AI before asking for more headcount,” 4 2025, cNBC; reproduces CEO Tobi L ¨utke’s internal memo making effective AI use “a fundamental expectation” of all employees. Accessed 2026-06-23. [Online]. Available: https://www.cnbc.com/2025/04/07/shopify-ceo-prove-ai-cant-d o-jobs-before-a...

  4. [4]

    Coinbase CEO urged engineers to use AI—then shocked them by firing those who wouldn’t,

    M. Quiroz-Gutierrez, “Coinbase CEO urged engineers to use AI—then shocked them by firing those who wouldn’t,” 8 2025, fortune; CEO Brian Armstrong mandated AI coding tools firmwide and dismissed engineers who did not adopt them. Accessed 2026-06-23. [Online]. Available: https://fortune.com/2025/08/25/coinbase-ceo-brian-armstrong-a i-coding-assistants-mand...

  5. [5]

    The Impact of AI on Developer Productivity: Evidence from GitHub Copilot

    S. Peng, E. Kalliamvakou, P. Cihon, and M. Demirer, “The impact of AI on developer productivity: Evidence from GitHub Copilot,”arXiv preprint arXiv:2302.06590, 2023

  6. [6]

    The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers,

    K. Z. Cui, M. Demirer, S. Jaffe, L. Musolff, S. Peng, and T. Salz, “The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers,”Man- agement Science, 2026

  7. [7]

    How much does AI impact development speed? An enterprise- based randomized controlled trial,

    E. Paradis, K. Grey, Q. Madison, D. Nam, A. Macvean, V . Meimand, N. Zhang, B. Ferrari-Church, and S. Chandra, “How much does AI impact development speed? An enterprise- based randomized controlled trial,” inInternational Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 2025, pp. 618–629

  8. [8]

    Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

    J. Becker, N. Rush, E. Barnes, and D. Rein, “Measuring the impact of early-2025 AI on experienced open-source developer productivity,”arXiv preprint arXiv:2507.09089, 2025

Show all 59 references
  1. [9]

    Developer productivity with and without GitHub Copi- lot: A longitudinal mixed-methods case study,

    V . Stray, E. G. Brandtzæg, V . T. Wivestad, A. Barbala, and N. B. Moe, “Developer productivity with and without GitHub Copi- lot: A longitudinal mixed-methods case study,”arXiv preprint arXiv:2509.20353, 2025

  2. [10]

    GitHub Copilot and developer productivity: An observational dose-response analysis,

    A. Heilman, A. Kyllo, and E. Murphy-Hill, “GitHub Copilot and developer productivity: An observational dose-response analysis,”arXiv preprint arXiv:2606.00438, 2026

  3. [11]

    ‘Maybe We Need Some More Examples:’ Individual and Team Drivers of Developer GenAI Tool Use,

    C. Miller, R. Choudhuri, M. Ulloa, S. Haniyur, R. DeLine, M.- A. Storey, E. Murphy-Hill, C. Bird, and J. L. Butler, “‘Maybe We Need Some More Examples:’ Individual and Team Drivers of Developer GenAI Tool Use,” inInternational Conference on Software Engineering (ICSE), 2026

  4. [12]

    Public memo announcing an organization-wide AI productivity mandate (a2×goal),

    Case-company CTO, “Public memo announcing an organization-wide AI productivity mandate (a2×goal),” 6 2025, documentary record; identifying details withheld

  5. [13]

    Guidelines for conducting and re- porting case study research in software engineering,

    P. Runeson and M. H ¨ost, “Guidelines for conducting and re- porting case study research in software engineering,”Empirical software engineering, vol. 14, no. 2, pp. 131–164, 2009

  6. [14]

    Revisiting event-study designs: robust and efficient estimation,

    K. Borusyak, X. Jaravel, and J. Spiess, “Revisiting event-study designs: robust and efficient estimation,”Review of Economic Studies, vol. 91, no. 6, pp. 3253–3285, 2024

  7. [15]

    What’s trending in difference-in-differences? A synthesis of the recent econometrics literature,

    J. Roth, P. H. Sant’Anna, A. Bilinski, and J. Poe, “What’s trending in difference-in-differences? A synthesis of the recent econometrics literature,”Journal of Econometrics, vol. 235, no. 2, pp. 2218–2244, 2023

  8. [16]

    Speed at the cost of quality: How Cursor AI increases short-term velocity and long-term complexity in open-source projects,

    H. He, C. Miller, S. Agarwal, C. K ¨astner, and B. Vasilescu, “Speed at the cost of quality: How Cursor AI increases short-term velocity and long-term complexity in open-source projects,” inInternational Conference on Mining Software Repositories (MSR), 2026

  9. [17]

    The SPACE of developer produc- tivity,

    N. Forsgren, M.-A. Storey, C. Maddila, T. Zimmermann, B. Houck, and J. Butler, “The SPACE of developer produc- tivity,”Communications of the ACM, vol. 64, no. 6, pp. 46–53, 2021

  10. [18]

    Is GitHub Copilot a substitute for human pair- programming? An empirical study,

    S. Imai, “Is GitHub Copilot a substitute for human pair- programming? An empirical study,” inInternational Confer- ence on Software Engineering (ICSE): Companion Proceedings, 2022, pp. 319–321

  11. [19]

    Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,

    P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,” inCHI Conference on Human Factors in Computing Systems (CHI) Extended Abstracts, 2022, pp. 1–7

  12. [20]

    The impact of generative AI on collaborative open-source software development: Evidence from GitHub Copilot,

    F. Song, A. Agarwal, and W. Wen, “The impact of generative AI on collaborative open-source software development: Evidence from GitHub Copilot,”arXiv preprint arXiv:2410.02091, 2024

  13. [21]

    Generative AI and the nature of work,

    M. Hoffmann, S. Boysel, F. Nagle, S. Peng, and K. Xu, “Generative AI and the nature of work,”Harvard Business School Strategy Unit Working Paper, no. 25-021, pp. 25–021, 2025

  14. [22]

    Who is using AI to code? Global diffusion and impact of generative AI,

    S. Daniotti, J. Wachs, X. Feng, and F. Neffke, “Who is using AI to code? Global diffusion and impact of generative AI,”Science, p. eadz9311, 2026

  15. [23]

    AI raises the productivity bar,

    L. Wu and B. Vasilescu, “AI raises the productivity bar,” Science, vol. 391, no. 6787, pp. 763–764, 2026

  16. [24]

    AI IDEs or autonomous agents? Measuring the Impact of Coding Agents on Software Development,

    S. Agarwal, H. He, and B. Vasilescu, “AI IDEs or autonomous agents? Measuring the Impact of Coding Agents on Software Development,” inInternational Conference on Mining Software Repositories (MSR) – Mining Challenge, 2026

  17. [25]

    AI agents and higher-order work,

    S. K. Sarkar, “AI agents and higher-order work,”SSRN Elec- tronic Journal, vol. 4, 2026

  18. [26]

    Code with me or for me? How increasing AI automation transforms developer workflows,

    V . Chen, A. Talwalkar, R. Brennan, and G. Neubig, “Code with me or for me? How increasing AI automation transforms developer workflows,” inCHI Conference on Human Factors in Computing Systems (CHI), 2026, pp. 1–19

  19. [27]

    Debt behind the AI boom: A large-scale empirical study of AI- generated code in the wild,

    Y . Liu, R. Widyasari, Y . Zhao, I. C. Irsan, J. Chen, and D. Lo, “Debt behind the AI boom: A large-scale empirical study of AI- generated code in the wild,”arXiv preprint arXiv:2603.28592, 2026

  20. [28]

    More Code, Less Understanding? On the Impact of AI Assis- tants on Developers’ Productivity and Code Ownership,

    A. Martin-Lopez, R. Tufano, E. Guglielmi, A. B. S ´anchez, A. L. Sanz, S. Scalabrino, R. Oliveto, S. Segura, and G. Bavota, “More Code, Less Understanding? On the Impact of AI Assis- tants on Developers’ Productivity and Code Ownership,”IEEE Transactions on Software Engineering, 2026

  21. [29]

    From technical debt to cognitive and intent debt: Rethinking software health in the age of AI,

    M.-A. Storey, “From technical debt to cognitive and intent debt: Rethinking software health in the age of AI,”arXiv preprint arXiv:2603.22106, 2026

  22. [30]

    Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time,

    R. M. Popescu, D. Gros, A. Botocan, R. Pandita, P. Devanbu, and M. Izadi, “Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time,” inInternational Conference on Mining Software Repositories (MSR), 2026

  23. [31]

    AI-assisted Programming May Decrease the Pro- ductivity of Experienced Developers by Increasing Maintenance Burden,

    F. Xu, P. K. Medappa, M. M. Tunc, M. Vroegindeweij, and J. C. Fransoo, “AI-assisted Programming May Decrease the Pro- ductivity of Experienced Developers by Increasing Maintenance Burden,”arXiv preprint arXiv:2510.10165, 2025

  24. [32]

    Echoes of AI: Investigating the downstream effects of AI assistants on software maintain- ability,

    M. Borg, D. Hewett, N. Hagatulah, N. Couderc, E. S ¨oderberg, D. Graham, U. Kini, and D. Farley, “Echoes of AI: Investigating the downstream effects of AI assistants on software maintain- ability,”Empirical Software Engineering, vol. 31, no. 6, p. 161, 2026

  25. [33]

    Expectations, outcomes, and chal- lenges of modern code review,

    A. Bacchelli and C. Bird, “Expectations, outcomes, and chal- lenges of modern code review,” inInternational Conference on Software Engineering (ICSE). IEEE, 2013, pp. 712–721

  26. [34]

    Code reviews do not find bugs. How the current code review best practice slows us down,

    J. Czerwonka, M. Greiler, and J. Tilford, “Code reviews do not find bugs. How the current code review best practice slows us down,” inInternational Conference on Software Engineering (ICSE), vol. 2. IEEE, 2015, pp. 27–28

  27. [35]

    A large-scale survey on the usability of AI programming assistants: Successes and chal- lenges,

    J. T. Liang, C. Yang, and B. A. Myers, “A large-scale survey on the usability of AI programming assistants: Successes and chal- lenges,” inInternational Conference on Software Engineering (ICSE), 2024, pp. 1–13

  28. [36]

    The impact of code review coverage and code review participa- 11 tion on software quality: A case study of the Qt, VTK, and ITK projects,

    S. McIntosh, Y . Kamei, B. Adams, and A. E. Hassan, “The impact of code review coverage and code review participa- 11 tion on software quality: A case study of the Qt, VTK, and ITK projects,” inInternational Conference on Mining Software Repositories (MSR), 2014, pp. 192–201

  29. [37]

    When code authors are agents: A large-scale study of human–agent collaboration in pull requests

    A. O. Njoku, Z. Sharafi, and F. Khomh, “When code authors are agents: A large-scale study of human–agent collaboration in pull requests.”

  30. [38]

    Collaborator or assistant? How AI coding agents partition work across pull request lifecycles,

    Y . Jo, S. Hassanet al., “Collaborator or assistant? How AI coding agents partition work across pull request lifecycles,” arXiv preprint arXiv:2605.08017, 2026

  31. [39]

    Comparing AI coding agents: A task-stratified analysis of pull request acceptance,

    G. Pinna, J. Gong, D. Williams, and F. Sarro, “Comparing AI coding agents: A task-stratified analysis of pull request acceptance,”arXiv preprint arXiv:2602.08915, 2026

  32. [40]

    On the use of agentic coding: An empirical study of pull requests on GitHub,

    M. Watanabe, H. Li, Y . Kashiwa, B. Reid, H. Iida, and A. E. Hassan, “On the use of agentic coding: An empirical study of pull requests on GitHub,”ACM Transactions on Software Engineering and Methodology, 2025

  33. [41]

    Coding agents in the wild: Failure modes and rejection patterns of ai-generated pull requests,

    M. Hindi, Y . Mahmood, L. Mohammed, S. Bouktif, and M. Me- diani, “Coding agents in the wild: Failure modes and rejection patterns of ai-generated pull requests,”IEEE Access, 2026

  34. [42]

    Where do AI coding agents fail? An Empirical Study of Failed Agentic Pull Requests in GitHub,

    R. Ehsani, S. Pathak, S. Rawal, A. A. Mujahid, M. M. Im- ran, and P. Chatterjee, “Where do AI coding agents fail? An Empirical Study of Failed Agentic Pull Requests in GitHub,” inInternational Conference on Mining Software Repositories (MSR) – Mining Challenge, 2026

  35. [43]

    These Aren’t the Reviews You’re Looking For: How Humans Review AI-Generated Pull Requests,

    K. Duma, P. Wr ´oblewski, J. Bobi´nska, J. Winiarska, and P. Przy- mus, “These Aren’t the Reviews You’re Looking For: How Humans Review AI-Generated Pull Requests,”arXiv preprint arXiv:2605.02273, 2026

  36. [44]

    Human-AI Synergy in Agentic Code Review,

    S. Zhong, S. Noei, Y . Zou, and B. Adams, “Human-AI Synergy in Agentic Code Review,”arXiv preprint arXiv:2603.15911, 2026

  37. [45]

    The productivity paradox of information tech- nology,

    E. Brynjolfsson, “The productivity paradox of information tech- nology,”Communications of the ACM, vol. 36, no. 12, pp. 66– 77, 1993

  38. [46]

    Informa- tion technology, workplace organization, and the demand for skilled labor: Firm-level evidence,

    T. F. Bresnahan, E. Brynjolfsson, and L. M. Hitt, “Informa- tion technology, workplace organization, and the demand for skilled labor: Firm-level evidence,”The Quarterly Journal of Economics, vol. 117, no. 1, pp. 339–376, 2002

  39. [47]

    Why are there still so many jobs? The history and future of workplace automation,

    D. H. Autor, “Why are there still so many jobs? The history and future of workplace automation,”Journal of Economic Perspectives, vol. 29, no. 3, pp. 3–30, 2015

  40. [48]

    Experimental evidence on the produc- tivity effects of generative artificial intelligence,

    S. Noy and W. Zhang, “Experimental evidence on the produc- tivity effects of generative artificial intelligence,”Science, vol. 381, no. 6654, pp. 187–192, 2023

  41. [49]

    Generative ai at work,

    E. Brynjolfsson, D. Li, and L. Raymond, “Generative ai at work,”The Quarterly Journal of Economics, vol. 140, no. 2, pp. 889–942, 2025

  42. [50]

    Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality,

    F. Dell’Acqua, E. McFowland III, E. R. Mollick, H. Lifshitz- Assaf, K. Kellogg, S. Rajendran, L. Krayer, F. Candelon, and K. R. Lakhani, “Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality...

  43. [51]

    From prompting to verification: How experience shapes vibe coding practices,

    A. Fawzy, A. Tahir, and K. Blincoe, “From prompting to verification: How experience shapes vibe coding practices,” arXiv preprint arXiv:2605.24521, 2026

  44. [52]

    Keeping Humans in the Driver’s Seat: A Conceptual Framework for GenAI and Competence Debt,

    C. M. L ¨uders, O. Kru ˙zycki, and C. Brandt, “Keeping Humans in the Driver’s Seat: A Conceptual Framework for GenAI and Competence Debt,” inInternational Conference on the F ounda- tions of Software Engineering (FSE), Companion Proceedings. ACM, 2026

  45. [53]

    Forsgren, J

    N. Forsgren, J. Humble, and G. Kim,Accelerate: The science of lean software and DevOps: Building and scaling high per- forming technology organizations. IT Revolution, 2018

  46. [54]

    Detecting and characterizing bots that commit code,

    T. Dey, S. Mousavi, E. Ponce, T. Fry, B. Vasilescu, A. Filip- pova, and A. Mockus, “Detecting and characterizing bots that commit code,” inInternational Conference on Mining Software Repositories (MSR), 2020, pp. 209–219

  47. [55]

    A ground-truth dataset and classification model for detecting bots in GitHub issue and PR comments,

    M. Golzadeh, A. Decan, D. Legay, and T. Mens, “A ground-truth dataset and classification model for detecting bots in GitHub issue and PR comments,”Journal of Systems and Software, vol. 175, p. 110911, 2021

  48. [56]

    Automating dependency updates in practice: An exploratory study on GitHub Depend- abot,

    R. He, H. He, Y . Zhang, and M. Zhou, “Automating dependency updates in practice: An exploratory study on GitHub Depend- abot,”IEEE Transactions on Software Engineering, vol. 49, no. 8, pp. 4004–4022, 2023

  49. [57]

    Mo- tivations, challenges, best practices, and benefits for bots and conversational agents in software engineering: A multivocal literature review,

    S. Lambiase, G. Catolino, F. Palomba, and F. Ferrucci, “Mo- tivations, challenges, best practices, and benefits for bots and conversational agents in software engineering: A multivocal literature review,”ACM Computing Surveys, vol. 57, no. 4, pp. 1–37, 2024

  50. [58]

    Difference-in-differences with multiple time periods,

    B. Callaway and P. H. Sant’Anna, “Difference-in-differences with multiple time periods,”Journal of Econometrics, vol. 225, no. 2, pp. 200–230, 2021

  51. [59]

    Modern code review: a case study at Google,

    C. Sadowski, E. S ¨oderberg, L. Church, M. Sipko, and A. Bac- chelli, “Modern code review: a case study at Google,” in International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), 2018, pp. 181–190. 12 APPENDIX Supplementary Materials AI Writ...

Pith tools

Reviewed July 3, 2026 · model on record in the stance chip above.