Pith. sign in

REVIEW 4 major objections 7 minor 50 references

Linking real AI coding chats to open-source histories shows vibe coding is heaviest in small, less collaborative repos, usually followed by commits, without broad code-quality decline.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 03:55 UTC pith:BFNBDN2O

load-bearing objection Solid large-scale map from SpecStory IDE chats to OSS outcomes; RQ2 before/after claims rest on a fragile adoption cutoff and self-selected sample, but the descriptive linkage and survey still earn referee time. the 4 major comments →

arxiv 2607.05677 v1 pith:BFNBDN2O submitted 2026-07-06 cs.SE cs.HC

From Conversation to Contribution: Characterizing Coding Agent in Open-Source Software

classification cs.SE cs.HC
keywords AI coding assistantsopen source softwarevibe codingdeveloper-AI chatOSS collaborationcode qualitypull requestsAI-assisted programming
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper sets out to show how developers' natural-language chats with AI coding assistants connect to later open-source work and collaboration, not just to isolated productivity metrics. The authors gathered thousands of real chat sessions from public repositories, tied each session to subsequent commits, issues, and pull requests, and surveyed developers whose chats appear in the data. They report that AI-related activity is heavier in smaller, younger, less collaborative projects; that Code Writing is the main chat purpose; and that nearly every session is followed by a commit. After first observed AI adoption, projects tend to gain active contributors and reduce committer concentration, while communication stays concentrated and the paper's quality and merge-rate signals do not broadly worsen. Developers say AI lowers contribution barriers but rate others' AI-generated code as more maintenance-heavy than their own, and many who would share chats still fear looking incompetent, burdening reviewers, or exposing ideas. A sympathetic reader cares because tool design and OSS governance need this process-level map if conversational AI is to be supported without eroding review capacity or trust.

Core claim

Across 13,360 AI conversation sessions linked to 1,356 open-source repositories plus a developer survey, AI-related commits are more prominent in smaller, less mature, and less collaborative repositories; after first observed AI adoption, active contributors rise and contributor concentration falls while communication remains highly concentrated; Code Writing dominates chat purpose and nearly all sessions are followed by commits; and the authors find no broad deterioration in their code-quality or pull-request merge signals, even as respondents perceive others' AI-generated code as harder to maintain than their own.

What carries the argument

The linkage of SpecStory-preserved IDE chat sessions to full repository histories, with the earliest observed chat defining an AI-adoption cutoff for before/after interrupted time-series and paired repository comparisons.

Load-bearing premise

The first SpecStory chat found in a public repository is treated as a valid AI-adoption moment, and repositories that commit those chat logs are treated as representative of AI-assisted open-source contribution.

What would settle it

A large matched sample of similar public OSS projects that use the same AI tools but never commit SpecStory histories, with independently dated adoption, shows no post-adoption rise in active contributors, no drop in contributor concentration, and clear worsening of quality or merge signals.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • OSS platforms can surface lightweight AI-purpose summaries or tags as review context without requiring full chat transcripts.
  • Maintainers can calibrate review effort for AI-assisted changes by risk and scope rather than treating every AI-touched commit the same.
  • Contribution guidelines can treat AI as a barrier-lowering aid for first contributions while setting clearer disclosure and responsibility norms.
  • Tool design can prioritize making private chat reasoning reviewable enough that design and debugging context does not disappear from the public project record.
  • Hosting platforms should plan capacity and project-discovery features around many small, AI-heavy repositories rather than only large collaborative ones.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If problem-solving migrates into private AI chats, OSS collaboration may shift from shared issue discussion toward post-hoc review of finished diffs.
  • The survey asymmetry between own and others' AI code suggests communities will resist high volumes of external AI-assisted PRs until responsibility and review expectations are explicit.
  • The early-heavy then declining AI-commit share may mean vibe coding often acts as a bootstrap accelerator rather than a permanent co-maintainer role.
  • Repositories that never commit chat logs may hide different risk profiles; any governance rule based only on visible SpecStory users will miss that population.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This paper links 13,360 SpecStory-preserved AI coding-assistant chat sessions (79,172 user messages) from 1,356 public OSS repositories to full GitHub development histories, and complements the mining with a 25-response developer survey. It reports that AI-related commit share is higher in smaller, less mature, and less collaborative repositories; that after first observed AI adoption, active contributors rise and contributor concentration falls (p<.001) while communication remains concentrated; that Code Writing dominates chat purpose and nearly all sessions are followed by commits; that observable code-quality and PR-merge signals show no broad deterioration; and that developers rate others’ AI-generated code as more maintenance-burdening than their own (p=.029) while viewing AI as lowering contribution barriers. The authors frame results as associations and discuss design/governance implications for vibe-coding in OSS.

Significance. If the characterizations hold under the stated selection constraints, the paper makes a timely and useful contribution: it is among the first large-scale studies to connect IDE-based multi-turn AI chat (not only agent-authored PRs or tool-adoption events) to subsequent OSS commits, issues, PRs, and collaboration metrics. Strengths include a non-trivial linked dataset, reuse of validated chat-purpose and project-type taxonomies with reported F1/accuracy, careful MSR hygiene (bot filtering, path-classifier manual checks, repository fixed-effect ITS with age controls, FDR within test families, and a mature validation cohort), and explicit association rather than causal language in the methods. The mixed-methods triangulation and concrete governance discussion (disclosure, review calibration, platform metadata) are practically valuable for the SE community even if external validity is limited to SpecStory-committing public repos.

major comments (4)
  1. §III-B1 and RQ2 sample construction: The central before/after claims (more active contributors, lower contributor concentration p<.001; issue/PR compositional shifts; ‘no broad quality deterioration’) rest on defining first observed SpecStory chat as the AI-adoption cutoff and restricting RQ2 to the 608 repos whose first chat is after GitHub creation. This is load-bearing and fragile: 632/1240 repos already have pre-publication chats, committing .specstory/history/ is itself a strong selection filter, and ~90% of the sample was created in 2025 or later (Threats §V-1). The n=114 validation cohort is still drawn from the same SpecStory-committing population, so it cannot separate ‘AI adoption effects’ from ‘dynamics of self-selected chat-log committers in young/solo repos.’ Please add sensitivity analyses (e.g., alternative cutoffs, intensity-based rather than first-chat adoption, comparis
  2. §IV-B2 / Table II, defect-related activity: The paper reports a significant rise in bug/fix commit share (9.6%→16.9%, p<.001) while arguing there is ‘no broad code-quality deterioration’ because raw bug/fix counts do not rise and ITS trends are non-significant, attributing the share increase to falling total commit volume (§IV-B1). That interpretation is plausible but currently under-supported as a headline claim: share-based quality signals move adversely while volume falls, and keyword-based bug/fix labeling is a coarse proxy. Either (i) strengthen the quality claim with additional independent signals (e.g., post-merge churn, review rework, static-analysis if available) or (ii) soften abstract/conclusion language to ‘no clear sustained increase in defect activity under our proxies, with share metrics confounded by volume decline,’ and report the share increase more prominently as a cav
  3. §III-A4 and §IV-C (RQ3): Survey claims that appear in the abstract—others’ AI code harder to maintain (p=.029), 68% willing to share chat, AI lowers barriers—are based on n=25 responses (4.2% of delivered invitations). The paper correctly notes consistency with prior SE survey rates and triangulates with mining, but the sample is too small and self-selected (public-email SpecStory users who opted in) to support abstract-level perception generalizations. Please demote survey findings in the abstract to more cautious wording, report exact item n and confidence intervals, and treat RQ3 as exploratory qualitative/descriptive support rather than confirmatory evidence of community-wide attitudes.
  4. §IV-B4 collaboration results: Main-cohort active contributors rise and concentration falls (p<.001), but the validation cohort shows the opposite direction for active-contributor share (87.9%→74.1%, p=.002). The paper notes this inconsistency briefly, yet the abstract still states that ‘after AI adoption, projects tended to show more active contributors and lower contributor concentration (p<.001)’ without the cohort caveat. Given that collaboration broadening is a central claim, the abstract and discussion should present the main/validation disagreement as a primary result limitation, not a parenthetical, and avoid implying a generalizable participation-broadening effect.
minor comments (7)
  1. Title inconsistency: arXiv id/header uses ‘Characterizing Coding Agent’ while the manuscript title uses ‘Characterizing Vibe Coding’; align title, abstract, and running heads.
  2. §III-A1 / data window: Collection is said to span September 2024 to March 2026 and references include 2026 venues; ensure all dates, arXiv citations, and ‘accessed’ notes are consistent for camera-ready and that future-dated industry links remain stable or are archived.
  3. Figure 3 caption and §IV-B1: Clarify that left-side cohort overlap is driven by young-repo composition so readers do not misread the pre-adoption trend as identical populations.
  4. Table I: Source vs total churn means are extremely right-skewed (e.g., Code Writing total mean 97,755 vs median 1,224); consider logging or adding IQR columns so readers do not over-interpret means.
  5. §III-B2 HHI definition is clear, but briefly state the [0,1] range and interpretation thresholds used when calling communication ‘highly concentrated’ in results.
  6. Replication package is mentioned but not linked in the provided text; add a stable URL/DOI and list which scripts reproduce ITS, FDR families, and chat labeling.
  7. Minor prose: ‘vibe-coding’ is sometimes hyphenated inconsistently; standardize terminology on first definition in §I.

Circularity Check

1 steps flagged

Observational OSS mining study: main before/after and correlation claims are independent measurements, not inputs renamed as predictions; only a minor near-tautology in the session-to-commit association given SpecStory sample construction.

specific steps
  1. self definitional [§III-B2 RQ1-1/RQ1-2; §IV-A2 Results (“Subsequent activity by chat purpose”)]
    "For each chat-history file identified through GitHub Code Search, we used the commit or ref in the GitHub blob URL to identify the repository state in which the artifact was observed. When the ref matched a commit in our collected history, we treated that commit as the artifact-containing commit for the session. ... We next examined development activity in the commit associated with each labeled chat session, referred to as subsequent activity. Nearly all sessions (98.9%) had subsequent activity"

    Sessions are sampled as SpecStory Markdown files already present in public Git history. Associating each session with that artifact-containing commit, then reporting that nearly all sessions have “subsequent activity,” largely restates sample construction: committed chat artifacts have commits by definition (failures only when the ref is missing from collected history). The substantive non-circular residue is the 96.1% mixed development-file co-change rate, not the 98.9% existence claim.

full rationale

This paper is a mixed-methods empirical characterization (SpecStory chat logs linked to GitHub histories plus a small survey), not a first-principles derivation. Core outcomes—AI-related commit share vs. size/contributors, ITS trends in commits/quality/CI, contributor HHI, PR/issue composition, and survey Likert comparisons—are measured quantities timed relative to a defined adoption cutoff, not algebraic restatements of fitted free parameters. Reuse of the chat-purpose taxonomy and project-type labels from prior overlapping-author work ([32], [34]) is standard labeling infrastructure with reported F1/accuracy against human checks; it does not force the repository-level dynamics or survey results. The only mild self-definitional flavor is the claim that nearly all chat sessions have “subsequent activity,” when sessions enter the sample as committed SpecStory artifacts and “subsequent activity” is operationalized via the associated artifact-containing commit—so existence of a linked commit is largely guaranteed by construction, with the non-tautological content being co-change of development files. That does not underwrite the paper’s central post-adoption or intensity claims. Score 1 reflects that minor descriptive circularity only; no fitted-input-as-prediction, uniqueness-from-self-citation, or load-bearing self-citation chain is present.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 1 invented entities

Load-bearing content is empirical measurement under domain assumptions about what SpecStory logs, GitHub timestamps, keyword bug labels, and a self-selected public sample mean. No new physical entities; free parameters are analysis cutoffs and model choices rather than fitted universal constants. The central characterization stands or falls on sample representativeness and adoption timing more than on invented theory.

free parameters (4)
  • Trivial-repository screening rules (excluded 47/1287)
    Manual/heuristic exclusion of toy/homework/template projects changes the analysis population; criteria follow prior work but remain analyst choices.
  • Validation cohort filter (n=114, ~5 contributors, ~3.5 years)
    Thresholds defining the 'more mature' cohort are chosen for robustness checks and affect which post-adoption patterns are called consistent.
  • Adoption month exclusion from pre/post comparisons
    Relative month 0 is dropped by design to avoid transition effects; this choice shapes ITS and rate comparisons.
  • Bug/fix and test path keyword lists
    Defect and test-touching measures depend on chosen message/path terms; misclassification shifts quality-signal conclusions.
axioms (5)
  • domain assumption Committed SpecStory Markdown under .specstory/history/ faithfully records developer–AI coding-assistant sessions for Copilot/Cursor/Claude Code.
    Data collection §III-A1 treats these artifacts as the AI conversation sample; incomplete logging or selective commit would bias intensity and purpose.
  • ad hoc to paper Earliest mapped AI-chat timestamp is a usable definition of first observed AI adoption for repository dynamics.
    §III-B1 defines adoption this way even when local use may predate GitHub or other tools leave no SpecStory trail.
  • domain assumption Keyword/path heuristics for bug/fix commits, bug issues, and test files are adequate proxies for defect and testing activity.
    RQ2-2 quality analysis follows prior MSR practice; proxy error can create false stability or false change.
  • domain assumption Paired pre/post and ITS associations after adoption can be interpreted as characterizations of AI-related dynamics without claiming full causality.
    Threats section states associations not causes; still, abstract/results language about 'after AI adoption' invites causal reading.
  • domain assumption LLM multi-label chat-purpose classifier (macro-F1 0.83) and project-type labels (90.5% accuracy on sample) are accurate enough for distributional claims.
    §III-A3; residual label noise especially affects purpose–commit association tests.
invented entities (1)
  • vibe-coding workflow (as operational study object) independent evidence
    purpose: Name the intent-driven multi-turn AI coding practice the paper characterizes in OSS.
    Term is taken from emerging usage rather than a new physical entity; operationalized via SpecStory chats plus repo activity. Not independently validated beyond this observational framing.

pith-pipeline@v1.1.0-grok45 · 23335 in / 3495 out tokens · 39888 ms · 2026-07-11T03:55:41.683808+00:00 · methodology

0 comments
read the original abstract

AI coding assistants such as GitHub Copilot and Cursor have evolved from code-suggestion tools into conversational collaborators, enabling vibe-coding workflows in which developers guide AI-generated code through natural-language dialogue. Although researchers have increasingly recognized the importance of AI coding agents and begun examining their impact on open-source development, a comprehensive understanding of how developers' chat-based interactions with AI relate to subsequent open-source development and collaboration remains limited. This hinders efforts to effectively design, evaluate, and govern AI-assisted open-source software development. To address this gap, we collected 13,360 AI conversation sessions comprising 79,172 user messages from 1,356 OSS repositories, linked them to repository development histories, and complemented this analysis with a targeted developer survey. We find heavier AI use in smaller, less mature, and less collaborative repositories. After AI adoption, projects tended to show more active contributors and lower contributor concentration (p < .001), although communication remained highly concentrated. Code Writing was the dominant chat purpose, and nearly all AI chat sessions were followed by subsequent commits. We find no broad deterioration in code-quality signals or pull request merging rates. However, developers perceive others' AI-generated code as harder to maintain than their own (p = .029) and view AI as lowering barriers to OSS contribution. While most developers (68%) are willing to share their chat, concerns remain around appearing incompetent, increasing reviewer burden, and exposing ideas to competitors. These findings provide a large-scale empirical characterization of AI-assisted OSS contribution and offer practical insights for designing and governing responsible vibe-coding practices in open-source development.

Figures

Figures reproduced from arXiv: 2607.05677 by Collin McMillan, Ningzhi Tang, Toby Jia-Jun Li, Yueke Zhang, Yu Huang, Zihan Fang.

Figure 1
Figure 1. Figure 1: Data collection pipeline insight into how OSS developers who actively use these tools perceive their effects on participation and project development. Our work focuses on these gaps. III. METHODOLOGY We collected 13,360 AI chat sessions from 1,356 OSS repositories and complemented them with each repository’s full development history. Using these data, we investigated the following research questions: • RQ1… view at source ↗
Figure 2
Figure 2. Figure 2: Chronological flow from AI-chat purpose to the file-type composition [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

50 extracted references · 50 canonical work pages · 7 internal anchors

  1. [1]

    Code with me or for me? how increasing ai automation transforms developer workflows,

    V . Chen, A. Talwalkar, R. Brennan, and G. Neubig, “Code with me or for me? how increasing ai automation transforms developer workflows,” inProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026, pp. 1–19

  2. [2]

    Accelerating software development with ai: exploring the impact of chatgpt and github copilot

    I. Solohubov, A. Moroz, M. Y . Tiahunova, H. H. Kyrychek, and S. Skrupsky, “Accelerating software development with ai: exploring the impact of chatgpt and github copilot.” inCTE, 2023, pp. 76–86

  3. [3]

    Programmers who use screen readers in the vibe coding era: Adaptation, empower- ment, and new accessibility landscape,

    N. Chen, L. K. Qiu, A. Z. Wang, Z. Wang, and Y . Yang, “Programmers who use screen readers in the vibe coding era: Adaptation, empower- ment, and new accessibility landscape,” inProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026, pp. 1–20

  4. [4]

    A review on vibe coding: Fundamentals, state-of-the-art, challenges and future directions,

    P. P. Ray, “A review on vibe coding: Fundamentals, state-of-the-art, challenges and future directions,”Authorea Preprints, 2025

  5. [5]

    The promises and perils of mining github,

    E. Kalliamvakou, G. Gousios, K. Blincoe, L. Singer, D. M. German, and D. Damian, “The promises and perils of mining github,” inProceedings of the 11th working conference on mining software repositories, 2014, pp. 92–101

  6. [6]

    Open source ai-based se tools: Opportunities and challenges of collaborative software learning,

    Z. Lin, W. Ma, T. Lin, Y . Zheng, J. Ge, J. Wang, J. Klein, T. F. Bissyand´e, Y . Liu, and L. Li, “Open source ai-based se tools: Opportunities and challenges of collaborative software learning,”ACM Transactions on Software Engineering and Methodology, vol. 34, no. 5, pp. 1–24, 2025

  7. [7]

    The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering

    H. Li, H. Zhang, and A. E. Hassan, “The rise of ai teammates in software engineering (se) 3.0: How autonomous coding agents are reshaping software engineering,”arXiv preprint arXiv:2507.15003, 2025

  8. [8]

    Microsoft turns to amazon for help with github’s ai-driven capacity issues,

    A. Stewart, “Microsoft turns to amazon for help with github’s ai-driven capacity issues,” Business Insider, Jun. 2026, accessed: 2026-06-22. [Online]. Available: https://www.businessinsider.com/ microsoft-github-amazon-ai-cloud-capacity-2026-6

  9. [9]

    The Impact of AI on Developer Productivity: Evidence from GitHub Copilot

    S. Peng, E. Kalliamvakou, P. Cihon, and M. Demirer, “The impact of ai on developer productivity: Evidence from github copilot,”arXiv preprint arXiv:2302.06590, 2023

  10. [10]

    Ai-assisted programming may decrease the productivity of ex- perienced developers by increasing maintenance burden,

    F. Xu, P. K. Medappa, M. M. Tunc, M. Vroegindeweij, and J. C. Fransoo, “Ai-assisted programming may decrease the productivity of ex- perienced developers by increasing maintenance burden,”arXiv preprint arXiv:2510.10165, 2025

  11. [11]

    Speed at the cost of quality: How cursor ai increases short-term velocity and long-term complexity in open-source projects,

    H. He, C. Miller, S. Agarwal, C. K ¨astner, and B. Vasilescu, “Speed at the cost of quality: How cursor ai increases short-term velocity and long-term complexity in open-source projects,”arXiv preprint arXiv:2511.04427, 2025

  12. [12]

    Ai ides or autonomous agents? measuring the impact of coding agents on software development,

    S. Agarwal, H. He, and B. Vasilescu, “Ai ides or autonomous agents? measuring the impact of coding agents on software development,”arXiv preprint arXiv:2601.13597, 2026

  13. [13]

    Investigating autonomous agent contributions in the wild: Activity patterns and code change over time,

    R. M. Popescu, D. Gros, A. Botocan, R. Pandita, P. Devanbu, and M. Izadi, “Investigating autonomous agent contributions in the wild: Activity patterns and code change over time,”arXiv preprint arXiv:2604.00917, 2026

  14. [14]

    Agentic Much? Adoption of Coding Agents on GitHub

    R. Robbes, T. Matricon, T. Degueule, A. Hora, and S. Zacchiroli, “Agentic much? adoption of coding agents on github,”arXiv preprint arXiv:2601.18341, 2026

  15. [15]

    When ai teammates meet code review: Collaboration signals shaping the integration of agent-authored pull requests,

    C. Nachuma and M. Zibran, “When ai teammates meet code review: Collaboration signals shaping the integration of agent-authored pull requests,”arXiv preprint arXiv:2602.19441, 2026

  16. [16]

    Large language models for software engineering: Sur- vey and open problems,

    A. Fan, B. Gokkaya, M. Harman, M. Lyubarskiy, S. Sengupta, S. Yoo, and J. M. Zhang, “Large language models for software engineering: Sur- vey and open problems,” in2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE). IEEE, 2023, pp. 31–53

  17. [17]

    A survey on large language models for code generation,

    J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim, “A survey on large language models for code generation,”ACM Transactions on Software Engineering and Methodology, vol. 35, no. 2, pp. 1–72, 2026

  18. [18]

    Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey

    J. Wang, Z. Zhang, Y . He, Z. Zhang, X. Song, Y . Song, T. Shi, Y . Li, H. Xu, K. Wuet al., “Enhancing code llms with reinforcement learning in code generation: A survey,”arXiv preprint arXiv:2412.20367, 2024

  19. [19]

    DPO-F+: Aligning Code Repair Feedback with Developers' Preferences

    Z. Fang, Y . Zhang, Y . Zhang, K. Leach, and Y . Huang, “Dpo-f+: Aligning code repair feedback with developers’ preferences,”arXiv preprint arXiv:2511.01043, 2025

  20. [20]

    Codeact-r: A cognitive simulation framework for human attention in code reading,

    Y . Zhang, Z. Fang, G. Trafton, D. Levin, K. Leach, and Y . Huang, “Codeact-r: A cognitive simulation framework for human attention in code reading,” in2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2025, pp. 3871–3875

  21. [21]

    Large language model-based agents for software engineering: A sur- vey,

    J. Liu, K. Wang, Y . Chen, X. Peng, Z. Chen, L. Zhang, and Y . Lou, “Large language model-based agents for software engineering: A sur- vey,”ACM Transactions on Software Engineering and Methodology, 2024

  22. [22]

    Enhancing software development practices with ai insights in high-tech companies,

    D. Ajiga, P. A. Okeleke, S. O. Folorunsho, and C. Ezeigweneme, “Enhancing software development practices with ai insights in high-tech companies,”Computer Science & IT Research Journal, vol. 5, no. 8, pp. 1897–1919, 2024

  23. [23]

    Enhancing software engineering with ai: Innovations, chal- lenges, and future directions,

    T. Abbas, S. A. Rathore, A. Turki, S. Khan, O. Alghushairy, and A. Daud, “Enhancing software engineering with ai: Innovations, chal- lenges, and future directions,”IET Software, vol. 2025, no. 1, p. 5691460, 2025

  24. [24]

    Measuring github copilot’s impact on productivity,

    A. Ziegler, E. Kalliamvakou, X. A. Li, A. Rice, D. Rifkin, S. Simister, G. Sittampalam, and E. Aftandilian, “Measuring github copilot’s impact on productivity,”Communications of the ACM, vol. 67, no. 3, pp. 54–63, 2024

  25. [25]

    An empirical evaluation of github copilot’s code suggestions,

    N. Nguyen and S. Nadi, “An empirical evaluation of github copilot’s code suggestions,” inProceedings of the 19th International Conference on Mining Software Repositories, 2022, pp. 1–5

  26. [26]

    Github copilot ai pair programmer: Asset or liability?

    A. M. Dakhel, V . Majdinasab, A. Nikanjam, F. Khomh, M. C. Desmarais, and Z. M. J. Jiang, “Github copilot ai pair programmer: Asset or liability?”Journal of Systems and Software, vol. 203, p. 111734, 2023

  27. [27]

    Asleep at the keyboard? assessing the security of github copilot’s code con- tributions,

    H. Pearce, B. Ahmad, B. Tan, B. Dolan-Gavitt, and R. Karri, “Asleep at the keyboard? assessing the security of github copilot’s code con- tributions,”Communications of the ACM, vol. 68, no. 2, pp. 96–105, 2025

  28. [28]

    Expectation vs. experi- ence: Evaluating the usability of code generation tools powered by large language models,

    P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experi- ence: Evaluating the usability of code generation tools powered by large language models,” inChi conference on human factors in computing systems extended abstracts, 2022, pp. 1–7

  29. [29]

    How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

    N. Tang, C. Chen, G. Xu, Y . Shi, Y . Huang, C. McMillan, T. Dong, and T. J.-J. Li, “How coding agents fail their users: A large-scale analysis of developer-agent misalignment in 20,574 real-world sessions,”arXiv preprint arXiv:2605.29442, 2026

  30. [30]

    The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot

    F. Song, A. Agarwal, and W. Wen, “The impact of generative ai on collaborative open-source software development: Evidence from github copilot,”arXiv preprint arXiv:2410.02091, 2024

  31. [31]

    Generative ai for pull request descriptions: Adoption, impact, and developer interventions,

    T. Xiao, H. Hata, C. Treude, and K. Matsumoto, “Generative ai for pull request descriptions: Adoption, impact, and developer interventions,” Proceedings of the ACM on Software Engineering, vol. 1, no. FSE, pp. 1043–1065, 2024

  32. [32]

    Programming by chat: A large-scale be- havioral analysis of 11,579 real-world ai-assisted ide sessions,

    N. Tang, C. Chen, Z. Fang, G. Xu, M. Dhakal, Y . Shi, C. McMillan, Y . Huang, and T. J.-J. Li, “Programming by chat: A large-scale be- havioral analysis of 11,579 real-world ai-assisted ide sessions,”arXiv preprint arXiv:2604.00436, 2026

  33. [33]

    Phantom: Curating github for engineered software projects using time- series clustering,

    P. Pickerill, H. J. Jungen, M. Ochodek, M. Ma ´ckowiak, and M. Staron, “Phantom: Curating github for engineered software projects using time- series clustering,”Empirical Software Engineering, vol. 25, no. 4, pp. 2897–2929, 2020

  34. [34]

    Con- tribution patterns in open source software for social good: Dynamics, individuals, and impact,

    Z. Fang, Y . Zhang, T. Zimmermann, D. Ford, and Y . Huang, “Con- tribution patterns in open source software for social good: Dynamics, individuals, and impact,”Proceedings of the ACM on Human-Computer Interaction, vol. 10, no. 2, pp. 1–26, 2026

  35. [35]

    Security developer studies with{GitHub}users: Exploring a convenience sam- ple,

    Y . Acar, C. Stransky, D. Wermke, M. L. Mazurek, and S. Fahl, “Security developer studies with{GitHub}users: Exploring a convenience sam- ple,” inThirteenth Symposium on Usable Privacy and Security (SOUPS 2017), 2017, pp. 81–95

  36. [36]

    Understanding skills for oss communities on github,

    J. T. Liang, T. Zimmermann, and D. Ford, “Understanding skills for oss communities on github,” inProceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2022, pp. 170–182

  37. [37]

    A matter of balance: Specialization, task variety, and individual learning in a software maintenance environment,

    S. Narayanan, S. Balasubramanian, and J. M. Swaminathan, “A matter of balance: Specialization, task variety, and individual learning in a software maintenance environment,”Management science, vol. 55, no. 11, pp. 1861–1876, 2009

  38. [38]

    When do changes induce fixes?

    J. ´Sliwerski, T. Zimmermann, and A. Zeller, “When do changes induce fixes?”ACM sigsoft software engineering notes, vol. 30, no. 4, pp. 1–5, 2005

  39. [39]

    Oops, my tests broke the build: An explorative analysis of travis ci with github,

    M. Beller, G. Gousios, and A. Zaidman, “Oops, my tests broke the build: An explorative analysis of travis ci with github,” in2017 IEEE/ACM 14th International conference on mining software repositories (MSR). IEEE, 2017, pp. 356–367

  40. [40]

    The im- pact of continuous integration on other software development practices: a large-scale empirical study,

    Y . Zhao, A. Serebrenik, Y . Zhou, V . Filkov, and B. Vasilescu, “The im- pact of continuous integration on other software development practices: a large-scale empirical study,” in2017 32nd IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2017, pp. 60–71

  41. [41]

    An empirical study on the survival rate of github projects,

    A. Ait, J. L. C. Izquierdo, and J. Cabot, “An empirical study on the survival rate of github projects,” inProceedings of the 19th International Conference on Mining Software Repositories, 2022, pp. 365–375

  42. [42]

    How do code changes evolve in different platforms? a mining-based investigation,

    M. Viggiato, J. Oliveira, E. Figueiredo, P. Jamshidi, and C. K ¨astner, “How do code changes evolve in different platforms? a mining-based investigation,” in2019 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 2019, pp. 218–222

  43. [43]

    Identifying reasons for software changes using historic databases,

    Mockus and V otta, “Identifying reasons for software changes using historic databases,” inProceedings 2000 international conference on software maintenance. IEEE, 2000, pp. 120–130

  44. [44]

    The missing links: bugs and bug-fix commits,

    A. Bachmann, C. Bird, F. Rahman, P. Devanbu, and A. Bernstein, “The missing links: bugs and bug-fix commits,” inProceedings of the eighteenth ACM SIGSOFT international symposium on Foundations of software engineering, 2010, pp. 97–106

  45. [45]

    It’s not a bug, it’s a feature: how misclassification impacts bug prediction,

    K. Herzig, S. Just, and A. Zeller, “It’s not a bug, it’s a feature: how misclassification impacts bug prediction,” in2013 35th international conference on software engineering (ICSE). IEEE, 2013, pp. 392–401

  46. [46]

    The social structure of free and open source software development,

    K. Crowston and J. Howison, “The social structure of free and open source software development,” 2005

  47. [47]

    Preliminary steps toward a general the- ory of internet-based collective-action in digital information commons: Findings from a study of open source software projects,

    C. M. Schweik and R. English, “Preliminary steps toward a general the- ory of internet-based collective-action in digital information commons: Findings from a study of open source software projects,”International Journal of the Commons, vol. 7, no. 2, 2013

  48. [48]

    A four-year study of student contributions to oss vs. oss4sg with a lightweight intervention,

    Z. Fang, M. Endres, T. Zimmermann, D. Ford, W. Weimer, K. Leach, and Y . Huang, “A four-year study of student contributions to oss vs. oss4sg with a lightweight intervention,” inProceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2023, pp. 3–15

  49. [49]

    Exploring solutions to tackle low-quality contributions on GitHub,

    C. Moraes, “Exploring solutions to tackle low-quality contributions on GitHub,” GitHub Community Discussion #185387, Jan. 2026, accessed: 2026-06-22. [Online]. Available: https://github.com/orgs/community/ discussions/185387

  50. [50]

    GitHub ponders kill switch for pull re- quests to stop AI slop,

    T. Register, “GitHub ponders kill switch for pull re- quests to stop AI slop,” The Register, Feb. 2026. [Online]. Available: https://www.theregister.com/software/2026/02/03/ github-ponders-kill-switch-for-pull-requests-to-stop-ai-slop/4334869