Pith. sign in

REVIEW 3 major objections 4 minor 65 references

A repository-hosted Agent Governance Manifest can make AI-made contributions reviewable, raising exact risk-label recovery from 40.5% to 97.4% in a controlled test.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 22:23 UTC pith:WAKI2D44

load-bearing objection A solid, transparent design-science paper whose headline evaluation is close to a manipulation check: putting the labels in the stimulus makes the labels recoverable, but the artifact and audit are still worth a serious referee. the 3 major comments →

arxiv 2607.15769 v1 pith:WAKI2D44 submitted 2026-07-17 cs.SE cs.CY

Making Agent-Mediated Contributions Governable: A Project-Level Governance Manifest for Open-Source AI Collaboration

classification cs.SE cs.CY
keywords open-source governanceAI agentsgeneration-verification asymmetryAgent Governance Manifestcontribution reviewevidence obligationsrisk zonesmaintainer authority
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that coding agents have widened a generation–verification asymmetry in open-source software: AI makes contributions cheap to produce but does not make them cheap to verify. Its central claim is that projects need project-side governability infrastructure—a repository-hosted rule set, instantiated as the Agent Governance Manifest (AGM)—that turns risk classification, evidence obligations, human confirmation, and review gates into contribution-level states that reviewers can recover. A diagnostic audit of 50 repositories found widespread general governance and agent-readable guidance but no project-wide arrangement satisfying all four governability functions. In a controlled reviewer-side evaluation, AGM-supported materials recovered the exact risk level in 37 of 38 outputs versus 15 of 37 for ordinary materials; in a contributor-side feasibility check, all 45 final evidence packages represented the core governance state correctly. If the mechanism holds, maintainers can shift from reconstructing a contribution's governance state to verifying it, while keeping final acceptance authority in human hands.

Core claim

The discovery is that governability is a distinct organizational function, separable from agent-readability and traceability, and that it can be externalized before review. AGM carries project rules into contribution-specific evidence packages: risk zones map changed files to evidence obligations; contributor-side agents prepare change summaries, test evidence, provenance notes, and missing-evidence reports; human contributors confirm declarations for high-risk changes; maintainer-side review packets expose risk, evidence, gate, and accountability states. The paper reports that this structured externalization made repository-defined risk levels recoverable in 97.4% of AGM-supported reviewer

What carries the argument

The Agent Governance Manifest (AGM), a repository-hosted boundary resource described in a human-readable Markdown document and a structured YAML file. It acts as a bidirectional governance contract: on the contributor side it allocates risk-sensitive evidence obligations and contributor-confirmation declarations; on the maintainer side it defines review gates and review-support artifacts such as risk summaries, missing-evidence reports, test-evidence summaries, and gate states. The load-bearing mechanism is the externalization of governance states—risk classification, evidence status, accountability, and gate state—into inspectable contribution-level artifacts before review.

Load-bearing premise

The load-bearing premise is that contributors and their agents prepare evidence packages honestly and that human contributors actually inspect what they confirm; a structurally valid package with invented test results or a rubber-stamped confirmation would pass AGM's gates and mislead maintainers.

What would settle it

Give AGM a live trial with contributors who have incentives to pad evidence and reviewers who rubber-stamp confirmations; if fabricated or unverified packages pass structural validation and are merged at rates comparable to ordinary materials, the central claim that AGM makes contributions governable would be refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Projects using AGM can expect higher-risk AI-mediated changes to be classified accurately at review time, reducing governance-risk under-classification from 59.5% of ordinary outputs to 2.6% of AGM-supported outputs.
  • Contributor-side agents can prepare structured evidence packages that humans confirm, while maintainer-review status remains outside the contributor workflow.
  • Maintainer review shifts from open-ended reconstruction to targeted verification, with missing or placeholder evidence surfaced by review-support output.
  • The same governance contract can be rendered in human-facing review interfaces such as risk cards, checklists, and gate-status panels without changing the underlying schema.
  • AGM can complement existing agent-readable instruction files and provenance records, which the paper positions as inputs rather than substitutes.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same bidirectional-contract logic could generalize beyond open-source to any workflow where human approval gates machine-generated output—CI/CD pipelines, scientific analysis scripts, or regulated documentation—with risk-zoned evidence obligations configured per domain.
  • Editorial inference: because AGM validates structure, not truth, its real-world value will depend on complementary assurance such as signing, attestation, or audit in adversarial settings; the paper explicitly leaves those mechanisms outside its scope.
  • Editorial inference: a testable extension is whether AGM changes maintainer behavior in live repositories—for example, reducing time-to-review or slowing the acceptance of low-quality AI contributions—which the controlled evaluation does not measure.
  • Editorial inference: the audit's finding that no repository coordinates all four functions suggests governability is not an emergent byproduct of mature governance; it requires deliberate design as a distinct layer.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper addresses the governance burden created by AI coding agents in open-source software. It proposes a three-layer framework (agent-readability, traceability, governability), reports a diagnostic audit of 50 GitHub repositories finding fragmented AI-governance cues but no project-wide governability arrangement, and develops the Agent Governance Manifest (AGM) as a repository-hosted governance resource. The evaluation has two parts: a within-participant reviewer-side study (15 participants, 75 outputs) reporting that AGM-supported materials improve exact risk-label recovery (37/38 vs. 15/37) and perceived review support, and a contributor-side feasibility check (15 participants, 45 tasks) reporting that all final packages represent the core governance state correctly with 41/45 passing strict structural validation.

Significance. If the central claim is established, the paper would make a useful conceptual and practical contribution to OSS governance: it names a distinct governability function, gives it a repository-hosted artifact form, and provides a bidirectional contributor/maintainer workflow. The audit is thoughtfully designed, the statistical reporting is transparent (participant-clustered bootstrap CIs, leave-one-out checks, pre-adjudication inter-coder reliability, disclosure of one AGM error), and the replication-package plan is a genuine strength. However, the headline quantitative support is partly circular: the AGM-supported condition contains the very governance states used as outcome measures. The paper's own boundary statements in §4.5 and §8.3 are honest about what AGM does not guarantee, but the title and abstract overstate what the controlled studies can support.

major comments (3)
  1. [§7.2, Tables S30–S31] The central objective contrast is confounded with information inclusion. AGM-supported materials include the repository-defined risk-zone reference, risk summaries, evidence indexes, contributor-confirmation declarations, and explicit gate states (Table S30). The outcome rubric codes whether the output names the same risk label, evidence status, and gate state (Table S31). The 97.4% vs. 40.5% difference therefore largely demonstrates that participants can read an answer that is literally present in the stimulus; it does not isolate AGM's specific governance structure (risk zoning, evidence packaging, confirmation gates) as the causal mechanism. A control condition presenting the same governance information in plain, non-AGM prose, or an analysis holding information constant while varying structure, is needed to support the claim that AGM's design, rather than simply telling reviewers the
  2. [§7.2, 'task-level pattern' paragraph] The text states that the condition difference is located in 'the structured externalization of repository-defined governance states.' This is an interpretation, not a demonstrated mechanism. The alternative explanation—that ordinary materials simply omitted the relevant risk-zone and gate information—is equally consistent with the data. The phrase 'structured externalization' presumes that the manifest's structure carries the effect. Since the design does not vary structure independently of information content, this specific attribution should be removed or explicitly flagged as untested.
  3. [§7.5 and Conclusion] The contributor-side feasibility check shows that cooperative participants can fill AGM templates and that agent output can be confirmed by humans. It does not test whether contributors prepare evidence honestly or whether maintainers actually inspect what they confirm. The paper appropriately acknowledges this boundary in §4.5 ('Projects seeking stronger guarantees against ignored rules, omitted evidence, or fabricated declarations require additional identity, signing, audit, attestation, cryptographic, or platform-level mechanisms') and §8.3. Given that acknowledgment, the abstract's and conclusion's phrasing—that AGM makes agent-mediated contributions 'governable'—should be tempered to 'governable under cooperative, non-adversarial conditions,' with the controlled feasibility clearly distinguished from field-level assurance.
minor comments (4)
  1. [Figure 4A] The panel reports 'Gate-state availability or correctness: 0.0% vs. 100.0%.' Under ordinary materials, gate states were not available at all, so 0.0% reflects non-observability, not incorrect judgment. The label should distinguish availability from correctness, and the text should note that the comparison conflates these two dimensions.
  2. [Abstract] The abstract reports 'exact risk-label recovery (37/38 vs. 15/37)' without noting that AGM-supported materials contained the risk labels in the stimulus. A short qualifier such as 'when the same governance information was supplied through the manifest structure' would prevent misreading.
  3. [§5.3, Table S10] Several legacy agreement statistics are reported with κ = 0.000, which can occur with low prevalence or skewed margins. Since these legacy variables were abandoned and replaced by the layered coding, the reporting is transparent, but a one-sentence explanation of why κ is uninformative in those rows would help.
  4. [§6.1, Finding 1] The finding that 'no repository in the audit satisfies the four-function criterion' is partly a consequence of the criterion's strictness (canonical, repository-visible, coordinating all four functions). The paper explains this, but it should be stated even more explicitly that the audit measures absence of a particular coordinating arrangement, not absence of governance intent or of individual mechanisms.

Circularity Check

1 steps flagged

The headline reviewer-side recovery result is largely an input-containment artifact: AGM-supported materials contain the exact risk labels, gate states, and reference-vs-observed packets that the outcome rubric then scores, so the 2.40x gain chiefly shows that supplying the answer makes it recoverable.

specific steps
  1. self definitional [§7.2; Table S30 and Table S31 in Supplementary S5–S6]
    "Among AGM-supported task-level reviewer-side outputs, 37 of 38 (97.4%) recovered the repository-defined risk level exactly, compared with 15 of 37 (40.5%) ordinary-material outputs. ... Repository-defined risk-zone reference: Not structured / Provided through AGM risk-zone rules and risk summaries. ... Maintainer-facing review packet: Not available / Available as a reference-vs-observed diagnostic report. ... Governance-gate state: Not available as an explicit state / Available as pass, needs-evidence, or blocked. ... Exact risk label: The final risk category in the reviewer-side output matche"

    The outcome measure is recovery of values that the AGM-supported condition explicitly supplies: the treatment includes risk-zone references, risk summaries, explicit gate states, contributor-confirmation declarations, and a reference-vs-observed review packet. The objective rubric then scores whether the reviewer-side output names those same values. Thus the 97.4% vs. 40.5% contrast primarily demonstrates that including the answer in the stimulus makes it recoverable; it does not isolate the contribution of AGM's governance structure (risk zoning, evidence packaging, confirmation gates) over simply telling reviewers the correct state. No control condition provides the same governance information in a plain, non-AGM format, so the central mechanism-level claim reduces to an input-containmen

full rationale

The paper has no meaningful self-citation chain: the cited prior work is external and the design is not justified by an author-imported uniqueness theorem. The 50-repository audit is a transparently coded diagnostic application of the paper's own four-function criterion, and applying a new criterion to find an absence is a legitimate, non-circular move even though the criterion is the authors' own. The contributor-side feasibility check (45/45 core states correct, 41/45 strict structural validation) is also not circular: it is a stated feasibility test of whether templates can be filled and human-confirmed, and the paper explicitly limits it to controlled conditions and acknowledges that stronger guarantees require identity, signing, audit, or platform mechanisms (§4.5, §8.3). The genuine circularity is confined to the headline reviewer-side comparison in §7.2: the treatment materials embed the very governance states used as outcome variables, so the reported 56.8-point gap is, by construction, mostly an effect of putting the answer into the stimulus rather than evidence for AGM's specific structural design. That makes the central quantitative support partially circular, but not the whole paper.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 3 invented entities

The central claims rest on: (1) the generation–verification asymmetry framing (domain assumption, cited); (2) the four-function governability criterion — a hand-chosen, post-hoc-refined definition that produces the 0/50 audit result; (3) the prototype's hand-chosen risk-zone calibration and author-assigned task reference labels, which define the evaluation's ground truth; (4) the honesty-of-evidence assumption, explicitly scoped out of the design (§4.5); and (5) author-built validators. No numeric fits or physics-like entities; the AGM artifact is the only concrete invented entity and it has a public, pinned, deterministic implementation.

free parameters (4)
  • Four-function governability criterion
    'Project-wide governability arrangement' requires canonical coordination of risk classification, evidence obligations, accountability states, and review gates. Adopted after adjudicating inter-coder disagreements; produces the 0/50 audit finding by construction (§5.3, S2).
  • Prototype risk-zone calibration = docs=low; tests/config=medium; core logic=high; auth/deps/workflows=critical
    Hand-chosen mapping in the AGM prototype; the paper states repositories define their own zones (§4.4). Evaluation reference labels derive from this calibration.
  • Task reference risk labels (T1-T5) = T1 low; T2 medium; T3 high; T4 critical; T5 critical
    Author-assigned ground truth for the exact risk-label recovery outcome (§5.4, S5). The main treatment effect measures recovery of these author-assigned labels.
  • Strict vs core structural validation thresholds = 41/45 strict; 45/45 core
    Author-specified rule set distinguishing schema precision from governance-state correctness; the review_gate.required boolean and command/result field requirements are hand-chosen (§7.5, S7).
axioms (4)
  • domain assumption Generation–verification asymmetry: AI lowers generation cost more than verification cost
    Framing premise asserted in §1/§3.1 and supported by cited empirical work; if maintainers could verify as cheaply as agents generate, governability infrastructure would be unnecessary.
  • domain assumption Contributors and their agents will prepare evidence honestly and contributors will actually confirm packages
    Load-bearing for the whole evidence-package mechanism; fabrication and ignored rules are excluded and deferred to 'additional identity, signing, audit, attestation, cryptographic, or platform-level mechanisms' (§4.5).
  • domain assumption Public GitHub artifacts are a valid operationalization of project governance
    The audit codes only public artifacts; private maintainer practices and hidden AI use are acknowledged as unobservable lower bounds (§5.2, §8.3).
  • domain assumption Agent outputs in the chosen reviewer-side environments represent reviewer-side behavior
    Reviewer-side outputs were produced by participants plus selected agents (Codex, Kilo Code, OpenCode, Continue); the paper states model-level attribution is out of scope (§5.4, S5).
invented entities (3)
  • AGM (Agent Governance Manifest) independent evidence
    purpose: Repository-hosted canonical governance resource carrying risk zones, evidence obligations, contributor-confirmation states, and maintainer review gates across the contribution workflow
    Public reference prototype pinned to commit c781a2f with deterministic validation scripts provides a falsifiable handle outside the paper; however, the evaluation of its effectiveness is author-run.
  • Project-side governability infrastructure no independent evidence
    purpose: Theorized organizational arrangement that allocates risk rules, evidence obligations, accountability states, and review gates
    Interpretive/theoretical construct; no falsifiable handle distinct from the paper's own coding scheme.
  • Governable boundary object no independent evidence
    purpose: Characterizes an evidence-bearing contribution interpretable across contributors, agents, maintainers, and platforms
    Theoretical framing concept applied to contributions; its existence claim reduces to the paper's own conceptual apparatus.

pith-pipeline@v1.3.0-alltime-deepseek · 43602 in / 19483 out tokens · 151503 ms · 2026-08-01T22:23:56.531467+00:00 · methodology

0 comments
read the original abstract

Generative AI and coding agents are intensifying a central governance tension in open-source software (OSS): they scale contribution generation faster than maintainers can assess risk, evidence, and accountability. Existing responses improve agent-readability and traceability, but project rules must also organize contribution-specific risk, evidence, accountability, and review-gate states. We theorize this organizational arrangement as project-side governability infrastructure. A diagnostic audit of 50 GitHub repositories finds widespread general governance artifacts, observable agent-readability, and fragmented AI-governance cues, but no project-wide arrangement that coordinates shared rules, preparation obligations, verification rights, and maintainer decision authority across AI-mediated contribution workflows. We develop the Agent Governance Manifest (AGM) as a repository-hosted boundary resource and bidirectional governance contract linking contributor-side evidence preparation with maintainer-side verification. In a controlled reviewer-side evaluation with 15 participants and 75 task-level outputs, AGM-supported materials improved exact risk-label recovery (37/38 vs. 15/37) and perceived review support (6.14 vs. 3.27 on a 1-7 scale). In a contributor-side feasibility check, 15 participants completed 45 tasks; all final packages represented the core governance state correctly, and 41 passed strict structural validation. The study develops a three-layer framework of agent-readability, traceability, and governability, theorizes agent-mediated contributions as governable boundary objects, and advances compliance-enabling digital innovation governance while preserving maintainer decision authority.

Figures

Figures reproduced from arXiv: 2607.15769 by Jinjin Gao, Ligang He, Luyang Li, Shufen Guo, Xiaoning Sun.

Figure 1
Figure 1. Figure 1: Three-layer framework for AI-mediated OSS governance. Agent-readability [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: AGM workflow across contribution preparation and maintainer review. Project [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Repository audit summary by governance-artifact category. Values indicate the [PITH_FULL_IMAGE:figures/full_fig_p030_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Evaluation summary. Panel A compares ordinary and AGM-supported materials [PITH_FULL_IMAGE:figures/full_fig_p035_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

65 extracted references · 7 canonical work pages

  1. [1]

    Brunswicker, S

    S. Brunswicker, S. Haefliger, Is There Collaboration in Open Col- laboration? The Role of Producers and Corporate Users in Open Source Software Development, Technovation 148 (2025) 103325.doi: 10.1016/j.technovation.2025.103325

  2. [2]

    M. E. Kowalski, L. A. de Vasconcelos Gomes, F. M. Borini, R. C. Bernardes, Business-to-business ecosystem smartification for manufac- turing: A definition, an integrative framework, and future directions, Technovation 151 (2026) 103425.doi:10.1016/j.technovation.2025. 103425. 47

  3. [3]

    Mager, K

    A. Mager, K. Mayer, R. Ridgway, The Politics of Open Digital Knowledge Infrastructures, Book chapter inThe Politics of Open Infrastructures, Open Book Publishers, accessed: 2026-06-24 (5 2026).doi:10.11647/O BP.0528.00

  4. [4]

    doi:10.1016/S0048-7333(03)00048-9

    S.O’Mahony, Guardingthecommons: Howcommunitymanagedsoftware projects protect their work, Research Policy 32 (7) (2003) 1179–1198. doi:10.1016/S0048-7333(03)00048-9. URLhttps://doi.org/10.1016/S0048-7333(03)00048-9

  5. [5]

    O’Mahony, F

    S. O’Mahony, F. Ferraro, The emergence of governance in an open source community, Academy of Management Journal 50 (5) (2007) 1079–1106. doi:10.5465/amj.2007.27169153. URLhttps://doi.org/10.5465/amj.2007.27169153

  6. [6]

    J. West, S. O’Mahony, The role of participation architecture in growing sponsored open source communities, Industry and Innovation 15 (2) (2008) 145–168.doi:10.1080/13662710801970142. URLhttps://doi.org/10.1080/13662710801970142

  7. [7]

    A. Fan, B. Gokkaya, M. Harman, M. Lyubarskiy, S. Sengupta, S. Yoo, J. M. Zhang, Large Language Models for Software Engineering: Survey and Open Problems, in: 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE), 2023, pp. 31–53.doi:10.1109/ICSE-FoSE59343.2023.00008

  8. [8]

    X. Hou, Y. Zhao, Y. Liu, Z. Yang, K. Wang, L. Li, X. Luo, D. Lo, J. Grundy, H. Wang, Large Language Models for Software Engineering: A Systematic Literature Review, ACM Trans. Softw. Eng. Methodol. 33 (8) (dec 2024).doi:10.1145/3695988

  9. [9]

    X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, D. Shao, N. Muennighoff, Y. Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, G. Neubig, Openhands: An open platform for ai software developers as generalist agents, in: Y. Yue, A. Garg, N. Peng, F. Sha, R. Yu (Eds.), Internat...

  10. [10]

    H. Li, H. Zhang, A. E. Hassan, The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering (2025).arXiv:2507.15003

  11. [11]

    H. Li, H. Zhang, A. E. Hassan, AIDev: Studying AI Coding Agents on GitHub (2026).arXiv:2602.09185

  12. [12]

    Gloaguen, N

    T. Gloaguen, N. Mündler, M. N. Mueller, V. Raychev, M. Vechev, Evaluating AGENTS.md: Are repository-level context files helpful for coding agents?, in: ICLR 2026 Workshop on Memory for LLM-Based Agentic Systems, 2026.arXiv:2602.11988

  13. [13]

    Mohsenimofidi, M

    S. Mohsenimofidi, M. Galster, C. Treude, S. Baltes, Context Engineering for AI Agents in Open-Source Software (2026).arXiv:2510.21413

  14. [14]

    Shepard, J

    A. Shepard, J. Albrecht, Probe-and-Refine Tuning of Repository Guid- ance for Coding Agents (2026).arXiv:2606.20512

  15. [15]

    URLhttps://agent-trace.dev/

    Cursor, Agent Trace, Open rfc specification, Cursor, accessed June 23, 2026 (1 2026). URLhttps://agent-trace.dev/

  16. [16]

    W. Yang, R. He, M. Zhou, Beyond Banning AI: A First Look at GenAI Governance in Open Source Software Communities (2026). arXiv: 2603.26487

  17. [17]

    M. L. Tushman, D. A. Nadler, Information processing as an integrating concept in organizational design, Academy of Management Review 3 (3) (1978) 613–624.doi:10.5465/amr.1978.4305791. URLhttps://doi.org/10.5465/amr.1978.4305791

  18. [18]

    AlMarzouq, V

    M. AlMarzouq, V. Grover, J. B. Thatcher, Taxing the Development Structure of Open Source Communities: An Information Processing View, Decision Support Systems 80 (2015) 27–41.doi:10.1016/j.dss. 2015.09.004

  19. [19]

    URLhttps://doi.org/10.5465/amr.1984.4277657 49

    R.L.Daft, K.E.Weick, Towardamodeloforganizationsasinterpretation systems, Academy of Management Review 9 (2) (1984) 284–295.doi: 10.5465/amr.1984.4277657. URLhttps://doi.org/10.5465/amr.1984.4277657 49

  20. [20]

    M. L. Markus, The governance of free/open source software projects: Monolithic, multidimensional, or configurational?, Journal of Manage- ment & Governance 11 (2) (2007) 151–163.doi:10.1007/s10997-007 -9021-x. URLhttps://doi.org/10.1007/s10997-007-9021-x

  21. [21]

    Shaikh, O

    M. Shaikh, O. Henfridsson, Governing open source software through coordination processes, Information and Organization 27 (2) (2017) 116– 135.doi:10.1016/j.infoandorg.2017.04.001. URLhttps://doi.org/10.1016/j.infoandorg.2017.04.001

  22. [22]

    Gousios, M

    G. Gousios, M. Pinzger, A. v. Deursen, An exploratory study of the pull-based software development model, in: Proceedings of the 36th International Conference on Software Engineering, ICSE 2014, Associa- tion for Computing Machinery, New York, NY, USA, 2014, pp. 345–355. doi:10.1145/2568225.2568260

  23. [23]

    Alami, R

    A. Alami, R. Pardo, M. L. Cohn, A. Wąsowski, Pull request governance in open source communities, IEEE Transactions on Software Engineering 48 (12) (2022) 4838–4856.doi:10.1109/TSE.2021.3128356

  24. [24]

    Linåker, E

    J. Linåker, E. Papatheocharous, T. Olsson, How to characterize the health of an open source software project? a snowball literature review of an emerging practice, in: Proceedings of the 18th International Sympo- sium on Open Collaboration, OpenSym ’22, Association for Computing Machinery, New York, NY, USA, 2022.doi:10.1145/3555051.3555067

  25. [25]

    Oliveira, T

    P. Oliveira, T. Conte, M. Gerosa, I. Steinmacher, Governance in Practice: How Open Source Projects Define and Document Roles (2026).arXiv: 2603.24879

  26. [26]

    arXiv:2509.16295

    M.Noori, M.Chakraborti, A.X.Zhang, S.Frey, Patternsinthetransition from founder-leadership to community governance of open source (2026). arXiv:2509.16295

  27. [27]

    S. K. Shah, Motivation, governance, and the viability of hybrid forms in open source software development, Management Science 52 (7) (2006) 1000–1014.doi:10.1287/mnsc.1060.0553. URLhttps://doi.org/10.1287/mnsc.1060.0553 50

  28. [28]

    L. Yin, M. Chakraborti, Y. Yan, C. Schweik, S. Frey, V. Filkov, Open source software sustainability: Combining institutional analysis and socio- technical networks, Proc. ACM Hum.-Comput. Interact. 6 (CSCW2) (nov 2022).doi:10.1145/3555129

  29. [29]

    Noori, M

    M. Noori, M. Chakraborti, A. X. Zhang, S. Frey, A human behavioral baseline for collective governance in software projects, in: NeurIPS 2025 Workshop on Algorithmic Collective Action, 2025.arXiv:2510.08956

  30. [30]

    E. D. L. Cruz, H. Le, K. Meduri, G. S. Nadella, H. Gonaygunta, Re- defining the programmer: Human-ai collaboration, llms, and security in modern software engineering, Computers, Materials and Continua 85 (2) (2025) 3569–3582.doi:10.32604/cmc.2025.068137

  31. [31]

    S. Peng, E. Kalliamvakou, P. Cihon, M. Demirer, The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (2023).arXiv: 2302.06590

  32. [32]

    H. He, C. Miller, S. Agarwal, C. Kästner, B. Vasilescu, Speed at the Cost of Quality? The Impact of LLM Agent Assistance on Software Development (2025).arXiv:2511.04427

  33. [33]

    F. Song, A. Agarwal, W. Wen, The Impact of Generative AI on Col- laborative Open-Source Software Development: Evidence from GitHub Copilot (2026).arXiv:2410.02091

  34. [34]

    Watanabe, H

    M. Watanabe, H. Li, Y. Kashiwa, B. Reid, H. Iida, A. E. Hassan, On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub, ACM Trans. Softw. Eng. Methodol.Just Accepted (mar 2026). doi:10.1145/3798166

  35. [35]

    Rahman, M

    S. Rahman, M. F. Rabbi, M. Zibran, A Task-Level Evaluation of AI Agents in Open-Source Projects (2026).arXiv:2602.02345

  36. [36]

    Branco, P

    R. Branco, P. Canelas, C. Gamboa, A. Fonseca, LGTM! Characteristics of Auto-Merged LLM-based Agentic PRs, metadata retained from provided BibTeX; public venue/identifier not verified. (2026)

  37. [37]

    Sawada, T

    S. Sawada, T. Shirai, Y. Kashiwa, K. Yamaguchi, H. Iwata, H. Iida, To What Extent Does Agent-generated Code Require Maintenance? An Empirical Study (2026).arXiv:2605.06464. 51

  38. [38]

    M. L. Siddiq, X. Zhao, V. C. Lopes, B. Casey, J. C. S. Santos, Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub (2026).arXiv:2601.00477

  39. [39]

    An Endless Stream of AI Slop

    S. Baltes, M. Cheong, C. Treude, "An Endless Stream of AI Slop": The Growing Burden of AI-Assisted Software Development (2026).arXiv: 2603.27249

  40. [40]

    Sen, GitHub Weighs Pull Request Kill Switch As AI Slop Floods Open Source, Open Source ForU, accessed: 2026-06-24 (2 2026)

    A. Sen, GitHub Weighs Pull Request Kill Switch As AI Slop Floods Open Source, Open Source ForU, accessed: 2026-06-24 (2 2026). URL https://www.opensourceforu.com/2026/02/github-weighs-p ull-request-kill-switch-as-ai-slop-floods-open-source/

  41. [41]

    Iyer, AI Coding Agents Broke the PR Pipeline

    A. Iyer, AI Coding Agents Broke the PR Pipeline. Validation Is How You Fix It., Signadot Blog, accessed: 2026-06-24 (5 2026). URL https://www.signadot.com/blog/ai-generated-code-crisi s/

  42. [42]

    A. Hora, R. Robbes, AI Policy, Disclosure, and Human in the Loop: How Are Contribution Guidelines Adapting to GenAI? (2026).arXiv: 2605.16706

  43. [43]

    URL https://docs.fedoraproject.org/en-US/council/policy/a i-contribution-policy/

    Fedora Project, AI-Assisted Contributions Policy, Official project policy, Fedora Council, accessed: 2026-06-24 (10 2025). URL https://docs.fedoraproject.org/en-US/council/policy/a i-contribution-policy/

  44. [44]

    Manita, A

    J. Manita, A. Amari, Regulating the Machine Contributor: Governance and Policy Alignment in Open Source (2026).arXiv:2606.14594

  45. [45]

    Kraishan, The AI Attribution Paradox: Transparency as Social Strategy in Open-Source Software Development (2025).arXiv:2512.0 0867

    O. Kraishan, The AI Attribution Paradox: Transparency as Social Strategy in Open-Source Software Development (2025).arXiv:2512.0 0867

  46. [46]

    Tufano, F

    R. Tufano, F. Pepe, F. Zampetti, A. Mastropaolo, O. Dabić, M. Di Penta, G. Bavota, Developers and Generative AI: A Study of Self-Admitted Usage in Open Source Projects, Empirical Software Engineering 31 (4) (2026) 108.doi:10.1007/s10664-026-10848-w. 52

  47. [47]

    G. V. Datla, A. Vurity, T. Dash, T. Ahmad, M. Adnan, S. Rafi, Exe- cutable Governance for AI: Translating Policies into Rules Using LLMs (2025).arXiv:2512.04408

  48. [48]

    Vispute, A

    N. Vispute, A. Kadam, Reasoning Provenance for Autonomous AI Agents: Structured Behavioral Analytics Beyond State Checkpoints and Execu- tion Traces (2026).arXiv:2603.21692

  49. [49]

    Y. Cai, W. Tang, C. Wen, S. Qin, Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents (2026).arXiv:2604.23374

  50. [50]

    Mayer, A

    A.-S. Mayer, A. Kostis, F. Strich, J. Holmström, Shifting dynamics: How generative ai as a boundary resource reshapes digital platform governance, Journal of Management Information Systems 42 (2) (2025) 400–430. arXiv:https://doi.org/10.1080/07421222.2025.2487312, doi:10.1080/07421222.2025.2487312

  51. [51]

    C. C.Su, N. K. Chan, Assembling platformgovernance asprivateordering in the age of generative ai: platform interdependence in policy evolution, Information, Communication & Society 29 (6) (2026) 1929–1953.arXiv: https://doi.org/10.1080/1369118X.2025.2513672 , doi:10.1080/ 1369118X.2025.2513672

  52. [52]

    Ghazawneh, O

    A. Ghazawneh, O. Henfridsson, Balancing platform control and external contribution in third-party development: The boundary resources model, Information Systems Journal 23 (2) (2013) 173–192.doi:10.1111/j.13 65-2575.2012.00406.x. URLhttps://doi.org/10.1111/j.1365-2575.2012.00406.x

  53. [53]

    Handler, Collaboration without consensus - free and open source software as boundary object, Conference abstract (2017)

    R. Handler, Collaboration without consensus - free and open source software as boundary object, Conference abstract (2017). URL https://socav.gu.se/digitalAssets/1642/1642685_all-abs tracts-170530.pdf

  54. [54]

    URLhttps://chaoss.community/kb/terminology/

    CHAOSS Community, CHAOSS Specific Terms, Community Knowledge Base, accessed: 2026-06-17 (2026). URLhttps://chaoss.community/kb/terminology/

  55. [55]

    53 URL https://www.cncf.io/blog/2025/10/22/lfx-insights-a-new -way-to-understand-open-source-projects/

    Linux Foundation, LFX Insights: A new way to understand open source projects, CNCF Blog, accessed: 2026-06-17 (10 2025). 53 URL https://www.cncf.io/blog/2025/10/22/lfx-insights-a-new -way-to-understand-open-source-projects/

  56. [56]

    Ivchenko, Community Health Metrics: Contributor Diversity, Bus Factor, and Sustainability Signals, Online research report, Zenodo (4 2026)

    O. Ivchenko, Community Health Metrics: Contributor Diversity, Bus Factor, and Sustainability Signals, Online research report, Zenodo (4 2026). URLhttps://zenodo.org/records/19476184

  57. [57]

    P. S. Adler, B. Borys, Two types of bureaucracy: Enabling and coercive, Administrative Science Quarterly 41 (1) (1996) 61–89. URLhttp://www.jstor.org/stable/2393986

  58. [58]

    Wouters, C

    M. Wouters, C. Wilderom, Developing performance-measurement sys- tems as enabling formalization: A longitudinal field study of a logistics department, Accounting, Organizations and Society 33 (4–5) (2008) 488– 516.doi:10.1016/j.aos.2007.05.002. URLhttps://doi.org/10.1016/j.aos.2007.05.002

  59. [59]

    R. L. Daft, R. H. Lengel, Organizational Information Requirements, Media Richness and Structural Design, Management Science 32 (5) (1986) 554–571.doi:10.1287/mnsc.32.5.554

  60. [60]

    A. G. L. Romme, J. Holmström, From theories to tools: Calling for research on technological innovation informed by design science, Techno- vation 121 (2023) 102692.doi:https://doi.org/10.1016/j.techno vation.2023.102692

  61. [61]

    Alvarez-Telena, M

    S. Alvarez-Telena, M. Diez-Fernandez, The three-ring architecture: Gov- erning agents in the era of on-platform organisations (2026).arXiv: 2606.07119

  62. [62]

    Santalo, I

    J. Santalo, I. Filatotchev, Strategic governance of blockchain platforms: From centralized to open source control systems, Long Range Planning 58 (4) (2025) 102539.doi:10.1016/j.lrp.2025.102539

  63. [63]

    Z. Feng, R. Milewicz, E. Murphy-Hill, T. Menezes, A. Serebrenik, I. Stein- macher, A. Sarma, Charting uncertain waters: A socio-technical roadmap for sustaining open source communities in the age of genai, ACM Trans. Softw. Eng. Methodol.Just Accepted (jan 2026).arXiv:2508.04921, doi:10.1145/3789210. 54

  64. [64]

    Hoffman, S

    T. Hoffman, S. Shambaugh, When AI Breaks the Systems Meant to Hear Us, O’Reilly Radar, metadata retained from provided BibTeX; public URL not verified. (3 2026)

  65. [65]

    MakingAgent-MediatedContributions Governable: A Project-Level Governance Manifest for Open-Source AI Collaboration

    M. Xu, J. Chen, Z. Zhang, Claw4science: A dataset and platform for the openclaw scientific agent ecosystem, bioRxiv (2026).arXiv:https: //www.biorxiv.org/content/early/2026/04/01/2026.03.30.7151 18.full.pdf,doi:10.64898/2026.03.30.715118. 55 SupplementaryMaterialfor“MakingAgent-MediatedContributions Governable: A Project-Level Governance Manifest for Open...