Pith. sign in

REVIEW 1 major objections 4 minor 3 cited by

Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development

T0 review · 1 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A survey of 574 developers finds that the most common view is that AI-generated code belongs to no one, while the second-largest group credits the prompter or employer, amid unresolved law.

desk verdict A transparent, well-scoped descriptive survey of developer views on GenAI copyright that earns its keep despite acknowledged sampling limits. read the letter →

arxiv 2411.10877 v5 pith:JQGQ5DVP submitted 2024-11-16 cs.SE cs.AI

classification cs.SEcs.AI
keywords generativeAIcopyrightsoftwarelicensingdevelopersurveyAI-generatedcodeownershipopen-sourcelegalperceptionsgeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a descriptive study of 574 software developers who use generative-AI coding tools, supplemented by seven follow-up interviews. It aims to capture, at a moment of legal uncertainty, what developers think about copyright and licensing for AI-generated code: who should own it, whether training on code should require permission or compensation, and what risks they worry about. The headline result is that opinion is split and confident: 44% of respondents say AI-generated code belongs to no one, 36% say it belongs to the prompter or their employer, and about three-quarters report being confident or very confident in their answer even though courts and regulators have not settled the question. The authors present the study as a resource for policymakers, arguing that regulation of this area should be informed by the views of the people who actually use the tools.

What carries the argument

The machinery that carries the study is the survey instrument paired with qualitative coding. The questionnaire has 26 closed and 7 open questions in three sections—current GenAI use, understanding and perception of copyright issues, and demographics—and the open-ended answers were coded independently by two researchers using an inductive open-coding process, with disagreements resolved by discussion. That combination produces the paper's result: a structured snapshot of percentages plus quotes that explain the reasoning behind them. The single question doing the most work is U11, the multiple-choice ownership question, together with its confidence and rationale follow-ups.

What would settle it

A replication survey drawn from a representative sample of developers outside GenAI-related open-source repositories—for example, a random sample of enterprise developers in proprietary settings—that found a different ownership distribution (say, a majority choosing the employer) would undercut the generalizing reading of the 44% and 36% figures, though it would not contradict the paper's description of its own respondents.

Watch

Extended reading notes

Core claim

The central claim is a descriptive one: developers' opinions on the copyright status of AI-generated code are diverse, context-dependent, and often held with high confidence despite unresolved law. When asked to select all that applied, the most common answer was that generated code belongs to no one and sits in the public domain (242 of 554 respondents, 44%), followed by the view that it belongs to the prompting developer or their employer (201, 36%); smaller groups assigned ownership to training-data creators (156, 28%), model creators (49, 9%), or said they did not know (57, 10%). The paper also finds that developers mostly treat AI-generated code as similar to other reused code, are indifferent or pleased when their prompts are reused, are context-dependent about whether their own code may be used in training data, and largely work in organizations with no formal process for documenting AI use, while only 12% of respondents reported any formal copyright training.

Load-bearing premise

The load-bearing premise is that developers who interacted with GenAI-related open-source repositories and then chose to answer the survey are a meaningful window onto 'developers' views' generally; if that pool overrepresents GenAI enthusiasts and open-source contributors, the headline percentages may not generalize.

Editorial extensions

If this is right

  • Organizations cannot assume a single coherent developer view on ownership; internal policies on AI-generated code will have to manage expectations that range from public domain to employer ownership.
  • Terms-of-service reading is the exception, not the rule, so contractual terms about output ownership are unlikely to reach most developers.
  • Most developers report no organizational process for documenting AI-generated code, which complicates any future compliance or provenance requirements.
  • Policies on training data will need to distinguish open-source from proprietary code, since developers' reactions depend heavily on that context.
  • If courts assign ownership one way or the other, a large share of current developers will be surprised, suggesting a need for guidance that connects legal outcomes to existing developer intuitions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the sample was recruited through GenAI-related open-source repositories, the 44% public-domain figure may overstate the view among enterprise and closed-source developers; a broader sample could shift the balance toward employer ownership.
  • The high confidence paired with unresolved law suggests that many developers are forming ownership intuitions from tool behavior and community norms rather than from legal sources, which may make later judicial or regulatory rulings harder to absorb.
  • If AI agents begin producing whole applications with less human editing, the 'tools are tools' reasoning that supports prompter ownership may weaken; the paper anticipates this tension but does not test it.
  • The documentation gap reported by developers implies that any provenance or disclosure mandate would require new tooling and workflows, not just new rules; the paper stops short of designing such mechanisms.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. This paper reports a mixed-methods study of 574 developers (plus 7 follow-up interviews) recruited from GitHub users who forked, starred, watched, or contributed to 30 GenAI-related repositories. The survey asks about developers' use of generative AI tools, their perceptions of copyright and licensing issues, and other legal concerns. The central descriptive results are that developers hold varied views on ownership of AI-generated code (44% choosing 'public domain,' 36% choosing the prompter or employer, 28% choosing training-data creators, 9% choosing model creators), that 75% are confident or very confident in their ownership opinions despite unresolved law, and that developers are concerned about data leakage and lack of documentation/process for GenAI usage. The paper presents 29 findings organized by three research questions, discusses implications for policy and practice, and makes its survey and analysis artifacts available in an online replication package.

Significance. The study is timely and policy-relevant, directly responding to calls from the U.S. Copyright Office for stakeholder perspectives. Its strengths include a detailed survey design process, transparent reporting of instruments and qualitative coding, and explicit acknowledgment of sampling limitations. The descriptive claims are directly supported by the data as presented: the percentages, denominators, and analysis procedures are internally consistent. The main substantive limitation is external validity: the sample is self-selected from GitHub users who showed interest in GenAI repositories, and the 1.9% response rate limits generalizability beyond this population. The authors appropriately disclaim statistical generalization in Section 7.3, but the abstract and introductory framing could be more careful. If the reported percentages are interpreted as describing only this surveyed population, the paper provides a valuable snapshot for lawmakers, regulators, and organizations.

major comments (1)
  1. [Section 5.4, Figure 9] The survey instrument (Figure 2, U11) lists 'prompt creator' as a single response option, but the results in Figure 9 and Section 5.4.2 report a combined category 'Belongs to prompter (or their employer)' with 201 selections (36%). The paper should clarify how 'employer' responses were derived, including whether this category merges the pre-specified 'prompt creator' option with 'Other' responses, and what the individual counts were. As written, this headline percentage conflates two conceptually distinct views (individual prompt authorship versus employer ownership), which directly affects the interpretation of the central ownership-distribution claim.
minor comments (4)
  1. [Abstract] The opening results sentence states that 'Our results show the benefits developers derive from GenAI...' without explicit qualification to the surveyed sample; consider adding a phrase such as 'among the surveyed developers' to avoid overgeneralization, given the sampling limitations acknowledged in Section 7.3.
  2. [Section 3.3] The description of the qualitative coding process is dense; summarizing the final codebook structure and the specific criteria used to merge codes would improve reproducibility and reader comprehension.
  3. [Section 4.2.2, Figure 5] Figure 5 displays counts of documentation types but does not show the denominator (n=83) in the caption; adding 'n=83' and clarifying that multiple selections were possible would prevent reader confusion about the base for the reported percentages.
  4. [Section 2.1] The legal background is thorough but U.S.-centric; an early sentence noting that the analysis is grounded in U.S. law would be helpful because the study sample is international and Section 7.3 later acknowledges this limitation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a descriptive survey study whose findings are direct summaries of self-reported participant responses.

full rationale

This paper is a descriptive empirical study, not a derivation. The central results (e.g., Finding 18: 44% of 554 familiar respondents selected that AI-generated code belongs to no one; Finding 19: 36% selected the prompter or employer; Finding 12: 41% were neutral about their code being used in training data) are direct aggregates of survey responses collected from 574 valid respondents. There is no fitted parameter later relabeled as a prediction, no model whose output is constrained by its input in a circular way, and no invoked uniqueness theorem or self-citation that supplies the content of a conclusion. The survey instrument was designed from legal and software-engineering expertise, but questionnaire construction is not circular reasoning: the conclusions report what respondents said, and the authors explicitly avoid claiming that the sample is statistically representative, acknowledging in Sections 7.1 and 7.3 that respondents may overrepresent frequent GenAI users and open-source-oriented developers and that the study does not claim generalizability. The acknowledged limitations concern external validity and construct validity, not circularity. Accordingly, no circular step can be exhibited, and the appropriate finding is no significant circularity with score 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The study introduces no free parameters or invented entities. Its inferential backbone rests on the representativeness of the sample and the truthfulness of self-reports, both of which are standard assumptions for survey research and are partially acknowledged as threats to validity in Section 7.

assumptions (4)
  • domain assumption Survey respondents' self-reported perceptions and practices reflect their actual views and, to some degree, their behavior.
    Section 3.3 and Section 7.1: the paper takes responses at face value and notes they are communicated perceptions, not verified practices.
  • domain assumption The GitHub-based sampling frame (fork, star, watch, or contribute to GenAI-related repos) identifies developers who use or are familiar with GenAI tools.
    Section 3.2 assumes these actions are 'highly correlated with familiarity with using GenAI tools', and acknowledges the sample may overrepresent frequent users.
  • domain assumption The set of 30 selected repositories covers the GenAI coding tool ecosystem.
    Section 3.2: the repositories were chosen via tag search plus three hand-picked repos (CodeLlama, GitHub Copilot, Codeium); this set may not represent all GenAI coding tools.
  • domain assumption Qualitative coding with two annotators and reconciliation yields reliable categories even without inter-rater agreement statistics.
    Section 3.3: they explicitly do not base analysis on inter-rater agreements because coding was inductive and multiple codes per response were allowed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development." pith.science (2026). https://pith.science/paper/JQGQ5DVP

@misc{pith2026241110877,
  author       = {Pith},
  title        = {Pith review of: Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JQGQ5DVP}},
  note         = {Machine review of arXiv:2411.10877}
}
read the original abstract

Despite the utility that Generative AI (GenAI) tools provide for tasks such as writing code, the use of these tools raises important legal questions and potential risks, particularly those associated with copyright law. As lawmakers and regulators engage with those questions, the views of users can provide relevant perspectives. In this paper, we provide: (1) a survey of 574 developers on the licensing and copyright aspects of GenAI for coding, as well as follow-up interviews; (2) a snapshot of developers' views at a time when GenAI and perceptions of it are rapidly evolving; and (3) an analysis of developers' views, yielding insights and recommendations that can inform future regulatory decisions in this evolving field. Our results show the benefits developers derive from GenAI, how they view the use of AI-generated code as similar to using other existing code, the varied opinions they have on who should own or be compensated for such code, that they are concerned about data leakage via GenAI, and much more, providing organizations and policymakers with valuable insights into how the technology is being used and what concerns stakeholders would like to see addressed.

Figures

Figures reproduced from arXiv: 2411.10877 by the authors.

Figure 1
Figure 1. Overview of the study methodology 3 Study Methodology The purpose of this study is to investigate developers’ perceptions about licensing and copyright issues related to the use of GenAI technology for software development. (Although we use the term “software development” throughout the article, our study focuses on the use of GenAI for code generation as the main task that can lead to copyright and licensing issues… view at source ↗
Figure 2
Figure 2. Survey design overview without permission a prompt that they created (U9); ownership and copyrightability of generated code (U11, U12, U13); and other, broader concerns about copyright/legal issues stemming from development tasks with AI-based tools (U14). • The “Demographics” section (8 questions) asked about respondents’ age (D1); experience with SE and AI (D2, D3); their primary programming languages (D4); whethe… view at source ↗
Figure 3
Figure 3. Software development tasks where developers reported using GenAI [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Instances of tool category uses reported by developers [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Information collected/stored during documentation processes [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Developer approaches to Terms of Service (ToS) for AI Tools [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Developer perceptions on copyright and GenAI [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Developer attitudes towards copyright and GenAI [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: Developer perceptions on output ownership [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring the Challenges and Opportunities of AI-assisted Codebase Generation

    cs.SE 2025-08 conditional novelty 6.0 of 10

    Developers prompting codebase-level AI assistants are often dissatisfied with generated code, citing missing functionality, poor code quality, and communication gaps, despite varied prompting strategies.

  2. CodeMirage: A Multi-Lingual Benchmark for Detecting AI-Generated and Paraphrased Source Code from Production-Level LLMs

    cs.SE 2025-05 conditional novelty 6.0 of 10

    CodeMirage is a ten-language, ten-LLM benchmark with original and paraphrased AI code, and it shows current AI-generated-code detectors drop sharply under cross-model and low-false-alarm settings.

  3. DevLicOps: A Framework for Mitigating Licensing Risks in AI-Generated Code

    cs.SE 2025-08 conditional novelty 4.0 of 10

    DevLicOps integrates license-compliance controls into the SDLC to reduce risk from AI-generated code, using policies, automated scans, manual audits, and indemnity-aware practices.

Reference graph

Works this paper leans on

136 extracted references · 60 canonical work pages · cited by 3 Pith papers

  1. [1]

    Letter from Robert J

    2023. Letter from Robert J. Kasunic, Associate Register of Copyrights and Director of Registration Policy and Practice, U.S. Copyright Office, to Van Lindberg, Taylor English Duma LLP (February 21, 2023), https://www.copyright.gov/ docs/zarya-of-the-dawn.pdf

  2. [2]

    Getty Images (US) Inc

    2023. Getty Images (US) Inc. v Stability AI Inc. 1:23-cv-00135 (D. Del.)

  3. [3]

    Online replication package

    2025. Online replication package. https://archive.softwareheritage.org/browse/origin/directory/?origin_url=https: //github.com/TStalnaker44/copyright_and_genai_tosem_25

  4. [4]

    Qualtrics

    [n.d.]. Qualtrics. https://www .qualtrics.com/. Accessed: 2023-21-06

  5. [5]

    § 102(a) [n

    17 U.S.C. § 102(a) [n. d.]. 17 U.S.C. § 102(a)

  6. [6]

    § 106 [n

    17 U.S.C. § 106 [n. d.]. 17 U.S.C. § 106

  7. [7]

    § 1202(b) [n

    17 U.S.C. § 1202(b) [n. d.]. 17 U.S.C. § 1202(b)

  8. [8]

    §102(b) [n

    17 U.S.C. §102(b) [n. d.]. 17 U.S.C. §102(b)

Show all 136 references
  1. [9]

    John Abascal, Stanley Wu, Alina Oprea, and Jonathan Ullman. 2023. TMI! Finetuned Models Leak Private Information from their Pretraining Data. Proc. Priv. Enhancing Technol. 2024 (2023), 202–223. https://api .semanticscholar.org/ CorpusID:259064156

  2. [10]

    Stability AI Ltd

    Andersen v. Stability AI Ltd. [n. d.]. Andersen v. Stability AI Ltd. 3:23-cv-00201, (N.D. Cal.)

  3. [11]

    Anthropic Models [n. d.]. Models. https://docs .anthropic.com/en/docs/about-claude/models. Accessed: 2024-17-10

  4. [12]

    Authors Guild, Inc. v. Google, Inc. 2015. Authors Guild, Inc. v. Google, Inc., 804 F.3d 202 (2d Cir. 2015)

  5. [13]

    OpenAI Inc

    Authors Guild v. OpenAI Inc. 2023. Authors Guild v. OpenAI Inc., 1:23-cv-08292, (S.D.N.Y.)

  6. [14]

    AutoGPT [n. d.]. AutoGPT. https://github.com/Significant-Gravitas/AutoGPT

  7. [15]

    Rozière et al

    B. Rozière et al. 2024. Code Llama: Open Foundation Models for Code. arXiv:2308.12950 [cs.CL]

  8. [16]

    Mahak Bandi. 2019. All About Open Source Licenses. https://fossa.com/blog/what-do-open-source-licenses-even- mean/. Accessed: 2023-24-09

  9. [17]

    Iain Barclay, Alun Preece, Ian Taylor, and Dinesh Verma. 2019. Towards Traceability in Data Ecosystems Using a Bill of Materials Model. arXiv preprint arXiv:1904.04253 (2019)

  10. [18]

    Shraddha Barke, Michael B James, and Nadia Polikarpova. 2023. Grounded Copilot: How Programmers Interact with Code-Generating Models. Proceedings of the ACM on Programming Languages 7, OOPSLA1 (2023), 85–111

  11. [19]

    Ashley Belanger. 2024. Air Canada Has to Honor a Refund Policy Its Chatbot Made Up. https://www.wired.com/ story/air-canada-chatbot-refund-policy

  12. [20]

    Blake Brittain. 2023. Authors sue Meta, Microsoft, Bloomberg in latest AI copyright clash. https://www.reuters.com/ legal/litigation/authors-sue-meta-microsoft-bloomberg-latest-ai-copyright-clash-2023-10-18/

  13. [21]

    Governor Newsom announces new initiatives to advance safe and responsible AI, protect Californians

    california-regulation 2024. Governor Newsom announces new initiatives to advance safe and responsible AI, protect Californians. https://www.gov.ca.gov/2024/09/29/governor-newsom-announces-new-initiatives-to-advance-safe- ACM Trans. Softw. Eng. Methodol., Vol. 1, No. 1, Article...

  14. [22]

    Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021. Extracting Training Data from Large Language Models. In 30th USENIX Security Symposium (USENIX Security 21) ...

  15. [23]

    codeium [n. d.]. codeium.vim. https://github .com/Exafunction/codeium.vim

  16. [24]

    cody [n. d.]. cody. https://github .com/sourcegraph/cody

  17. [25]

    Consumer Protections for Artificial Intelligence

    colorado-bill 2024. Consumer Protections for Artificial Intelligence. https://leg .colorado.gov/bills/sb24-205

  18. [26]

    European Commission. 2024. The EU copyright legislation. https://digital-strategy.ec.europa.eu/en/policies/copyright- legislation

  19. [27]

    Kattiana Constantino, Mauricio Souza, Shurui Zhou, Eduardo Figueiredo, and Christian Kästner. 2023. Perceptions of open-source software developers on collaborations: An interview and survey study. Journal of Software: Evolution and Process 35, 5 (2023), e2393

  20. [28]

    Feder Cooper and James Grimmelmann

    A. Feder Cooper and James Grimmelmann. 2024. The Files Are in the Computer: Copyright, Memorization, and Generative AI. arXiv preprint arXiv:2404.12590 (2024)

  21. [29]

    CopyrightCatcher [n. d.]. Introducing CopyrightCatcher, the first Copyright Detection API for LLMs. Accessed: March 22, 2024. https://www.patronus.ai/blog/introducing-copyright-catcher

  22. [30]

    Council of the European Union. 2024. Article 53: Obligations for Providers of General-Purpose AI Models. https: //artificialintelligenceact.eu/article/53/

  23. [31]

    Council of the European Union. 2024. Proposal for a Regulation of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) and amending certain Union legislative acts. https://digital-strategy .ec.europa.e...

  24. [32]

    cptX [n. d.]. cptX. https://github .com/maxim-saplin/cptX

  25. [33]

    Carys J Craig. 2024. THE AI-Copyright Trap. A vailable at SSRN, https:// papers.ssrn.com/ sol3/ papers.cfm?abstract_id = 4905118 (2024)

  26. [34]

    GitHub [n

    DOE 1 et al v. GitHub [n. d.]. DOE 1 et al v. GitHub. 4:22-cv-06823, (N.D. Cal.)

  27. [35]

    Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. 2024. Do Membership Inference Attacks Work on Large Language Models? arXiv preprint arXiv:2402.07841 (2024)

  28. [36]

    Christof Ebert and Panos Louridas. 2023. Generative AI for software practitioners. IEEE Software 40, 4 (2023), 30–38

  29. [37]

    Emilio Ferrara. 2023. Fairness and Bias in Artificial Intelligence: A Brief Survey of Sources, Impacts, and Mitigation Strategies. Sci 6, 1 (2023), 3

  30. [38]

    figma [n. d.]. Figma. https://www .figma.com/

  31. [39]

    Anna Filippova and Hichang Cho. 2016. The effects and antecedents of conflict in free and open source software development. In Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work & Social Computing . 705–716

  32. [40]

    Lothar Fritsch, Aws Jaber, and Anis Yazidi. 2022. An overview of artificial intelligence used in malware. InSymposium of the Norwegian AI Society . Springer, 41–51

  33. [41]

    Gemini Team et al. 2023. Gemini: A Family of Highly Capable Multimodal Models. arXiv:2312.11805 [cs.CL]

  34. [42]

    Generative Artificial Intelligence and Copyright Law [n. d.]. Generative Artificial Intelligence and Copyright Law. https://crsreports.congress.gov/product/pdf/LSB/LSB10922

  35. [43]

    Gervais, Noam Shemtov, Haralambos Marmanis, and Catherine Zaller Rowland

    Daniel J. Gervais, Noam Shemtov, Haralambos Marmanis, and Catherine Zaller Rowland. 2024. The Heart of the Matter: Copyright, AI Training, and LLMs. A vailable at SSRN, https:// papers.ssrn.com/ sol3/ papers.cfm?abstract_id =4963711 (2024)

  36. [44]

    GitHub. [n. d.]. GitHub Copilot. Retrieved March 22, 2024, from https://copilot .github.com

  37. [45]

    GitHub REST API documentation [n. d.]. GitHub REST API documentation. https://docs .github.com/en/rest

  38. [46]

    Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Stamos, and Michael Tomz. 2024. How persuasive is AI-generated propaganda? PNAS Nexus 3, 2 (2024), pgae034

  39. [47]

    gpt-pilot [n. d.]. gpt-pilot. https://github .com/Pythagora-io/gpt-pilot

  40. [48]

    gptengineer [n. d.]. gpt-engineer. https://github .com/gpt-engineer-org/gpt-engineer

  41. [49]

    grammarly [n. d.]. Transforming How the World Communicates Through AI. https://www.grammarly.com/ai

  42. [50]

    Groves, Floyd J

    Robert M. Groves, Floyd J. Fowler Jr., Mick P. Couper, James M. Lepkowski, Eleanor Singer, and Roger Tourangeau

  43. [51]

    Andres Guadamuz. 2024. A Scanner Darkly: Copyright Liability and Exceptions in Artificial Intelligence Inputs and Outputs. GRUR International 73, 2 (2024), 111–127

  44. [52]

    Maanak Gupta, CharanKumar Akiri, Kshitiz Aryal, Eli Parker, and Lopamudra Praharaj. 2023. From ChatGPT to ThreatGPT: Impact of Generative AI in Cybersecurity and Privacy. IEEE Access (2023). ACM Trans. Softw. Eng. Methodol., Vol. 1, No. 1, Article 1. Publication date: January ...

  45. [53]

    Runzhi He, Hao He, Yuxia Zhang, and Minghui Zhou. 2023. Automating dependency updates in practice: An exploratory study on github dependabot. IEEE Transactions on Software Engineering 49, 8 (2023), 4004–4022

  46. [54]

    Peter Henderson, Jieru Hu, Mona Diab, and Joelle Pineau. 2024. Rethinking Machine Learning Benchmarks in the Context of Professional Codes of Conduct. In Proceedings of the Symposium on Computer Science and Law . 109–120

  47. [55]

    Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A Lemley, and Percy Liang. 2023. Foundation Models and Fair Use. arXiv preprint arXiv:2303.15715 (2023)

  48. [56]

    Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, et al. 2023. Metagpt: Meta programming for multi-agent collaborative framework. arXiv preprint arXiv:2308.00352 (2023)

  49. [57]

    Stephen Hood. 2023. llamafile: bringing LLMs to the people, and to your own computer. https://future.mozilla.org/ builders/news_insights/introducing-llamafile/. Accessed: 2024-09-11

  50. [58]

    Hongsheng Hu, Zoran Salcic, Lichao Sun, Gillian Dobbie, Philip S Yu, and Xuyun Zhang. 2022. Membership Inference Attacks on Machine Learning: A Survey. ACM Computing Surveys (CSUR) 54, 11s (2022), 1–37

  51. [59]

    Yu Huang, Denae Ford, and Thomas Zimmermann. 2021. Leaving my fingerprints: Motivations and challenges of contributing to OSS for social good. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 1020–1032

  52. [60]

    HuggingChat [n. d.]. HuggingChat. https://huggingface .co/chat/. Accessed: 2024-08-11

  53. [61]

    Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A Choquette-Choo, and Nicholas Carlini. 2023. Preventing Generation of Verbatim Memorization in Language Models Gives a False Sense of Privacy. In Proceedings of the 16th ...

  54. [62]

    Jing Jiang, David Lo, Xinyu Ma, Fuli Feng, and Li Zhang. 2017. Understanding inactive yet available assignees in GitHub. Information and Software Technology 91 (2017), 44–55

  55. [63]

    Mitchell Joblin, Sven Apel, Claus Hunsen, and Wolfgang Mauerer. 2017. Classifying developers into core and peripheral: An empirical study on count and network metrics. In 2017 IEEE/ACM 39th International Conference on Software Engineering (ICSE). IEEE, 164–174

  56. [64]

    Theodoros Karathanasis. 2023. EU Copyright Directive: A ‘Nightmare’ for Generative AI Researchers and Developers? https://ai-regulation.com/eu-copyright-directive-a-nightmare-for-gai/

  57. [65]

    Kitchenham and Shari Lawrence Pfleeger

    Barbara A. Kitchenham and Shari Lawrence Pfleeger. 2002. Principles of Survey Research Part 2: Designing a Survey. ACM SIGSOFT Software Engineering Notes 27, 1 (2002), 18–20

  58. [66]

    Kitchenham and Shari Lawrence Pfleeger

    Barbara A. Kitchenham and Shari Lawrence Pfleeger. 2002. Principles of Survey Research: Part 3: Constructing a Survey Instrument. ACM SIGSOFT Software Engineering Notes 27, 2 (2002), 20–24

  59. [67]

    Kitchenham and Shari Lawrence Pfleeger

    Barbara A. Kitchenham and Shari Lawrence Pfleeger. 2002. Principles of Survey Research Part 4: Questionnaire Evaluation. ACM SIGSOFT Software Engineering Notes 27, 3 (2002), 20–23

  60. [68]

    Kitchenham and Shari Lawrence Pfleeger

    Barbara A. Kitchenham and Shari Lawrence Pfleeger. 2002. Principles of Survey Research: Part 5: Populations and Samples. ACM SIGSOFT Software Engineering Notes 27, 5 (2002), 17–20

  61. [69]

    Kitchenham and Shari Lawrence Pfleeger

    Barbara A. Kitchenham and Shari Lawrence Pfleeger. 2003. Principles of Survey Research Part 6: Data Analysis. ACM SIGSOFT Software Engineering Notes 28, 2 (2003), 24–27

  62. [70]

    Paul Krill. [n. d.]. GitHub survey finds nearly all developers using AI coding tools. https://www .infoworld.com/ article/3489925/github-survey-finds-nearly-all-developers-using-ai-coding-tools .html

  63. [71]

    Logan Kugler. 2024. Who Owns AI’s Output? Commun. ACM (2024). https://api .semanticscholar.org/CorpusID: 273132452

  64. [72]

    Jose Antonio Lanz. [n. d.]. AI Art Wars: Japan Says AI Model Training Doesn’t Violate Copyright. https://finance.yahoo.com/news/ai-art-wars-japan-says-185350499.html

  65. [73]

    Feder Cooper, and James Grimmelmann

    Katherine Lee, A. Feder Cooper, and James Grimmelmann. 2023. Talkin’ ’Bout AI Generation: Copyright and the Generative-AI Supply Chain. arXiv preprint arXiv:2309.08133 (2023)

  66. [74]

    Feder Cooper, James Grimmelmann, and Daphne Ippolito

    Katherine Lee, A. Feder Cooper, James Grimmelmann, and Daphne Ippolito. 2023. The Devil Is in the Training Data. https://genlaw.org/explainers/training-data.html

  67. [75]

    Lemley and Bryan Casey

    Mark A. Lemley and Bryan Casey. 2020. Fair Learning. Tex. L. Rev. 99 (2020), 743–785

  68. [76]

    Jiawei Li and Iftekhar Ahmed. 2023. Commit message matters: Investigating impact and evolution of commit message quality. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 806–817

  69. [77]

    Jenny T Liang, Chenyang Yang, and Brad A Myers. 2024. A large-scale survey on the usability of AI programming assistants: Successes and challenges. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering. 1–13

  70. [78]

    Jenny T Liang, Thomas Zimmermann, and Denae Ford. 2022. Understanding skills for OSS communities on GitHub. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 170–182. ACM Trans. Softw. Eng. M...

  71. [79]

    Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Kailong Wang

  72. [80]

    Llama.cpp [n. d.]. Llama.cpp. https://github .com/ggerganov/llama.cpp. Accessed: 2024-08-11

  73. [81]

    Nicola Lucchi. 2023. ChatGPT: A Case Study on Copyright Challenges for Generative Artificial Intelligence Systems. European Journal of Risk Regulation (2023), 1–23

  74. [82]

    Lijia Ma, Xingchen Xu, and Yong Tan. 2024. Crafting Knowledge: Exploring the Creative Mechanisms of Chat-Based Search Engines. arXiv preprint arXiv:2402.19421 (2024)

  75. [83]

    Vahid Majdinasab, Michael Joshua Bishop, Shawn Rasheed, Arghavan Moradidakhel, Amjed Tahir, and Foutse Khomh

  76. [84]

    Vahid Majdinasab, Amin Nikanjam, and Foutse Khomh. 2024. Trained Without My Consent: Detecting Code Inclusion in Language Models Trained on Code. arXiv preprint arXiv:2402.09299 (2024)

  77. [85]

    Tom Malley. 2023. AI Have a Deal: Driver uses ChatGPT hack to get dealer to agree to sell new car for $1 in ‘legally binding deal’ in blow for AI rollout. https://www.the-sun.com/motors/9888857/driver-uses-ai-loophole-buy-new- car-1

  78. [86]

    In 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)

    Assessing the Security of GitHub Copilot’s Generated Code-A Targeted Replication Study. In 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 435–444

  79. [87]

    Shiona McCallum. 2023. ChatGPT banned in Italy over privacy concerns. https://www .bbc.com/news/technology- 65139406

  80. [88]

    Meta. 2024. Introducing Meta Llama 3: The most capable openly available LLM to date. https://ai.meta.com/blog/meta- llama-3/

  81. [89]

    Eira May. [n. d.]. Where developers feel AI coding tools are working—and where they’re missing the mark. https://stackoverflow.blog/2024/09/23/where-developers-feel-ai-coding-tools-are-working-and-where-they- re-missing-the-mark/

  82. [90]

    Joao Pedro Moraes, Ivanilton Polato, Igor Wiese, Filipe Saraiva, and Gustavo Pinto. 2021. From one to hundreds: multi-licensing in the JavaScript ecosystem. Empirical Software Engineering 26 (2021), 1–29

  83. [91]

    Seth Neel and Peter Chang. 2023. Privacy Issues in Large Language Models: A Survey. arXiv preprint arXiv:2312.06717 (2023)

  84. [92]

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model Cards for Model Reporting. In Proceedings of the conference on fairness, accountability, and transparency. 220–229

  85. [93]

    United States Copyright Office. 2023. Artificial Intelligence and Copyright. https://www .federalregister.gov/ documents/2023/08/30/2023-18624/artificial-intelligence-and-copyright

  86. [94]

    United States Copyright Office. 2023. Copyright Registration Guidance: Works Containing Material Generated by Artificial Intelligence. 16190 Federal Register, Vol. 88, No. 51

  87. [95]

    Newsroom. 2023. WormGPT: New AI Tool Allows Cybercriminals to Launch Sophisticated Cyber Attacks. https: //thehackernews.com/2023/07/wormgpt-new-ai-tool-allows.html

  88. [96]

    United States Copyright Office. 2024. Copyright and Artificial Intelligence Part 1: Digital Replicas. https: //www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-1-Digital-Replicas-Report .pdf

  89. [97]

    United States Copyright Office. 2025. Copyright and Artificial Intelligence Part 2: Copyrightability. https: //www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report .pdf

  90. [98]

    United States Copyright Office. 2024. Artificial Intelligence Study. https://www .copyright.gov/policy/artificial- intelligence/

  91. [99]

    Oobabooga [n. d.]. Oobabooga Text Generation WebUI. https://github .com/oobabooga/text-generation-webui. Accessed: 2024-08-11

  92. [100]

    OpenAI. [n. d.]. ChatGPT https://openai.com/blog/chatgpt. Last accessed: March 2024

  93. [101]

    Ollama [n. d.]. Ollama. https://ollama .com/. Accessed: 2024-08-11

  94. [102]

    Abraham Naftali Oppenheim. 2000. Questionnaire design, interviewing and attitude measurement . Bloomsbury Publishing

  95. [103]

    Yin Minn Pa Pa, Shunsuke Tanizaki, Tetsui Kou, Michel Van Eeten, Katsunari Yoshioka, and Tsutomu Matsumoto

  96. [104]

    OpenAI. 2024. OpenAI safety update. https://openai .com/index/openai-safety-update/

  97. [105]

    Kitchenham

    Shari Lawrence Pfleeger and Barbara A. Kitchenham. 2001. Principles of Survey Research: Part 1: Turning Lemons into Lemonade. ACM SIGSOFT Software Engineering Notes 26, 6 (2001), 16–18

  98. [106]

    Li et al

    R. Li et al. 2023. StarCoder: may the source be with you. arXiv preprint arXiv:2305.06161 (2023). ACM Trans. Softw. Eng. Methodol., Vol. 1, No. 1, Article 1. Publication date: January 2025. Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for...

  99. [107]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning . PMLR, 28492–28518

  100. [108]

    Perplexity [n. d.]. What is Perplexity? https://www.perplexity.ai/hub/faq/what-is-perplexity. Accessed: 2024-08-11

  101. [109]

    Responsible AI. 2022. Big Science Open Rail-M License https://www.licenses.ai/blog/2022/8/26/bigscience-open-rail- m-license

  102. [110]

    Aresh Sarkari. 2024. Exploring Uncensored LLM Model – Dolphin 2.9 on Llama-3-8b. https://askaresh.com/2024/05/ 02/exploring-uncensored-llm-model-dolphin-2-9-on-llama-3-8b/

  103. [111]

    Sk Golam Saroar and Maleknaz Nayebi. 2023. Developers’ perception of GitHub Actions: A survey analysis. In Proceedings of the 27th International Conference on Evaluation and Assessment in Software Engineering . 121–130

  104. [112]

    Asha Rajbhoj, Akanksha Somase, Piyush Kulkarni, and Vinay Kulkarni. 2024. Accelerating Software Development Using Generative AI: ChatGPT Case Study. In Proceedings of the 17th Innovations in Software Engineering Conference . 1–11

  105. [113]

    June 10, 2023

    The Yomieru Shinbun. June 10, 2023. Intellectual Property Plan Signals Reversal on AI Policy. Japan News

  106. [114]

    Donna Spencer. 2009. Card sorting: Designing usable categories . Rosenfeld Media

  107. [115]

    Richard Stallman. [n. d.]. Why Upgrade to GPLv3. https://www .gnu.org/licenses/rms-why-gplv3

  108. [116]

    Agnia Sergeyuk, Yaroslav Golubev, Timofey Bryksin, and Iftekhar Ahmed. 2024. Using AI-Based Coding Assistants in Practice: State of Affairs, Perceptions, and Ways Forward. CoRR abs/2406.07765 (2024)

  109. [117]

    tabby [n. d.]. tabby. https://github .com/TabbyML/tabby

  110. [118]

    Tabnine [n. d.]. The AI code assistant you control. https://www .tabnine.com/. Accessed: 2024-25-10

  111. [119]

    FOSSA Editorial Team. 2021. Open Source Software Licenses 101: The AGPL License. https://fossa.com/blog/open- source-software-licenses-101-agpl-license/

  112. [120]

    Trevor Stalnaker, Nathan Wintersgill, Oscar Chaparro, Massimiliano Di Penta, Daniel M German, and Denys Poshy- vanyk. 2024. BOMs Away! Inside the Minds of Stakeholders: A Comprehensive Study of Bills of Materials for Software Systems. In Proceedings of the 46th IEEE/ACM Intern...

  113. [121]

    Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2024. Jailbroken: How does LLM safety training fail? Advances in Neural Information Processing Systems 36 (2024)

  114. [122]

    Whisper [n. d.]. Whisper. https://github .com/openai/whisper. Accessed: 2024-11-12

  115. [123]

    The Law Doesn’t Work Like a Computer

    Nathan Wintersgill, Trevor Stalnaker, Laura A. Heymann, Oscar Chaparro, and Denys Poshyvanyk. 2024. “The Law Doesn’t Work Like a Computer”: Exploring Software Licensing Issues Faced by Legal Practitioners. In Proceedings of the ACM on Software Engineering , Vol. 1. ACM New Yor...

  116. [124]

    Microsoft Corporation [n

    The New York Times Company v. Microsoft Corporation [n. d.]. The New York Times Company v. Microsoft Corpo- ration, No. 1:23-cv-11195 (S.D.N.Y., filed Dec. 27, 2023), https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_ Dec2023.pdf

  117. [125]

    XAI. 2024. Open Release of Grok-1. https://x .ai/blog/grok-os

  118. [126]

    Boming Xia, Tingting Bi, Zhenchang Xing, Qinghua Lu, and Liming Zhu. 2023. An Empirical Study on Software Bill of Materials: Where We Stand and the Road Ahead. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 2630–2642

  119. [127]

    Zhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi, Dongsun Kim, Donggyun Han, and David Lo. 2024. Unveiling Memorization in Code Models. In 2024 IEEE/ACM 46th International Conference on Software Engineering (ICSE) . IEEE Computer Society, 856–856

  120. [128]

    Scott Wu. 2024. Meet Devin: The World’s First AI Software Engineer. https://www.cognition-labs.com/introducing- devin

  121. [129]

    Yuanshun Yao, Xiaojun Xu, and Yang Liu. 2023. Large Language Model Unlearning. CoRR abs/2310.10683 (2023)

  122. [130]

    Rui-Jie Yew. 2024. Break It ’Til You Make It: An Exploration of the Ramifications of Copyright Liability Under a Pre-training Paradigm of AI Development. In Proceedings of the Symposium on Computer Science and Law . 64–72

  123. [131]

    Yang et al

    Z. Yang et al. 2023. Gotcha! This Model Uses My Code! Evaluating Membership Leakage Risks in Code Models. arXiv preprint arXiv:2310.01166 (2023)

  124. [132]

    Deborah Yao. 2023. One Year On, GitHub Copilot Adoption Soars. https://aibusiness.com/companies/one-year-on- github-copilot-adoption-soars. AI Business (27 6 2023). Accessed: March 22, 2024. https://aibusiness.com/companies/ one-year-on-github-copilot-adoption-soars

  125. [136]

    Sheng Zhang and Hui Li. 2023. Code Membership Inference for Detecting Unauthorized Data Use in Code Pre-trained Language Models. arXiv preprint arXiv:2312.07200 (2023). ACM Trans. Softw. Eng. Methodol., Vol. 1, No. 1, Article 1. Publication date: January 2025

  126. [2009]

    Survey Methodology, 2nd edition. Wiley

  127. [2023]

    In Proceedings of the 16th Cyber Security Experimentation and Test Workshop

    An attacker’s dream? exploring the capabilities of chatgpt for developing malware. In Proceedings of the 16th Cyber Security Experimentation and Test Workshop . 10–18

  128. [2024]

    In Proceedings of the 4th International Workshop on Software Engineering and AI for Data Quality in Cyber-Physical Systems/Internet of Things

    A Hitchhiker’s Guide to Jailbreaking ChatGPT via Prompt Engineering. In Proceedings of the 4th International Workshop on Software Engineering and AI for Data Quality in Cyber-Physical Systems/Internet of Things . 12–21

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.