{"id":"cda3cd5c-1086-4bb0-bc11-09351e091b03","arxiv_id":"2411.10877","paper_version":5,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A plurality of developers would place AI-generated code in the public domain, most see it as similar to reusing existing code, and few document AI usage or have copyright training.","lead":"A survey of 574 software developers maps how they use AI coding tools and what they think about copyright, licensing, and ownership of AI-generated code. The findings give regulators and companies a real-world snapshot of developer practices and concerns as AI copyright law is still forming.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified.","rationale":"The paper's central contribution is a descriptive account of what 574 self-selected developers reported. The survey instrument, response processing, and quantitative summaries are transparent, and the replication package supports verification. The reader's identified weakest assumption, sampling representativeness, is the only plausible threat to broader claims, and it is disclosed and mitigated by the paper's careful wording: the abstract says 'a survey of 574 developers' and 'a snapshot of developers' views,' and Section 7.3 explicitly disclaims generalizability. I also checked for internal issues: the ownership question allowed multiple selections, but the paper reports per-option percentages and acknowledges overlap in Finding 22; the confidence statistic is explicitly self-reported; and the seven interviews are used only as supplementary context. No missing proof or circular reasoning appears in the empirical chain from responses to findings. The proposed concrete test would usefully quantify the non-response bias and could strengthen the external-validity discussion, but it is not necessary for the paper's stated descriptive claims. Therefore the ACCEPT verdict stands without modification.","tokens_in":40527,"tokens_out":10359,"duration_ms":108791,"concrete_test":"Using the replication package, compare the 574 respondents against the 30,659 invitees on GitHub-observable covariates (e.g., account age, number of public repositories, follower count, number of starred or forked GenAI repositories) and re-weight the headline ownership percentages using propensity score weights. If the reweighted 44% public-domain and 36% prompter/employer figures shift by more than 5 percentage points, the generalizing language in the abstract and conclusions should be explicitly conditioned on the sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No significant objection identified. The central claim is explicitly about the 574 surveyed respondents, and the reported percentages (44% public domain, 36% prompter/employer, 75% confident) are directly supported by the survey data. The weakest point is external validity: the sample is a self-selected subset of GitHub users interested in GenAI repositories, with a 1.9% response rate from 30,659 invitees, so the numeric percentages may not generalize to all developers. This limitation is acknowledged in Sections 7.1 and 7.3, and the paper does not claim statistical generalizability. Because the descriptive claims are not undermined by the acknowledged sampling limits, no load-bearing concern changes the verdict.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a mixed-methods study of 574 developers (plus 7 follow-up interviews) recruited from GitHub users who forked, starred, watched, or contributed to 30 GenAI-related repositories. The survey asks about developers' use of generative AI tools, their perceptions of copyright and licensing issues, and other legal concerns. The central descriptive results are that developers hold varied views on ownership of AI-generated code (44% choosing 'public domain,' 36% choosing the prompter or employer, 28% choosing training-data creators, 9% choosing model creators), that 75% are confident or very confident in their ownership opinions despite unresolved law, and that developers are concerned about data leakage and lack of documentation/process for GenAI usage. The paper presents 29 findings organized by three research questions, discusses implications for policy and practice, and makes its survey and analysis artifacts available in an online replication package.","tokens_in":40609,"tokens_out":6636,"duration_ms":67192,"significance":"The study is timely and policy-relevant, directly responding to calls from the U.S. Copyright Office for stakeholder perspectives. Its strengths include a detailed survey design process, transparent reporting of instruments and qualitative coding, and explicit acknowledgment of sampling limitations. The descriptive claims are directly supported by the data as presented: the percentages, denominators, and analysis procedures are internally consistent. The main substantive limitation is external validity: the sample is self-selected from GitHub users who showed interest in GenAI repositories, and the 1.9% response rate limits generalizability beyond this population. The authors appropriately disclaim statistical generalization in Section 7.3, but the abstract and introductory framing could be more careful. If the reported percentages are interpreted as describing only this surveyed population, the paper provides a valuable snapshot for lawmakers, regulators, and organizations.","major_comments":[{"comment":"The survey instrument (Figure 2, U11) lists 'prompt creator' as a single response option, but the results in Figure 9 and Section 5.4.2 report a combined category 'Belongs to prompter (or their employer)' with 201 selections (36%). The paper should clarify how 'employer' responses were derived, including whether this category merges the pre-specified 'prompt creator' option with 'Other' responses, and what the individual counts were. As written, this headline percentage conflates two conceptually distinct views (individual prompt authorship versus employer ownership), which directly affects the interpretation of the central ownership-distribution claim.","section":"Section 5.4, Figure 9"}],"minor_comments":[{"comment":"The opening results sentence states that 'Our results show the benefits developers derive from GenAI...' without explicit qualification to the surveyed sample; consider adding a phrase such as 'among the surveyed developers' to avoid overgeneralization, given the sampling limitations acknowledged in Section 7.3.","section":"Abstract"},{"comment":"The description of the qualitative coding process is dense; summarizing the final codebook structure and the specific criteria used to merge codes would improve reproducibility and reader comprehension.","section":"Section 3.3"},{"comment":"Figure 5 displays counts of documentation types but does not show the denominator (n=83) in the caption; adding 'n=83' and clarifying that multiple selections were possible would prevent reader confusion about the base for the reported percentages.","section":"Section 4.2.2, Figure 5"},{"comment":"The legal background is thorough but U.S.-centric; an early sentence noting that the analysis is grounded in U.S. law would be helpful because the study sample is international and Section 7.3 later acknowledges this limitation.","section":"Section 2.1"}],"recommendation":"minor_revision","confidential_remarks":"The paper is methodologically sound for a descriptive survey, and the central claims are supported by the data as presented. The only substantive issue I see is the need to clarify how the 'prompter or employer' category in Figure 9 was constructed from the survey's response options; this is a local fix that does not undermine the overall contribution. I do not see any concerns about novelty or citation practices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the key takeaway: this is the first large-scale descriptive snapshot of developer opinions on copyright and licensing for AI-generated code, and the paper delivers exactly that. The survey of 574 developers, with follow-up interviews, is clearly described, the analysis is careful, and the authors are appropriately modest about what they can claim. The central findings, such as the split between those who think AI output belongs to no one (44%) and those who think it belongs to the prompter or employer (36%), are directly supported by the data they report. That is real value for policymakers and companies trying to understand how developers think in a legally unsettled area.\n\nWhat the paper does well: the survey design is transparent and grounded in both legal scholarship and SE survey best practices; the open-ended coding is done by two annotators with reconciliation; a replication package is available; and the threats-to-validity section is unusually honest. Notably, the authors explicitly state they do not claim statistical generalizability. That honesty is worth respecting.\n\nThe soft spots, in proportion: the sample is the big one. Participants were self-selected from GitHub users who forked, starred, watched, or contributed to 30 GenAI-related repos, and the response rate was 1.9% (574 of 30,659). That means the percentages likely overrepresent open-source enthusiasts and GenAI users, and underrepresent enterprise and closed-source developers. The paper acknowledges this, and because the claims are phrased as reported views of these respondents, the descriptive conclusions hold. But anyone citing the 44% figure as a general population estimate would overread it. Minor: there is no inter-rater reliability metric, though the authors explain why they chose an inductive approach, and the follow-up interviews number only seven, so they are illustrative at best.\n\nWho is this for? SE researchers studying GenAI adoption, legal scholars and policymakers working on copyright regulation, and companies drafting AI usage policies. It deserves a serious referee; the methodology is solid for what it claims, and the topic is timely. The main thing I would ask of the authors is to tighten the abstract so readers cannot walk away thinking this sample represents all developers. In a review, I would accept with minor revisions.","headline":"A transparent, well-scoped descriptive survey of developer views on GenAI copyright that earns its keep despite acknowledged sampling limits.","tokens_in":41118,"tokens_out":1497,"would_cite":true,"duration_ms":15850,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 574 developers finds that the most common view is that AI-generated code belongs to no one, while the second-largest group credits the prompter or employer, amid unresolved law.","keywords":["generative AI","copyright","software licensing","developer survey","AI-generated code ownership","open-source software","legal perceptions","code generation"],"falsifier":"A replication survey drawn from a representative sample of developers outside GenAI-related open-source repositories—for example, a random sample of enterprise developers in proprietary settings—that found a different ownership distribution (say, a majority choosing the employer) would undercut the generalizing reading of the 44% and 36% figures, though it would not contradict the paper's description of its own respondents.","tokens_in":40368,"feed_emoji":"⚖️","tokens_out":6145,"duration_ms":60760,"temperature":0.7,"pith_summary":"This paper reports a descriptive study of 574 software developers who use generative-AI coding tools, supplemented by seven follow-up interviews. It aims to capture, at a moment of legal uncertainty, what developers think about copyright and licensing for AI-generated code: who should own it, whether training on code should require permission or compensation, and what risks they worry about. The headline result is that opinion is split and confident: 44% of respondents say AI-generated code belongs to no one, 36% say it belongs to the prompter or their employer, and about three-quarters report being confident or very confident in their answer even though courts and regulators have not settled the question. The authors present the study as a resource for policymakers, arguing that regulation of this area should be informed by the views of the people who actually use the tools.","feed_headline":"44% of developers say AI code belongs to no one","feed_subtitle":"A 574-developer survey finds confident but conflicting views on copyright while the law is still unsettled.","key_machinery":"The machinery that carries the study is the survey instrument paired with qualitative coding. The questionnaire has 26 closed and 7 open questions in three sections—current GenAI use, understanding and perception of copyright issues, and demographics—and the open-ended answers were coded independently by two researchers using an inductive open-coding process, with disagreements resolved by discussion. That combination produces the paper's result: a structured snapshot of percentages plus quotes that explain the reasoning behind them. The single question doing the most work is U11, the multiple-choice ownership question, together with its confidence and rationale follow-ups.","core_discovery":"The central claim is a descriptive one: developers' opinions on the copyright status of AI-generated code are diverse, context-dependent, and often held with high confidence despite unresolved law. When asked to select all that applied, the most common answer was that generated code belongs to no one and sits in the public domain (242 of 554 respondents, 44%), followed by the view that it belongs to the prompting developer or their employer (201, 36%); smaller groups assigned ownership to training-data creators (156, 28%), model creators (49, 9%), or said they did not know (57, 10%). The paper also finds that developers mostly treat AI-generated code as similar to other reused code, are indifferent or pleased when their prompts are reused, are context-dependent about whether their own code may be used in training data, and largely work in organizations with no formal process for documenting AI use, while only 12% of respondents reported any formal copyright training.","pith_inferences":["Because the sample was recruited through GenAI-related open-source repositories, the 44% public-domain figure may overstate the view among enterprise and closed-source developers; a broader sample could shift the balance toward employer ownership.","The high confidence paired with unresolved law suggests that many developers are forming ownership intuitions from tool behavior and community norms rather than from legal sources, which may make later judicial or regulatory rulings harder to absorb.","If AI agents begin producing whole applications with less human editing, the 'tools are tools' reasoning that supports prompter ownership may weaken; the paper anticipates this tension but does not test it.","The documentation gap reported by developers implies that any provenance or disclosure mandate would require new tooling and workflows, not just new rules; the paper stops short of designing such mechanisms."],"forward_implications":["Organizations cannot assume a single coherent developer view on ownership; internal policies on AI-generated code will have to manage expectations that range from public domain to employer ownership.","Terms-of-service reading is the exception, not the rule, so contractual terms about output ownership are unlikely to reach most developers.","Most developers report no organizational process for documenting AI-generated code, which complicates any future compliance or provenance requirements.","Policies on training data will need to distinguish open-source from proprietary code, since developers' reactions depend heavily on that context.","If courts assign ownership one way or the other, a large share of current developers will be surprised, suggesting a need for guidance that connects legal outcomes to existing developer intuitions."],"supporting_citations":[{"why":"The closest prior large-scale survey of AI programming assistants, whose usability findings motivate this study's focus on developer experience.","marker":"[77]"},{"why":"Supplies the qualitative open-coding method used to analyze the survey's open-ended responses.","marker":"[114]"},{"why":"The public call for comments that frames the study as a resource for copyright policymaking.","marker":"[93]"},{"why":"Maps the generative-AI supply chain and copyright questions, supporting the paper's point that ownership answers may be case-specific.","marker":"[73]"},{"why":"The pending lawsuit alleging license violations from training on open-source code, a central piece of the unresolved legal backdrop.","marker":"[34]"},{"why":"The fair-use precedent that respondents' views on training data are interpreted against.","marker":"[12]"},{"why":"The European Union's AI regulation, used to show the non-U.S. legal landscape the authors compare their findings with.","marker":"[31]"}],"fun_headline_variants":["44% of devs say AI code belongs to no one","Devs' top pick for AI code ownership: no one","Survey: Devs confident on AI copyright, but views vary","AI code copyright: 44% see public domain, 36% see themselves","Most devs treat AI code like reused code, survey finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that developers who interacted with GenAI-related open-source repositories and then chose to answer the survey are a meaningful window onto 'developers' views' generally; if that pool overrepresents GenAI enthusiasts and open-source contributors, the headline percentages may not generalize.","fun_headline_variants_meta":{"raw":{"variants":["44% of devs say AI code belongs to no one","Devs' top pick for AI code ownership: no one","Survey: Devs confident on AI copyright, but views vary","AI code copyright: 44% see public domain, 36% see themselves","Most devs treat AI code like reused code, survey finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000737,"raw_usage":{"total_tokens":3286,"prompt_tokens":930,"completion_tokens":2356,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":2280}},"tokens_in":546,"tokens_out":2356,"duration_ms":18668,"temperature":1.0,"reasoning_tokens":2280,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:11:15.588856+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication survey drawn from a representative sample of developers outside GenAI-related open-source repositories—for example, a random sample of enterprise developers in proprietary settings—that found a different ownership distribution (say, a majority choosing the employer) would undercut the generalizing reading of the 44% and 36% figures, though it would not contradict the paper's description of its own respondents.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The closest prior large-scale survey of AI programming assistants, whose usability findings motivate this study's focus on developer experience."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The public call for comments that frames the study as a resource for copyright policymaking."}],"review_version":1}