{"id":"6ac97d08-3e9c-4581-8d46-060522d59957","arxiv_id":"2605.16706","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Of 118 AI policies in popular GitHub repos, 78% allow GenAI contributions, 51% require disclosure, and 74% require a human in the loop.","lead":"Researchers examined 1,000 popular GitHub projects and found 118 explicit AI contribution policies. Most allow AI-generated code if contributors disclose it and remain responsible, mapping how open source is adapting to generative AI.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The strongest claim is purely descriptive of the recovered sample. The recovery procedure is imperfect (keyword + manual filter), yet the percentages are conditioned on that sample and the paper supplies a public dataset, concrete policy excerpts, and clear RQ framing. The reader's weakest assumption is therefore a real methodological boundary, not a load-bearing flaw that would reverse or materially weaken the reported 78/51/74 figures. No further concern rises to the level of changing the ACCEPT verdict.","tokens_in":10938,"tokens_out":427,"duration_ms":3961,"concrete_test":"Independently re-run the keyword filter (LLM|generative AI|GenAI|AI|AI agents) on the same top-1,000 seart-ghs snapshot, manually re-label the resulting candidates into the three RQ1 buckets (welcome / permitted / discouraged), and check whether the 78%/22% split shifts by more than 5 percentage points; if it does not, the headline descriptive claim is stable under re-coding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is a transparent descriptive count over the 118 policies it actually recovered: 78% allow GenAI, 51% require disclosure, 74% require human-in-the-loop. The reader's weakest assumption correctly notes that keyword search over CONTRIBUTING.md / DEVELOPING.md / AI_POLICY.md (plus single-link external docs) can miss policies that use different wording or live only in issue templates, wikis, or CODE_OF_CONDUCT. That incompleteness is real, but it does not undermine the reported percentages among the recovered set; the paper never claims the 118 are an exhaustive census of all AI policies among the top-1,000 repos, only that these are the policies found by the stated procedure. Selection bias toward popular GitHub projects and unreported inter-rater reliability are standard limitations of the genre and are already acknowledged. No internal inconsistency or arithmetic error is visible in the three RQs.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This short empirical paper studies how popular open-source projects are adapting contribution guidelines to generative AI. From the top 1,000 starred GitHub repositories (filtered for activity and non-forks), the authors recover 118 AI policies via keyword search over CONTRIBUTING.md / DEVELOPING.md / dedicated AI_POLICY files (plus single-link external docs) and manual false-positive filtering. They report three descriptive findings among those 118 policies: 78% allow GenAI contributions (51% explicitly welcome, 27% permit without encouragement) while 22% discourage them; 51% require disclosure of AI assistance; and 74% require a human in the loop. The discussion highlights vague guidance language, disclosure practices (including commit trailers and multi-level AI involvement), and maintainer concerns about AI slop, with implications for practitioners and researchers. A public dataset is provided.","tokens_in":11173,"tokens_out":1236,"duration_ms":19438,"significance":"If the counts hold, the paper supplies a timely first snapshot of how contribution guidelines are responding to GenAI at scale. The three headline percentages, the welcome/permit/discourage split, and the extensive quoted examples (Tables II–IV) give maintainers and researchers a concrete baseline rather than anecdote. Strengths include an explicit search pipeline, manual review of 158 candidates, transparent category illustrations, and a released dataset. The work is observational and does not claim causal inference or exhaustiveness; its value is as an early map of policy practice and as a seed for follow-on studies of disclosure artifacts, AI slop, and autonomous-agent contributions.","major_comments":[{"comment":"Section II.B and Results (Table I, Findings 1–3): The central percentages rest on manual coding of 158 candidates into true/false positives and of the 118 policies into welcome/permitted/discouraged, disclosure-required, and human-in-the-loop. The manuscript does not report inter-rater reliability, a coding protocol, or whether multiple raters were used. Because these labels are the load-bearing results, the paper should either (a) report agreement metrics / dual coding on a sample or (b) make the single-coder process and decision rules fully explicit (e.g., decision tree for “welcome” vs “permitted,” and for borderline disclosure language such as “significant/substantial/meaningful”). Without that, reproducibility of the 78%/51%/74% figures is weaker than the rest of the design.","section":"II.B, III, Table I"},{"comment":"Section II.B (Detecting Repositories with AI Policy): The recovery procedure is keyword-driven over a fixed set of guideline files and single-link external docs. Policies that use different wording, live only in issue/PR templates, wikis, CODE_OF_CONDUCT, or community forums are invisible. The paper correctly reports percentages among the recovered 118 rather than among all 1,000 repos, but the abstract and Finding 1 can still be read as characterizing “open source AI policies” more broadly. A short, explicit bound—that 118 is a lower-bound recovered set under the stated procedure, with possible bias if restrictive policies use different language or locations—should appear in II.B and V so the central claim is not over-generalized.","section":"II.B, V, Abstract"}],"minor_comments":[{"comment":"Abstract vs body: the abstract collapses “welcome” and “permitted” into “allow AI-assisted contributions,” while Table I and RQ1 keep the three-way split. Align the abstract wording with the body (or briefly note the split) so readers are not surprised by Table I.","section":"Abstract, Table I"},{"comment":"Section II.A: selection requires “at least one commit in 2026.” Given the arXiv date this is coherent, but a one-sentence justification (e.g., ensuring ongoing maintenance at snapshot time) would help readers outside the snapshot window.","section":"II.A"},{"comment":"Section IV.A: the observation that many policies use vague terms (“significant,” “substantial,” “meaningful”) is useful; consider quantifying how often such hedges appear among the 60 disclosure-requiring policies, even if only as a rough count.","section":"IV.A"},{"comment":"Threats (V) is brief. Besides popularity and platform scope, briefly note single-coder risk and keyword incompleteness (even if expanded under major comments) so the limitations section matches the design.","section":"V"},{"comment":"Tables II–IV are strong; ensure every quoted policy is linked to a stable commit or archive URL in the dataset so examples remain checkable if upstream files change.","section":"Tables II–IV, [19]"},{"comment":"Minor wording: “Overall,wefindthatthe majorityoftheanalyzedAIpolicies” (Introduction, after RQ3) needs spacing; a few similar run-ons appear in the camera-ready text.","section":"Introduction"}],"recommendation":"minor_revision","confidential_remarks":"Solid short empirical paper for a venue that values timely MSR-style snapshots. The two major points (coding reliability and bounding the recovered set) are fixable without new data collection and do not undermine the descriptive core. I would not block on them if the authors add a short reliability/protocol note and a clearer lower-bound framing. Fit for a short paper / early-results track is good; not a full-length theory or causal study."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a straightforward, useful first empirical snapshot. The new result is the three counts over the 118 policies they actually recovered from the top-1,000 starred GitHub repos: 78% allow GenAI contributions (51% welcome + 27% permit), 22% discourage, 51% require disclosure, and 74% require a human in the loop. No prior work I know of has quantified those three dimensions across a large sample of contribution guidelines. The method is transparent—keyword search over CONTRIBUTING/DEVELOPING/AI_POLICY files plus single-link external docs, then manual filtering of 158 candidates—and they ship the dataset. Tables II–IV give enough real quotes that you can see how the categories were applied. The discussion of vague language (“significant,” “substantial”), disclosure trailers, and AI-slop countermeasures is practical for maintainers and sets up follow-on work.\n\nSoft spots are real but proportionate to a short descriptive paper. The search can miss policies that use different wording or live only in issue templates, wikis, or codes of conduct; the paper never claims the 118 are an exhaustive census of the top-1,000, only what the stated procedure found. Popular-GitHub selection bias and unreported inter-rater reliability are standard limitations of the genre and already noted. No arithmetic or internal inconsistency shows up in the three RQs. Circularity burden is essentially zero—pure observational tallies.\n\nThis is for people who write or study contribution guidelines, maintainers drafting AI policies, and researchers tracking GenAI’s effect on open source. It does not rewrite SE theory; it supplies a reusable baseline. I would bring it to reading group as a short, concrete data point. A serious editor should send it to peer review rather than desk-reject; the evidence is sharp enough for referee time even if revisions tighten the completeness discussion. I would cite the percentages and the dataset if I were writing on AI governance or contribution guidelines in the next year.","headline":"Clean first map of GenAI contribution policies in popular OSS: 78% allow, 51% require disclosure, 74% require human-in-the-loop, with a public 118-policy dataset.","tokens_in":11665,"tokens_out":512,"would_cite":true,"duration_ms":7803,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Most open-source AI policies allow GenAI contributions, but demand disclosure and a human in the loop.","keywords":["AI policy","generative AI","contribution guidelines","AI disclosure","human in the loop","open source","GitHub","AI slop"],"falsifier":"A re-scan of the same 1,000 repositories that also examines issue templates, wikis, codes of conduct, and non-standard filenames, or that uses broader synonym lists, would show whether many more (or fewer) policies exist and whether the 78 / 51 / 74 percentages hold.","tokens_in":11867,"feed_emoji":"⚙️","tokens_out":606,"duration_ms":5251,"temperature":0.7,"pith_summary":"Open-source projects are writing explicit rules for generative-AI help in pull requests and issues. From the top 1,000 starred GitHub repositories the authors recovered 118 such AI policies and classified them. Nearly four in five allow AI-assisted work, yet more than half require the contributor to say so, and three-quarters insist that a human still owns, understands, and can defend the change. The paper treats these three requirements—permission, disclosure, and human oversight—as the practical core of how maintainers are adapting contribution guidelines to the sudden flood of AI-generated code. Readers who maintain or contribute to open source can use the numbers and examples as a snapshot of current community norms and as a checklist for writing clearer policies of their own.","feed_headline":"78% of open-source AI policies allow GenAI code","feed_subtitle":"But half demand disclosure and three-quarters keep a human responsible for every PR","key_machinery":"The classification of each recovered AI policy along three binary axes—stance toward GenAI (welcome / permitted / discouraged), disclosure required or not, and human-in-the-loop required or not—derived from contribution-guideline and AI_POLICY files in the top 1,000 starred repositories.","core_discovery":"Among 118 AI policies found in popular GitHub repositories, 78 percent allow contributions generated with generative AI (51 percent welcome it, 27 percent merely permit it) while 22 percent explicitly discourage or forbid it; 51 percent require disclosure of AI assistance; and 74 percent require a human in the loop who remains responsible for the submitted code.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["78% of GitHub AI policies allow GenAI code contributions","51% of open-source AI policies require GenAI disclosure","74% keep a human responsible for every GenAI PR","Most AI policies permit GenAI but demand disclosure and humans","22% of open-source projects ban or discourage GenAI code"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Keyword search of standard guideline filenames for a short list of AI-related terms, followed by manual review of the hits, is assumed to recover essentially all of the AI policies that exist among those 1,000 repositories.","fun_headline_variants_meta":{"raw":{"variants":["78% of GitHub AI policies allow GenAI code contributions","51% of open-source AI policies require GenAI disclosure","74% keep a human responsible for every GenAI PR","Most AI policies permit GenAI but demand disclosure and humans","22% of open-source projects ban or discourage GenAI code"]},"model":"grok-4.5","effort":"low","cost_usd":0.00338,"raw_usage":{"total_tokens":1070,"prompt_tokens":760,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":33800000,"prompt_tokens_details":{"text_tokens":760,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":242,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":760,"tokens_out":68,"duration_ms":2510,"temperature":1.0,"reasoning_tokens":242,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T11:10:30.995400+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A re-scan of the same 1,000 repositories that also examines issue templates, wikis, codes of conduct, and non-standard filenames, or that uses broader synonym lists, would show whether many more (or fewer) policies exist and whether the 78 / 51 / 74 percentages hold.","supporting_citations":[],"review_version":2}