{"id":"d430a718-0a96-4e62-9dc5-78b83b0daa0d","arxiv_id":"2603.27130","paper_version":3,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In real-world repositories, AI-assisted and human-written code differ only modestly on code-level metrics, while commit size, stability, duplication, and language-specific security show clearer patterns.","lead":"This paper measures AI-assisted code versus human-written code at large scale in real-world repositories, across code structure, style, security, and commit behavior. It reports that real-world AI–human differences on code-level metrics are smaller than lab studies suggested, and adds first measurements of duplication, commit size, and post-commit stability.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Central claim rests on unverifiable AI-vs-human labeling of real-world code; manuscript text is too corrupted to audit the detector or bias controls.","rationale":"The Reader correctly identifies labeling reliability as the weakest assumption and correctly notes that the corrupted manuscript prevents auditing methods, effect sizes, and artifacts, yielding LOW confidence and a CONDITIONAL verdict. No stronger internal inconsistency is visible from the abstract and garbled body; the concern is empirical validity of the measurement, not circular math. The concrete test above is the minimal check that would settle whether the “small real-world gap” is an artifact of noisy or biased attribution. Until that check (or equivalent public data/code) is available, the verdict should remain CONDITIONAL rather than ACCEPT or REJECT. No formal verification or parameter-free derivation is claimed, so none offsets the labeling risk.","tokens_in":17088,"tokens_out":519,"duration_ms":7120,"concrete_test":"Recover a clean PDF/source of the paper and the public labeling pipeline (or re-implement the stated detector). On a held-out set of ≥200 commits with independent provenance (e.g., Copilot/Cursor telemetry or developer self-report), compute precision/recall and re-run the main code-level tables after (a) removing low-confidence labels and (b) treating human-edited AI code as a third class. If effect sizes reverse or grow beyond the paper’s “rather small” characterization, the central contrast with lab settings fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that real-world AI–human differences on code-level metrics are small (unlike lab settings), plus first large-scale commit/duplication results. Every comparative result requires reliable labels of which code is AI-assisted vs human-written. The supplied full text is largely unreadable (garbled tokens, mixed physics arXiv stamp), so the attribution method, ground-truth validation, false-positive/negative rates, and handling of human-edited AI snippets cannot be checked. If the detector systematically tags only stereotypically “AI-looking” fragments, or if humans heavily revise AI output before commit, measured gaps shrink by construction and the “small real-world difference” conclusion does not follow. Labeling validity is therefore load-bearing for the entire measurement study; without recoverable methodology and public attribution artifacts, the claim remains untestable from the given materials.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper claims to present the first large-scale measurement of AI-assisted versus human-written code in real-world repositories, comparing a broad suite of code-level metrics (structural/graph complexity, coding style, security quality, duplication) and commit-level metrics (size, frequency, post-commit stability). From the abstract and recoverable framing, the central empirical claim is that real-world AI–human differences on code-level metrics are rather small—contrasting more pronounced gaps reported in lab/synthetic settings—while also reporting first large-scale observations on duplication and commit dynamics, with language-dependent security variation and a discussion of practical implications for AI-assisted programming.","tokens_in":17228,"tokens_out":826,"duration_ms":8515,"significance":"If the labeling and sampling are sound, this would be a high-value empirical contribution for software engineering: it moves AI-code evaluation from lab benchmarks to production repositories and supplies multi-metric, multi-language evidence that could recalibrate expectations about complexity, security, and maintenance cost of AI-assisted code. The breadth of metrics (code + commit) and the explicit contrast with lab findings are strengths in principle. However, the supplied manuscript body is severely corrupted (encoding damage, unreadable token streams, and an apparent mix with unrelated physics.app-ph material), so the claimed scale, methods, and results cannot be audited from the materials provided. Significance therefore remains conditional on a recoverable, verifiable manuscript.","major_comments":[{"comment":"Load-bearing labeling validity cannot be assessed. Every comparative claim (small real-world AI–human gaps on complexity, style, security, duplication, and commit metrics) requires a reliable method for attributing code/commits as AI-assisted vs human-written. The garbled full text does not recover the detector/provenance pipeline, ground-truth validation, precision/recall or FPR/FNR, thresholds, or treatment of human-edited AI snippets. If labels systematically tag only stereotypically AI-looking fragments or miss heavily revised AI output, measured gaps shrink by construction and the central contrast with lab settings does not follow. This must be fully specified, validated, and preferably released as artifacts.","section":null},{"comment":"Manuscript integrity failure blocks scientific review. Large portions of the body are unreadable (garbled tokens, broken equations/tables, and an arXiv stamp for physics.app-ph 2603.27133v2 mixed into the stream). Sections that should contain study design, sampling filters, metric definitions, statistical tests, and result tables cannot be reconstructed. Without a clean, complete manuscript, the claimed large-scale measurements and the “small real-world difference” conclusion are not verifiable; resubmission of an intact PDF/source is required before any accept/reject decision on the science.","section":null},{"comment":"Sampling and confounding controls are not recoverable. Free parameters noted in the design (repository/commit filters, metric aggregation, outlier handling) and potential confounds (project maturity, language mix, developer skill, tool adoption era) are load-bearing for generalizing “real-world” results. The corrupted text does not allow checking whether AI-labeled and human-labeled cohorts are matched or whether differences are explained by non-AI factors. These design choices must be stated with sensitivity analyses.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":"The supplied full-text stream is not a usable manuscript (severe encoding corruption plus apparent contamination by unrelated physics content). I cannot confirm that the authors’ actual PDF matches the abstract’s claims. Recommend the editor require a clean source PDF and, if possible, public attribution labels/code before external review proceeds. Scope is appropriate for an empirical SE venue if methods are recoverable; novelty of “first large-scale real-world” should be checked against concurrent AI-code measurement work once the text is readable. Confidence in any scientific judgment is low until the body is intact."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing to know: the abstract sells a first large-scale real-world comparison of AI-assisted versus human code across code-level and commit-level metrics, and the headline result—that real-world code-level gaps are small, unlike lab work—would be a useful baseline for SE and AI-coding tools. The second thing: the full text we were given is largely unreadable (garbled tokens, mixed physics arXiv stamp), so we cannot check methods, effect sizes, or artifacts from the materials in hand.\n\nWhat is actually new, if the study is as framed, is the scale and the metric suite: structural/graph complexity, style, security quality with language variance, plus under-covered pieces like duplication rate and commit size/frequency/post-commit stability, all in real repositories rather than synthetic tasks. That package is not a restatement of lab benchmarks. The design is a straightforward comparative measurement study; circularity risk is the usual SE kind (metrics and detectors encoding the intuitions you hope to measure), not mathematical self-fitting.\n\nThe soft spot is load-bearing and not minor. Every comparative result depends on reliable AI-vs-human labels in the wild. If the detector tags only stereotypically AI-looking fragments, or if humans heavily edit AI output before commit, measured gaps shrink by construction and the “small real-world difference” claim does not follow. Sampling filters, classifier cutoffs, and aggregation choices are free parameters. From the garbled text we cannot audit ground-truth validation, FP/FN rates, or bias controls. That is a data/methods audit problem, not a reason to dismiss the research question.\n\nWho this is for: empirical SE and AI-for-code people who need real-repo baselines, not theory. A serious editor should send a clean version to referees—the question and the claimed contrast with lab results are important enough—provided the authors ship recoverable methodology, public attribution artifacts, and bias checks. I would not cite from the abstract alone. Bring a cleaned PDF and data release to reading group; until then, treat the central claim as untestable from what we have.","headline":"Useful real-world AI-vs-human code measurement on paper, but the supplied manuscript is too corrupted to audit the labeling that carries every claim.","tokens_in":17925,"tokens_out":534,"would_cite":false,"duration_ms":11381,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"In real-world repositories, AI-assisted code differs only modestly from human-written code on structure, style, and security—unlike the larger gaps reported in lab studies.","keywords":["AI-generated code","LLM-assisted programming","real-world repositories","code quality metrics","security quality","commit analysis","software measurement","code duplication"],"falsifier":"Re-run the same metric suite on the same repositories with an independent, higher-precision provenance or detection method for AI-assisted code; if the code-level AI–human gaps become large and consistent, or if commit-level patterns reverse, the “small real-world difference” claim fails.","tokens_in":17919,"feed_emoji":"💻","tokens_out":641,"duration_ms":10640,"temperature":0.7,"pith_summary":"This paper argues that the true picture of AI-generated code only emerges when it is measured inside actual production repositories, not only on synthetic lab benchmarks. The authors run a large-scale comparison of AI-assisted versus human-written code across both code-level properties (complexity, style, security, duplication) and commit-level behavior (size, frequency, and how stable the change stays after it lands). Their central result is that real-world differences on code-level metrics are small, which contrasts with earlier lab findings that painted AI code as more sharply distinct. They also report new measurements—duplication rates, commit sizes, and post-commit stability—and finer language-by-language variation in security quality. A sympathetic reader cares because these numbers shape how teams should review, test, and govern AI-assisted contributions once they are already mixed into live codebases.","feed_headline":"Real-world AI code barely differs from human code","feed_subtitle":"Large-scale repo study finds small gaps on complexity, style, and security—unlike lab benchmarks","key_machinery":"A dual-level measurement design that attributes code and commits as AI-assisted or human-written in real repositories, then compares them on a fixed suite of code-level metrics (structure, graph complexity, style, security, duplication) and commit-level metrics (size, frequency, post-commit stability), including language-stratified cuts.","core_discovery":"When AI-assisted and human-written code are measured side by side in real-world repositories, differences on code-level metrics such as structural and graph complexity, coding style, and security quality are rather small, in contrast to more pronounced gaps seen in laboratory settings; the study further supplies first large-scale evidence on code duplication and on commit size, frequency, and post-commit stability.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Real-world AI code barely differs from human code on complexity","AI-human code gaps shrink in actual repositories, not labs","Tiny real-repo differences between AI and human code metrics","AI-assisted code mirrors human code on style and security in the wild","Large-scale study finds real AI code almost matches human code"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The study’s labels that mark which code and commits are AI-assisted versus human-written in real repositories must be accurate enough that mislabeling does not shrink or invent the measured gaps.","fun_headline_variants_meta":{"raw":{"variants":["Real-world AI code barely differs from human code on complexity","AI-human code gaps shrink in actual repositories, not labs","Tiny real-repo differences between AI and human code metrics","AI-assisted code mirrors human code on style and security in the wild","Large-scale study finds real AI code almost matches human code"]},"model":"grok-4.5","effort":"low","cost_usd":0.005196,"raw_usage":{"total_tokens":1486,"prompt_tokens":834,"num_sources_used":0,"completion_tokens":89,"cost_in_usd_ticks":51960000,"prompt_tokens_details":{"text_tokens":834,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":563,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":834,"tokens_out":89,"duration_ms":6765,"temperature":1.0,"reasoning_tokens":563,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T17:12:49.367391+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same metric suite on the same repositories with an independent, higher-precision provenance or detection method for AI-assisted code; if the code-level AI–human gaps become large and consistent, or if commit-level patterns reverse, the “small real-world difference” claim fails.","supporting_citations":[],"review_version":2}