{"id":"c95c3b67-b7b9-4774-8e7c-30ab42a57268","arxiv_id":"2506.17833","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In a 40-person practitioner survey, AI tools were reported to help most with small code snippets and to degrade architecture quality when applied to large, complex problems.","lead":"Using a 40-person survey and a one-developer Tetris experiment, this paper reports that AI tools raise self-rated productivity but hurt solution quality on large, complex problems. The headline productivity claim is not backed by statistical tests, and the survey's 35% benchmark comes from the authors' own experiment.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Paper's own counts contradict the headline: only 18/40 (45%) reported productivity above the 35% reference, not an 'overwhelming majority'; the central productivity claim is internally unsupported.","rationale":"I read the paper in good faith: it provides a clear GQM structure, a transparent description of a small survey, and it makes raw data available. The qualitative findings about decomposition, integration effort, and architectural erosion are plausible and consistent with prior work. However, the central quantitative productivity claim can be checked directly against the paper's own reported numbers, and the check fails. The reader's weakest assumption pointed to self-reported productivity and the absence of objective measures or a control group; I agree with that, but I found an even more immediate problem: the paper's internal counts contradict the 'overwhelming majority' and 'over 35%' phrasing. Only 45% reported above 35%, so the abstract's strong conclusion is not supported even by the authors' own descriptive results. This is not a disagreement with external consensus; it is an internal inconsistency in the reported evidence. A single re-analysis of the raw survey responses with an exact binomial test would settle whether the headline claim can be salvaged. Since the central claim fails this check, the REJECT verdict is appropriate; the paper could become acceptable if reframed as a descriptive perception study without the 'significant' and 'overwhelming' language, and with proper statistical reporting.","tokens_in":11862,"tokens_out":5049,"duration_ms":52268,"concrete_test":"Using the raw survey data linked in Section 4.1 (github.com/AItoolsSE/CAIN-AI-documents), reproduce the counts in Figure 6 for the self-reported productivity question. Compute an exact binomial two-sided 95% confidence interval and a one-sided exact binomial test for the proportion of participants who selected 'Much higher' or 'Higher' (i.e., improvement over the 35% reference) against a null proportion of 0.5. With the reported counts of 18/40, the 95% CI is approximately 29% to 62% and the one-sided p-value is about 0.75, so the data do not support an 'overwhelming majority' or a statistically significant majority claiming >35% productivity gain. If the raw data confirm these counts, the abstract's 'significantly increase' claim loses its quantitative basis; if the counts differ, the discrepancy itself warrants correction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that AI tools 'significantly increase productivity' rests on a specific quantitative assertion: in the introduction the authors call it an 'overwhelming majority' reporting a productivity increase 'over 35%', and the conclusion repeats that 'the increase in productivity of over 35% is a clear and strong sign'. The survey's own reported data, Figure 6 and Section 4.3(i), give 1 'much higher', 17 'higher', 14 'around 35%', 7 'lower', and 1 'zero' out of N=40. Thus only 18/40 (45%) reported an increase over the 35% reference; 22/40 (55%) did not. Calling 45% an 'overwhelming majority' is internally inconsistent, not merely a sampling concern. Furthermore, that 35% reference value comes from the authors' own single-engineer Tetris substudy (Section 3.2, M3.1), and participants were asked to rate their improvement relative to that value, so the scale is anchored to an n=1 measurement. No confidence interval or significance test is reported anywhere; 'significant' is used colloquially. The related claim that AI users 'outperform those who do not use such tools' is also unsupported because the survey sampled only AI-tool users and contains no comparison group. The most load-bearing weakness is that the headline quantitative result is contradicted by the paper's own reported frequencies, so the abstract's conclusion does not follow from the presented data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a two-part study of AI tools in software engineering: a single-engineer, six-week Tetris project comparing development with and without AI assistance, and a survey of 40 practitioners who already use AI tools. The authors conclude that AI tools significantly increase engineer productivity (quantified as 'over 35%'), that productivity gains diminish as project complexity grows, and that adopting AI-generated code snippets does not significantly erode software architecture, while large AI-generated blocks are of significantly lower quality. The paper argues that non-users of AI tools will become less competitive and that architects remain necessary for problem decomposition and solution integration.","tokens_in":12112,"tokens_out":4757,"duration_ms":41753,"significance":"If the claims were supported, the paper would provide field evidence on the productivity and architectural effects of AI-assisted development, a topic of active interest. The study has strengths: the GQM structure is systematic, the raw survey results and Tetris code are made publicly available, and the distinction between small snippets and large AI-generated blocks is a useful hypothesis. However, the central quantitative productivity claim is contradicted by the paper's own reported frequencies, and the survey design—a convenience sample of 40 self-selected AI users with self-reported productivity measures and no control group—cannot support the causal and comparative statements made in the abstract and conclusion. The architecture claims are inferred from majority agreement without statistical tests. The paper is therefore best seen as an exploratory descriptive study whose headline conclusions overstate the evidence.","major_comments":[{"comment":"The abstract and introduction state that an 'overwhelming majority' of participants reported a productivity increase 'over 35%', but the paper's own data in Figure 6 and Section 4.3(i) show only 18 of 40 respondents (45%) chose 'higher' or 'much higher' relative to the 35% reference; 14 chose 'around 35%', 7 'lower', and 1 'zero'. Thus 55% of respondents did not report an increase over 35%. The conclusion in Section 5 that 'the increase in productivity of over 35% is a clear and strong sign' is therefore not supported by the presented frequencies. Moreover, no confidence interval, effect size, or significance test is reported anywhere; the word 'significant' is used colloquially rather than statistically. The central quantitative claim of the paper is internally inconsistent with its own data.","section":"Section 4.3(i) and Figure 6"},{"comment":"The survey's productivity scale is anchored to a reference value of 'up to 35% improvement' taken from the authors' own Tetris substudy, which is a single-engineer, six-week project (Section 3.2). Participants were asked to rate their gains relative to that specific value, and the same 35% value is then used in Section 4.3(i) and Section 5 as the threshold for 'over 35%' increases. This makes the quantitative productivity headline partially circular: the instrument is calibrated with the same n=1 measurement that is later used as the benchmark for the claim. The 35% figure is also presented as a finding of the substudy without any uncertainty quantification, so it cannot serve as a reliable anchor for survey interpretation.","section":"Section 3.2 (M3.1) and Section 4.2 (M2.1)"},{"comment":"The introduction and conclusion claim that individuals using AI tools 'significantly outperform those who do not use such tools'. However, the survey (Section 3.3) recruited only professionals who already use AI tools; there is no comparison group of non-users, and no objective productivity measure. The only comparative evidence comes from the Tetris substudy, which involves a single engineer and three versions, without repeated measures or statistical analysis. These claims are therefore beyond what the study design can support.","section":"Section 3.3 and Section 5"},{"comment":"The claim that adopting AI-generated snippets 'does NOT lead to a significant erosion of software architecture' is based on agreement counts: in Figure 10-b, 25 of 40 respondents (62.5%) agree or strongly agree that snippets are of comparable maintenance quality, while 11 are neutral and 4 disagree. This is a majority opinion, but it is not a statistical demonstration of 'no significant negative influence'; the absence of a negative effect cannot be inferred from majority agreement without a hypothesis test or a confidence interval on the proportion. The same issue applies to Figures 10-c and 11-b. The abstract's wording ('no significant negative influences') overstates the evidence.","section":"Section 4.3(ii) and Figure 10-b"}],"minor_comments":[{"comment":"The text says '79% of the surveyed participants experience the same, higher or much higher productivity,' but the counts in Figure 6 (1+17+14=32) give 32/40 = 80%, not 79%.","section":"Section 4.2"},{"comment":"The text states that 'over 35% of the surveyed participants' consider integration not to be a challenge, but Figure 9-c shows 1 strongly agree and 10 agree (11/40 = 27.5%), not over 35%.","section":"Section 4.2, Figure 9-c"},{"comment":"There are typos and style issues: 'summirized' should be 'summarized', and 'Inversion a)' appears three times where 'version a)' is meant.","section":"Section 3.2"},{"comment":"Reference [1] is incomplete: it lists 'Openai and ashley pilipiszyn.' with a URL but no publication year or proper author formatting.","section":"References"}],"recommendation":"reject","confidential_remarks":"The paper's raw data and code availability are commendable, but the internal inconsistency between the 'overwhelming majority' claim and the reported 45% count is a serious correctness issue that cannot be fixed by presentation changes alone; the survey design also cannot support the comparative claims about non-users. I recommend rejection, though a revised manuscript that reframes the results as exploratory and softens the causal language could be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take on arXiv:2506.17833 (Amasanti and Jahić). The paper is a survey of 40 AI-tool-using practitioners plus a three-way Tetris building exercise. The raw data and code are on GitHub, and the survey results are presented in full figures. That transparency is real credit.\n\nWhat is actually new: the dataset itself and the specific 40-person sample. The conceptual findings — AI helps early-stage boilerplate and small snippets, degrades on large complex tasks, architects are needed for decomposition and integration — are consistent with prior work (including Waseem et al., which they cite). So as a measurement of practitioner perceptions, it is a modest but legitimate addition.\n\nThe soft spots are where the reader's report lands, and the stress test is right. The abstract and conclusion claim AI tools \"significantly increase productivity\" and that an \"overwhelming majority\" saw over 35% increase. Their own Figure 6 gives 1 much higher + 17 higher = 18/40 (45%), and 14 around 35%, 7 lower, 1 zero. So 55% did not report over 35%. Calling that an overwhelming majority is internally inconsistent, not just a sampling issue. They also have no control group, so the jump from \"users report gains\" to \"non-users are left behind\" is unsupported. The 35% reference value comes from their own single-engineer Tetris experiment, and the survey question was anchored to that value, so the quantitative headline is partly built into the instrument. No confidence intervals or significance tests appear anywhere; \"significant\" is used informally.\n\nThat said, the paper is not a mess. The qualitative claims about decomposition and integration are well supported by the open-ended responses, and the authors are transparent about the 21% who saw lower/no gains and note they could not find correlating factors. The limitation is not in the data collection but in the interpretation: the paper overreaches from a descriptive self-report survey to causal productivity claims.\n\nRecommendation: I would send this to peer review rather than desk reject — it's a real dataset, timely topic, and the flaws are fixable with a rewrite that honestly reports 45% and drops the \"significant/overwhelming\" language. A serious referee could push it to a decent descriptive study. Right now, the abstract should not be trusted as written.","headline":"The paper has a real survey dataset and sensible qualitative findings, but its headline claim of an 'overwhelming majority' seeing over 35% productivity gains is contradicted by its own reported counts.","tokens_in":12645,"tokens_out":2059,"would_cite":false,"duration_ms":19566,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI tools raise software engineers' productivity, but only when problems are broken into small, well-scoped pieces.","keywords":["software architecture","AI","productivity","quality","architecture erosion","survey study","AI-generated code","software engineering"],"falsifier":"Conduct a controlled experiment with two matched teams on the same greenfield and maintenance tasks, one using AI tools and one not, measuring actual cycle time, defect rates, and architectural metrics; if the AI team's measured productivity gain is not significantly above zero, the paper's central claim would not survive.","tokens_in":11603,"feed_emoji":"🤖","tokens_out":5894,"duration_ms":55810,"temperature":0.7,"pith_summary":"This paper asks whether AI coding tools deliver real productivity gains without quietly degrading software architecture. Based on a survey of 40 practitioners who already use AI tools, the authors conclude that AI tools significantly increase productivity—most respondents reported gains at or above 35%—and that adopting small AI-generated code snippets does not seriously erode architectural quality. The benefits, however, shrink as problems grow: large AI-generated blocks are harder to maintain, more likely to break existing code, and judged of significantly lower quality than human-written code. The practical conclusion is that AI pays off when engineers decompose problems into small prompts and integrate the results themselves.","feed_headline":"AI lifts coding productivity by a third — if problems stay small","feed_subtitle":"Survey of 40 practitioners finds small AI snippets preserve code quality, while large AI blocks make code harder to maintain.","key_machinery":"The organizing mechanism is a size-and-complexity boundary rather than a single identity. The study's machinery is the Goal-Question-Metric (GQM) framework used to structure both the Tetris substudy and the survey, combined with grounded-theory open and axial coding of free-text responses. These tools generate the paper's central contrast: small, focused AI-generated snippets preserve cohesion and maintainability, while large AI-generated blocks degrade logical organization. The 35% productivity reference value comes from one engineer's Tetris substudy and anchors the survey's ordinal productivity scale.","core_discovery":"The paper's central claim is that AI assistance is a net productivity positive in software engineering, but only under decomposition: the same tools that generate excellent small snippets produce structurally weak large blocks. The authors arrive at this from a two-part study: a three-week Tetris project built in three parallel versions (no AI, AI pair, multi-agent AI setup), which produced the 35% productivity reference, and a survey of 40 industry practitioners. On self-reported productivity, 79% of respondents matched or exceeded the 35% reference, and 45% said their gains exceeded 35%. On architecture, the survey found no significant erosion of cohesion, coupling, or logical partitioning when adopting small AI-generated snippets, but clear agreement that large AI-generated blocks are of significantly lesser quality, harder to maintain, and more prone to breaking existing code. The authors read this as evidence that architects remain essential: their job is to decompose problems and integrate AI-generated solutions.","pith_inferences":["The paper does not claim, but its data suggest, that productivity studies should measure greenfield and maintenance work separately, since AI appears to help the former and hinder the latter.","A controlled experiment with matched teams could test whether the self-reported 35% gain survives objective measurement; the paper's own Tetris substudy is a single-engineer data point, so the survey reference value is itself an assumption to verify.","The survey's 'no architectural erosion for snippets' finding is about perceived cohesion and coupling, so an automated code-quality analysis on real repositories would be a natural next test.","The reported tendency of AI to write tests that pass existing code rather than validate requirements implies that test generation quality, not just code quality, may degrade as problem size grows."],"forward_implications":["Engineers who use AI tools can expect the largest productivity gains in early-stage work such as generating skeletons, boilerplate, and basic features, where most respondents reported gains matching or exceeding 35%.","AI-generated code is not a uniform category: small snippets are comparable to human code in cohesion and coupling, so teams can adopt them without significant architectural erosion.","Large AI-generated blocks and AI-driven changes to established code bases carry real costs, including lower maintainability, more frequent breaking of existing code, and longer processing times, so they need human architectural review.","Since 21% of respondents saw no productivity improvement, AI benefits are not automatic; prompt skill and problem decomposition mediate outcomes.","The skills that matter shift toward decomposition and integration: architects who break problems into AI-sized pieces and integrate the results are the ones who capture the productivity gains."],"supporting_citations":[{"why":"It supplies the adoption baseline (76.6% of developers using or planning to use AI) that motivates the study.","marker":"[24]"},{"why":"It justifies the difficulty of measuring productivity in software engineering and the study's choice of self-reported measures.","marker":"[3]"},{"why":"It provides the Goal-Question-Metric approach used to structure both substudies' design.","marker":"[23]"},{"why":"It supplies the methodology for collecting valid software engineering data that GQM builds on.","marker":"[4]"},{"why":"It provides the survey-design guidelines followed when constructing the industry questionnaire.","marker":"[15]"},{"why":"It is the earlier case study of ChatGPT as a software engineering tool whose reported strengths and limits frame the survey questions.","marker":"[27]"},{"why":"It documents a recent high-performing AI debugger result (98.2% Pass@1) cited as evidence of the productivity potential of current tools.","marker":"[30]"},{"why":"It establishes the state of practice for LLMs in software architecture, motivating the architecture-erosion research question.","marker":"[12]"}],"fun_headline_variants":["AI boosts coder productivity, but only on small tasks","Survey: AI coding help shines on snippets, fails on big blocks","AI code tools: big speed gains, but quality drops with size","AI assistance: productivity up, but complex problems need humans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The productivity conclusion rests on taking the self-reported 'subjective feeling of productivity' of 40 self-selected practitioners who already use AI tools as a valid measure of real productivity, with no control group and no objective output metric.","fun_headline_variants_meta":{"raw":{"variants":["AI boosts coder productivity, but only on small tasks","Survey: AI coding help shines on snippets, fails on big blocks","AI code tools: big speed gains, but quality drops with size","AI assistance: productivity up, but complex problems need humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1435,"prompt_tokens":897,"completion_tokens":538,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":467}},"tokens_in":513,"tokens_out":538,"duration_ms":4979,"temperature":1.0,"reasoning_tokens":467,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:59:50.314998+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a controlled experiment with two matched teams on the same greenfield and maintenance tasks, one using AI tools and one not, measuring actual cycle time, defect rates, and architectural metrics; if the AI team's measured productivity gain is not significantly above zero, the paper's central claim would not survive.","supporting_citations":[{"cited_title":"https://stackoverflow.co/labs/2024-developer-survey- insights-for-ai-ml/ (2024), accessed: 08/09/2024","cited_arxiv_id":null,"evidence_quote":"It supplies the adoption baseline (76.6% of developers using or planning to use AI) that motivates the study."},{"cited_title":"In: 2010 Fifth International Conference on Software Engineering Advances","cited_arxiv_id":null,"evidence_quote":"It justifies the difficulty of measuring productivity in software engineering and the study's choice of self-reported measures."},{"cited_title":"IEEE Transactions on Software EngineeringSE-10(6), 728–738 (1984)","cited_arxiv_id":null,"evidence_quote":"It supplies the methodology for collecting valid software engineering data that GQM builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the survey-design guidelines followed when constructing the industry questionnaire."},{"cited_title":"ChatGPT as a Software Development Bot: A Project-based Study","cited_arxiv_id":"2310.13648","evidence_quote":"It is the earlier case study of ChatGPT as a software engineering tool whose reported strengths and limits frame the survey questions."}],"review_version":1}