{"id":"ce0ad675-7a1d-4ba6-9a6b-56af73fa9745","arxiv_id":"2607.13156","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a controlled industrial experiment with 49 developers, design-system-aware AI assistance reduced front-end implementation time by 46.7–69.4% and raised task completeness from 68% to 96% compared with manual and design-system-only workflows.","lead":"This paper runs a controlled industrial experiment at a Brazilian tech company comparing front-end developers who work manually, with a design system, or with a design-system-aware AI assistant. It reports that the AI group finished screens 46–69% faster with higher completeness, though the underlying statistics and data are not fully reported.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central significance claim is not checkable as written: the manual group is tiny and first-cycle-only while AI/DS pool both cycles, and no group-equivalence or statistical-test details are reported.","rationale":"The reader's weakest_assumption is group equivalence, and I agree that this is the most load-bearing concern. The paper's central claim is a causal one: AI assistance causes large productivity gains. For that to hold, the groups must be comparable except for the intervention. The paper gives no evidence of this, and the design creates a structural confound: the manual baseline is much smaller and was run only in the first cycle, while the AI/DS analyses pool both cycles. Any unmeasured difference between cycles (e.g., task familiarity, tool improvements, participant motivation) would be entangled with the manual-vs-AI comparison. The absence of any inferential statistics makes it impossible to assess whether the reported p<0.05 is credible. This is not a matter of disagreeing with the consensus; it is a correctness risk internal to the study design. The concrete test I propose, re-running on Cycle 1 only, directly addresses the confound. If the effect survives, the concern is mitigated; if not, the central claim collapses. Since the reader's verdict is already CONDITIONAL and this concern is the same one, the verdict should remain unchanged, but the paper should be required to report the requested data and analysis before acceptance.","tokens_in":5395,"tokens_out":3618,"duration_ms":48459,"concrete_test":"Obtain the raw per-participant dataset and re-run the central comparison using only Cycle 1, where Manual, DS, and AI groups coexisted, stratifying by stack and seniority. If the AI-vs-manual difference is no longer statistically significant or the effect size drops materially from the reported 46.7–69.4% range, the pooled analysis is confounded by the cycle/group design.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 3.1) that AI-assisted development 'significantly reduced completion time' (p<0.05) rests on the three groups being comparable. The manuscript reports that participants were 'balanced by their technical stack and seniority level' (Section 2.1), but provides no random-assignment scheme, no per-stack cell sizes, and no baseline equivalence data (e.g., experience, prior tool familiarity). Critically, the Manual group had only 7 participants and ran only in the first cycle, whereas DS and AI groups pool both cycles (Sections 2.1 and 3.1). If the Manual group was disproportionately slower, or the AI group more experienced, the reported 46.7–69.4% reductions would reflect participant composition rather than the tool. The paper reports no test name, test statistic, degrees of freedom, or per-group sample sizes, so the p<0.05 claim cannot be independently checked. The Limitations section (4.3) does not mention this threat. This is the load-bearing assumption on which the central causal claim depends.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a controlled between-subjects experiment conducted at Zup Innovation (a large Brazilian enterprise) comparing three front-end development workflows: manual development, Design-System-only development, and DS-aware AI-assisted development, across Angular, iOS, and Android stacks. Based on two experimental cycles, the authors claim that AI assistance significantly reduced time-to-delivery by 46.7%–69.4% relative to manual development, increased task completeness from 68% (manual) to 85% (DS-only) to 96% (AI-assisted), and reduced performance variability. The paper also reports qualitative break-pattern analysis as evidence of reduced workflow friction and cognitive load. The central statistical claim (Section 3.1) is that AI-assisted development achieved p<0.05 compared with both manual and DS-only conditions, but the manuscript provides no test statistics, confidence intervals, effect sizes, or per-stack sample sizes.","tokens_in":5659,"tokens_out":3587,"duration_ms":43767,"significance":"If the findings withstand scrutiny, this would be a valuable industrial empirical contribution: it directly measures a DS-aware AI assistant in a real enterprise setting across multiple technology stacks, uses a controlled task with incremental submission checkpoints, and provides stack-specific lessons (Angular benefits from complexity reduction, iOS from variability smoothing, Android from incremental gains). The study has low circularity risk because the outcome is a direct measurement with no fitted parameters or outcome-defining equations. However, the central claim is currently not checkable from the reported statistics, and the group-equivalence assumption is insufficiently supported. The paper would be much stronger if it reported full inferential details, per-cell sample sizes, a Cycle-1-only analysis, and explicit discussion of the non-random assignment threat.","major_comments":[{"comment":"The central claim that AI assistance 'significantly reduced completion time' (p<0.05) cannot be independently verified because no test name, test statistic, degrees of freedom, exact p-value, confidence interval, effect size, or per-cell sample size is reported. The only p-value appears in prose. Please report the exact test used (e.g., Mann-Whitney U, Welch t), its statistic, exact p-value, effect size (e.g., Cliff's delta or Hedges' g), and per-stack per-group n. In addition, because the Manual group (n=7) ran only in Cycle 1 while DS and AI groups pool both cycles (Section 2.1), any pooled comparison confounds group with cycle. Provide a Cycle-1-only analysis and a cycle-equivalence check (e.g., DS/AI means by cycle).","section":"Section 3.1 and Table 3"},{"comment":"Group equivalence is load-bearing for the causal interpretation. The paper states participants were 'balanced by their technical stack and seniority level' but reports no random assignment procedure, no per-stack cell sizes, and no baseline measures (e.g., years of experience, prior familiarity with StackSpot AI or the DS). If assignment was not random, the observed 46.7%–69.4% reductions could reflect participant composition rather than the tool. Provide a demographic/experience table, the assignment mechanism, and equivalence tests between groups. This threat is not acknowledged in Section 4.3, which is a notable omission.","section":"Section 2.1"},{"comment":"Task completeness was assessed by specialists, but the manuscript gives no scoring rubric, number of reviewers, blinding procedure, or inter-rater reliability. Table 2 reports only group means (68%, 85%, 96%) with no variance, confidence interval, or test statistic. The claim of 'higher average completeness' is therefore unsupported as presented. Please specify how completeness was operationalized, whether multiple specialists independently rated the deliverables, and provide a statistical comparison (including per-stack and per-cycle breakdowns).","section":"Sections 2.3 and 3.1 (Table 2)"},{"comment":"Break patterns are used to infer reduced workflow friction and cognitive load, but no quantitative analysis is reported. The cited ranges overlap (AI 0–90 minutes vs. DS 60–110 minutes), and the claim of 'shorter and less frequent breaks' is not accompanied by any table, distribution, or test. If this secondary evidence is retained, report break frequency and duration summaries by group and stack, and clearly label this component as exploratory rather than confirmatory.","section":"Section 3.3"}],"minor_comments":[{"comment":"Tables 1 and 3 present the same productivity means and SDs. Consolidate them or cross-reference to avoid redundancy.","section":"Tables 1 and 3"},{"comment":"'R3 dataset' is mentioned without definition. Clarify whether R3 refers to a specific cycle, task section, or subgroup.","section":"Section 3.3"},{"comment":"The text contains a formatting artifact: 'followingRe' in 'investigates the followingRe' should read 'the following research question'.","section":"Section 1"},{"comment":"The experiment is described as 'controlled' but no random assignment is stated. Use 'quasi-experiment' or explicitly describe the assignment mechanism.","section":"Section 2.1"},{"comment":"The abstract reports percentages without indicating they are relative reductions. Clarify that 46.7%, 57.9%, and 69.4% are reductions relative to manual time-to-delivery.","section":"Abstract/Introduction"},{"comment":"Exact p-values should replace the single 'p<0.05' statement, and confidence intervals for the mean reductions would materially improve interpretability.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important industrial question and the direct measurement design is a strength, but the statistical reporting is currently insufficient to support the headline p<0.05 claim. The most urgent fixes are: full inferential statistics, per-cell sample sizes, a Cycle-1-only analysis, and a transparent discussion of the non-random assignment and small manual group. If these are addressed, the paper could make a solid SBES contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about this paper. First, the experimental design is genuinely new: a controlled three-condition comparison (manual, DS-only, DS-aware AI) run with 49 professional developers in a real company across Angular, iOS, and Android. That combination doesn't appear in the cited literature, which is dominated by case studies and reviews. Second, the headline claim — \"significantly reduced completion time (p<0.05)\" — cannot be checked from the paper as written. There are no test statistics, no per-stack sample sizes, no confidence intervals, and the only p-value appears in prose.\n\nWhat the paper does well: the task protocol is sensible (incremental submission, specialist review, timer that continues on returns), the break-pattern analysis is a nice indirect proxy for workflow friction, and the DS-only control condition helps separate the effect of the AI from the effect of the design system alone. The limitations section honestly acknowledges single-company and two-screen scope. The direction of results is consistent across all three stacks, which lends some credibility even with the statistical gaps.\n\nThe main soft spot is group comparability. The manual group had 7 participants and ran only in cycle 1; the DS and AI groups pool both cycles. The paper says participants were \"balanced by technical stack and seniority\" but gives no assignment scheme, no baseline equivalence data, and no per-cell sizes. If the manual group happened to be slower or less familiar with the mockup format, the 46–69% reductions would partly reflect that. This is the load-bearing assumption for the causal claim, and the Limitations section does not mention it. That said, this is a reporting problem, not a fatal design flaw. It can be fixed with per-cycle breakdowns, test details, effect sizes, and raw data. The two self-citations are used as background and don't drive the outcome, so that's minor.\n\nWho this is for: people studying AI-assisted development in industry and anyone designing similar controlled evaluations. A serious referee should see it, but the revision bar should be explicit about the missing statistical details.\n\nI'd send it to peer review, with a clear request for the missing reporting. If the authors supply that, the paper would be a solid empirical data point.","headline":"Genuinely new three-condition industrial comparison, but the headline p-value is unverifiable as written — a reporting problem, not a fatal one.","tokens_in":6103,"tokens_out":1946,"would_cite":true,"duration_ms":16759,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A design-system-aware AI assistant cut front-end task completion time by 46.7% to 69.4% in an industrial experiment with 49 professional developers.","keywords":["Design Systems","AI-assisted development","front-end productivity","time-to-delivery","task completeness","controlled experiment","Angular","iOS"],"falsifier":"Re-run the experiment with random assignment to condition, equal group sizes in every cycle, and a pre-test of developer proficiency on a comparable task; if the time reduction disappears or loses statistical significance after controlling for proficiency, the central claim collapses.","tokens_in":5304,"feed_emoji":"⚡","tokens_out":2584,"duration_ms":27170,"temperature":0.7,"pith_summary":"This paper tries to establish that AI tools which are explicitly aware of an enterprise design system substantially accelerate front-end development and improve design consistency, beyond what the design system alone achieves. It reports a controlled experiment at a large Brazilian company where developers implemented two high-fidelity screens across Angular, iOS, and Android, either manually, with a design system, or with a design-system-aware AI assistant. The central claim is that AI assistance reduced average time-to-delivery by roughly half to two-thirds, raised task completeness from 68% (manual) and 85% (design-system-only) to 96%, and shrank performance variability. A sympathetic reader should care because direct industrial evidence on AI plus design-system workflows has been scarce, and the results suggest concrete productivity and consistency gains for real-world front-end teams.","feed_headline":"AI with design-system context cuts front-end time by 46 to 69 percent","feed_subtitle":"Industrial experiment across Angular, iOS, and Android shows faster, more consistent delivery and higher fidelity.","key_machinery":"The central mechanism is a controlled industrial experiment with an incremental-submission protocol: developers submit each implemented screen section as they finish, and specialists review in real time, recording exact completion times and rejecting non-conforming submissions until correct. The intervention is a design-system-aware AI assistant—an internal tool given the enterprise design system as context—contrasted against manual work and design-system-only work. This setup is what lets the authors attribute differences in time, completeness, and variability to the AI tool rather than to differences in task difficulty.","core_discovery":"In a controlled between-subjects experiment conducted over two cycles with 49 professional developers, the use of a design-system-aware AI assistant significantly reduced completion time compared to both a manual baseline and design-system-only development (p<0.05). Average time-to-delivery fell from 536 to 164 minutes in Angular (69.4%), from 593 to 316 minutes in iOS (46.7%), and from 388 to 163 minutes in Android (57.9%); reductions against the design-system-only condition ranged from 15.5% to 24.4%. AI assistance also raised average task completeness to 96%, versus 85% with the design system alone and 68% manually, and reduced the standard deviation of completion times across all stacks.","pith_inferences":["The paper's results suggest that the key driver of the productivity gain is not generic AI code completion, but the embedding of organizational design knowledge into the tool's context; if true, similar gains might be replicable with any AI assistant that is given sufficient design-system context, not just the specific tool studied.","The measured time reductions include natural work pauses, so the reported 'time-to-delivery' may overstate pure coding speed gains; a more precise estimate of active development time could change the magnitude of the claimed effect.","The sharp reduction in break durations (from up to 195 minutes manually to at most 90 minutes with AI) is an indirect cognitive-load signal; directly measuring mental effort with physiological or self-report instruments would test the paper's implicit claim about reduced cognitive load.","The experiment used a single two-screen task, so the long-term effects on maintenance, onboarding, and design-governance compliance remain untested; a longitudinal study with varied task types would clarify whether the gains persist in production workflows."],"forward_implications":["If the central claim holds, organizations with enterprise design systems can expect measurable time-to-delivery reductions of roughly 15–25% over design-system-only workflows, and 46–69% over fully manual workflows, when adopting a context-aware AI assistant.","Engineering leaders can make stack-specific adoption decisions: the evidence suggests Angular benefits most from complexity reduction, iOS from variability smoothing, and Android from incremental gains on an already efficient baseline.","Task completeness rising to 96% supports the idea that design-system-aware AI can improve visual fidelity and reduce rework caused by non-compliant implementations.","Lower standard deviations in completion time imply more predictable delivery schedules, which matters for planning in regulated or large-portfolio environments.","The break-pattern results, if reliable, imply reduced workflow friction and potentially lower cognitive load during design-to-code translation, which could improve developer experience and sustainability."],"fun_headline_variants":["DS-aware AI cuts front-end time by 46-69% in industrial test","AI with design system context speeds dev by 46-69%","Design-system-aware AI lifts completeness to 96% and cuts time","Industrial experiment: DS-aware AI cuts dev time up to 69%","AI with DS context: 46-69% faster front-end delivery"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the three groups were equivalent in developer skill, motivation, and working conditions, because participants were balanced but not randomly assigned, and the manual baseline came from a smaller group that ran only in the first cycle.","fun_headline_variants_meta":{"raw":{"variants":["DS-aware AI cuts front-end time by 46-69% in industrial test","AI with design system context speeds dev by 46-69%","Design-system-aware AI lifts completeness to 96% and cuts time","Industrial experiment: DS-aware AI cuts dev time up to 69%","AI with DS context: 46-69% faster front-end delivery"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1153,"prompt_tokens":684,"completion_tokens":469,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":373}},"tokens_in":428,"tokens_out":469,"duration_ms":4711,"temperature":1.0,"reasoning_tokens":373,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:01:34.175199+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the experiment with random assignment to condition, equal group sizes in every cycle, and a pre-test of developer proficiency on a comparable task; if the time reduction disappears or loses statistical significance after controlling for proficiency, the central claim collapses.","supporting_citations":[],"review_version":1}