{"id":"c215b1cf-313d-4ffb-8798-bf6338b0a156","arxiv_id":"2411.17194","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a 160-participant experiment, structured designer prompts produced higher-rated AI-generated urban pocket garden images than freeform prompts, but group differences were only marginally significant.","lead":"This paper tests whether urban designers still matter when AI image generators create public space designs from text prompts. In a 160-person experiment, designs made with structured designer guidance scored higher on beauty and intent-alignment than fully free AI outputs, though the statistical evidence was weak.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Designer involvement is conflated with prompt structure and mask predefinition; the paper's own 2x2 design is analyzed as a one-factor 'freedom' variable, so the headline causal claim is underidentified.","rationale":"Good-faith reading: the paper is an exploratory human-AI collaboration study with a sensible toolkit (fine-tuned Stable Diffusion, knowledge graphs, public participation platform). The strongest claim, however, requires causal attribution to designer involvement. The study's design actually realizes a 2x2 factorial on prompt structure and mask structure (freeform/freeform, structured/freeform, freeform/structured, structured/structured), yet the analysis collapses to a three-level 'freedom' factor and one pairwise medium comparison. This leaves the causal interpretation unidentified: structured prompts, predefined masks, or their interaction could produce the observed differences even if human designers added nothing. The reader's weakest assumption flagged this same conflation; I agree. Two further issues (marginal p-values and no released data/code) reinforce but do not replace this concern. A verdict of UNVERDICTED is appropriate: the paper may still be a useful proof-of-concept, but its central conclusion about the role of designers cannot be accepted on the current evidence. The minimal corrective step is to analyze the existing 2x2 design and, ideally, add an arm where structured prompts are authored without designer input.","tokens_in":7914,"tokens_out":7070,"duration_ms":67809,"concrete_test":"Reanalyze the existing four-cell data as a 2x2 factorial ANOVA with PromptType (structured/freeform) and MaskType (structured/freeform) as crossed factors, testing main effects and interaction for both aesthetic and alignment scores; if the structured-prompt main effect is not significant, the medium-freedom contrast is not evidence for designer involvement.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim that 'designer involvement significantly enhances quality and alignment' is not identified by the manipulation. In 'Levels of Designer Involvement', the four experimental cells are: freeform prompt + freeform mask (high freedom), structured prompt + freeform mask (medium), freeform prompt + structured mask (medium), and structured prompt + structured mask (low). This is a complete 2x2 factorial on PromptType x MaskType, but the analysis collapses cells into an ordinal 'degree of freedom' factor. The only significant contrast, structured-prompt/freeform-mask vs freeform-prompt/structured-mask (p=.013/.026), differs in both factors simultaneously; no main effect or interaction is reported. The low-vs-high comparison, which most directly tracks 'designer involvement', is only marginal (p=.055 aesthetic, .069 alignment) and also changes both prompt structure and mask structure together. Consequently the abstract's 'designers significantly improve' is not supported: the data are equally consistent with structured prompts alone (or mask predefinition alone) improving outputs, independent of any human designer. There is also no arm in which a designer directly supervises generation or refinement beyond pre-authored scaffolds, so 'designer involvement' as a causal agent is never tested. A 2x2 factorial ANOVA on the existing cells would separate the components; until that is reported, the headline causal attribution cannot stand.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an experimental study of how varying levels of designer involvement affect the aesthetic quality and goal alignment of AIGC-generated urban pocket-garden images. Using a fine-tuned Stable Diffusion model with soft inpainting and a public participation WebUI, the authors recruited 160 participants and collected 1–10 aesthetic and alignment ratings under conditions that differ in prompt structure (freeform vs. structured) and mask predefinition (freeform vs. structured). The stated headline finding is that designer involvement significantly improves the quality and alignment of AIGC-generated images, with the main value of designers lying in providing structured prompts and predefined modification areas. The paper also reports secondary analyses comparing participants with and without design backgrounds, and discusses implications for the future role of urban designers in AI-assisted participatory planning.","tokens_in":8253,"tokens_out":2175,"duration_ms":21233,"significance":"If the central claim were adequately supported, the study would be a useful empirical contribution to human–AI collaboration in urban design and public participation, with practical implications for how AIGC tools should be integrated into participatory workflows. The study's strengths include a relatively large participant sample (160), a priori power analysis, a realistic experimental setup with multiple urban scenarios, and a concrete system (fine-tuned diffusion model plus WebUI) that others could adapt. However, the current statistical and design issues mean the headline causal claim is not established; the paper's contribution at this stage is more of a demonstrated proof-of-concept for a participatory AIGC pipeline than a rigorous test of the role of designers.","major_comments":[{"comment":"The abstract and the 'Results and discussion' section state that designer involvement 'significantly enhances the quality and alignment of AIGC-generated images,' but the reported inferential statistics do not support the word 'significantly.' The main effect of guidance degree on aesthetic scores has p = .064 and on alignment scores p = .075 (Table 1), and the Tukey HSD comparison between the low- and high-freedom groups yields p = .055 and .069. All of these exceed the conventional .05 threshold, and the manuscript itself uses 'marginally significant' in the Results section. The conclusions should be reworded to reflect marginal or suggestive evidence, or the authors must provide additional analyses that meet the stated significance criterion.","section":"Results and discussion (Impact of AIGC on the Role of Designers); Table 1"},{"comment":"The design is a 2x2 factorial (PromptType × MaskType), but the analysis collapses the four cells into an ordinal 'degree of freedom' factor. This aggregation conflates the two manipulations and prevents identification of which component (prompt structure or mask predefinition) drives the observed quality differences. The only significant contrast reported, 'structured prompt & freeform mask' versus 'freeform prompt & structured mask' (p = .013 and .026), differs in both factors simultaneously, so it cannot be attributed to either factor alone. A two-way ANOVA with PromptType and MaskType as factors, including their interaction, should be reported to separate the effects. Until this is done, the claim that 'structured prompts' (rather than mask predefinition, or some interaction) are the effective ingredient is not supported.","section":"Research Design (Levels of Designer Involvement); Result Analysis with different degrees of freedom"},{"comment":"Hypothesis 2 frames 'designer involvement' as the causal variable, but the manipulation operationalizes designer involvement only through pre-authored prompt modules and pre-defined mask areas. There is no experimental condition in which a designer directly supervises, refines, or iterates on the generation process. The data are therefore equally consistent with an interpretation in which prompt format or mask predefinition alone (independent of any human designer) improve outputs. The causal attribution to 'designers' is underidentified. The authors should either add an experimental arm with direct designer interaction, or reframe the claim as an effect of structured input scaffolding rather than of designer involvement per se.","section":"Research Design (Levels of Designer Involvement); Hypothesis 2"}],"minor_comments":[{"comment":"The text says 'the significance level of difference in mean aesthetic scores between various degrees of freedom was 0.065' while Table 1 reports .064 for aesthetic scores; please ensure consistency between the text and table values.","section":"Result Analysis with different degrees of freedom"},{"comment":"The heading 'Result Analysis with different degrees of freedom' appears twice in the paper, and the phrase 'Comparison of guiding methods in moderate freedom' is followed by a figure labeled 'Bloxplot' (a typo for 'Boxplot'). Please correct the repeated heading and the figure label.","section":"Result Analysis with different degrees of freedom"},{"comment":"The term 'polishing' is used interchangeably with 'refinement' in different sections; the 'Extended Analysis' refers to 'refinement' while the results section uses 'polishing.' Please unify the terminology to avoid confusion.","section":"Difference analysis before and after polishing of the same guiding method"}],"recommendation":"major_revision","confidential_remarks":"The paper's topic fits a human-computer interaction or design-research venue, and the system-building effort is commendable. However, the editorial process should require the authors to either strengthen the statistical evidence or substantially temper the causal claims; the current gap between 'marginally significant' results and 'significantly enhances' language is a reproducibility concern. If the authors can provide the requested 2x2 factorial analysis and clarify the construct of designer involvement, the manuscript could become a solid empirical contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, it has a genuine, reasonably sized experiment: 160 participants, a fine-tuned Stable Diffusion model for pocket gardens, and one clean significant contrast—under the 'moderate guidance' condition, structured prompts beat structured masks on both aesthetic and alignment scores (p = .013 and .026). That is a useful, citable result about what makes text-to-image generation work in participatory design. Second, the abstract and conclusion claim that 'designer involvement significantly enhances' output quality. The reported statistics do not support that causal claim. The ANOVA table shows the main effect of 'guidance degree' is marginal (p = .064 for aesthetics, .075 for alignment), and the low-vs-high freedom comparison is also marginal (p = .055 and .069). The paper's own results section honestly calls these 'marginally significant'—then the conclusion forgets that.\n\nThe deeper problem is identification. The experimental design is a 2x2 factorial on prompt structure (freeform vs structured) and mask structure (freeform vs structured), but the analysis collapses the cells into an ordinal 'degree of freedom' factor. The only highly significant contrast, structured-prompt/freeform-mask versus freeform-prompt/structured-mask, changes both factors simultaneously, so you cannot attribute the effect to 'designer involvement' as a whole. It could be prompt structure alone, mask predefinition alone, or an interaction. A proper 2x2 ANOVA on the existing cells would separate these components. Until that is reported, the headline should be weakened to something like 'structured prompt guidance is associated with higher perceived quality' rather than 'designers significantly improve.'\n\nWhat the paper does well: it addresses a real, current question about designer roles in AIGC-mediated public participation; the scenario (street pocket gardens) is sensible; the sample size is justified; and the significant prompt-vs-mask contrast is a legitimate empirical finding. The authors also built a working platform and fine-tuned a model, which is concrete. But evaluation is self-rated by the participants about their own generated images—not blind, no expert assessment, no inter-rater reliability—and no data or code are released. The professional-background analysis is mostly descriptive, and the interaction with guidance degree was non-significant, so the claims about designers and public users should be modest.\n\nWho is this for? People working on participatory planning tools, text-to-image applications in design, or human-AI collaboration will get something from the experimental setup and the prompt-vs-mask result. The paper deserves a serious referee, but the referee should insist on the 2x2 analysis and language that matches the p-values. It is not ready as is; it is a solid revision candidate.","headline":"A real experiment on prompt structure in AI-generated urban design, but the headline claim about 'designer involvement' is not identified by the analysis.","tokens_in":8675,"tokens_out":2254,"would_cite":false,"duration_ms":23052,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When designers structure prompts and masks, AI-generated urban garden designs score higher and align better with user intent.","keywords":["AIGC","urban design","public participation","Stable Diffusion","text-to-image generation","designer involvement","pocket garden","soft inpainting"],"falsifier":"Run the same three-scene pocket-garden task with a fourth arm in which non-designers write their own structured prompt modules and define their own masks, or use structured prompts generated by a generic prompt-optimizer instead of by designers; if that arm scores as high as the intensive-guidance group, the paper's claim that designers specifically improve output collapses.","tokens_in":2,"feed_emoji":"🏙️","tokens_out":6566,"duration_ms":119894,"temperature":0.7,"pith_summary":"The paper tries to establish that urban designers are not made obsolete by text-to-image AI but become more valuable as their role shifts from drafting to structuring. Using a Stable Diffusion model fine-tuned for street-garden scenes, the authors had 160 participants generate pocket-garden designs under three levels of designer involvement, then scored the results on aesthetic quality and alignment with intent. Its headline claim is that designer involvement—especially structured prompt modules plus predefined mask areas—significantly improves AIGC output quality, while fully freeform generation yields lower and more variable scores. A sympathetic reader would take the practical message to be: in participatory urban design, designers should intervene early by framing and constraining the generation task, not by fixing images afterward.","feed_headline":"Designer guidance boosts AI-generated urban design scores","feed_subtitle":"Experiment with 160 participants: structured prompts and predefined masks beat freeform input on quality and alignment.","key_machinery":"The mechanism carrying the argument is the experimental contrast among three designer-involvement levels, operationalized by two controls: prompt format (freeform vs structured prompt modules prepared by designers) and mask format (freeform vs designer-predefined modification areas), applied within a fine-tuned Stable Diffusion pipeline with soft inpainting and two domain knowledge graphs encoded as tags. These controls convert 'designer involvement' into measurable inputs, and the 1-10 aesthetic and alignment ratings from 160 participants convert output quality into comparison data. The structured-prompt control is the load-bearing component: at moderate involvement it produced significantly higher scores than predefined masks, indicating that how design intent is phrased matters more than where edits are allowed.","core_discovery":"On its own terms, the paper's discovery is that the quality and goal-alignment gap in AIGC-generated urban designs is driven mainly by upstream designer guidance. In the experiment, low-freedom conditions (designer-defined masks plus structured prompt modules) produced the best mean aesthetic and alignment scores; high-freedom conditions (freeform masks and freeform prompts) produced the lowest and most scattered scores. At moderate involvement, structured prompts with freeform masks outperformed freeform prompts with structured masks, and this difference was statistically significant. Post-hoc polishing under freeform prompts raised scores but not significantly, which the paper reads as evidence that early structuring beats later refinement. The paper concludes that designers' value in the AIGC era lies in translating public intentions into structured prompts and bounded modification areas, with professional background playing no significant interaction with guidance effects.","pith_inferences":["Editorial inference: the manipulation equates designer involvement with prompt and mask structure, so the causal claim about designers is only as strong as the assumption that freeform input is a fair proxy for 'no designer'; a direct arm where designers personally supervise generation would test whether the human, rather than the format, carries the effect.","Editorial inference: the quality gap between structured and freeform prompts may shrink as text-to-image models become better at instruction following, which would push designers' protected role toward tasks the models still miss, such as spatial strategy and cultural context.","Editorial inference: the same factorial design could be applied to other generative outputs (3D massing, plan layouts, section drawings) and to other participant populations to test whether structured prompts generalize or are specific to pocket-garden aesthetics.","Editorial inference: because the authors did not measure whether participants' design intent itself changed under structured prompts, a remaining question is whether structured guidance improves alignment by better capturing intent or by steering participants toward designer-approved intentions."],"forward_implications":["If designer involvement lifts quality through structured prompts and predefined masks, participatory AIGC workflows should place designers at the input stage, before generation, rather than only in post-processing.","The finding that non-professionals scored higher under low-freedom conditions implies that deep designer involvement can make public participation more effective, not less.","Structured prompts being more effective than structured masks at moderate involvement points to prompt-design guidance as the highest-yield intervention for AI-assisted urban design.","If post-hoc refinement alone does not significantly rescue freeform outputs, then investment in upstream guidance may be more cost-effective than iterative polishing."],"supporting_citations":[{"why":"Provides the latent diffusion architecture used to generate and edit street-view images.","marker":"Rombach et al., 2022"},{"why":"Supplies the text-to-visual alignment model the paper relies on for translating prompts into imagery.","marker":"Radford et al., 2021"},{"why":"Establishes the participatory placemaking-with-AI precedent this experiment extends.","marker":"Kim et al., 2022"},{"why":"Documents the Stable Diffusion, MidJourney, and DALL-E 2 generators the paper situates its tool choice among.","marker":"Borji, 2022"},{"why":"Grounds the public-participation motivation and the role of communities in the planning process.","marker":"Lane, 2005"},{"why":"Supports the use of natural language as a carrier of sense-of-place that AIGC can translate into visuals.","marker":"Wartmann & Purves, 2018"},{"why":"Frames the space-to-place shift that motivates designing for spatial experience.","marker":"Seamon & Sowers, 2008"},{"why":"Defines the concept of place the design-element knowledge graph is built around.","marker":"Najafi & Shariff, 2011"}],"fun_headline_variants":["Designer guidance lifts AI urban design scores","Structured prompts beat freeform in AI urban design","Human-AI collaboration improves urban design quality","AI urban design thrives with early designer structure","Designer input boosts AI-generated urban spaces"],"cache_read_input_tokens":10880,"weakest_assumption_plain":"The argument assumes that designer involvement is fully captured by structured prompt modules and predefined mask areas, so that freeform prompts with freeform masks count as 'no designer involvement'—if the quality difference actually comes from prompt format alone, the designer-necessity claim weakens.","fun_headline_variants_meta":{"raw":{"variants":["Designer guidance lifts AI urban design scores","Structured prompts beat freeform in AI urban design","Human-AI collaboration improves urban design quality","AI urban design thrives with early designer structure","Designer input boosts AI-generated urban spaces"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1384,"prompt_tokens":908,"completion_tokens":476,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":408}},"tokens_in":524,"tokens_out":476,"duration_ms":5289,"temperature":1.0,"reasoning_tokens":408,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:23:16.946720+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three-scene pocket-garden task with a fourth arm in which non-designers write their own structured prompt modules and define their own masks, or use structured prompts generated by a generic prompt-optimizer instead of by designers; if that arm scores as high as the intensive-guidance group, the paper's claim that designers specifically improve output collapses.","supporting_citations":[],"review_version":1}