{"id":"4da03878-716a-4d99-8735-09c95fbfb1bb","arxiv_id":"2505.11666","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A GenAI-assisted system that decomposes reference product images into design features and lets consumers compose those features into new product designs, improving engagement and exploration in a 24-user study.","lead":"DesignFromX is a prototype system that lets consumers pick parts of product photos, get AI-generated descriptions of those parts, and combine the described features into a new product image. In a 24-person user study, people using it explored more design features, generated more design variants, and reported more enjoyment and less effort than with a manual prompting baseline.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'features explored' outcome is measured differently across conditions, confounding the headline comparison.","rationale":"The reader's weakest_assumption concerned the validity and completeness of the eight-category feature taxonomy. My concern is related but distinct and more load-bearing: even granting the taxonomy, the way 'features explored' was counted differs systematically between the two conditions. DesignFromX records selections from a presented list; the baseline requires manual prompt articulation followed by experimenter coding. This asymmetry makes the central quantitative claim uninterpretable as evidence about exploration support. The reader did not explicitly identify this measurement-invariance confound, though their concern about the system's own ontology is adjacent. The paper may still have useful system contributions, but the primary evaluation cannot be trusted without re-analysis or raw data. Since the reported statistics cannot settle the question and no code or data are provided, the appropriate verdict is UNVERDICTED pending a condition-blind re-scoring of the key outcome or equivalent evidence. This is a stronger position than the reader's CONDITIONAL because it targets the construct validity of the main dependent measure rather than an auxiliary assumption about taxonomy completeness.","tokens_in":27123,"tokens_out":5244,"duration_ms":56685,"concrete_test":"Re-score 'features explored' from the raw logs with condition-blind annotators. For every participant in both conditions, assemble the full list of generated images and all text prompts actually used for image generation, stripped of any reference to DesignFromX's interface or to the eight category labels. Have two independent raters count the number of distinct design features instantiated in each prompt or image, using a shared ontology (e.g., the paper's own eight categories). Compute the same Wilcoxon signed-rank comparison on these blind-coded counts. If the difference is no longer significant, the headline result is an artifact of the measurement instrument; if it remains significant, the concern is resolved.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The primary quantitative support for the central claim—that DesignFromX increases the number of design features explored (Table 1, p=0.023)—rests on a non-comparable measurement of the dependent variable. In the DesignFromX condition, 'features explored' is counted from the selections users make in the system's own eight-category feature table (Section 4.3 and 5.4.2). In the baseline condition, the count is obtained post hoc by analyzing the free-text prompts participants wrote to generate images (Section 5.4.2: 'we classified the design features identified by participants by analyzing the prompts'). These two procedures measure different constructs: the system presents a visible checklist that can be clicked in one step, while the baseline requires participants to spontaneously recall and articulate features under time pressure. The higher count in DesignFromX may therefore reflect the salience and low cost of selecting items from the displayed taxonomy rather than a genuine increase in the number of features the participant actively explored. The same confound affects the category-level comparisons (mechanism p=0.038, ergonomic p=0.014) and the interpretation of Figure 8. This measurement invariance issue is more fundamental than taxonomy completeness: even if the eight categories are a perfect representation of product features, the two conditions are still measured with different instruments, so the reported difference cannot be attributed to the system's support for exploration.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents DesignFromX, a GenAI-based design support system that lets consumers explore product design space by clicking on components in reference images, receiving LLM-generated feature analyses organized into an eight-category taxonomy, and composing selected features into a base design via DALL-E image editing, with optional 3D model generation. The authors report a formative study (N=8), a module-quality evaluation (segmentation accuracy, feature-analysis ratings), and a within-subject user study (N=24) comparing DesignFromX with a manual-prompt baseline that shares the same interface and underlying image-generation model. They report that DesignFromX significantly increased the number of features explored and the number of new designs generated, and that participants reported higher enjoyment, immersion, and transparency, with lower effort. The abstract additionally claims that DesignFromX lowers frustration, although the reported frustration comparison is not significant.","tokens_in":27374,"tokens_out":4703,"duration_ms":48873,"significance":"The contribution is timely: consumer participation in product design through generative AI is an active HCI topic, and the proposed workflow—interactive segmentation, LLM-driven feature decomposition, and visual feature composition—is a plausible way to reduce articulation barriers for novice users. The two-phase evaluation, including expert ratings of module outputs and a counterbalanced within-subject comparison, is a serious empirical effort. The paper also provides concrete interaction examples and failure cases (Fig. 10) and discusses trade-offs such as user control versus exploration in a balanced way. However, the headline quantitative claims rest on a non-invariant measure and on uncorrected multiple comparisons, and one reported significant result (effort) is contradicted by the direction of the means as reported. The central claim is therefore plausible but not yet established.","major_comments":[{"comment":"The dependent variable 'number of features explored' is not measured in the same way in the two conditions. In DesignFromX, the count comes from selections in the system's own feature table (pop-up C in Fig. 3), whereas in the baseline it is derived post hoc from free-text prompts (Section 5.4.2: 'we classified the design features identified by participants by analyzing the prompts'). The DesignFromX interface makes the eight feature categories salient and selection a one-click action, while the baseline requires spontaneous recall and articulation. The reported difference (8.333 vs 6.625, p=0.023) may therefore reflect the measurement instrument rather than a genuine difference in exploration behavior. This measurement-invariance issue is more fundamental than taxonomy completeness: even if the eight categories are a perfect representation, the two conditions are still measured with different instruments. Please use condition-blind coding of interaction logs or of the final products for both conditions, and report inter-rater reliability.","section":"5.4.2, Table 1"},{"comment":"The abstract states that DesignFromX 'lowers the barriers and frustration,' and the conclusion repeats the engagement/enjoyment framing, but the NASA-TLX frustration comparison is not significant (DesignFromX M=1.917, SD=1.176; Baseline M=2.167, SD=0.963; p=0.265; r=-0.271), and the other workload subscales except effort are also non-significant. The claim of frustration reduction is not supported by the reported data and should be removed or reframed as a directional trend.","section":"Abstract, 6.2.3, Appendix A.3"},{"comment":"Many Wilcoxon signed-rank tests are reported (Tables 1-2, Figures 9-11), but no correction for multiple comparisons (e.g., FDR or Bonferroni) is applied. With roughly 20 tests, p-values such as 0.042 and 0.038 would not survive correction, and even p=0.023 is not robust. Please report the total number of comparisons, apply a correction or pre-registered hypotheses, and interpret marginal effects accordingly.","section":"5.4.4, Tables 1-2, Figures 9-11"},{"comment":"The effort result is internally inconsistent. The text says DesignFromX significantly reduced effort and reports a negative effect size (r=-0.729), which the paper defines as DesignFromX having a lower score. However, the reported means are DesignFromX M=3.042 vs Baseline M=2.471 on the NASA-TLX effort subscale, where higher scores conventionally mean more effort. Either the scale direction is reversed or the conclusion is the opposite; please clarify the coding and verify the reported comparison.","section":"6.2.3, Figure 11"},{"comment":"The significant increase in the number of new designs generated (81.667 vs 66.25, p=0.042) is potentially confounded with time on task: time spending is marginally non-significant (25.694 vs 20.885 minutes, p=0.056), and minutes per image generated is not significant (p=0.229). If participants simply used DesignFromX longer, a higher raw count of generated images would follow. Please report a rate-based analysis (designs per minute) or otherwise control for time on task.","section":"Table 2"}],"minor_comments":[{"comment":"There is an unfinished placeholder, 'nnnn period.', after the participant recruitment paragraph; please remove or complete it.","section":"5.3.1"},{"comment":"Typos remain, including 'desgin' in Section 3.3, 'atheistic' in Section 2.2, and 'metrices' in Appendix A.3; a copyedit pass is needed.","section":"Throughout"},{"comment":"The questionnaire items described as 'qualitative data' are Likert-scale ratings (quantitative self-report measures); please use the term 'subjective ratings' or distinguish these from the interview data.","section":"5.4.3"},{"comment":"The SUS figure reports only descriptive results for DesignFromX with no baseline comparison and no explicit scale in the caption; please state the SUS scoring range, sample size, and whether a statistical comparison was performed.","section":"Figure 12"},{"comment":"The caption of Figure 13 does not clearly distinguish the self-reported and expert-reported panels; please label each panel and state which rows correspond to which rater group.","section":"6.3.2, Figure 13"}],"recommendation":"major_revision","confidential_remarks":"The system contribution and the two-phase study design are solid, but the main quantitative evidence needs repair before acceptance. The measurement-invariance issue in the 'features explored' metric is the key review point, and the effort-result direction error and frustration overclaim must be fixed. I would be willing to look at a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's the quick take on DesignFromX. The system is a reasonable integration of SAM2, GPT-4o feature analysis under a fixed eight-category taxonomy, and DALL-E 2 composition. The workflow—click a reference image, get a structured feature list, compose selected features into a target product—is genuinely new as a consumer-facing design tool, and the formative study and design goals are sensible. The module evaluation (IoU 0.839, label accuracy 0.90, expert-rated feature descriptions above 4/5) is real evidence that the pieces work.\n\nThe soft spot is the headline quantitative claim. The paper argues DesignFromX increases the number of features explored (8.33 vs 6.63, p=0.023), but the two conditions count this differently. In DesignFromX, features explored are the items users click in the system's own feature table—a visible checklist that costs one click. In the baseline, the count comes from post hoc analysis of the free-text prompts users wrote to describe features. Those are different instruments measuring different constructs. The higher count in DesignFromX could just reflect the salience and low cost of selecting from a displayed taxonomy, not a real increase in exploration. That confound also affects the category-level comparisons (mechanism, ergonomic) and makes the 'features explored' result uninterpretable as evidence for the system's support. The self-reported UX results (enjoyment, immersion, exploration support) are not confounded the same way, though the multiple Wilcoxon tests without correction and p-values hovering just under 0.05 mean I'd treat those as suggestive rather than conclusive. The abstract's claim that the system lowers frustration is also not supported—the frustration item was not significant (p=0.265), and the body reports it as non-significant. That's an overclaim in the abstract, not a fatal flaw.\n\nThe paper is otherwise honest. The limitations section openly discusses lack of fine control, granularity of features, and GenAI understanding. No code or data are released, which matters for reproducibility but is not unusual for an HCI systems paper.\n\nWho should read it: anyone building GenAI-based design support tools, and anyone thinking about how to evaluate 'exploration' in systems that impose their own ontology. The measurement invariance issue is a useful cautionary tale for the field. It deserves a serious referee; I'd send it to review, but I'd require the authors to either fix the measurement (e.g., count written descriptions of features in both conditions) or soften the claim substantially.","headline":"A useful integration of off-the-shelf GenAI components for consumer design, with a genuine measurement confound undercutting its headline 'features explored' result.","tokens_in":27869,"tokens_out":2505,"would_cite":true,"duration_ms":24470,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that clicking on parts of reference images and composing the detected features into one's own product lets consumers explore more design options than manually prompting an image generator.","keywords":["User Interface Design","Generative AI","Product Design","Design Space Exploration","Feature Composition","Consumer-Driven Design","Human-AI Collaboration","Image Segmentation"],"falsifier":"A think-aloud study in which novice consumers freely describe what they would change about a product, without being shown the eight-category list, would settle the taxonomy question: if more than about a quarter of the changes they name cannot be mapped onto color, style, texture, shape, structure, mechanism, electronic, or ergonomic features, the system's feature analysis and its 'features explored' metric are biased toward the system's own ontology rather than consumers' actual preferences.","tokens_in":26969,"feed_emoji":"🎨","tokens_out":8736,"duration_ms":78499,"temperature":0.7,"pith_summary":"The paper claims that ordinary consumers can meaningfully explore a product's design space if the tool decomposes visual references for them. DesignFromX lets a user click on a part of a reference photo, see that part described as a set of design features, and then brush selected features onto their own product image while a generative model visualizes the result. In a 24-person comparison against a manual-prompting baseline, users of DesignFromX explored more design features, generated more candidate designs, and reported higher enjoyment, higher immersion, and lower effort. The payoff, if the claim holds, is that consumers can shape early-stage product designs without learning design terminology or prompt engineering.","feed_headline":"Consumers explore more designs with click-and-compose AI","feed_subtitle":"In a 24-person study, the tool beat manual prompting on features explored, designs generated, enjoyment, and effort.","key_machinery":"The central mechanism is the feature-composition pipeline. A click-driven segmentation model (SAM 2) extracts the component a user points to; a large-language-model agent (GPT-4o) analyzes that component into eight named design features—color, style, texture, shape, structure, mechanism, electronic, and ergonomic; and a second language-model step rewrites the initial design description, which an image-editing model (DALL·E 2) uses to visualize the updated product. This pipeline converts a visual reference into structured, composable attributes, so users can iterate by adding features from different products without writing prompts from scratch.","core_discovery":"DesignFromX's central claim is that a click-to-compose pipeline lets consumers explore more of a product's design space than manually describing reference features and prompting an image generator. In the within-subject user study (N=24), DesignFromX significantly increased the number of design features explored (M=8.333 vs 6.625, p=0.023) and the number of new designs generated (M=81.667 vs 66.25, p=0.042), and significantly raised self-reported enjoyment (p=0.002) and immersion (p=0.003) while lowering effort (p=0.037). The paper reads these results as evidence that lowering the barrier to feature analysis and prompt formulation reduces frustration and supports consumer-driven design space exploration.","pith_inferences":["A direct extension would be to replace the fixed eight-category taxonomy with an open-ended vocabulary generated per product, which would test whether the taxonomy itself or the act of structured decomposition drives the exploration gains.","The logged chains of features users select could be aggregated across participants to reveal preference patterns, effectively turning the tool into a lightweight market-research instrument for designers.","The study measured exploration and experience, not manufacturability; coupling feature composition with parametric CAD or physical constraints would test whether broader exploration translates into buildable products.","Several participants wanted to edit or refine the system's suggested features and prompts, suggesting that a hybrid mode with direct prompt or parameter control might combine the exploration benefits with the controllability expert designers need."],"forward_implications":["Consumers will explore more functional features—especially mechanism and ergonomics—because the system translates visual references into selectable terms they would not otherwise know how to name.","Novices can generate many more candidate designs per session, giving them a larger design space to choose from and increasing satisfaction with their final design.","Design support tools can offload prompt formulation from the user to an LLM chain, reducing the effort needed to drive text-to-image models.","Because expert ratings of final designs were similar across conditions, the system's advantage lies in the process and experience of exploration rather than in producing objectively better products.","Treating design as incremental local edits rather than full regeneration lets consumers build step by step and keep earlier choices intact."],"supporting_citations":[{"why":"Supplies the click-based segmentation model that isolates a user-selected component from a reference image.","marker":"[78]"},{"why":"Supplies the language-model agent that analyzes the segmented component and produces the design-feature descriptions.","marker":"[72]"},{"why":"Supplies the image-generation and editing model used to visualize the composed designs.","marker":"[73]"},{"why":"Provides the aesthetic design-feature taxonomy (shape, color, texture) that the feature analysis module follows.","marker":"[22]"},{"why":"Provides the functional-feature categories (structure, mechanism, electronics) used by the feature analysis module.","marker":"[53]"},{"why":"Justifies adding style as an aesthetic design feature because it shapes consumer perception.","marker":"[21]"},{"why":"Justifies including ergonomics as a functional design feature because it affects user experience.","marker":"[7]"},{"why":"The Creativity Support Index that the paper adapts into its Design Support Index for measuring user experience.","marker":"[17]"},{"why":"The NASA-TLX workload instrument used to compare task load between DesignFromX and the baseline.","marker":"[42]"},{"why":"Provides the human-reported metrics used to evaluate the quality of the feature-analysis and composition modules.","marker":"[57]"}],"fun_headline_variants":["Click-to-compose AI helps consumers explore more product designs","AI design tool clicks features to expand consumer exploration","Feature-clicking AI lets consumers explore more design space","Click-to-compose tool boosts design exploration in consumers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the fixed eight-category feature list—color, style, texture, shape, structure, mechanism, electronic, and ergonomic—is a valid and sufficiently complete map of the design features consumers care about, so that counting those categories measures how much design space was explored.","fun_headline_variants_meta":{"raw":{"variants":["Click-to-compose AI helps consumers explore more product designs","AI design tool clicks features to expand consumer exploration","Feature-clicking AI lets consumers explore more design space","Click-to-compose tool boosts design exploration in consumers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000787,"raw_usage":{"total_tokens":3425,"prompt_tokens":852,"completion_tokens":2573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":2510}},"tokens_in":468,"tokens_out":2573,"duration_ms":16357,"temperature":1.0,"reasoning_tokens":2510,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:50:02.671198+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A think-aloud study in which novice consumers freely describe what they would change about a product, without being shown the eight-category list, would settle the taxonomy question: if more than about a quarter of the changes they name cannot be mapped onto color, style, texture, shape, structure, mechanism, electronic, or ergonomic features, the system's feature analysis and its 'features explored' metric are biased toward the system's own ontology rather than consumers' actual preferences.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the language-model agent that analyzes the segmented component and produces the design-feature descriptions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the image-generation and editing model used to visualize the composed designs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the functional-feature categories (structure, mechanism, electronics) used by the feature analysis module."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The NASA-TLX workload instrument used to compare task load between DesignFromX and the baseline."}],"review_version":1}