{"id":"f8f186cb-0f7f-4bc1-9065-48a588b27335","arxiv_id":"2506.15873","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"DeckFlow is a multimodal generative-AI canvas that uses goal, action, and cluster cards to support decomposition and exploration, and users preferred it for open-ended creative tasks over a chat baseline.","lead":"DeckFlow is a new infinite-canvas tool that lets people create images, text, and audio with generative AI by breaking big goals into connected cards and exploring many output variations at once. In user studies, people preferred it over a chat-based AI interface for open-ended creative tasks, though the evidence is qualitative and the code is not yet publicly available.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Open-ended preference may be driven by DeckFlow's automatic prompt variation, not its canvas; ChatFlow lacks this, so the central attribution is uncontrolled.","rationale":"The reader's weakest_assumption correctly identifies this confound: DeckFlow automatically constructs three prompt variations, including a SuperPrompt-augmented row, while ChatFlow relies on user-authored prompts. I agree that automatic prompt engineering could dominate the open-ended preference results. I am not proposing to reject the paper: the design framework, qualitative findings, and transferable affordances are valuable, and the paper explicitly acknowledges limitations in Section 6.4, including that the baseline is just one comparison point. However, the headline claim depends on isolating the interface contribution from the prompt-engineering contribution. A targeted follow-up comparing ChatFlow with the same automatic prompt-variation capability is inexpensive and would settle attribution. Until then, CONDITIONAL remains the right verdict, and I see no reason to move it.","tokens_in":16509,"tokens_out":2455,"duration_ms":26208,"concrete_test":"Modify ChatFlow to include a 'variations' control that, on demand, generates the same three prompt families DeckFlow uses (direct concatenation, GPT-4-Vision rewrite, SuperPrompt augmentation) and displays them as candidate prompts before image generation. Run the open-ended task from Table 1 with the same population and rating scales, with a sample size powered to detect the observed preference gap. If ChatFlow-with-variations closes the preference gap, the interface affordances are not the primary cause of DeckFlow's advantage; if the gap persists, the interface explanation survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that DeckFlow improves outcomes for open-ended creative tasks (Section 7). The evidence is a within-subjects comparison in which DeckFlow was universally preferred over ChatFlow on open-ended tasks (Section 5.3.2). However, the two conditions differ in at least two ways at once: the canvas/action-card/cluster interface, and the automatic construction of three prompt variants per Action Card generation, including a row produced by a specialized prompt-augmentation LLM (SuperPrompt) selected to improve creative quality (Sections 3.3.5 and 3.1). ChatFlow users had to author or request prompts themselves; no automatic variation or augmentation was provided (Section 4.1). Thus the preference could reflect better or simpler prompt engineering rather than the spatial, dataflow, or clustering affordances the paper claims to validate. The phrase 'improves outcomes' is also stronger than what was measured: ratings and interviews capture subjective preference, not objective outcome quality. This is a confound, not an internal inconsistency, but it is load-bearing because it determines whether the paper's design implications are about canvas interactions or about embedding prompt-engineering models in any interface.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"DeckFlow is a multimodal generative AI tool built on an infinite canvas of cards, Action Cards, Goal Cards, and Clusters, with the goal of addressing task decomposition, specification decomposition, and generative space exploration for text, image, and audio creation. The paper motivates these three design problems through a literature review and describes the DeckFlow implementation in detail. It reports two user studies: a within-subjects comparative study (n=16) against a conversational baseline (ChatFlow) for text-to-image tasks, and a multimodal behavioral study (n=7) with no baseline. The claimed contribution is that DeckFlow supports diverse workflows and improves outcomes for open-ended creative tasks, while performing similarly to ChatFlow for closed-ended tasks.","tokens_in":16686,"tokens_out":3341,"duration_ms":35031,"significance":"If the central claim is accepted, the paper would provide a useful open-source system and a well-structured vocabulary for thinking about generative AI interfaces. The three design problems are clearly articulated and the system's affordances (Goal Cards, Action Cards with labeled ports, Clusters, generative grids) are concrete and transferable. The paper also ships a reproducible implementation and an honest effort to compare against a strengthened conversational baseline. However, the evaluation evidence is weaker than the conclusions drawn: the comparative ratings are descriptive only, the multimodal study lacks any baseline, and the open-ended preference is confounded by an automatic prompt-engineering component that is absent in the baseline. These issues undermine the generalizable claims about interface affordances.","major_comments":[{"comment":"The conclusion in Section 7 that DeckFlow \"improves outcomes for open-ended creative tasks\" is not supported by the evidence presented. Figure 9 shows rating distributions, and the text reports qualitative preferences, but no inferential statistics (e.g., Wilcoxon signed-rank test), effect sizes, or confidence intervals are reported. The phrase \"universally preferred\" (Section 5.3.2) is also stronger than what the data can establish in a sample of 16 participants. Please either add appropriate statistical analyses or substantially weaken the claims to describe subjective preference in this sample.","section":"§5.3.2, Figure 9, §7"},{"comment":"The central attribution of open-ended task preference to DeckFlow's interface is confounded with automatic prompt engineering. DeckFlow automatically constructs three prompt variants per Action Card generation, including a row based on the SuperPrompt LLM (Section 3.3.5), whereas ChatFlow relies on user-authored or user-requested prompts (Section 4.1). Participants in DeckFlow therefore did not have to write prompts themselves and received systematically varied, potentially higher-quality prompts. It is unclear whether the measured preference reflects the canvas, cards, and clusters, or simply the embedded prompt-variation system. This is not controlled for anywhere in the study design and is not discussed in the threats to validity (Section 6.4). Please add a control condition or analyze usage logs to separate these factors, or explicitly reframe the contribution of the evaluation as showing the combined system, not the interface affordances per se.","section":"§4.1 vs. §3.3.5"},{"comment":"The multimodal behavioral study (n=7) has no baseline condition, yet the conclusion section states that users \"decompose open-ended creative tasks which involve multiple modalities in similar, structured ways\" and makes generalizable claims about multimodal generative space exploration. Without a comparison condition, these are descriptive observations about how people interact with this specific tool, not evidence that DeckFlow's design is responsible for the observed behaviors. Please temper the generalizable claims or add a comparison condition.","section":"§4.2, §7"},{"comment":"The statement that \"similar performance in closed-ended tasks\" is based on the absence of a clear favorite in the rating distributions, but no equivalence test or inferential comparison is provided. The paper should either support this claim with appropriate statistical procedures for evaluating similarity/equivalence or phrase it as an observation about participant ratings rather than an established finding.","section":"§5.3.1"}],"minor_comments":[{"comment":"The sentence \"To refer to participants, we use the format nB =3\" is garbled and appears to be a leftover from an earlier draft; it should be removed or clarified so that the participant naming scheme is clear.","section":"§5.1"},{"comment":"The phrase \"least interrogated output\" uses \"interrogated\" in an unusual way; consider replacing it with e.g., \"least examined\" or \"least used output\".","section":"§5.4.2"},{"comment":"The caption \"(16 open-ended, 16 close-ended)\" would be clearer if it stated that these are the number of task instances completed by the 16 participants in each condition, since the wording is easy to misread as 32 distinct participants.","section":"Figure 9 caption"},{"comment":"In the quote from PA 12, \"since there are 3 roles every time\" should likely read \"3 rows every time\" to match the interface terminology used elsewhere in the paper.","section":"§5.3.3"}],"recommendation":"major_revision","confidential_remarks":"The system contribution is real and the paper is generally well written, but the headline claim of improved creative outcomes is currently built on a confounded comparison and descriptive statistics. I would encourage the editor to ask for either a controlled follow-up that separates automatic prompt variation from the canvas affordances, or a revision that narrows the paper's claims to what the current evidence can support. The open-source release is a valuable part of the contribution and should be preserved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on generative-AI interfaces. The core contribution is a clean three-problem taxonomy (task decomposition, specification decomposition, generative space exploration) and a working system that maps each problem to concrete affordances: Goal Cards that auto-decompose a prompt into labeled Action Card ports, Clusters for example-based spec, and three-row grid generation with prompt variants. The taxonomy is genuinely useful synthesis, not just repackaging; the Goal Card decomposition is the most novel piece and seems to have real traction with users. The paper also reports honest, fairly rich qualitative data: distinct workflow patterns (top-down, one-card iteration, divide-and-conquer), label-type categories, and user quotes that include frustration and confusion, not just praise. The literature review is thorough and the two studies are well described.\n\nWhere it gets soft: the headline claim that DeckFlow \"improves outcomes for open-ended creative tasks\" leans on a within-subjects comparison where the two conditions differ in two ways at once. DeckFlow automatically constructs three prompt variants per generation, one via a prompt-augmentation LLM; ChatFlow users wrote their own prompts. So the observed preference could be driven by automatic prompt engineering rather than the canvas, clusters, or action cards. That is a load-bearing confound because the paper's design implications are about canvas interactions. It is not fatal: the open-ended-task preference is still a real user-reported result, the closed-ended task provides some comparison, and the multimodal study adds behavioral evidence that doesn't depend on the baseline. But the abstract's \"significant improvements\" overstates what descriptive ratings and interviews can support. The small homogeneous sample (16 and 7 CS/EE students) and lack of inferential statistics are limitations the authors themselves acknowledge; I'd call them expected for an HCI systems paper, not disqualifying. One more minor issue: the multimodal study has no baseline at all, so it can only describe behavior, not validate the design.\n\nWho is this for? HCI and creative-AI researchers, and anyone building canvas-based generative tools. It deserves a serious referee: the system is substantial, the qualitative analysis is careful, and the design framework will be citable even if the comparative claim needs tempering. My recommendation: engage with it, but push the authors to either add a condition controlling for prompt variation or rewrite the conclusion to claim \"users preferred DeckFlow in open-ended tasks\" rather than \"DeckFlow improves outcomes.\"","headline":"Solid systems paper with a real confound: DeckFlow's open-ended-task win over ChatFlow mixes canvas interactions with automatic prompt variation, so the headline attribution is not settled by the data as presented.","tokens_in":17210,"tokens_out":597,"would_cite":true,"duration_ms":8510,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A card-based infinite canvas with automatic goal decomposition improves open-ended generative AI creation over chat, while matching it on closed-ended tasks.","keywords":["generative AI","prompt engineering","text generation","image generation","audio generation","infinite canvas","multimodal"],"falsifier":"Run the same open-ended image tasks with the chat baseline augmented to generate the same three automatic prompt variations per request, or with DeckFlow's variation rows disabled; if the preference gap in outcome or usability vanishes, the improvement is the built-in prompt generation, not the canvas.","tokens_in":16281,"feed_emoji":"🎨","tokens_out":7474,"duration_ms":70831,"temperature":0.7,"pith_summary":"The paper argues that generative AI tools fail users in three ways: they do not let people split a large creative goal into connected subtasks, combine several partial specifications for one artifact, or see and steer many outputs at once. DeckFlow answers all three with an infinite canvas of cards: a Goal Card breaks a user's sentence into labeled slots, an Action Card combines the cards attached to those slots and generates three rows of three image variants, and Clusters interpret groups of outputs back into text for the next round. In a 16-person comparison against a conversational chat baseline using the same image model, participants rated DeckFlow better for open-ended creation such as 'make a picture you would hang in your dining room,' while closed-ended reproduction tasks came out roughly even. A follow-up 7-person study added audio and found that users still specified mainly through text, generated mostly images, and reacted most strongly to audio. The paper's claim is that visual, decomposable interaction gives users more control and better creative outcomes than conversation for open-ended tasks.","feed_headline":"Canvas beats chat for open-ended AI image creation","feed_subtitle":"A card-and-cluster canvas helps users split prompts, explore variations, and keep control; the study shows the split improves creative…","key_machinery":"The load-bearing object is the Action Card, a function-like card with labeled sockets that reifies a decomposed specification: each label names a feature (style, subject, lighting) and any text, image, or audio card connected to that socket constrains that feature. The Goal Card initializes this decomposition automatically from a high-level prompt, and the Cluster converts a group of outputs back into a textual description, so visual results can become future input. When triggered, the Action Card produces three prompt-variant rows of three outputs each, creating a breadth axis for exploring the generative space directly on the canvas; the paper argues that this combination is what lets users iterate by modifying one slot at a time rather than re-prompting from scratch.","core_discovery":"The central claim is that the three design problems—task decomposition, specification decomposition, and generative space exploration—can be jointly addressed by a card-based dataflow canvas, and that this materially improves open-ended generative creation over a conversational interface. The mechanism is that a Goal Card is automatically split into an Action Card whose labeled ports accept text, image, or audio cards; triggering the card builds three prompt variants (a literal concatenation, a coherent LLM rewrite, and an aesthetics-oriented rewrite) and emits three outputs per variant; Clusters reinterpret selected outputs back into textual descriptions that feed the next iteration. The evaluation argues that this supports distinct user workflows, yields comparable results on replication tasks, and is preferred on both outcome and usability in open-ended tasks, where the generative space matters most.","pith_inferences":["Editorial extension: the study does not isolate automatic prompt generation, because DeckFlow builds three prompt variations while the chat baseline only uses user-authored prompts; part of the preference gap may be prompt-engineering quality rather than canvas affordances, testable by swapping identical prompt-generation logic into the chat baseline.","Editorial extension: if Goal Card decomposition is the active ingredient, the same decomposition could be offered inside a chat interface as structured follow-up suggestions, predicting an improvement in open-ended tasks without any canvas at all.","Editorial extension: the finding that audio produced the strongest emotional responses but the least precise specification suggests a division of labor across modalities—text to constrain, images to explore, audio to feel—that interface designers could exploit instead of treating modalities as interchangeable.","Editorial extension: several participants said they would use Clusters and Action Cards more with practice, so the short sessions may understate the tool's benefits; a longitudinal study would test whether the initial learning investment pays off."],"forward_implications":["If DeckFlow's evaluation is right, designers of generative AI tools can improve open-ended creative tasks with spatial, decomposable canvases even when the underlying generative model is identical to a chat tool's.","Automatically generated prompt variants are a viable way to give novices a productive breadth of outputs; at least one variant row was valued for its creativity despite lower prompt adherence.","The Goal Card pattern is a reusable scaffold: novices who do not know how to start can have the model split a high-level goal into labeled feature slots and fill those slots incrementally.","Clusters that reinterpret a group of images into text provide a working bridge from visual intent to prompt language, including discovering concepts the user could not name.","In multimodal generation, text remains the dominant and clearest input channel even when image and audio input exist, so designers should keep a precise text fallback in every modality."],"supporting_citations":[{"why":"The Promptify canvas that lays out batches of image results around a revised prompt; it supplies the comparison point for generative space exploration that DeckFlow extends.","marker":"[7]"},{"why":"ChainForge's prompt-template fields that can be swept or frozen; it motivates the Action Card's labeled-slot approach to specification decomposition.","marker":"[9]"},{"why":"Sensecape, the closest prior canvas where inputs and outputs are movable cards and outputs can be repurposed as inputs; DeckFlow builds on this task-decomposition approach.","marker":"[10]"},{"why":"The Stable Audio model is the openly accessible audio generator used to add DeckFlow's audio modality in the second study.","marker":"[22]"},{"why":"The Creativity Support Index survey that the paper adapts to measure perceived support for text, image, and audio modalities.","marker":"[27]"},{"why":"The ChatGPT product that the conversational ChatFlow baseline is modeled after; it anchors the state-of-practice comparison condition.","marker":"[28]"}],"fun_headline_variants":["Card canvas beats chat for open-ended AI creation","DeckFlow: Split tasks, explore space, keep control","Beyond chat: A canvas for iterative AI generation","Cards and clusters: Better than chat for generative design"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"DeckFlow automatically creates three prompt variations, including an aesthetics-augmented row, while the conversational baseline only uses the user's own prompts, and the comparison treats any resulting difference as due to the interface rather than to that automatic prompt engineering.","fun_headline_variants_meta":{"raw":{"variants":["Card canvas beats chat for open-ended AI creation","DeckFlow: Split tasks, explore space, keep control","Beyond chat: A canvas for iterative AI generation","Cards and clusters: Better than chat for generative design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000651,"raw_usage":{"total_tokens":2941,"prompt_tokens":857,"completion_tokens":2084,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":2030}},"tokens_in":473,"tokens_out":2084,"duration_ms":16267,"temperature":1.0,"reasoning_tokens":2030,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:29:43.432670+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same open-ended image tasks with the chat baseline augmented to generate the same three automatic prompt variations per request, or with DeckFlow's variation rows disabled; if the preference gap in outcome or usability vanishes, the improvement is the built-in prompt generation, not the canvas.","supporting_citations":[{"cited_title":"Promptify: Text-to-image generation through interactive prompt exploration with large language models,","cited_arxiv_id":null,"evidence_quote":"The Promptify canvas that lays out batches of image results around a revised prompt; it supplies the comparison point for generative space exploration that DeckFlow extends."},{"cited_title":"Chain- forge: A visual toolkit for prompt engineering and llm hypothesis testing,","cited_arxiv_id":null,"evidence_quote":"ChainForge's prompt-template fields that can be swept or frozen; it motivates the Action Card's labeled-slot approach to specification decomposition."},{"cited_title":"Sensecape: Enabling multilevel exploration and sensemaking with large language models,","cited_arxiv_id":null,"evidence_quote":"Sensecape, the closest prior canvas where inputs and outputs are movable cards and outputs can be repurposed as inputs; DeckFlow builds on this task-decomposition approach."},{"cited_title":"Stable audio open,","cited_arxiv_id":null,"evidence_quote":"The Stable Audio model is the openly accessible audio generator used to add DeckFlow's audio modality in the second study."},{"cited_title":"Quantifying the creativity support of digital tools through the creativity support index,","cited_arxiv_id":null,"evidence_quote":"The Creativity Support Index survey that the paper adapts to measure perceived support for text, image, and audio modalities."}],"review_version":1}