{"id":"4b38a403-c1c7-4196-89e4-3928ac87b6ca","arxiv_id":"2411.13422","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper introduces light prompting and tangible prompt fragments as ways to explore the latent space of diffusion models, reframing prompt engineering as craft.","lead":"This paper describes three design projects that use shadows, cards, and body movement to interact with AI image generators, and proposes 'prompt craft' as a more fitting frame than 'prompt engineering.' It matters because it offers practical design patterns for making generative AI interfaces more exploratory, tangible, and inclusive.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cardshark's explicitly omitted workshop outcomes leave the accessibility and inclusiveness claims untestable, so the central reframing to 'prompt craft' is not yet supported by comparative evidence.","rationale":"The reader's weakest assumption identified the same core issue: the generalizable claims rest on anecdotal, self-assessed observations from one workshop and two exhibition settings, with no comparison to conventional prompting interfaces. The Cardshark project is the clearest case because it is the only project with a defined user group (people with Young Onset Dementia) and an explicit accessibility goal, yet its outcomes are explicitly not reported. The Discussion then uses this project as evidence for inclusiveness, creating a direct gap between evidence and claim. A concrete test would settle whether the concern lands by either producing the missing comparative data or forcing a reframing of the contribution as unvalidated design speculation. I considered whether the lack of outcome data is acceptable for a pictorial, since the authors frame their contributions as strong concepts rather than empirical findings. However, the claim 'demonstrate that it is possible to make interactions with generative AI inclusive' is an empirical claim, and the paper's own limitation statement makes verification impossible. The reader's UNVERDICTED status is therefore appropriate: the work is promising and honest, but not enough to accept as empirically grounded. No machine-checked proofs or systematic evaluations are present, and the open-source Shadowplay code does not cover Cardshark's interaction or outcomes. The concern is not about internal inconsistency or authorial intent; it is about the sufficiency of evidence for a central, transferable knowledge claim.","tokens_in":8900,"tokens_out":2942,"duration_ms":33777,"concrete_test":"Re-run or fully report the Cardshark workshop with a minimal outcome protocol: count facilitator interventions, measure autonomous interaction time, record participant-selected saved images, and use a short post-session expression or engagement measure. Run the same protocol with a matched text-prompt-only interface as a baseline. If Cardshark does not outperform or match the baseline on engagement or expressive outcomes, the inclusive-interface claim fails. If outcomes cannot be collected, the strong concept should be explicitly reframed as a design proposal rather than an empirically supported finding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that 'prompt craft' is a productive reframing supported by three projects and transferable as strong concepts or intermediate knowledge. The load-bearing inference is that the Cardshark project demonstrates that tangible prompt fragments make generative AI accessible and inclusive. However, the Cardshark section explicitly states: 'In this pictorial there is not the scope or space to report on the successes and failures of the workshop for participants.' The Discussion nevertheless claims that 'the embodied interactions of Shadowplay and the tangible interactions of Cardshark demonstrate that it is possible to make interactions with generative AI inclusive for a wide range of people' and that the system 'empowered our participants to operate the system autonomously.' These are empirical outcome claims, but the only support is anecdotal: a charity representative's comment and the observation that some participants interacted extensively. No comparison to a conventional text-prompt interface is provided, so the specific design strategies (light prompting, fragment cards, prompt arena) cannot be distinguished from novelty, facilitator presence, or the general appeal of AI image generation. The concept of 'soft edges' of the latent space is likewise metaphorical and not operationalized, making the materiality contribution difficult to verify or falsify. The paper is internally consistent and honest about its status, but the central reframing stands on an evidence gap: the most direct demonstration of the claimed benefit is intentionally omitted, and the remaining demonstrations are not sufficient to establish transferability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports three research-through-design projects using Stable Diffusion: Shadowplay V1 (body and shadow input with 'light prompting'), Cardshark (tangible prompt-fragment cards for a workshop with people living with Young Onset Dementia), and Shadowplay V2 (real-time StreamDiffusion with dynamic prompts, movement-controlled parameters, and reactive audio). From these cases the authors propose 'prompt craft' as a reframing of prompt engineering, with two contributions: a perspective on the materiality of diffusion models and a craft-like method for navigating their latent possibility space, plus interaction design strategies such as light prompting and prompt fragments. The outcomes are modestly framed as strong concepts or intermediate knowledge rather than confirmed results.","tokens_in":9033,"tokens_out":3551,"duration_ms":39988,"significance":"If the framing is accepted, the paper offers useful direction for designing interfaces to generative AI, particularly for tangible and embodied interaction and for broadening access beyond text-prompt interfaces. The work is explicitly self-positioned as intermediate knowledge, it is transparent about its exploratory status, and it ships an open-source codebase for Shadowplay. Its main limitation is evidential: the strongest claims about inclusivity and empowerment are supported by anecdote and self-assessment, and the Cardshark section explicitly declines to report workshop outcomes. The contribution is therefore plausible and generative, but not yet demonstrated at the level the Discussion sometimes claims.","major_comments":[{"comment":"The Discussion states that 'the embodied interactions of Shadowplay and the tangible interactions of Cardshark demonstrate that it is possible to make interactions with generative AI inclusive for a wide range of people' and that Cardshark 'empowered our participants to operate the system autonomously.' However, the Cardshark section explicitly says 'In this pictorial there is not the scope or space to report on the successes and failures of the workshop for participants.' The supports offered are one charity representative's comment and the observation that some participants interacted extensively; no comparison against conventional text-prompt interfaces or against facilitator-free conditions is given. Since the inclusivity and autonomy claims are load-bearing for the accessibility argument, the authors should either report systematic workshop outcomes (e.g., observed autonomy, engagement, participant feedback) or revise the Discussion to present these as promising indications rather than demonstrations.","section":"Cardshark / Discussion"},{"comment":"The materiality contribution rests on the notion of 'soft edges' to the model's latent possibility space, but the term is used metaphorically and is not given observable criteria. As written, it is difficult to verify, falsify, or apply: a designer cannot tell from the paper which behaviors of a diffusion model count as soft edges or how to recognize them. Please specify the concept operationally, for instance by tying it to observable model behaviors such as prompt weighting sensitivity, seed variance, or interpolation continuity, or explicitly position it as an open conceptual metaphor rather than part of the proposed method.","section":"Discussion"},{"comment":"The abstract and conclusion name a 'method for a craft-like navigation of the latent space' as a central contribution, but the Prompt Crafting in Practice section says it is 'tempting to offer specific steps and guidance' and declines to do so. The reader is given a retrospective narrative of Cardshark's development rather than an articulable method that others could adopt or test. If the contribution is a method, the paper should state its constituent heuristics or phases, even at a high level; if it is intended only as an illustrative case, the contribution should be renamed accordingly.","section":"Prompt Crafting in Practice / Discussion"}],"minor_comments":[{"comment":"The text 'developed Cardshark around the to the purpose of the workshop' contains a malformed phrase; it should presumably read 'around the purpose of the workshop.'","section":"Cardshark"},{"comment":"The sentence 'We developed a practice of oscillation between modes... as we discovered the 'soft edged' of the model's latent possibility space' should read 'soft edges'.","section":"Cardshark"},{"comment":"The phrase 'led us to develop the the prompt arena and fragment cards' contains a duplicated definite article.","section":"Cardshark"},{"comment":"The sentence 'To produce a engaging exhibition experience' should read 'an engaging exhibition experience.'","section":"Shadowplay V2"},{"comment":"The comparison of light prompting to 'the quantization that takes place during model development' may be unclear to readers unfamiliar with diffusion models; a brief explanation of what is being compared would help.","section":"Discussion"},{"comment":"The figure captions are informative, but the main text does not always point readers to the specific figures showing the installation setups and the signal-processing UI; adding explicit references would improve navigability.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about its intermediate-knowledge status, which I credit, and it does not appear to suffer from internal inconsistency. My main concern is that the Discussion's empirical-sounding claims about inclusivity and empowerment outrun the evidence reported in the Cardshark section, which explicitly omits participant outcomes. A revision that either supplies those outcomes or recalibrates the language would make the paper's contributions match its evidence, and I would then see it as a publishable pictorial-style contribution. No concerns about duplication or research misconduct arose during my reading."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. It is a practice-based design pictorial, not an empirical paper, and judged on that basis it is a decent piece of work with two genuinely new things in it: 'light prompting' — constraining a diffusion model's latent space by feeding it high-contrast monochrome shadows rather than relying on text — and 'prompt fragments,' the tangible cards in Cardshark. The reframing from prompt engineering to prompt craft is more than a slogan; it gives designers a vocabulary for iterative, material-like engagement with generative models. The paper is also honest: it explicitly says the materiality metaphor is not new, and it openly states there is no room to report workshop outcomes.\n\nThe soft spots are real but not fatal for what this is. The Discussion leans on Cardshark to claim that tangible interactions 'demonstrate that it is possible to make interactions with generative AI inclusive for a wide range of people,' yet the workshop successes and failures are deliberately omitted. That is a load-bearing evidence gap: the only support is a charity representative's comment and the observation that some participants interacted extensively. Without a comparison to a conventional prompt interface, or even a systematic account of who did what, the inclusiveness claim is a strong hypothesis, not a demonstrated result. Likewise, 'soft edges' of the latent space is a useful metaphor but is never operationalized. These weaknesses matter more in the abstract than in the pictorial's own intermediate-knowledge frame, but the authors should tone down 'demonstrate' and say 'suggest.'\n\nOn the academic side: the citation pattern is reasonable, leaning on the authors' own prior work, which is acceptable in a Research through Design programme, and the GitHub link for Shadowplay is a genuine reproducibility gesture. The three projects are described in enough detail to be plausible, and the craft-like process of meta-prompt development is itself a useful transferable insight.\n\nWho this is for: interaction designers and HCI researchers working on generative AI interfaces. It deserves a serious referee, not a desk reject. A referee should push for either systematic reporting of the Cardshark workshop or a careful rewording of the empirical claims.","headline":"A genuinely useful design-research pictorial with two new techniques (light prompting, prompt fragments), but the inclusiveness claims outrun the reported evidence.","tokens_in":9638,"tokens_out":1572,"would_cite":true,"duration_ms":16492,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that interacting with diffusion-based AI image generation is best understood as prompt craft: iterative, embodied, material exploration of a latent possibility space, not prompt engineering.","keywords":["prompt craft","prompt engineering","diffusion models","latent space","tangible interaction","embodied interaction","generative AI","materiality"],"falsifier":"A controlled study comparing a prompt-craft interface (fragment cards or light prompting) with a conventional prompt text box for the same image-generation tasks would settle the claim: if users do not produce more novel, more consistent, or more personally meaningful outputs with the craft interface, the proposed advantage of prompt craft weakens.","tokens_in":8619,"feed_emoji":"🎨","tokens_out":6318,"duration_ms":61646,"temperature":0.7,"pith_summary":"This paper argues that working with diffusion-based image generators such as Stable Diffusion is better understood as prompt craft than as prompt engineering. The authors claim that interacting with these models is an iterative, embodied, material exploration: designers do not write commands so much as cultivate and constrain a latent possibility space, the space of all images the model can produce. To support the claim, they report three projects from 2024: Shadowplay V1, which uses participants' shadows and bright light ('light prompting') to shape the model's outputs; Cardshark, which uses physical prompt-fragment cards placed in a lightbox arena; and Shadowplay V2, which exploits real-time generation for embodied control of prompts, diffusion amount, and reactive audio. Their contributions are a perspective on the materiality of diffusion models with a craft-like method for navigating the latent space, and interaction design strategies for interfaces that support such navigation. If the claim holds, the design of generative-AI interfaces should shift from optimizing text prompts toward creating tangible, responsive contexts that expose and constrain the model's possibility space.","feed_headline":"AI image generation is a craft, not an engineering task","feed_subtitle":"Three design projects show that steering diffusion models means exploring and constraining a latent space, not optimizing prompts.","key_machinery":"The load-bearing mechanism is the latent possibility space of a diffusion model, the high-dimensional internal space of images the model can generate from random noise. The paper's central method is prompt craft, an iterative oscillation between creative exploration and systematic testing that finds and constrains a useful region of that space. Supporting techniques include light prompting (using shadows as high-contrast monochrome image input), prompt fragments with adjustable weightings, meta-prompt development, and real-time modulation of diffusion amount and dynamic text prompts.","core_discovery":"The central discovery is that the latent space of a diffusion model behaves as a workable material with 'soft edges,' and that this materiality can be navigated by craft-like, iterative practices. Across three installations, the researchers found that deliberately constraining the input—through high-contrast monochrome shadows, curated prompt fragments, and meta-prompts—makes the difference between chaotic output and a usable, engaging experience. The paper demonstrates that model choice, seed behavior, prompt terms, and their weightings co-evolve in a back-and-forth process of creative experimentation and systematic batch-testing. On this basis, the authors propose prompt craft: a method and mindset in which the practitioner shapes the latent possibility space rather than simply writing more elaborate textual instructions.","pith_inferences":["If prompt craft is a genuine material practice, then comparable latent-space navigation techniques should emerge for other generative modalities (audio, video, 3D), where users can constrain generation through embodied or tangible inputs rather than text.","A testable extension would be to measure whether craft-like interfaces increase users' sense of authorship and iterative discovery compared with conventional prompt text boxes in controlled studies.","The 'soft edges' metaphor points toward a design principle: interfaces that make latent-space constraints visible (for example, by showing how prompt weighting changes output) may help users develop mental models of the model's materiality.","The paper's anecdotal evidence invites replication: the Cardshark fragment-card interaction could be evaluated with other user groups to see whether the accessibility benefits generalise beyond the reported workshop."],"forward_implications":["AI interfaces should move beyond text boxes and sliders to tangible, embodied, and real-time controls that let users feel the model's possibility space.","Techniques like light prompting give designers a local, context-specific way to constrain aesthetics and content without modifying the underlying model.","Craft-like prompt development makes generative-AI interaction accessible to people who struggle with written language, including the workshop participants with young onset dementia.","Higher generation frame rates open up real-time embodied interactions, but the resulting complexity requires more sophisticated control interfaces and reactive feedback.","The paper's outcomes function as intermediate design knowledge, meant to inspire other designers rather than prescribe fixed steps."],"supporting_citations":[{"why":"Establishes the prior Entoptic Media Camera project that Shadowplay builds on, including the role of uncertainty in AI image-to-image generation.","marker":"[1]"},{"why":"Provides the emergence-focused strategies for practice-based design research that structure the paper's approach.","marker":"[5]"},{"why":"Defines 'strong concepts' as intermediate knowledge, the framing the paper uses for its contributions.","marker":"[7]"},{"why":"Supplies the StreamDiffusion real-time generation pipeline that makes Shadowplay V2's dynamic interactions possible.","marker":"[8]"},{"why":"Establishes the latent diffusion model and latent space concept that the paper's notion of a possibility space depends on.","marker":"[10]"}],"fun_headline_variants":["Diffusion models are material, and prompt craft shapes it","From prompt engineering to prompt craft, a new mindset","Latent space has soft edges, so craft your prompts","Craft beats engineering for steering AI image models","Prompt craft: navigating the latent space, not optimizing words"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that anecdotal, self-assessed observations from one workshop and two exhibition settings are enough to establish the value and transferability of prompt craft.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion models are material, and prompt craft shapes it","From prompt engineering to prompt craft, a new mindset","Latent space has soft edges, so craft your prompts","Craft beats engineering for steering AI image models","Prompt craft: navigating the latent space, not optimizing words"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000143,"raw_usage":{"total_tokens":1111,"prompt_tokens":827,"completion_tokens":284,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":206}},"tokens_in":443,"tokens_out":284,"duration_ms":4198,"temperature":1.0,"reasoning_tokens":206,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:24:56.969851+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled study comparing a prompt-craft interface (fragment cards or light prompting) with a conventional prompt text box for the same image-generation tasks would settle the claim: if users do not produce more novel, more consistent, or more personally meaningful outputs with the craft interface, the proposed advantage of prompt craft weakens.","supporting_citations":[],"review_version":1}