REVIEW 4 major objections 6 minor 17 references
Expanding the Generative AI Design Space through Structured Prompting and Multimodal Interfaces
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper argues that replacing free-form prompts with structured multimodal panels lets novice users like small business owners communicate brand context to generative AI, improving alignment and preserving creative control.
desk verdict A genuine design contribution with a clearly deferred evaluation; the abstract overclaims what the paper actually shows. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is ACAI's 'super prompt': a synthesized multimodal representation built in the processing layer from all three input panels. The panels do the conceptual work—Branding supplies colors, typography, values, and uploaded assets; Audience and Goals supplies objectives, audience segments, and emotional tone; Inspiration Board supplies reference images whose visual features are automatically segmented and annotated, with adjustable sliders for attributes like depth of field and light source. The 'super prompt' fuses textual attributes and image-derived descriptors into one structured instruction set that a multimodal large language model (a model that processes both text and images) converts into an Ad Brief with Summary, Background, and Foreground sections. Because users can edit inputs and regenerate, ACAI's role is to externalize tacit brand knowledge rather than to make the creative decision.
What would settle it
Give the same small-business owners their own brand assets and a fixed design task, then have each work once through ACAI and once through a conventional text-prompt interface. If brand-alignment ratings by independent judges, or the number of regeneration cycles needed, are not clearly better with ACAI, the central claim that structured panels improve alignment and control would be contradicted.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a diagnosis plus a design response. The diagnosis, from a formative study of six UK small-business owners, is that prompt-based generative systems fail novices in three specific ways: brand intuition is hard to articulate as text; generated content comes out generic and off-brand; and single-shot outputs offer no path for fine-grained refinement. The design response, ACAI, replaces the free-text prompt with a panel-based input layer whose three sections encode branding, audience and goals, and visual inspiration; a multimodal language model combines these into a 'super prompt' that generates a structured Ad Brief. The paper's central claim is that this structured, multimodal interaction approach can foreground user-defined context, improve alignment, and enhance co-creative control for novice creative workflows, with evaluation explicitly deferred to a companion paper.
Load-bearing premise
The load-bearing premise is that the design requirements drawn from six interviews, once built into ACAI, will actually deliver the claimed improvements in alignment and creative control—an assumption that remains untested because the companion evaluation is not included here.
Editorial extensions
If this is right
- If the claim holds, novice users can shift from trial-and-error prompt writing to guided panel filling, lowering the cognitive load of articulating creative intent.
- Brand assets and style references become fixed constraints in generation, so outputs are more likely to preserve logos, typefaces, palettes, and tone across iterations.
- The Ad Brief as an intermediate artefact lets users review and approve the concept before any visual is produced, keeping a human decision point in the loop.
- The three-panel structure offers a reusable template for other domains where users have tacit context but no design vocabulary.
Reading between the lines
- A direct within-subjects comparison between ACAI and a conventional text-prompt baseline, using the same users and brand materials, would be the clean way to test whether the claimed alignment improvement is real; the paper defers that evaluation.
- If the panel approach works, it could be implemented as a prompt-engineering layer on top of existing text-to-image or text-to-content APIs, without modifying the underlying generative model.
- The Inspiration Board's segmentation-and-annotation interaction may shift the burden from linguistic articulation to visual selection, which is promising for novices but could introduce new usability load that the paper does not measure.
- Adapting the same structure to other non-expert creative contexts, such as local event posters or interior design, would clarify whether the benefit is brand-specific or a general property of structured prompting.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes a formative study with six small business owners (SBOs) in the UK, identifies three challenges in using generative AI for advertising (articulating brand intuition, limited fine-grained adjustment, and generic outputs), derives two design requirements (DR1 multimodal input and interactive style selection; DR2 affordances and constraints for brand consistency), and presents ACAI, a multimodal generative AI prototype with Branding, Audience & Goals, and Inspiration Board panels. The abstract and conclusion claim that structured interfaces can foreground user-defined context, improve alignment, and enhance co-creative control; however, the manuscript explicitly defers evaluation of ACAI to a companion paper (reference [9]) and reports no user testing, baseline comparison, or outcome measures for the prototype itself. The paper is positioned as a design and system contribution for a CHI workshop.
Significance. If the claimed improvements are eventually validated, the paper offers a useful design direction for novice-oriented generative tools: replacing free-form prompting with structured panel inputs, multimodal references, and an inspiration board with segmented annotations and sliders. The qualitative findings about prompt articulation difficulties, brand misalignment, and lack of editability are plausible and consistent with prior HCI work on promptability and the gulf of envisioning. The main strength is a concrete, well-specified prototype and a clear mapping from study findings to design requirements. However, the central claim that ACAI demonstrably improves alignment and co-creative control is not supported in this manuscript; as written, the contribution is a design exploration plus a formative study, not an empirical demonstration. The paper is transparent about deferring evaluation, which is creditworthy, but transparency does not substitute for evidence.
major comments (4)
- [Abstract and Section 6 (Conclusion)] The abstract states that the work contributes to HCI research by 'showing how structured interfaces can foreground user-defined context, improve alignment, and enhance co-creative control,' and Section 6 repeats that ACAI 'demonstrates how structured interfaces can foreground user-defined context to improve both alignment and promptability.' Yet Section 6 also defers the evaluation to reference [9], and the present manuscript contains no user test, baseline comparison, or outcome measure for ACAI. The word 'showing' and 'demonstrates' overstate what the evidence supports; the manuscript should either include an evaluation of ACAI or rephrase these claims as design proposals ('is designed to', 'may improve') rather than demonstrated results. This is load-bearing because the abstract's stated contribution depends on the claimed improvement.
- [Section 4.1, Output Generation Layer] ACAI generates a textual Ad Brief rather than the final visual advertisement, as the authors state: 'the generated Ad Brief functions as a guide for downstream visual production, whether by a designer or a future generative model.' The conclusion's claim of improved alignment and co-creative control for novice creative workflows is therefore asserted one step removed from the artifact users would actually deploy. The paper should explicitly acknowledge this gap when claiming alignment improvements, and ideally discuss how the Ad Brief is expected to translate into final visuals without losing the user's intent.
- [Section 3 and Section 3.2] The design requirements are derived from a six-participant formative study using reflexive thematic analysis. Presenting DR1 and DR2 as provisional design implications is reasonable for a formative study, but the paper should avoid implying saturation or generalizability. In particular, the claim in Section 3.2 that 'AI tools should allow users to embed key elements of their brand identity' is presented as a general requirement even though it rests on a small, demographically narrow sample. The scope and exploratory nature of the study should be stated more prominently, and the requirements should be framed as hypotheses for future validation rather than established design mandates.
- [Section 4.2 (walkthrough)] The walkthrough illustrates how ACAI directly implements DR1 and DR2, which makes the statement that ACAI 'addresses' the identified challenges partly definitional. This is not a circular derivation in the formal sense, but it means the paper's contribution is constrained to the design logic rather than empirical validation. The authors should explicitly note that the prototype's success at addressing the challenges is an open question pending the evaluation in [9].
minor comments (6)
- [Section 2.1] There is a typo: 'While these systems cproduce compelling outputs' should read 'While these systems produce compelling outputs.'
- [Section 3.1, Finding 3] The phrase 'a output that cannot be directly edited' should be 'an output that cannot be directly edited.'
- [Section 4.1 and 4.2] The capitalization of 'Ad brief' is inconsistent; it appears as 'Ad Brief' in Section 4.1 and 'Ad brief' in several places in Section 4.2. Please standardize.
- [Section 4.2] The sentence beginning 'She specifies Montserrat and Barlow as the preferred typefaces (Figure 3B), and uncertain about a preferred visual aesthetic...' appears garbled; it should be rewritten, e.g., 'She specifies Montserrat and Barlow as the preferred typefaces (Figure 3B). Since she is uncertain about a preferred visual aesthetic...'.
- [References] Reference [9] contains a line break in the arXiv URL ('https://arxiv.org/abs/\n2503.'); ensure the link is rendered as a single valid URL in the final version.
- [Table 1] The age ranges in Table 1 are inconsistent and overlapping (e.g., '18–40', '24–40'); clarify the age brackets or report exact ages.
Circularity Check
ACAI's 'addresses these challenges' claim is a definitional consequence of deriving DR1/DR2 from the study and building ACAI to implement them; evaluation is deferred to a same-author companion.
-
self definitional
[Section 3.2 (Design Requirements) -> Section 4.1 (System Overview) -> Section 6 (Conclusion)]
"ACAI addresses these challenges through a panel-based interface that enables users to specify branding elements, define audience goals, and select visual inspirations, facilitating more directed and brand-consistent co-creation."
The three challenges in Section 3.1 (prompt articulation, fine-grained adjustment, brand specificity) are translated into design requirements DR1 and DR2 in Section 3.2; Section 4 then says 'In response to these requirements, we developed ACAI' and describes panels that implement DR1/DR2. The conclusion's claim that ACAI 'addresses these challenges through a panel-based interface' is therefore entailed by construction: the features named (branding elements, audience goals, visual inspirations) are the requirements derived from the challenges. No outcome measure, baseline, or user test of alignment or co-creative control appears in this manuscript, so the abstract's 'showing how structured interfaces can ...
full rationale
The paper contains no equations, no fitted parameters, and no quantitative prediction, so the strongest circularity patterns do not apply. The formative study is a qualitative interview with six SBOs, and the design requirements are the authors' translation of those findings. The main circularity risk is that the prototype is claimed to 'address' the challenges because it was built to satisfy requirements that were themselves written from those challenges: challenge -> DR1/DR2 -> ACAI panels -> 'ACAI addresses these challenges'. This is a mild design-level tautology, not a statistical or mathematical reduction. The self-citation [9] is used transparently as a pointer to a companion evaluation, not as an imported uniqueness theorem or fitted parameter; it is not load-bearing in the sense of proving a derived result, but it does mean the evaluative part of the central contribution is not contained in this manuscript. No renaming of known results or ansatz-smuggling is present. The score of 3 reflects one partially definitional step while acknowledging that the design contribution is self-contained as a proposal and the paper is candid about deferring evaluation.
Assumptions & free parameters
assumptions (3)
- domain assumption Qualitative themes from six SBOs generalize to the broader small business owner population.
- domain assumption The Gemini 1.5 Pro multimodal model can reliably parse structured inputs and generate useful ad briefs.
- ad hoc to paper The 'super prompt' synthesis preserves user intent and brand constraints effectively.
Cite this review
Pith. "Pith review of Expanding the Generative AI Design Space through Structured Prompting and Multimodal Interfaces." pith.science (2026). https://pith.science/paper/OWDQNPAR
@misc{pith2026250414320,
author = {Pith},
title = {Pith review of: Expanding the Generative AI Design Space through Structured Prompting and Multimodal Interfaces},
year = {2026},
howpublished = {\url{https://pith.science/paper/OWDQNPAR}},
note = {Machine review of arXiv:2504.14320}
}
read the original abstract
Text-based prompting remains the predominant interaction paradigm in generative AI, yet it often introduces friction for novice users such as small business owners (SBOs), who struggle to articulate creative goals in domain-specific contexts like advertising. Through a formative study with six SBOs in the United Kingdom, we identify three key challenges: difficulties in expressing brand intuition through prompts, limited opportunities for fine-grained adjustment and refinement during and after content generation, and the frequent production of generic content that lacks brand specificity. In response, we present ACAI (AI Co-Creation for Advertising and Inspiration), a multimodal generative AI tool designed to support novice designers by moving beyond traditional prompt interfaces. ACAI features a structured input system composed of three panels: Branding, Audience and Goals, and the Inspiration Board. These inputs allow users to convey brand-relevant context and visual preferences. This work contributes to HCI research on generative systems by showing how structured interfaces can foreground user-defined context, improve alignment, and enhance co-creative control in novice creative workflows.
Figures
Reference graph
Works this paper leans on
-
[9]
Nimisha Karnatak, Adrien Baranes, Rob Marchant, Triona Butler, and Kristen Olson. 2025. ACAI for SBOs: AI Co-creation for Advertising and Inspiration for Small Business Owners. arXiv:2503.06729 [cs.HC] https://arxiv.org/abs/2503. 06729
arXiv 2025
-
[1]
Bruce Croft
Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani, and W. Bruce Croft
-
[2]
Seungho Baek, Hyerin Im, Jiseung Ryu, Juhyeong Park, and Takyeon Lee. 2023. PromptCrafter: Crafting Text-to-Image Prompt through Mixed-Initiative Dia- logue with LLM. arXiv preprint arXiv:2307.08985 (2023). https://arxiv.org/abs/ 2307.08985
arXiv 2023
-
[3]
Joshua Ballard. 2016. 6 Hats Worn By Business Owners — paradoxmar- keting.io. https://paradoxmarketing.io/capabilities/knowledge-management/ insights/6-hats-worn-by-business-owners/. [Accessed 02-09-2024]
work page 2016
-
[4]
Kuldeep Bhalerao, Anuj Kumar, Arya Kumar, and Purvi Pujari. 2022. A study of barriers and benefits of artificial intelligence adoption in small and medium enterprise. Academy of Marketing Studies Journal 26 (2022), 1–6
work page 2022
-
[5]
Annie Chen. 2005. Context-Aware Collaborative Filtering System: Predicting the User’s Preferences in Ubiquitous Computing. In CHI ’05 Extended Abstracts on Human Factors in Computing Systems . ACM, 1110–1111
work page 2005
-
[6]
Fabio Clarizia, Francesco Colace, Massimo De Santo, Marco Lombardi, Francesco Pascale, and Domenico Santaniello. 2019. A Context-Aware Chatbot for Tourist Destinations. In 15th International Conference on Signal-Image Technology and Internet-Based Systems (SITIS) . 348–354
work page 2019
-
[7]
Anind K. Dey. 2001. Understanding and Using Context. Personal and Ubiquitous Computing 5, 1 (2001), 4–7
work page 2001
Show all 17 references
-
[8]
Xianzhe Fan, Zihan Wu, Chun Yu, Fenggui Rao, Weinan Shi, and Teng Tu. 2024. ContextCam: Bridging Context Awareness with Creative Human-AI Image Co- Creation. In Proceedings of the 2024 CHI Conference on Human Factors in Comput- ing Systems. Association for Computing Machinery....
2024
-
[10]
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu. 2023. Aligning text-to-image models using human feedback. arXiv preprint arXiv:2302.12192 (2023)
2023 arXiv
-
[11]
Meredith Ringel Morris. 2024. Prompting considered harmful. Commun. ACM 67, 12 (2024), 28–30
2024
-
[12]
June 2024
United Nations. June 2024. Micro-,Small and Medium-Sized Enterprises Report. https://www.un.org/sites/un2.un.org/files/globalmsmesreport2024.pdf
2024
-
[13]
Quinn, and Karthik Ramani
Jingyu Shi, Rahul Jain, Seungguen Chi, Hyungjun Doh, Hyunggun Chi, Alexan- der J. Quinn, and Karthik Ramani. 2025. CARING-AI: Towards Authoring Context- aware Augmented Reality Instruction through Generative Artificial Intelligence. In Proceedings of the 2025 CHI Conference on...
2025
-
[14]
Hari Subramonyam, Roy Pea, Christopher Pondoc, Maneesh Agrawala, and Colleen Seifert. 2024. Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMs. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–19
2024
-
[15]
JD Zamfirescu-Pereira, Richmond Y Wong, Bjoern Hartmann, and Qian Yang
-
[2021]
ACM Transactions on Information Systems 39, 3 (2021), 29:1–29:30
Context-Aware Target Apps Selection and Recommendation for Enhancing Personal Mobile Assistants. ACM Transactions on Information Systems 39, 3 (2021), 29:1–29:30
2021
-
[2023]
In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems
Why Johnny can’t prompt: how non-AI experts try (and fail) to design LLM prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–21
2023
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.