{"id":"63daf9b7-940c-460b-b62c-1c59fea67a63","arxiv_id":"2504.14320","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured, multimodal interface for generative AI that guides novice users to produce brand-aligned advertising briefs, with no direct evaluation in this paper.","lead":"This paper introduces ACAI, a multimodal AI tool that helps small business owners create advertising briefs through structured panels for brand identity, audience, and visual inspiration rather than free-form prompts. The work is grounded in interviews with six UK business owners, but the tool itself is not tested in this paper; an evaluation appears in a separate companion paper.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that ACAI improves alignment and co-creative control is not evidenced in this manuscript; evaluation is deferred to reference [9], so 'showing how structured interfaces can ... improve alignment' overstates what the paper demonstrates.","rationale":"The reader's weakest_assumption correctly identifies the central issue: the paper claims improvements in alignment and co-creative control without presenting an evaluation, instead deferring to a companion paper. My reading confirms this and adds precision: the prototype produces an Ad Brief rather than final visual assets, so the claimed benefit to advertisement outcomes is even more indirect. This is not an ad hominem critique; the authors are transparent about the deferral and about ACAI's current output format. The object-level design rationale is plausible and grounded in related work, and the formative study, while small, is acceptable as motivation for design requirements. However, the strongest claim as worded in the abstract and conclusion goes beyond what the manuscript can support. Since the reader's verdict was already CONDITIONAL and my concern aligns with the reader's weakest assumption, no change in the verdict is needed; the condition should be that the deferred evaluation in [9] actually demonstrates the claimed improvements, or that the paper's claims be revised accordingly.","tokens_in":7564,"tokens_out":2742,"duration_ms":27312,"concrete_test":"Retrieve companion paper [9] (arXiv:2503.06729) and check whether it reports a user evaluation with outcome measures for alignment and perceived control, ideally including a baseline condition such as free-form prompting or a conventional text-to-image interface. If [9] contains no such comparison, the current paper's central claim should be weakened from 'showing' improved alignment to 'proposing a design intended to improve' alignment. If [9] does provide a controlled comparison, the concern is resolved and the conditional verdict can be revisited.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is that directly implementing design requirements DR1 and DR2 in ACAI yields the claimed improvements in alignment and co-creative control for novice users. The paper does not supply evidence for this: Section 3 reports a six-participant formative study that identifies challenges, Section 4 describes the ACAI prototype, and Section 6 explicitly defers evaluation to a companion paper [9]. No user testing, baseline comparison, or outcome measures appear in the present manuscript. Consequently, the abstract and conclusion's phrase 'showing how structured interfaces can foreground user-defined context, improve alignment, and enhance co-creative control' is a design conjecture rather than a demonstrated result. There is also a gap between the claimed contribution and the artifact: ACAI generates a textual 'Ad Brief' (Section 4.1, Output Generation Layer), not the final visual advertisement, so improvements in advertisement alignment are asserted one step away from the output users would actually deploy. This is not an internal inconsistency, but it is a real correctness risk for the central claim as worded. The paper is transparent about the deferral, which is creditworthy, but transparency about missing evidence does not supply the evidence itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a formative study with six small business owners (SBOs) in the UK, identifies three challenges in using generative AI for advertising (articulating brand intuition, limited fine-grained adjustment, and generic outputs), derives two design requirements (DR1 multimodal input and interactive style selection; DR2 affordances and constraints for brand consistency), and presents ACAI, a multimodal generative AI prototype with Branding, Audience & Goals, and Inspiration Board panels. The abstract and conclusion claim that structured interfaces can foreground user-defined context, improve alignment, and enhance co-creative control; however, the manuscript explicitly defers evaluation of ACAI to a companion paper (reference [9]) and reports no user testing, baseline comparison, or outcome measures for the prototype itself. The paper is positioned as a design and system contribution for a CHI workshop.","tokens_in":7813,"tokens_out":2919,"duration_ms":27362,"significance":"If the claimed improvements are eventually validated, the paper offers a useful design direction for novice-oriented generative tools: replacing free-form prompting with structured panel inputs, multimodal references, and an inspiration board with segmented annotations and sliders. The qualitative findings about prompt articulation difficulties, brand misalignment, and lack of editability are plausible and consistent with prior HCI work on promptability and the gulf of envisioning. The main strength is a concrete, well-specified prototype and a clear mapping from study findings to design requirements. However, the central claim that ACAI demonstrably improves alignment and co-creative control is not supported in this manuscript; as written, the contribution is a design exploration plus a formative study, not an empirical demonstration. The paper is transparent about deferring evaluation, which is creditworthy, but transparency does not substitute for evidence.","major_comments":[{"comment":"The abstract states that the work contributes to HCI research by 'showing how structured interfaces can foreground user-defined context, improve alignment, and enhance co-creative control,' and Section 6 repeats that ACAI 'demonstrates how structured interfaces can foreground user-defined context to improve both alignment and promptability.' Yet Section 6 also defers the evaluation to reference [9], and the present manuscript contains no user test, baseline comparison, or outcome measure for ACAI. The word 'showing' and 'demonstrates' overstate what the evidence supports; the manuscript should either include an evaluation of ACAI or rephrase these claims as design proposals ('is designed to', 'may improve') rather than demonstrated results. This is load-bearing because the abstract's stated contribution depends on the claimed improvement.","section":"Abstract and Section 6 (Conclusion)"},{"comment":"ACAI generates a textual Ad Brief rather than the final visual advertisement, as the authors state: 'the generated Ad Brief functions as a guide for downstream visual production, whether by a designer or a future generative model.' The conclusion's claim of improved alignment and co-creative control for novice creative workflows is therefore asserted one step removed from the artifact users would actually deploy. The paper should explicitly acknowledge this gap when claiming alignment improvements, and ideally discuss how the Ad Brief is expected to translate into final visuals without losing the user's intent.","section":"Section 4.1, Output Generation Layer"},{"comment":"The design requirements are derived from a six-participant formative study using reflexive thematic analysis. Presenting DR1 and DR2 as provisional design implications is reasonable for a formative study, but the paper should avoid implying saturation or generalizability. In particular, the claim in Section 3.2 that 'AI tools should allow users to embed key elements of their brand identity' is presented as a general requirement even though it rests on a small, demographically narrow sample. The scope and exploratory nature of the study should be stated more prominently, and the requirements should be framed as hypotheses for future validation rather than established design mandates.","section":"Section 3 and Section 3.2"},{"comment":"The walkthrough illustrates how ACAI directly implements DR1 and DR2, which makes the statement that ACAI 'addresses' the identified challenges partly definitional. This is not a circular derivation in the formal sense, but it means the paper's contribution is constrained to the design logic rather than empirical validation. The authors should explicitly note that the prototype's success at addressing the challenges is an open question pending the evaluation in [9].","section":"Section 4.2 (walkthrough)"}],"minor_comments":[{"comment":"There is a typo: 'While these systems cproduce compelling outputs' should read 'While these systems produce compelling outputs.'","section":"Section 2.1"},{"comment":"The phrase 'a output that cannot be directly edited' should be 'an output that cannot be directly edited.'","section":"Section 3.1, Finding 3"},{"comment":"The capitalization of 'Ad brief' is inconsistent; it appears as 'Ad Brief' in Section 4.1 and 'Ad brief' in several places in Section 4.2. Please standardize.","section":"Section 4.1 and 4.2"},{"comment":"The sentence beginning 'She specifies Montserrat and Barlow as the preferred typefaces (Figure 3B), and uncertain about a preferred visual aesthetic...' appears garbled; it should be rewritten, e.g., 'She specifies Montserrat and Barlow as the preferred typefaces (Figure 3B). Since she is uncertain about a preferred visual aesthetic...'.","section":"Section 4.2"},{"comment":"Reference [9] contains a line break in the arXiv URL ('https://arxiv.org/abs/\\n2503.'); ensure the link is rendered as a single valid URL in the final version.","section":"References"},{"comment":"The age ranges in Table 1 are inconsistent and overlapping (e.g., '18–40', '24–40'); clarify the age brackets or report exact ages.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-style design paper whose central evaluative claim is deferred to a companion paper. The manuscript is honest about this deferral, and the design contributions are plausible and clearly presented. However, for a full archival venue the abstract and conclusion must be brought in line with the evidence actually presented, and the authors should either add a small evaluation of ACAI or explicitly recast the contribution as a design proposal. The paper should not be rejected outright because the gap is fixable through rephrasing and framing changes, but it is too large to be handled as a minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this before reading: it is a workshop paper that describes a prototype (ACAI) and a six-person formative study, but it does not evaluate the prototype. The abstract and conclusion say ACAI 'shows how structured interfaces can ... improve alignment, and enhance co-creative control.' That is not what the manuscript demonstrates. Section 6 explicitly sends you to companion paper [9] for the evaluation. So the central claim is a design conjecture, not a result. The paper is transparent about this, which is creditworthy, but transparency does not supply the evidence.\n\nWhat is actually new: the three-panel structured input system (Branding, Audience & Goals, Inspiration Board), the idea of extracting design annotations and adjustable sliders from reference images, and the application to small business owners (SBOs) as a user group. The formative study, though small, surfaces three plausible challenges: prompt articulation, lack of fine-grained control, and generic brand-misaligned outputs. The design requirements DR1 and DR2 follow from the interviews in a straightforward way. For a CHI workshop paper, this is a legitimate design research contribution. The demo walkthrough with the fictional brand Arrivl is clear and gives a concrete sense of how the interface works.\n\nSoft spots: the n=6 sample is thin for generalizing design requirements, and the authors do not overstate that; the limitation is implicit. The bigger issue is the gap between the claims and the tested artifact. ACAI outputs a textual 'Ad Brief,' not the final visual advertisement. So even if the brief aligns with brand and goals, the claim about advertisement alignment is one step removed. Also, since the prototype was built directly from the requirements, saying it 'addresses' those requirements is partially definitional. That is acceptable if framed as design rationale, less so if framed as validation. There are small copy issues (e.g., 'cproduce' in Section 2.1, odd age ranges in Table 1), but nothing damaging.\n\nThe citation pattern looks fine: PromptCrafter, ContextCam, and the gulf of envisioning are appropriate anchors. The self-citation to [9] is legitimate; the companion paper is the evaluation.\n\nWho is this for? Anyone in HCI working on prompt alternatives for novice users, especially in the creator economy. It deserves a serious referee, but the referee should insist that the abstract and conclusion be revised to match the evidence: this is a design proposal with a formative study, not a demonstrated improvement. If the companion paper delivers the evaluation, that would change the picture.\n\nRecommendation: engage with it; send it to peer review with a request for softened claims and a pointer to the evaluation.","headline":"A genuine design contribution with a clearly deferred evaluation; the abstract overclaims what the paper actually shows.","tokens_in":8265,"tokens_out":1419,"would_cite":false,"duration_ms":15996,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that replacing free-form prompts with structured multimodal panels lets novice users like small business owners communicate brand context to generative AI, improving alignment and preserving creative control.","keywords":["structured prompting","multimodal interfaces","novice users","small business owners","generative AI","co-creation","brand alignment","user interface design"],"falsifier":"Give the same small-business owners their own brand assets and a fixed design task, then have each work once through ACAI and once through a conventional text-prompt interface. If brand-alignment ratings by independent judges, or the number of regeneration cycles needed, are not clearly better with ACAI, the central claim that structured panels improve alignment and control would be contradicted.","tokens_in":7417,"feed_emoji":"🎨","tokens_out":8278,"duration_ms":67999,"temperature":0.7,"pith_summary":"This paper claims that the main barrier novice users face with generative AI is not model capability but the prompt interface. Working from interviews with six small-business owners, it argues that users know their brand intuition but cannot express it as text, receive generic misaligned outputs, and lose control because outputs cannot be finely adjusted. To show a way out, the paper presents ACAI, a prototype whose three structured panels let users supply brand assets, audience goals, and reference images; those inputs are fused into a 'super prompt' that produces an ad brief. The contribution is a design claim: structured multimodal interfaces can foreground user-defined context, improve alignment, and strengthen co-creative control in novice workflows.","feed_headline":"Structured panels, not prompts, can align AI with novice brand vision","feed_subtitle":"A three-panel interface distills brand assets, audience goals, and visual references into one 'super prompt' behind every ad brief.","key_machinery":"The load-bearing mechanism is ACAI's 'super prompt': a synthesized multimodal representation built in the processing layer from all three input panels. The panels do the conceptual work—Branding supplies colors, typography, values, and uploaded assets; Audience and Goals supplies objectives, audience segments, and emotional tone; Inspiration Board supplies reference images whose visual features are automatically segmented and annotated, with adjustable sliders for attributes like depth of field and light source. The 'super prompt' fuses textual attributes and image-derived descriptors into one structured instruction set that a multimodal large language model (a model that processes both text and images) converts into an Ad Brief with Summary, Background, and Foreground sections. Because users can edit inputs and regenerate, ACAI's role is to externalize tacit brand knowledge rather than to make the creative decision.","core_discovery":"On the paper's own terms, the discovery is a diagnosis plus a design response. The diagnosis, from a formative study of six UK small-business owners, is that prompt-based generative systems fail novices in three specific ways: brand intuition is hard to articulate as text; generated content comes out generic and off-brand; and single-shot outputs offer no path for fine-grained refinement. The design response, ACAI, replaces the free-text prompt with a panel-based input layer whose three sections encode branding, audience and goals, and visual inspiration; a multimodal language model combines these into a 'super prompt' that generates a structured Ad Brief. The paper's central claim is that this structured, multimodal interaction approach can foreground user-defined context, improve alignment, and enhance co-creative control for novice creative workflows, with evaluation explicitly deferred to a companion paper.","pith_inferences":["A direct within-subjects comparison between ACAI and a conventional text-prompt baseline, using the same users and brand materials, would be the clean way to test whether the claimed alignment improvement is real; the paper defers that evaluation.","If the panel approach works, it could be implemented as a prompt-engineering layer on top of existing text-to-image or text-to-content APIs, without modifying the underlying generative model.","The Inspiration Board's segmentation-and-annotation interaction may shift the burden from linguistic articulation to visual selection, which is promising for novices but could introduce new usability load that the paper does not measure.","Adapting the same structure to other non-expert creative contexts, such as local event posters or interior design, would clarify whether the benefit is brand-specific or a general property of structured prompting."],"forward_implications":["If the claim holds, novice users can shift from trial-and-error prompt writing to guided panel filling, lowering the cognitive load of articulating creative intent.","Brand assets and style references become fixed constraints in generation, so outputs are more likely to preserve logos, typefaces, palettes, and tone across iterations.","The Ad Brief as an intermediate artefact lets users review and approve the concept before any visual is produced, keeping a human decision point in the loop.","The three-panel structure offers a reusable template for other domains where users have tacit context but no design vocabulary."],"supporting_citations":[{"why":"It identifies the 'gulf of envisioning' between user intent and prompt wording that motivates ACAI's structured input design.","marker":"[14]"},{"why":"It shows that non-AI experts fail at prompt design, supplying the novice-user problem the paper addresses.","marker":"[15]"},{"why":"It frames 'promptability' as a core concern for generative systems, which the paper uses to position structured interfaces.","marker":"[11]"},{"why":"It demonstrates user-captured visual context guiding image generation, a precedent for the Inspiration Board interaction.","marker":"[8]"},{"why":"It introduces a dialogue-driven prompt crafting interface, the closest alternative design the paper distinguishes itself from.","marker":"[2]"},{"why":"It underlines the limitations of text-only prompting in text-to-image models, supporting the need for multimodal input.","marker":"[10]"},{"why":"It is the companion evaluation of ACAI with small business owners, to which this paper defers validation of its claimed benefits.","marker":"[9]"}],"fun_headline_variants":["Panels beat prompts: ACAI aligns AI to novice brand vision","Structured input, not prompts, helps novices steer AI ads","ACAI: a panel-based way for novices to control AI creativity","Three panels replace prompts for better AI ad alignment","For novices, structured interfaces outperform free-text prompts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the design requirements drawn from six interviews, once built into ACAI, will actually deliver the claimed improvements in alignment and creative control—an assumption that remains untested because the companion evaluation is not included here.","fun_headline_variants_meta":{"raw":{"variants":["Panels beat prompts: ACAI aligns AI to novice brand vision","Structured input, not prompts, helps novices steer AI ads","ACAI: a panel-based way for novices to control AI creativity","Three panels replace prompts for better AI ad alignment","For novices, structured interfaces outperform free-text prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1328,"prompt_tokens":894,"completion_tokens":434,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":349}},"tokens_in":510,"tokens_out":434,"duration_ms":3787,"temperature":1.0,"reasoning_tokens":349,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:50:38.490978+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same small-business owners their own brand assets and a fixed design task, then have each work once through ACAI and once through a conventional text-prompt interface. If brand-alignment ratings by independent judges, or the number of regeneration cycles needed, are not clearly better with ACAI, the central claim that structured panels improve alignment and control would be contradicted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It frames 'promptability' as a core concern for generative systems, which the paper uses to position structured interfaces."}],"review_version":1}