{"id":"b4835049-b1ab-4b5d-8fe6-a65db5ed1352","arxiv_id":"2506.03741","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A widget-based canvas for AI writing prompts was rated as more creativity-supportive and less mentally demanding than a conversational chat UI in a lab study (N=18) and a follow-up field study (N=10).","lead":"PromptCanvas is a new writing interface that turns AI prompts into movable, customizable widgets placed on an infinite canvas. In two small studies, writers rated it higher on creativity support and lower on mental demand and frustration than a chat-style interface.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The relative-support claim rests on a single lab comparison whose baseline is underspecified (temperature, system prompts, output completeness) and whose widget condition also adds non-widget features, so the CSI/TLX gains may not be attributable to dynamic widgets.","rationale":"The reader's CONDITIONAL verdict is appropriate. The most load-bearing assumption for the central claim is that the lab comparison isolates the dynamic-widget interface from other differences. I agree with the reader that baseline fairness is the key risk, and I sharpen it: the baseline's backend behavior (temperature, system prompts, output completeness) is not reported, and the experimental condition adds direct text editing, canvas navigation, and history—features absent from the chat condition. These are not merely cosmetic; they are part of the system's interaction model and could alone explain lower mental demand and higher CSI (e.g., direct editing reduces the need to re-prompt; persistent widgets preserve context that chat does not). The paper's own participant quotes show the static UI 'cut off' previous text, suggesting the baseline may be a particularly weak instantiation of a chat UI rather than a representative one. The field study cannot rescue this because it has no baseline. The concrete test would settle whether the widget concept itself carries the effect; if it does not, the paper should be reframed. None of this impugns the authors' integrity; it is a standard internal-validity concern in system comparisons. The system is novel and the study transparently reports limitations, so the verdict remains CONDITIONAL rather than REJECT.","tokens_in":23737,"tokens_out":10044,"duration_ms":93614,"concrete_test":"Run a three-arm within-subject study with matched backend settings (gpt-4o-2024-08-06, temperature 1.06, same system prompts): (A) current PromptCanvas, (B) PromptCanvas with widgets disabled but the same canvas/text editor/history (or equivalently, a chat-in-canvas control), and (C) the current chat baseline. If A does not significantly beat B on CSI or NASA-TLX, the RQ2/RQ3 attribution to dynamic widgets is unsupported and the paper should be reframed as a system-level comparison. Also report the baseline's system prompts and temperature.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result depends entirely on the lab study (N=18, §5.3), but the two conditions differ in many ways beyond the widget concept. PromptCanvas (§4.1) includes an infinite canvas, a directly editable text editor with history and word counts, persistent interactive widgets, and a rephrasing backend whose prompt explicitly requires returning the complete modified text (§4.2.4–4.2.5). The baseline (§5.2) is a custom chat UI with no direct editing of generated text (only copy), and the paper reports neither the baseline's system prompts nor its sampling temperature; for PromptCanvas the temperature is stated as 1.06 (§4.2). If the baseline used a different temperature or allowed partial/truncated regeneration, the observed CSI and NASA-TLX differences (§6.2, §6.3) could be inflated. Participant P12's report that 'some of my previous texts were being cut off' in the static UI (§6.1.6) is consistent with a backend difference rather than an inherent property of chat UIs. The field study (§5.4) has no baseline, so it cannot confirm relative superiority as the abstract claims. Consequently, RQ2/RQ3 ('Do dynamic widgets ... improve ...?') are not answered by this design; only a system-level comparison is shown, and that only if the custom baseline is accepted as representative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PromptCanvas, a workspace for LLM-based creative writing in which prompts are decomposed into interactive 'dynamic widgets' placed on an infinite canvas. Widgets can be system-suggested, user-prompted, or manually created, and each widget controls an attribute of the generated text (tone, names, plot elements, etc.). The authors report a within-subject lab study (N=18) comparing PromptCanvas with a custom chat-based 'conversational UI', finding significantly higher Creativity Support Index scores and lower NASA-TLX mental demand and frustration for PromptCanvas, alongside qualitative reports of improved exploration and control. A two-week field study (N=10) without a baseline reports high CSI and SUS scores. The paper frames these results as answering RQ2/RQ3, i.e., that dynamic widgets improve creativity support and reduce cognitive load relative to conversational UIs.","tokens_in":23980,"tokens_out":3108,"duration_ms":34716,"significance":"If the comparative claims held, PromptCanvas would be a useful design pattern for LLM writing tools: modular, persistent widgets that make prompt attributes visible and editable could support decomposition of tasks, iterative exploration, and a greater sense of control. The paper's strengths include a within-subject, counterbalanced lab design using the same underlying LLM (gpt-4o-2024-08-06) in both conditions, standardized instruments (CSI, NASA-TLX, SUS), Bonferroni-Holm correction, and a clearly described system implementation. The qualitative analysis is rich and gives concrete insight into how users appropriate widget-based interfaces. However, the central comparative claim is undermined by confounds between conditions and by an insufficiently specified baseline, so the significance of the specific RQ2/RQ3 conclusions is currently provisional.","major_comments":[{"comment":"The baseline conversational UI is not specified to the same level as PromptCanvas. For PromptCanvas the paper states a sampling temperature of 1.06 (§4.2) and gives the full system prompts for each backend service (Tables 4–8 referenced in §4.2.1–4.2.5), including an explicit instruction that natural-language editing must return the complete modified text (§4.2.5). For the baseline, the paper reports neither the temperature nor the system prompt, nor how regeneration/rephrasing handles the previous message. Without matched parameters, the significant differences in CSI and NASA-TLX (§6.2, §6.3) could be driven by backend configuration rather than by the dynamic-widget concept. Please report the baseline's prompts, temperature, and output-completeness behavior, or run a comparison with matched settings.","section":"§5.2 and §4.2"},{"comment":"The two conditions differ in many features beyond dynamic widgets: PromptCanvas includes an infinite canvas, a directly editable text editor with history and word counts, a widget panel, and a dedicated 'rephrase based on widgets' action, while the baseline chat UI offers only copy and edit-message functionality. Consequently, RQ2 and RQ3 ('Do dynamic widgets ... improve creativity support / reduce cognitive load?') are not answered by this design; at best the study shows a system-level comparison. To support the widget-specific claim, the authors should either add a condition that isolates the widgets (e.g., a chat UI augmented with the same editor and history features, or a widget-based UI without the canvas), or explicitly reframe RQ2/RQ3 as comparisons of PromptCanvas as a whole against a chat UI and temper the causal language accordingly.","section":"§5.2, §4.1, and §6.2–6.3"},{"comment":"The abstract states that the field study (N=10) 'confirmed these results', but the field study has no baseline condition and only measures PromptCanvas itself (CSI and SUS, §5.4.1, Table 3). It cannot confirm relative superiority over a conversational UI. The relative claim rests entirely on the 18-participant lab comparison. The authors should either add a comparative baseline to the field study or revise the abstract and §5.4 to state that the field study provided further qualitative and usability evidence for PromptCanvas, not confirmation of the comparative advantage.","section":"Abstract and §5.4"},{"comment":"Participant P12 is quoted as saying that in the static UI 'some of my previous texts were being cut off' when regenerating text. This suggests a functional deficiency or different regeneration behavior in the baseline, rather than an inherent property of chat UIs. If the baseline's regeneration replaced only part of the message or truncated it, this would bias the mental-demand and frustration ratings in favor of PromptCanvas. The manuscript should clarify how the baseline handled regeneration and, if the behavior was not the intended one, treat this as a confound rather than as evidence about the interface paradigm.","section":"§6.1.6"}],"minor_comments":[{"comment":"For the significant NASA-TLX differences, only p-values are reported. Please add effect sizes (e.g., Cohen's d or rank-biserial correlation) and confidence intervals for the CSI and TLX comparisons, especially given the modest sample size.","section":"§6.3.1"},{"comment":"The paper describes the baseline as 'designed according to the design and interaction principles of ChatGPT' but does not provide a screenshot annotation of its regeneration/edit behavior. A short description of what happens when a user edits a message or regenerates would help readers assess the fairness of the baseline.","section":"§5.2 and Figure 12"},{"comment":"The limitations section mentions sample size and generalizability but does not acknowledge the confounds between the two conditions or the lack of a field-study baseline. Consider adding these as explicit limitations.","section":"§7.1"},{"comment":"Reference [1] is formatted incorrectly ('Philip T. Kortum Aaron Bangor' should be 'Aaron Bangor, Philip T. Kortum, and James T. Miller'). Also, references [37] and [38] both list the same Luminate paper; one duplicate should be removed.","section":"References"},{"comment":"The stacked percentage bars in Figures 14 and 16 sum to over 100% within some rows (e.g., 'Which tool made you feel hurried or rushed' shows 39% + 39% + 33% = 111%). Please clarify the response format or adjust the visualization so readers can correctly interpret the preference data.","section":"Figures 14 and 16"}],"recommendation":"major_revision","confidential_remarks":"The paper would be strengthened by repositioning itself as a system-and-qualitative-insights paper, with the quantitative comparative claims explicitly scoped to the specific PromptCanvas implementation versus a particular chat UI. The current abstract overclaims relative to the field study, and the confounds between conditions make the causal statements about 'dynamic widgets' hard to defend. If the authors can provide a cleaner comparison (e.g., matched baseline prompts/parameters and an additional condition that holds non-widget features constant) or carefully rewrite the claims, the contribution could be publishable. The participating author's own prior work appears adequately cited; I did not see an issue there."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The gist: PromptCanvas is a solid, well-scoped system paper. It brings DynaVis-style dynamic widgets from visualization into creative writing and evaluates the result with a reasonably designed within-subject lab study against a ChatGPT-like chat UI. The main claim, that the widget canvas beats the chat UI on perceived creativity support and mental demand, looks credible as a system-level result. It is not, despite the RQ wording, a clean demonstration that dynamic widgets are the cause.\n\nWhat is actually new: the domain transfer, plus the empirical comparison. The paper is honest about its ingredients—canvas UIs, direct manipulation, structured prompting, and LLM suggestions all appear in prior work, and they cite that work. The engineering is described in enough detail to reconstruct the system: five backend services, JSON schemas, and explicit prompt templates in the supplementary material. The study uses the same model in both conditions, counterbalanced order, Bonferroni-Holm correction, and standardized instruments (CSI, NASA-TLX, SUS). The qualitative data are genuinely useful; they give the mechanism: widgets act as visible, persistent prompts, they break the text into pieces, and they preserve context across iterations. The prompt-count difference (4 vs 11) is a solid behavioral outcome.\n\nSoft spots. The comparison is 18 convenience participants, all self-report except prompt count. More importantly, PromptCanvas and the baseline differ on several dimensions at once: infinite canvas, direct editing with history, word counts, a widget panel, and a rephrasing backend whose system prompt explicitly demands the complete modified text. The baseline is a custom chat UI with copy-only output, and the paper does not report its system prompt or sampling temperature, while PromptCanvas states 1.06. P12's complaint that previous texts got cut off in the baseline suggests part of the advantage may come from the backend's completion behavior, not from widgets. So the RQ2/RQ3 conclusions should be softened to \"the PromptCanvas system as a whole outperformed this particular chat UI.\" The two-week field study has no baseline; calling it a confirmation of relative superiority is too strong—it confirms acceptance and usability. No effect sizes, no released code or data.\n\nI don't think any of this sinks the paper. The baseline is not a strawman; it is just under-specified. A serious reviewer should ask for the missing baseline details, effect sizes, and a more careful causal attribution, and the authors should temper the abstract. The paper deserves a real review.\n\nMy recommendation: send it to peer review, and if you work on LLM writing interfaces, cite it as a useful data point. I'd bring it to reading group.","headline":"A solid system-level evaluation of dynamic widgets for creative writing, with an under-specified baseline and overclaimed causal framing; deserves peer review.","tokens_in":24508,"tokens_out":2290,"would_cite":true,"duration_ms":23830,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PromptCanvas turns prompts into persistent, adjustable widgets on an infinite canvas, and in a lab study this interface outperformed a conversational UI on the Creativity Support Index while reducing mental demand and frustration.","keywords":["dynamic widgets","prompting","large language models","human-AI co-creation","creativity support","cognitive load","creative writing","conversational UI"],"falsifier":"Run a larger preregistered within-subject study in which the chat baseline is augmented with persistent prompt history, editable previous messages, and one-click suggestion options; if the roughly 20-point Creativity Support Index gap and the NASA-TLX differences disappear or reverse, the claim that dynamic widgets themselves improve creative writing support would be falsified.","tokens_in":23529,"feed_emoji":"🧩","tokens_out":9321,"duration_ms":98834,"temperature":0.7,"pith_summary":"The paper aims to show that how writers prompt a large language model is itself a design problem: prompts that are turned into persistent, adjustable widgets on an infinite canvas support creativity and reduce cognitive load compared with ordinary chat turns. In a within-subject lab study with 18 participants, PromptCanvas beat a ChatGPT-style conversational UI on every factor of the Creativity Support Index, with an overall score of 82.09 versus 61.65 (p = 0.005), and on NASA-TLX it produced significantly lower mental demand and frustration. A two-week field study with 10 participants yielded a similar overall creativity score (79.73) and a System Usability Scale average of 86.50, suggesting the effect is not limited to a single lab session. A sympathetic reader would take this as evidence that widget-based decomposition of prompts is a better interface pattern for human-AI co-writing than free-form chat.","feed_headline":"Widget canvas beats chat UI for AI-assisted writing","feed_subtitle":"Users scored it 82 vs 62 on the Creativity Support Index, with less mental demand and more exploration.","key_machinery":"The load-bearing object is the dynamic widget: a small interactive panel with a title, a current value, and a list of alternative values, tied to one attribute of the text under revision. It carries the argument by converting an ephemeral prompt into a persistent, manipulable interface object: the label-value pairs of all active widgets are sent together with the draft to the same language model that generates the text, so the user's specifications become reusable state rather than one-shot chat history. The surrounding mechanism is the infinite canvas and the widget panel, which let users generate, arrange, cluster, and discard widgets without losing the workflow, and the rephrasing pipeline that streams revised text back into the editor.","core_discovery":"PromptCanvas's central claim is that dynamic widgets make prompting composable: each widget represents one attribute of the text (such as tone, length, or a character name), carries a current value and a set of alternatives, and can be created from system suggestions, a user prompt, or an empty double-click on the canvas. Once placed on the canvas, active widgets are converted into label-value pairs and sent to the language model to rephrase the draft, so the prompt does not evaporate after one exchange but remains visible and editable as the writing evolves. The paper reports that this design outperformed a conversational baseline in the lab, with an overall Creativity Support Index of 82.09 against 61.65, significant differences favoring PromptCanvas on mental demand (1.89 vs 3.06, p = 0.02) and frustration (1.28 vs 2.17, p = 0.03), and far fewer prompts needed (4.0 vs 11.1, p = 0.0006). Eighty-nine percent of participants preferred PromptCanvas, and the two-week field study echoed the creativity results, with participants also using the canvas for programming and multilingual writing.","pith_inferences":["A factorial follow-up could separate the effects of persistence, spatial layout, and suggestion content; this study does not isolate which component of the widget format carries the creativity gain.","The widget pattern could be extended to code generation and image generation, where attributes such as model, seed, or style could become widgets on a canvas instead of arguments buried in a prompt.","Because widget values and layouts are structured, they could be logged and shared as reusable prompt workflows, turning prompt engineering from throwaway utterances into compositional, named components."],"forward_implications":["If PromptCanvas is right, chat is not the default interface for LLM writing support; persistent widgetized prompts are a more effective pattern for open-ended creative tasks.","Users needing roughly one-third the number of prompts (4.0 vs 11.1) implies widget canvases can compress iterative refinement into fewer, more targeted model calls.","Lower mental demand and frustration on NASA-TLX implies the visual persistence of prompt variables reduces metacognitive load, not just user preference.","The field-study use of widgets for programming and non-English writing suggests the pattern generalizes beyond creative writing to other LLM-assisted workflows."],"supporting_citations":[{"why":"It supplies the dynamic-widget concept that PromptCanvas transplants from visualization editing to text prompting.","marker":"[42]"},{"why":"It frames the metacognitive demands of LLM prompting that persistent visible widgets are designed to reduce.","marker":"[40]"},{"why":"It defines the dearth-of-the-author problem that motivates moving away from static text-field prompting.","marker":"[24]"},{"why":"It provides the structured-generation design space and task topics that the widgets and the lab tasks draw on.","marker":"[38]"},{"why":"It provides the Creativity Support Index factor structure used to measure the headline creativity comparison.","marker":"[13]"},{"why":"It supplies the direct manipulation principle behind turning prompts into editable interface objects.","marker":"[35]"}],"fun_headline_variants":["Widget canvas lifts AI writing creativity over chat","PromptCanvas widgets beat chat UI in creativity test","Interactive widgets outperform chat for AI-assisted writing","Canvas widgets reduce mental load in AI writing","Composable widgets enhance AI writing exploration"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparative claim rests on the 18-person lab study and on the assumption that the ChatGPT-style chat interface built as the baseline is a fair, representative control rather than a deliberately plain one; the two-week field study had no baseline, so it cannot independently support the relative advantage.","fun_headline_variants_meta":{"raw":{"variants":["Widget canvas lifts AI writing creativity over chat","PromptCanvas widgets beat chat UI in creativity test","Interactive widgets outperform chat for AI-assisted writing","Canvas widgets reduce mental load in AI writing","Composable widgets enhance AI writing exploration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000151,"raw_usage":{"total_tokens":1191,"prompt_tokens":925,"completion_tokens":266,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":200}},"tokens_in":541,"tokens_out":266,"duration_ms":3652,"temperature":1.0,"reasoning_tokens":200,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:56:09.846621+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a larger preregistered within-subject study in which the chat baseline is augmented with persistent prompt history, editable previous messages, and one-click suggestion options; if the roughly 20-point Creativity Support Index gap and the NASA-TLX differences disappear or reverse, the claim that dynamic widgets themselves improve creative writing support would be falsified.","supporting_citations":[{"cited_title":"Carroll, Celine Latulipe, Richard Fung, and Michael Terry","cited_arxiv_id":null,"evidence_quote":"It provides the Creativity Support Index factor structure used to measure the headline creativity comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the direct manipulation principle behind turning prompts into editable interface objects."}],"review_version":1}