{"id":"848f86e1-65ec-4dae-b9d5-da6f17ee54b7","arxiv_id":"2412.02357","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Users preferred dynamically generated prompt-refinement controls over a fixed preset list when steering AI explanations, reporting more control and lower context-providing barriers, despite difficulty predicting option effects.","lead":"This paper tests whether AI chat tools should generate new control buttons based on what you type, rather than showing the same fixed options each time. In a 16-person lab study, users preferred the generated, context-specific controls and said they made explanations easier to steer, though predicting how each control changes the answer remained hard.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Static PRC baseline is confounded: Section 4.3 excludes participant-requested controls and offers only six generic options, so the Dynamic PRC preference may reflect option coverage rather than dynamic generation; a Static PRC-Plus condition is needed to isolate the mechanism.","rationale":"The reader's weakest assumption identifies exactly the concern I find most load-bearing: whether the Static PRC preset list is a fair, non-degenerate baseline. The paper's own Section 4.3 admits that several options requested in the formative survey were excluded because they required dynamic generation or interpretation, and Appendix A.3 shows a small, generic set. This means the comparison does not isolate 'dynamic generation' as the causal factor; it compares a task-adaptive generator against a fixed, limited list. The Task 4 result in Section 6.2.2, where Static PRC significantly outperformed Dynamic PRC, reinforces that the static list is not uniformly appropriate and that task-specific adequacy matters. I considered the statistical concerns raised by the reader (Mann-Whitney U on paired data, multiple comparisons) and they are real, but they affect the strength of the quantitative evidence rather than the internal validity of the comparison itself. The baseline confound threatens the central claim more directly: if the static list is weak, the observed preference for Dynamic PRC could be explained without invoking any benefit of dynamic generation. The proposed Static PRC-Plus condition is a concrete way to settle this by expanding the static arm to include all operationally static controls and expert-selected options, while keeping the same protocol and using the correct paired test. Until that check is run, the conditional verdict is appropriate: the qualitative findings are suggestive and the system is well described, but the empirical claim of dynamic advantage is not yet isolated from baseline coverage. I therefore agree with the reader and keep the verdict unchanged.","tokens_in":33267,"tokens_out":4568,"duration_ms":53414,"concrete_test":"Run a follow-up within-subjects study (n=16, same tasks/protocol) with three arms: Dynamic PRC, original Static PRC, and Static PRC-Plus, where Static PRC-Plus augments Appendix A.3 with every formative-survey-requested control that can be operationalized as a fixed UI element (including controls like 'Talk to me like a data scientist' and any other requested options flagged as too specific), and also adds controls selected by two independent HCI experts as broadly applicable to the six tasks. Use Wilcoxon signed-rank tests (paired data) with a pre-registered primary outcome of the Section 6.1.1 preference item. If Dynamic PRC no longer significantly beats Static PRC-Plus, the original dynamic advantage is due to baseline coverage; if it still does, the dynamic-generation mechanism is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that Dynamic PRC is preferred because it affords more control and lowers context barriers—rests on the contrast with Static PRC. That contrast is confounded. Section 4.3 states the Static PRC option list was derived from the formative survey but explicitly omits options that 'required Dynamic PRC generation or interpretation' and keeps only 'generally applicable' controls. Appendix A.3 lists six generic controls (expertise, length, role, type, start, tone). Dynamic PRC, by contrast, generates 3-5 task-specific controls per prompt. Thus the experiment varies two things at once: (1) static versus dynamically generated controls, and (2) generic versus task-adapted option content. The paper provides no evidence that the six generic options are applicable or useful for the six specific tasks. Section 6.2.2 even reports Task 4 (Python code explanation) where Static PRC significantly outperformed Dynamic PRC (U=12.0, p=0.0355), suggesting adequacy is task-dependent. If the static list is an unusually weak or mismatched baseline, the observed 'preference for Dynamic PRC' and the significant effectiveness result (U=73.5, p=0.0245) would be an artifact of baseline weakness, not a validation of dynamic generation. The paper's own limitation section acknowledges no baseline chat condition, but the more immediate threat is within the two-condition comparison: the Static PRC arm is not a controlled instantiation of 'static middleware'; it is a particular, possibly unrepresentative option set. Without validating or augmenting that set, the abstract's causal claim is underdetermined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the challenge of prompting generative AI for comprehension tasks (explaining spreadsheet formulas, code, text) by proposing 'Dynamic Prompt Refinement Control' (Dynamic PRC), which uses an LLM-based Option Module to generate task-specific UI controls for refining prompts, compared with 'Static Prompt Refinement Control' (Static PRC), which offers a fixed preset list of controls. The paper reports a formative survey (n=38) that motivates two design goals (direct control and adaptive/dynamic control), a within-subjects user study (n=16) across six comprehension tasks, and qualitative interviews. The main results are a stated participant preference for Dynamic PRC, higher perceived effectiveness of Dynamic PRC for controlling AI output, and several qualitative themes: Dynamic PRC lowers barriers to providing context, improves perceived control, offers guidance, and encourages exploration/reflection, while making reasoning about the effects of controls more difficult. The paper concludes with design implications for future dynamic prompt middleware systems.","tokens_in":33530,"tokens_out":4547,"duration_ms":49560,"significance":"If the quantitative claims are robust, the paper makes a useful contribution to HCI research on generative AI interaction by articulating a new design space (dynamic prompt middleware), providing a concrete implementation with full system prompts in appendices, and offering qualitative findings that are well supported by participant quotes and coherent thematic analysis. The design implications in Section 7 (control over how options are applied, leveraging user context, direct manipulation) are actionable and plausible. The paper explicitly frames its design goals as aligning with prior work on controllable and adaptable explanations, which is a strength. However, the central quantitative claim of superiority for Dynamic PRC is currently undermined by an inappropriate statistical test for paired data and by a confounded baseline condition, so the paper's main evidence rests on qualitative insights and descriptive statistics.","major_comments":[{"comment":"The manuscript reports Mann-Whitney U tests for the within-subjects comparison of Dynamic PRC and Static PRC. This is statistically inappropriate because each participant contributed data to both conditions, so the observations are paired, not independent. The correct test is the Wilcoxon signed-rank test (or a paired t-test if assumptions hold). The reported significant results for effectiveness (U=73.5, p=0.0245) and for needing more control (U=190.5, p=0.0174) may not remain significant under a paired analysis. Notably, the median effectiveness rating is identical (6.0 vs 6.0) in the two conditions, and the significance rests on distributional differences; a paired test is needed to establish whether the within-subject direction is consistent. Because these p-values are the only quantitative support for the claim that Dynamic PRC is 'significantly more effective,' this is a load-bearing issue.","section":"§5.3, §6.2.1"},{"comment":"The Static PRC baseline is confounded with the content of the controls. Section 4.3 states that the Static PRC preset list was derived from the formative survey but excludes options that 'required Dynamic PRC generation or interpretation' and keeps only 'generally applicable' controls; Appendix A.3 confirms that the list contains six generic controls (expertise, length, role, type, start, tone). Dynamic PRC, in contrast, generates 3-5 task-specific controls per prompt. Thus the experiment varies both the generation mechanism (static vs dynamic) and the content (generic vs task-adapted). The observed preference for Dynamic PRC may reflect the poor fit of the generic preset list rather than the value of dynamic generation. This is highlighted by the Task 4 result in Section 6.2.2, where Static PRC significantly outperformed Dynamic PRC (U=12.0, p=0.0355), suggesting that the static list can be adequate for some tasks. Without a Static PRC-Plus condition with a richer, task-specific preset list (or at least evidence that the six generic controls are applicable to the six study tasks), the central claim that Dynamic PRC is preferred because of dynamic generation is not fully supported.","section":"§4.3, §6.2.2"},{"comment":"The one significant task-level result (Task 4, Dynamic PRC worse than Static PRC, U=12.0, p=0.0355) is reported without correction for multiple comparisons across six tasks and multiple questionnaire items. With this many tests, one significant result at p≈0.036 is expected by chance. The interpretation that 'the options provided in Static PRC are an effective set of refinements for code comprehension tasks' goes beyond what the data can support. The authors should either apply a multiple-comparison correction (e.g., Bonferroni or FDR) or explicitly label these as exploratory and refrain from task-specific claims.","section":"§6.2.2"},{"comment":"The key outcome of the study — participant preference for Dynamic PRC over Static PRC (median=2.0, mean=2.81 on a 1-7 scale where 1 is Dynamic) — is reported descriptively without any inferential test. Given that the abstract and conclusion state that 'Results show a preference for the Dynamic PRC approach,' this central claim should be supported by a statistical test, e.g., a Wilcoxon signed-rank test against the neutral value (4) or a sign test on the paired preference direction. As written, the preference claim rests on descriptive statistics and is not backed by significance testing.","section":"§6.1.1"}],"minor_comments":[{"comment":"The Python code snippet contains stray closing braces (e.g., 'import pandas as pd }' and similar after each line) that appear to be formatting artifacts; these should be corrected for clarity.","section":"§5.2, Task B1"},{"comment":"The text states that Dynamic PRC was 'significantly more effective' than Static PRC, but the medians are identical (6.0 vs 6.0) and the difference appears only in the means and the (inappropriate) U test. Please clarify that the result is distributional and not a median shift.","section":"§6.2.1"},{"comment":"The protocol description says 'We performed and report Mann-Whitney U Tests'; for a within-subjects design this should be a paired test. If the authors intend to retain Mann-Whitney U, they need to justify the independence assumption, which is not met here.","section":"§5.3"},{"comment":"The terms 'mental load,' 'mental demand,' and 'cognitive effort' are used somewhat interchangeably; for consistency with the NASA-TLX instrument, the authors should align their terminology.","section":"§6.3"},{"comment":"The sentence in §6.1.1 that 'the affordances that Dynamic PRC brought ... led to participants responding that they preferred' uses causal language that may overstate what self-report data can establish; suggest tempering to 'participants reported preferring.'","section":"§6.4.2"},{"comment":"The thematic analysis mentions 'negotiated-agreement [37, 51]' but does not report inter-rater reliability statistics (e.g., Cohen's kappa); including a brief reliability measure would strengthen the qualitative claims.","section":"§6.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is situated in an active line of research and the authors cite several of their own prior works; this is not in itself problematic. The main concerns are methodological: the inappropriate Mann-Whitney U test for paired data and the confounded Static PRC baseline. These are fixable through re-analysis and more careful framing, though the baseline confound may require additional data (e.g., a Static PRC-Plus condition) to fully resolve. The qualitative findings are valuable and well-presented, and the implementation detail is a strength. I recommend major revision rather than rejection because the central design contribution is sound and the evidence, if properly analyzed, may still support a more modest version of the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline is: this paper is worth reading for its system design and qualitative findings, but the central quantitative claim—that dynamic generation of controls is better than a fixed list—is underdetermined because the static baseline is a specific, possibly unrepresentative option set, not a fair instantiation of 'static middleware.'\n\nWhat's new: applying dynamic prompt middleware (generated UI controls that refine prompts) to comprehension and learning tasks, with a two-tier inline/session design. That specific combination isn't in the cited prior work. BISCUIT does it for notebooks, DynaVis for visualization edits, and MacNeil's prompt middleware is a static mapping. The formative survey (n=38) gives useful data on user desire for control, and the qualitative analysis (n=16) is the strongest part. The themes—lowered context barriers, guidance, exploration, but also difficulty reasoning about option effects—are coherent and supported by direct quotes. The paper also openly acknowledges the missing baseline chat condition and the fact that system prompt effects weren't evaluated.\n\nThe soft spots are real. Section 4.3 explicitly says the Static PRC option set was limited to options that were 'generally applicable'; anything that required dynamic generation or interpretation was excluded. So the comparison varies both the generation mechanism and the coverage/relevance of the options. The observed preference for Dynamic PRC could just mean that task-adapted options beat a fixed set of six generic ones, not that dynamic generation per se helps. The significant reversal on Task 4 (Python code) suggests that the static list is not uniformly weak, which is informative but also underscores task-dependence. The stats also need work: Mann-Whitney U treats paired within-subjects data as unpaired; Wilcoxon signed-rank is the obvious fix. Multiple comparisons uncorrected. That said, these issues don't sink the qualitative story; they just mean the abstract's causal claim is stronger than the evidence supports.\n\nMy take: if the authors reanalyze with appropriate tests, add a Static PRC-Plus condition that matches the option coverage, or at least validate the static list's adequacy on the tasks, the paper would be much more convincing. As it stands, it's a solid design and a useful qualitative contribution, but the central claim needs qualification.\n\nI'd send it to peer review—the idea is timely and the system is described in enough detail to build on—but I'd expect major revision on the baseline and statistics.\n\nRecommendation: engage with it, but read Section 4.3 carefully before trusting the abstract.","headline":"Worth reading for the system design and qualitative findings, but the central causal claim about dynamic generation is underdetermined because the Static PRC baseline is a specific, possibly unrepresentative option set.","tokens_in":34098,"tokens_out":2689,"would_cite":true,"duration_ms":30025,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Users preferred dynamically generated prompt refinement controls over a fixed preset list, reporting more control, lower barriers to providing context, and more exploration of comprehension tasks, though the effects of controls on the…","keywords":["prompt middleware","dynamic UI generation","user control","generative AI","comprehension tasks","explanation interfaces","human-AI interaction","prompt refinement"],"falsifier":"Run the same within-subjects study with a Static PRC list that is carefully validated by domain experts as well-matched to the specific tasks (e.g., a preset list distilled from the Dynamic PRC system's own most common generations), and measure both preference and objective explanation quality; if the dynamic advantage disappears or reverses, the reported preference is an artifact of the particular static list rather than the generation mechanism.","tokens_in":33048,"feed_emoji":"🎛️","tokens_out":3143,"duration_ms":37549,"temperature":0.7,"pith_summary":"This paper argues that prompt middleware—interfaces that help users construct prompts—can be improved by generating the control elements themselves on the fly, tailored to the user's current prompt and task. The authors built a system whose UI options (radio buttons, checkboxes, text boxes) are produced by a language model based on the user's input, and compared it to a version offering a static, preset list of generally useful refinements. In a formative survey (n=38) and a controlled within-subjects study (n=16) with comprehension tasks such as explaining spreadsheet formulas, Python code, and text passages, participants preferred the dynamic version for affording more control and lowering barriers to providing context, while also finding it more mentally demanding and harder to reason about. The central claim is that dynamically generating prompt refinements can improve the user experience of generative AI workflows, and the paper derives design implications for future such systems.","feed_headline":"AI-crafted prompt controls beat fixed menus in user study","feed_subtitle":"16 users preferred AI-generated controls for more control, lower context barriers, and more exploration.","key_machinery":"The central object is the 'Option Module'—an LLM agent that takes the user's prompt, conversation history, and current session options as input, and returns a set of prompt options as JSON conforming to a TypeScript schema. These options are rendered as radio buttons, checkboxes, and free-text fields with initial values; a serializer converts the selected options into textual prompt refinements that are appended to the prompt sent to a separate Chat Module, allowing deterministic additions, removals, or modifications of the user's original prompt. A two-tier structure separates inline options (regenerated for every user input) from session options (applied to all subsequent prompts), giving users both task-specific tailoring and persistent preferences.","core_discovery":"The paper's central claim is that dynamic prompt middleware—where an LLM analyzes the user's prompt and generates a set of refinement options rendered as GUI controls—gives users more effective control over AI responses than a static, pre-defined list of refinements, for comprehension and learning tasks. The evidence is a preference for the Dynamic PRC approach in a within-subjects study: participants reported that it afforded more control, lowered the barrier to providing context, and encouraged exploration and reflection, at the cost of greater perceived complexity and difficulty in predicting how individual options would affect the output. The paper presents this as a trade-off between standardized but predictable support (Static PRC) and adaptive but less predictable support (Dynamic PRC), with users favoring the adaptive form.","pith_inferences":["A plausible extension is that dynamic generation could benefit not just comprehension tasks but any LLM-steering interface, including code generation, data analysis, and creative writing; the evaluated mechanism is general and modular.","The paper measures perceived control and preference, not objective explanation quality or learning outcomes; a testable extension would examine whether dynamic options actually improve comprehension test scores or task performance.","The static baseline's fairness is the main internal-validity concern: if the preset list happened to be poorly matched to the study tasks, the observed dynamic advantage would be an artifact; future work should validate static option sets against expert-derived alternatives.","The reported difficulty of predicting option effects suggests that dynamically generated controls could be paired with XAI-style transparency, such as showing exactly how each option changed the prompt or the response, which the paper's own design implications begin to outline."],"forward_implications":["Dynamic prompt middleware can improve the user experience of generative AI workflows by increasing perceived control and reducing the effort needed to express context.","Users can generate new controls through natural language or by directly editing JSON, enabling them to create refinements not anticipated by the system.","The approach encourages exploration and reflection on comprehension tasks, suggesting potential as a 'tool for thought' or metacognitive scaffold.","Design implications include giving users control over how and where the AI applies options, leveraging user context and data to generate helpful controls, and supporting direct manipulation of generated options.","The remaining barrier—reasoning about the effects of generated controls on final output—points toward integrating explainability mechanisms such as response diffs or more transparent option application."],"supporting_citations":[{"why":"Documents the burden of elaborating task context in prompting, motivating the lower-barrier goal of the Dynamic PRC approach.","marker":"[11]"},{"why":"Identifies the fuzzy abstraction matching problem, which the generated options are designed to help users overcome.","marker":"[34]"},{"why":"Defines the concept of prompt middleware as GUI affordances for manipulating prompts, the category the paper extends.","marker":"[36]"},{"why":"Provides design goals that explanations be controllable and adaptable, which the paper adopts as its own D.1 and D.2.","marker":"[65]"},{"why":"Shows a precedent for dynamically generated ephemeral UIs in code generation, which the paper adapts to comprehension tasks.","marker":"[8]"},{"why":"Demonstrates dynamically synthesized UI widgets for visualization editing, a related instance of generated control surfaces.","marker":"[62]"},{"why":"Describes the metacognitive demands of prompting, which the paper's exploration-and-reflection finding builds on.","marker":"[60]"}],"fun_headline_variants":["Adaptive prompt controls win over preset lists in study","AI-generated prompt refinements preferred for better control","Dynamic controls give users more control over AI explanations","Users favor adaptive prompt controls despite added complexity","Study: dynamic prompt controls improve user control and exploration"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the Static PRC preset list is a fair, non-degenerate baseline; if the fixed options were unusually ill-suited to the study tasks, the observed Dynamic PRC advantage would reflect baseline weakness rather than the value of dynamic generation.","fun_headline_variants_meta":{"raw":{"variants":["Adaptive prompt controls win over preset lists in study","AI-generated prompt refinements preferred for better control","Dynamic controls give users more control over AI explanations","Users favor adaptive prompt controls despite added complexity","Study: dynamic prompt controls improve user control and exploration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000597,"raw_usage":{"total_tokens":2813,"prompt_tokens":988,"completion_tokens":1825,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":1753}},"tokens_in":604,"tokens_out":1825,"duration_ms":15352,"temperature":1.0,"reasoning_tokens":1753,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:32:56.825908+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same within-subjects study with a Static PRC list that is carefully validated by domain experts as well-matched to the specific tasks (e.g., a preset list distilled from the Dynamic PRC system's own most common generations), and measure both preference and objective explanation quality; if the dynamic advantage disappears or reverses, the reported preference is an artifact of the particular static list rather than the generation mechanism.","supporting_citations":[],"review_version":1}