{"id":"a5b8d1c0-b54c-4d6c-bb8e-f6315101665f","arxiv_id":"2506.11767","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A clickable interface with predefined commands improves teachers' perceived usability and reduces workload compared with a ChatGPT-style chat interface for course outline creation.","lead":"This paper tests two new interface designs that help teachers build course outlines with ChatGPT. In a 20-person study, a button-based interface with predefined commands felt easier to use and lighter on workload than plain ChatGPT.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The control condition is not shown to match the treatment arms on model, system prompt, or context injection, so the reported SUS/RTLX advantage may reflect prompt scaffolding rather than interface design.","rationale":"The reader's weakest_assumption correctly identifies the most load-bearing concern: the control condition is not demonstrably matched to the treatment arms on model, system prompt, or context injection. This is not a minor implementation detail; it goes directly to whether the study isolates interface design from LLM prompt configuration. The paper itself motivates the control by saying direct ChatGPT could not guarantee the same model, so the authors clearly recognized the need for model equivalence, but they do not report whether the same system prompt and context-injection behavior used in Section 3.3 were applied to the open-webui control. If they were not, the reported differences could be explained by the control generating less structured, less context-aware outputs rather than by the UI per se. I agree with the reader that this warrants a conditional verdict rather than a rejection: the raw data are shared, the within-subject design is reasonable, and the direction of the effect is plausible, but the central causal claim about interface design is not yet supported without matching the prompt pipeline. The concrete test I propose would settle the concern: either inspect the control configuration or run an ablation where the control receives identical model, system prompt, and context injection, with only the input modality differing. If the effect persists under those conditions, the interface attribution is much stronger. If not, the paper's headline claim must be weakened. This does not move the verdict away from CONDITIONAL, so the recommendation is UNCHANGED.","tokens_in":7584,"tokens_out":3791,"duration_ms":39011,"concrete_test":"Ask the authors for the exact open-webui configuration and the system prompt used for the control condition, or run an ablation: configure open-webui with the same gpt-4o-2024-08-06 model, the same Section 3.3 system prompt, and the same automatic course-outline context injection, but present commands as typed text instead of buttons. If UI Predefined no longer shows a significant SUS/RTLX advantage over this prompt-matched control, the original effect is due to prompt scaffolding rather than interface design.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the abstract and Section 5.1, is that UI Predefined significantly outperforms the standard ChatGPT interface in usability and workload. For this comparison to be valid, the only difference between UI Predefined and the control should be the interface. Section 3.4 explains that the control is an open-webui-based 'ChatGPT replica' because direct ChatGPT could not guarantee the same model or allow interaction monitoring, but the paper never documents the control's model version, system prompt, or context injection. Section 3.3 describes a tailored system prompt for the proposed UIs that defines the persona, delimits components, specifies the course-outline curation task, and fixes the output format, and it also says the current course outline and user commands were always included in prompts. If the open-webui control did not receive the same gpt-4o-2024-08-06 model, the same system prompt, and the same automatic inclusion of the course outline, then the 17.75-point SUS difference and 1.05-point RTLX difference attributed to UI design could instead be caused by the treatment arm having a structured output contract and full context while the control chat only gets whatever the participant types. That would undermine the causal attribution in Section 5.1, because the advantage would be prompt engineering, not the direct-manipulation interface itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two user interfaces, UI Predefined and UI Open, grounded in direct manipulation principles, to reduce the prompt-engineering burden faced by educators when using LLMs for curriculum development. A within-subjects user study with 20 participants compares these UIs against an open-webui-based ChatGPT replica using the System Usability Scale (SUS) and the NASA Raw Task Load Index (RTLX). The results show that UI Predefined achieves the highest SUS score (86.75) and lowest RTLX score (2.25), followed by UI Open (70.75, 3.00) and the ChatGPT replica (69.00, 3.30). Statistical tests using the Wilcoxon signed-rank test are reported for the two winning comparisons, and the authors conclude that UI Predefined significantly outperforms both ChatGPT and UI Open in usability and workload.","tokens_in":7825,"tokens_out":3584,"duration_ms":33339,"significance":"If the control condition is shown to match the treatment arms on model, system prompt, and context provisioning, the study would provide a useful empirical demonstration that GUI-style direct manipulation can improve perceived usability and reduce workload for LLM interaction in a practical educational task. The paper uses standard, externally validated instruments (SUS, NASA RTLX), randomizes the order of interface use, and makes raw data available, which are strengths that support reproducibility. The contribution is moderate but relevant to HCI research on LLM interfaces and to the learning-technology community, as it addresses an actual pain point (complex prompt engineering) for a specific user group (educators).","major_comments":[{"comment":"The causal attribution in the abstract and Section 5.1 is undermined by an undocumented confound in the control condition. Section 3.3 describes a tailored system prompt for the proposed UIs that defines the persona, delimits prompt components, specifies the course-outline curation task, fixes an output format, and always includes the current course outline and user commands. Section 3.4 describes the control as an open-webui-based ChatGPT replica, but the manuscript never states whether the control used the same model (gpt-4o-2024-08-06), the same system prompt, or the same automatic inclusion of the course outline and user commands. If the control lacked these prompt and context provisions, the reported SUS advantage of 17.75 points and RTLX advantage of 1.05 points (Tables 1 and 2) could arise from prompt engineering or context injection rather than from the direct-manipulation interface itself. The authors should document the control configuration in detail, or run an additional control condition that holds model, system prompt, and context constant while varying only the interface, and then re-analyze the data.","section":"3.3-3.4, 5.1"},{"comment":"The statistical evidence for the central workload claim is weaker than reported. Tables 1 and 2 present p-values only as inequalities and do not report confidence intervals, effect sizes, or any correction for multiple comparisons. For the NASA RTLX comparison between UI Predefined and ChatGPT, the reported p<0.032 would not survive a Bonferroni correction for the three pairwise comparisons (corrected threshold 0.017), and the comparison between UI Predefined and UI Open (p<0.022) would also not survive. The abstract and Section 5.1 claim significantly reduced task load for UI Predefined, but as reported this claim is not robust. The authors should report exact p-values, effect sizes (e.g., matched rank-biserial correlation), and either a multiple-comparison correction or a pre-specified analysis plan.","section":"Tables 1-2, 5.1"}],"minor_comments":[{"comment":"The comparison between UI Open and ChatGPT has no reported p-value, yet Section 5.1 describes it as a non-significant improvement; please either provide the test statistic or explicitly state that the test was not performed.","section":"Table 2"},{"comment":"The caption says 'The dotted lines are mean,' but the figure legend could clarify which dotted line corresponds to which interface to avoid ambiguity.","section":"Figure 6"},{"comment":"The reference lists 'EC-TEL 2020' but the publication year appears to be 2024; please correct the conference or year.","section":"Reference [15]"},{"comment":"There is a typo in 'UI Predefined' written as 'UIPredefined' in the sentence 'Our findings revealed that the UIPredefined significantly outperformed ChatGPT.'","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The control-condition confound is the central issue and is likely addressable by providing supplementary documentation of the open-webui setup. If the authors can confirm that the control received identical model, system prompt, and context injection, the paper may become publishable after the statistical reporting is improved. The topic fits the journal's scope and the paper has several strengths, including the use of standard instruments and the availability of raw data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper does something useful—it builds two Direct Manipulation interfaces for LLM-assisted curriculum development and compares them head-to-head. The internal comparison between UI Predefined and UI Open is the cleanest result and it holds. The headline comparison against the ChatGPT-style control is weaker than the authors claim, because the control condition is not shown to match the treatment arms on model, system prompt, or context injection. That's a documentation gap, but it means the 17.75-point SUS difference could plausibly come from prompt scaffolding rather than from the interface itself.\n\nWhat's new: DirectGPT already applied DM principles to LLM UIs, so this is not a paradigm shift. The addition here is a domain-specific instantiation for curriculum development, a careful mapping onto Shneiderman's four pillars, and a study that contrasts a button-driven interface with a more open command interface. That contrast is informative: UI Predefined beat UI Open on SUS and RTLX with p<0.02, and the effect sizes are plausible. The authors also share raw data and a demo, which makes the work easier to build on than many HCI papers in this area.\n\nSoft spots, in order of severity. First, the control condition is under-specified. Section 3.4 says they used open-webui to replicate ChatGPT, but never documents the model version, system prompt, or whether the course outline was automatically injected into the context. Section 3.3 shows the treatment arms had a tailored system prompt and automatic context inclusion. If the control lacked those, the comparison is not interface-vs-interface; it is prompt-engineered-vs-raw-chat. The paper should document the control configuration or run a matched control condition. Second, the sample is 20, which is small but reasonable for a within-subjects design. The paper reports p-values but no multiple-comparison correction, and Figure 5 lacks error bars. Minor, but worth noting. Third, the abstract states that UI Predefined significantly outperformed ChatGPT without qualifying the control confound; a more cautious phrasing would attribute the result to the UI plus its prompt setup.\n\nThe central internal comparison holds up: the two UIs share the same back-end and prompt structure, so the difference between them is genuinely about interface design. This paper is worth a serious referee; the main revision request would be to make the control condition transparent and, ideally, to add a matched prompt condition. I would cite this in future work on LLM interface design, and a reading group could use it as a case study in control-condition design. My recommendation: send it to peer review with a request for a thorough revision on the control documentation.","headline":"Useful internal comparison between two DM-based LLM interfaces; headline control comparison is confounded by undocumented prompt configuration.","tokens_in":8291,"tokens_out":3701,"would_cite":true,"duration_ms":32914,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Clickable, predefined commands outperform free-form chat for LLM-assisted curriculum design.","keywords":["User-centered design","LLM user interface","Curriculum development","Direct manipulation","System Usability Scale","NASA task load","Human-AI collaboration","Course outline generation"],"falsifier":"A replication that captures the raw prompts, system prompt, and model configuration for all three arms would settle the claim: if a standard chat interface given the same system prompt and output format closes the SUS and NASA RTLX gaps, the reported advantage is due to prompt engineering rather than interface design. Alternatively, an in-the-wild study with educators revising real courses under deadline pressure could falsify the lab result if the usability gains disappear.","tokens_in":7441,"feed_emoji":"🎓","tokens_out":6509,"duration_ms":57382,"temperature":0.7,"pith_summary":"This paper tries to establish that the main barrier to educators using large language models is not model capability but the chat interface, and that interfaces built on Direct Manipulation principles remove that barrier. It reports that a UI with curated, clickable commands (UI Predefined) significantly beat both a ChatGPT-like control and a more flexible open-command UI on usability and workload in a 20-participant study. Concretely, UI Predefined scored 86.75 on the System Usability Scale versus 69.00 for the control and 70.75 for UI Open, and lowered NASA RTLX task load to 2.25 from 3.30 and 3.00. The paper argues this is evidence that reducing reliance on complex prompt engineering, rather than teaching prompt skills, is the productive path for educator-AI collaboration in curriculum development.","feed_headline":"Clickable-command UI beats ChatGPT for course design","feed_subtitle":"In a 20-teacher study, predefined commands cut task load and lifted usability over open chat.","key_machinery":"The carrying mechanism is the application of Direct Manipulation to LLM interaction: (DM1) continuous representation of the object of interest, realized as an interactive course-outline table that stays visible; (DM2) physical actions such as clicking, checking, dragging, and dropping instead of typing syntax; (DM3) rapid, incremental, reversible operations with undo/redo and loading effects confined to changed sections; and (DM4) recognition of available commands through menus and buttons rather than recall of prompt syntax. UI Predefined operationalizes this with curated command groups, while UI Open adds a dynamic command chat box with drag-to-localize and save/reuse. Behind the interface, a system prompt defines persona, task, delimiters, and output format, and the UI engine translates user actions into prompts that always include the current course outline.","core_discovery":"The central claim is that interface design is the decisive factor in making LLM-assisted curriculum development usable by educators. The paper claims UI Predefined, an interface that replaces free-form prompts with four groups of expert-curated clickable commands acting on an always-visible interactive course outline, outperformed the ChatGPT-like control on both measured dimensions: usability (SUS 86.75 vs 69.00, p < 0.009) and workload (NASA RTLX 2.25 vs 3.30, p < 0.03). It also claims UI Open, which keeps the same outline view but adds a chat box for dynamic, reusable, and drag-and-drop commands, ranked between the two, with greater flexibility but a steeper learning curve and no statistically significant advantage over ChatGPT. The authors read the results as evidence that applying Direct Manipulation to LLM interaction, with a human check in the loop, is an effective design strategy for educator-LLM collaboration.","pith_inferences":["A testable extension is to run the same three conditions with the ChatGPT-like control matched to the proposed UIs on system prompt, model version, and output formatting; this would separate interface effects from prompt-engineering effects, which the current report does not fully distinguish.","The same design pattern likely transfers to other structured authoring tasks, such as assessments, syllabi, and reports, where the object of interest is a visible document with a stable schema, though the paper only tests curriculum outlines.","The 20-participant sample with 1-21 years of teaching experience suggests the usability gap may be smaller for educators already fluent in prompt engineering; screening by prompt skill could reveal an interaction.","A field deployment with real course revision deadlines, rather than a lab task, would test whether the workload reduction survives time pressure and context switching."],"forward_implications":["Educators can produce course outlines with less typing and lower cognitive load when LLM commands are predefined and applied directly to the document, rather than expressed as chat prompts.","The non-significant gap between UI Open and the control suggests that flexibility alone does not buy usability; structure and guidance carry the benefit.","A hybrid interface that combines predefined commands with the ability to save custom ones, as the paper proposes, would target the flexibility cost seen in UI Open.","Institutional adoption of LLM tools for curriculum work may hinge on providing expert-curated command sets and visible output formatting, not on prompt-training educators."],"supporting_citations":[{"why":"Defines the four Direct Manipulation traits that serve as the design foundation for both proposed UIs.","marker":"[20]"},{"why":"Presents a prior Direct Manipulation LLM interface whose usability results and evaluation approach the paper builds on.","marker":"[14]"},{"why":"Supplies the System Usability Scale interpretation thresholds used to classify UI Predefined as Excellent and the other conditions as OK.","marker":"[3]"},{"why":"Surveys cognitive workload measurement and grounds the choice of the NASA RTLX questionnaire.","marker":"[11]"},{"why":"Informs the system-prompt structure and output-format logic used by both proposed UIs in the experiment.","marker":"[17]"},{"why":"Provides the Wilcoxon signed-rank test used to compute the reported significance of SUS and task-load differences.","marker":"[24]"}],"fun_headline_variants":["Clickable commands beat open chat for lesson planning","Predefined UI cuts workload for educators using LLMs","UI design trumps prompts: study shows clickable wins","Study: Clickable-command UI tops ChatGPT for course design"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes the ChatGPT-like control differed from the two new interfaces only in interface design, with the same underlying system prompt, model settings, and output formatting, but the paper does not document that those were matched.","fun_headline_variants_meta":{"raw":{"variants":["Clickable commands beat open chat for lesson planning","Predefined UI cuts workload for educators using LLMs","UI design trumps prompts: study shows clickable wins","Study: Clickable-command UI tops ChatGPT for course design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1257,"prompt_tokens":961,"completion_tokens":296,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":231}},"tokens_in":577,"tokens_out":296,"duration_ms":3618,"temperature":1.0,"reasoning_tokens":231,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:03:12.610528+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that captures the raw prompts, system prompt, and model configuration for all three arms would settle the claim: if a standard chat interface given the same system prompt and output format closes the SUS and NASA RTLX gaps, the reported advantage is due to prompt engineering rather than interface design. Alternatively, an in-the-wild study with educators revising real courses under deadline pressure could falsify the lab result if the usability gains disappear.","supporting_citations":[{"cited_title":"Computer16(08), 57–69 (1983)","cited_arxiv_id":null,"evidence_quote":"Defines the four Direct Manipulation traits that serve as the design foundation for both proposed UIs."},{"cited_title":"In: Proceedings of the CHI Conference on Human Factors in Computing Systems","cited_arxiv_id":null,"evidence_quote":"Presents a prior Direct Manipulation LLM interface whose usability results and evaluation approach the paper builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the System Usability Scale interpretation thresholds used to classify UI Predefined as Excellent and the other conditions as OK."},{"cited_title":"ACM Computing Surveys55(13s), 1–39 (2023)","cited_arxiv_id":null,"evidence_quote":"Surveys cognitive workload measurement and grounds the choice of the NASA RTLX questionnaire."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Informs the system-prompt structure and output-format logic used by both proposed UIs in the experiment."},{"cited_title":"Encyclopedia of Biostatistics8(2005)","cited_arxiv_id":null,"evidence_quote":"Provides the Wilcoxon signed-rank test used to compute the reported significance of SUS and task-load differences."}],"review_version":1}