{"id":"d91506a1-6ad3-4ca2-8027-d859ca2c9c85","arxiv_id":"2502.04983","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An interactive no-code system combining element-level modular LLM sessions, drawing-based graphical proxies, and automatic sliders creates 2D scenes faster than Cursor Composer in a 10-participant study.","lead":"MoGraphGPT is a graphical tool that lets people create 2D interactive scenes, like simple games, by typing instructions and drawing on a canvas, with separate AI modules for each object and a central module that wires interactions together. It reports that users finish tasks faster and with less effort than with Cursor Composer, an AI coding assistant.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The controlled evaluation does not test the claim's scope: its three tasks are one- or two-element animations with no user interaction, so the abstract's 'complex interactive scenes with multiple visual elements' is not yet supported.","rationale":"The reader's weakest assumption concerns fairness of the baseline (hidden code, non-counterbalanced order), and that is a legitimate methodological risk. My concern is different but related: even if the comparison were perfectly fair, the controlled study would still not test the full central claim because its tasks lack user interaction and contain at most two elements. This is a scope mismatch between the abstract's 'complex 2D interactive scenes with multiple visual elements' and the empirical evidence. The open-ended study provides encouraging feasibility evidence for larger, interactive scenes, but it does not compare against a baseline, so it cannot support the word 'significantly.' I therefore keep the same conditional verdict: the paper is a solid systems contribution, but the headline claim should either be narrowed to the evaluated task class or supported by a controlled comparison that includes a multi-element interactive task. I do not think this concern rises to rejection, because the reported effects are large and the open-ended results are consistent with the system's intended value; however, the current evidence does not justify accepting the broadest formulation of the central claim.","tokens_in":21579,"tokens_out":7672,"duration_ms":92736,"concrete_test":"Add a fourth controlled task (or replace Task 3) with at least three elements and at least one genuine user interaction, e.g., an arrow-key-controlled player that collects items, triggers score changes, and handles collision with an obstacle, while other elements move independently. Run the same within-subject protocol with counterbalanced order, and compare completion time, prompt count, prompt length, and the same questionnaire items. If MoGraphGPT's advantage remains significant on this task, the scope concern is resolved; if not, the abstract should be narrowed to non-interactive, few-element scenes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims MoGraphGPT 'significantly improves easiness, controllability, and refinement in creating complex 2D interactive scenes with multiple visual elements.' The significance tests in Section 6.6, however, are run only on the three fixed tasks in Section 6.3: Task 1 is a single fish moving point-to-point, Task 2 is a single fish moving along a curved path, and Task 3 is a sun and an earth with self-rotation plus orbit. None of these tasks involve user input (keyboard, mouse, touch, or other interaction), none involve more than two elements, and none involve collision, scoring, conditional logic, or the kind of multi-element coordination that the modular architecture is designed to support. Thus the Wilcoxon results support a narrower claim: MoGraphGPT reduces time and prompt effort for simple one- or two-element animations under the study protocol. The open-ended study (Section 7) does include scenes with 4-8 elements and real user interactions, but it is not comparative and reports only SUS scores and qualitative feedback, so it cannot establish 'significantly improves' for that task class. If the controlled tasks had included a genuinely multi-element interactive scenario, the observed advantage might shrink or disappear, because the benefit of element-level modularization and graphical control is most relevant exactly when several elements interact and when users must specify behaviors. As reported, the central claim in the abstract generalizes beyond the evidence actually collected.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents MoGraphGPT, a system for creating 2D interactive scenes using modular large language models (LLMs) combined with graphical control. The system maintains independent LLM sessions for each scene element and a central session for interactions, and provides graphical proxies (points, lines, curves, regions), direct manipulation of element positions, and automatically generated sliders for parameter refinement. The authors report a formative content analysis of online tutorials, a within-subjects comparative study with 10 participants against Cursor Composer on three fixed tasks, and an open-ended usability study with 6 participants. The abstract claims that MoGraphGPT significantly improves easiness, controllability, and refinement in creating complex interactive scenes with multiple visual elements in a coding-free manner.","tokens_in":21798,"tokens_out":4788,"duration_ms":46877,"significance":"The system is a thoughtful and potentially valuable contribution to end-user development of interactive visuals. The element-level modularization with context sharing is a plausible response to the identified limitations of conversational LLMs, and the graphical control features (drawing proxies, direct manipulation, sliders) directly address usability gaps. The comparative study shows large, statistically significant differences on objective measures and all subjective items for the three tasks used, and the open-ended study demonstrates the system's expressiveness in creating diverse games and animations. The strengths are the clear implementation, the concrete modular architecture, and the attempt at a controlled comparison with a state-of-the-art baseline. However, the generalizability of the empirical claims is limited by the task scope, sample size, and design choices, which are detailed in the major comments.","major_comments":[{"comment":"The abstract claims that MoGraphGPT significantly improves easiness, controllability, and refinement in creating complex 2D interactive scenes with multiple visual elements. The controlled evaluation, however, uses three tasks that are one- or two-element animations with no user interaction: Task 1 is a single fish moving between two points, Task 2 is a single fish moving along a drawn curve, and Task 3 is a sun and earth with self-rotation and orbit. None of these tasks involve keyboard, mouse, touch, collision, scoring, or more than two elements. Thus the Wilcoxon results reported in Section 6.6 support a narrower claim about simple animations, not the abstract's scope. The open-ended study (Section 7) contains scenes with 4–8 elements and real interactions, but it is not comparative and reports only SUS and qualitative feedback, so it cannot support the claim of significant improvement for that task class.","section":"Section 6.3 and 6.6"},{"comment":"The procedure always presents MoGraphGPT before Cursor Composer for each task. The authors state this was to avoid learning effects of their system, but it introduces a systematic order confound: participants may carry over task familiarity, prompt strategies, or fatigue from the first condition to the second. This makes it impossible to attribute the observed differences solely to the tool. A balanced within-subject design or a between-subjects component is needed to support the causal claim of improved performance with MoGraphGPT.","section":"Section 6.4"},{"comment":"The baseline setup for Cursor Composer hides the code view and instructs participants to focus only on text input, context selection, and result rendering. Since Cursor is fundamentally a code editor, removing its code view may disadvantage it relative to MoGraphGPT, which is designed to hide code. This asymmetry on the primary workflow could inflate the subjective and objective differences. The authors should either justify this setup or discuss it explicitly as a threat to fairness in the comparison.","section":"Section 6.1"}],"minor_comments":[{"comment":"The comparative study uses a convenience sample of 10 participants, and the paper does not report effect sizes or confidence intervals alongside the p-values. Adding Cohen's d or equivalent would help readers gauge the magnitude of the observed differences.","section":"Section 6.2"},{"comment":"The paper conducts multiple Wilcoxon signed-rank tests on subjective questionnaire items without any correction for multiple comparisons. Reporting adjusted p-values or clearly labeling the results as exploratory would be more appropriate.","section":"Section 6.5"},{"comment":"The open-ended study reports an overall SUS score of 85 from 6 participants, but no variance or per-item statistics are shown, and 3 of the 6 participants had already taken part in the comparative study. This should be acknowledged as a limitation when interpreting the usability result.","section":"Section 7"},{"comment":"Table 1 contains typographical errors such as 'Interacitve' in the tool names and missing definite/indefinite articles in the descriptions; these should be corrected in a final polish.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The reader's take aligns with my own assessment: the system is interesting and the evaluation is underpowered for the claim as stated. The paper would benefit from either narrowing the abstract's claim to the tested task scope or adding a controlled evaluation that includes interactive multi-element scenes. There is no derivational circularity; the concerns are about the breadth of the empirical evidence. I do not see evidence that the authors are hiding limitations—Section 8.3 explicitly acknowledges context-loss problems—but the gap between the abstract and the controlled tasks is substantial and should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a genuine integrative contribution to LLM-based creative tools, but the abstract's 'complex interactive scenes with multiple visual elements' is broader than what the comparative evaluation actually tested. You should treat the system as addressing independent element refinement plus graphical grounding, not as proven for complex interactive scenes.\n\nWhat's new: per-element LLM sessions with a central session that consumes distilled class context, direct on-canvas graphical proxies (point, line, curve, region) injected into prompts, and auto-generated sliders for parameter tuning. Each piece is known, but the combination in one tool for 2D scene creation is not in the cited prior work. The comparison against Cursor Composer is reasonable, and the quantitative results are large: about 70% less time, 69% fewer prompts, 89% shorter prompts, with Wilcoxon p<0.01. The open-ended study shows real users building games with 4-8 elements and actual keyboard/mouse interaction, which is the strongest evidence for the system's expressiveness.\n\nNow the soft spots, in order of severity. First, the controlled study's three tasks are a point-to-point fish, a fish on a curved path, and a sun-earth orbit. None involve user input, collision, scoring, or more than two elements. So the significant differences support a narrower claim: faster and cheaper creation of simple one- or two-element animations. The 'complex interactive scenes' language in the abstract leans on the open-ended study, which is not comparative. Second, the procedure always presented MoGraphGPT first and hid Cursor's code view, so the baseline may have been harder than usual; order effects and a small convenience sample (n=10) further temper the effect sizes. Third, no code or data are released. These are fixable concerns, not fatal ones; the direction of the effect is plausible and the qualitative feedback is consistent.\n\nThe paper is for HCI folks working on generative tools; it deserves a serious referee. I'd ask for a revision that either adds a controlled task with genuine multi-element interaction or rewrites the central claim to match the evidence, and that addresses ordering and artifact release. Conditional accept, leaning positive.","headline":"A solid integrative HCI systems paper whose modular-LLM design is plausible and useful, but the controlled study's simple tasks don't support the abstract's 'complex interactive scenes' claim.","tokens_in":22340,"tokens_out":2436,"would_cite":true,"duration_ms":25453,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MoGraphGPT claims that splitting LLM code generation by element, with a central module for interactions plus graphical control, makes creating 2D interactive scenes significantly easier, and the user study reports large reductions in…","keywords":["Code Generation","Modularization","Large Language Models","ChatGPT","Graphical Control","Interactive Scenes","2D Games","Natural Language Interface"],"falsifier":"Run the same three tasks with tool order counterbalanced and the baseline's generated code visible to participants, as in normal use of an AI code editor; if the completion time, prompt count, prompt length, and rating advantages of MoGraphGPT shrink or disappear, the reported improvement would be traced to ordering and hidden code rather than to modularization and graphical control.","tokens_in":21362,"feed_emoji":"🎮","tokens_out":5530,"duration_ms":55538,"temperature":0.7,"pith_summary":"MoGraphGPT is a system for creating 2D interactive scenes, such as small games and animated demos, without writing code. The paper's claim is that its design, one LLM session per scene element plus a central session that wires interactions, fixes three failures of ordinary ChatGPT-style coding: generated code for one element accidentally changes others, spatial information is hard to express in text, and fine-tuning effects demands endless prompt rewording. To back this, the system adds graphical control, letting users draw points, lines, curves, and regions, reference them by name in prompts, and adjust parameters with automatically generated sliders. In a within-subject study against Cursor Composer, the paper reports significantly faster completions with significantly fewer and shorter prompts, and higher subjective ratings for easiness, controllability, and refinement.","feed_headline":"Modular chatbots build 2D games with fewer, shorter prompts","feed_subtitle":"A GUI that gives each scene element its own LLM session and adds sliders cuts task time about 70% versus Cursor.","key_machinery":"The load-bearing object is a two-level module hierarchy: individual LLM modules per element plus one central LLM module, linked by a context information repository that stores each element class's variables and functions. The individual modules generate and update class code independently, so refining one element does not rewrite another element's behavior. The central module reads the repository to generate interaction code, but it is instructed to keep element-specific variable and function definitions inside their own classes and to call them from the central code, preserving separation while retaining context.","core_discovery":"The paper introduces element-level context-aware modularization for LLM-based scene creation. Each scene element gets its own LLM module that generates and maintains a class for that element's properties and behaviors, while a central LLM module instantiates all elements, coordinates communication, and scripts interactions, calling functions that remain defined inside each element's class. A context information repository passes each class's variables and functions to the central module, so the central module knows what is available without entangling the code. On top of this, MoGraphGPT provides graphical controls: users draw point, line, curve, and region proxies that are labeled P1, L1, C1, and R1 and can be mentioned in prompts, and sliders extracted from generated variables enable precise parameter tuning. The paper's central claim is that this combination makes coding-free creation of complex 2D interactive scenes easier, more controllable, and easier to refine than a state-of-the-art AI coding tool.","pith_inferences":["The same element-level split could apply to any multi-component LLM code generation beyond 2D scenes, such as web applications or IoT automation, where independent modules with explicit interfaces would prevent cross-component breakage.","The hidden-code interface trades away a learning path; a version that reveals code after successful generations could turn the tool from a scene generator into a programming tutor, but the current evaluation does not test that.","Because the authors note context loss with long sessions, a scaling test that increases element count and measures how often the central module generates correct interactions would show where the modularization starts to fail.","Adding geometric primitives such as circles and rectangles, plus automatic background segmentation, which participants requested, would likely extend the same graphical proxy idea to richer spatial constraints."],"forward_implications":["Non-programmers can create complete 2D scenes with four to eight elements in ten to thirty minutes, as demonstrated by the open-ended study.","Modifying one element's behavior no longer rewrites another element's behavior, because each element's variables and functions live in its own class and the central module only calls them.","Spatial specifications such as exact positions, paths, and regions can be expressed by drawing and then referencing P1, L1, C1, or R1 in a prompt, avoiding laborious coordinate guesswork.","Effect parameters like speed, radius, and amplitude can be tuned with sliders instead of re-prompting the model with comparative words.","Task completion time, number of prompts, and prompt length all drop significantly compared with Cursor Composer on the three benchmark tasks."],"supporting_citations":[{"why":"Serves as the baseline system in the comparative evaluation that the central claim depends on.","marker":"[14]"},{"why":"Closest prior system for creating interactive worlds; the paper positions its LLM-based modular approach against DrawTalking's rule-based interactions.","marker":"[49]"},{"why":"DirectGPT motivates input-side graphical control; the paper extends this to multi-element dynamic scenes with parameter sliders.","marker":"[38]"},{"why":"ChatDev supplies the agent-level modularization baseline that the paper contrasts with its element-level modularization.","marker":"[47]"},{"why":"Keyframer shows LLM-driven animation from text; the paper distinguishes its multi-element interaction control from single-animation prompting.","marker":"[57]"},{"why":"LogoMotion provides a comparison point for content-aware code generation that lacks explicit control over individual elements.","marker":"[34]"},{"why":"CoLadder supports hierarchical decomposition for programmers; the paper adapts the decomposition idea for non-programmers creating interactive scenes.","marker":"[75]"},{"why":"Supplies the design framework for object-oriented interaction with LLMs that the element-module design extends.","marker":"[28]"},{"why":"Provides the content-analysis methodology used to identify the three challenges the system addresses.","marker":"[20]"}],"fun_headline_variants":["Modular LLMs plus sliders make 2D scene coding-free","Each scene element gets its own LLM for precise control","Graphical controls tame LLM output for interactive 2D scenes","One LLM per object: coding-free 2D scenes","Split LLM per element cuts interactive scene coding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the baseline was not artificially handicapped, yet every participant used MoGraphGPT first and the baseline's generated code was hidden, so part of the measured advantage may be practice or setup rather than the system itself.","fun_headline_variants_meta":{"raw":{"variants":["Modular LLMs plus sliders make 2D scene coding-free","Each scene element gets its own LLM for precise control","Graphical controls tame LLM output for interactive 2D scenes","One LLM per object: coding-free 2D scenes","Split LLM per element cuts interactive scene coding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000841,"raw_usage":{"total_tokens":3656,"prompt_tokens":925,"completion_tokens":2731,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":2646}},"tokens_in":541,"tokens_out":2731,"duration_ms":19563,"temperature":1.0,"reasoning_tokens":2646,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T20:42:42.513994+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three tasks with tool order counterbalanced and the baseline's generated code visible to participants, as in normal use of an AI code editor; if the completion time, prompt count, prompt length, and rating advantages of MoGraphGPT shrink or disappear, the reported improvement would be traced to ordering and hidden code rather than to modularization and graphical control.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Serves as the baseline system in the comparative evaluation that the central claim depends on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"DirectGPT motivates input-side graphical control; the paper extends this to multi-element dynamic scenes with parameter sliders."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ChatDev supplies the agent-level modularization baseline that the paper contrasts with its element-level modularization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the content-analysis methodology used to identify the three challenges the system addresses."}],"review_version":1}