{"id":"63d28e95-02fb-4d10-9bb2-29891d9d8384","arxiv_id":"2504.14071","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A minimal 1-parameter AI music plugin was rated usable and acceptable by expert composers, but the study finds the single control insufficient for professional creative goals.","lead":"MMM-Cubase (MMM-C) is a one-parameter AI plugin that helps composers generate music inside the Cubase DAW. In a study of 18 expert composers, it scored as usable and easy to operate, but users reported that the single temperature knob gave them too little control over the output.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"TAM scores sit near the neutral midpoint, so the abstract's 'positive acceptance' is an overstatement; the acceptance component of the central claim is not supported by the paper's own statistics.","rationale":"The reader's verdict of CONDITIONAL is appropriate, and this stress-test does not move it. The reader's weakest_assumption focuses on attrition from 34 onboarded to 18 participants; that is a genuine generalizability concern, but it is less decisive for the central claim because the 16 non-completers mainly never started the study, and there is no evidence they tried the tool and quit in frustration. The TAM overstatement is more directly load-bearing: it is an internal inconsistency between the reported numbers and the paper's own summary, it is checkable from data already in the paper, and it concerns the 'positive acceptance' part of the abstract's central claim. The usability component is supported by SUS means of 71–76, and the insufficiency claim is supported by controllability and qualitative results, so these parts stand. The required fix is textual and statistical: test TAM scores against the neutral midpoint, report the tests, and soften the acceptance wording accordingly. That is exactly the kind of revision a CONDITIONAL verdict should request.","tokens_in":13320,"tokens_out":7488,"duration_ms":67938,"concrete_test":"Re-analyze the TAM subscale data reported in §6.4 with one-sample Wilcoxon signed-rank tests against the neutral midpoint of 3, separately for perceived ease of use and perceived usefulness in each task, and report exact p-values, 95% confidence intervals, and effect sizes. If the means are not significantly above 3 (or effect sizes are small), the abstract and Discussion must be revised to say 'neutral-to-slightly-positive' or 'mixed' rather than 'positive acceptance'; if they are significantly above 3, the acceptance claim survives this specific objection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Discussion claim 'positive usability and acceptance scores,' but the TAM data in §6.4 are 3.12–3.26/5 for perceived ease of use and 3.42–3.68/5 for perceived usefulness, with SDs around 0.5–0.96. On a 5-point Likert scale, 3 is the neutral point; these means are within a fraction of a point of neutral. The paper itself concedes in §6.4 that participant scores 'are above average, however not significant enough to be conclusive,' yet the Discussion and abstract convert that into 'positive' acceptance. No test against the neutral midpoint, no confidence intervals, and no multiple-comparison correction are reported; the only p-value in the quantitative results (task 1 vs task 3 friendliness, p=0.03) is an uncorrected pairwise comparison. As a result, the 'positive acceptance' component of the central claim is an overstatement of the measurements. The '1-parameter design is not enough' component is better supported by the controllability items (ease of control 5.23/10; desire for more control 9.54/10) and by the qualitative themes, but the acceptance claim as written should be downgraded.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a mixed-methods evaluation of MMM-Cubase (MMM-C), a \"1-parameter\" plugin interface to the MMM generative music model, integrated into the Cubase DAW. Eighteen expert composers (8 hobbyist, 10 professional) completed tasks involving arrangement, variation, and original composition. The authors measure usability (SUS, CSI, custom controllability items), user experience (qualitative coding of open-ended questions), and technology acceptance (TAM scales), and they report quantitative and qualitative results. The central claims are that MMM-C receives positive usability and acceptance scores, that a 1-parameter design is insufficient for expert composers' goal-directed work, and that no significant differences exist between hobbyist and professional groups.","tokens_in":13518,"tokens_out":2219,"duration_ms":21022,"significance":"If the claims were fully supported, the study would be a valuable addition to the human-AI co-creation literature, particularly for its use of expert composers, its integration into a professional DAW, and its combination of standardized instruments (SUS, CSI, TAM) with qualitative coding. The finding that a minimal interface to a powerful generative model is usable for exploration but insufficient for controlled composition is a plausible and practically relevant contribution, and the study's methodology could serve as a benchmark for future evaluations of richer MMM interfaces such as Calliope. However, the strength of the acceptance claim is not backed by the reported statistics, and the attrition pattern raises questions about selection bias. The qualitative and controllability data do support the \"1-parameter is not enough\" conclusion, but the acceptance component of the abstract and Discussion needs substantial qualification.","major_comments":[{"comment":"The claims of \"positive acceptance\" in the Abstract and \"acceptance levels are positive\" in the Discussion are not supported by the reported TAM data. Perceived ease of use ranges from 3.12 to 3.26 and perceived usefulness from 3.42 to 3.68 on a 5-point scale, with standard deviations around 0.5–0.96; 3 is the neutral midpoint, and the means lie within a fraction of a point of it. The paper itself concedes in §6.4 that the scores are \"not significant enough to be conclusive.\" To support the acceptance claim, the authors should either report a formal test against the neutral midpoint (e.g., one-sample Wilcoxon or t-test) with confidence intervals, or rephrase the claim to \"neutral-to-slightly-positive\" and clearly label it as inconclusive. As written, the abstract overstates the results.","section":"§6.4, Abstract, §7 Discussion"},{"comment":"The manuscript reports that 18 of 34 participants who completed onboarding actually took part in the study, a 47% attrition rate, but provides no analysis of whether the non-completers differ from completers in ways that could bias the results. If participants dropped out because they found the tool unusable or frustrating, the reported SUS and TAM scores would be systematically inflated. The authors should compare demographic or early-task data (e.g., onboarding survey responses) between completers and non-completers, or at minimum discuss this as a distinct limitation and temper the strength of the quantitative claims accordingly.","section":"§6, opening paragraph"},{"comment":"The only inferential statistic reported in the quantitative results is the pairwise comparison of task-based friendliness scores between Task 1 and Task 3 (p=0.03), which is presented without correction for multiple comparisons and without an effect size or confidence interval. Given the number of comparisons implicitly made across SUS, user-friendliness, TAM, and one-value ratings, this p-value is likely to be a false positive. The authors should either apply a multiple-comparison correction, report effect sizes, or explicitly label this finding as exploratory. This does not affect the main usability conclusion (SUS scores are in the acceptable range), but the current reporting overstates the evidentiary value of this single comparison.","section":"§6.2, friendliness comparison"}],"minor_comments":[{"comment":"The word \"wiskers\" should be \"whiskers\" in the figure caption footnote, and \"frustation\" should be \"frustration.\"","section":"§6.2, bottom paragraph"},{"comment":"The CSI instrument is modified substantially (single item per factor, 5-point scale instead of 10-point, collaboration omitted), and the resulting scores are compared to published CSI benchmarks; the authors should explicitly note that the modified CSI is not directly comparable to the original instrument's normative ranges.","section":"§4.4, CSI description"},{"comment":"The phrase \"20h to 30h of effort\" for a participant seems unusually high and may be a typo; if not, the figure deserves justification since it affects the attrition interpretation.","section":"§5, study procedure"},{"comment":"Two of the 18 participants are described as \"Enthusiasts\" rather than experts; this should be acknowledged as a slight deviation from the stated participant definition, and its effect on the expert-composer claims should be briefly discussed.","section":"§6.1, demographics"},{"comment":"The paper uses both \"experiment\" and \"study\" to describe the methodology; since the design is not a controlled experiment, the term \"study\" should be used consistently.","section":"§8, Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper stems from the group that developed MMM and its interface, so there is a potential conflict of interest in evaluating their own system; this is not disqualifying, but the manuscript should explicitly disclose it and should address how the authors' dual role might have influenced the qualitative coding and interpretation. The journal should also consider whether the attrition and the borderline TAM results are adequately handled before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a real, honest empirical study: the authors built a one-parameter Cubase plugin on top of the MMM generative model, recruited 18 expert composers from Steinberg's beta pool, and ran a mixed-methods evaluation with standard instruments. Second, the abstract says 'positive usability and acceptance scores,' but the TAM data do not support the acceptance half of that sentence. The means sit at 3.1–3.7 on a 5-point scale, within a few decimal points of neutral, and the paper itself admits the results are 'not significant enough to be conclusive.' That is an overstatement, not a fatal flaw, but it should be fixed.\n\nWhat is actually new: this is the first TAM measurement for an AI music co-creation plugin, and the first evaluation of MMM-C with expert composers. The design is sound: SUS scores (71–76) support the usability claim, and the qualitative coding is well done. The central finding—that a single temperature knob is not enough for expert composers' goal-directed work—is well supported by the controllability scores (5.23/10 ease of control, 9.54/10 desire for more control) and by the quotes about feeling a lack of parameters. That part of the discussion holds up.\n\nThe soft spots are real but proportionate. The attrition from 34 onboarded to 18 participants is never analyzed; if the non-completers left because the tool was frustrating, the positive scores are inflated. The modified CSI (shortened to 5 items, collaboration removed) and the custom controllability scale are used without any validation discussion. The only significant pairwise comparison (task 1 vs task 3 friendliness, p=0.03) is uncorrected and could easily be a false positive. And the authors are the system developers, which is a bias to acknowledge, though it does not make the study circular.\n\nWho is this for? Researchers designing interfaces for generative music systems, especially anyone working on parameter design or co-creative tools. It is a useful baseline for the MMM line of work and a reasonable comparison point for Louie et al. and Magenta Studio. It deserves a serious peer review, but the revision should tone down the abstract, analyze or at least discuss the dropouts, report test statistics with confidence intervals, and address the psychometric shortcuts.\n\nIf I were the editor, I would send it out. The empirical core is valuable and the overclaims are correctable.","headline":"A credible first evaluation of a minimal one-parameter AI music plugin with expert composers, but the acceptance claim is overstated and the 47% attrition is unanalyzed.","tokens_in":14115,"tokens_out":1510,"would_cite":true,"duration_ms":13894,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One dial is not enough for expert AI music co-creation","keywords":["human-AI co-creation","generative music","usability evaluation","technology acceptance model","creativity support index","expert composers","MMM-Cubase","parameter design"],"falsifier":"Re-run the study with all onboarded participants retained (or with exit interviews for dropouts); if SUS/TAM scores fall below acceptable thresholds or controllability stops tracking usefulness, the central claim would be overturned. Alternatively, run the same protocol on a multi-parameter MMM interface; if experts still report the same steering difficulty, then parameter count is not the binding constraint.","tokens_in":13094,"feed_emoji":"🎹","tokens_out":5073,"duration_ms":41941,"temperature":0.7,"pith_summary":"This paper evaluates whether a deliberately minimal interface to a powerful generative music model is enough for expert composers working in a real digital audio workstation. The authors built MMM-Cubase (MMM-C), a plugin that exposes exactly one control — the model's temperature — inside Cubase, and ran a three-part mixed-method study with 18 hobbyist and professional composers. The central claim is that a 1-parameter design is not sufficient for goal-directed generative music co-creation by experts, even when the underlying model is expressive and highly controllable. The evidence supports positive usability and acceptance overall, with users enjoying exploration and surprise, while controllability and predictability lag behind. The paper matters because it provides a baseline for how much interface control a co-creative music AI actually needs.","feed_headline":"One dial is not enough for expert AI music co-creation","feed_subtitle":"Experts found the minimal MMM-C plugin easy and fun but hard to steer, so richer controls should follow.","key_machinery":"The central object is MMM-Cubase (MMM-C), a plugin that wraps the Multi-Track Music Machine (MMM), a transformer-based generative model for multi-track symbolic music, and exposes a single temperature parameter (0–100%, default 50%) inside the Cubase DAW. The user selects bars of MIDI, adjusts temperature, and triggers generation; the model uses surrounding vertical and horizontal context to infill tracks or bars. The evaluative machinery is a three-part mixed-method assemblage: SUS and task-based user-friendliness for usability, a shortened Creativity Support Index plus custom controllability questions for user experience, and the Technology Acceptance Model with open-ended qualitative coding for acceptance. The single-parameter design is what does the argumentative work: it isolates the question of how much control a co-creative interface must expose, and the gap between high usability and low controllability is the paper's key finding.","core_discovery":"On the paper's own terms, a single-knob interface is a passable exploratory tool but a poor steering wheel for expert composers. Quantitative scores were mostly positive: SUS scores across tasks were 71–76 (acceptable), TAM perceived usefulness was 3.42–3.68 and ease of use 3.12–3.26 on a 5-point scale, and CSI factors showed enjoyment highest (3.85/5) with expressiveness lowest (2.85/5). Controllability was the weak point: ease of control averaged 5.23/10 while desire for more control averaged 9.54/10. Qualitative coding converged on the same picture — users found the interaction easy but described difficulty steering the system, a lack of parameters, repeated generation to find acceptable output, and non-determinism as both a source of surprise and a source of frustration. The paper concludes that a 1-parameter design is not enough for generative music co-creation in the case of expert composers and finds no significant difference between hobbyists and professionals.","pith_inferences":["The paper's own attrition (34 onboarded, 18 completed at least one task) implies that the favorable scores may overstate the tool's appeal; a replication that tracks why the 16 dropped out would test this directly.","Because users said they repeatedly generated and curated outputs, the practical bottleneck may be predictability more than parameter count; adding previews, seeds, or explicit conditioning could improve control without adding many knobs.","The one-parameter baseline suggests a testable design principle: for expert composers, control requirements scale with the ambition of the task, so original composition may need more parameters than arrangement does.","No users voiced fear of work replacement in this study, which hints that expert adoption barriers for creative AI are more about controllability and trust than about displacement anxiety."],"forward_implications":["Exposing more of MMM's existing attribute controls (note density, polyphony, duration, style) should raise expressiveness and acceptance scores above the baseline reported here.","Even a one-parameter interface can serve expert composers as an exploration and inspiration tool, helping with writer's block and generating ideas the composer would not have written.","The negative controllability findings imply that interface complexity for co-creative AI should be matched to the target user's required quality bar, not just to the model's expressive capacity.","Because no hobbyist/professional difference appeared, future interface studies may not need to treat expertise level as a controlling factor for basic usability.","The same mixed-method protocol can benchmark richer interfaces (e.g., Calliope) and transfer to other DAWs or to generative tasks such as language and visual in-painting."],"supporting_citations":[{"why":"Supplies the MMM generative model and its track/bar in-filling and attribute controls, which MMM-C reduces to one parameter.","marker":"[Ens and Pasquier, 2020a]"},{"why":"Baseline showing semantic steering controls improve novice-AI music co-creation; the paper extends this comparison to expert composers and multi-track output.","marker":"[Louie et al., 2020]"},{"why":"Defines the System Usability Scale used to produce the acceptable SUS scores.","marker":"[Brooke, 1996]"},{"why":"Defines the Creativity Support Index whose factor scores quantify expressiveness and enjoyment.","marker":"[Cherry and Latulipe, 2014]"},{"why":"Defines the Technology Acceptance Model constructs that yield the perceived usefulness and ease-of-use scores.","marker":"[Davis, 1989]"},{"why":"Supports the argument that highly encapsulated systems hide parameters and hurt visibility, corroborating the controllability findings.","marker":"[Bray et al., 2017]"},{"why":"Provides the triangular mixed-methods design that merges quantitative and qualitative strands.","marker":"[Creswell et al., 2003]"},{"why":"Supplies a comparable professional-DAW generative music evaluation whose ease-of-use results anchor the discussion.","marker":"[Roberts et al., 2019]"}],"fun_headline_variants":["One-knob AI music tool fails expert composers","Expert composers need more than a single AI dial","Single-parameter AI music interface leaves experts wanting","MMM-C's one control: too simple for pros"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the 18 participants who completed at least one task fairly represent expert composers, so the 16 who dropped out after onboarding did not take their dissatisfaction with them.","fun_headline_variants_meta":{"raw":{"variants":["One-knob AI music tool fails expert composers","Expert composers need more than a single AI dial","Single-parameter AI music interface leaves experts wanting","MMM-C's one control: too simple for pros"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00061,"raw_usage":{"total_tokens":2853,"prompt_tokens":975,"completion_tokens":1878,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":1816}},"tokens_in":591,"tokens_out":1878,"duration_ms":12914,"temperature":1.0,"reasoning_tokens":1816,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:56:21.889161+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the study with all onboarded participants retained (or with exit interviews for dropouts); if SUS/TAM scores fall below acceptable thresholds or controllability stops tracking usefulness, the central claim would be overturned. Alternatively, run the same protocol on a multi-parameter MMM interface; if experts still report the same steering difficulty, then parameter count is not the binding constraint.","supporting_citations":[{"cited_title":"Novice-AI music co-creation via AI-steering tools for deep generative mod- els","cited_arxiv_id":null,"evidence_quote":"Baseline showing semantic steering controls improve novice-AI music co-creation; the paper extends this comparison to expert composers and multi-track output."},{"cited_title":"Sus: A ’quick and dirty’ us- ability scale","cited_arxiv_id":null,"evidence_quote":"Defines the System Usability Scale used to produce the acceptable SUS scores."},{"cited_title":"Quantifying the creativity support of digital tools through the creativity support index","cited_arxiv_id":null,"evidence_quote":"Defines the Creativity Support Index whose factor scores quantify expressiveness and enjoyment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the argument that highly encapsulated systems hide parameters and hurt visibility, corroborating the controllability findings."},{"cited_title":"Magenta studio: Augmenting creativity with deep learning in ableton live","cited_arxiv_id":null,"evidence_quote":"Supplies a comparable professional-DAW generative music evaluation whose ease-of-use results anchor the discussion."}],"review_version":1}