{"id":"497c04aa-2fb3-4436-aeb6-7fced41109b6","arxiv_id":"2411.16919","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A listening experiment with 43 participants suggests that novelty and surprise ratings of unfinished AI-generated music are statistically indistinguishable, so a two-attribute (value and originality) creativity assessment may suffice for proto-artifacts in computational co-creation.","lead":"In a listening test with 43 participants, researchers asked people to rate unfinished AI-generated music on novelty, surprise, and value. They claim novelty and surprise are rated so similarly that two dimensions, value and originality, are enough to judge creative potential during human-AI co-creation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's two-attribute conclusion rests on treating a non-significant Tukey novelty–surprise difference as evidence of equivalence; without a pre-specified equivalence bound or Bayes factor, the central reduction is unsupported.","rationale":"Read in good faith, the paper asks whether a two-attribute model can replace Boden's three-attribute model for evaluating unfinished co-creative artifacts. Their experiment compares mean ratings of value, novelty, and surprise on a Likert scale. The key empirical bridge is the observed similarity of novelty and surprise ratings. The reader's weakest assumption identifies exactly this bridge: the paper moves from 'we could not reject equal means' (Tukey p=0.119 overall) to 'novelty and surprise are not discernible' and then to 'two attributes suffice.' This is the canonical absence-of-evidence-as-evidence-of-absence problem, and it is load-bearing because every conclusion in Sections 5 and 6 leans on it. No equivalence margin, TOST, or Bayes factor is reported; the paper even concedes in Section 5 that observed proximity 'can result from the experimental conditions,' yet proceeds to the two-attribute recommendation. There is no machine-checked proof or released data that could independently support the claim; reproducibility is limited by the absence of data/code. I agree with the reader's assessment and would keep the REJECT verdict, with the clear path to revision being equivalence-based inference, effect-size reporting, and data release.","tokens_in":10045,"tokens_out":3508,"duration_ms":36476,"concrete_test":"Ask the authors for the raw per-participant ratings of the 20 stimuli (or re-analyze if released) and perform an equivalence analysis on the novelty–surprise difference with a pre-registered bound δ, for example ±0.3 on the 6-point scale or a scale-unit fraction justified as negligible; compute the 90% confidence interval and/or a Bayes factor (e.g., scaled JZS prior) overall and within training subgroups. If the interval is not fully within ±δ and the Bayes factor does not favor the null/equivalence model, the claim that value + novelty suffices is not supported and the verdict should be revised to exploratory only.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim—that value + novelty suffices and Boden's surprise can be dropped—depends entirely on the claim that subjects' novelty and surprise ratings are not discernible. The supporting evidence is only a set of Tukey HSD p-values: p=0.119 overall and p=0.642/0.438/0.370 across training groups (Section 4.2, Table 1). A non-significant difference is not evidence of no difference; it can also arise from insufficient power or high variance. The paper has no pre-specified equivalence margin, no two-one-sided tests, and no Bayes factor, so the interval of plausible novelty–surprise differences has not been shown to be practically negligible. The conclusion in Section 6 that 'a two-attributes definition of creativity could account for Boden's three-attributes definition' therefore does not follow from the inferential machinery used. Even granting the statistical claim, the paper substitutes 'originality' for Boden's 'novelty' in the proposed two-attribute model without measuring originality independently, a second gap in the same argumentative chain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an active listening experiment (N=43 after excluding three inconsistent raters) in which participants judged unfinished musical pieces generated by the multi-NEA computational co-creative system on three attributes: value, novelty, and surprise. A one-way ANOVA and Tukey HSD tests showed that value ratings differed significantly from novelty and surprise ratings, while the novelty–surprise difference was not statistically significant (overall p = 0.119; subgroup p values 0.642, 0.438, 0.370). On this basis, the authors argue that a two-attribute definition of creativity (value plus originality) suffices to assess proto-artifacts in computational co-creation, potentially replacing Boden's three-attribute model. They also report training-level effects and discuss implications for human and agent-based assessment of unfinished artifacts.","tokens_in":10236,"tokens_out":4588,"duration_ms":44199,"significance":"If the central claim were established, the paper would offer a practically useful simplification of creativity metrics for evaluating the large number of unfinished artifacts produced in computational co-creative processes. The study is one of the few attempts to empirically relate Boden's and Corazza's attribute definitions in a concrete generative-music setting, and it draws attention to an important evaluation problem. However, the significance is currently limited by the statistical and conceptual gaps described in the major comments; the evidence does not yet support the strong reduction claimed in the abstract.","major_comments":[{"comment":"The paper's central inference that novelty and surprise are non-discernible rests solely on non-significant Tukey HSD p-values (overall novelty–surprise p = 0.119; subgroup p = 0.642, 0.438, 0.370). A non-significant difference does not constitute evidence of equivalence, especially with only 43 participants and 20 stimuli from a single system. No equivalence test (e.g., two one-sided tests), Bayes factor, effect size, or pre-specified equivalence margin is provided. The claim in Sections 5 and 6 that novelty and surprise are 'not discernible' or 'not statistically distinguishable' therefore does not follow from the inferential machinery used. At minimum, the authors should add an equivalence analysis or explicitly rephrase the conclusion as 'no difference was detected' rather than 'non-discernible.'","section":"§4.2, Table 1"},{"comment":"The proposed two-attribute model replaces novelty with originality: Section 6 suggests 'the dimensions of value and originality (rather than Corazza's effectiveness and originality).' However, the listening session in Section 4.1 only asked participants 'how novel does it seem to you?'; no originality rating was collected. The conceptual mapping from novelty ratings to originality is made post hoc and is not validated. Even if novelty and surprise were empirically indistinguishable, it does not follow that originality and novelty are the same construct, so the specific recommendation of a value-plus-originality model is not directly tested by this dataset.","section":"§4.1 and §6"},{"comment":"The abstract states that a two-attribute definition 'suffices to assess unfinished work leading to innovative products,' whereas Section 6 only says a two-attribute definition 'could account for' Boden's model and that it 'is not invalidated.' The stronger claim of sufficiency is not supported by the study design, which involved one generative system (multi-NEA), one musical style, and 20 stimuli. The generalization to proto-artifacts in CCC processes across domains overreaches. The authors should either restrict their claim to the experimental context or provide additional evidence that the reduction holds across systems and domains.","section":"Abstract and §6"},{"comment":"The Discussion acknowledges that 'one could argue that the observed proximity between the novelty and surprise concepts can result from the experimental conditions,' but this caveat is not carried into the conclusions. Because the experiment asked participants to rate novelty and surprise on the same piece immediately after listening, the non-significant difference could reflect a method artifact, such as scale-use consistency or task demands, rather than conceptual equivalence. The authors should address this alternative explanation, for example with item-level analyses, response-time data, or a manipulation check.","section":"§5"}],"minor_comments":[{"comment":"The layout of Table 1 is ambiguous: the header 'Attribute pair' lists three pairs, but the subsequent columns repeat 'mean sd p' without making clear which mean/sd/p corresponds to which pair. Please restructure the table so each attribute pair has its own set of columns.","section":"Table 1"},{"comment":"The captions for Figures 1 and 2 do not describe the axis labels or the scale used; please add explicit axis labels, units, and sample sizes to make the figures self-contained.","section":"Figures 1 and 2"},{"comment":"Reference [12] contains a typo: 'Four pppperspectives' should be 'Four perspectives.'","section":"References"},{"comment":"The correspondence email in the header contains stray symbols ('/envel⌢pe-⌢penjsal@illinois.edu'); please correct it.","section":"Author block"},{"comment":"The exclusion rule (a difference of more than 3 points on all three attributes for duplicated control stimuli) is not reported as a pre-registered criterion. Please provide additional details on how many control comparisons were performed and how many additional subjects, if any, showed inconsistent responses.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central claim is currently stronger than the evidence supports, and the inferential gap around non-significance is the main obstacle. The paper may become publishable if the authors either add a proper equivalence analysis and measure originality separately, or substantially weaken the claim to an exploratory finding. The authors' own caveats in Section 5 are more careful than the abstract, which suggests the framing can be adjusted without changing the underlying data. The paper fits the scope of a computational creativity / co-creativity venue, but the current statistical support is too thin for acceptance as is."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent empirical comparison of Boden's three-attribute and Corazza's two-attribute creativity frameworks on unfinished musical proto-artifacts. The comparison is new as far as I know, and the paper is clearly written and honest about its limitations. But the central conclusion—that value + originality suffices and surprise can be dropped—does not follow from the statistics. The evidence for 'non-discernibility' of novelty and surprise is just a Tukey p-value of 0.119 overall (and 0.64/0.44/0.37 across training groups). A non-significant difference is not evidence of no difference, especially with N=43 and 20 stimuli. No equivalence bound, no TOST, no Bayes factor, no effect size. So the reduction claim is unsupported.\n\nWhat the paper does well: the conceptual framing around proto-artifacts and dynamic creativity is useful; the discussion of novelty vs. surprise, and the citation to Xu et al. (2021), is relevant. The experimental procedure is described clearly enough to be reproduced, except for some missing details on the NEA system. The finding that value separates from novelty/surprise, and that training affects value more than novelty/surprise, is interesting and plausible.\n\nThe soft spots beyond the equivalence issue: the authors measure 'novelty' and 'surprise' but then conclude that 'originality' can absorb both. Originality is a third construct, never measured. That is a gap in the argument. Also, three participants were excluded post hoc for inconsistent responses on control stimuli; that is defensible but should be reported as a robustness check rather than just a cleanup. Likert scales are treated as interval; minor in this context. No data or code released, which undercuts independent reanalysis.\n\nThe Discussion does acknowledge that the novelty–surprise proximity might be an artifact of the experimental conditions, and the conclusion uses 'could account' in places. The abstract, however, says 'suffices,' which overstates the evidence.\n\nBottom line: useful as a preliminary study, but the headline claim needs equivalence testing and, ideally, a direct measure of originality. I would send it to peer review—the question is timely and the empirical comparison is new—but I would expect revision before acceptance. For a reading group, it is a good case study in inference; I would bring it up. I would not cite it in my own work unless I needed an example of the non-significance fallacy.","headline":"A clearly written exploratory study whose headline claim—that two attributes suffice—rests on treating a non-significant novelty–surprise difference as equivalence; worth a referee but not acceptance as is.","tokens_in":10755,"tokens_out":3371,"would_cite":false,"duration_ms":29106,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that judging unfinished computational co-creative artifacts requires only value and novelty, not Boden's three attributes.","keywords":["computational co-creativity","creativity assessment","proto-artifacts","dynamic creativity","novelty","surprise","value","music co-creation"],"falsifier":"A within-subject replication with a pre-specified equivalence bound (or a Bayes factor) on the novelty–surprise difference would settle the claim: if the difference is found to be non-negligible, or if expert listeners can reliably separate novelty from surprise in a classification task, the two-attribute reduction is falsified. Until such a test is run, the non-significant Tukey p-values remain the main evidence.","tokens_in":9824,"feed_emoji":"🎵","tokens_out":5573,"duration_ms":47694,"temperature":0.7,"pith_summary":"This paper asks whether judging the creative merit of unfinished work produced during human–AI co-creation needs the full three-attribute definition of creativity (value, novelty, surprise) or only two. It reports an active listening experiment in which 43 participants with varied musical training rated short unfinished musical pieces generated by a system called NEA. Statistical comparisons of the ratings showed that value was scored distinctly from novelty and surprise, while novelty and surprise received statistically indistinguishable ratings. The authors conclude that a two-attribute model—value plus originality—suffices to assess proto-artifacts in computational co-creative processes, which would simplify evaluation as generative assistants produce growing numbers of intermediate artifacts.","feed_headline":"Measuring unfinished AI-made music may need just two scales","feed_subtitle":"Listeners in a 43-person study could not tell novelty from surprise, so value plus novelty may cover all three.","key_machinery":"The machinery is the paired statistical comparison of creativity attributes in a repeated-measures setting: each subject rates every proto-artifact on three six-step Likert scales (value, novelty, surprise), and one-way ANOVA followed by Tukey HSD tests determine which attributes are statistically separable. The load-bearing identity is the apparent collapse of novelty and surprise into one perceived dimension: because their mean ratings across all participants and across each training subgroup were statistically indistinguishable, the paper treats originality as able to stand in for both, reducing the attribute space from three to two. The New Electronic Assistant (NEA), a generative music system trained on classical and pop melodic styles, supplies the unfinished pieces that instantiate the proto-artifacts being evaluated.","core_discovery":"The paper's central claim is that, when people assess unfinished artifacts produced in a computational co-creative process, their appraisals of novelty and surprise are not discernible, so a two-attribute definition of creativity (value plus originality) can account for Boden's three-attribute definition (value, novelty, surprise). The evidence comes from a listening experiment in which subjects rated one-minute pieces generated by the New Electronic Assistant (NEA) on three Likert scales. A one-way ANOVA found an overall difference among the three score sets (F(2)=14.9, p<0.005), but Tukey HSD tests showed that value differed from novelty (p=0.001) and from surprise (p=0.0018), whereas novelty and surprise did not differ significantly (p=0.119); the same pattern held within low, mid, and high musical training subgroups. The paper presents this as empirical support for Corazza's dynamic definition of creativity and recommends using value and originality as the two operational dimensions.","pith_inferences":["Extension: the paper treats a non-significant p-value as evidence that novelty and surprise are interchangeable; an equivalence test with a pre-specified bound, or a Bayes factor, would turn that statistical non-difference into a formal claim of equivalence.","Extension: since the coupling of novelty and surprise is described as a general cognitive pattern, the two-attribute reduction may carry over to other co-creative domains such as design, writing, or video-game content, not just music.","Extension: an automated estimator could approximate originality by comparing a generated artifact against its training corpus, leaving value as the only dimension that needs human judgment, a division of labor the paper points toward but does not develop.","Extension: the result suggests a minimal evaluation protocol for generative co-creative systems—ask only 'is this worth pursuing?' and 'is this unlike what came before?'—which could be tested as a replacement for longer creativity questionnaires."],"forward_implications":["If two attributes suffice, large-scale evaluations of computational co-creative processes can cut the number of rating scales per artifact from three to two, saving effort in human studies.","Human and AI estimators could filter the growing stream of unfinished artifacts produced by generative assistants using only value and originality.","Because domain expertise showed up mainly in value ratings, a two-attribute model preserves the most expert-sensitive dimension while dropping a non-discernible one.","The empirical support for Corazza's dynamic definition gives practitioners a practical measurement vocabulary for judging intermediate creative work."],"supporting_citations":[{"why":"Corazza's dynamic definition of creativity as potential originality and effectiveness; the two-attribute model the paper adopts.","marker":"[10]"},{"why":"Boden's three-attribute definition of creativity (value, novelty, surprise) that the paper aims to reduce.","marker":"[11]"},{"why":"Runco and Jaeger's standard definition of creativity, the source of the originality/effectiveness pairing.","marker":"[9]"},{"why":"Xu et al.'s dissociation of novelty and surprise as coupled but distinct cognitive operations, used to interpret the non-discernibility.","marker":"[33]"},{"why":"Maguire et al.'s finding that surprise judgments depend on expertise, invoked to explain training-level effects.","marker":"[34]"},{"why":"Kantosalo et al.'s two-dimensional co-creativity assessment model, cited as a precedent for value-plus-novelty frameworks.","marker":"[35]"}],"fun_headline_variants":["Novelty and surprise indistinguishable in music evaluation","Two metrics instead of three for unfinished AI art","Value and novelty cover creativity in proto-artifacts","Shed surprise: value and novelty judge AI co-creation","In AI co-creation, surprise merges into novelty"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a statistically non-significant difference between novelty and surprise ratings counts as evidence that people do not distinguish the two; that step presupposes an equivalence threshold or a Bayesian comparison that the paper does not specify.","fun_headline_variants_meta":{"raw":{"variants":["Novelty and surprise indistinguishable in music evaluation","Two metrics instead of three for unfinished AI art","Value and novelty cover creativity in proto-artifacts","Shed surprise: value and novelty judge AI co-creation","In AI co-creation, surprise merges into novelty"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000289,"raw_usage":{"total_tokens":1662,"prompt_tokens":880,"completion_tokens":782,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":705}},"tokens_in":496,"tokens_out":782,"duration_ms":7416,"temperature":1.0,"reasoning_tokens":705,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:44:17.230198+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A within-subject replication with a pre-specified equivalence bound (or a Bayes factor) on the novelty–surprise difference would settle the claim: if the difference is found to be non-negligible, or if expert listeners can reliably separate novelty from surprise in a classification task, the two-attribute reduction is falsified. Until such a test is run, the non-significant Tukey p-values remain the main evidence.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Corazza's dynamic definition of creativity as potential originality and effectiveness; the two-attribute model the paper adopts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Boden's three-attribute definition of creativity (value, novelty, surprise) that the paper aims to reduce."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Runco and Jaeger's standard definition of creativity, the source of the originality/effectiveness pairing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Xu et al.'s dissociation of novelty and surprise as coupled but distinct cognitive operations, used to interpret the non-discernibility."},{"cited_title":"Maguire, P","cited_arxiv_id":null,"evidence_quote":"Maguire et al.'s finding that surprise judgments depend on expertise, invoked to explain training-level effects."},{"cited_title":"Kantosalo, P","cited_arxiv_id":null,"evidence_quote":"Kantosalo et al.'s two-dimensional co-creativity assessment model, cited as a precedent for value-plus-novelty frameworks."}],"review_version":1}