REVIEW 4 major objections 5 minor 35 references
The Dynamic Creativity of Proto-artifacts in Generative Computational Co-creation
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that judging unfinished computational co-creative artifacts requires only value and novelty, not Boden's three attributes.
desk verdict A clearly written exploratory study whose headline claim—that two attributes suffice—rests on treating a non-significant novelty–surprise difference as equivalence; worth a referee but not acceptance as is. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the paired statistical comparison of creativity attributes in a repeated-measures setting: each subject rates every proto-artifact on three six-step Likert scales (value, novelty, surprise), and one-way ANOVA followed by Tukey HSD tests determine which attributes are statistically separable. The load-bearing identity is the apparent collapse of novelty and surprise into one perceived dimension: because their mean ratings across all participants and across each training subgroup were statistically indistinguishable, the paper treats originality as able to stand in for both, reducing the attribute space from three to two. The New Electronic Assistant (NEA), a generative music system trained on classical and pop melodic styles, supplies the unfinished pieces that instantiate the proto-artifacts being evaluated.
What would settle it
A within-subject replication with a pre-specified equivalence bound (or a Bayes factor) on the novelty–surprise difference would settle the claim: if the difference is found to be non-negligible, or if expert listeners can reliably separate novelty from surprise in a classification task, the two-attribute reduction is falsified. Until such a test is run, the non-significant Tukey p-values remain the main evidence.
Extended reading notes
Core claim
The paper's central claim is that, when people assess unfinished artifacts produced in a computational co-creative process, their appraisals of novelty and surprise are not discernible, so a two-attribute definition of creativity (value plus originality) can account for Boden's three-attribute definition (value, novelty, surprise). The evidence comes from a listening experiment in which subjects rated one-minute pieces generated by the New Electronic Assistant (NEA) on three Likert scales. A one-way ANOVA found an overall difference among the three score sets (F(2)=14.9, p<0.005), but Tukey HSD tests showed that value differed from novelty (p=0.001) and from surprise (p=0.0018), whereas novelty and surprise did not differ significantly (p=0.119); the same pattern held within low, mid, and high musical training subgroups. The paper presents this as empirical support for Corazza's dynamic definition of creativity and recommends using value and originality as the two operational dimensions.
Load-bearing premise
The load-bearing premise is that a statistically non-significant difference between novelty and surprise ratings counts as evidence that people do not distinguish the two; that step presupposes an equivalence threshold or a Bayesian comparison that the paper does not specify.
Editorial extensions
If this is right
- If two attributes suffice, large-scale evaluations of computational co-creative processes can cut the number of rating scales per artifact from three to two, saving effort in human studies.
- Human and AI estimators could filter the growing stream of unfinished artifacts produced by generative assistants using only value and originality.
- Because domain expertise showed up mainly in value ratings, a two-attribute model preserves the most expert-sensitive dimension while dropping a non-discernible one.
- The empirical support for Corazza's dynamic definition gives practitioners a practical measurement vocabulary for judging intermediate creative work.
Reading between the lines
- Extension: the paper treats a non-significant p-value as evidence that novelty and surprise are interchangeable; an equivalence test with a pre-specified bound, or a Bayes factor, would turn that statistical non-difference into a formal claim of equivalence.
- Extension: since the coupling of novelty and surprise is described as a general cognitive pattern, the two-attribute reduction may carry over to other co-creative domains such as design, writing, or video-game content, not just music.
- Extension: an automated estimator could approximate originality by comparing a generated artifact against its training corpus, leaving value as the only dimension that needs human judgment, a division of labor the paper points toward but does not develop.
- Extension: the result suggests a minimal evaluation protocol for generative co-creative systems—ask only 'is this worth pursuing?' and 'is this unlike what came before?'—which could be tested as a replacement for longer creativity questionnaires.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an active listening experiment (N=43 after excluding three inconsistent raters) in which participants judged unfinished musical pieces generated by the multi-NEA computational co-creative system on three attributes: value, novelty, and surprise. A one-way ANOVA and Tukey HSD tests showed that value ratings differed significantly from novelty and surprise ratings, while the novelty–surprise difference was not statistically significant (overall p = 0.119; subgroup p values 0.642, 0.438, 0.370). On this basis, the authors argue that a two-attribute definition of creativity (value plus originality) suffices to assess proto-artifacts in computational co-creation, potentially replacing Boden's three-attribute model. They also report training-level effects and discuss implications for human and agent-based assessment of unfinished artifacts.
Significance. If the central claim were established, the paper would offer a practically useful simplification of creativity metrics for evaluating the large number of unfinished artifacts produced in computational co-creative processes. The study is one of the few attempts to empirically relate Boden's and Corazza's attribute definitions in a concrete generative-music setting, and it draws attention to an important evaluation problem. However, the significance is currently limited by the statistical and conceptual gaps described in the major comments; the evidence does not yet support the strong reduction claimed in the abstract.
major comments (4)
- [§4.2, Table 1] The paper's central inference that novelty and surprise are non-discernible rests solely on non-significant Tukey HSD p-values (overall novelty–surprise p = 0.119; subgroup p = 0.642, 0.438, 0.370). A non-significant difference does not constitute evidence of equivalence, especially with only 43 participants and 20 stimuli from a single system. No equivalence test (e.g., two one-sided tests), Bayes factor, effect size, or pre-specified equivalence margin is provided. The claim in Sections 5 and 6 that novelty and surprise are 'not discernible' or 'not statistically distinguishable' therefore does not follow from the inferential machinery used. At minimum, the authors should add an equivalence analysis or explicitly rephrase the conclusion as 'no difference was detected' rather than 'non-discernible.'
- [§4.1 and §6] The proposed two-attribute model replaces novelty with originality: Section 6 suggests 'the dimensions of value and originality (rather than Corazza's effectiveness and originality).' However, the listening session in Section 4.1 only asked participants 'how novel does it seem to you?'; no originality rating was collected. The conceptual mapping from novelty ratings to originality is made post hoc and is not validated. Even if novelty and surprise were empirically indistinguishable, it does not follow that originality and novelty are the same construct, so the specific recommendation of a value-plus-originality model is not directly tested by this dataset.
- [Abstract and §6] The abstract states that a two-attribute definition 'suffices to assess unfinished work leading to innovative products,' whereas Section 6 only says a two-attribute definition 'could account for' Boden's model and that it 'is not invalidated.' The stronger claim of sufficiency is not supported by the study design, which involved one generative system (multi-NEA), one musical style, and 20 stimuli. The generalization to proto-artifacts in CCC processes across domains overreaches. The authors should either restrict their claim to the experimental context or provide additional evidence that the reduction holds across systems and domains.
- [§5] The Discussion acknowledges that 'one could argue that the observed proximity between the novelty and surprise concepts can result from the experimental conditions,' but this caveat is not carried into the conclusions. Because the experiment asked participants to rate novelty and surprise on the same piece immediately after listening, the non-significant difference could reflect a method artifact, such as scale-use consistency or task demands, rather than conceptual equivalence. The authors should address this alternative explanation, for example with item-level analyses, response-time data, or a manipulation check.
minor comments (5)
- [Table 1] The layout of Table 1 is ambiguous: the header 'Attribute pair' lists three pairs, but the subsequent columns repeat 'mean sd p' without making clear which mean/sd/p corresponds to which pair. Please restructure the table so each attribute pair has its own set of columns.
- [Figures 1 and 2] The captions for Figures 1 and 2 do not describe the axis labels or the scale used; please add explicit axis labels, units, and sample sizes to make the figures self-contained.
- [References] Reference [12] contains a typo: 'Four pppperspectives' should be 'Four perspectives.'
- [Author block] The correspondence email in the header contains stray symbols ('/envel⌢pe-⌢penjsal@illinois.edu'); please correct it.
- [§4.2] The exclusion rule (a difference of more than 3 points on all three attributes for duplicated control stimuli) is not reported as a pre-registered criterion. Please provide additional details on how many control comparisons were performed and how many additional subjects, if any, showed inconsistent responses.
Circularity Check
No significant circularity: the two-attribute conclusion is an empirical inference from external ratings, not a derivation from its own inputs.
full rationale
The paper contains no fitted parameters, no equations that reduce to their own inputs, and no load-bearing self-citations. The central claim that value and novelty/originality suffice for assessing proto-artifacts is supported by an active-listening experiment in which 43 subjects rated musical stimuli on three externally defined attributes (value, novelty, surprise). These ratings are independent of the paper's conclusion; the conclusion is not used to construct the ratings or select the stimuli. The statistical comparison of novelty versus surprise uses Tukey HSD p-values, and the paper interprets non-significance as non-discernibility. That is a genuine inferential weakness—absence of evidence is not evidence of absence without an equivalence bound or Bayesian analysis—but it is a correctness or statistical-reasoning concern, not circularity. The paper even frames the inference cautiously ('one could cautiously argue that two-attribute models of creativity could suffice') and explicitly acknowledges alternative explanations, such as the unfinished nature of the stimuli confounding subjective assessment. The later mapping of originality onto novelty and surprise is a post-hoc conceptual interpretation, not a definitional equivalence built into the method. References to Corazza and Boden are standard external frameworks, and no cited uniqueness theorem or prior work by the present authors is used to force the conclusion. The derivation chain is therefore self-contained with respect to circularity, and the appropriate score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Six-point Likert responses can be treated as interval data for ANOVA and Tukey tests.
- ad hoc to paper A non-significant p-value for the difference between novelty and surprise ratings implies the two attributes are non-discernible.
- domain assumption The 18 excerpts generated by one multi-NEA system trained on classical and pop styles are representative of proto-artifacts in computational co-creativity generally.
- domain assumption Self-reported musical training on a 1 to 6 scale meaningfully partitions subjects into low, mid, and high expertise groups.
Cite this review
Pith. "Pith review of The Dynamic Creativity of Proto-artifacts in Generative Computational Co-creation." pith.science (2026). https://pith.science/paper/AI5YEEK7
@misc{pith2026241116919,
author = {Pith},
title = {Pith review of: The Dynamic Creativity of Proto-artifacts in Generative Computational Co-creation},
year = {2026},
howpublished = {\url{https://pith.science/paper/AI5YEEK7}},
note = {Machine review of arXiv:2411.16919}
}
read the original abstract
This paper explores the attributes necessary to determine the creative merit of intermediate artifacts produced during a computational co-creative process (CCC) in which a human and an artificial intelligence system collaborate in the generative phase of a creative project. In an active listening experiment, subjects with diverse musical training (N=43) judged unfinished pieces composed by the New Electronic Assistant (NEA). The results revealed that a two-attribute definition based on the value and novelty of an artifact (e.g., Corazza's effectiveness and novelty) suffices to assess unfinished work leading to innovative products, instead of Boden's classic three-attribute definition of creativity (value, novelty, and surprise). These findings reduce the creativity metrics needed in CCC processes and simplify the evaluation of the numerous unfinished artifacts generated by computational creative assistants.
Figures
Reference graph
Works this paper leans on
-
[1]
W. Huang, H. Zheng, Architectural drawings recog- nition and generation through machine learning, in: Proceedings of the 38th Annual Conference of the Association for Computer Aided Design in Archi- tecture (ACADIA), CumInCad, 2018, pp. 156–165. doi:10.52842/conf.acadia.2018.156
-
[2]
H. Osone, J.-L. Lu, Y. Ochiai, Buncho: ai supported story co-creation via unsupervised multitask learn- ing to increase writers’ creativity in japanese, in: Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, 2021, pp. 1–10
work page 2021
-
[3]
M. Avdeeff, Artificial intelligence & popular mu- sic: Skygge, flow machines, and the audio un- canny valley, Arts 8 (2019) 130. doi: 10.3390/ arts8040130
work page 2019
-
[4]
V. Volz, J. Schrum, J. Liu, S. M. Lucas, A. Smith, S. Risi, Evolving mario levels in the latent space of a deep convolutional generative adversarial network, in: Proceedings of the genetic and evolutionary computation conference, 2018, pp. 221–228
work page 2018
-
[5]
A. Jordanous, A standardised procedure for eval- uating creative systems: Computational creativity evaluation based on what it is to be creative, Cog- nitive Computation 4 (2012) 246–279
work page 2012
- [6]
-
[7]
C. Lamb, D. G. Brown, C. L. Clarke, Evaluating com- putational creativity: An interdisciplinary tutorial, ACM Computing Surveys (CSUR) 51 (2018) 1–34
work page 2018
-
[8]
H. Herndon, M. Dryhurst, Latent visions, promptism and the future of ai art with ad- verb [audio podcast episode], NPR, 2021. URL: https://interdependence.fm/episodes/ latent-visions-promptism-and-the-future-of-ai-art-with-adverb
work page 2021
Show all 35 references
-
[9]
M. A. Runco, G. J. Jaeger, The standard definition of creativity, Creativity research journal 24 (2012) 92–96
2012
-
[10]
G. E. Corazza, Potential originality and effective- ness: The dynamic definition of creativity, Creativ- ity research journal 28 (2016) 258–267
2016
-
[11]
M. A. Boden, Creativity and art: Three roads to surprise, Oxford University Press, 2010
2010
-
[12]
Jordanous, Four pppperspectives on computa- tional creativity in theory and in practice, Connec- tion Science 28 (2016) 194–216
A. Jordanous, Four pppperspectives on computa- tional creativity in theory and in practice, Connec- tion Science 28 (2016) 194–216
2016
-
[13]
T. Lubart, How can computers be partners in the creative process: classification and commentary on the special issue, International Journal of Human- Computer Studies 63 (2005) 365–369
2005
-
[14]
N. M. Davis, Human-computer co-creativity: Blend- ing human and computational creativity, in: Ninth Artificial Intelligence and Interactive Digital Enter- tainment Conference, 2013, pp. 9–12
2013
-
[15]
Miller, E
G. Miller, E. Galanter, K. Pribram, Plans and the Structure of Behavior, Martino Publishing, USA, 1960
1960
-
[16]
Carver, M
C. Carver, M. Scheier, Attention and Self-Regulation : A Control-Theory Approach to Human Behavior, New York: Springer-Verlag, 1981
1981
-
[17]
Wallas, The art of thought, volume 10, Harcourt, Brace, 1926
G. Wallas, The art of thought, volume 10, Harcourt, Brace, 1926
1926
-
[18]
Sadler-Smith, Wallas’ four-stage model of the cre- ative process: More than meets the eye?, Creativity Research Journal 27 (2015) 342–352
E. Sadler-Smith, Wallas’ four-stage model of the cre- ative process: More than meets the eye?, Creativity Research Journal 27 (2015) 342–352
2015
-
[19]
Csikszentmihalyi, Flow and the psychology of discovery and invention, HarperPerennial, New York 39 (1997)
M. Csikszentmihalyi, Flow and the psychology of discovery and invention, HarperPerennial, New York 39 (1997)
1997
-
[20]
D. K. Simonton, Creativity and discovery as blind variation: Campbell’s (1960) bvsr model after the half-century mark, Review of General Psychology 15 (2011) 158–174
1960
-
[21]
T. B. Ward, S. M. Smith, R. A. Finke, Creative cogni- tion, in: R. J. Sternberg (Ed.), Handbook of Creativ- ity, Cambridge University Press, 1998, p. 189–212. doi:10.1017/CBO9780511807916.012
1998 doi
-
[22]
Amabile, Componential theory of creativity, Har- vard Business School Boston, MA, 2011
T. Amabile, Componential theory of creativity, Har- vard Business School Boston, MA, 2011
2011
-
[23]
L.-C. Yang, A. Lerch, On the evaluation of gen- erative models in music, Neural Computing and Applications 32 (2020) 4773–4784
2020
-
[24]
M. I. Stein, Creativity and culture, The journal of psychology 36 (1953) 311–322
1953
-
[25]
M. A. Boden, The creative mind: Myths and mecha- nisms, Routledge, 2004
2004
-
[26]
G. A. Wiggins, A preliminary framework for de- scription, analysis and comparison of creative sys- tems, Knowledge-Based Systems 19 (2006) 449–458. doi:10.1016/j.knosys.2006.04.009, creative Systems
2006 doi
-
[27]
Grace, M
K. Grace, M. L. Maher, Expectation-based models of novelty for evaluating computational creativity, in: Computational Creativity, Springer, 2019, pp. 195–209
2019
-
[28]
M. E. Q. Gonzalez, et al., Creativity: Surprise and abductive reasoning, Semiotica 2005 (2005) 325– 342
2005
-
[29]
R. W. Weisberg, On the usefulness of “value” in the definition of creativity, Creativity Research Journal 27 (2015) 111–124. doi:10.1080/10400419.2015. 1030320
2015
-
[30]
V. P. Glăveanu, Creativity as a sociocultural act, The Journal of Creative Behavior 49 (2015) 165–180
2015
-
[31]
Heinich, A pragmatic redefinition of value (s): Toward a general model of valuation, Theory, Cul- ture & Society 37 (2020) 75–94
N. Heinich, A pragmatic redefinition of value (s): Toward a general model of valuation, Theory, Cul- ture & Society 37 (2020) 75–94
2020
-
[32]
Dewey, Theory of valuation., International ency- clopedia of unified science (1939)
J. Dewey, Theory of valuation., International ency- clopedia of unified science (1939)
1939
-
[33]
H. A. Xu, A. Modirshanechi, M. P. Lehmann, W. Ger- stner, M. H. Herzog, Novelty is not surprise: Human exploratory and adaptive behavior in sequential decision-making, PLOS Computational Biology 17 (2021) e1009070
2021
-
[34]
Maguire, P
R. Maguire, P. Maguire, M. T. Keane, Making sense of surprise: an investigation of the factors influenc- ing surprise judgments., Journal of Experimental Psychology: Learning, Memory, and Cognition 37 (2011) 176
2011
-
[35]
Kantosalo, P
A. Kantosalo, P. T. Ravikumar, K. Grace, T. Takala, Modalities, styles and strategies: An interaction framework for human-computer co-creativity., in: ICCC, 2020, pp. 57–64
2020
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.