Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Large Language Models show both individual and collective creativity comparable to humans

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Across 13 creative tasks, top LLMs rank at the 52nd percentile of humans and 10 responses equal a group of 8-10 people.

desk verdict A serious multi-domain benchmark with a novel collective-creativity metric, but the headline figures hinge on an unvalidated rater equating. read the letter →

arxiv 2412.03151 v1 pith:VA36U5PF submitted 2024-12-04 cs.AI

classification cs.AI
keywords largelanguagemodelscreativitydivergentthinkingproblemsolvingcreativewritingcollectivehumanbenchmarkingfutureofwork
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Using 13 creative tasks across divergent thinking, problem solving, and creative writing, this paper asks whether large language models are as creative as the people who will soon work beside them. It finds that the best models, Claude and GPT-4, rank at the 52nd percentile of a large human sample, with models strongest on divergent thinking and problem solving and weakest on creative writing. The paper goes further than single-response benchmarks by sampling each model repeatedly and showing that one model asked ten times is equivalent, in top-idea production, to a group of eight to ten humans. Because a typical workplace brainstorming group is smaller than that, the result suggests that LLMs could serve as near-human creative teammates rather than merely routine-task tools.

What carries the argument

The load-bearing instrument is the collective-creativity equivalence metric. For each task, the authors split a model's responses into groups of ten, pool each group with responses from N randomly drawn human participants, and bootstrap this pooling 1000 times; the N at which humans and the model contribute half of the top ten responses becomes the 'number of humans the model equals.' The individual-level comparison is carried by the same percentile-ranking procedure, and both rest on ratings produced by five trained judges per round using the Consensual Assessment Technique, with the two rating rounds put on one scale by z-score linear equating anchored on repeated GPT-3.5 responses.

What would settle it

Have the same panel rate a random sample of human and LLM responses from both rounds in one sitting; if the directly observed percentile ranks or human-equivalence numbers diverge noticeably from the linearly equated values, the central claim weakens. A simpler check is to re-estimate the collective-creativity metric without the anchor responses and see whether GPT-4 and Claude still land between 8 and 10 humans.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that five LLMs tested against 467 humans on 13 bespoke creative tasks produce responses that human judges, blind to authorship, rate as roughly average-to-slightly-above human: all models together average the 46th percentile, with Claude and GPT-4 at the 52nd. In divergent thinking and problem solving the models sit at the 55th and 59th percentiles, while in creative writing they fall to the 25th. When the same model is asked ten times and its responses are pooled with those of varying numbers of humans, the model's share of the top-rated ideas matches that of 8-10 humans; asking for more responses yields a linear exchange rate in which roughly two additional LLM responses add as much as one additional human. The paper reads this as evidence that LLM creativity is already comparable to individual human creativity and, when sampled repeatedly, to the collective output of a small human group.

Load-bearing premise

The headline comparisons assume that the two groups of human raters used a common internal scale, so a linear z-score rescaling is enough to make Round 1 human scores and Round 2 LLM scores directly comparable.

Editorial extensions

If this is right

  • Workplace brainstorming with fewer than ten people can now plausibly include an LLM as a source of top-idea generation, not just a clerical aid.
  • In divergent thinking and problem-solving tasks, asking a model several times before selecting the best idea is a cheap way to match a small human group's best output.
  • Creative writing remains the domain where human superiority is clear; organizations should not expect LLMs to replace human writers on emotional or memorable messaging.
  • Because the paper replicates the finding that LLM outputs are less diverse than human outputs, teams that lean on LLMs alone should expect idea homogenisation and should keep human variability in the pipeline.
  • Single-task creativity studies will keep producing inconsistent verdicts; a multi-domain, multi-dimension battery is needed to say anything general about machine creativity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the two rating rounds differ by more than a linear shift, the 52nd percentile and the 8-10 human equivalence would move; a single-round re-rating of a common sample would show how much.
  • The equivalence to 8-10 humans is measured against a nominal group of independently working humans; interactive brainstorming groups behave differently, so the real-world team size an LLM replaces could be larger or smaller.
  • An untested extension is whether the 2-responses-per-human exchange rate holds beyond fifty responses or saturates as the model's own output diversity becomes the limiting factor.
  • Since temperature changes diversity more than rated creativity, sampling strategy (how many responses, at what temperature, with what prompt variation) may matter more than any single generation setting for collective creative output.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper benchmarks five large language models (GPT-3.5, GPT-4, Claude, Qwen, SparkDesk) against 467 human participants on 13 creativity tasks spanning divergent thinking, problem solving, and creative writing. Human raters scored all responses following the Consensual Assessment Technique, and the authors report percentiles of individual LLM performance in the human distribution, with the best models (Claude, GPT-4) at the 52nd percentile. They then introduce a collective-creativity metric that pools multiple LLM responses with varying numbers of human responders, and claim that one LLM asked 10 times is equivalent to a group of 8--10 humans, and that roughly two additional LLM responses add as much as one extra human. Domain-specific results and analyses of diversity, temperature effects, and demographic differences are also reported.

Significance. If the central results hold, this is a useful multi-domain benchmark with a novel operationalization of collective creativity that goes beyond single-response comparisons. The study has notable strengths: the tasks are not taken from existing online datasets, reducing contamination concerns; ratings are provided by trained judges with reported inter-rater reliability; the data and code are made available; and the authors transparently discuss limitations such as the non-representative human sample, the nominal-group framing, and the absence of prompt engineering. The collective-creativity equivalence is a falsifiable empirical claim rather than a derivation, and the paper clearly separates the bootstrapped pooling procedure from the linear slope fitted to the aggregate results. These strengths make the paper a potentially valuable reference point for debates about LLM creativity in applied settings.

major comments (3)
  1. [Methods: Rating procedure (pp. 22-23); Figs 2, 6, 7] The z-score linear equating that aligns Round 1 (human) and Round 2 (LLM) ratings is the backbone of every human-LLM comparison, but the manuscript provides no equating diagnostics. No anchor means, standard deviations, scatterplots, or equating error are reported for the GPT-3.5 v1 anchor responses. The anchor set is small (about 50 responses per task) and comes from a model that is generally weaker than humans, so the anchor ratings are likely concentrated at the low end of the scale; yet the collective-creativity analysis selects the top 10 responses from a pooled set, which is exactly the upper tail where anchor data are thinnest. If rater severity differs nonlinearly between rounds, the set of responses entering the top 10 would change, shifting both the 52nd-percentile claim and the 8--10-human equivalences. Please report anchor distributions per task, provide equating diagnostics, and include a robustness check (e.g., re-rating a subset of human responses in Round 2, using a different equating method, or nonparametric equating).
  2. [Results Part 1 and Part 2; Figs 2, 6, 7; Table S12] The headline numbers -- the 52nd percentile for Claude/GPT-4, the collective equivalence of 8--10 humans, and the slope of 0.52 used for 'two additional LLM responses equal one extra human' -- are reported as point estimates without confidence intervals or dispersion measures. The percentile ranks vary widely across the 13 tasks (e.g., 25th percentile in creative writing vs. 55th and 59th percentiles in divergent thinking and problem solving), so the average percentile is not characterized by its mean alone. In Fig. 7 and Table S13, the linear relationship is fitted without reporting standard errors, R-squared, or residual diagnostics; this is particularly relevant because the slope is used directly in the abstract and discussion. Please provide per-task distributions, bootstrap confidence intervals for the equivalence points, and regression fit diagnostics for the linear slopes.
  3. [Statistical Analysis; Tables S4, S5, S16-S22] The paper runs dozens of independent-sample t-tests and ANOVAs without any multiple-comparison correction. While the central percentile and collective-creativity claims are not based on these individual significance tests, the domain-level conclusions (e.g., 'LLMs excel in divergent thinking and problem solving' and specific task-by-task superiority claims) rely on patterns across many comparisons, and several would likely not survive a Benjamini-Hochberg or Bonferroni correction. Please either apply a correction to the main task-level comparisons or explicitly label the tables as exploratory and focus the interpretation on effect sizes and consistency across tasks.
minor comments (6)
  1. [Discussion (p. 17) vs Results (p. 8)] The percentile for GPT-3.5 is reported as the 37th percentile in Results and as the 38th percentile in Discussion; please make these consistent.
  2. [Table 1, reference 15] The author name 'can der Maas' should read 'van der Maas'.
  3. [Reference 39] The author name 'A. uncdogan' appears to be a corrupted rendering; please verify the spelling against the published source.
  4. [Results: 'When one LLM is asked 10 times' (p. 14) and Table S12] For GPT-4 and Claude the overall human contribution at N=10 is 48.71% and 50.03%, respectively, so the reported equivalence '10 humans' is a rounded or interpolated value; please state explicitly how the equivalence point is derived from the bootstrapped percentages and whether values are rounded.
  5. [Effect of temperature (p. 12)] The statement that 'temperature should be better interpreted as a parameter for diversity rather than a parameter for creativity' would benefit from a brief report of the effect sizes (partial eta-squared) for the non-significant creativity ANOVAs, not only the significant diversity results.
  6. [Figures 6 and 7] The figures show mean collective-creativity values without error bars or confidence bands; please indicate in the captions whether error bars are omitted because they were not computed or because they were too small to display.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical benchmark whose headline numbers are computed from independent human ratings, with no load-bearing self-citation or definitional identity between inputs and conclusions.

full rationale

The paper's central claims are empirical comparisons against external human judgments. The 52nd-percentile result is obtained by locating the mean LLM response scores in the human performance distribution (Results, Fig. 2), which is a direct benchmarking computation, not a quantity fitted to reproduce itself. The collective-creativity equivalence is explicitly an operationalization: the Methods state that the metric 'is measured by the number of humans that contribute the same proportion of top responses as an LLM when it is asked 10 times,' and the value 8–10 humans is then estimated by bootstrap sampling of human participants and LLM response groups. This is a defined measurement, not a derivation that assumes its own conclusion. The claim that 'two additional LLM responses equal one extra human' is a linear-regression slope fitted to the computed collective-creativity values, and is presented as a summary of those empirical values rather than as an independent prediction. The z-score equating between the two rating rounds uses GPT-3.5 v1 responses rated in both rounds as anchors; this is a calibration step and, although the paper reports no equating diagnostics, it is not circular because the anchor responses are actual rated responses and the final comparisons are not constrained to match the anchors. The paper cites prior work for diversity and brainstorming effects, but these citations are not load-bearing for the headline creativity comparison, and no uniqueness theorem or self-citation is invoked to forbid alternative interpretations. The 'collective creativity' framing is an operational definition, and the Discussion explicitly acknowledges it as a limitation ('Another limitation of our study is the operationalisation of the measurement of LLMs’ collective creativity'), which further confirms it is not being passed off as an independent derivation. Overall, the derivation chain is self-contained: human ratings, LLM outputs, and a transparent metric define the results.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper does not introduce new physical or conceptual entities. It relies on measurement assumptions about creativity tasks, the human sample, and cross-round rating equating. The only fitted quantities relevant to the headline are the linear slope for additional LLM responses and the equating constants.

free parameters (2)
  • Collective creativity linear slope = 0.52 equivalent humans per additional LLM response (inverse ~1.92 responses per human)
    The headline 'two additional LLM responses equal one extra human' is the inverse of a linear regression slope fitted to the collective creativity data in Figure 7 and Table S13.
  • Z-score equating constants = Not reported
    Round 2 LLM ratings were linearly equated onto Round 1 human ratings using GPT-3.5 v1 responses as anchor items; the mean and SD used for equating are estimated from data and their uncertainty is not propagated into any comparison.
assumptions (4)
  • domain assumption The 13 constructed tasks validly measure creativity across divergent thinking, problem solving, and creative writing.
    The tasks follow PISA 2022 and the Consensual Assessment Technique, but task validity is assumed rather than independently established beyond face validity.
  • domain assumption The distribution of creativity in the human sample approximates the general adult population.
    The sample is 467 Chinese Master's applicants in a high-stakes exam; demographic analyses show little difference, but this does not establish representativeness for all adults.
  • domain assumption Ratings from the two rater panels can be equated linearly via z-score transformation.
    Round 1 rated humans, Round 2 rated LLMs; the paper equates them using GPT-3.5 v1 responses rated in both rounds, assuming linear comparability of rater severity.
  • domain assumption LLM outputs collected under the stated temperatures and web portal defaults reflect model creativity rather than prompting artifacts.
    Minimal prompt engineering is a deliberate design, but for Claude, Qwen, and SparkDesk the exact system prompts and decoding settings are not fully specified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Language Models show both individual and collective creativity comparable to humans." pith.science (2026). https://pith.science/paper/VA36U5PF

@misc{pith2026241203151,
  author       = {Pith},
  title        = {Pith review of: Large Language Models show both individual and collective creativity comparable to humans},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VA36U5PF}},
  note         = {Machine review of arXiv:2412.03151}
}
read the original abstract

Artificial intelligence has, so far, largely automated routine tasks, but what does it mean for the future of work if Large Language Models (LLMs) show creativity comparable to humans? To measure the creativity of LLMs holistically, the current study uses 13 creative tasks spanning three domains. We benchmark the LLMs against individual humans, and also take a novel approach by comparing them to the collective creativity of groups of humans. We find that the best LLMs (Claude and GPT-4) rank in the 52nd percentile against humans, and overall LLMs excel in divergent thinking and problem solving but lag in creative writing. When questioned 10 times, an LLM's collective creativity is equivalent to 8-10 humans. When more responses are requested, two additional responses of LLMs equal one extra human. Ultimately, LLMs, when optimally applied, may compete with a small group of humans in the future of work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis

    cs.HC 2025-05 conditional novelty 6.0 of 10

    A meta-analysis of 28 studies finds no average creativity gap between GenAI and humans, a small boost when humans collaborate with GenAI, and a large drop in idea diversity in those collaborations.

Reference graph

Works this paper leans on

71 extracted references · 59 canonical work pages · cited by 1 Pith paper

  1. [1]

    Eloundou, S

    T. Eloundou, S. Manning, P. Mishkin, D. Rock, Gpts are gpts: An early look at the labor market impact potential of large language models. arXiv:2303.10130 [econ.GN] (2023)

  2. [2]

    Jobs of tomorrow: Large language models and jobs

    World Economic Forum, “Jobs of tomorrow: Large language models and jobs” (2023); https://www.weforum.org/publications/jobs-of-tomorrow-large-language-models-and- jobs

  3. [3]

    Bubeck, V

    S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y.T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M.T. Ribeiro, Y. Zhang, Sparks of Artificial General Intelligence: Early experiments with GPT -4. arXiv:2303.12712 [cs.CL] (2023)

  4. [4]

    The impact of artificial intelligence on the future of workforces in the European Union and the United States of America

    The White House, “The impact of artificial intelligence on the future of workforces in the European Union and the United States of America” (2022); https://www.whitehouse.gov/cea/written-materials/2022/12/05/the-impact-of-artificial- intelligence

  5. [5]

    The Impact of Large Language Multi-Modal Models on the Future of Job Market

    T. Singh, The impact of large language multi-modal models on the future of job market. arXiv:2304.06123 [cs.CY] (2023)

  6. [6]

    The origins of creativity

    E. Picciuto, P. Carruthers, “The origins of creativity” in The Philosophy of Creativity: New Essays, E. S. Paul, S. B. Kaufman, Eds. (Oxford University Press, 2014), pp. 199 - 223

  7. [7]

    B. A. Hennessey, T. M. Amabile, Creativity. Annual Review of Psychology 61, 569–598 (2010)

  8. [8]

    M. A. Runco, G. J. Jaeger, The standard definition of creativity. Creativity Research Journal 24, 92–96 (2012)

Show all 71 references
  1. [9]

    J. P. Guilford, A. T.Vaughan , Factors that aid and hinder creativity. Teachers College Record 63, 1-13 (1962)

  2. [10]

    E. P. Torrance, Torrance Test on Creative Thinking: Norms-Technical Manual Research Edition (Personnel Press, 1966)

  3. [11]

    Is creativity domain specific?

    J. Baer, “Is creativity domain specific?” in Cambridge Handbook of Creativity, J. C. Kaufman, R. J. Sternberg, Eds. (Cambridge University Press, 2010), pp. 321-341

  4. [12]

    R. J. Sternberg, Creativity or creativities? International Journal of Human -Computer Studies 63, 370-382 (2005)

  5. [13]

    Haase, P

    J. Haase, P. H. P. Hanel, Artificial muses: Generative artificial intelligence chatbots have risen to human-level creativity. Journal of Creativity 33, 100066 (2023)

  6. [14]

    Koivisto, S

    M. Koivisto, S. Grassini, Best humans still outperform artificial intelligence in a creative divergent thinking task. Sci. Rep. 13, 13601 (2023)

  7. [15]

    Stevenson, I

    C. Stevenson, I. Smal, M. Baas, R. Grasman, H. van der Maas, Putting GPT -3’s creativity to the (Alternative Uses) Test. arXiv:2206.08932 [cs.AI] (2022)

  8. [16]

    Cropley, Is artificial intelligence more creative than humans? ChatGPT and the Divergent Association Task

    D. Cropley, Is artificial intelligence more creative than humans? ChatGPT and the Divergent Association Task. Learning Letters 2, 13 (2023)

  9. [17]

    Girotra, L

    K. Girotra, L. Meincke, C. Terwiesch, K. T. Ulrich, Ideas are dimes a dozen: Large language models for idea generation in innovation. (2023); http://dx.doi.org/10.2139/ssrn.4526071. 26

  10. [18]

    E. E. Guzik, C. Byrge, C. Gilde, The originality of machines: AI takes the Torrance Test. Journal of Creativity 33, 100065 (2023)

  11. [19]

    M. I. Vicente -Yagüe-Jara, O. López -Martínez, V. Navarro -Navarro, F. Cuéllar - Santiago, Escritura, creatividad e inteligencia artificial. ChatGPT en el contexto universitario. Comunicar: Revista Científica de Comunicación y Educación 31, 47-57 (2023)

  12. [20]

    Chakrabarty, P

    T. Chakrabarty, P. Laban, D. Agarwal, S. Muresan, C. S. Wu, Art or artifice? Large language models and the false promise of creativity , in Proceedings of the CHI Conference on Human Factors in Computing Systems (Association for Computing Machinery, 2024), pp. 1-34

  13. [21]

    Prompting a large language model to generate diverse motivational messages: A comparison with human-written messages

    S. R. Cox, A. Abdul, W. T. Ooi, “Prompting a large language model to generate diverse motivational messages: A comparison with human-written messages” in Proceedings of the 11th International Conference on Human -Agent Interaction (Association for Computing Machinery, 2023), p...

  14. [22]

    Y. Tian, A. Ravichander, L. Qin, R. L. Bras, R. Marjieh, N. Peng, Y. Choi, T. L. Griffiths, F. Brahman, MacGyver: Are Large Language Models Creative Problem Solvers? In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Lingui...

  15. [23]

    Summers -Stay, C

    D. Summers -Stay, C. R. Voss, S. M. Lukin, Brainstorm, then select: a generative language model improves its creativity score, In AAAI-23 Workshop on Creative AI Across Modalities (Association for the Advancement of Artificial Intelligence, 2023)

  16. [24]

    Orwig, E

    W. Orwig, E. R. Edenbaum , J. D. Greene, D. L. Schacter, The language of creativity: Evidence from humans and large language models. J. Creat. Behav. 58, 128-136 (2024)

  17. [25]

    S., Dayan, P., & Stevenson, C, Characterising the creative process in humans and large language models

    Nath, S. S., Dayan, P., & Stevenson, C, Characterising the creative process in humans and large language models. arXiv:2405.00899 [cs.HC] (2024)

  18. [26]

    creativity

    H. Chen, N. Ding, Probing the “creativity” of large language models: Can models produce divergent semantic association? in Findings of the Association for Computational Linguistics: EMNLP 2023 (Association for Computational Linguistics, 2023), pp. 12881-12888

  19. [27]

    Castelo, Z

    N. Castelo, Z. Katona, P. Li, Miklos Sarvary, How AI outperforms humans at creative idea generation (2024); http://dx.doi.org/10.2139/ssrn.4751779

  20. [28]

    K. F. Hubert, K. N. Awa, D. L. Zabelina, The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks. Sci. Rep. 14, 3440 (2024)

  21. [29]

    Marco, J

    G. Marco, J. Gonzalo, L. Rello, Transformers can outperform humans in short creative writing tasks (2023); http://dx.doi.org/10.2139/ssrn.4673692

  22. [30]

    Grassini, M

    S. Grassini, M. Koivisto, Understanding how personality traits, experiences, and attitudes shape negative bias toward AI-generated artworks. Sci. Rep. 14, 4113 (2024)

  23. [31]

    Bellemare-Pepin, F

    A. Bellemare-Pepin, F. Lespinasse, P. Thölke, Y. Harel, K. Mathewson, J. A. Olson, Y. Bengio, K. Jerbi, Divergent creativity in humans and large language models. arXiv.2405.13012 [cs.CL] (2024)

  24. [32]

    Gómez -Rodríguez, P

    C. Gómez -Rodríguez, P. Williams, A confederacy of models: A comprehensive evaluation of LLMs on creative writing, In Findings of the Association for Computational Linguistics: EMNLP 2023 (Association for Computational Linguistics, 2023), pp. 14504-14528

  25. [33]

    C. Si, D. Yang, T. Hashimoto, Can LLMs generate novel research ieas? A large -scale human study with 100+ NLP researchers. arXiv:2409.04109 [cs.CL] (2024). 27

  26. [34]

    C. Deng, Y. Zhao, X. Tang, M. Gerstein, A. Cohan, Investigating data contamination in modern benchmarks for large language models. arXiv:2311.09783 [cs.CL] (2023)

  27. [35]

    Balloccu, P

    S. Balloccu, P. Schmidtová, M. Lango, O. Dušek, Leak, cheat, repeat: Data contamination and evaluation malpractices in closed -source LLMs. arXiv:2402.03927 [cs.CL] (2024)

  28. [36]

    C. Li, J. Flanigan, Task contamination: Language models may not be few-shot anymore, In Proceedings of the AAAI Conference on Artificial Intelligence (Association for the Advancement of Artificial Intelligence, 2024), pp. 18471-18480

  29. [37]

    T. M. Amabile, Social psychology of creativity: A consensual assessment technique. Journal of Personality and Social Psychology 43, 997–1013 (1982)

  30. [38]

    I. J. Hoever, D. van Knippenberg, W. P. van Ginkel, H. G. Barkema, Fostering team creativity: Perspective taking as key to unlocking diversity’s potential. J. Appl. Psychol. 97, 982 (2012)

  31. [39]

    O. A. Acar, A. uncdogan, D. van Knippenberg, K. R. Lakhani, Collective creativity and innovation: An interdisciplinary review, integration, and research agenda. J. Manag. 50, 2119-2151 (2024)

  32. [40]

    M. E. Manske, G. A. Davis, Effects of simple instructional biases upon performance in the unusual uses test. The Journal of General Psychology 79, 25-33 (1968)

  33. [41]

    E. F. Rietzschel, B. A. Nijstad, W. Stroebe, in The Oxford handbook of group creativity and innovation, P. B. Paulus & B. A. Nijstad, Ed. (Oxford University Press, 2019), pp. 179-197

  34. [42]

    A. R. Doshi, O. P. Hauser, Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Adv. 10, eadn5290 (2024)

  35. [43]

    Homogenizing effect of large language model on creativity: An empirical comparison of human and ChatGPT writing,

    K. Moon, “Homogenizing effect of large language model on creativity: An empirical comparison of human and ChatGPT writing,” thesis, Washington, DC (2024)

  36. [44]

    B. R. Anderson, Jash Hemant Shah, M. Kreminski, Homogenization effects of large language models on human creative ideation, in Proceedings of the 16th Conference on Creativity & Cognition (Association for Computing Machinery, 2024), pp. 413-425

  37. [45]

    A. F. Osborn, Applied Imagination: Principles and Procedures of Creative Thinking (Scribner, 1953)

  38. [46]

    Ananiadou, M

    K. Ananiadou, M. Claro, 21st century skills and Competences for new millennium learners in OECD Countries. (OECD, 2009)

  39. [47]

    M. A. Runco, AI can only produce artificial creativity. Journal of Creativity 33, 100063 (2023)

  40. [48]

    S. Noy, W. Zhang, Experimental evidence on the productivity effects of generative artificial intelligence. Science 381, 187-192 (2023)

  41. [49]

    B. C. Lee, J. J. Chung, An empirical investigation of the impact of ChatGPT on creativity. Nat. Hum. Behav. 8, 1906-1914 (2024)

  42. [50]

    Why Johnny can’t prompt: how non -AI experts try (and fail) to design LLM prompts

    J. D. Zamfirescu -Pereira, R. Y. Wong, B. Hartmann, Q. Yang, “Why Johnny can’t prompt: how non -AI experts try (and fail) to design LLM prompts” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Association for Computing Machinery, 2023), pp. 1-21

  43. [51]

    H. Nori, Y. T. Lee, S. Zhang, D. Carignan, R. Edgar, N. Fusi, N. King, J. Larson, Y. Li, W. Liu, R. Luo, S. M. Mckinney, R. O. Ness, H. Poon, T. Qin, N. Usuyama, C. White, E. Horvitz, Can generalist foundation models outcompete special-purpose tuning? Case study in Medicine. a...

  44. [52]

    AI prompt engineering isn’t the future

    O. A. Acar, “AI prompt engineering isn’t the future” (Harvard Business Review, 2023); https://hbr.org/2023/06/ai-prompt-engineering-isnt-the-future

  45. [53]

    J. Zhuo, S. Zhang, X. Fang, H. Duan, D. Lin, K. Chen, ProSA: Assessing and understanding the prompt sensitivity of LLMs. arXiv:2410.12405 [cs.CL] (2024)

  46. [54]

    D. W. Taylor, P. C. Berry, C. H. Block, Does group participation when using brainstorming facilitate or inhibit creative thinking? Administrative science quarterly 3, 23-47 (1958)

  47. [55]

    Mullen, C

    B. Mullen, C. Johnson, E. Salas, Productivity loss in brainstorming groups: A meta - analytic integration. Basic and applied social psychology 12, 3-23 (1991)

  48. [56]

    Faure, Beyond brainstorming: Effects of different group procedures on selection of ideas and satisfaction with the process

    C. Faure, Beyond brainstorming: Effects of different group procedures on selection of ideas and satisfaction with the process. J. Creat. Behav. 38, 13-34 (2004)

  49. [57]

    The potentially large effects of artificial intelligence on economic growth

    J. Hatzius, J. Briggs, D. Kodnani, G. Pierdomenico, “The potentially large effects of artificial intelligence on economic growth” (Goldman Sachs, 2023); https://static.poder360.com.br/2023/03/Global-Economics-Analyst_-The-Potentially- Large-Effects-of-Artificial-Intelligence-o...

  50. [58]

    P Partia l η² Task 1 M 4.417 4.883 4.687 4.83 0.896 0.4 44 0.012 Creativity Mean SD 1.112 1.479 1.241 1.273 n 25 74 52 79 Task 1 M 5.991 6.602 6.482 6.559 0.684 0.5 63 0.009 Creativity Max SD 1.736 1.932 2.062 1.856 n 25 74 52 79 Statis tics Arts and humanities STEM Programme ...

  51. [59]

    P Partia l η² Task 2 M 5.02 4.94 5.041 5.13 0.267 0.8 49 0.003 Creativity Mean SD 0.9 1.089 1.001 1.609 n 44 57 51 83 Task 2 M 6.643 6.556 6.66 6.803 0.184 0.9 07 0.002 Creativity Max SD 1.834 1.975 1.662 2.288 n 44 57 51 83 Scientific Problem Solving Task 1 Creativity Statis ...

  52. [60]

    P Partia l η² M 1.895 1.942 1.992 1.952 0.535 0.6 59 0.007 SD 0.369 0.359 0.355 0.399 n 44 58 51 84 Task 2 Creativity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,2

  53. [61]

    P Partia l η² M 2.309 2.345 2.392 2.376 0.107 0.9 56 0.001 SD 0.648 0.893 0.695 0.832 n 44 58 50 84 Task 3 Creativity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,2

  54. [62]

    P Partia l η² M 2.741 2.634 2.706 2.557 0.582 0.6 27 0.007 SD 0.869 0.781 0.787 0.906 n 44 58 51 84 75 Social Problem Solving Task 1 Flexibility Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,2

  55. [63]

    P Partia l η² M 2.28 2.257 2.346 2.333 0.333 0.8 02 0.004 SD 0.614 0.598 0.59 0.55 n 25 74 52 78 Task 2 Creativity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,2

  56. [64]

    P Partia l η² M 1.992 2.081 2.077 2.218 0.755 0.5 2 0.01 SD 0.692 0.794 0.703 0.828 n 25 74 52 78 Task 3 Creativity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,2

  57. [65]

    P Partia l η² M 2.904 2.676 2.529 2.747 1.314 0.2 71 0.017 SD 0.837 0.813 0.891 0.806 n 25 74 51 76 Keyword-prompted Writing Creativity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,2

  58. [66]

    P Partia l η² M 2.467 2.361 2.447 2.432 0.24 0.8 68 0.003 SD 0.789 0.669 0.644 0.69 n 24 70 51 76 Diversity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,1

  59. [71]

    P Partia l η² M 2.294 2.282 2.358 2.249 0.62 0.6 02 0.004 SD 0.572 0.595 0.673 0.601 n 62 120 96 140 77 Fig. S1. Densities of the creativity ratings for LLM responses in the divergent thinking tasks. 78 Fig. S2. Densities of the creativity ratings for LLM responses in the prob...

  60. [74]

    P Partia l η² M 3.27 3.224 3.346 3.356 0.449 0.7 18 0.008 SD 0.757 0.728 0.687 0.597 n 20 58 41 59 Emoji-prompted Writing Task 1 Diversity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,1

  61. [77]

    P Partia l η² M 2.565 2.7 2.62 2.531 0.445 0.7 21 0.007 SD 0.799 0.744 0.845 0.751 n 34 48 41 58 Task 3 Creativity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,1

  62. [85]

    P Partia l η² 76 M 2.572 2.612 2.537 2.516 0.197 0.8 98 0.003 SD 0.725 0.682 0.754 0.654 n 36 51 38 64 Creative Advert Writing Creativity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,4

  63. [97]

    P Partia l η² M 3.04 3.109 3.173 3.031 0.421 0.7 38 0.006 SD 0.702 0.724 0.721 0.748 n 40 44 45 72 Task 2 Creativity Statis tics Arts and humanities STEM Programme subject Other social science subjects F(3,1

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.