REVIEW 4 major objections 5 minor 2 cited by
Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read GenAI matches humans on creativity, aids teams, cuts diversity
desk verdict First meta-analysis on GenAI and creativity: transparent and useful, but the headline null effect rests on a pseudo-replication problem and the publication-bias correction reverses it; still deserves peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying machinery is a random-effects meta-analysis of standardized effect sizes using Hedges' $g$. Effect sizes are computed from each eligible experiment either directly or by converting $t$-values, $F$-values, means, standard deviations, and $\beta$ coefficients with standard formulas, then pooled with a random-effects estimator designed for heterogeneous study sets. Heterogeneity is quantified with $Q$, $I^2$, and $\tau^2$, and moderator structure is probed with subgroup analyses, meta-regressions, leave-one-out sensitivity checks, regression tests for funnel asymmetry, and trim-and-fill bias adjustments. The design organizes the 127 effect sizes into three comparison families—human versus GenAI (RQ1), human plus GenAI versus human (RQ2a), and idea diversity in collaboration (RQ2b)—which is what lets a single study base produce three distinct answers.
What would settle it
Recompute the three pooled effects separately for each creativity-measure family: expert ratings, self-reports, semantic-distance scores, and rule-based scores. If the subgroup estimates differ enough to flip the sign or significance of any headline $g$ (for example, if the $g = 0.27$ collaboration gain disappears under expert ratings), then the different measures are not measuring one construct and the pooled numbers are not a single creativity effect.
Extended reading notes
Core claim
The paper's central claim is that the current empirical literature supports a three-part description of generative AI and creativity. First, in direct comparisons, GenAI-generated ideas are rated about as creative as human ideas (pooled Hedges' $g = -0.05$; 95% CI $[-0.26, 0.16]$), so average parity, not superhuman creativity, is the norm. Second, human-GenAI collaboration raises creative performance relative to unaided humans ($g = 0.27$; 95% CI $[0.02, 0.53]$), a small but robust effect that survives leave-one-out deletion and bias adjustments. Third, the same collaboration lowers the diversity of ideas ($g = -0.86$; 95% CI $[-1.33, -0.40]$), an effect the authors interpret as a homogenization risk. The authors also show that these averages hide real moderators: GPT-4 outperforms humans when generating alone, laypeople gain most from collaboration, and creative-writing tasks favor humans while divergent-thinking batteries favor GenAI. These results are the basis for the paper's conclusion that GenAI should be treated as an augmentative support, not a creative replacement.
Load-bearing premise
The analysis assumes that the different creativity measurements in the 28 studies—expert ratings, self-assessments, novelty scales, semantic distance, cosine similarity, and rule-based flexibility scores—are close enough to the same underlying construct that averaging their effect sizes into one pooled $g$ is meaningful.
Editorial extensions
If this is right
- The pooled parity estimate ($g = -0.05$, CI $[-0.26, 0.16]$) means claims that GenAI is broadly more creative than humans are not supported by current evidence; any advantage is model-specific, such as GPT-4's $g = 0.50$.
- The significant collaboration gain ($g = 0.27$) supports deploying GenAI as an ideation aid; the effect survives leave-one-out deletion and trim-and-fill adjustment.
- The large diversity drop ($g = -0.86$, CI $[-1.33, -0.40]$) implies that using GenAI in creative work can homogenize outputs, a cost that should be weighed alongside the performance gain.
- Heterogeneity results imply that creativity is not a single skill: GenAI does better on divergent-thinking batteries such as alternate uses and consequences tasks but worse on creative writing, and laypeople benefit more than academics.
- The diversity estimate rests on six observations from four studies; its direction is robust to removing any single study, but its magnitude is not precise.
Reading between the lines
- A joint interpretation the paper stops short of: the same GenAI that raises average creativity in collaboration could be doing so by anchoring people to a narrower idea space, so the performance gain and the diversity loss may be two sides of one mechanism rather than independent effects.
- A testable extension: measure whether the collaboration gain shrinks as participants' domain expertise rises, ideally holding the GenAI model fixed; the paper's moderator pattern (laypeople gain, academics do not) predicts exactly that gradient.
- A caution about generalizing: the diversity estimate is built on six observations from four studies, all text-based; extending the same meta-analytic comparison to image, music, and organizational ideation would show whether the homogenization effect is a general property of GenAI or a text-task property.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a systematic literature review and meta-analysis (PRISMA 2020) of 28 studies examining whether GenAI matches human creative performance (RQ1), whether humans collaborating with GenAI outperform humans working alone on creative tasks (RQ2a), and whether such collaboration reduces idea diversity (RQ2b). Using random-effects meta-analysis with Hedges' g, the authors report no significant difference between GenAI and human creativity (g = -0.05), a significant positive collaboration effect (g = 0.27), and a significant negative diversity effect (g = -0.86). They also report subgroup analyses by model, participant background, and task type, along with robustness checks including leave-one-out, Egger's test, trim-and-fill, and influence diagnostics. Code and data are provided in a public repository.
Significance. If the three headline estimates withstand reanalysis, the paper would provide a useful quantitative synthesis of a fast-moving literature, with practical implications for human-AI collaboration. The authors deserve credit for making code and data publicly available, for following a structured PRISMA workflow, and for reporting multiple heterogeneity analyses. However, the current statistical implementation has two load-bearing weaknesses: a unit-of-analysis problem from treating 52 observations of one study as independent in RQ1, and a publication-bias analysis that reverses the RQ1 null result yet is described as raising no substantial concerns. These issues must be resolved before the central claims can be accepted.
major comments (4)
- [Section 3.2, Table 2] In the RQ1 analysis, HAI15 (Sun et al.) contributes 52 of the 100 observations, and the statement in Section 3.2 that 'each study contributes less than 1.1% weight' refers to per-observation weights. Because these 52 observations share one protocol, one participant pool, and one modeling pipeline, treating them as independent units inflates precision and can dominate the pooled g = -0.048. The leave-one-out analysis removes single observations and therefore cannot detect dependence on the entire HAI15 study. Please re-estimate RQ1 with study-level aggregation or a three-level/multivariate meta-analysis with cluster-robust variance estimation, and report both observation-level and study-level estimates.
- [Section 3.2.2, Section 2.3] Egger's test in Section 3.2.2 is significant (z = -3.363, p = 0.001) and the trim-and-fill procedure imputes 24 studies, shifting the pooled estimate to g = 0.364 with 95% CI [0.131, 0.596], thereby reversing the direction of the RQ1 effect. This contradicts the statement in Section 2.3 that the bias analyses 'revealed no substantial concerns regarding bias.' The paper must reconcile this contradiction, present the bias-corrected estimate prominently, and temper the conclusion that GenAI and humans show 'parity' in creative performance.
- [Section 2.2, Section 3.2] The meta-analysis pools effect sizes across different operational definitions of creativity: creativity scales, originality scales, novelty scales, semantic distance, cosine similarity, and flexibility scores, rated by self, laypeople, experts, rule-based, or AI evaluators. With I2 ≈ 99% for RQ1, the pooled estimate is not interpretable as measuring a single latent construct unless the authors demonstrate that measurement type does not moderate the effects. The manuscript states that a subgroup analysis by construct found no substantial heterogeneity but omits it from the main text; please include these analyses or a measurement-type moderator in the paper or appendix.
- [Section 3.4] The RQ2b result (g = -0.86) is based on only 6 observations from 4 studies. Although the leave-one-out analysis is reassuring, the evidence base is small and the confidence interval is wide. The paper should more explicitly frame this finding as preliminary and should avoid definitive language such as 'consistent and statistically meaningful reduction' without acknowledging the small number of contributing studies.
minor comments (5)
- [Table 3] The table header contains the typo 'anaysis'; it should read 'analysis.'
- [Section 3.1] The paper reports 28 included studies but states that RQ1 uses 21 studies, RQ2a 12, and RQ2b 4; since some studies appear in multiple analyses, please clarify the overlap and the relation between the 28 studies and the 127 observation-level effect sizes.
- [Figure 1 and Section 2.1] The PRISMA flowchart and the eligibility criteria list different exclusion categories (e.g., 'insufficient study design' vs. 'insufficient statistics' vs. 'insufficient data reporting'), and the numbering in the text is inconsistent. Please align the flowchart, the numbered criteria, and the descriptions in the text.
- [Section 2.1] The screening and eligibility decisions were performed by one person (the first author). This is a potential source of bias and should be acknowledged as a limitation; PRISMA typically recommends dual screening.
- [Section 3.2 and Section 3.3] The phrases 'each study contributes less than 1.1% weight' and 'no single study contributes more than 5.1% weight' refer to per-observation weights, not study-level weights. Rephrase to avoid misleading readers into thinking each study is a single unit.
Circularity Check
No significant circularity: the meta-analytic estimates are standard aggregations of externally reported effect sizes.
full rationale
The paper is a systematic review and meta-analysis. Its headline results (g = -0.05, g = 0.27, g = -0.86) are weighted averages of effect sizes computed from statistics reported in 28 external studies, using standard formulas (Cohen's d and Hedges' g correction) and a DerSimonian-Laird random-effects model. No parameter is fitted to a subset of data and then relabeled as a prediction: the pooled estimates are descriptive summaries of the included effect sizes, not derived quantities that presuppose the conclusions. The only self-citations (e.g., Feuerriegel et al. for definitions of GenAI and NLP-based text analysis) appear as background references and are not load-bearing for the meta-analytic estimates. No uniqueness theorem, ansatz, or prior work by the same authors is invoked to force the choice of model or to forbid alternatives. The paper's robustness checks are standard meta-analytic diagnostics. Concern about non-independence caused by counting 52 observations from one study (HAI15) as separate units is a serious statistical validity issue, but it is not circularity: the estimates still reduce to the extracted input data, which is the normal and transparent behavior of a meta-analysis. The derivation chain is therefore self-contained with respect to the external evidence it aggregates.
Assumptions & free parameters
assumptions (4)
- standard math Random-effects meta-analysis with the DerSimonian-Laird estimator adequately models the between-study heterogeneity.
- domain assumption The creativity constructs across studies (creativity scale, originality, novelty, semantic distance, flexibility, cosine similarity) are commensurable enough to pool into one Hedges' g.
- domain assumption The three-database, title-based search plus manual additions captures the relevant population of experiments.
- domain assumption Egger's test and trim-and-fill are valid diagnostics for publication bias in this literature.
Cite this review
Pith. "Pith review of Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis." pith.science (2026). https://pith.science/paper/EZMAHNPG
@misc{pith2026250517241,
author = {Pith},
title = {Pith review of: Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/EZMAHNPG}},
note = {Machine review of arXiv:2505.17241}
}
read the original abstract
Generative artificial intelligence (GenAI) is increasingly used to support a wide range of human tasks, yet empirical evidence on its effect on creativity remains scattered. Can GenAI generate ideas that are creative? To what extent can it support humans in generating ideas that are both creative and diverse? In this study, we conduct a meta-analysis to evaluate the effect of GenAI on the performance in creative tasks. For this, we first perform a systematic literature search, based on which we identify n = 28 relevant studies (m = 8214 participants) for inclusion in our meta-analysis. We then compute standardized effect sizes based on Hedges' g. We compare different outcomes: (i) how creative GenAI is; (ii) how creative humans augmented by GenAI are; and (iii) the diversity of ideas by humans augmented by GenAI. Our results show no significant difference in creative performance between GenAI and humans (g = -0.05), while humans collaborating with GenAI significantly outperform those working without assistance (g = 0.27). However, GenAI has a significant negative effect on the diversity of ideas for such collaborations between humans and GenAI (g = -0.86). We further analyze heterogeneity across different GenAI models (e.g., GPT-3.5, GPT-4), different tasks (e.g., creative writing, ideation, divergent thinking), and different participant populations (e.g., laypeople, business, academia). Overall, our results position GenAI as an augmentative tool that can support, rather than replace, human creativity-particularly in tasks benefiting from ideation support.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
Faster, Higher, Stronger? The Impact of GenAI on Knowledge Work Productivity - Evidence from the Field
GenAI support cut task time across all three knowledge-work tasks but improved output quality only for packaging and creation, while lowering quality and widening quality gaps in knowledge acquisition.
-
AutoSynthesis: An agentic system for automated meta-analysis
An LLM agent pipeline can perform an end-to-end meta-analysis, but its claim of close agreement with expert meta-analyses rests on a single, co-authored benchmark and unreleased code.
Reference graph
Works this paper leans on
-
[1]
Jaan Aru. 2025. Artificial intelligence and the internal processes of creativity.The Journal of Creative Behavior(2025). doi:10.1002/jocb.1530
-
[2]
Joshua Ashkinaze, Julia Mendelsohn, Li Qiwei, Ceren Budak, and Eric Gilbert. 2024. How AI ideas affect the creativity, diversity, and evolution of human ideas: Evidence from a large, dynamic experiment. https://arxiv.org/abs/2401.13481
arXiv 2024
-
[3]
Michael Borenstein, Larry V. Hedges, Julian P. T. Higgins, and Hannah R. Rothstein. 2021.Introduction to meta-analysis. John Wiley & Sons, Ltd, West Sussex, UK. doi:10.1002/9780470743386
-
[4]
Kevin J. Boudreau and Karim R. Lakhani. 2013. Using the crowd as an innovation partner.Harvard Business Review91, 4 (2013), 60–69
work page 2013
-
[5]
Lane, Miaomiao Zhang, Vladimir Jacimovic, and Karim R
Léonard Boussioux, Jacqueline N. Lane, Miaomiao Zhang, Vladimir Jacimovic, and Karim R. Lakhani. 2024. The crowdless future? Generative AI and creative problem-solving.Organization Science35, 5 (2024), 1589–1607. doi:10. 1287/orsc.2023.18430
arXiv 2024
-
[6]
Erik Brynjolfsson, Danielle Li, and Lindsey Raymond. 2025. Generative AI at work.The Quarterly Journal of Economics 140, 2 (2025), 889–942. doi:10.1093/qje/qjae044
-
[7]
Noah Castelo, Zsolt Katona, Peiyao Li, and Miklos Sarvary. 2024. How AI outperforms humans at creative idea generation.A vailable at SSRN 4751779(2024). doi:10.2139/ssrn.4751779
-
[8]
Gary Charness and Daniela Grieco. 2024. Creativity and AI.A vailable at SSRN 4686415(2024). doi:10.2139/ssrn.4686415
Show all 76 references
-
[9]
Zenan Chen and Jason Chan. 2024. Large language model in creative work: The role of collaboration modality and user expertise.Management Science70, 12 (2024), 9101–9117. doi:10.1287/mnsc.2023.03014
2024
-
[10]
2013.Statistical power analysis for the behavioral sciences(2nd ed.)
Jacob Cohen. 2013.Statistical power analysis for the behavioral sciences(2nd ed.). Routledge, New York. doi:10.4324/ 9780203771587
2013
-
[11]
Zheyuan Kevin Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, and Tobias Salz. 2024. The effects of generative ai on high skilled work: Evidence from three field experiments with software developers.A vailable at SSRN 4945566 (2024). doi:10.2139/ssrn.4945566
2024 doi
-
[12]
María-Isabel de Vicente-Yagüe-Jara, Olivia López-Martínez, Verónica Navarro-Navarro, and Francisco Cuéllar-Santiago
-
[13]
Deeks, Julian P
Jonathan J. Deeks, Julian P. T. Higgins, Douglas G. Altman, and on behalf of the Cochrane Statistical Methods Group
-
[14]
Rebecca DerSimonian and Nan Laird. 1986. Meta-analysis in clinical trials.Controlled Clinical Trials7, 3 (1986), 177–188. doi:10.1016/0197-2456(86)90046-2
1986 doi
-
[15]
Doshi, Sen Chai, and Matthias Troebinger
Anil R. Doshi, Sen Chai, and Matthias Troebinger. 2024. How experience moderates the impact of generative AI ideas on the research process.A vailable at SSRN 5013086(2024). doi:10.2139/ssrn.5013086
2024 doi
-
[16]
Doshi and Oliver P
Anil R. Doshi and Oliver P. Hauser. 2024. Generative AI enhances individual creativity but reduces the collective diversity of novel content.Science Advances10, 28 (2024), eadn5290. doi:10.1126/sciadv.adn5290
2024 doi
-
[18]
Sue Duval and Richard Tweedie. 2000. Trim and fill: a simple funnel-plot–based method of testing and adjusting for publication bias in meta-analysis.Biometrics56, 2 (2000), 455–463. doi:10.1111/j.0006-341x.2000.00455.x
2000
-
[19]
Matthias Egger, George Davey Smith, Martin Schneider, and Christoph Minder. 1997. Bias in meta-analysis detected by a simple, graphical test.bmj315, 7109 (1997), 629–634. doi:10.1136/bmj.315.7109.629
1997 doi
-
[20]
Anja Eisenreich, Julian Just, Daniela Gimenez-Jimenez, and Johann Füller. 2024. Revolution or inflated expectations? Exploring the impact of generative AI on ideation in a practical sustainability context.Technovation138 (2024), 103123. doi:10.1016/j.technovation.2024.103123
2024
-
[21]
Stefan Feuerriegel, Jochen Hartmann, Christian Janiesch, and Patrick Zschech. 2024. Generative AI.Business & Information Systems Engineering66, 1 (2024), 111–126. doi:10.1007/s12599-023-00834-7
2024 doi
-
[22]
Robertson, Steve Rathje, Jochen Hartmann, Saif M
Stefan Feuerriegel, Abdurahman Maarouf, Dominik Bär, Dominique Geissler, Jonas Schweisthal, Nicolas Pröllochs, Claire E. Robertson, Steve Rathje, Jochen Hartmann, Saif M. Mohammad, Oded Netzer, Alexandra A. Siegel, Barbara Plank, and Jay J. van Bavel. 2025. Using natural langu...
2025 doi
-
[23]
Carlos Gómez-Rodríguez and Paul Williams. 2023. A confederacy of models: A comprehensive evaluation of LLMs on creative writing. https://arxiv.org/abs/2310.08433v1
2023 arXiv
-
[24]
Simone Grassini and Mika Koivisto. 2025. Artificial creativity? Evaluating AI against human performance in creative interpretation of visual stimuli.International Journal of Human–Computer Interaction41, 7 (2025), 4037–4048. doi:10. 1080/10447318.2024.2345430 Generative AI and...
2025
-
[25]
Forward flow
Kurt Gray, Stephen Anderson, Eric Evan Chen, John Michael Kelly, Michael S. Christian, John Patrick, Laura Huang, Yoed N. Kenett, and Kevin Lewis. 2019. “Forward flow”: A new measure to quantify free thought and predict creativity. American Psychologist74, 5 (2019), 539–554. d...
2019 doi
-
[26]
Matthew Grimes, Georg von Krogh, Stefan Feuerriegel, Floor Rink, and Marc Gruber. 2023. From Scarcity to Abundance: Scholars and Scholarship in an Age of Generative Artificial Intelligence.Academy of Management Journal66, 6 (2023), 1617–1624. doi:10.5465/amj.2023.4006
2023
-
[27]
Christensen, Philip R
Joy Paul Guilford, Paul R. Christensen, Philip R. Merrifield, and Robert C. Wilson. 1978. Alternate uses. (1978). doi:10.1037/t06443-000
1978 doi
-
[28]
Steffen Herbold, Annette Hautli-Janisz, Ute Heuer, Zlata Kikteva, and Alexander Trautsch. 2023. A large-scale comparison of human-written versus ChatGPT-generated essays.Scientific Reports13, 1 (2023), 18617. doi:10.1038/ s41598-023-45644-9
2023
-
[29]
Hubert, Kim N
Kent F. Hubert, Kim N. Awa, and Darya L. Zabelina. 2024. The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks.Scientific Reports14, 1 (2024), 3440. doi:10.1038/s41598-024-53303-w
2024 doi
-
[30]
Mete Ismayilzada, Debjit Paul, Antoine Bosselut, and Lonneke van der Plas. 2024. Creativity in AI: Progresses and challenges. https://arxiv.org/abs/2410.17218v4
2024 arXiv
-
[31]
Mete Ismayilzada, Claire Stevenson, and Lonneke van der Plas. 2024. Evaluating creative short story generation in humans and large language models. https://arxiv.org/abs/2411.02316v5
2024 arXiv
-
[32]
Dan Jackson and Jack Bowden. 2016. Confidence intervals for the between-study variance in random-effects meta- analysis using generalised heterogeneity statistics: should we use unequal tails?BMC Medical Research Methodology 16, 1 (2016), 118. doi:10.1186/s12874-016-0219-y
2016 doi
-
[33]
Olson, Johnny Nahas, Denis Chmoulevitch, Simon J
Jay A. Olson, Johnny Nahas, Denis Chmoulevitch, Simon J. Cropper, and Margaret E. Webb. 2021. Naming unrelated words predicts creativity.Proceedings of the National Academy of Sciences118, 25 (2021), e2022340118. doi:10.1073/ pnas.2022340118
2021
-
[34]
Jennifer Haase and Paul H.P. Hanel. 2023. Artificial muses: Generative artificial intelligence chatbots have risen to human-level creativity.Journal of Creativity33, 3 (2023), 100066. doi:10.1016/j.yjoc.2023.100066
2023
-
[35]
Plucker, Ronald A
Jonathan A. Plucker, Ronald A. Beghetto, and Gayle T. Dow and. 2004. Why isn’t creativity more important to educational psychologists? Potentials, pitfalls, and future directions in creativity research.Educational Psychologist39, 2 (2004), 83–96. doi:10.1207/s15326985ep3902{_}1
2004 doi
-
[36]
Anuj Kapoor and Madhav Kumar. 2025. Frontiers: Generative AI and personalized video advertisements.Marketing Science(2025). doi:10.1287/mksc.2023.0494
2025
-
[37]
Mika Koivisto and Simone Grassini. 2023. Best humans still outperform artificial intelligence in a creative divergent thinking task.Scientific Reports13, 1 (2023), 13601. doi:10.1038/s41598-023-40858-3
2023 doi
-
[38]
Larry V. Hedges. 1981. Distribution theory for glass’s estimator of effect size and related estimators.Journal of Educational Statistics6, 2 (1981), 107–128. doi:10.3102/10769986006002107
1981 doi
-
[39]
Byung Cheol Lee and Jaeyeon Chung. 2024. An empirical investigation of the impact of ChatGPT on creativity.Nature Human Behaviour8, 10 (2024), 1906–1914. doi:10.1038/s41562-024-01953-1
2024 doi
-
[40]
Runco and Garrett J
Mark A. Runco and Garrett J. Jaeger and. 2012. The standard definition of creativity.Creativity Research Journal24, 1 (2012), 92–96. doi:10.1080/10400419.2012.650092
2012
-
[41]
Jack McGuire, David de Cremer, and Tim van de Cruys. 2024. Establishing the importance of co-creation and self-efficacy in creative collaboration with artificial intelligence.Scientific Reports14, 1 (2024), 18525. doi:10.1038/s41598-024-69423-2
2024 doi
-
[42]
Brewis, Fortune Nwaiwu, Deshan Sumanathilaka, Fernando Alva-Manchego, and Joanna Demaree-Cotton
Peidong Mei, Deborah N. Brewis, Fortune Nwaiwu, Deshan Sumanathilaka, Fernando Alva-Manchego, and Joanna Demaree-Cotton. 2025. If ChatGPT can do it, where is my creativity? Generative AI boosts performance but diminishes experience in creative writing.Computers in Human Behavi...
2025
-
[43]
Noah Bohren, Rustamdjan Hakimov, and Rafael Lalive. 2024. Creative and strategic capabilities of Generative AI: Evidence from large-scale experiments: IZA Discussion Papers. https://hdl.handle.net/10419/305744
2024
-
[44]
Edenbaum, Joshua D
William Orwig, Emma R. Edenbaum, Joshua D. Greene, and Daniel L. Schacter. 2024. The language of creativity: Evidence from humans and large language models.The Journal of Creative Behavior58, 1 (2024), 128–136. doi:10. 1002/jocb.636
2024
-
[45]
Page, Joanne E
Matthew J. Page, Joanne E. McKenzie, Patrick M. Bossuyt, Isabelle Boutron, Tammy C. Hoffmann, Cynthia D. Mulrow, Larissa Shamseer, Jennifer M. Tetzlaff, Elie A. Akl, Sally E. Brennan, Roger Chou, Julie Glanville, Jeremy M. Grimshaw, Asbjørn Hróbjartsson, Manoj M. Lalu, Tianjin...
-
[46]
Page, Joanne E
Matthew J. Page, Joanne E. McKenzie, Patrick M. Bossuyt, Isabelle Boutron, Tammy C. Hoffmann, Cynthia D. Mulrow, Larissa Shamseer, Jennifer M. Tetzlaff, Elie A. Akl, Sally E. Brennan, Roger Chou, Julie Glanville, Jeremy M. Grimshaw, 20 Niklas Holzner, Sebastian Maier, and Stef...
-
[47]
Satu Parjanen. 2012. Experiencing creativity in the organization: From individual creativity to collective creativity. Interdisciplinary Journal of Information, Knowledge & Management7 (2012), 109–128. doi:10.28945/1580
2012 doi
-
[48]
Peterson and Steven P
Robert A. Peterson and Steven P. Brown. 2005. On the use of beta coefficients in meta-analysis.Journal of Applied Psychology90, 1 (2005), 175–181. doi:10.1037/0021-9010.90.1.175
2005 doi
-
[49]
Rosnow and Robert Rosenthal
Ralph L. Rosnow and Robert Rosenthal. 1996. Computing contrasts, effect sizes, and counternulls on other people’s published data: General procedures for research consumers.Psychological Methods1, 4 (1996), 331. doi:10.1037/1082- 989X.1.4.331
1996 doi
- [50]
-
[51]
Solve Sæbø and Helge Brovold. 2024. On the stochastics of human and artificial creativity. https://arxiv.org/abs/2403. 06996v1
2024
-
[52]
Max Schemmer, Patrick Hemmer, Maximilian Nitsche, Niklas Kühl, and Michael Vössing. 2022. A meta-analysis of the utility of explainable artificial intelligence in human-AI decision-making. InAAAI/ACM Conference on AI, Ethics, and Society (AIES). doi:10.1145/3514094.3534128
2022
-
[53]
Max Schemmer, Niklas Kuehl, Carina Benz, Andrea Bartos, and Gerhard Satzger. 2023. Appropriate reliance on AI advice: Conceptualization and the effect of explanations. InInternational Conference on Intelligent User Interfaces (IUI). doi:10.1145/3581641.3584066
2023
-
[54]
Suebsarn Ruksakulpiwat, Lalipat Phianhasin, Chitchanok Benjasirisan, Kedong Ding, Anuoluwapo Ajibade, Ayanesh Kumar, and Cassie Stewart. 2024. Assessing the efficacy of ChatGPT versus human researchers in identifying relevant studies on mHealth interventions for improving medi...
2024 doi
-
[55]
Yu Song, Longchao Huang, Lanqin Zheng, Mengya Fan, and Zehao Liu. 2025. Interactions with generative AI chatbots: unveiling dialogic dynamics, students’ perceptions, and practical competencies in creative problem-solving. International Journal of Educational Technology in High...
2025 doi
-
[56]
Liz Sanders SonicRim. 2001. Collective creativity.Design6, 3 (2001), 1–6. http://www.echo.iat.sfu.ca/library/sanders_ 01_collective_creativity.pdf
2001
-
[57]
Jonathan A. C. Sterne, Jelena Savović, Matthew J. Page, Roy G. Elbers, Natalie S. Blencowe, Isabelle Boutron, Christo- pher J. Cates, Hung-Yuan Cheng, Mark S. Corbett, Sandra M. Eldridge, Jonathan R. Emberson, Miguel A. Hernán, Sally Hopewell, Asbjørn Hróbjartsson, Daniela R. ...
2019
-
[58]
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. 2024. Can LLMs generate novel research ideas? A large-scale human study with 100+ NLP researchers. https://arxiv.org/abs/2409.04109
2024 arXiv
-
[59]
Yuan Sun, Eunchae Jang, Fenglong Ma, and Ting Wang. 2024. Generative AI in the wild: Prospects, challenges, and strategies. InCHI Conference on Human Factors in Computing Systems (CHI). doi:10.1145/3613904.3642160
2024
-
[60]
Paletz and Kaiping Peng
Susannah B.F. Paletz and Kaiping Peng. 2008. Implicit theories of creativity across cultures: Novelty and appropriateness in two product domains.Journal of Cross-Cultural Psychology39, 3 (2008), 286–302. doi:10.1177/0022022108315112
2008 doi
-
[61]
Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone. 2024. When combinations of humans and AI are useful: A systematic review and meta-analysis.Nature Human Behaviour8, 12 (2024), 2293–2303. doi:10.1038/s41562-024-02024-1
2024 doi
-
[62]
Luning Sun, Yuzhuo Yuan, Yuan Yao, Yanyan Li, Hao Zhang, Xing Xie, Xiting Wang, Fang Luo, and David Stillwell
-
[63]
Wolfgang Viechtbauer and Mike W.-L. Cheung. 2010. Outlier and influence diagnostics for meta-analysis.Research Synthesis Methods1, 2 (2010), 112–125. doi:10.1002/jrsm.11
2010 doi
-
[64]
Kelly, Saumya Pareek, Qiushi Zhou, and Eduardo Velloso
Samangi Wadinambiarachchi, Ryan M. Kelly, Saumya Pareek, Qiushi Zhou, and Eduardo Velloso. 2024. The effects of generative AI on design fixation and divergent thinking. InCHI Conference on Human Factors in Computing Systems (CHI). doi:10.1145/3613904.3642919
2024
-
[65]
Su, Zhun Deng, Michael Qizhe Xie, Hannah Brown, and Kenji Kawaguchi
Haonan Wang, James Zou, Michael Mozer, Anirudh Goyal, Alex Lamb, Linjun Zhang, Weijie J. Su, Zhun Deng, Michael Qizhe Xie, Hannah Brown, and Kenji Kawaguchi. 2024. Can AI be as creative as humans? https://arxiv.org/ abs/2401.01623v4
2024 arXiv
-
[66]
Emily Wenger and Yoed Kenett. 2025. We’re different, we’re the same: Creative homogeneity across LLMs. https: //arxiv.org/abs/2501.19361v1 Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis 21
2025 arXiv
-
[67]
Wolfgang Viechtbauer. 2010. Conducting meta-analyses in R with the metafor package.Journal of Statistical Software 36, 3 (2010), 1–48. doi:10.18637/jss.v036.i03
2010 doi
-
[68]
Zhikun Wu, Thomas Weber, and Florian Müller. 2025. One does not simply meme alone: Evaluating co-creativity between LLMs and humans in the generation of humor. InInternational Conference on Intelligent User Interfaces (IUI). doi:10.1145/3708359.3712094
2025
-
[69]
Cheng Xu, Shuhao Guan, Derek Greene, and M-Tahar Kechadi. 2024. Benchmark data contamination of large language models: A survey. https://arxiv.org/abs/2406.04244v1
2024 arXiv
-
[70]
Man Zhang, Ying Li, Yang Peng, Yijia Sun, Wenxin Guo, Huiqing Hu, Shi Chen, and Qingbai Zhao. 2025. AI delivers creative output but struggles with thinking processes. https://arxiv.org/abs/2503.23327v1
2025 arXiv
-
[71]
Jiexin Zheng, Ka Chau Wong, Jiali Zhou, and Tat Koon Koh. 2024. Large language model in ideation for product innovation: An exploratory comparative study.A vailable at SSRN 4729982(2024). doi:10.2139/ssrn.4729982
2024 doi
-
[72]
Wilson, Joy P
Robert C. Wilson, Joy P. Guilford, and Paul R. Christensen. 1953. The measurement of individual differences in originality.Psychological Bulletin50, 5 (1953), 362–370. doi:10.1037/h0060857
1953 doi
-
[73]
Wenbo Zou and Feng Zhu. 2025. Generative AI adoption in human creative tasks: Experimental evidence.A vailable at SSRN 5196748(2025). doi:10.2139/ssrn.5196748 22 Niklas Holzner, Sebastian Maier, and Stefan Feuerriegel Acknowledgment of AI Usage & Overview of tools usedDuring t...
2025 doi
-
[77]
Eric Zhou and Dokyun Lee. 2024. Generative artificial intelligence, human creativity, and art.PNAS Nexus3, 3 (2024), pgae052. doi:10.1093/pnasnexus/pgae052
2024 doi
-
[2019]
InCochrane Handbook for Systematic Reviews of Interventions
Analysing data and undertaking meta-analyses: 10. InCochrane Handbook for Systematic Reviews of Interventions. John Wiley & Sons, Ltd, 241–284. doi:10.1002/9781119536604.ch10
-
[2023]
doi:10.3916/C77-2023-04
Writing, creativity, and artificial intelligence: ChatGPT in the university context.Comunicar: Media Education Research Journal31, 77 (2023), 45–54. doi:10.3916/C77-2023-04
2023 doi
-
[2024]
https://arxiv.org/ abs/2412.03151
Large language models show both individual and collective creativity comparable to humans. https://arxiv.org/ abs/2412.03151
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.