{"id":"37f65953-3eb9-40b9-8539-bb65ac3514db","arxiv_id":"2412.03151","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The best LLMs match the average human on overall creativity, excel in divergent thinking and problem solving, lag in creative writing, and when sampled repeatedly match a small human group.","lead":"This study benchmarked five large language models against 467 humans on 13 creativity tasks across three domains, finding the best models at the 52nd percentile of human performance. It also quantifies collective creativity: one LLM asked 10 times produces ideas comparable to a group of 8 to 10 humans.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline percentile and collective-equivalence numbers rest on an unvalidated z-score equating between two rater rounds; the GPT-3.5 v1 anchors are small and low-scoring, leaving the upper tail where 'top 10' is decided unverified.","rationale":"The paper's central claim requires that the human ratings (Round 1) and LLM ratings (Round 2) are on a common scale. The only bridge is the z-score linear equating using GPT-3.5 v1 responses, rated in both rounds. This is not a minor statistical footnote: every headline number—the 52nd percentile, the '8-10 humans' equivalence, the 1/0.52 slope—is a function of where LLM scores land in the human distribution. The equating is especially fragile at the upper tail because the anchor responses come from the weakest model and are mostly low-scoring; there is almost no anchor information about how Round 2 raters treat the kinds of high-quality, top-10 responses that drive the collective-creativity analysis. A nonlinear difference in rater severity—for example, Round 2 raters being more discriminating among strong responses—would shift the equated scores of top LLM responses and alter the composition of the top 10. The authors do not report equating diagnostics, anchor sample sizes per task, or any alternative linking (e.g., equipercentile, IRT). This is an addressable concern: the data are public, so the check we propose is straightforward. I do not see an internal inconsistency in the paper; the concern is an unvalidated assumption, not a demonstrated error. The other limitations the reader notes (sample representativeness, operationalization of collective creativity, multiple comparisons) are real but secondary; the equating is the single point without which the headline comparisons do not exist. Hence the CONDITIONAL verdict is appropriate, and no change to the reader's verdict is needed.","tokens_in":51706,"tokens_out":7305,"duration_ms":64189,"concrete_test":"Download the public data (figshare 24878421). For each task, take the GPT-3.5 v1 responses rated in both rounds and compute the equipercentile (rank-based) linking function between Round 1 and Round 2 ratings; apply that function to Round 2 LLM ratings instead of the z-score linear equation. Recompute the LLM percentile ranks (Fig. 2) and the collective-creativity values (Figs. 6-7) under the equipercentile equating. If the best-LLM percentile changes by more than about 3 points, or the equivalent human group size by more than 1, the linear equating assumption is not robust. Additionally, report the anchor-score correlation and a loess fit; a clear nonlinear pattern would confirm the concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All human-LLM comparisons in Figs 2, 6, and 7 depend on equating Round 1 (humans) and Round 2 (LLMs) ratings. The Methods ('Rating procedure') state that z-score linear equating was applied using GPT-3.5 v1 responses as anchors, with truncation, but no diagnostics are reported: no anchor means/SDs per round, no scatterplot, no equating error. The anchor set is small (about 50 responses per task) and comes from the weakest model, so it is concentrated at the low end of the scale. The collective-creativity metric selects the top 10 responses from a pooled set; this is exactly the score range where the anchor data are thinnest and where linear equating is most likely to fail if Round 2 raters are more (or less) severe for high-creativity outputs than for typical ones. A nonlinear rater effect would change which responses enter the top 10 and therefore move both the 52nd percentile and the '8-10 humans' equivalences. In the absence of equating diagnostics or a robustness check, the central claim is conditional.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks five large language models (GPT-3.5, GPT-4, Claude, Qwen, SparkDesk) against 467 human participants on 13 creativity tasks spanning divergent thinking, problem solving, and creative writing. Human raters scored all responses following the Consensual Assessment Technique, and the authors report percentiles of individual LLM performance in the human distribution, with the best models (Claude, GPT-4) at the 52nd percentile. They then introduce a collective-creativity metric that pools multiple LLM responses with varying numbers of human responders, and claim that one LLM asked 10 times is equivalent to a group of 8--10 humans, and that roughly two additional LLM responses add as much as one extra human. Domain-specific results and analyses of diversity, temperature effects, and demographic differences are also reported.","tokens_in":51887,"tokens_out":6298,"duration_ms":57902,"significance":"If the central results hold, this is a useful multi-domain benchmark with a novel operationalization of collective creativity that goes beyond single-response comparisons. The study has notable strengths: the tasks are not taken from existing online datasets, reducing contamination concerns; ratings are provided by trained judges with reported inter-rater reliability; the data and code are made available; and the authors transparently discuss limitations such as the non-representative human sample, the nominal-group framing, and the absence of prompt engineering. The collective-creativity equivalence is a falsifiable empirical claim rather than a derivation, and the paper clearly separates the bootstrapped pooling procedure from the linear slope fitted to the aggregate results. These strengths make the paper a potentially valuable reference point for debates about LLM creativity in applied settings.","major_comments":[{"comment":"The z-score linear equating that aligns Round 1 (human) and Round 2 (LLM) ratings is the backbone of every human-LLM comparison, but the manuscript provides no equating diagnostics. No anchor means, standard deviations, scatterplots, or equating error are reported for the GPT-3.5 v1 anchor responses. The anchor set is small (about 50 responses per task) and comes from a model that is generally weaker than humans, so the anchor ratings are likely concentrated at the low end of the scale; yet the collective-creativity analysis selects the top 10 responses from a pooled set, which is exactly the upper tail where anchor data are thinnest. If rater severity differs nonlinearly between rounds, the set of responses entering the top 10 would change, shifting both the 52nd-percentile claim and the 8--10-human equivalences. Please report anchor distributions per task, provide equating diagnostics, and include a robustness check (e.g., re-rating a subset of human responses in Round 2, using a different equating method, or nonparametric equating).","section":"Methods: Rating procedure (pp. 22-23); Figs 2, 6, 7"},{"comment":"The headline numbers -- the 52nd percentile for Claude/GPT-4, the collective equivalence of 8--10 humans, and the slope of 0.52 used for 'two additional LLM responses equal one extra human' -- are reported as point estimates without confidence intervals or dispersion measures. The percentile ranks vary widely across the 13 tasks (e.g., 25th percentile in creative writing vs. 55th and 59th percentiles in divergent thinking and problem solving), so the average percentile is not characterized by its mean alone. In Fig. 7 and Table S13, the linear relationship is fitted without reporting standard errors, R-squared, or residual diagnostics; this is particularly relevant because the slope is used directly in the abstract and discussion. Please provide per-task distributions, bootstrap confidence intervals for the equivalence points, and regression fit diagnostics for the linear slopes.","section":"Results Part 1 and Part 2; Figs 2, 6, 7; Table S12"},{"comment":"The paper runs dozens of independent-sample t-tests and ANOVAs without any multiple-comparison correction. While the central percentile and collective-creativity claims are not based on these individual significance tests, the domain-level conclusions (e.g., 'LLMs excel in divergent thinking and problem solving' and specific task-by-task superiority claims) rely on patterns across many comparisons, and several would likely not survive a Benjamini-Hochberg or Bonferroni correction. Please either apply a correction to the main task-level comparisons or explicitly label the tables as exploratory and focus the interpretation on effect sizes and consistency across tasks.","section":"Statistical Analysis; Tables S4, S5, S16-S22"}],"minor_comments":[{"comment":"The percentile for GPT-3.5 is reported as the 37th percentile in Results and as the 38th percentile in Discussion; please make these consistent.","section":"Discussion (p. 17) vs Results (p. 8)"},{"comment":"The author name 'can der Maas' should read 'van der Maas'.","section":"Table 1, reference 15"},{"comment":"The author name 'A. uncdogan' appears to be a corrupted rendering; please verify the spelling against the published source.","section":"Reference 39"},{"comment":"For GPT-4 and Claude the overall human contribution at N=10 is 48.71% and 50.03%, respectively, so the reported equivalence '10 humans' is a rounded or interpolated value; please state explicitly how the equivalence point is derived from the bootstrapped percentages and whether values are rounded.","section":"Results: 'When one LLM is asked 10 times' (p. 14) and Table S12"},{"comment":"The statement that 'temperature should be better interpreted as a parameter for diversity rather than a parameter for creativity' would benefit from a brief report of the effect sizes (partial eta-squared) for the non-significant creativity ANOVAs, not only the significant diversity results.","section":"Effect of temperature (p. 12)"},{"comment":"The figures show mean collective-creativity values without error bars or confidence bands; please indicate in the captions whether error bars are omitted because they were not computed or because they were too small to display.","section":"Figures 6 and 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to draw attention given the strong practical claims about LLM creativity. The central results are conditional on the validity of the z-score equating between the two rating rounds, and the missing equating diagnostics are a real gap that, in my view, prevents acceptance at the current stage. The authors should be encouraged to provide the diagnostics and a robustness analysis; the rest of the manuscript is well within the scope and the data-sharing practices are a strength. I would suggest that the editor consider sending the revised version to a statistician or psychometrician to assess the equating and the collective-creativity metric."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: This is the most thorough multi-task benchmark of LLM creativity I've seen—13 new tasks, three domains, 467 humans, and a genuinely new collective-creativity comparison. The main results are plausible, but the two-round rating equating is load-bearing and unvalidated, so the percentile and collective-equivalence figures should be read as conditional.\n\nWhat the paper does well: it goes beyond single-task studies by constructing 13 tasks across divergent thinking, problem solving, and creative writing, and by benchmarking five LLMs against a large high-stakes human sample using the Consensual Assessment Technique with five raters per round. The new collective metric—asking an LLM 10 times and comparing its top-10 ideas against groups of increasing human size—is relevant to real usage. The authors also check temperature effects, text diversity, and demographic differences, and they are upfront about the non-representative sample and nominal group setting. Data are on figshare, and Table 1 is a useful summary of prior work.\n\nThe biggest soft spot is the equating of Round 1 (humans + GPT-3.5 v1) and Round 2 (all LLMs) ratings. The paper uses z-score linear equating with GPT-3.5 v1 responses as anchors, but reports no diagnostics: no anchor means/SDs per round, no scatterplot, no equating error, and no robustness check. The anchor set is small (about 50 responses per task) and comes from the weakest model, so it is concentrated at the low end of the scale. The collective metric selects top-10 responses, exactly the range where anchor data are thinnest. A nonlinear rater-severity difference would shift which responses enter the top 10 and thus move both the 52nd percentile and the 8-10 human equivalences. This is fixable—report the equating diagnostics and rerun under alternative equating assumptions—but without it the headline claims are conditional.\n\nOther, more minor concerns: the human sample is a Master's admission applicant pool (Chinese, mostly female, high-stakes), not a general population; the authors test demographics and find little, which helps. They also run many t-tests without multiple-comparison correction, so a few significant differences are likely false positives. The collective-creativity numbers depend on operational choices (10 responses, top-10 count, fitted linear slope) and should get a sensitivity analysis. None of this undermines the overall pattern—LLMs at or near human average, strong in divergent thinking, weak in creative writing—which is consistent with prior work.\n\nThis paper is for anyone working on AI evaluation, creativity measurement, or the future of work. It deserves a serious referee. I'd recommend conditional acceptance with a request for equating diagnostics and sensitivity analyses.","headline":"A serious multi-domain benchmark with a novel collective-creativity metric, but the headline figures hinge on an unvalidated rater equating.","tokens_in":52449,"tokens_out":3251,"would_cite":true,"duration_ms":30717,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 13 creative tasks, top LLMs rank at the 52nd percentile of humans and 10 responses equal a group of 8-10 people.","keywords":["large language models","creativity","divergent thinking","problem solving","creative writing","collective creativity","human benchmarking","future of work"],"falsifier":"Have the same panel rate a random sample of human and LLM responses from both rounds in one sitting; if the directly observed percentile ranks or human-equivalence numbers diverge noticeably from the linearly equated values, the central claim weakens. A simpler check is to re-estimate the collective-creativity metric without the anchor responses and see whether GPT-4 and Claude still land between 8 and 10 humans.","tokens_in":51484,"feed_emoji":"🎨","tokens_out":6282,"duration_ms":59513,"temperature":0.7,"pith_summary":"Using 13 creative tasks across divergent thinking, problem solving, and creative writing, this paper asks whether large language models are as creative as the people who will soon work beside them. It finds that the best models, Claude and GPT-4, rank at the 52nd percentile of a large human sample, with models strongest on divergent thinking and problem solving and weakest on creative writing. The paper goes further than single-response benchmarks by sampling each model repeatedly and showing that one model asked ten times is equivalent, in top-idea production, to a group of eight to ten humans. Because a typical workplace brainstorming group is smaller than that, the result suggests that LLMs could serve as near-human creative teammates rather than merely routine-task tools.","feed_headline":"LLMs match humans in creativity and equal 8-10 people in groups","feed_subtitle":"The best models sit at the 52nd human percentile on 13 tasks, so creative work may no longer be exclusively human.","key_machinery":"The load-bearing instrument is the collective-creativity equivalence metric. For each task, the authors split a model's responses into groups of ten, pool each group with responses from N randomly drawn human participants, and bootstrap this pooling 1000 times; the N at which humans and the model contribute half of the top ten responses becomes the 'number of humans the model equals.' The individual-level comparison is carried by the same percentile-ranking procedure, and both rest on ratings produced by five trained judges per round using the Consensual Assessment Technique, with the two rating rounds put on one scale by z-score linear equating anchored on repeated GPT-3.5 responses.","core_discovery":"On its own terms, the paper establishes that five LLMs tested against 467 humans on 13 bespoke creative tasks produce responses that human judges, blind to authorship, rate as roughly average-to-slightly-above human: all models together average the 46th percentile, with Claude and GPT-4 at the 52nd. In divergent thinking and problem solving the models sit at the 55th and 59th percentiles, while in creative writing they fall to the 25th. When the same model is asked ten times and its responses are pooled with those of varying numbers of humans, the model's share of the top-rated ideas matches that of 8-10 humans; asking for more responses yields a linear exchange rate in which roughly two additional LLM responses add as much as one additional human. The paper reads this as evidence that LLM creativity is already comparable to individual human creativity and, when sampled repeatedly, to the collective output of a small human group.","pith_inferences":["If the two rating rounds differ by more than a linear shift, the 52nd percentile and the 8-10 human equivalence would move; a single-round re-rating of a common sample would show how much.","The equivalence to 8-10 humans is measured against a nominal group of independently working humans; interactive brainstorming groups behave differently, so the real-world team size an LLM replaces could be larger or smaller.","An untested extension is whether the 2-responses-per-human exchange rate holds beyond fifty responses or saturates as the model's own output diversity becomes the limiting factor.","Since temperature changes diversity more than rated creativity, sampling strategy (how many responses, at what temperature, with what prompt variation) may matter more than any single generation setting for collective creative output."],"forward_implications":["Workplace brainstorming with fewer than ten people can now plausibly include an LLM as a source of top-idea generation, not just a clerical aid.","In divergent thinking and problem-solving tasks, asking a model several times before selecting the best idea is a cheap way to match a small human group's best output.","Creative writing remains the domain where human superiority is clear; organizations should not expect LLMs to replace human writers on emotional or memorable messaging.","Because the paper replicates the finding that LLM outputs are less diverse than human outputs, teams that lean on LLMs alone should expect idea homogenisation and should keep human variability in the pipeline.","Single-task creativity studies will keep producing inconsistent verdicts; a multi-domain, multi-dimension battery is needed to say anything general about machine creativity."],"supporting_citations":[{"why":"Supplies the standard definition of creativity as novelty plus usefulness that guides the rating dimensions.","marker":"[7]"},{"why":"Restates the standard definition the paper uses to justify rating ideas on novelty and usefulness.","marker":"[8]"},{"why":"Provides the Consensual Assessment Technique: independent judges unaware of authorship rate creative products, the core of the rating procedure.","marker":"[37]"},{"why":"Prior benchmark showing LLM stories pass at low rates versus professionals; the creative-writing comparison the paper extends and qualifies.","marker":"[20]"},{"why":"Establishes that typical real-world brainstorming groups are smaller than ten people, the practical yardstick for the collective-creativity equivalence.","marker":"[45]"},{"why":"Earlier evidence that generative AI raises individual creativity but lowers collective diversity; the paper replicates this diversity deficit and uses it to frame a limitation.","marker":"[42]"}],"fun_headline_variants":["LLMs hit human-level creativity, score 52nd percentile","AI creativity matches humans: best models equal 8-10 people","LLMs rank 52nd percentile in human creativity tests","Two extra AI answers equal one human in creative tasks","LLMs collectively equal a group of 8-10 humans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline comparisons assume that the two groups of human raters used a common internal scale, so a linear z-score rescaling is enough to make Round 1 human scores and Round 2 LLM scores directly comparable.","fun_headline_variants_meta":{"raw":{"variants":["LLMs hit human-level creativity, score 52nd percentile","AI creativity matches humans: best models equal 8-10 people","LLMs rank 52nd percentile in human creativity tests","Two extra AI answers equal one human in creative tasks","LLMs collectively equal a group of 8-10 humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000514,"raw_usage":{"total_tokens":2474,"prompt_tokens":902,"completion_tokens":1572,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":1488}},"tokens_in":518,"tokens_out":1572,"duration_ms":9692,"temperature":1.0,"reasoning_tokens":1488,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:42:29.994750+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have the same panel rate a random sample of human and LLM responses from both rounds in one sitting; if the directly observed percentile ranks or human-equivalence numbers diverge noticeably from the linearly equated values, the central claim weakens. A simpler check is to re-estimate the collective-creativity metric without the anchor responses and see whether GPT-4 and Claude still land between 8 and 10 humans.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the standard definition of creativity as novelty plus usefulness that guides the rating dimensions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Restates the standard definition the paper uses to justify rating ideas on novelty and usefulness."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Consensual Assessment Technique: independent judges unaware of authorship rate creative products, the core of the rating procedure."},{"cited_title":"Chakrabarty, P","cited_arxiv_id":null,"evidence_quote":"Prior benchmark showing LLM stories pass at low rates versus professionals; the creative-writing comparison the paper extends and qualifies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that typical real-world brainstorming groups are smaller than ten people, the practical yardstick for the collective-creativity equivalence."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier evidence that generative AI raises individual creativity but lowers collective diversity; the paper replicates this diversity deficit and uses it to frame a limitation."}],"review_version":1}