{"id":"745150e1-40d2-4d27-ae65-ba8678771c7f","arxiv_id":"2510.20091","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An evaluation of 17 LLMs across nine creativity tasks shows that creativity is fragmented: novelty scores correlate weakly or negatively with quality and diversity, and the proprietary-model advantage largely disappears on divergent thinking.","lead":"CreativityPrism is a new benchmark that scores 17 large language models on nine creativity tasks across three domains, using 20 metrics grouped into quality, novelty, and diversity. It finds that proprietary models lead on creative writing and logical reasoning, but the gap shrinks on divergent thinking, and strong performance in one creativity dimension rarely predicts another.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The fragmentation claim rests on 17-model correlations with no uncertainty; the cited r=-0.25 is not significantly different from zero, so the central conclusion is statistically under-supported.","rationale":"The paper does valuable work assembling nine tasks and validating some judge-based metrics against human data. However, the abstract's central claim is a correlation-structure claim. In §5.3, the authors form one vector per metric from 17 models and report Pearson r values. At n=17, the sampling distribution is wide; a reported r of −0.25 is within one standard error of zero. Without confidence intervals, significance tests, or correction for the dozens/hundreds of pairwise comparisons, the specific 'negative correlation' example cannot be distinguished from noise. The non-independence of model checkpoints compounds this. Because the reader already assigned CONDITIONAL, my check does not move the verdict; it identifies a sharper justification for the condition: before accepting the fragmentation claim, the authors should provide an artifact and bootstrap/permutation tests. I also credit the paper for acknowledging LLM-judge bias in the limitations and for using existing human validation where available; the TTCT and CS4 gaps remain, but they are secondary to the missing statistical support for the headline.","tokens_in":35297,"tokens_out":5704,"duration_ms":53347,"concrete_test":"Release the per-metric normalized score matrix for all 17 models (or a public artifact) and recompute the §5.3 pairwise Pearson correlations with 95% bootstrap CIs and Benjamini-Hochberg-adjusted p-values. Specifically, test whether the Surprise–Divergence@0 interval excludes zero and whether any cross-domain novelty correlations survive adjustment. If the intervals all include zero, the central fragmentation claim is not established by the current data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in §5.3 — that novelty metrics correlate weakly or negatively across domains — is a statistical claim, but the paper provides no uncertainty quantification. With M=17 models, even a true correlation of zero produces sampling noise with standard error ≈1/√(17−3)≈0.27; the one negative correlation highlighted in the text (Creative Short Story 'Surprise' vs NeoCoder 'Divergence@0', r=-0.25) has a 95% CI that spans roughly −0.66 to +0.25. The paper reports no p-values, confidence intervals, or multiple-comparison correction for the full matrix, and the model set is not an independent sample (OLMo-2-7B/13B, Qwen-2.5-7B/32B/72B, DeepSeek-V3/R1, multiple Claude/GPT checkpoints), so the effective N is smaller than 17. This matters because the headline 'fragmentation' conclusion is exactly the kind of claim that can be produced by noise when comparing many correlations. The LLM-judge validation concern raised by the reader is real — TTCT has no human annotations and CS4 has only 15 stories with r=0.55 — but it compounds rather than replaces this problem: without human-scored recalibration, both the absolute scores and the correlation structure could shift.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CreativityPrism, a benchmark and analysis framework for evaluating LLM creativity. It consolidates nine tasks in three domains (divergent thinking, creative writing, logical reasoning) and organizes twenty metrics into three dimensions (quality, novelty, diversity). The authors evaluate 17 proprietary and open-weight LLMs, aggregate scores by dimension and domain, and report a correlation analysis across metrics. The central claims are that proprietary models outperform open models in creative writing and logical reasoning but not in divergent thinking, and that creativity is fragmented: high performance on one dimension or domain does not reliably transfer, especially for novelty metrics, which show weak or negative correlations with other metrics. The paper also validates automatic LLM judges for six of the nine tasks, with varying levels of supporting evidence.","tokens_in":35557,"tokens_out":4418,"duration_ms":40928,"significance":"The benchmark addresses a real and timely gap: most LLM creativity evaluations are task-specific, single-domain, or rely on expensive human evaluation. CreativityPrism is useful as a structured consolidation of existing tasks and metrics, and the three-dimensional taxonomy is a sensible organizing principle. The paper also makes an explicit effort to check LLM-as-a-judge reliability, which is more than many benchmark papers do. If the results hold, the fragmentation claim would be an important caution against reporting a single creativity score. However, the two main supports for that claim — the LLM-judge validation and the cross-metric correlation analysis — are currently not strong enough to carry the conclusions, so the significance is conditional on the revisions described below.","major_comments":[{"comment":"The central fragmentation conclusion is a statistical claim about correlations among model performance vectors, but the paper provides no uncertainty quantification. For M=17 models, the standard error of a correlation under the null is about 1/sqrt(17−3) ≈ 0.27, so the headline example (Creative Short Story 'Surprise' vs NeoCoder 'Divergence@0', r=−0.25) has a 95% interval of roughly (−0.66, +0.26). That interval includes zero, and many other correlations in the 17×17 matrix are likely consistent with zero after accounting for multiple comparisons. In addition, several rows are from the same model family (Qwen, DeepSeek, Claude, GPT), so the effective independent sample size is smaller than 17. The paper should report confidence intervals or significance tests (ideally with a multiple-comparison adjustment) before claiming that novelty metrics are weakly or negatively correlated across","section":"§5.3, Figures 3–4"},{"comment":"The LLM-as-a-judge validation is too thin to support the ranking and correlation structure. For TTCT the paper states 'we have no human annotations at all' and instead reports Pearson correlations between Qwen2.5-72B and GPT-4.1 (0.50–0.69). That measures agreement between two LLM judges, not validity against human judgment, and the lower values are not obviously acceptable for a metric used to rank 17 models. For CS4, the human–LLM correlation is 0.55 on 15 stories; this is a small and modest alignment. For TTCW, the appendix (E.1.4) mentions a correlation threshold of 0.2 while the main text (D) reports accuracy values; the two descriptions are not consistent. These issues affect six of nine tasks, so they are load-bearing for the benchmark's conclusions; the manuscript's own acknowledgement that more human annotation is being collected confirms the validation is incomplete.","section":"§4.2, Appendix D, Appendix E.2.5, E.8"},{"comment":"The abstract states that frontier LLMs 'offer no significant advantage in divergent thinking,' but §5.2 claims 'more than 10% in each domain' and Table 2 shows Claude3-Sonnet at 0.833 in divergent thinking versus the best open-weight model (Qwen2.5-72B) at 0.731, a gap of about 0.10 (roughly 14% relative). If 'no significant advantage' is meant statistically, no significance test is reported; if it is meant practically, the table does not support it. The claim should be reconciled with the reported numbers and, if retained, supported by an appropriate test.","section":"Abstract vs. §5.2 and Table 2"}],"minor_comments":[{"comment":"The phrase 'by a .10 (or 15%) lead' is unclear and appears garbled. Please use consistent absolute or relative numbers and specify which models are being compared.","section":"Abstract"},{"comment":"Table 3 lists 19 models (including OLMo2-13B-SFT and OLMo2-13B-DPO), while Table 2 reports 17 models, and Appendix E.1.5/E.9.5 include Claude3-Opus, which does not appear in Table 2. The model inventory should be harmonized so that every reported result has a corresponding model entry.","section":"Table 3 / Appendix E"},{"comment":"DeepSeek-R1 and DeepSeek-V3 are labeled as 'Gemini' in the Family column; this is a typo.","section":"Table 3"},{"comment":"The 'more than 20% in each dimension' and 'more than 10% in each domain' claims should state whether they are absolute or relative gaps and give the exact best-model comparisons from Table 2.","section":"§5.2"},{"comment":"The table header in the Creative Short Story results uses 'Surprisal' and 'N-gram Diversity,' while the text refers to a 'novelty score'; please unify the metric names.","section":"E.4.6"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: CreativityPrism is a well-documented benchmark suite that could become a standard tool for comparing LLM creativity, but the paper's central claim—that creativity is fragmented across dimensions—rests on a correlation analysis that is statistically under-powered, and the headline numbers in the abstract are not fully consistent with Table 2.\n\nWhat's genuinely useful: the three-dimensional taxonomy (quality/novelty/diversity) is a reasonable organizing scheme, and compiling nine tasks with 20 metrics across three domains is real work. The task documentation is unusually detailed, including prompts, model outputs, and judge-validation attempts. The empirical observation that quality and diversity correlate while novelty doesn't is interesting if it holds; it's the kind of result that would motivate a multi-task benchmark.\n\nWhere it gets soft. First, the abstract says frontier models 'offer no significant advantage in divergent thinking,' but Table 2 shows Claude3-Sonnet at 0.833 vs Mistral-7B at 0.758—about 10% relative. The §5.2 text says 'more than 10% in each domain,' which only barely holds for best-vs-best and not for average gaps. That's a real inconsistency. Second, the fragmentation claim in §5.3 is built on Pearson correlations across 17 models, with no confidence intervals or p-values. The example they highlight, r=-0.25, is not significantly different from zero given that sample size; the effective N is smaller because the models are not independent (multiple checkpoints of the same family). The stress-test note is right: this claim needs error bars or it should be stated as a hypothesis. Third, the LLM-judge validation is thin where it matters most: TTCT has no human annotations, just correlation between two LLMs, and CS4 has n=15 with r=0.55. That's not enough to establish that the judge rankings are trustworthy enough to support the cross-task correlation structure. There are also minor internal inconsistencies—8 vs 9 tasks, 17 vs 19 models, DeepSeek labeled 'Gemini' in Table 3—that suggest the paper isn't ready in its current form.\n\nNone of this is fatal. The framework is worth building on, and the authors are honest about limitations. But the load-bearing fragmentation conclusion needs statistical backing, and the judge validation needs more human data.\n\nWho this is for: anyone working on LLM evaluation or creativity benchmarks. I'd send it to a serious referee, expecting major revisions. I wouldn't cite it yet without checking the numbers, and I'd want to see code and data released first.","headline":"A useful benchmark assembly with a load-bearing fragmentation claim that the current numbers don't yet support.","tokens_in":36148,"tokens_out":3533,"would_cite":false,"duration_ms":28037,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that LLM creativity is fragmented: a model's strong performance in one creative dimension or domain rarely predicts strong performance in another, so a single score cannot measure machine creativity.","keywords":["creativity evaluation","large language models","benchmark","divergent thinking","creative writing","logical reasoning","novelty","diversity"],"falsifier":"Collect human creativity ratings on a sample of outputs from all nine tasks (especially the one with no human annotations at all), replace the AI-judge scores with human scores, and recompute the cross-metric correlation matrix; if the weak and negative novelty correlations disappear, the fragmentation claim would collapse.","tokens_in":35130,"feed_emoji":"🧩","tokens_out":4494,"duration_ms":42810,"temperature":0.7,"pith_summary":"The paper sets out to build a scalable, cross-domain way to measure creativity in large language models. It proposes a framework with nine tasks spanning divergent thinking, creative writing, and logical reasoning, and scores each model on three dimensions—quality, novelty, and diversity. Evaluated on 17 models, the framework shows that proprietary models lead open models by about 10–15% overall but not in divergent thinking. The central finding is that creativity does not transfer: quality and diversity correlate with each other, while novelty correlates weakly or even negatively (e.g., story surprise vs. coding divergence at r = -0.25). The authors conclude that meaningful assessment of LLM creativity requires multiple tasks, domains, and dimensions rather than any single aggregate.","feed_headline":"Novelty barely correlates across LLM creativity metrics","feed_subtitle":"Seventeen models, nine tasks: quality and diversity track together, but novelty does not—so no single score can rank machine creativity.","key_machinery":"CreativityPrism, a benchmark and analysis suite. It organizes nine existing tasks (e.g., Alternative Uses Test, creative short-story generation, constrained code generation, creative math) into three domains and classifies twenty task-specific metrics into three dimensions: quality (does the output satisfy task requirements), novelty (is it rare compared to existing content), and diversity (how varied are multiple outputs). The analysis machinery is a correlation matrix over model performance vectors: for each metric, the normalized scores of all 17 models are stacked into a vector, and Pearson correlations between metric pairs reveal which dimensions travel together. That matrix is the load","core_discovery":"The central claim, stated on the paper's own terms, is that creativity in LLMs is multidimensional and fragmented. High performance in one creative dimension or domain rarely generalizes to others; in particular, novelty metrics often show weak or negative correlations with other metrics, because 'novelty' means different things in different tasks—being surprising in a short story is not the same as solving a coding problem in an unprecedented way. From this, the paper argues that a holistic benchmark spanning three domains and three creativity dimensions is necessary, and that an 'overall' creativity score should be treated only as a coarse comparison aid, not as a measure of a single under","pith_inferences":["A testable extension: because the open–proprietary gap is smallest in divergent thinking, post-training with divergent-thinking tasks might be where open models can catch up most cheaply.","The weak novelty dimension suggests that 'novelty' is not one capability but a family of task-specific behaviors; a model could be trained to inflate one novelty metric without gaining another.","If human ratings later show the AI judges are biased, the specific rankings would shift, but the fragmentation structure might survive only if the bias is uniform across tasks; checking this with per-task human data is a natural next step."],"forward_implications":["A single-task or single-domain evaluation of LLM creativity will misreport model capability, because strong performance in one domain does not predict performance in another.","Novelty should be measured and reported per task, since novelty metrics across tasks share little variance and can even be negatively correlated.","Models strong on quality tend to be strong on diversity, so diversity can serve as a rough proxy for quality, but neither predicts novelty.","Frontier proprietary models lead in creative writing and logical reasoning by roughly 0.10–0.15 normalized points, yet show no clear advantage in divergent thinking, suggesting those domains are undertrained.","The benchmark's three-dimension scores—quality, novelty, diversity—should be reported separately; the overall score is only a coarse summary."],"fun_headline_variants":["Novelty in LLMs doesn't travel with quality or diversity","Frontier LLMs still trail in divergent thinking, benchmark shows","Creativity isn't one thing: quality, novelty, diversity part ways","New benchmark: LLM creativity splits into non-overlapping skills","Cross-domain test: novelty metrics clash with practical creativity"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The rankings and the fragmentation conclusion rest on the assumption that the AI judge's creativity scores align with human judgments of creativity, an alignment that is only thinly validated for some of the nine tasks.","fun_headline_variants_meta":{"raw":{"variants":["Novelty in LLMs doesn't travel with quality or diversity","Frontier LLMs still trail in divergent thinking, benchmark shows","Creativity isn't one thing: quality, novelty, diversity part ways","New benchmark: LLM creativity splits into non-overlapping skills","Cross-domain test: novelty metrics clash with practical creativity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000317,"raw_usage":{"total_tokens":1647,"prompt_tokens":778,"completion_tokens":869,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":782}},"tokens_in":522,"tokens_out":869,"duration_ms":8748,"temperature":1.0,"reasoning_tokens":782,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T08:30:22.546701+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect human creativity ratings on a sample of outputs from all nine tasks (especially the one with no human annotations at all), replace the AI-judge scores with human scores, and recompute the cross-metric correlation matrix; if the weak and negative novelty correlations disappear, the fragmentation claim would collapse.","supporting_citations":[],"review_version":1}