REVIEW 3 major objections 5 minor 32 references
How do Humans and Language Models Reason About Creativity? A Comparative Analysis
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Language models can predict human creativity ratings more accurately than human experts can, but they do so by collapsing the three rated facets of creativity into a single dimension, especially when given example solutions with ratings.
desk verdict Fine-grained human/LLM creativity comparison with a promising but untested homogenization claim; the joint-prompt confound needs an ablation before that interpretation is justified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the facet-based originality rating protocol: each solution is rated for originality plus remoteness, uncommonness, and cleverness on five-point Likert scales, and raters write one-to-two-sentence explanations of their originality score. The explanations are then coded by two LLMs (GPT-4o and Claude-3.5-sonnet) for five linguistic markers—comparative, causal/analytical, perceptual, past/future, and cleverness—using zero-shot prompts adapted from Rathje et al. The argument is carried by comparing Pearson correlations among facets and between each facet and originality across conditions (with versus without examples) and across populations (humans versus LLMs), with Fisher's z tests for significance. The key identity at stake is the view of originality as an aggregation of facets: if facet correlations approach 1.0, the facets no longer carry independent information, which the paper terms homogenization.
What would settle it
Have human annotators label the same set of expert and LLM explanations for the five linguistic markers and compute agreement (e.g., Cohen's kappa) with GPT-4o and Claude-3.5-sonnet ratings. Low kappa (below roughly 0.6) on comparative or causal/analytical markers would undercut the claim that no-example experts rely more on memory retrieval. Alternatively, if an independent LLM evaluation with modified prompts or different examples shows facet–originality correlations below 0.8, then homogenization is not a fixed property of LLM creativity evaluation.
Extended reading notes
Core claim
The paper claims that LLM evaluators of creativity have stronger predictive validity but weaker construct validity than human experts. In Study 1, experts who rated design-problem solutions without examples used more comparative language and emphasized uncommonness, suggesting memory-based retrieval; those given examples shifted weight toward cleverness. In Study 2, GPT-4o-mini and Claude-3.5-haiku prioritized remoteness and uncommonness, and giving them example solutions raised their correlation with ground-truth originality from roughly 0.6–0.67 to 0.74–0.76. Yet examples also made the three facets nearly perfectly correlated with originality (above 0.99), meaning the models no longer distinguished between 'clever,' 'remote,' and 'uncommon'—the facets homogenized into a single originality judgment. The paper argues these patterns reveal diverging evaluation strategies: humans weigh facets differently depending on context, while LLMs reduce all facets to semantic distance from prior knowledge, especially when shown examples.
Load-bearing premise
The study assumes that the LLM-generated labels for linguistic markers in the explanations—comparative, causal/analytical, perceptual, past/future, and cleverness—are valid measures of the cognitive processes raters used, yet these labels were not validated against human-annotated gold standards for these specific categories and prompts.
Editorial extensions
If this is right
- Automated LLM creativity scoring can achieve high agreement with human averages while failing to model the conceptual distinctions human raters make; agreement on scores does not imply agreement on the basis of evaluation.
- Low-temperature LLM evaluators produce less diverse explanation styles and more rigid, analytical justifications than human experts, even when their numeric ratings match.
- Including few-shot example solutions improves LLM accuracy but amplifies facet homogenization, so benchmark gains in predictive accuracy can conceal a loss in construct validity.
- Construct-level evaluation of AI judges—testing whether facets stay separable—is needed alongside accuracy metrics before deploying LLMs as reviewers of STEM proposals or papers.
Reading between the lines
- A testable extension is to probe LLM evaluators with 'trap' responses in which one facet is high and another low; if facet correlations remain near 1.0, the model is likely computing a single semantic-similarity score rather than three distinct constructs.
- The contrast between humans shifting to cleverness and LLMs collapsing facets hints that LLM pretraining proxies 'originality' by distance from common responses, not by task-specific insight; this could be tested by fine-tuning on creativity-annotation data and re-running the facet analysis.
- If the linguistic-marker labels are accepted, the comparative-language result offers a cheap way to detect memory-based evaluation in human raters, which could be used to calibrate rater training or flag AI-generated justifications.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports two studies comparing how human experts and LLMs evaluate the originality of responses to engineering design problems (DPT). Raters provide originality scores, brief written explanations, and ratings for three facets of originality—uncommonness, remoteness, and cleverness—either with or without example solutions and ratings. Study 1 (72 Prolific participants with STEM degrees) finds that no-example experts use more comparative and causal/analytical language than example experts, and that the presence of examples changes some facet correlations. Study 2 runs a parallel protocol with GPT-4o-mini and Claude-3.5-Haiku using a structured prompt, finding that examples improve LLM accuracy against ground-truth originality scores but that facet-originality correlations rise to values near 0.99 in the example condition. The paper interprets this as a homogenization of the facets in LLM evaluation and argues that LLM scores have stronger predictive validity but weaker construct validity than human scores.
Significance. If the central claims survive scrutiny, the paper is a useful contribution: it moves beyond accuracy-only benchmarks for LLM creativity evaluation, offers a parallel human–LLM design, uses multiple LLM families, reports both Pearson and Kappa statistics, and makes code and data available. The fine-grained facet protocol and the attention to explanation text are valuable. However, the paper's headline claim about LLM facet homogenization is currently confounded by the joint-format prompt used for LLM ratings, and the linguistic-marker analyses rely on LLM annotations without human validation for the specific categories. The significance is therefore conditional on resolving these issues.
major comments (3)
- [Study 2 Methods and Results; Figure 10; Future Work] The central evidence for facet homogenization comes from correlations among LLM ratings that are generated in a single API call with a fixed format requiring ORIGINALITY, UNCOMMON, REMOTE, CLEVER, and EXPLANATION in one response. With temperature set to zero, the model is not producing independent measurements of four constructs; the near-1 correlations in the example condition (Figure 2 and Figure 11) may reflect response formatting, anchoring on the first rating, or prompt-induced consistency rather than a collapse of the underlying evaluation constructs. Cohen's Kappa, which is computed on the same jointly generated ratings, does not address output independence. The authors' own Future Work section acknowledges that the results 'may have been driven in part by the structure of the prompt,' but this possibility is not tested. Human participants, by contrast, rated originality first and then rated the facets in a separate step after writing explanations, so the human and LLM procedures are not directly comparable. A prompt-ablation design with separate calls per facet, or at least randomized output order, is needed before the claimed trade-off between predictive and construct validity is established.
- [Study 1 Methods (linguistic markers) and Study 2 Results] The linguistic-marker analyses in both studies use GPT-4o and Claude-3.5-Sonnet as annotators for past/future focus, perceptual details, causal/analytical language, comparative language, and cleverness, with no human-annotated gold standard for these specific prompts and categories. The paper cites Rathje et al. (2024) for the general validity of LLM psycholinguistic rating, but that validation does not cover the precise constructs used here (e.g., 'comparative language,' 'causal/analytical markers') nor the short, domain-specific explanations collected in this study. This assumption is load-bearing for the cognitive-process claims—e.g., that no-example experts rely more on memory retrieval because they produce more comparative language. I request either a human-annotation reliability check on a sample of explanations or a clear statement that the cognitive-process interpretation is provisional pending such validation.
- [Study 1 and Study 2 Methods (ground-truth originality)] The criterion for 'true originality scores' is described as factor scores from Patterson et al. (2025), a manuscript under review. The current paper does not describe how these factor scores were derived, how items were selected or deduplicated, or how the ground-truth ratings relate to the newly collected human ratings. Since the conclusion that LLM originality scores have stronger predictive validity depends on this criterion, readers cannot currently verify the accuracy gain from examples. Please provide a description of the factor-scoring procedure and a direct link to the scoring data, or state clearly which parts of the dataset are available.
minor comments (5)
- [Study 1 Results] The text reports 'U = 75076.5, p < 0.5' for Claude-3.5-Sonnet's causal/analytical comparison; this appears to be a typo for p > 0.05 or p = 0.5, and the exact p value should be reported.
- [Throughout] Model naming is inconsistent: the text refers to 'GPT-4 O3' in Study 1 Methods and 'GPT-4O' elsewhere, and 'CLAUDE-3.5-SONNET' alternates with 'Claude-3.5-Sonnet.' Please unify the model names.
- [Figure 7 prompt] The prompt in Figure 7 says 'Casual / analytical markers' while the text uses 'causal/analytical'; the typo could affect prompt interpretation and should be corrected.
- [Figures] Figure numbering is confusing: 'Figure 4' appears to refer to different displays in Study 1 Results and Study 2 Results, and the caption for the linguistic-marker comparison appears late in the supplementary material. Please renumber and cross-check all figure references.
- [Study 2 Results] The paper reports many pairwise Fisher's z tests without a multiple-comparison correction; because the correlation matrix contains ten pairs per model per condition, I recommend reporting corrected p values or noting which comparisons survive correction.
Circularity Check
No significant circularity: the paper's comparisons are externally grounded in human ratings and ground-truth originality scores.
full rationale
The paper's central claims are descriptive comparisons between human and LLM ratings, not derivations in which an output is equivalent to an input by construction. The ground-truth originality scores come from independent expert ratings in Patterson et al. (2025), and the LLM originality ratings are zero-shot predictions correlated against that external benchmark; no parameter is fitted and then renamed as a prediction. The facet-correlation analysis is a direct summary of the collected ratings, and the 'homogenization' claim is an interpretation of those observed correlations rather than a result forced by the measurement procedure. The paper also includes human raters as a separate baseline, so the human-LLM comparison does not reduce to self-citation or to LLM-generated data alone. The most serious concern is methodological rather than circular: the Study 2 prompt (Figure 10) asks the LLM to output ORIGINALITY, UNCOMMON, REMOTE, CLEVER, and EXPLANATION in a single response, which could inflate facet correlations through response anchoring. The paper itself flags this possibility in the Future Work section: 'it remains possible that our results may have been driven in part by the structure of the prompt.' A prompt-format confound is a validity threat, not a circularity: the observed correlations are real outputs, and the paper's interpretation of them as weaker construct validity is an inference that could be wrong without being circular. Reused data and self-citations (e.g., Patterson et al., Orwig et al.) are load-bearing only as sources of stimuli and methods, not as a substitute for the paper's own empirical comparisons. No step in the claimed derivation chain equates a prediction with an input, fits a parameter and then reports it as a finding, or imports a uniqueness claim from the authors' prior work.
Assumptions & free parameters
assumptions (3)
- domain assumption Factor scores from Patterson et al. (2025) are valid ground-truth originality labels for DPT responses.
- domain assumption LLM ratings of linguistic markers (comparative, causal/analytical, perceptual, past/future, cleverness) accurately measure these constructs in explanations.
- domain assumption Written explanations of originality scores reflect the rater's underlying evaluation process.
Cite this review
Pith. "Pith review of How do Humans and Language Models Reason About Creativity? A Comparative Analysis." pith.science (2026). https://pith.science/paper/BFDVFHGO
@misc{pith2026250203253,
author = {Pith},
title = {Pith review of: How do Humans and Language Models Reason About Creativity? A Comparative Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/BFDVFHGO}},
note = {Machine review of arXiv:2502.03253}
}
abstract
Creativity assessment in science and engineering is increasingly based on both human and AI judgment, but the cognitive processes and biases behind these evaluations remain poorly understood. We conducted two experiments examining how including example solutions with ratings impact creativity evaluation, using a finegrained annotation protocol where raters were tasked with explaining their originality scores and rating for the facets of remoteness (whether the response is "far" from everyday ideas), uncommonness (whether the response is rare), and cleverness. In Study 1, we analyzed creativity ratings from 72 experts with formal science or engineering training, comparing those who received example solutions with ratings (example) to those who did not (no example). Computational text analysis revealed that, compared to experts with examples, no-example experts used more comparative language (e.g., "better/worse") and emphasized solution uncommonness, suggesting they may have relied more on memory retrieval for comparisons. In Study 2, parallel analyses with state-of-the-art LLMs revealed that models prioritized uncommonness and remoteness of ideas when rating originality, suggesting an evaluative process rooted around the semantic similarity of ideas. In the example condition, while LLM accuracy in predicting the true originality scores improved, the correlations of remoteness, uncommonness, and cleverness with originality also increased substantially -- to upwards of $0.99$ -- suggesting a homogenization in the LLMs evaluation of the individual facets. These findings highlight important implications for how humans and AI reason about creativity and suggest diverging preferences for what different populations prioritize when rating.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
amabile1982social APACrefauthors Amabile, T M. APACrefauthors \ 1982 . Social psychology of creativity: A consensual assessment technique. Social psychology of creativity: A consensual assessment technique. Journal of personality and social psychology 43 5 997
work page 1982
-
[2]
boiko2023emergent APACrefauthors Boiko, D A. , MacKnight, R. \ Gomes, G. APACrefauthors \ 2023 . Emergent autonomous scientific research capabilities of large language models Emergent autonomous scientific research capabilities of large language models . arXiv preprint arXiv:2304.05332
arXiv 2023
-
[3]
, Mann, B
brown2020language APACrefauthors Brown, T. , Mann, B. , Ryder, N. , Subbiah, M. , Kaplan, J D. , Dhariwal, P. others APACrefauthors \ 2020 . Language models are few-shot learners Language models are few-shot learners . Advances in neural information processing systems 33 1877--1901
2020
-
[4]
chiang2024chatbot APACrefauthors Chiang, W L. , Zheng, L. , Sheng, Y. , Angelopoulos, A N. , Li, T. , Li, D. others APACrefauthors \ 2024 . Chatbot arena: An open platform for evaluating llms by human preference Chatbot arena: An open platform for evaluating llms by human preference . arXiv preprint arXiv:2403.04132
arXiv 2024
-
[5]
Cseh2019 APACrefauthors Cseh, G M. \ Jeffries, K K. APACrefauthors \ 2019 . A scattered CAT: A critical evaluation of the consensual assessment technique for creativity research A scattered cat: A critical evaluation of the consensual assessment technique for creativity research . Psychology of Aesthetics, Creativity, and the Arts 13 2 159--166 . APACrefU...
- [6]
-
[7]
Diedrich2015 APACrefauthors Diedrich, J. , Benedek, M. , Jauk, E. \ Neubauer, A C. APACrefauthors \ 2015 . Are creative ideas novel and useful? Are creative ideas novel and useful? Psychology of Aesthetics, Creativity, and the Arts 9 1 35--40 . APACrefURL https://doi.org/10.1037/a0038688 APACrefURL APACrefDOI doi:10.1037/a0038688 APACrefDOI
-
[8]
gilhooly2007divergent APACrefauthors Gilhooly, K J. , Fioratou, E. , Anthony, S H. \ Wynn, V. APACrefauthors \ 2007 . Divergent thinking: Strategies and executive involvement in generating novel uses for familiar objects Divergent thinking: Strategies and executive involvement in generating novel uses for familiar objects . British Journal of Psychology 9...
work page 2007
Show all 32 references
-
[9]
\ Krenn, M
gu2024interesting APACrefauthors Gu, X. \ Krenn, M. APACrefauthors \ 2024 . Interesting scientific idea generation using knowledge graphs and LLMs: Evaluations with 100 research group leaders Interesting scientific idea generation using knowledge graphs and llms: Evaluations w...
2024 arXiv
-
[10]
, Huang, Y
huang2025large APACrefauthors Huang, S. , Huang, Y. , Liu, Y. , Luo, Z. \ Lu, W. APACrefauthors \ 2025 . Are large language models qualified reviewers in originality evaluation? Are large language models qualified reviewers in originality evaluation? Information Processing & M...
2025
-
[11]
\ Mak \'o , C
illessy2020automation APACrefauthors Ill \'e ssy, M. \ Mak \'o , C. APACrefauthors \ 2020 . Automation and Creativity In Work Automation and creativity in work . Intersections 6 2 112--129
2020
-
[12]
, Peng, Z
lin2024evaluating APACrefauthors Lin, E. , Peng, Z. \ Fang, Y. APACrefauthors \ 2024 . Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications Evaluating and enhancing large language models for novelty assessment in scholarly publication...
2024 arXiv
-
[13]
lu2024ai APACrefauthors Lu, C. , Lu, C. , Lange, R T. , Foerster, J. , Clune, J. \ Ha, D. APACrefauthors \ 2024 . The ai scientist: Towards fully automated open-ended scientific discovery The ai scientist: Towards fully automated open-ended scientific discovery . arXiv preprin...
2024 arXiv
-
[14]
, Maliakkal, N T
luchini2025automated APACrefauthors Luchini, S A. , Maliakkal, N T. , DiStefano, P V. , Laverghetta Jr, A. , Patterson, J D. , Beaty, R E. \ Reiter-Palmon, R. APACrefauthors \ 2025 . Automated scoring of creative problem solving with large language models: A comparison of orig...
2025
-
[15]
, Acar, S
organisciak2023beyond APACrefauthors Organisciak, P. , Acar, S. , Dumas, D. \ Berthiaume, K. APACrefauthors \ 2023 . Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models Beyond semantic distance: Automated scoring of div...
2023
-
[16]
, Beaty, R E
orwig2024creative APACrefauthors Orwig, W. , Beaty, R E. , Benedek, M. \ Schacter, D L. APACrefauthors \ 2024 . Creative Evaluation: The Role of Memory in Novelty & Effectiveness Judgements Creative evaluation: The role of memory in novelty & effectiveness judgements . Creativ...
2024
-
[17]
ouyang2022training APACrefauthors Ouyang, L. , Wu, J. , Jiang, X. , Almeida, D. , Wainwright, C. , Mishkin, P. others APACrefauthors \ 2022 . Training language models to follow instructions with human feedback Training language models to follow instructions with human feedback...
2022
-
[18]
, Bowman, S R
panickssery2024llm APACrefauthors Panickssery, A. , Bowman, S R. \ Feng, S. APACrefauthors \ 2024 . Llm evaluators recognize and favor their own generations Llm evaluators recognize and favor their own generations . arXiv preprint arXiv:2404.13076
2024 arXiv
-
[19]
, Schoenegger, P
park2024diminished APACrefauthors Park, P S. , Schoenegger, P. \ Zhu, C. APACrefauthors \ 2024 . Diminished diversity-of-thought in a standard large language model Diminished diversity-of-thought in a standard large language model . Behavior Research Methods 1--17
2024
-
[20]
, Pronchick, J
Patterson2025 APACrefauthors Patterson, J D. , Pronchick, J. , Panchanadikar, R. , Fuge, M. , van Hell, J G. , Miller, S R. Beaty, R E. APACrefauthors \ 2025 . CAP: The Creativity Assessment Platform for Open Tests and Automated Scoring. Cap: The creativity assessment platform...
2025
-
[21]
, Mirea, D M
rathje2024gpt APACrefauthors Rathje, S. , Mirea, D M. , Sucholutsky, I. , Marjieh, R. , Robertson, C E. \ Van Bavel, J J. APACrefauthors \ 2024 . GPT is an effective tool for multilingual psychological text analysis Gpt is an effective tool for multilingual psychological text ...
2024
-
[22]
schmidgall2025agent APACrefauthors Schmidgall, S. , Su, Y. , Wang, Z. , Sun, X. , Wu, J. , Yu, X. Barsoum, E. APACrefauthors \ 2025 . Agent Laboratory: Using LLM Agents as Research Assistants Agent laboratory: Using llm agents as research assistants . arXiv preprint arXiv:2501.04227
2025 arXiv
-
[23]
schmidt2011creativity APACrefauthors Schmidt, A L. \ . APACrefauthors \ 2011 . Creativity in science: Tensions between perception and practice Creativity in science: Tensions between perception and practice . Creative Education 2 05 435
2011
-
[24]
, Yang, D
si2024can APACrefauthors Si, C. , Yang, D. \ Hashimoto, T. APACrefauthors \ 2024 . Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers . arXiv pr...
2024 arXiv
-
[25]
APACrefauthors \ 2008
silvia2008another APACrefauthors Silvia, P J. APACrefauthors \ 2008 . Another look at creativity and intelligence: Exploring higher-order models and probable confounds Another look at creativity and intelligence: Exploring higher-order models and probable confounds . Personali...
2008
-
[26]
, Winterstein, B P
silvia2008assessing APACrefauthors Silvia, P J. , Winterstein, B P. , Willse, J T. , Barona, C M. , Cram, J T. , Hess, K I. Richard, C A. APACrefauthors \ 2008 . Assessing creativity with divergent thinking tasks: exploring the reliability and validity of new subjective scorin...
2008
-
[27]
APACrefauthors \ 2004
simonton2004creativity APACrefauthors Simonton, D K. APACrefauthors \ 2004 . Creativity in science: Chance, logic, genius, and zeitgeist Creativity in science: Chance, logic, genius, and zeitgeist . Cambridge Univ Pr
2004
-
[28]
, Ward, T B
smith1995creative APACrefauthors Smith, S M. , Ward, T B. \ Finke, R A. APACrefauthors \ 1995 . The creative cognition approach The creative cognition approach . MIT press
1995
-
[29]
\ Pennebaker, J W
tausczik2010psychological APACrefauthors Tausczik, Y R. \ Pennebaker, J W. APACrefauthors \ 2010 . The psychological meaning of words: LIWC and computerized text analysis methods The psychological meaning of words: Liwc and computerized text analysis methods . Journal of langu...
2010
-
[30]
tsegaye2019antecedent APACrefauthors Tsegaye, W. , Su, Q. \ Malik, M. APACrefauthors \ 2019 . The antecedent impact of culture and economic growth on nationscreativity and innovation capability The antecedent impact of culture and economic growth on nationscreativity and innov...
2019
-
[31]
, Yang, S
wang2025behind APACrefauthors Wang, J. , Yang, S. \ Long, H. APACrefauthors \ 2025 . Behind the scores: Unraveling rater judgment in subjective creativity assessments. Behind the scores: Unraveling rater judgment in subjective creativity assessments. Psychology of Aesthetics, ...
2025
-
[32]
, Downey, D
wang-etal-2024-scimon APACrefauthors Wang, Q. , Downey, D. , Ji, H. \ Hope, T. APACrefauthors \ 2024 08 . S ci MON : Scientific Inspiration Machines Optimized for Novelty S ci MON : Scientific inspiration machines optimized for novelty . L W. Ku, A. Martins \ V. Srikumar\ ( ),...
2024
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.