Pith. sign in

REVIEW 15 cited by

Assessing and Understanding Creativity in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12491 v1 pith:YXG2TLFX submitted 2024-01-23 cs.CL cs.AI

classification cs.CLcs.AI
keywords creativityllmsassessinglanguageoriginalitycreativeelaborationfindings
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In the field of natural language processing, the rapid development of large language model (LLM) has attracted more and more attention. LLMs have shown a high level of creativity in various tasks, but the methods for assessing such creativity are inadequate. The assessment of LLM creativity needs to consider differences from humans, requiring multi-dimensional measurement while balancing accuracy and efficiency. This paper aims to establish an efficient framework for assessing the level of creativity in LLMs. By adapting the modified Torrance Tests of Creative Thinking, the research evaluates the creative performance of various LLMs across 7 tasks, emphasizing 4 criteria including Fluency, Flexibility, Originality, and Elaboration. In this context, we develop a comprehensive dataset of 700 questions for testing and an LLM-based evaluation method. In addition, this study presents a novel analysis of LLMs' responses to diverse prompts and role-play situations. We found that the creativity of LLMs primarily falls short in originality, while excelling in elaboration. Besides, the use of prompts and the role-play settings of the model significantly influence creativity. Additionally, the experimental results also indicate that collaboration among multiple LLMs can enhance originality. Notably, our findings reveal a consensus between human evaluations and LLMs regarding the personality traits that influence creativity. The findings underscore the significant impact of LLM design on creativity and bridges artificial intelligence and human creativity, offering insights into LLMs' creativity and potential applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    LLM outputs are usually plausible but cover far less of the human response space than matched human samples do, with the biggest gap at the periphery of the human distribution.

  2. Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction

    cs.LG 2025-04 conditional novelty 6.0 of 10

    On four minimal graph-construction tasks, multi-token training (teacherless or diffusion) produces more diverse and original outputs than next-token training, and random seed prefixes can replace temperature as a dive...

  3. Dynamic Reinforcement Learning for Actors

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A reinforcement learning update that adjusts each neuron's input-output sensitivity using TD error can replace external exploration noise and backpropagation through time in small actor-critic tasks.

  4. Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A fine-tuned 14B LLM judge, trained with scenario-based prompts and controlled instruction generation, approaches GPT-4's human-agreement performance, and the paper documents why scaling distillation data can fail.

  5. We're Different, We're the Same: Creative Homogeneity Across LLMs

    cs.CY 2025-01 conditional novelty 6.0 of 10

    Across three divergent-thinking tests, responses from seven LLM families were substantially more similar to one another than responses from 102 humans were to one another.

  6. Can Hallucinations Help? Boosting LLMs for Drug Discovery

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Adding hallucinated molecule descriptions to prompts improves ROC-AUC for several LLMs on molecular property prediction, with GPT-4o-generated text giving the largest consistent gains.

  7. EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new escape-room benchmark measures creative reasoning in language-model agents, and the EscapeAgent framework (Foresight plus Reflection) reduces hints and steps by up to 40%.

  8. Evaluating Creativity and Deception in Large Language Models: A Simulation Framework for Multi-Agent Balderdash

    cs.MA 2024-11 conditional novelty 6.0 of 10

    A Balderdash simulation framework shows LLMs generate plausible fake definitions but fail to reason over game rules or adapt strategy, with the effect strongest on rare words.

  9. Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    A generative AI framework that reads plant structure-function literature, generates hypotheses, and produces a lab-validated pollen-based adhesive with measured shear strength.

  10. Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach

    cs.CL 2025-04 conditional novelty 5.0 of 10

    A reference-based Likert scoring method with analyze-rate prompting improves LLM judges' agreement with human creativity rankings on the TTCW benchmark, though the headline result is partly fitted to the test set.

  11. PerPO: Perceptual Preference Optimization via Discriminative Rewarding

    cs.AI 2025-02 conditional novelty 5.0 of 10

    PerPO trains multimodal LLMs by ranking their candidate answers with deterministic visual rewards (IoU, edit distance) and using the reward differences as margins in listwise preference optimization.

  12. THiNK: Can Large Language Models Think-aloud?

    cs.CL 2025-05 reject novelty 4.0 of 10

    THiNK uses a multi-agent, feedback-driven loop of problem revision and GPT-4O-based Bloom's Taxonomy scoring to measure and improve higher-order thinking in LLMs on math word problems.

  13. AI Awareness

    cs.AI 2025-04 accept novelty 4.0 of 10

    A review arguing that AI awareness is a measurable, four-dimensional functional capacity (metacognition, self, social, situational) that current LLMs partially exhibit and that both improves AI and creates safety risks.

  14. Do LLMs Agree on the Creativity Evaluation of Alternative Uses?

    cs.AI 2024-11 conditional novelty 4.0 of 10

    Four LLMs show high agreement when scoring and ranking alternative uses and do not favor their own outputs, but the accuracy benchmark is derived from the generation prompts rather than human judgment.

  15. Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery

    cs.AI 2025-05 accept novelty 3.0 of 10

    A perspective review argues that LLMs should be deeply integrated into all stages of science, with human oversight and clear metrics, to become creative engines.

Pith tools