Pith. sign in

REVIEW 5 cited by

Evaluating Creative Short Story Generation in Humans and Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.02316 v5 pith:X2R7W6K2 submitted 2024-11-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords creativityllmsstoriescreativehumanratingsshortstory-writing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Story-writing is a fundamental aspect of human imagination, relying heavily on creativity to produce narratives that are novel, effective, and surprising. While large language models (LLMs) have demonstrated the ability to generate high-quality stories, their creative story-writing capabilities remain under-explored. In this work, we conduct a systematic analysis of creativity in short story generation across 60 LLMs and 60 people using a five-sentence cue-word-based creative story-writing task. We use measures to automatically evaluate model- and human-generated stories across several dimensions of creativity, including novelty, surprise, diversity, and linguistic complexity. We also collect creativity ratings and Turing Test classifications from non-expert and expert human raters and LLMs. Automated metrics show that LLMs generate stylistically complex stories, but tend to fall short in terms of novelty, surprise and diversity when compared to average human writers. Expert ratings generally coincide with automated metrics. However, LLMs and non-experts rate LLM stories to be more creative than human-generated stories. We discuss why and how these differences in ratings occur, and their implications for both human and artificial creativity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    LLM outputs are usually plausible but cover far less of the human response space than matched human samples do, with the biggest gap at the periphery of the human distribution.

  2. Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis

    cs.HC 2025-05 conditional novelty 6.0 of 10

    A meta-analysis of 28 studies finds no average creativity gap between GenAI and humans, a small boost when humans collaborate with GenAI, and a large drop in idea diversity in those collaborations.

  3. CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity

    cs.CL 2025-10 conditional novelty 5.0 of 10

    An evaluation of 17 LLMs across nine creativity tasks shows that creativity is fragmented: novelty scores correlate weakly or negatively with quality and diversity, and the proprietary-model advantage largely disappea...

  4. IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations

    cs.CL 2025-09 conditional novelty 5.0 of 10

    IDEAlign uses a pick-the-odd-one-out triplet task to measure idea-level similarity between LLMs and expert human annotations, and shows LLM judges using this protocol outperform lexical and vector-based baselines.

  5. Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An integrated storytelling framework that generates text, scene graphs, images, and sound together reports higher expert-rated coherence than a sequential baseline, but the evaluation is small and qualitative.

Pith tools