Pith. sign in

REVIEW 6 cited by

A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08433 v1 pith:LTM55UO3 submitted 2023-10-12 cs.CL cs.CY

classification cs.CLcs.CY
keywords llmshumanscoherenceconfederacycreativehumorstylewriting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We evaluate a range of recent LLMs on English creative writing, a challenging and complex task that requires imagination, coherence, and style. We use a difficult, open-ended scenario chosen to avoid training data reuse: an epic narration of a single combat between Ignatius J. Reilly, the protagonist of the Pulitzer Prize-winning novel A Confederacy of Dunces (1980), and a pterodactyl, a prehistoric flying reptile. We ask several LLMs and humans to write such a story and conduct a human evalution involving various criteria such as fluency, coherence, originality, humor, and style. Our results show that some state-of-the-art commercial LLMs match or slightly outperform our writers in most dimensions; whereas open-source LLMs lag behind. Humans retain an edge in creativity, while humor shows a binary divide between LLMs that can handle it comparably to humans and those that fail at it. We discuss the implications and limitations of our study and suggest directions for future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers

    cs.AR 2026-04 unverdicted novelty 6.0 of 10

    Benchmarking on four edge platform configurations shows hardware accelerators improve LLM inference efficiency and reveals trade-offs in power use, device size, and token throughput for constrained deployments.

  2. If You Had to Pitch Your Ideal Software -- Evaluating Large Language Models to Support User Scenario Writing for User Experience Experts and Laypersons

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Laypeople using an LLM writing assistant produced user scenarios rated as high in structure and clarity as those written by UX experts.

  3. Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection

    cs.CL 2025-08 conditional novelty 5.0 of 10

    ROSI bakes the refusal direction into a model's weight matrices via a rank-one update, raising refusal and jailbreak robustness with minimal measured utility cost.

  4. Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    A generative AI framework that reads plant structure-function literature, generates hypotheses, and produces a lab-validated pollen-based adhesive with measured shear strength.

  5. GPTFootprint: Increasing Consumer Awareness of the Environmental Impacts of LLMs

    cs.HC 2025-05 conditional novelty 5.0 of 10

    An eco-feedback browser extension for ChatGPT raises user awareness of energy and water use, but a nine-participant study finds limited effect on query frequency.

  6. A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations

    cs.CL 2025-07 conditional novelty 4.0 of 10

    PATTR adds a target-length penalty to the Type-Token Ratio, producing a lexical diversity score with tunable, reduced short-text bias for LLM synthetic data.

Pith tools