Pith. sign in

REVIEW 4 cited by

AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.04265 v2 pith:LKIKI6HV submitted 2024-10-05 cs.CL

classification cs.CL
keywords creativityindextexthumanllmsauthorsaverageeven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Creativity has long been considered one of the most difficult aspect of human intelligence for AI to mimic. However, the rise of Large Language Models (LLMs), like ChatGPT, has raised questions about whether AI can match or even surpass human creativity. We present CREATIVITY INDEX as the first step to quantify the linguistic creativity of a text by reconstructing it from existing text snippets on the web. CREATIVITY INDEX is motivated by the hypothesis that the seemingly remarkable creativity of LLMs may be attributable in large part to the creativity of human-written texts on the web. To compute CREATIVITY INDEX efficiently, we introduce DJ SEARCH, a novel dynamic programming algorithm that can search verbatim and near-verbatim matches of text snippets from a given document against the web. Experiments reveal that the CREATIVITY INDEX of professional human authors is on average 66.2% higher than that of LLMs, and that alignment reduces the CREATIVITY INDEX of LLMs by an average of 30.1%. In addition, we find that distinguished authors like Hemingway exhibit measurably higher CREATIVITY INDEX compared to other human writers. Finally, we demonstrate that CREATIVITY INDEX can be used as a surprisingly effective criterion for zero-shot machine text detection, surpassing the strongest existing zero-shot system, DetectGPT, by a significant margin of 30.2%, and even outperforming the strongest supervised system, GhostBuster, in five out of six domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing

    cs.CL 2025-06 conditional novelty 7.0 of 10

    Reader evaluations of AI versus human stories split into two measurable preference profiles, surface-focused and holistic, that track reader expertise and explain why prior studies disagree.

  2. GuessBench: Sensemaking Multimodal Creativity in the Wild

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A Minecraft-based benchmark shows vision-language models often fail to decode player-built creations, with accuracy falling sharply for rare concepts and low-resource languages.

  3. Dynamic Reinforcement Learning for Actors

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A reinforcement learning update that adjusts each neuron's input-output sensitivity using TD error can replace external exploration noise and backpropagation through time in small actor-critic tasks.

  4. CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity

    cs.CL 2025-10 conditional novelty 5.0 of 10

    An evaluation of 17 LLMs across nine creativity tasks shows that creativity is fragmented: novelty scores correlate weakly or negatively with quality and diversity, and the proprietary-model advantage largely disappea...

Pith tools