Pith. sign in

REVIEW 9 cited by

We're Different, We're the Same: Creative Homogeneity Across LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.19361 v1 pith:3BR2D3GO submitted 2025-01-31 cs.CY cs.AIcs.CLcs.LG

We're Different, We're the Same: Creative Homogeneity Across LLMs

classification cs.CY cs.AIcs.CLcs.LG
keywords creativellmsresponsesoutputscreativityotheracrossassistants
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Numerous powerful large language models (LLMs) are now available for use as writing support tools, idea generators, and beyond. Although these LLMs are marketed as helpful creative assistants, several works have shown that using an LLM as a creative partner results in a narrower set of creative outputs. However, these studies only consider the effects of interacting with a single LLM, begging the question of whether such narrowed creativity stems from using a particular LLM -- which arguably has a limited range of outputs -- or from using LLMs in general as creative assistants. To study this question, we elicit creative responses from humans and a broad set of LLMs using standardized creativity tests and compare the population-level diversity of responses. We find that LLM responses are much more similar to other LLM responses than human responses are to each other, even after controlling for response structure and other key variables. This finding of significant homogeneity in creative outputs across the LLMs we evaluate adds a new dimension to the ongoing conversation about creativity and LLMs. If today's LLMs behave similarly, using them as a creative partners -- regardless of the model used -- may drive all users towards a limited set of "creative" outputs.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios

    cs.CR 2026-05 unverdicted novelty 7.0

    LLM agents exhibit persistent attack-selection biases as fixed traits independent of success rates, with a bias momentum effect that resists steering and yields no performance gain.

  2. Ex Ante Evaluation of AI-Induced Idea Diversity Collapse

    cs.AI 2026-05 unverdicted novelty 7.0

    Frontier LLMs generate creative ideas with excess population-level crowding below human-relative parity across tasks, but targeted generation protocols can reduce it.

  3. Large Language Models Align with the Human Brain during Creative Thinking

    q-bio.NC 2026-04 unverdicted novelty 7.0

    LLMs show scaling and training-dependent alignment with human brain responses in creativity-related networks during divergent thinking tasks, measured via RSA on fMRI data.

  4. Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework

    cs.CL 2025-09 conditional novelty 7.0

    Proposes a task taxonomy for functional diversity in LLM outputs, validates it via user study, introduces targeted sampling to boost diversity only where needed, and presents evidence that the diversity-quality tradeo...

  5. The One-Word Census: Answer-Choice Conformity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0

    Forty-four language models asked to name one thing per category converge on the same modal answers far more than people do, with newest flagships most conformist and persona-tuned models most divergent.

  6. The One-Word Census: Answer-Choice Conformity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0

    Across 31 open one-word categories, 44 LMs converge extremely (often >80% on one answer), with newest flagships most conformist and persona-tuned models most divergent.

  7. Optimization Is Not All You Need

    cs.AI 2026-07 unverdicted novelty 6.0

    A critical essay argues that LLM alignment enforces linguistic norms through scalar optimization while lacking the interpretive capacity to distinguish error from invention, narrowing what generated language can be.

  8. Introducing multiplex semantic networks as multifaceted representations of creative associative knowledge across multilingual samples

    cs.CL 2026-06 unverdicted novelty 6.0

    Multiplex semantic networks assembled from verbal fluency, free association, sentence-chain and narrative tasks capture non-redundant aspects of semantic organization and improve ridge-regression prediction of individ...

  9. Optimization Is Not All You Need

    cs.AI 2026-07 unverdicted novelty 5.0

    Optimization can measure how improbable generated text is but cannot tell whether that unlikelihood is error or invention, yet it now sets the protocols of legitimate language.