Pith. sign in

REVIEW 19 cited by

Does Writing with Language Models Reduce Content Diversity?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05196 v3 pith:RLY4B2GA submitted 2023-09-11 cs.CL cs.CYcs.HCcs.LG

classification cs.CLcs.CYcs.HCcs.LG
keywords diversitycontentmodelwritingdiverseinstructgptmodelsdifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have led to a surge in collaborative writing with model assistance. As different users incorporate suggestions from the same model, there is a risk of decreased diversity in the produced content, potentially limiting diverse perspectives in public discourse. In this work, we measure the impact of co-writing on diversity via a controlled experiment, where users write argumentative essays in three setups -- using a base LLM (GPT3), a feedback-tuned LLM (InstructGPT), and writing without model help. We develop a set of diversity metrics and find that writing with InstructGPT (but not the GPT3) results in a statistically significant reduction in diversity. Specifically, it increases the similarity between the writings of different authors and reduces the overall lexical and content diversity. We additionally find that this effect is mainly attributable to InstructGPT contributing less diverse text to co-written essays. In contrast, the user-contributed text remains unaffected by model collaboration. This suggests that the recent improvement in generation quality from adapting models to human feedback might come at the cost of more homogeneous and less diverse content.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language Models Agree With Each Other, Not With Readers

    cs.IR 2026-07 accept novelty 7.0 of 10

    Across 18 model arms, model-model excess agreement (+0.093 median) is 2.3x human-human agreement (+0.040), against a naturalistic uninstructed reader baseline.

  2. More Is Not More: What Matters for Diversity in LLM Opinions?

    cs.CL 2026-05 conditional novelty 7.0 of 10

    Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.

  3. Measuring Non-Adversarial Reproduction of Training Data in Large Language Models

    cs.CL 2024-11 conditional novelty 7.0 of 10

    In non-adversarial settings, popular LLMs reproduce 7-15% of characters from online sources on average, versus far less for humans; worst-case generations can match 100% of their content verbatim.

  4. The One-Word Census: Answer-Choice Conformity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Forty-four language models asked to name one thing per category converge on the same modal answers far more than people do, with newest flagships most conformist and persona-tuned models most divergent.

  5. A framework for single and multi-agent human-AI curiosity ecosystems

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A toy framework models curiosity as an ecosystem where agents' inquiry weights drift with experience and shared knowledge stocks shape collective discovery.

  6. Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).

  7. Value Drifts: Tracing Value Alignment During LLM Post-Training

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Value alignment in LLMs is set largely during supervised fine-tuning; standard preference-optimization datasets carry too little stance contrast to re-align it, but with engineered contrast algorithms differ (DPO ampl...

  8. The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices

    cs.HC 2025-09 conditional novelty 6.0 of 10

    LLM-based agents asked to choose among a person's own Facebook likes select more popular and less diverse pages, reducing both interpersonal distinctiveness and intrapersonal diversity.

  9. The Anatomy of Speech Persuasion: Linguistic Shifts in LLM-Modified Speeches

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GPT-4o increases emotional lexicon and uses more questions and exclamations when asked to strengthen speeches, but follows a surface style rather than human-like persuasive argumentation.

  10. Measuring Diversity in Synthetic Datasets

    cs.CL 2025-02 conditional novelty 6.0 of 10

    DCScore measures dataset diversity as the sum of self-classification probabilities under a softmax similarity matrix, and the paper shows it tracks generation temperature, human judgment, and LLM rankings.

  11. Content-Driven Local Response: Supporting Sentence-Level and Message-Level Mobile Email Replies With and Without AI

    cs.HC 2025-02 conditional novelty 6.0 of 10

    Content-Driven Local Response, a mobile email UI that attaches optional sentence-level and message-level AI support to the incoming message, reduced typing and errors while offering flexible AI involvement in a 126-us...

  12. Understanding Design Fixation in Generative AI

    cs.HC 2025-02 conditional novelty 6.0 of 10

    Generative AI models exhibit a design fixation phenomenon that limits the diversity and originality of their design outputs, according to a small lab study and a proposed theoretical framework.

  13. Redefining Research Crowdsourcing: Incorporating Human Feedback with LLM-Powered Digital Twins

    cs.HC 2025-05 conditional novelty 5.0 of 10

    A study of an LLM-powered 'digital twin' system for crowd workers shows modest accuracy on Likert-scale surveys, with caveats around threshold tuning and evaluation contamination.

  14. Can Generative Agent-Based Modeling Replicate the Friendship Paradox in Social Media Simulations?

    cs.SI 2025-02 conditional novelty 5.0 of 10

    LLM-driven agent simulations of social media reproduce the Friendship Paradox and its variants, driven mainly by infrequent connections to highly popular agents.

  15. One world, one opinion? The superstar effect in LLM responses

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Across ten languages, LLMs consistently name a small set of figures such as Einstein, Shakespeare, and Turing for each profession, revealing a 'superstar effect' that may narrow cultural representation.

  16. Benchmarking Linguistic Diversity of Large Language Models

    cs.CL 2024-12 conditional novelty 5.0 of 10

    State-of-the-art LLMs generate less linguistically diverse text than humans on creative tasks, and several training and deployment choices systematically shift lexical, syntactic, and semantic diversity.

  17. Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

    cs.CL 2026-02 conditional novelty 4.0 of 10

    Quality-constrained entropy maximization yields simple DPO-like objectives that increase LLM output diversity while preserving or slightly improving quality, with theoretical guarantees under tuned temperature conditions.

  18. A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations

    cs.CL 2025-07 conditional novelty 4.0 of 10

    PATTR adds a target-length penalty to the Type-Token Ratio, producing a lexical diversity score with tunable, reduced short-text bias for LLM synthetic data.

  19. Analysis of LLMs vs Human Experts in Requirements Engineering

    cs.SE 2025-01 reject novelty 4.0 of 10

    In a 50-participant study, LLM-generated requirements were rated more aligned on average than human expert documents, but the design gave the LLM a higher minimum feature count.

Pith tools