Pith. sign in

REVIEW 18 cited by

Does Writing with Language Models Reduce Content Diversity?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05196 v3 pith:RLY4B2GA submitted 2023-09-11 cs.CL cs.CYcs.HCcs.LG

Does Writing with Language Models Reduce Content Diversity?

classification cs.CL cs.CYcs.HCcs.LG
keywords diversitycontentmodelwritingdiverseinstructgptmodelsdifferent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models (LLMs) have led to a surge in collaborative writing with model assistance. As different users incorporate suggestions from the same model, there is a risk of decreased diversity in the produced content, potentially limiting diverse perspectives in public discourse. In this work, we measure the impact of co-writing on diversity via a controlled experiment, where users write argumentative essays in three setups -- using a base LLM (GPT3), a feedback-tuned LLM (InstructGPT), and writing without model help. We develop a set of diversity metrics and find that writing with InstructGPT (but not the GPT3) results in a statistically significant reduction in diversity. Specifically, it increases the similarity between the writings of different authors and reduces the overall lexical and content diversity. We additionally find that this effect is mainly attributable to InstructGPT contributing less diverse text to co-written essays. In contrast, the user-contributed text remains unaffected by model collaboration. This suggests that the recent improvement in generation quality from adapting models to human feedback might come at the cost of more homogeneous and less diverse content.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Language Models Agree With Each Other, Not With Readers

    cs.IR 2026-07 accept novelty 7.0

    Across 18 model arms, model-model excess agreement (+0.093 median) is 2.3x human-human agreement (+0.040), against a naturalistic uninstructed reader baseline.

  2. Four Types of LLM Reliance and Their Predictors Among Undergraduate Writers: A Mixed-Methods Study at a Minority-Serving R1 University

    cs.CY 2026-06 unverdicted novelty 7.0

    Four types of LLM reliance (Strategic 34.3%, Instrumental 30.9%, Dialogic 30.4%, Dependent 4.5%) were identified among undergraduates, with AI literacy predicting type and value/cost beliefs predicting intensity.

  3. Four Types of LLM Reliance and Their Predictors Among Undergraduate Writers: A Mixed-Methods Study at a Minority-Serving R1 University

    cs.CY 2026-06 unverdicted novelty 7.0

    The paper classifies undergraduate LLM reliance into four types and shows that AI literacy predicts type while value beliefs predict intensity, with implications for outcome measurement.

  4. Before and After Temperature: A Distributional View of Creative LLM Generation

    cs.CL 2026-05 unverdicted novelty 7.0

    A per-token feature from temperature-induced changes in LLM token distributions predicts within-prompt creativity rank at Spearman rho 0.918 vs LLM judges and 0.870 vs humans, outperforming perplexity, entropy, top-1 ...

  5. More Is Not More: What Matters for Diversity in LLM Opinions?

    cs.CL 2026-05 conditional novelty 7.0

    Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.

  6. The One-Word Census: Answer-Choice Conformity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0

    Across 31 open one-word categories, 44 LMs converge extremely (often >80% on one answer), with newest flagships most conformist and persona-tuned models most divergent.

  7. The One-Word Census: Answer-Choice Conformity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0

    Forty-four language models asked to name one thing per category converge on the same modal answers far more than people do, with newest flagships most conformist and persona-tuned models most divergent.

  8. A framework for single and multi-agent human-AI curiosity ecosystems

    cs.AI 2026-07 conditional novelty 6.0

    A single- and multi-agent framework models curiosity as a weighted value of immediate uncertainty reduction, cost, delayed return, and open-question value, with weights that drift with experience and ecology.

  9. "I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise

    cs.CL 2026-06 unverdicted novelty 6.0

    Decan (D_Ca_n = C × a_n) measures text diversity as progressive conditional surprise from base LM log-probabilities, scoring 0.846 OCA on McDiv benchmark and detecting monotonic diversity drop across base→SFT→DPO→RLVR stages.

  10. AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing

    cs.CL 2026-05 unverdicted novelty 6.0

    Analysis of news text in 34 languages shows cross-lingual convergence on AI-associated lemmas and increased prevalence of top AI-overused items after ChatGPT's release.

  11. Annotations Mitigate Post-Training Mode Collapse

    cs.CL 2026-05 unverdicted novelty 6.0

    Annotation-anchored training reduces semantic diversity collapse in post-trained language models by a factor of six compared to standard supervised fine-tuning while preserving instruction-following and improving with scale.

  12. Human Thinking under Plural LLM Assistance: Mathematical Problem Solving and Open-Ended Writing

    cs.HC 2026-04 unverdicted novelty 6.0

    Two controlled experiments show multi-agent LLM configurations with both tutors and peers deliver higher learning gains and less homogeneous outputs than single-LLM tutoring in math problem-solving and essay writing.

  13. Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders

    cs.CL 2026-02 conditional novelty 6.0

    Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).

  14. Value Drifts: Tracing Value Alignment During LLM Post-Training

    cs.CL 2025-10 conditional novelty 6.0

    Value alignment in LLMs is set largely during supervised fine-tuning; standard preference-optimization datasets carry too little stance contrast to re-align it, but with engineered contrast algorithms differ (DPO ampl...

  15. Double-Edged Sword or Sharp Tool? Designing and Evaluating Triadic LLM-Teacher Collaboration for K-12 Writing at Scale

    cs.AI 2026-05 unverdicted novelty 5.0

    A two-year deployment across 120 schools shows that LLM-teacher collaboration improves K-12 writing quality via labor division, with a ceiling effect from excessive LLM linguistic expansion.

  16. Human Thinking under Plural LLM Assistance: Mathematical Problem Solving and Open-Ended Writing

    cs.HC 2026-04 unverdicted novelty 5.0

    Plural LLM setups (expert+peer in math; role-specialized pair in writing) improve post-task math performance and preserve writing idea diversity better than single-assistant or no-AI baselines.

  17. A framework for single and multi-agent human-AI curiosity ecosystems

    cs.AI 2026-07 conditional novelty 4.0

    A toy framework models curiosity as an ecosystem where agents' inquiry weights drift with experience and shared knowledge stocks shape collective discovery.

  18. Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

    cs.CL 2026-02 conditional novelty 4.0

    Quality-constrained entropy maximization yields simple DPO-like objectives that increase LLM output diversity while preserving or slightly improving quality, with theoretical guarantees under tuned temperature conditions.