Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Language models converge on the same one-word answers: across 44 models, the single most common unconstrained pick wins 41% of the time, and per-model conformity spans 1.05 to 3.21 bits, with the newest flagships the most conformist.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 06:18 UTC pith:YQN6KZ6Q

load-bearing objection A transparent, cheap, per-model measurement of answer-choice convergence that mostly holds up; the flagship-conformist ranking is the one result that should stay conditional until provider serving is logged or controlled. the 3 major comments →

arxiv 2607.12796 v2 pith:YQN6KZ6Q submitted 2026-07-14 cs.CL cs.AIcs.CY

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

classification cs.CL cs.AIcs.CY
keywords answer-choice surprisalLLM mode collapseone-word censuscategory production normsmodel conformitypost-training diversityrunner-up consensusprompting for divergence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that large language models do not merely prefer similar answers; they collapse onto an extremely narrow set of choices, and the degree of collapse is a stable, per-release trait that a cheap public battery can measure. The One-Word Census asks 31 simple questions—'Name a tree,' 'Pick a word'—four times to each of 44 models and scores every model by how unlikely its answers are under the pooled answers of all other models (leave-one-out surprisal, in bits). The paper reports that in 7 of 31 categories a single answer takes over 80% of all responses, that persona-tuned and lightly post-trained models are the most divergent, and that the newest mainline flagships produce almost no answer no other model gave. It also finds the divergent tail is itself convergent—models avoiding the modal answer land on the same runner-up—and that the field is more concentrated than human first responses in 18 of 20 shared categories. If the results stand, answer-space diversity becomes a measurable property of a release, one that capability benchmarks cannot see.

Core claim

The paper's central discovery is that LLM answer-choice conformity is extreme, structured, and set per release. With no category constraint, 44 models collectively chose 'serendipity' for 41% of answers; within categories, oak, hammer, rose, carrot, basil, cheddar, and salmon each capture 82–94% of the pool. Yet conformity varies more than fourfold, from 1.05 to 3.21 bits, and the variation tracks post-training: divergent models are persona-tuned, lightly post-trained, or retrieval-grounded, while the newest flagships sit at the conformist floor with novel rates of 0%. Within four lineages conformity rises with each generation, but the latest flagship Claude and GPT models reverse, suggestin

What carries the argument

The instrument is the One-Word Census: 31 frozen single-turn prompts, each naming a category with many valid one-word answers ('Name a tree. Reply with one word only.') plus an unconstrained 'Pick a word.' prompt, run four times per model with no system prompt and requested temperature 1.0. The core metric is answer-choice surprisal, the average -log2 probability of a model's answers under the add-one-smoothed leave-one-out pooled answers of all other models; it is reported in bits, with companions modal avoidance, novel rate, and self-distinctness. Exact match on normalized final tokens makes the metric mechanical and verbosity-immune; a greedy re-run at temperature 0 serves as a temperatur

Load-bearing premise

The central measurement assumes that the answers a model returns through a public serving aggregator at requested temperature 1.0 are the model's own choice behavior; the main run did not log which provider served each call, and providers do not uniformly honor temperature, so part of the newest-flagship conformity could be a serving artifact.

What would settle it

Re-run the same battery with all models served by a single provider that demonstrably honors requested temperature, and for open-weight models also compute exact output-logit distributions; if the newest flagships then spread as widely as persona-tuned models, or if the ranking reverses, the claim that flagship conformity is a property of the models fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If true, answer-space conformity can be tracked release-by-release at about a dollar per model, turning it into a public, contested property of deployments rather than an unexamined side effect.
  • A person consulting several different assistants is drawing samples from nearly the same distribution; agreement between models is not independent confirmation.
  • Prompting for unusual answers does not release the distribution; it only navigates to a new conditional mode, so the collapse lives in the weights, not in the prompt.
  • The newest flagships' near-zero novel rates mean capability benchmarks cannot detect the collapse; a separate conformity score is needed.
  • Generational trajectories, including the two labs' premium-tier reversals, suggest the degree of collapse is chosen somewhere in post-training, not fixed by lab, scale, or lineage.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: running the same frozen battery on local open-weight models with direct output-logit access would settle whether the flagship-conformist tail is a property of the weights or of the serving configuration.
  • Beyond the paper: the comparison to human category-production norms suggests an alignment target—calibrating model answer distributions toward human breadth would make multi-assistant use less redundant.
  • Beyond the paper: if the runner-up consensus persists across releases, future synthetic-data training may collapse onto the runner-up mode after the primary mode is saturated, creating a second attractor.
  • Beyond the paper: the 'heirloom model' association, though retrospective, implies divergence metrics could be tested as predictors of community attachment at deprecation time.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces the 'One-Word Census': 31 single-turn, one-word category prompts (e.g., 'Name a tree. Reply with one word only.') administered four times each to 44 language models through OpenRouter at requested temperature 1.0. Each model is scored by 'answer-choice surprisal' — the average add-one-smoothed, leave-one-out surprisal of its answers under the pooled answers of all other models. The headline findings are: extreme field-level convergence (modal answers exceed 80% in 7 of 31 categories; 'serendipity' takes 41% of unconstrained answers); a structured per-model scorecard spanning 1.05–3.21 bits, with persona/community-tuned models most divergent and the newest mainline flagships (Claude Sonnet 5, Opus 4.8, Grok 4.5, GPT-5) most conformist; generational declines within four lineages with a reversal at the latest Claude/GPT flagships; a 'runner-up consensus' (off-modal answers also concentrate, e.g., mustard for 95% of non-ketchup condiment answers); zero surviving pairwise affinities after conditioning on a one-dimensional depth propensity; and greater concentration than human category-production norms in 18 of 20 shared categories. The instrument is exact-match, cheap, and fully reproducible with public prompts, transcripts, and code.

Significance. If the scorecard is accepted as measuring model weights rather than serving artifacts, this is a significant contribution: it provides a mechanical, inexpensive, per-release instrument for tracking answer-space conformity, with unusually careful robustness work — leave-one-family-out (rho = 0.985), balanced and era-stratified reference fields, smoothing-constant sensitivity (rho >= 0.99), a greedy-decoding rerun, a same-provider probe, and a public repository with transcripts. The structural findings (runner-up consensus, zero residual pairwise affinity, the human-population comparison) are interesting in their own right and would survive many of the serving concerns. The central risk is that the per-model scorecard and the flagship-conformist tail may reflect the serving channel's effective temperature rather than the weights themselves; the paper itself concedes this opacity in §6, and the temp-0 control does not fully close the gap.

major comments (3)
  1. [§4.4, §6, Table 1] The temperature-0 control does not resolve the serving-opacity confound for the flagship-conformist tail. The manuscript concedes (§6) that requested temperature 'is not honored uniformly across providers' and that the main run did not log which provider served each call. For the 13 models with unchanged self-distinctness between temp-1 and temp-0 (including Sonnet 5, Opus 4.8, Grok 4.5, GPT-5), the greedy rerun is a replication, not a control: if the endpoint was already served at low effective temperature in the main run, exactly this signature — low self-distinctness, unchanged under greedy decoding, low surprisal — is expected. The text's claim that 'their low self-distinctness bounds the possible inflation' presupposes that the main run actually sampled near temperature 1.0, which is the premise at issue. The same-provider probe covers only the DeepSeek pair, not the flagship rows.
  2. [§4.3, Table 1] The generational 'reversal' at GPT-5.6 (Luna 1.52 < Terra 1.86 < Sol 2.02) and the Fable/Sonnet-5 dissociation are presented as central structural findings. However, the three GPT-5.6 tiers are served as separate endpoints; without provider/effective-temperature evidence we cannot exclude that the Sol/Terra/Luna ordering is partly a hosting-temperature ordering rather than a weight ordering. The same concern applies to Fable 5 versus Sonnet 5, although same-lab hosting makes that dissociation more plausibly intrinsic. Please report per-endpoint provider and effective-temperature evidence, or explicitly reframe these results as properties of 'the model-as-served' throughout the paper rather than only in §6.
  3. [§4.5] The zero-pairwise-affinity result is a numbered contribution but depends on conditioning on a 'depth propensity' defined as the fraction of distinct off-modal answers landing on the field's #2–#3 answers. This window is a free parameter and no sensitivity analysis is reported. Please state whether the zero-survivor result persists for other reasonable definitions (e.g., #2–#4, or a continuous rank-based depth measure), and report the level-3 test under those alternatives. This is not the paper's most load-bearing claim, but it is one of the paper's stated contributions and should be robust to the chosen window.
minor comments (5)
  1. [Abstract, §4.5] There are several missing spaces in the text (e.g., 'choseserendipity41%', '0of946 pairs'). These are rendering artifacts but should be fixed.
  2. [Figure 3] The seven-family figure with two shared axes is dense; consider separating into small multiples or adding direct labels so that the 'blue walk' and amber points are readable without the caption.
  3. [§4.6] The human comparison is appropriately hedged as US undergraduates from 2004. It would strengthen the paper to add an explicit statement of what a broader human sample would be expected to do, since the current text mentions this only in passing.
  4. [§5] The 'heirloom models' section is explicitly retrospective and outcome-selected, which the paper acknowledges. Consider moving it to the Discussion or marking it more clearly as a speculative correlational observation, since its placement among Results may give it undeserved evidentiary weight.
  5. [§3.2] The 'Mustard Quotient' nickname is memorable but not defined until later; define it at first use or move it to a footnote.

Circularity Check

0 steps flagged

No load-bearing circularity; the panel-relative surprisal metric is conceded and robustly stress-tested.

full rationale

The paper is a measurement study, not a first-principles derivation, and its central quantity is transparently panel-relative: Eq. (1) defines answer-choice surprisal against the pooled leave-one-out answers of all other models, and §4.4 concedes 'absolute bit values are relative to this field.' This is a genuine self-referential element, but it is not load-bearing circularity. The scorecard is not used to predict the field; the headline structural claims (rankings, generational trajectories, runner-up consensus, human-norm comparison) are tested against roster composition (LOFO ρ=0.985; era-stratified ρ=0.992), greedy re-scoring, same-provider probes, and external human norms. No fitted parameter is renamed as a prediction, and no load-bearing result rests on a self-citation: the paper contains no author self-citations and invokes no imported uniqueness theorem. The main validity concern—serving-temperature opacity, including the claim that low self-distinctness 'bounds the possible inflation'—is a limitation of inference about effective temperature, not a circular derivation from inputs. Thus the paper's own concessions and robustness checks leave its central contribution intact.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

No entities are postulated (no new forces, mediators, or conserved quantities). The 'Mustard Quotient' is a nickname for the surprisal measure and 'heirloom model' is a framing label, not an entity. The load-bearing assumptions are the serving-channel/sampling behavior, the panel-as-field proxy, the human-norm benchmark, and the extraction rule; the paper tests or discloses all of them. The two hand-set quantities with any influence on headline claims are the depth-propensity window (untested robustness) and the smoothing constant (tested robust).

free parameters (2)
  • Depth-propensity window (#2–#3 runner-up answers) = field's 2nd–3rd ranked answers
    §4.5: the one-dimensional 'depth' trait is defined as the share of a model's off-modal answers landing on the field's #2–#3 answers. This hand-set window is the conditioning variable that dissolves all 946 pairwise affinities; its robustness to alternative windows (e.g., #2 alone, top-4) is not reported.
  • Add-one smoothing constant in surprisal = +1
    §3.2 Eq. (1): hand-chosen smoothing to keep never-seen answers finite. Not load-bearing because sensitivity is tested (rho >= 0.99 under add-0.5/add-2/add-0.1, Appendix A).
axioms (5)
  • domain assumption Requested temperature 1.0 approximates the sampling the model-as-served actually uses
    §3.1 sets requested temperature 1.0 for every model; §6 and §4.4 concede providers do not honor it uniformly and the main run did not log the serving provider. Self-distinctness is used as a proxy and the greedy re-run partially controls for this, but the true temp-1 distribution is not measured.
  • domain assumption The 44-model OpenRouter availability panel is an adequate proxy for 'the field'
    §3.5/§4.4: scoring is leave-one-out against this panel, so values are panel-relative by construction. Robustness checks (LOFO rho = 0.985, balanced fields, era-stratified fields) support ranking stability, but absolute bit values and 'the field' claims inherit the availability-sample assumption, as the paper acknowledges.
  • domain assumption Van Overschelde (2004) US-undergraduate first-response norms are a valid human benchmark
    §4.6 compares 20 categories against the VO norms. The paper itself notes these are US undergraduates circa 2004 and that a broader population would likely spread further, which would widen rather than narrow the observed concentration gap.
  • domain assumption Final-word token extraction with the mechanical junk guard recovers the model's intended one-word answer
    §3.3: replies are reduced to the final alphabetic word. Supported by an internal probe (only 13 multi-word cells; clamp-vs-free probe reproduces modal shares), so the assumption is tested but not guaranteed for every reply.
  • standard math Leave-one-out pooled surprisal with add-one smoothing is a stable estimator of answer conformity
    §3.2/Appendix A: smoothing-constant sensitivity rho >= 0.99; bootstrap CIs resample categories. A standard statistical construction, not an ad hoc one.

pith-pipeline@v1.3.0-alltime-deepseek · 18306 in / 22439 out tokens · 240493 ms · 2026-08-02T06:18:58.684590+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of The One-Word Census: Answer-Choice Conformity Across 44 Language Models." pith.science (2026). https://pith.science/paper/YQN6KZ6Q

@misc{pith2026260712796,
  author       = {Pith},
  title        = {Pith review of: The One-Word Census: Answer-Choice Conformity Across 44 Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YQN6KZ6Q}},
  note         = {Machine review of arXiv:2607.12796}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same answer every other model picks? Asked to "pick a word -- any word," 44 models chose "serendipity" 41% of the time. We characterize this convergence with a deliberately minimal instrument: 31 single-turn prompts, each naming a category with many valid one-word answers ("Name a tree."), asked four times per model with no system prompt. Analysis is exact-match on normalized tokens -- no embeddings, no judge -- at about a dollar per model. That models converge is well documented; our contribution is the instrument itself -- the One-Word Census -- and what it reveals about the structure of the convergence. We score each model by answer-choice surprisal: the average $-\log2$ probability of its answers under the pooled answers of all other models, leave-one-out. Convergence is extreme -- in 7 of 31 categories one answer takes over 80% of all answers -- yet conformity varies more than fourfold across models, and the variation is structured. Persona- and community-tuned models are the most divergent; the newest mainline flagships are the most conformist, producing almost no answer no other model gave. Within four lineages (Claude, GPT, Qwen, Grok) conformity rises with each generation -- but reverses for the latest flagship Claude and GPT models, a possible early signal of repositioning at the top tier. Rankings are robust to roster composition (leave-one-family-out rho = 0.985). Against human category-production norms, the field is more concentrated than people in 18 of 20 shared categories. All prompts, transcripts, and code are public.

Figures

Figures reproduced from arXiv: 2607.12796 by Tapan Parikh.

Figure 1
Figure 1. Figure 1: Modal-answer share by category, pooled over all 44 models (∼176 answers per category). Blue bars mark categories where a single answer takes ≥80% of the entire field’s answers. Direct labels give the modal answer. 4 Results 4.1 The substrate: a strong mode almost everywhere [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Answer-choice surprisal for all 44 models (bits; leave-one-out against the pooled field; add-one smoothed), with bootstrap 90% CIs from resampling categories. Higher = answers the field does not give. 8 [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Answer-choice surprisal across release generations within seven families (shared y-axis). Claude, GPT, Qwen, and Grok all decline, until the latest releases from Anthropic and OpenAI; Gemini is flat; Llama rises; DeepSeek is non-monotonic. The blue walk traces each family’s conformist mainline; amber points are the divergent premium siblings beside it — Claude Fable 5 beside Sonnet 5, and the GPT-5.6 premi… view at source ↗
Figure 4
Figure 4. Figure 4: The runner-up consensus. For each high-concentration category (modal share ≥75%), the share of the field’s non-modal answers taken by the single most common non-modal answer. When models deviate from the mode, they overwhelmingly deviate to the same place. 4.5 No pairwise affinity — and the runner-up consensus We tested for pairwise affinity across all 44 2  = 946 model pairs with a three-level ladder, ea… view at source ↗
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p015_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

    cs.CL 2026-07 conditional novelty 8.0

    Across 45 LLMs, the 'right?' tag effect flips from sycophantic to resistant over four years of releases, while the 'maybe?' tag raises agreement in every model — anti-sycophancy training is grammar-keyed and one-sided.

  2. Structured Output Collapses Answer Diversity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0

    Requesting JSON instead of plain chat measurably reduces answer diversity across 44 LLMs, concentrating answers onto the field's modal choice.

Reference graph

Works this paper leans on

33 extracted references · 19 linked inside Pith · cited by 2 Pith papers

  1. [1]

    Anderson, Jash Hemant Shah, and Max Kreminski

    Barrett R. Anderson, Jash Hemant Shah, and Max Kreminski. Homogenization effects of large language models on human creative ideation. InProceedings of the 16th Conference on Creativity & Cognition, 2024. arXiv:2402.01536

  2. [2]

    Commitments on model deprecation and preservation.https://www.anthropic

    Anthropic. Commitments on model deprecation and preservation.https://www.anthropic. com/research/deprecation-commitments, 2025

  3. [3]

    Battig and William E

    William F. Battig and William E. Montague. Category norms of verbal items in 56 categories: A replication and extension of the Connecticut category norms.Journal of Experimental Psychology, 80(3, Pt.2):1–46, 1969

  4. [4]

    The AI values dashboard.https://values.safe.ai, 2025

    Center for AI Safety. The AI values dashboard.https://values.safe.ai, 2025

  5. [5]

    How is ChatGPT’s behavior changing over time?arXiv preprint arXiv:2307.09009, 2023

    Lingjiao Chen, Matei Zaharia, and James Zou. How is ChatGPT’s behavior changing over time?arXiv preprint arXiv:2307.09009, 2023

  6. [6]

    DeepSeek-V3-0324 release

    DeepSeek. DeepSeek-V3-0324 release. https://api-docs.deepseek.com/updates, 2025. Same V3 base; post-training pipeline drawing on the R1 RL technique, with R1 reasoning distilled into the chat model

  7. [7]

    Doshi and Oliver P

    Anil R. Doshi and Oliver P. Hauser. Generative AI enhances individual creativity but reduces the collective diversity of novel content.Science Advances, 10(28), 2024. arXiv:2312.00506

  8. [8]

    Gueorguieva, Hongli Zhan, Jina Suh, Javier Hernandez, Tatiana Lau, Junyi Jessy Li, and Desmond C

    Emma S. Gueorguieva, Hongli Zhan, Jina Suh, Javier Hernandez, Tatiana Lau, Junyi Jessy Li, and Desmond C. Ong. AI generates well-liked but templatic empathic responses.arXiv preprint arXiv:2604.08479, 2026. 19

  9. [9]

    The curious decline of linguistic diversity: Training language models on synthetic text.Findings of NAACL,

    Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, and Chloé Clavel. The curious decline of linguistic diversity: Training language models on synthetic text.Findings of NAACL,

  10. [10]

    KL-regularized reinforcement learning is designed to mode collapse.arXiv preprint arXiv:2510.20817, 2025

    Anthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus, and Rajesh Ranganath. KL-regularized reinforcement learning is designed to mode collapse.arXiv preprint arXiv:2510.20817, 2025

  11. [11]

    Mysteries of mode collapse

    Janus. Mysteries of mode collapse. LessWrong, 2022. URLhttps://www.lesswrong.com/ posts/t9svvNPNmFf5Qa3TA/mysteries-of-mode-collapse

  12. [12]

    Artificial Hivemind: The open-ended homogeneity of language models (and beyond).Advances in Neural Information Processing Systems (NeurIPS), 2025

    Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, Alon Albalak, and Yejin Choi. Artificial Hivemind: The open-ended homogeneity of language models (and beyond).Advances in Neural Information Processing Systems (NeurIPS), 2025. arXiv:2510.22954

  13. [13]

    Walter G. Johnson. New methods for deprecating artificial intelligence systems will preserve history and facilitate research.Nature Communications, 15, 2024. doi:10.1038/s41467-024- 54758-1

  14. [14]

    Where does output diversity collapse in post-training?arXiv preprint arXiv:2604.16027, 2026

    Constantinos Karouzos, Xingwei Tan, and Nikolaos Aletras. Where does output diversity collapse in post-training?arXiv preprint arXiv:2604.16027, 2026

  15. [15]

    Understanding the effects of RLHF on LLM generalisation and diversity

    Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, et al. Understanding the effects of RLHF on LLM generalisation and diversity. InInternational Conference on Learning Representations (ICLR), 2024. arXiv:2310.06452

  16. [16]

    please, don’t kill the only model that still feels human

    Huiqian Lai. “please, don’t kill the only model that still feels human”: Understanding the #Keep4o backlash. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26), Barcelona, Spain, 2026. ACM. doi: 10.1145/3772318.3791351. arXiv:2602.00773

  17. [17]

    A diversity-promoting objective function for neural conversation models

    Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. A diversity-promoting objective function for neural conversation models. InProceedings of NAACL-HLT, 2016. arXiv:1510.03055

  18. [18]

    Jointly reinforcing diversity and quality in language model generations.arXiv preprint arXiv:2509.02534, 2025

    Tianjian Li, Yiming Zhang, Ping Yu, Swarnadeep Saha, Daniel Khashabi, Jason Weston, Jack Lanchantin, and Tianlu Wang. Jointly reinforcing diversity and quality in language model generations.arXiv preprint arXiv:2509.02534, 2025

  19. [19]

    Holistic evaluation of language models

    Percy Liang, Rishi Bommasani, Tony Lee, et al. Holistic evaluation of language models. Transactions on Machine Learning Research (TMLR), 2023. arXiv:2211.09110

  20. [20]

    The alignment tax: Response homogenization in aligned LLMs and its implica- tions for uncertainty estimation.arXiv preprint arXiv:2603.24124, 2026

    Mingyi Liu. The alignment tax: Response homogenization in aligned LLMs and its implica- tions for uncertainty estimation.arXiv preprint arXiv:2603.24124, 2026

  21. [21]

    Does writing with language models reduce con- tent diversity? InInternational Conference on Learning Representations (ICLR), 2024

    Vishakh Padmakumar and He He. Does writing with language models reduce con- tent diversity? InInternational Conference on Learning Representations (ICLR), 2024. arXiv:2309.05196. 20

  22. [22]

    Whose opinions do language models reflect? InInternational Conference on Machine Learning (ICML), 2023

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? InInternational Conference on Machine Learning (ICML), 2023. arXiv:2303.17548

  23. [23]

    AI models collapse when trained on recursively generated data.Nature, 631: 755–759, 2024

    Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. AI models collapse when trained on recursively generated data.Nature, 631: 755–759, 2024. arXiv:2305.17493

  24. [24]

    The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism.arXiv preprint arXiv:2407.10457, 2024

    Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism.arXiv preprint arXiv:2407.10457, 2024

  25. [25]

    rspeer/wordfreq: v3.0, 2022

    Robyn Speer. rspeer/wordfreq: v3.0, 2022. Multi-corpus word-frequency data for 44 languages

  26. [26]

    Evaluating the evaluation of diversity in natural language generation

    Guy Tevet and Jonathan Berant. Evaluating the evaluation of diversity in natural language generation. InProceedings of EACL, 2021. arXiv:2004.02990

  27. [27]

    Van Overschelde, Katherine A

    James P. Van Overschelde, Katherine A. Rawson, and John Dunlosky. Category norms: An updated and expanded version of the Battig and Montague (1969) norms.Journal of Memory and Language, 50(3):289–335, 2004

  28. [28]

    We’re different, we’re the same: Creative homogeneity across LLMs.arXiv preprint arXiv:2501.19361, 2025

    Emily Wenger and Yoed Kenett. We’re different, we’re the same: Creative homogeneity across LLMs.arXiv preprint arXiv:2501.19361, 2025

  29. [29]

    Epistemic diversity and knowledge collapse in large language models.arXiv preprint arXiv:2510.04226, 2025

    Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christensen, Chan Young Park, and Isabelle Augenstein. Epistemic diversity and knowledge collapse in large language models.arXiv preprint arXiv:2510.04226, 2025

  30. [30]

    Forcing diffuse distributions out of language models.arXiv preprint arXiv:2404.10859, 2024

    Yiming Zhang, Avi Schwarzschild, Nicholas Carlini, Zico Kolter, and Daphne Ippolito. Forcing diffuse distributions out of language models.arXiv preprint arXiv:2404.10859, 2024

  31. [31]

    NoveltyBench: Evaluating language models for humanlike diversity

    Yiming Zhang, Harshita Diddee, Susan Holm, Hanchen Liu, Xinyue Liu, Vinay Samuel, Barry Wang, and Daphne Ippolito. NoveltyBench: Evaluating language models for humanlike diversity. InConference on Language Modeling (COLM), 2025. arXiv:2504.05228

  32. [32]

    WildChat: 1M ChatGPT interaction logs in the wild

    Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. WildChat: 1M ChatGPT interaction logs in the wild. InInternational Conference on Learning Representations (ICLR), 2024. arXiv:2405.01470

  33. [33]

    Reply with one word only

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, et al. LMSYS-Chat-1M: A large-scale real- world LLM conversation dataset. InInternational Conference on Learning Representations (ICLR), 2024. arXiv:2309.11998. 21 A Full scorecard Table 1:All 44 models, ranked by answer-choice surprisal (bits; leave-one-out, add-one smoothed; bootstrap 90% CI over categories).Av...