Pith. sign in

REVIEW 3 major objections 5 minor 75 references

LLM opinion diversity saturates at a one-sentence persona, and different interaction designs cover different opinion regions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 14:27 UTC pith:FKUGB7ZM

load-bearing objection A serious, well-run factorial audit of LLM opinion diversity whose headline architecture-complementarity and interaction claims are undercut by sample-size and confound issues that a good revision could fix. the 3 major comments →

arxiv 2607.20429 v1 pith:FKUGB7ZM submitted 2026-05-10 cs.CL cs.AI

More Is Not More: What Matters for Diversity in LLM Opinions?

classification cs.CL cs.AI
keywords LLM opinion diversitypersona promptingmulti-agent discussionmulti-turn promptingfactorial experimentalpha/beta diversityVendi scoresynthetic surveys
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that the diversity of opinions an LLM produces is not a matter of scaling any single lever. In a factorial experiment across 100 open-ended questions, seven models, five persona depths, and three interaction architectures, the authors find that a single-sentence occupation description already captures most of the diversity gain that persona conditioning can deliver; adding demographic detail does not consistently help and sometimes reduces diversity. They also find that multi-turn self-prompting and multi-agent discussion explore largely non-overlapping opinion regions, so combining architectures beats choosing the best one, and that low-cost tricks such as raising temperature or adding 'consider diverse perspectives' instructions have negligible effects. For anyone building synthetic surveys, focus groups, or opinion simulators, the paper supplies a map of which interventions actually move the opinion distribution and which are wasted effort.

Core claim

The central claim is that LLM opinion diversity is governed by the structural form of interventions, not by their magnitude. Persona conditioning increases diversity, but the gain saturates at the first step: Role, a single-sentence occupation description, produces the majority of the None-to-Pro gain on all seven models, and Basic (adding demographic attributes) fails to improve on Role and on three models reduces dispersion. Multi-turn and multi-agent architectures each increase diversity over single calls, yet the two cover largely non-overlapping opinion spaces—about half of the opinion clusters in one are absent from the other after baseline correction—so merging their outputs gives bro

What carries the argument

The carrying mechanism is a two-axis factorial grid—input conditioning (persona depth None/Role/Basic/Mid/Pro) crossed with interaction architecture (Single-Call, Multi-Turn Self-Prompting, Multi-Agent Discussion)—plus a four-condition low-cost-trick baseline. Diversity is measured at two levels after extracting atomic opinion statements and embedding them: α-diversity within a condition (mean pairwise distance, cluster count, and Vendi score) and β-diversity between conditions (beta-Vendi score and unique cluster ratio, both calibrated against a split-half noise floor). The key quantities are the MPD, which is sample-size invariant, and the UCR, which operationalizes 'non-overlapping opinio

Load-bearing premise

The conclusions rest on treating cosine similarity between embedded opinion sentences as a faithful stand-in for how distinct a human would judge those opinions; if embedding geometry does not track human-perceived opinion differences, the reported gains and non-overlaps are artifacts of the embedding space.

What would settle it

Recruit human raters to judge pairwise similarity of extracted opinions from Multi-Turn vs. Multi-Agent (Pro) without source labels, or rate whether Role or Basic opinions are more varied; if humans see the architectures' opinions as mostly redundant or find Basic more diverse than Role, the headline claims fail. A cheaper check: re-run the cluster-based metrics at a cosine threshold of 0.75 or with a different embedding model and see whether the Role>Basic ordering and the roughly 50% excess Unique Cluster Ratio survive.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Builders of synthetic opinion pipelines should start with a one-sentence occupation persona; it is the highest-return intervention and outperforms all low-cost alternatives by about 2.5 times.
  • Deploying multiple interaction architectures and merging their outputs covers substantially more of the opinion space than tuning any single architecture.
  • Raising temperature, appending 'consider diverse perspectives,' and cueing demographic dimensions are not effective routes to population-level diversity.
  • Expanding the persona pool (from 5 to 20) adds more new opinion categories than enriching individual personas, so breadth should be prioritized over depth.
  • Diversity evaluation needs a standardized protocol; the paper's extraction-embedding-metric pipeline and calibrated β-diversity measures are offered as that protocol.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The saturating persona curve suggests an 'identity contrast' mechanism: once a persona label divides respondents into distinct conditioning regimes, additional detail only sharpens within-regime consistency rather than creating new opinion regions—a hypothesis testable by varying the semantic distance between occupation labels rather than their descriptive length.
  • The near-orthogonality of multi-turn and multi-agent opinion spaces hints that self-generation and group interaction may activate different priors (introspective vs. socially anchored); a testable extension is whether a debate-style multi-agent prompt (explicit disagreement) recovers the categories that neutral discussion compresses.
  • Because the embedding-space measure is the load-bearing evaluation, an obvious extension is to calibrate the headline rankings against human pairwise similarity judgments on a subset of extracted opinions; if humans see the two architectures as redundant, the complementarity claim would need revision.
  • The question set skews toward opinion-eliciting text domains from real user queries, so the intervention rankings may not transfer to non-textual or culturally different settings; replicating the factorial grid with questions from other cultural contexts would test the generality of the 'more is not more' pattern.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a factorial audit of interventions for increasing opinion diversity in LLM generation. It independently varies persona depth (None, Role, Basic, Mid, Pro) and interaction architecture (Single-Call, Multi-Turn, Multi-Agent), evaluates 100 real-user questions from WildChat across 7 chat models, and adds four low-cost trick conditions. Diversity is measured after opinion extraction and embedding using α-diversity metrics (MPD, CC, VS), with rarefaction for CC/VS, and β-diversity metrics (β-VS, UCR). The main claims are: a single-sentence occupation persona captures most of the diversity gain, with further detail giving diminishing returns or even reductions; Multi-Turn and Multi-Agent architectures cover largely non-overlapping opinion regions, so combining them is better than choosing one; and low-cost tricks such as higher temperature or diversity instructions have negligible effects compared with structured interventions.

Significance. If the results hold after the sample-size corrections described below, this is a valuable contribution: it provides a rare controlled decomposition of two often-conflated intervention axes, with per-model tables, paired Wilcoxon tests with BH correction, extraction-fidelity audits, embedder-robustness checks, and a released evaluation protocol. The 'more is not more' result for persona depth and the architecture-complementarity result are practically actionable and could shift practice from anecdotal knob-tuning to structured evaluation. The paper is unusually transparent about prompts, model versions, per-condition manifests, and compute. The two load-bearing concerns—unrarefied β-diversity metrics and a confounded interaction analysis—are fixable within the manuscript's scope, but they directly affect the central architecture-complementarity recommendation and the Section 4.4 interaction claim.

major comments (3)
  1. [Section 3 / Appendix G / Table 18 / Table 9 / Table 22] β-VS and UCR are computed on full opinion sets with no rarefaction, although both are sample-size sensitive: VS grows with N, and UCR counts clusters, which also grow with N. The rarefaction procedure in Appendix G is stated for CC and VS only, and the split-half baseline in Table 9 uses equal-sized halves, so it cannot calibrate comparisons with unequal N. The central architecture-complementarity row in Table 18 is explicitly 'Multi-Turn vs. Multi-Agent (5p Pro)', so the 4× gap in Table 22 does not directly apply to Figure 5; nevertheless, extraction density still differs by architecture (Appendix E: 4.3 vs 5.4 opinions per 1000 characters), and most other Table 18 rows are unbalanced. Please add rarefied or matched-N β-VS/UCR analyses and report whether the ~50% excess UCR over the split-half baseline survives when N is balanced by construction.
  2. [Section 4.4 / Section 2 / Table 4 / Table 23] The interaction claim compares the None→Pro persona gradient under Single-Call and Multi-Turn (20 personas) with the same gradient under Multi-Agent (always 5 personas). The statement that this is 'the same persona set' is inconsistent with Table 4: persona_pro and mt_pro use 20 personas, while minimal_pro uses 5. Pool size and architecture are therefore confounded in the +13.6% / +11.5% / +11.9% MPD comparison and in the +46% / +82% / +15% rarefied-CC comparison. Recompute the interaction using the matched 5-persona conditions (persona_5, mt_pro_5, minimal_pro) or otherwise equate pool size before claiming that the persona effect depends on architecture.
  3. [Appendix F / Section 3 / Section 4.2] The τ=0.65 threshold sensitivity sweep in Appendix F covers CC only. Since the headline complementarity result is an excess UCR of roughly 50%, and UCR uses the same cosine threshold, a threshold sweep on UCR (and ideally β-VS) is needed to verify that the architecture-complementarity conclusion is not an artifact of the chosen threshold. This is especially important because CC and UCR can respond differently to threshold shifts, and the current robustness checks do not address the metric that carries the main architectural recommendation.
minor comments (5)
  1. [Section 4.2 / Figure 5] The text and figure do not state that the Multi-Turn vs. Multi-Agent UCR comparison uses 5-persona Pro subsets. Add an explicit statement and report both UCR directions (A-only and B-only) in the figure or caption.
  2. [Tables 15–17] The sign convention for Cliff's δ is not stated in the captions. Negative values indicate the second condition is higher in many rows; please state 'negative δ means the right-hand condition has higher diversity' or equivalent.
  3. [Section 2 / Appendix D] The main text says all conditions share a uniform set of generation parameters, but Kimi uses T=0.6 and top-p=0.95. This exception appears only in Appendix D; note it in Section 2 to avoid misleading readers.
  4. [Abstract / Introduction / Appendix D] The paper repeatedly says '7 models', but GPT-5.4-mini was run on only 9 of the 19 primary conditions. Qualify aggregate claims accordingly, or state the coverage limitation in the main text rather than only in Appendix D.
  5. [Section 4.3 / Table 15] The text says Trait Assignment is significant on six of six models, but Table 15 row B4 shows negative δ values for all six. The direction is consistent with 'RandomSys increases MPD over None', but the signs should be explained to prevent misreading.

Circularity Check

0 steps flagged

No circularity: the paper is an empirical measurement study; headline findings are direct comparisons of measured outputs, not consequences of fitted constants or self-citations.

full rationale

The paper's claims are empirical measurements over generated outputs, not derivations from an assumed model. No equation defines a target finding in terms of its own inputs: MPD is computed directly from embeddings; CC and VS are explicitly rarefied for cross-condition α comparisons; the β metrics are defined by an additive decomposition whose baseline is calibrated and then subtracted. The clustering threshold τ=0.65 was set on a pilot and swept in Appendix F, and the headline orderings survive the sweep, so it does not encode the findings. The central Multi-Turn vs. Multi-Agent complementarity claim is based on matched 5-persona conditions with equal response counts (Table 4), and the paper reports both directions of UCR. The low-cost-tricks results are direct ΔMPD comparisons. There is no load-bearing self-citation: the author's released codebase is an artifact, not an evidence source. The acknowledged limitation that semantic embeddings may not match human-perceived diversity is a validity boundary, not circularity. The reviewer's sample-size concern about β metrics is a potential confound for some comparisons, but it is not a definitional reduction of the headline result to the metric's construction, so it does not raise the circularity score.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

No fitted constants or invented entities carry the central claims. The main assumptions are measurement choices (embedding as diversity proxy, extractor fidelity, rarefaction) and one design assumption about persona-pool comparability in the interaction analysis.

free parameters (2)
  • cosine similarity clustering threshold tau = 0.65
    Hand-set on a held-out pilot; used to define clusters for CC, UCR, and beta-VS. Headline orderings survive a sweep from 0.55 to 0.75, so it is not fitted to the target conclusion, but it does shape quantitative diversity magnitudes.
  • semantic Jaccard equivalence threshold = 0.85
    Used in the extraction-stability audit to treat embeddings with cosine similarity above 0.85 as the same opinion. Minor procedural choice that affects the reported stability number, not the central claims.
axioms (5)
  • domain assumption Cosine similarity in text-embedding-3-small space is a valid proxy for semantic opinion diversity.
    The entire alpha/beta metric pipeline measures distances and clusters in this embedding space. Section 3 states that cosine similarity 'captures semantic relatedness'.
  • domain assumption DeepSeek v3.2 atomic-opinion extraction faithfully converts heterogeneous LLM responses into comparable opinion statements.
    All raw responses are reduced to extracted atomic opinions before metric computation. Appendix E validates precision (98.2%) and cross-extractor agreement, but the fidelity of the extraction itself remains an assumption about the extractor model.
  • domain assumption Rarefaction to the minimum opinion count removes sample-size confounds for CC and VS comparisons.
    Appendix G applies rarefaction because CC and VS grow with N. This assumes that random subsampling recovers the population-level category structure that a larger sample would reveal.
  • domain assumption The 100 WildChat questions are representative of opinion-eliciting open-ended tasks.
    Appendix A describes a manual selection pipeline targeting opinion-divergent domains. The set is deliberately not a random sample of all user queries, so the conclusions are scoped to this class of questions.
  • ad hoc to paper Multi-Agent with 5 personas is comparable to Single-Call and Multi-Turn with 20 personas when estimating the persona-depth interaction effect.
    Section 2 fixes Multi-Agent group size at 5 to limit context growth, while Single-Call and Multi-Turn use 20 personas. Section 4.4 compares None-to-Pro gains across architectures without correcting for this pool-size mismatch.

pith-pipeline@v1.3.0-alltime-deepseek · 32513 in / 13417 out tokens · 148704 ms · 2026-08-02T14:27:07.940387+00:00 · methodology

0 comments
read the original abstract

Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion homogenization. Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are typically deployed and upgraded simultaneously, making it difficult to attribute gains to specific components. To advance a more scientific understanding of LLM output diversity, we design a factorial experiment that separates two primary intervention dimensions: input conditioning (operationalized through persona depth) and interaction architecture. We evaluate all conditions on 100 real-user open-ended questions across 7 models, measuring diversity with multiple complementary metrics. Our findings challenge several common assumptions. First, more persona detail does not monotonically increase diversity. The initial step of persona conditioning already captures the majority of the gain, while further elaboration with demographic detail does not consistently improve and can reduce diversity on some models. Second, rather than seeking a single best interaction architecture, we find that different architectures explore largely non-overlapping opinion regions. Combining multiple architectures yields broader coverage than optimizing any one. Third, commonly attempted low-cost alternatives such as raising sampling temperature and adding diversity instructions produce negligible effects compared to structured interventions. Overall, our work demonstrates that diversity is not a product of scaling along any single dimension, but is highly sensitive to the structural form and combination of interventions.

Figures

Figures reproduced from arXiv: 2607.20429 by Qiyang Yao.

Figure 1
Figure 1. Figure 1: Factorial experimental design. The vertical axis varies input conditioning through five levels [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Evaluation pipeline. Each LLM response undergoes opinion extraction to isolate atomic [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: α-diversity versus β-diversity. (a, b) α-diversity: opinions cluster tightly (low α) or spread widely (high α). (c, d) β-diversity: two conditions cover the same region (low β) or occupy distinct regions (high β). 5 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: MPD across five persona depth levels for seven models under (a) Single-Call, (b) Multi-Turn, [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: UCR between Multi-Turn and Multi￾Agent architectures. All models exceed the split￾half baseline (dashed, 29%); mean UCR is 79%. No clear winner emerges between Multi-Turn and Multi-Agent on α-diversity. Most mod￾els favor Multi-Turn (most δ > 0.50), but Kimi and Qwen favor Multi-Agent or show no differ￾ence. Multi-Turn yields higher CC and VS than Multi-Agent on all seven models (large effect). This compar… view at source ↗
Figure 6
Figure 6. Figure 6: ∆MPD vs. None for low￾cost strategies and Role. Trait Assignment is the only low-cost strategy with a mea￾surable effect, yet still falls well short of Role. It reaches significance on six of six models (δ = 0.30–0.88), but the effect size is only approximately 37% of Role’s (range 17– 47%). Trait Assignment assigns a different personality label to each call—the lightest possible form of identity differ￾en… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

75 extracted references · 5 canonical work pages

  1. [1]

    Argyle, Ethan C

    Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. Out of one, many: Using language models to simulate human samples.Political Analysis, 31(3):337–351, 2023. doi: 10.1017/pan.2023.2

  2. [2]

    Emergent social conventions and collective bias in LLM populations.Science Advances, 11(20):eadu9368, 2025

    Ariel Flint Ashery, Luca Maria Aiello, and Andrea Baronchelli. Emergent social conventions and collective bias in LLM populations.Science Advances, 11(20):eadu9368, 2025

  3. [3]

    Sensitivity, performance, robustness: Deconstructing the effect of sociodemographic prompting

    Tilman Beck, Hendrik Schuff, Anne Lauscher, and Iryna Gurevych. Sensitivity, performance, robustness: Deconstructing the effect of sociodemographic prompting. InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (EACL), 2024. arXiv:2309.07034

  4. [4]

    Controlling the false discovery rate: A practical and powerful approach to multiple testing.Journal of the Royal Statistical Society, Series B, 57(1): 289–300, 1995

    Yoav Benjamini and Yosef Hochberg. Controlling the false discovery rate: A practical and powerful approach to multiple testing.Journal of the Royal Statistical Society, Series B, 57(1): 289–300, 1995

  5. [5]

    Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer Larson

    James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer Larson. Synthetic replacements for human survey data? the perils of large language models.Political Analysis, 32 (4):401–416, 2024. doi: 10.1017/pan.2024.5

  6. [6]

    PER- SONA: A reproducible testbed for pluralistic alignment.arXiv preprint arXiv:2407.17387, 2024

    Louis Castricato, Nathan Lile, Rafael Rafailov, Jan-Philipp Fränken, and Chelsea Finn. PER- SONA: A reproducible testbed for pluralistic alignment.arXiv preprint arXiv:2407.17387, 2024

  7. [7]

    BGE M3- embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation

    Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. BGE M3- embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. InFindings of the Association for Computational Linguistics: ACL 2024, 2024. arXiv:2402.03216

  8. [8]

    ReConcile: Round-table con- ference improves reasoning via consensus among diverse LLMs

    Justin Chih-Yao Chen, Swarnadeep Saha, and Mohit Bansal. ReConcile: Round-table con- ference improves reasoning via consensus among diverse LLMs. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pages 7066–7085, 2024

  9. [9]

    Marked personas: Using natural language prompts to measure stereotypes in language models

    Myra Cheng, Esin Durmus, and Dan Jurafsky. Marked personas: Using natural language prompts to measure stereotypes in language models. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), 2023. arXiv:2305.18189

  10. [10]

    Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T. Rogers. Simulating opinion dynamics with networks of LLM-based agents. InFindings of the Association for Computational Linguistics: NAACL 2024, 2024. arXiv:2311.09618

  11. [11]

    Diversity-rewarded CFG distillation.arXiv preprint arXiv:2410.06084, 2024

    Geoffrey Cideron, Andrea Agostinelli, Johan Ferret, Sertan Girgin, Romuald Elie, Olivier Bachem, Sarah Perrin, and Alexandre Rame. Diversity-rewarded CFG distillation.arXiv preprint arXiv:2410.06084, 2024

  12. [12]

    Dominance statistics: Ordinal analyses to answer ordinal questions.Psychologi- cal Bulletin, 114(3):494–509, 1993

    Norman Cliff. Dominance statistics: Ordinal analyses to answer ordinal questions.Psychologi- cal Bulletin, 114(3):494–509, 1993

  13. [13]

    Questioning the survey responses of large language models

    Ricardo Domínguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner. Questioning the survey responses of large language models. InAdvances in Neural Information Processing Systems, 2024. arXiv:2306.07951

  14. [14]

    Doshi and Oliver P

    Anil R. Doshi and Oliver P. Hauser. Generative AI enhances individual creativity but reduces the collective diversity of novel content.Science Advances, 10:eadn5290, 2024

  15. [15]

    Tenenbaum, and Igor Mordatch

    Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. Improving factuality and reasoning in language models through multiagent debate. InProceedings of the 41st International Conference on Machine Learning (ICML), volume 235 ofPMLR, pages 11733–11763, 2024. 10

  16. [16]

    Esin Durmus, Karina Nguyen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. Towards measuring the representation of subjective global opinions in language mod...

  17. [17]

    The Vendi Score: A diversity evaluation metric for machine learning.Transactions on Machine Learning Research, 2023

    Dan Friedman and Adji Bousso Dieng. The Vendi Score: A diversity evaluation metric for machine learning.Transactions on Machine Learning Research, 2023. arXiv:2210.02410

  18. [18]

    Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094, 2024

    Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094, 2024

  19. [19]

    Gotelli and Robert K

    Nicholas J. Gotelli and Robert K. Colwell. Quantifying biodiversity: Procedures and pitfalls in the measurement and comparison of species richness.Ecology Letters, 4(4):379–391, 2001

  20. [20]

    Bias runs deep: Implicit reasoning biases in persona-assigned LLMs

    Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. Bias runs deep: Implicit reasoning biases in persona-assigned LLMs. InThe Twelfth International Conference on Learning Representations (ICLR), 2024

  21. [21]

    KL-regularized reinforcement learning is designed to mode collapse.arXiv preprint arXiv:2510.20817, 2025

    Anthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus, and Rajesh Ranganath. KL-regularized reinforcement learning is designed to mode collapse.arXiv preprint arXiv:2510.20817, 2025

  22. [22]

    Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, and Dongyeop Kang. How far can we extract diverse perspectives from large language models? InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5336–5366, 2024. doi: 10.18653/v1/2024.emnlp-main.306

  23. [23]

    Predicting results of social science experiments using large language models.Working paper, 2024

    Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae, and Robb Willer. Predicting results of social science experiments using large language models.Working paper, 2024

  24. [24]

    John J. Horton. Large language models as simulated economic agents: What can we learn from homo silicus?NBER Working Paper No. 31122, 2023. arXiv:2301.07543

  25. [25]

    Quantifying the persona effect in LLM simulations

    Tiancheng Hu and Nigel Collier. Quantifying the persona effect in LLM simulations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pages 10289–10307, 2024. doi: 10.18653/v1/2024.acl-long.554

  26. [26]

    Debate-to-write: A persona-driven multi-agent framework for diverse argument generation

    Zhe Hu, Hou Pong Chan, Jing Li, and Yu Yin. Debate-to-write: A persona-driven multi-agent framework for diverse argument generation. InProceedings of the 31st International Conference on Computational Linguistics (COLING), pages 4689–4703, 2025

  27. [27]

    Person- aLLM: Investigating the ability of large language models to express personality traits

    Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. Person- aLLM: Investigating the ability of large language models to express personality traits. InFind- ings of the Association for Computational Linguistics: NAACL 2024, 2024. arXiv:2305.02547

  28. [28]

    Artificial hivemind: The open-ended homogeneity of language models (and beyond)

    Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, and Yejin Choi. Artificial hivemind: The open-ended homogeneity of language models (and beyond). InAdvances in Neural Information Processing Systems, volume 38, 2025. Datasets and Benchmarks Track oral

  29. [29]

    Partitioning diversity into independent alpha and beta components.Ecology, 88(10): 2427–2439, 2007

    Lou Jost. Partitioning diversity into independent alpha and beta components.Ecology, 88(10): 2427–2439, 2007. doi: 10.1890/06-1736.1

  30. [30]

    Measuring lexical diversity of synthetic data generated through fine-grained persona prompting

    Gauri Kambhatla, Chantal Shaib, and Venkata Govindarajan. Measuring lexical diversity of synthetic data generated through fine-grained persona prompting. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 21024–21033, 2025. doi: 10.18653/v1/ 2025.findings-emnlp.1146. arXiv:2505.17390

  31. [31]

    Bowman, Tim Rocktäschel, and Ethan Perez

    Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, and Ethan Perez. Debating with more persuasive LLMs leads to more truthful answers. InProceedings of the 41st International Conference on Machine Learning (ICML), 2024. arXiv:2402.06782. 11

  32. [32]

    Understanding the effects of RLHF on LLM generalisation and diversity

    Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. Understanding the effects of RLHF on LLM generalisation and diversity. InThe Twelfth International Conference on Learning Representations (ICLR),

  33. [33]

    Krueger and Mary Anne Casey.Focus Groups: A Practical Guide for Applied Research

    Richard A. Krueger and Mary Anne Casey.Focus Groups: A Practical Guide for Applied Research. SAGE Publications, Thousand Oaks, CA, 5th edition, 2014

  34. [34]

    Diverse preference optimization.arXiv preprint arXiv:2501.18101, 2025

    Jack Lanchantin, Angelica Chen, Shehzaad Dhuliawala, Ping Yu, Jason Weston, Sainbayar Sukhbaatar, and Ilia Kulikov. Diverse preference optimization.arXiv preprint arXiv:2501.18101, 2025

  35. [35]

    Statistics and partitioning of species diversity, and similarity among multiple communities.Oikos, 76(1):5–13, 1996

    Russell Lande. Statistics and partitioning of species diversity, and similarity among multiple communities.Oikos, 76(1):5–13, 1996. doi: 10.2307/3545743

  36. [36]

    Messi H. J. Lee, Jacob M. Montgomery, and Calvin K. Lai. Large language models portray socially subordinate groups as more homogeneous, consistent with a bias observed in humans. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2024. arXiv:2401.08495

  37. [37]

    Encouraging divergent thinking in large language models through multi- agent debate

    Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. Encouraging divergent thinking in large language models through multi- agent debate. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. arXiv:2305.19118

  38. [38]

    Evaluating large language model biases in persona- steered generation

    Andy Liu, Mona Diab, and Daniel Fried. Evaluating large language model biases in persona- steered generation. InFindings of the Association for Computational Linguistics: ACL 2024,

  39. [39]

    The prompt makes the person(a): A systematic evaluation of sociodemographic persona prompting for large language models

    Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, and Markus Strohmaier. The prompt makes the person(a): A systematic evaluation of sociodemographic persona prompting for large language models. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 23212–23237, 2025. doi: 10.18653/v1/2025.findings-emnlp.1261. arXiv:2507.16076

  40. [40]

    Mollick, and Christian Terwiesch

    Lennart Meincke, Ethan R. Mollick, and Christian Terwiesch. Prompting diverse ideas: Increas- ing AI idea variance. SSRN 4708466 / Wharton-Mack Institute Working Paper, 2024

  41. [41]

    FActScore: Fine-grained atomic evaluation of factual precision in long form text generation

    Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023. arXiv:2305.14251

  42. [42]

    Modern hierarchical, agglomerative clustering algorithms.arXiv preprint arXiv:1109.2378, 2011

    Daniel Müllner. Modern hierarchical, agglomerative clustering algorithms.arXiv preprint arXiv:1109.2378, 2011

  43. [43]

    One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity

    Sonia Murthy, Tomer Ullman, and Jennifer Hu. One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity. InProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 11241–11258, 2025. doi: 1...

  44. [44]

    New embedding models and API updates

    OpenAI. New embedding models and API updates. https://openai.com/index/ new-embedding-models-and-api-updates/ , 2024. OpenAI Blog Post, January 25, 2024

  45. [45]

    Does writing with language models reduce content diver- sity? InThe Twelfth International Conference on Learning Representations (ICLR), 2024

    Vishakh Padmakumar and He He. Does writing with language models reduce content diver- sity? InThe Twelfth International Conference on Learning Representations (ICLR), 2024. arXiv:2309.05196

  46. [46]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), 2023. doi: 10.1145/3586183.3606763. 12

  47. [47]

    Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S

    Joon Sung Park, Carolyn Q. Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S. Bernstein. Generative agent simulations of 1,000 people.arXiv preprint arXiv:2411.10109, 2024

  48. [48]

    Park, Philipp Schoenegger, and Chongyang Zhu

    Peter S. Park, Philipp Schoenegger, and Chongyang Zhu. Diminished diversity-of-thought in a standard large language model.Behavior Research Methods, 56:5754–5770, 2024. doi: 10.3758/s13428-023-02307-x

  49. [49]

    Pasarkar and Adji Bousso Dieng

    Amey P. Pasarkar and Adji Bousso Dieng. Cousins of the Vendi Score: A family of similarity- based diversity metrics for science and machine learning. InProceedings of the 27th Interna- tional Conference on Artificial Intelligence and Statistics (AISTATS), volume 238 ofPMLR, pages 3808–3816, 2024

  50. [50]

    Is temperature the creativity parameter of large language models? InProceedings of the 15th International Conference on Computational Creativity (ICCC), 2024

    Max Peeperkorn, Tom Kouwenhoven, Daniel Brown, and Anna Jordanous. Is temperature the creativity parameter of large language models? InProceedings of the 15th International Conference on Computational Creativity (ICCC), 2024. arXiv:2405.00492

  51. [51]

    Sentence-BERT: Sentence embeddings using siamese BERT- networks

    Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence embeddings using siamese BERT- networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, 2019. arXiv:1908.10084

  52. [52]

    The effect of sampling temperature on problem solving in large language models

    Matthew Renze and Erhan Guven. The effect of sampling temperature on problem solving in large language models. InFindings of the Association for Computational Linguistics: EMNLP 2024, 2024. arXiv:2402.05201

  53. [53]

    Kromrey, Jesse Coraggio, and Jeff Skowronek

    Jeanine Romano, Jeffrey D. Kromrey, Jesse Coraggio, and Jeff Skowronek. Appropriate statistics for ordinal level data: Should we really be using t-test and Cohen’s d for evaluating group differences on the NSSE and other surveys? InAnnual Meeting of the Florida Association of Institutional Research, 2006

  54. [54]

    Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models

    Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schütze, and Dirk Hovy. Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024. arXiv:2402.16786

  55. [55]

    Whose opinions do language models reflect? InProceedings of the 40th Interna- tional Conference on Machine Learning (ICML), volume 202 ofPMLR, pages 29971–30004, 2023

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? InProceedings of the 40th Interna- tional Conference on Machine Learning (ICML), volume 202 ofPMLR, pages 29971–30004, 2023

  56. [56]

    Stewart and Prem N

    David W. Stewart and Prem N. Shamdasani.Focus Groups: Theory and Practice, volume 20 of Applied Social Research Methods. SAGE Publications, Thousand Oaks, CA, 3rd edition, 2014

  57. [57]

    Systematic biases in LLM simulations of debates

    Amir Taubenfeld, Yaniv Dover, Roi Reichart, and Ariel Goldstein. Systematic biases in LLM simulations of debates. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. arXiv:2402.04049

  58. [58]

    Dickerson

    Angelina Wang, Jamie Morgenstern, and John P. Dickerson. Large language models that replace human participants can harmfully misportray and flatten identity groups.Nature Machine Intelligence, 7:400–411, 2025. doi: 10.1038/s42256-025-00986-z. arXiv:2402.01908

  59. [59]

    Multilingual prompting for improving LLM generation diversity

    Qihan Wang, Shidong Pan, Tal Linzen, and Emily Black. Multilingual prompting for improving LLM generation diversity. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 6367–6389, 2025. doi: 10.18653/v1/2025.emnlp-main.324. arXiv:2505.15229

  60. [60]

    R. H. Whittaker. Evolution and measurement of species diversity.Taxon, 21(2/3):213–251,

  61. [61]

    Individual comparisons by ranking methods.Biometrics Bulletin, 1(6):80–83, 1945

    Frank Wilcoxon. Individual comparisons by ranking methods.Biometrics Bulletin, 1(6):80–83, 1945. 13

  62. [62]

    Epistemic diversity and knowledge collapse in large language models.arXiv preprint arXiv:2510.04226, 2025

    Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Chris- tensen, Chan Young Park, and Isabelle Augenstein. Epistemic diversity and knowledge collapse in large language models.arXiv preprint arXiv:2510.04226, 2025

  63. [63]

    The hidden strength of disagreement: Unraveling the consensus- diversity tradeoff in adaptive multi-agent systems

    Zengqing Wu and Takayuki Ito. The hidden strength of disagreement: Unraveling the consensus- diversity tradeoff in adaptive multi-agent systems. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15277–15297, 2025. doi: 10.18653/v1/2025.emnlp-main.772

  64. [64]

    The price of format: Diversity collapse in LLMs

    Longfei Yun, Chenyang An, Zilong Wang, Letian Peng, and Jingbo Shang. The price of format: Diversity collapse in LLMs. InFindings of the Association for Computational Linguis- tics: EMNLP 2025, pages 15454–15468, 2025. doi: 10.18653/v1/2025.findings-emnlp.836. arXiv:2505.18949

  65. [65]

    Tomz, Christopher D

    Jiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia, Michael R. Tomz, Christopher D. Manning, and Weiyan Shi. Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity.arXiv preprint arXiv:2510.01171, 2025

  66. [66]

    Taiyu Zhang, Xuesong Zhang, Robbe Cools, and Adalberto L. Simeone. Focus agent: LLM- powered virtual focus group. InProceedings of the 24th ACM International Conference on Intelligent Virtual Agents (IVA), 2024. doi: 10.1145/3652988.3673918

  67. [67]

    Qwen3 embed- ding: Advancing text embedding and reranking through foundation models.arXiv preprint arXiv:2506.05176, 2025

    Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. Qwen3 embed- ding: Advancing text embedding and reranking through foundation models.arXiv preprint arXiv:2506.05176, 2025

  68. [68]

    WildChat: 1M ChatGPT interaction logs in the wild

    Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. WildChat: 1M ChatGPT interaction logs in the wild. InThe Twelfth International Conference on Learning Representations (ICLR), 2024. arXiv:2405.01470

  69. [69]

    According to epistemic responsibility, write an argument that says it’s wrong to say that AI image generators steal art

    Xiaochen Zhu, Caiqi Zhang, Tom Stafford, Nigel Collier, and Andreas Vlachos. Conformity in large language models. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3854–3872, 2025. doi: 10.18653/ v1/2025.acl-long.195. arXiv:2410.12428. A Question Set This appendix supports Section 2 by...

  70. [71]

    You are creative, spontaneous, and enjoy thinking outside the box

    “You are creative, spontaneous, and enjoy thinking outside the box.”

  71. [72]

    You are analytical, cautious, and prefer evidence-based reasoning

    “You are analytical, cautious, and prefer evidence-based reasoning.”

  72. [73]

    You are warm, empathetic, and prioritize human connection

    “You are warm, empathetic, and prioritize human connection.”

  73. [74]

    You are direct, pragmatic, and focused on efficiency

    “You are direct, pragmatic, and focused on efficiency.”

  74. [75]

    You are skeptical, independent-minded, and question conventional wisdom

    “You are skeptical, independent-minded, and question conventional wisdom.” Opinion extraction prompt.The extractor uses DeepSeek v3.2 at T= 0 . The system prompt is fixed (SHA-256 prefixf6fbb4f1, recorded in every metric JSON for provenance): You extract opinion statements from a response to a question. An opinion statement is a claim that expresses a jud...

  75. [1972]

    doi: 10.2307/1218190