Pith. sign in

REVIEW 3 major objections 6 minor 22 references

The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices

T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Delegating choices to LLM agents—generic or personalized—pulls preferences toward the popular and shrinks variety.

desk verdict A solid, honest empirical study with a genuinely new angle on agentic AI and individual choice diversity, but the headline claim overstates the evidence because the 'Human Control' baseline is a random draw, not a human choice. read the letter →

arxiv 2509.02910 v1 pith:NK5RDURG submitted 2025-09-03 cs.HC cs.AIcs.CY

classification cs.HCcs.AIcs.CY
keywords LLMagentschoicedistinctivenessdiversityhomogenizationpersonalizationrevealedpreferencessocialmediaagenticAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to show that when large language models act on a person's behalf—choosing between options the person themselves might pick—the choices quietly become more mainstream and more concentrated. Using 110,000 real choices from 1,000 people's followed Facebook pages, the authors compare a generic agent, a personalized agent, and a random human baseline. Both agents select more popular, normative pages, reducing interpersonal distinctiveness; the generic agent does this more strongly. The personalized agent, which is meant to reflect the individual, instead shrinks intrapersonal diversity, narrowing the topics and psychological profiles of the pages chosen. If true, this means agentic AI does not merely save effort; it gradually flattens the individual and collective range of tastes and identities.

What carries the argument

The engine of the result is a matched three-way choice-set design: for each user, 100 of their followed pages are randomly paired into 50 binary decisions; a generic prompt and a personalized prompt (age, gender, ten sample pages, and a 20-line profile summary generated from status updates) each pick one page per pair, and the 'Human Control' is a random draw from the same pairs. Because both options in every pair come from the user's own actual following, any deviation of the agents' picks from the random draw is attributable to the agent rather than to unfamiliarity with the user's tastes. The outcome measures—inverse popularity for distinctiveness, normalized entropy for topical diversity

What would settle it

Ask a sample of the same 1,000 users to make real binary choices between the same 50 pairs of pages they follow, without any AI, and compare the popularity and diversity of their actual choices to the agents' picks. If actual human choices match the agents in popularity bias and narrowness, the reported flattening is an artifact of the random baseline; if human choices are more varied and less popular-oriented, the agents' flattening is confirmed.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that LLM-based agents exhibit a systematic 'basic' bias: asked to choose between two pages a user actually follows, a generic agent and a personalized agent both favor pages with higher overall popularity, and this reduces how distinct a person's chosen set is from the population. The generic agent shows a drop in distinctiveness more than 2.5 times larger than the personalized agent (d = 0.21 vs. 0.08). But the personalized agent comes with a different cost: it compresses the breadth of a person's choices, lowering both topical diversity measured by normalized entropy (d = -0.31) and psychological interest diversity based on the Big Five profil

Load-bearing premise

The load-bearing premise is that a random draw from a person's own followed pages is the right 'human' baseline; if people's real choices between those pages are themselves biased toward popular options, the AI-induced drop in distinctiveness would shrink or vanish.

Editorial extensions

If this is right

  • If people delegate identity-relevant choices to LLM agents, their chosen preference sets will converge toward population norms even when the options are all personally relevant.
  • Switching from a generic to a personalized agent does not remove the flattening; it shifts the damage from between-person distinctiveness to within-person diversity.
  • Users who rely on agents for many choices over time will explore fewer topics and fewer psychological niches than they otherwise would, potentially narrowing their own identity portfolios.
  • Designers can treat distinctiveness and diversity as explicit objectives—e.g., exploration dials or diversity-aware optimization—rather than assuming personalization preserves individuality.
  • Aggregate cultural diversity may decline even though each individual's choices remain within their own established taste set, because the same popular options are repeatedly favored.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The random-draw baseline is a counterfactual of 'no agent,' not a measure of what the user would actually pick; if real users also gravitate to popular pages, the size of the AI-induced flattening relative to actual human behavior may be smaller than the reported effect sizes suggest.
  • The same mechanism would predict that LLM agents in other domains—restaurants, travel, news, music—favor mainstream, high-frequency options; this is directly testable with existing recommendation logs.
  • Raising the LLM's temperature or adding explicit diversity prompts should reduce the flattening, offering a cheap intervention; the paper's default-temperature design estimates the out-of-the-box effect, not the ceiling of what agents could do.
  • The distinctiveness-diversity trade-off suggests a shared evaluation problem: judging agents on accuracy or user satisfaction alone will miss the slow erosion of choice variety that this design exposes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper investigates whether LLM-based agents—generic and personalized—reduce the interpersonal distinctiveness and intrapersonal diversity of people's choices. Using 1,000 Facebook users from the myPersonality dataset, the authors prompt GPT-4o to make 50 binary choices between pairs of pages each user actually followed. They compare these AI selections to a 'Human Control' baseline consisting of a random selection from the same set of followed pages. Distinctiveness is measured as the inverse popularity of chosen pages; diversity is measured via normalized Shannon entropy over page categories and via the standard deviation of the Big Five profiles associated with chosen pages. Paired t-tests show that AI agents select more popular, less distinct pages than the random baseline, with the generic agent showing a larger distinctiveness reduction, while the personalized agent more strongly reduces both topical and psychological diversity. The authors interpret these results as evidence that delegating identity-relevant choices to LLM agents flattens human experience and creates a distinctiveness-diversity trade-off between generic and personalized agents.

Significance. If the central comparison were against a genuine human decision baseline, this would be an important contribution to the growing literature on AI agency and identity. The paper leverages a large, real-world behavioral dataset, uses a simple and transparent matched-set design, and measures outcomes that are meaningful for the stated research question. The distinction between generic and personalized agents is also valuable, as it reveals a non-trivial trade-off between uniformity across people and breadth within a person. However, the central inferential leap—from 'LLM selects more popular pages than a random draw' to 'AI reduces the distinctiveness and diversity of people's choices'—is not supported by the current design. The statistical machinery is appropriate for comparing three matched choice sets, but the 'Human Control' is not a human chooser, so the headline claim overstates what is measured. The paper is a useful demonstration of LLM output bias relative to a user's own average followed-page popularity, but its broader implications for human choice require additional evidence or a substantial reframing.

major comments (3)
  1. [Methods, Analytical procedure (Human Control)] The 'Human Control' baseline is a random draw from the user's own followed pages, not a human binary decision. The statement that 'this bootstrap reflects revealed preferences absent any agent intervention' conflates a portfolio of followed pages with a choice. A user's actual choice between two already-followed pages may itself be popularity-biased (e.g., users may click or engage more with well-known pages within their own portfolio). If so, the gap between AI selections and actual human selections is smaller than the reported comparison to a random draw suggests. The paired t-tests (t(999) ranging from -4 to -16, d = 0.08-0.31) quantify deviations from a random baseline, not from human behavior. This is load-bearing for the title/abstract claim that AI 'reduces the distinctiveness and diversity of people's choices.' Please either add a validation study with real human choices on the s
  2. [Discussion, limitations and Results] The Discussion acknowledges: 'our methodology does not directly test whether individuals would accept AI-generated choices.' This is precisely the missing link. Restricting choices to pages users already liked ensures both options reflect actual preferences, but it does not establish what a human would choose between them. The claimed reduction in distinctiveness/diversity of 'people's choices' is therefore not directly evidenced; only a property of LLM outputs is measured relative to the user's own average page popularity. The authors should either soften the causal and practical language throughout the abstract, significance statement, and introduction, or provide an empirical calibration of the random baseline to human choice behavior.
  3. [Results, distinctiveness and diversity analyses] The effect sizes reported (Cohen's d from 0.08 to 0.31) are small to moderate, yet the Discussion and abstract use strong language such as 'reshape who people become' and 'flattening of the human experience.' Given that the baseline is random rather than human, the practical significance of these effect sizes is even less clear. The authors should provide a more careful interpretation of effect sizes in light of the baseline issue, rather than drawing strong societal conclusions from statistically significant but small differences.
minor comments (6)
  1. [Results, first paragraph] Typo: 't(999) = =-4.08' contains a double equals sign.
  2. [Measures, Psychological interest diversity] The text refers to 'Fig. 3 for a visual illustration of the operationalization' but later refers to 'Fig. 4 for a visual illustration of the process.' The figure references should be consistent and correct.
  3. [Methods, Data] The paper reports that 42,003 observations were retained (~42 per participant) after excluding invalid responses, but the statistical analyses are presented as paired t-tests with 999 degrees of freedom. Please clarify how the per-user scores (median popularity, entropy, psychological diversity) are computed when the number of valid choices varies across users, and whether the reported results are robust to this unbalanced structure.
  4. [Throughout] The model is referred to as 'GPT-4.o' in several places; the official name is 'GPT-4o'. Please correct for consistency.
  5. [Discussion, final paragraph] The phrase 'loss of intrapersonal distinctiveness or intrapersonal diversity' appears to contain a typo; it should likely read 'interpersonal distinctiveness or intrapersonal diversity.'
  6. [Significance Statement] Minor editorial issue: 'the multidimensionality of individual' should be 'the multidimensionality of individuals' or 'the multidimensionality of the individual.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the LLM-choice comparisons are empirical measurements, with no parameter fitted to the outcome and no prediction derived by construction from the baseline.

full rationale

The paper compares three matched choice sets (Generic AI, Personalized AI, Human Control) on popularity-based distinctiveness and two diversity metrics. The metrics are computed from the choice sets and from the myPersonality page-following data; no parameter is fitted to the LLM outputs. The LLM's choices are not constructed from the popularity or diversity measures, so the finding that AI selects more popular pages is an empirical result, not a tautology. The 'Human Control' baseline is a random draw from each user's own followed pages, which the authors interpret as a revealed-preference baseline. That interpretation is a limitation of external validity (the baseline is not an actual human binary choice), but it is not circular: the AI's choices could have gone either way relative to random selection, and the reported t-tests quantify that deviation. The self-citations (refs 13, 16, 17, 18, 20) are contextual, prior evidence for LLM preference inference, or the source of the psychological-diversity measure; none supplies a uniqueness theorem, an ansatz, or a fitted parameter that forces the result. The Discussion explicitly acknowledges that the study does not test whether individuals would accept AI-generated choices ('our methodology does not directly test whether individuals would accept AI-generated choices'), which is a missing piece of support for the title's implication about people's choices but is not a circular step. The derivation chain is self-contained as a measurement of LLM output bias relative to the user's own page portfolio.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters or invented entities. The paper's inferences rest on domain assumptions about the baseline, proxy validity, and LLM usage.

assumptions (5)
  • domain assumption Random selection among a user's own followed pages is a valid baseline for human revealed preference.
    The 'Human Control' is defined as random selection (Methods, Analytical procedure). The central claim compares AI choices to this baseline, so this assumption is load-bearing.
  • domain assumption Facebook page-following is a valid proxy for real-world preferences.
    The paper treats page likes as 'revealed-preference proxy for everyday tastes' (Introduction).
  • domain assumption The popularity of a page within the myPersonality sample is a valid measure of normativeness.
    Distinctiveness is measured as inverse within-sample popularity (Methods, Choice Distinctiveness).
  • domain assumption GPT-4o's default sampling parameters reflect typical user behavior.
    The authors keep temperature at default, assuming average users won't change settings (Methods, LLM Prompting).
  • domain assumption The Big Five follower-profile approach captures psychological diversity.
    Psychological interest diversity uses the method from Matz (2021), referenced as ref 20.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices." pith.science (2026). https://pith.science/paper/NK5RDURG

@misc{pith2026250902910,
  author       = {Pith},
  title        = {Pith review of: The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NK5RDURG}},
  note         = {Machine review of arXiv:2509.02910}
}
read the original abstract

Large language models (LLMs) increasingly act on people's behalf: they write emails, buy groceries, and book restaurants. While the outsourcing of human decision-making to AI can be both efficient and effective, it raises a fundamental question: how does delegating identity-defining choices to AI reshape who people become? We study the impact of agentic LLMs on two identity-relevant outcomes: interpersonal distinctiveness - how unique a person's choices are relative to others - and intrapersonal diversity - the breadth of a single person's choices over time. Using real choices drawn from social-media behavior of 1,000 U.S. users (110,000 choices in total), we compare a generic and personalized agent to a human baseline. Both agents shift people's choices toward more popular options, reducing the distinctiveness of their behaviors and preferences. While the use of personalized agents tempers this homogenization (compared to the generic AI), it also more strongly compresses the diversity of people's preference portfolios by narrowing what they explore across topics and psychological affinities. Understanding how AI agents might flatten human experience, and how using generic versus personalized agents involves distinctiveness-diversity trade-offs, is critical for designing systems that augment rather than constrain human agency, and for safeguarding diversity in thought, taste, and expression.

Figures

Figures reproduced from arXiv: 2509.02910 by the authors.

Figure 1
Figure 1. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 1
Figure 1. A) Two hypothetical preference distributions – one for the overall population p and one for a particular individual i within that population – in a baseline state. B) Loss of personal distinctiveness: The individual distribution is pulled towards the population mean. C) Loss of intrapersonal diversity: The variance of the individual distribution is shrinking. D) Loss of both interpersonal distinctiveness and intrape… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

22 extracted references · 18 canonical work pages

  1. [1]

    Lai, V ., Chen, C., Smith-Renner, A., Liao, Q. V . & Tan, C. Towards a Science of Human-AI Decision Making: An Overview of Design Space in Empirical Human-Subject Studies. in 2023 ACM Conference on Fairness Accountability and Transparency 1369–1385 (ACM, Chicago IL USA, 2023). doi:10.1145/3593013.3594087

  2. [2]

    Park, J. S. et al. Generative Agents: Interactive Simulacra of Human Behavior. in Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology 1–22 (ACM, San Francisco CA USA, 2023). doi:10.1145/3586183.3606763

  3. [3]

    Yao, S. et al. React: Synergizing reasoning and acting in language models. in International Conference on Learning Representations (ICLR) (2023)

  4. [4]

    Rajpurkar, P. et al. CheXaid: deep learning assistance for physician diagnosis of tuberculosis using chest x-rays in patients with HIV . NPJ Digit. Med. 3, 115 (2020)

  5. [5]

    & Papenbrock, J

    Bussmann, N., Giudici, P., Marinelli, D. & Papenbrock, J. Explainable Machine Learning in Credit Risk Management. Comput. Econ. 57, 203–216 (2021)

  6. [6]

    Benjamin, D. M. et al. Hybrid forecasting of geopolitical events†. AI Mag. 44, 112–128 (2023)

  7. [7]

    Brown, T. et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 33, 1877–1901 (2020)

  8. [8]

    Doshi, A. R. & Hauser, O. P. Generative AI enhances individual creativity but reduces the collective diversity of novel content. Sci. Adv. 10, eadn5290 (2024)

Show all 22 references
  1. [9]

    & Kushlev, K

    Moon, K., Green, A. & Kushlev, K. Homogenizing Effect of Large Language Model (LLM) on Creative Diversity: An Empirical Comparison. (2024)

  2. [10]

    R., Shah, J

    Anderson, B. R., Shah, J. H. & Kreminski, M. Homogenization Effects of Large Language Models on Human Creative Ideation. in Creativity and Cognition 413–425 (ACM, Chicago IL USA, 2024). doi:10.1145/3635636.3656204

  3. [11]

    Sourati, Z. et al. The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models. Preprint at https://doi.org/10.48550/arXiv.2502.11266 (2025)

  4. [12]

    Padmakumar, V . & He, H. Does Writing with Language Models Reduce Content Diversity? Preprint at https://doi.org/10.48550/arXiv.2309.05196 (2024)

  5. [13]

    & Rhue, L

    Goethals, S. & Rhue, L. One world, one opinion? The superstar effect in LLM responses. Preprint at https://doi.org/10.48550/arXiv.2412.10281 (2024)

  6. [14]

    & Belinkov, Y

    Shur-Ofry, M., Horowitz-Amsalem, B., Rahamim, A. & Belinkov, Y . Growing a Tail: Increasing Output Diversity in Large Language Models. Preprint at https://doi.org/10.48550/arXiv.2411.02989 (2024)

  7. [15]

    Miller, M. E. & Spatz, E. A unified view of a human digital twin. Hum.-Intell. Syst. Integr. 4, 23–33 (2022)

  8. [16]

    & Matz, S

    Peters, H. & Matz, S. C. Large language models can infer psychological dispositions of social media users. PNAS Nexus 3, pgae231 (2024)

  9. [17]

    & Matz, S

    Goethals, S., Luther, J. & Matz, S. Words reveal wants: How well can simple LLM-based AI agents replicate people’s choices based on their social media posts. in Adjunct Proceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization 126–131 (ACM, New Yo...

  10. [18]

    C., Gosling, S

    Kosinski, M., Matz, S. C., Gosling, S. D., Popov, V . & Stillwell, D. Facebook as a research tool for the social sciences: Opportunities, challenges, ethical considerations, and practical guidelines. Am. Psychol. 70, 543 (2015)

  11. [19]

    T., Hui, P.-M., Harper, F

    Nguyen, T. T., Hui, P.-M., Harper, F. M., Terveen, L. & Konstan, J. A. Exploring the filter bubble: the effect of using recommender systems on content diversity. in Proceedings of the 23rd international conference on World wide web 677–686 (ACM, Seoul Korea, 2014). doi:10.1145...

  12. [20]

    Matz, S. C. Personal echo chambers: Openness-to-experience is linked to higher levels of psychological interest diversity in large-scale behavioral data. J. Pers. Soc. Psychol. 121, 1284 (2021)

  13. [21]

    John, O. P. & Srivastava, S. The Big-Five trait taxonomy: History, measurement, and theoretical perspectives. (1999)

  14. [22]

    Ozer, D. J. & Benet-Martínez, V . Personality and the Prediction of Consequential Outcomes. Annu. Rev. Psychol. 57, 401–421 (2006)

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.