{"id":"c3e04b42-e03d-4e78-b974-1db5ce0b7062","arxiv_id":"2509.02910","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM-based agents asked to choose among a person's own Facebook likes select more popular and less diverse pages, reducing both interpersonal distinctiveness and intrapersonal diversity.","lead":"This paper tested whether AI agents make people's choices less unique and less varied. Using 110,000 real Facebook page choices from 1,000 people, it found that AI agents pick more popular, less diverse options than a random baseline.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'Human Control' baseline is a random draw from the user's own followed pages, not a human choice, so the claim that AI reduces the distinctiveness of people's choices overstates what is measured.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing concern: the 'Human Control' condition is a random draw from the user's own followed pages, not a human choice. This is the point on which the central claim depends. The paper's own framing treats the random bootstrap as reflecting 'revealed preferences absent any agent intervention,' but revealed preferences are the fact that the user followed both pages, not a prediction of what the user would choose if forced to pick between them. If actual human choices are popularity-biased within the user's own portfolio, the AI-versus-human difference in distinctiveness would be attenuated; if human choices also tend to be topically concentrated, the diversity comparisons would change as well. The statistical results are internally consistent as a comparison of LLM outputs to a random baseline; the problem is construct validity of the 'human baseline,' which is a necessary condition for the title claim. This does not require rejecting the paper outright, because the random-baseline comparison is a meaningful and clean null model, and the authors could either reframe the findings accordingly or collect real human choices. The reader's conditional verdict already captures this need for revision, so no further verdict change is warranted. I agree with the reader rather than partially or fully disagreeing: the same worry is the most load-bearing, and no additional independent concern outweighs it.","tokens_in":6703,"tokens_out":4140,"duration_ms":53528,"concrete_test":"Obtain actual binary choices from a sample of participants using the same 50-pair design: for each participant, present the same pairs of their own followed pages and ask them to pick one. Compute the median page popularity and the two diversity metrics (topical entropy, psychological interest diversity) on the human-chosen sets, then rerun the paired comparisons against Generic AI and Personalized AI. If the human-chosen sets are more popular or less diverse than the random bootstrap, the reported effect sizes shrink; if they are indistinguishable from random, the current conclusions hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The decisive control in the paper is a random sample of the user's own followed pages (Methods, 'Analytical procedure'). The paper calls this 'Human Control' and asserts that 'this bootstrap reflects revealed preferences absent any agent intervention.' That conflates a user's portfolio of followed pages with a choice. A real binary choice between two already-followed pages is not a random draw; it may itself be popularity-biased (e.g., users may disproportionately click or engage with well-known pages within their own portfolio). If so, the gap between AI selections and actual human selections is smaller than the reported comparison to a random draw suggests. The paired t-tests (t(999) range from -4 to -16, d 0.08-0.31) quantify deviations from a random baseline, not from human behavior. The Discussion's caveat that 'both options reflect actual preferences' addresses whether options are valid preference indicators, not whether the baseline is a human decision process. The title and abstract claim that AI 'reduces the distinctiveness and diversity of people's choices' therefore overstates the evidence. The finding is secure as a statement about LLM output bias relative to the user's own average page popularity, but it is not yet a demonstration of an effect on human choice. The authors' own acknowledged limitation ('our methodology does not directly test whether individuals would accept AI-generated choices') is precisely where the load-bearing assumption sits.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether LLM-based agents—generic and personalized—reduce the interpersonal distinctiveness and intrapersonal diversity of people's choices. Using 1,000 Facebook users from the myPersonality dataset, the authors prompt GPT-4o to make 50 binary choices between pairs of pages each user actually followed. They compare these AI selections to a 'Human Control' baseline consisting of a random selection from the same set of followed pages. Distinctiveness is measured as the inverse popularity of chosen pages; diversity is measured via normalized Shannon entropy over page categories and via the standard deviation of the Big Five profiles associated with chosen pages. Paired t-tests show that AI agents select more popular, less distinct pages than the random baseline, with the generic agent showing a larger distinctiveness reduction, while the personalized agent more strongly reduces both topical and psychological diversity. The authors interpret these results as evidence that delegating identity-relevant choices to LLM agents flattens human experience and creates a distinctiveness-diversity trade-off between generic and personalized agents.","tokens_in":7026,"tokens_out":5609,"duration_ms":68481,"significance":"If the central comparison were against a genuine human decision baseline, this would be an important contribution to the growing literature on AI agency and identity. The paper leverages a large, real-world behavioral dataset, uses a simple and transparent matched-set design, and measures outcomes that are meaningful for the stated research question. The distinction between generic and personalized agents is also valuable, as it reveals a non-trivial trade-off between uniformity across people and breadth within a person. However, the central inferential leap—from 'LLM selects more popular pages than a random draw' to 'AI reduces the distinctiveness and diversity of people's choices'—is not supported by the current design. The statistical machinery is appropriate for comparing three matched choice sets, but the 'Human Control' is not a human chooser, so the headline claim overstates what is measured. The paper is a useful demonstration of LLM output bias relative to a user's own average followed-page popularity, but its broader implications for human choice require additional evidence or a substantial reframing.","major_comments":[{"comment":"The 'Human Control' baseline is a random draw from the user's own followed pages, not a human binary decision. The statement that 'this bootstrap reflects revealed preferences absent any agent intervention' conflates a portfolio of followed pages with a choice. A user's actual choice between two already-followed pages may itself be popularity-biased (e.g., users may click or engage more with well-known pages within their own portfolio). If so, the gap between AI selections and actual human selections is smaller than the reported comparison to a random draw suggests. The paired t-tests (t(999) ranging from -4 to -16, d = 0.08-0.31) quantify deviations from a random baseline, not from human behavior. This is load-bearing for the title/abstract claim that AI 'reduces the distinctiveness and diversity of people's choices.' Please either add a validation study with real human choices on the s","section":"Methods, Analytical procedure (Human Control)"},{"comment":"The Discussion acknowledges: 'our methodology does not directly test whether individuals would accept AI-generated choices.' This is precisely the missing link. Restricting choices to pages users already liked ensures both options reflect actual preferences, but it does not establish what a human would choose between them. The claimed reduction in distinctiveness/diversity of 'people's choices' is therefore not directly evidenced; only a property of LLM outputs is measured relative to the user's own average page popularity. The authors should either soften the causal and practical language throughout the abstract, significance statement, and introduction, or provide an empirical calibration of the random baseline to human choice behavior.","section":"Discussion, limitations and Results"},{"comment":"The effect sizes reported (Cohen's d from 0.08 to 0.31) are small to moderate, yet the Discussion and abstract use strong language such as 'reshape who people become' and 'flattening of the human experience.' Given that the baseline is random rather than human, the practical significance of these effect sizes is even less clear. The authors should provide a more careful interpretation of effect sizes in light of the baseline issue, rather than drawing strong societal conclusions from statistically significant but small differences.","section":"Results, distinctiveness and diversity analyses"}],"minor_comments":[{"comment":"Typo: 't(999) = =-4.08' contains a double equals sign.","section":"Results, first paragraph"},{"comment":"The text refers to 'Fig. 3 for a visual illustration of the operationalization' but later refers to 'Fig. 4 for a visual illustration of the process.' The figure references should be consistent and correct.","section":"Measures, Psychological interest diversity"},{"comment":"The paper reports that 42,003 observations were retained (~42 per participant) after excluding invalid responses, but the statistical analyses are presented as paired t-tests with 999 degrees of freedom. Please clarify how the per-user scores (median popularity, entropy, psychological diversity) are computed when the number of valid choices varies across users, and whether the reported results are robust to this unbalanced structure.","section":"Methods, Data"},{"comment":"The model is referred to as 'GPT-4.o' in several places; the official name is 'GPT-4o'. Please correct for consistency.","section":"Throughout"},{"comment":"The phrase 'loss of intrapersonal distinctiveness or intrapersonal diversity' appears to contain a typo; it should likely read 'interpersonal distinctiveness or intrapersonal diversity.'","section":"Discussion, final paragraph"},{"comment":"Minor editorial issue: 'the multidimensionality of individual' should be 'the multidimensionality of individuals' or 'the multidimensionality of the individual.'","section":"Significance Statement"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and interesting question, and the dataset is unusually rich for this kind of study. However, the central comparison is not against a human decision maker, and the title/abstract overstate the findings. I believe a major revision is warranted: the authors should either collect a real human binary-choice baseline (even on a subset of pairs) to calibrate the random baseline, or substantially reframe the claims as being about LLM selection relative to the user's own portfolio. The current manuscript, as written, would mislead readers about the strength of the evidence. The informal title may also be a poor fit for a serious journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper is worth reading, and the new angle is real. Prior homogenization work has mostly been about LLMs making the collective output more similar. This moves to agentic delegation and looks at what happens inside one person's portfolio of choices, and it separates generic from personalized agents. That's a meaningful extension, and the dataset is substantial: 1,000 people, 110,000 real Facebook page choices. The operationalizations are clear, the paired t-tests are appropriate for the per-person aggregates, and the Discussion is honest about limits.\n\nWhat the paper does well: it shows that both generic and personalized GPT-4.o agents, when choosing between pages a user actually follows, pick more popular pages than a random draw from the same user's pages. It also shows a distinctiveness–diversity trade-off: the personalized agent preserves distinctiveness better than the generic one but compresses within-person topical and psychological diversity more. Effect sizes are small (d 0.08–0.31), but for a behavioral intervention on preferences that's not disqualifying.\n\nThe soft spot is exactly where the stress-test points: the Human Control baseline is a bootstrap random draw from the user's own followed pages, not a human decision between the same pairs. The paper calls that 'revealed preferences absent any agent intervention,' but a portfolio of followed pages is not a choice process. Users who follow popular pages might well click popular pages when forced to choose between two already-followed options. So the measured gap between AI and random says something about LLM output bias relative to the user's average page popularity, not necessarily about what the human would have chosen. The authors themselves concede in the limitations that they don't directly test whether people would accept AI-generated choices, and that concession sits right on top of the load-bearing assumption. The title and abstract, promising 'reduces the distinctiveness and diversity of people's choices,' go beyond what is actually demonstrated. That's a real gap, but it's fixable: either collect actual human choices for a subsample of the pairs, or reframe the claims as relative to a random baseline and discuss what that means.\n\nWho gets value: people working on agentic AI, algorithmic recommendation, and the social consequences of LLM delegation. It deserves a serious referee. My recommendation is peer review with major revision, and the reviewers should make the baseline issue non-negotiable. A clean version of this finding—whether framed as LLM bias relative to random, or with real human choices—would be a useful contribution.","headline":"A solid, honest empirical study with a genuinely new angle on agentic AI and individual choice diversity, but the headline claim overstates the evidence because the 'Human Control' baseline is a random draw, not a human choice.","tokens_in":7474,"tokens_out":1124,"would_cite":true,"duration_ms":15991,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Delegating choices to LLM agents—generic or personalized—pulls preferences toward the popular and shrinks variety.","keywords":["LLM agents","choice distinctiveness","choice diversity","homogenization","personalization","revealed preferences","social media","agentic AI"],"falsifier":"Ask a sample of the same 1,000 users to make real binary choices between the same 50 pairs of pages they follow, without any AI, and compare the popularity and diversity of their actual choices to the agents' picks. If actual human choices match the agents in popularity bias and narrowness, the reported flattening is an artifact of the random baseline; if human choices are more varied and less popular-oriented, the agents' flattening is confirmed.","tokens_in":6630,"feed_emoji":"🤖","tokens_out":7116,"duration_ms":75730,"temperature":0.7,"pith_summary":"This paper sets out to show that when large language models act on a person's behalf—choosing between options the person themselves might pick—the choices quietly become more mainstream and more concentrated. Using 110,000 real choices from 1,000 people's followed Facebook pages, the authors compare a generic agent, a personalized agent, and a random human baseline. Both agents select more popular, normative pages, reducing interpersonal distinctiveness; the generic agent does this more strongly. The personalized agent, which is meant to reflect the individual, instead shrinks intrapersonal diversity, narrowing the topics and psychological profiles of the pages chosen. If true, this means agentic AI does not merely save effort; it gradually flattens the individual and collective range of tastes and identities.","feed_headline":"AI agents make your choices more 'basic'","feed_subtitle":"Between your own followed pages, both generic and personalized agents picked more popular options and narrowed variety.","key_machinery":"The engine of the result is a matched three-way choice-set design: for each user, 100 of their followed pages are randomly paired into 50 binary decisions; a generic prompt and a personalized prompt (age, gender, ten sample pages, and a 20-line profile summary generated from status updates) each pick one page per pair, and the 'Human Control' is a random draw from the same pairs. Because both options in every pair come from the user's own actual following, any deviation of the agents' picks from the random draw is attributable to the agent rather than to unfamiliarity with the user's tastes. The outcome measures—inverse popularity for distinctiveness, normalized entropy for topical diversity","core_discovery":"On its own terms, the paper's central claim is that LLM-based agents exhibit a systematic 'basic' bias: asked to choose between two pages a user actually follows, a generic agent and a personalized agent both favor pages with higher overall popularity, and this reduces how distinct a person's chosen set is from the population. The generic agent shows a drop in distinctiveness more than 2.5 times larger than the personalized agent (d = 0.21 vs. 0.08). But the personalized agent comes with a different cost: it compresses the breadth of a person's choices, lowering both topical diversity measured by normalized entropy (d = -0.31) and psychological interest diversity based on the Big Five profil","pith_inferences":["The random-draw baseline is a counterfactual of 'no agent,' not a measure of what the user would actually pick; if real users also gravitate to popular pages, the size of the AI-induced flattening relative to actual human behavior may be smaller than the reported effect sizes suggest.","The same mechanism would predict that LLM agents in other domains—restaurants, travel, news, music—favor mainstream, high-frequency options; this is directly testable with existing recommendation logs.","Raising the LLM's temperature or adding explicit diversity prompts should reduce the flattening, offering a cheap intervention; the paper's default-temperature design estimates the out-of-the-box effect, not the ceiling of what agents could do.","The distinctiveness-diversity trade-off suggests a shared evaluation problem: judging agents on accuracy or user satisfaction alone will miss the slow erosion of choice variety that this design exposes."],"forward_implications":["If people delegate identity-relevant choices to LLM agents, their chosen preference sets will converge toward population norms even when the options are all personally relevant.","Switching from a generic to a personalized agent does not remove the flattening; it shifts the damage from between-person distinctiveness to within-person diversity.","Users who rely on agents for many choices over time will explore fewer topics and fewer psychological niches than they otherwise would, potentially narrowing their own identity portfolios.","Designers can treat distinctiveness and diversity as explicit objectives—e.g., exploration dials or diversity-aware optimization—rather than assuming personalization preserves individuality.","Aggregate cultural diversity may decline even though each individual's choices remain within their own established taste set, because the same popular options are repeatedly favored."],"supporting_citations":[{"why":"Supplies the 1,000 users' followed pages, status updates, demographics, and personality scores from which all choice sets and baselines are built.","marker":"18"},{"why":"Shows LLMs can infer psychological dispositions from social media, justifying the personalized agent's profile-based prompts.","marker":"16"},{"why":"Demonstrates simple LLM agents can replicate people's choices from social media posts, grounding the personalized agent design.","marker":"17"},{"why":"Provides the psychological interest diversity metric used to measure intrapersonal diversity.","marker":"20"},{"why":"Establishes LLMs as probabilistic models that favor statistically likely continuations, the mechanism behind the popularity bias.","marker":"7"},{"why":"Prior evidence that generative AI reduces collective creative diversity, the between-person effect this paper extends to individual choice.","marker":"8"},{"why":"Shows LLM responses concentrate on a small set of 'superstar' answers, direct support for the popularity-seeking tendency in choice.","marker":"13"}],"fun_headline_variants":["AI agents push choices toward the mainstream","Generic AI flattens your taste more than personalized","LLM agents make you less unique and less varied","Personalized AI narrows your interests more than generic"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that a random draw from a person's own followed pages is the right 'human' baseline; if people's real choices between those pages are themselves biased toward popular options, the AI-induced drop in distinctiveness would shrink or vanish.","fun_headline_variants_meta":{"raw":{"variants":["AI agents push choices toward the mainstream","Generic AI flattens your taste more than personalized","LLM agents make you less unique and less varied","Personalized AI narrows your interests more than generic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000532,"raw_usage":{"total_tokens":2411,"prompt_tokens":769,"completion_tokens":1642,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":1594}},"tokens_in":513,"tokens_out":1642,"duration_ms":12495,"temperature":1.0,"reasoning_tokens":1594,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T11:15:40.989351+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask a sample of the same 1,000 users to make real binary choices between the same 50 pairs of pages they follow, without any AI, and compare the popularity and diversity of their actual choices to the agents' picks. If actual human choices match the agents in popularity bias and narrowness, the reported flattening is an artifact of the random baseline; if human choices are more varied and less popular-oriented, the agents' flattening is confirmed.","supporting_citations":[{"cited_title":"C., Gosling, S","cited_arxiv_id":null,"evidence_quote":"Supplies the 1,000 users' followed pages, status updates, demographics, and personality scores from which all choice sets and baselines are built."},{"cited_title":"& Matz, S","cited_arxiv_id":null,"evidence_quote":"Shows LLMs can infer psychological dispositions from social media, justifying the personalized agent's profile-based prompts."},{"cited_title":"& Matz, S","cited_arxiv_id":null,"evidence_quote":"Demonstrates simple LLM agents can replicate people's choices from social media posts, grounding the personalized agent design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the psychological interest diversity metric used to measure intrapersonal diversity."},{"cited_title":"One world, one opinion? The superstar effect in LLM responses","cited_arxiv_id":"2412.10281","evidence_quote":"Shows LLM responses concentrate on a small set of 'superstar' answers, direct support for the popularity-seeking tendency in choice."}],"review_version":1}