Pith. sign in

REVIEW 16 cited by

LLM Generated Persona is a Promise with a Catch

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.16527 v1 pith:CSMOOQQR submitted 2025-03-18 cs.CL cs.AIcs.CYcs.SI

LLM Generated Persona is a Promise with a Catch

classification cs.CL cs.AIcs.CYcs.SI
keywords personagenerationpersonassignificantanalysisbiasesgeneratedincluding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The use of large language models (LLMs) to simulate human behavior has gained significant attention, particularly through personas that approximate individual characteristics. Persona-based simulations hold promise for transforming disciplines that rely on population-level feedback, including social science, economic analysis, marketing research, and business operations. Traditional methods to collect realistic persona data face significant challenges. They are prohibitively expensive and logistically challenging due to privacy constraints, and often fail to capture multi-dimensional attributes, particularly subjective qualities. Consequently, synthetic persona generation with LLMs offers a scalable, cost-effective alternative. However, current approaches rely on ad hoc and heuristic generation techniques that do not guarantee methodological rigor or simulation precision, resulting in systematic biases in downstream tasks. Through extensive large-scale experiments including presidential election forecasts and general opinion surveys of the U.S. population, we reveal that these biases can lead to significant deviations from real-world outcomes. Our findings underscore the need to develop a rigorous science of persona generation and outline the methodological innovations, organizational and institutional support, and empirical foundations required to enhance the reliability and scalability of LLM-driven persona simulations. To support further research and development in this area, we have open-sourced approximately one million generated personas, available for public access and analysis at https://huggingface.co/datasets/Tianyi-Lab/Personas.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale

    cs.CL 2026-06 unverdicted novelty 7.0

    Introduces GenAI agent framework for auditing personalization algorithms via synthetic accounts with fixed personas, applied to X post-2024 election showing amplification of toxic and right-leaning content varying by ...

  2. Adaptive Querying with AI Persona Priors

    stat.ML 2026-05 unverdicted novelty 7.0

    A persona-induced latent variable model with LLM response distributions enables closed-form Bayesian updates and finite-mixture predictions for scalable adaptive querying of user-dependent quantities.

  3. Adaptive Querying with AI Persona Priors

    stat.ML 2026-05 unverdicted novelty 7.0

    A persona-induced latent variable model with LLM-generated priors enables scalable adaptive item selection with closed-form Bayesian updates for accurate user-specific predictions.

  4. Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys

    cs.AI 2026-04 unverdicted novelty 7.0

    A method using predicted rectification difficulty for optimal human sample allocation in LLM-augmented surveys captures 61-79% of theoretical efficiency gains and reduces MSE by 11% on two datasets without pilot data.

  5. Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

    cs.CL 2026-07 unverdicted novelty 6.0

    Persona-driven generations by LLMs in MCQA tasks exhibit instability that differs systematically by model family, size, domain, and prompt format.

  6. Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits

    physics.soc-ph 2026-05 accept novelty 6.0

    Minor perturbations in persona format, instruction framing, and network structure shift cooperation by up to 76 percentage points and polarization metrics consistently, showing that LLM social simulations require per-...

  7. Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys

    cs.AI 2026-04 accept novelty 6.0

    Predicting question-level rectification difficulty from text and allocating human labels by a square-root rule recovers most of the hybrid human–LLM survey efficiency gains without pilot data.

  8. Sentipolis: Emotion-Aware Agents for Social Simulations

    cs.AI 2026-01 unverdicted novelty 6.0

    Sentipolis equips LLM agents with continuous PAD emotional states, dual-speed dynamics, and memory coupling to improve emotional continuity and grounded behavior in social simulations.

  9. Graph-Based Alternatives to LLMs for Human Simulation

    cs.CL 2025-11 conditional novelty 6.0

    GEMS formulates close-ended human-behavior simulation as link prediction on a heterogeneous graph and matches or exceeds LLM performance with three orders of magnitude fewer parameters across three datasets and three ...

  10. Fund2Persona: A Framework for Building and Refining Financial Advisor Personas from Fund Disclosure Data

    cs.CL 2026-06 unverdicted novelty 5.0

    Fund2Persona creates and refines LLM personas for financial advisors from fund disclosure data, outperforming generic baselines on holdings reconstruction and commentary alignment.

  11. Fund2Persona: A Framework for Building and Refining Financial Advisor Personas from Fund Disclosure Data

    cs.CL 2026-06 unverdicted novelty 5.0

    Fund2Persona grounds and refines financial advisor personas in fund disclosures, holdings transitions, and manager commentary, showing improved performance on reconstruction and alignment tasks over generic baselines.

  12. The $\textit{Silicon Society}$ Cookbook: Design Space of LLM-based Social Simulations

    cs.MA 2026-04 unverdicted novelty 5.0

    The base LLM choice dominates simulation outcomes in LLM-based social networks, while other design parameters show either additive or complex interactive effects.

  13. Temperature and Persona Shape LLM Agent Consensus With Minimal Accuracy Gains in Qualitative Coding

    cs.CL 2025-07 unverdicted novelty 5.0

    Temperature and persona variations shape consensus speed in LLM multi-agent coding but produce no robust accuracy gains over single agents on human-annotated tutoring transcripts.

  14. Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation

    cs.CL 2026-07 conditional novelty 4.0

    Anamnesis packages backstory-conditioned LLM personas into an interactive open-source survey platform that better matches real human opinion distributions than demographic-list prompting on ATP and New Yorker tasks.

  15. Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language

    cs.CL 2026-06 unverdicted novelty 4.0

    Modifying nationality and language parameters in English-centric personas for mental health dialogues introduces clinical inconsistencies across languages and causes LLM judges to perform inaccurately on non-English d...

  16. The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences

    cs.CL 2025-09 unverdicted novelty 3.0

    The paper reduces a broad set of prompt engineering techniques to six core approaches and applies them to life sciences use cases while addressing common LLM pitfalls.