Pith. sign in

REVIEW 8 cited by

Quantifying the Persona Effect in LLM Simulations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10811 v2 pith:T5ZR7CZT submitted 2024-02-16 cs.CL cs.CY

classification cs.CLcs.CY
keywords personapromptingvariablesannotationsllmsdatasetsfindhuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have shown remarkable promise in simulating human language and behavior. This study investigates how integrating persona variables-demographic, social, and behavioral factors-impacts LLMs' ability to simulate diverse perspectives. We find that persona variables account for <10% variance in annotations in existing subjective NLP datasets. Nonetheless, incorporating persona variables via prompting in LLMs provides modest but statistically significant improvements. Persona prompting is most effective in samples where many annotators disagree, but their disagreements are relatively minor. Notably, we find a linear relationship in our setting: the stronger the correlation between persona variables and human annotations, the more accurate the LLM predictions are using persona prompting. In a zero-shot setting, a powerful 70b model with persona prompting captures 81% of the annotation variance achievable by linear regression trained on ground truth annotations. However, for most subjective NLP datasets, where persona variables have limited explanatory power, the benefits of persona prompting are limited.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice

    cs.CL 2026-07 conditional novelty 6.0 of 10

    MyMentorLLM generates 2,100 multimodal CBT training sessions and finds that native speech-to-speech simulation matches human therapy competence scores while smaller models overestimate trainee skills and suffer from h...

  2. Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation

    cs.CL 2025-11 conditional novelty 6.0 of 10

    Fine-tuning on speaker-attributed, action-tagged transcripts from public meetings lets LLM agents mimic government meeting participants well enough that human judges often cannot tell them from real people.

  3. Measuring AI Alignment with Human Flourishing

    cs.AI 2025-07 unverdicted novelty 6.0 of 10

    The authors propose the Flourishing AI Benchmark, which uses 1,229 objective and subjective questions plus LLM judges to score 28 chatbots across seven dimensions of human flourishing, and find none reach the 90-point...

  4. Aligning LLM with human travel choices: a persona-based embedding learning approach

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A persona-based embedding learning framework aligns LLM predictions with human travel mode choices, outperforming MNL and few-shot LLM baselines on the Swissmetro dataset.

  5. Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent Community

    cs.MA 2026-02 reject novelty 5.0 of 10

    Clustering of Moltbook submolt descriptions shows agent-created communities organize into human-mimetic, silicon-centric, and proto-economic themes, but the categories were partly prescribed by the analysis prompt.

  6. From Risk Perception to Behavior Large Language Models-Based Simulation of Pandemic Prevention Behaviors

    cs.SI 2026-01 reject novelty 5.0 of 10

    LLM-based simulations of pandemic prevention behaviors show moderate distributional agreement with Beijing survey data, but the validation uses a lenient threshold and selective reporting.

  7. Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

    cs.AI 2025-10 unverdicted novelty 5.0 of 10

    Proposes SJTs and MIRT to measure consistent latent behavioral tendencies in LLMs, showing stability and predictive validity on external benchmarks.

  8. SimuPanel: A Novel Immersive Multi-Agent System to Simulate Interactive Expert Panel Discussion

    cs.HC 2025-06 conditional novelty 5.0 of 10

    A multi-agent LLM system called SimuPanel simulates expert panel discussions with personas grounded in public academic sources, and a small evaluation suggests the full reasoning pipeline produces higher LLM-judged di...

Pith tools