Pith. sign in

REVIEW 8 cited by

LLM-based Human Simulations Have Not Yet Been Reliable

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.08579 v3 pith:YMGRICRU submitted 2025-01-15 cs.CL

classification cs.CL
keywords humanllm-basedsimulationsdesigndiscrepanciesframeworklimitationsllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant discrepancies between their outcomes and authentic human actions. Our investigation begins with a systematic review of LLM-based human simulations in social, economic, policy, and psychological contexts, identifying their common frameworks, recent advances, and persistent limitations. This review reveals that such discrepancies primarily stem from inherent limitations of LLMs and flaws in simulation design, both of which are examined in detail. Building on these insights, we propose a systematic solution framework that emphasizes enriching data foundations, advancing LLM capabilities, and ensuring robust simulation design to enhance reliability. Finally, we introduce a structured algorithm that operationalizes the proposed framework, aiming to guide credible and human-aligned LLM-based simulations. To facilitate further research, we provide a curated list of related literature and resources at https://github.com/Persdre/awesome-llm-human-simulation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    LLMs match condition-level patterns in a noodle purchase survey but fail to replicate distributional structure, with no model beating a pooled human baseline for purchase quantities.

  2. Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    ConsumerSimBench evaluates 13 LLMs on reconstructing crowd reactions from 1,553 Chinese social-media topics using 23,122 auditable yes-no criteria, finding maximum coverage of 47.8% by Gemini-3.1-Pro.

  3. The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies

    cs.CL 2025-09 conditional novelty 7.0 of 10

    A systematic audit of LLM-based AI societies finds that 89.7% of 39 studies violate at least one of six PIMMUR validity principles, with reproductions showing that many claimed collective behaviors disappear when cont...

  4. Simulating Human Memory with Language Models

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Language models show superior memory to humans on psych experiments but can be adjusted via prompting and compaction to forget more human-like, yielding better user simulators.

  5. The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    LLMs organize prompted social roles along a dominant, stable, and causally steerable granularity axis in representation space that runs from micro to macro levels.

  6. Talking to an AI Mirror: Designing Self-Clone Chatbots for Enhanced Engagement in Digital Mental Health Support

    cs.HC 2025-09 conditional novelty 6.0 of 10

    Self-clone chatbots that mirror a user's support style showed higher emotional and cognitive engagement than a generic counselor chatbot, but only among the subgroup who found the clone believable.

  7. Prompt Optimization for User Simulation in Conversational Recommender Systems: A Multi-Objective Framework

    cs.IR 2026-05 unverdicted novelty 5.0 of 10

    A multi-objective prompt optimization framework for LLM user simulators in conversational recommender systems improves behavioral alignment with human patterns over baselines.

  8. We Need Strong Preconditions For Using Simulations In Policy

    cs.CY 2026-04 unverdicted novelty 4.0 of 10

    Societal-scale LLM agent simulations for policy need three preconditions: avoid neutral treatment of marginalized population simulations, require population participation, ensure accountability, plus development and d...

Pith tools