Pith. sign in

REVIEW 4 cited by

LLM-based Human Simulations Have Not Yet Been Reliable

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.08579 v3 pith:YMGRICRU submitted 2025-01-15 cs.CL

classification cs.CL
keywords humanllm-basedsimulationsdesigndiscrepanciesframeworklimitationsllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant discrepancies between their outcomes and authentic human actions. Our investigation begins with a systematic review of LLM-based human simulations in social, economic, policy, and psychological contexts, identifying their common frameworks, recent advances, and persistent limitations. This review reveals that such discrepancies primarily stem from inherent limitations of LLMs and flaws in simulation design, both of which are examined in detail. Building on these insights, we propose a systematic solution framework that emphasizes enriching data foundations, advancing LLM capabilities, and ensuring robust simulation design to enhance reliability. Finally, we introduce a structured algorithm that operationalizes the proposed framework, aiming to guide credible and human-aligned LLM-based simulations. To facilitate further research, we provide a curated list of related literature and resources at https://github.com/Persdre/awesome-llm-human-simulation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Talking to an AI Mirror: Designing Self-Clone Chatbots for Enhanced Engagement in Digital Mental Health Support

    cs.HC 2025-09 conditional novelty 6.0 of 10

    Self-clone chatbots that mirror a user's support style showed higher emotional and cognitive engagement than a generic counselor chatbot, but only among the subgroup who found the clone believable.

  2. PUB: An LLM-Enhanced Personality-Driven User Behaviour Simulator for Recommender System Evaluation

    cs.IR 2025-06 conditional novelty 6.0 of 10

    PUB uses LLM-inferred Big Five personality traits to generate synthetic recommender-system interactions that the authors say mimic real Amazon behavior and preserve algorithm performance rankings.

  3. SimuPanel: A Novel Immersive Multi-Agent System to Simulate Interactive Expert Panel Discussion

    cs.HC 2025-06 conditional novelty 5.0 of 10

    A multi-agent LLM system called SimuPanel simulates expert panel discussions with personas grounded in public academic sources, and a small evaluation suggests the full reasoning pipeline produces higher LLM-judged di...

  4. PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading

    cs.CL 2025-06 reject novelty 5.0 of 10

    MAS traders using Reddit sentiment from PulseReddit beat traditional baselines in the reported bull-market backtests, but the gains are small, most runs lose money, and the evaluation has critical flaws.

Pith tools