REVIEW 4 cited by
LLM-based Human Simulations Have Not Yet Been Reliable
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant discrepancies between their outcomes and authentic human actions. Our investigation begins with a systematic review of LLM-based human simulations in social, economic, policy, and psychological contexts, identifying their common frameworks, recent advances, and persistent limitations. This review reveals that such discrepancies primarily stem from inherent limitations of LLMs and flaws in simulation design, both of which are examined in detail. Building on these insights, we propose a systematic solution framework that emphasizes enriching data foundations, advancing LLM capabilities, and ensuring robust simulation design to enhance reliability. Finally, we introduce a structured algorithm that operationalizes the proposed framework, aiming to guide credible and human-aligned LLM-based simulations. To facilitate further research, we provide a curated list of related literature and resources at https://github.com/Persdre/awesome-llm-human-simulation.
Forward citations
Cited by 4 Pith papers
-
Talking to an AI Mirror: Designing Self-Clone Chatbots for Enhanced Engagement in Digital Mental Health Support
Self-clone chatbots that mirror a user's support style showed higher emotional and cognitive engagement than a generic counselor chatbot, but only among the subgroup who found the clone believable.
-
PUB: An LLM-Enhanced Personality-Driven User Behaviour Simulator for Recommender System Evaluation
PUB uses LLM-inferred Big Five personality traits to generate synthetic recommender-system interactions that the authors say mimic real Amazon behavior and preserve algorithm performance rankings.
-
SimuPanel: A Novel Immersive Multi-Agent System to Simulate Interactive Expert Panel Discussion
A multi-agent LLM system called SimuPanel simulates expert panel discussions with personas grounded in public academic sources, and a small evaluation suggests the full reasoning pipeline produces higher LLM-judged di...
-
PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading
MAS traders using Reddit sentiment from PulseReddit beat traditional baselines in the reported bull-market backtests, but the gains are small, most runs lose money, and the evaluation has critical flaws.
Discussion (0). Sign in to comment.