Pith. sign in

REVIEW 4 cited by

Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.23426 v1 pith:MVOAYAL4 submitted 2024-10-30 cs.CL

classification cs.CL
keywords llmsreliabilitysimulationsimulationsllm-basedsocialapplicationslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) are increasingly employed for simulations, enabling applications in role-playing agents and Computational Social Science (CSS). However, the reliability of these simulations is under-explored, which raises concerns about the trustworthiness of LLMs in these applications. In this paper, we aim to answer ``How reliable is LLM-based simulation?'' To address this, we introduce TrustSim, an evaluation dataset covering 10 CSS-related topics, to systematically investigate the reliability of the LLM simulation. We conducted experiments on 14 LLMs and found that inconsistencies persist in the LLM-based simulated roles. In addition, the consistency level of LLMs does not strongly correlate with their general performance. To enhance the reliability of LLMs in simulation, we proposed Adaptive Learning Rate Based ORPO (AdaORPO), a reinforcement learning-based algorithm to improve the reliability in simulation across 7 LLMs. Our research provides a foundation for future studies to explore more robust and trustworthy LLM-based simulations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Vision language models prompted with a low vision participant's vision profile and one example response reach only 70% agreement with that participant's held-out image answers.

  2. LLMs are Introvert

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A psychology-inspired prompting method (SIP-CoT with emotion-guided memory) makes LLM agents reproduce human-like attitudes and behaviors more closely in social simulations, but the evaluation lacks error bars, a name...

  3. Simulating Students with Large Language Models: A Review of Architecture, Mechanisms, and Role Modelling in Education with Generative AI

    cs.CY 2025-11 conditional novelty 4.0 of 10

    LLM-based simulated students can produce plausible classroom dialogue and support low-stakes education experiments, but their fidelity is limited and validation standards are lacking.

  4. Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A unified review of AI applications in spectroscopy, organizing forward and inverse tasks across MS, NMR, IR, Raman, and UV-Vis, with a curated resource repository.

Pith tools