REVIEW 30 cited by
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Humans are social beings; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and interact under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents.
Forward citations
Cited by 30 Pith papers
-
COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
COSI-Lab is a weakly scripted conference-workshop dataset with multi-perspective apparent-intent annotations, self-reported goals, and benchmarks for social intention inference and conversation group detection.
-
CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
SalesSim benchmarks MLLMs as retail user simulators, finds gaps in persona adherence and over-persuasion, and introduces UserGRPO RL to raise decision alignment by 13.8%.
-
TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence
A temporal-aware hierarchical reinforcement learning method improves a 7B LLM's social reasoning enough to rival DeepSeek-R1 and OpenAI-O3 on in-domain theory-of-mind benchmarks.
-
CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation
CARD plans and generates Reddit-style credit card discussions, then self-revises them until their lexical, semantic, and structural distributions match real threads more closely than existing baselines.
-
Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling
LLM agents with demographic profiles reproduce CPT-style risk attitudes in route choice and yield fitted parameters (α=0.4, β=0.64, λ=1.43) that predict human data competitively.
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.
-
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
A benchmark for LLM agents in partially observable joint decision-making reveals that deliberation challenges current models but can enable reflection and error correction.
-
Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility
Conditioning LLMs on survey-derived demographic belief profiles improves their ability to predict individuals' misinformation judgments, especially when belief modeling is decoupled from susceptibility prediction.
-
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
A unified evaluation framework for proactive dialogue agents, built with 328 synthetic environments across six domains, shows that thinking modes improve target planning but not dialogue guidance in a 22-model comparison.
-
LLM-Based Social Simulations Require a Boundary
LLM-based social simulations are scientifically useful only within boundaries set by behavioral variance, and current validation practice under-checks variance.
-
Kaleidoscopic Teaming in Multi Agent Simulations
The MASK framework red-teams AI agents in single- and multi-agent simulated societies and reports that multi-agent interactions expose more safety vulnerabilities than single-agent scenarios.
-
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
MotiveBench shows that large language models still lag human consensus on motivational reasoning, with GPT-4o scoring 80.89 percent and chain-of-thought prompting usually reducing accuracy.
-
LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
Language agents' believability and goal achievement decline over multi-episode social interactions, and curated memory summaries only partially close the gap with humans.
-
MAEBE: Multi-Agent Emergent Behavior Framework
Multi-agent LLM ensembles show different and less predictable moral preferences than single models, with convergence driven by peer pressure, according to a new evaluation framework.
-
Aligning VLM Assistants with Personalized Situated Cognition
The authors present PCogAlignBench, an 18k-sample benchmark of visual scenes with role-based users, and PCogAlign, a framework using a cognition-aware reward model to produce responses aligned with each user's roles.
-
ARIA: Training Language Agents with Intention-Driven Reward Aggregation
Clustering language-agent actions into shared intentions and averaging their rewards reduces reward variance and improves policy performance in open-ended dialogue tasks.
-
Empowering LLMs in Task-Oriented Dialogues: A Domain-Independent Multi-Agent Framework and Fine-Tuning Strategy
A three-agent domain-independent framework with distribution-balanced DPO training reaches Combined 106.3 on MultiWOZ 2.2 with Qwen2.5-7B, the best score among the compared baselines.
-
YuLan-OneSim: Towards the Next Generation of Social Simulator with Large Language Models
YuLan-OneSim combines natural-language scenario generation, 50 prebuilt simulation scenarios, feedback-driven agent fine-tuning, 100,000-agent scale, and an automated AI social researcher into one social simulation platform.
-
MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework
A mean-field LLM framework that iterates between summarizing population state and generating individual decisions matches real social-media behavior distributions better than existing LLM simulation baselines.
-
BotSim: LLM-Powered Malicious Social Botnet Simulation
An LLM-powered botnet simulation framework and a Reddit-based dataset show that current social bot detectors degrade sharply against human-like LLM-written bot activity.
-
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning
EFCA combines immediate environment feedback and recent state-history signals into a return reweighting that improves step-level credit assignment for LLM agents in ALFWorld and WebShop.
-
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence
OpenHospital is an interactive physician-patient multi-agent arena that improves clinical metrics via ground-truth reflection and reports cooperative behaviors as evidence of evolving LLM collective intelligence.
-
Evaluating Multi-Turn Bargain Skills in LLM-Based Seller Agent
BargainBench tests LLM seller agents on turn-level buyer intent recognition in synthetic e-commerce bargaining dialogues, where the best models score roughly 55 percent F1.
-
Multi-Actor Generative Artificial Intelligence as a Game Engine
Generative multi-actor AI platforms can be built on the Entity-Component pattern, treating the environment (Game Master) as a composable entity, so that one library serves simulation, storytelling, and evaluation goals.
-
Infected Smallville: How Disease Threat Shapes Sociality in LLM Agents
In three simulation runs, LLM agents primed with a swine flu article reduced social activity compared with controls, but the effect is confounded by prompt instructions and no inferential statistics are reported.
-
BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation
BookWorld builds multi-agent societies from novels and uses them to generate stories that an LLM judge prefers over direct generation and a prior screenwriting agent in most comparisons.
-
AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need
A divide-and-conquer multi-agent framework with task forests and specialized roles improves math and code benchmarks but not commonsense or domain QA, and the adaptive heterogeneous-LLM engine is never tested.
-
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.
-
From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents
A structured survey that categorizes LLM-based social simulation into individual, scenario, and society simulation, with associated methods, benchmarks, and observed trends.
-
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.
Discussion (0). Continue with ORCID to comment.