REVIEW 5 cited by
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) are increasingly leveraged to empower autonomous agents to simulate human beings in various fields of behavioral research. However, evaluating their capacity to navigate complex social interactions remains a challenge. Previous studies face limitations due to insufficient scenario diversity, complexity, and a single-perspective focus. To this end, we introduce AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios. Drawing on Dramaturgical Theory, AgentSense employs a bottom-up approach to create 1,225 diverse social scenarios constructed from extensive scripts. We evaluate LLM-driven agents through multi-turn interactions, emphasizing both goal completion and implicit reasoning. We analyze goals using ERG theory and conduct comprehensive experiments. Our findings highlight that LLMs struggle with goals in complex social scenarios, especially high-level growth needs, and even GPT-4o requires improvement in private information reasoning. Code and data are available at \url{https://github.com/ljcleo/agent_sense}.
Forward citations
Cited by 5 Pith papers
-
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
SAGE evaluates LLMs' higher-order social cognition by having a simulated 'sentient' user update its emotion score during multi-turn supportive dialogues, and reports that this score correlates with psychology-informed...
-
Multi-Agent Simulator Drives Language Models for Legal Intensive Interaction
MASER generates synthetic interactive legal dialogues from real judgments and uses them to fine-tune LLMs that outperform GPT-4o on a new interactive complaint-drafting benchmark.
-
PIORS: Personalized Intelligent Outpatient Reception based on Large Language Model with Multi-Agents Medical Scenario Simulation
A fine-tuned LLM receptionist trained on simulated patient conversations outperformed GPT-4o and other baselines in virtual outpatient triage tests.
-
Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant
Auto-SLURP is a new end-to-end benchmark for LLM multi-agent personal assistants, and the best current framework still fails a majority of user requests.
-
From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents
A structured survey that categorizes LLM-based social simulation into individual, scenario, and society simulation, with associated methods, benchmarks, and observed trends.
Discussion (0). Continue with ORCID to comment.