{"work":{"id":"47b9c80b-5140-4a3e-9d18-d298c3ffac79","openalex_id":"https://openalex.org/W4387800437","doi":"10.48550/arxiv.2310.11667","arxiv_id":"2310.11667","raw_key":null,"title":"SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents","authors":null,"authors_text":"Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap","year":2023,"venue":"cs.AI","abstract":"Humans are social beings; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and interact under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents.","external_url":"https://arxiv.org/abs/2310.11667","cited_by_count":3,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2310.11667","created_at":"2026-05-10T09:43:47.715950+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Sotopia: Interactive evaluation for social intelligence in language agents","render_title":"Sotopia: Interactive evaluation for social intelligence in language agents"},"hub":{"state":{"work_id":"47b9c80b-5140-4a3e-9d18-d298c3ffac79","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":28,"external_cited_by_count":3,"distinct_field_count":9,"first_pith_cited_at":"2024-10-09T11:01:29+00:00","last_pith_cited_at":"2026-07-07T11:34:10+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T21:29:42.775057+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":1}],"polarity_counts":[{"context_polarity":"background","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}