A new 21-game benchmark scores LLM social reasoning with rule-decided outcomes, and SPaRTan, a self-reflection loop, transfers playbook lessons across games.
Acta Universitatis Sapientiae, Informatica , volume=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
A new 21-game benchmark scores LLM social reasoning with rule-decided outcomes, and SPaRTan, a self-reflection loop, transfers playbook lessons across games.