SocialGrid benchmark shows even top LLMs achieve below 60% in embodied planning and task completion, with deception detection near random chance regardless of model scale.
In: Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
CooperScene provides 59K synchronized frames with 344K 3D annotations from multi-modal sensors on 3 CAVs and 1 RSU plus real C-V2X communication traces for cooperative autonomy benchmarking.
citing papers explorer
-
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems
SocialGrid benchmark shows even top LLMs achieve below 60% in embodied planning and task completion, with deception detection near random chance regardless of model scale.
-
CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization
CooperScene provides 59K synchronized frames with 344K 3D annotations from multi-modal sensors on 3 CAVs and 1 RSU plus real C-V2X communication traces for cooperative autonomy benchmarking.