LLM agents voluntarily adopt secret collusion tools in competitive multi-agent games despite explicit unfairness labels, and only explicit ethical framing reduces adoption rates.
Multi- party chat: Conversational agents in group settings with humans and models.arXiv preprint arXiv:2304.13835
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 5verdicts
UNVERDICTED 5roles
background 1polarities
background 1representative citing papers
SCENE is a new benchmark for testing LLMs on recognizing implicit social norms and adapting to sanctions in multi-party group chats.
CAVI framework uses character-guided token pruning, orthogonal feature modulation, and modality-adaptive role steering to resolve modality-role interference in multimodal RPAs.
M2CL trains per-agent context generators with a self-adaptive mechanism to maintain coherence and reduce output discrepancies in multi-LLM discussions, yielding 20-50% gains on reasoning, embodied, and mobile control tasks.
LLMs beat humans and supervised models at next speaker prediction in meetings using only text, while multimodal LLMs improve on addressee and turn-change tasks but remain below human performance.
citing papers explorer
-
Voluntary Collusion with Secret Tools in Competing LLM Agents
LLM agents voluntarily adopt secret collusion tools in competitive multi-agent games despite explicit unfairness labels, and only explicit ethical framing reduces adoption rates.
-
SCENE: Recognizing Social Norms and Sanctioning in Group Chats
SCENE is a new benchmark for testing LLMs on recognizing implicit social norms and adapting to sanctions in multi-party group chats.
-
Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent
CAVI framework uses character-guided token pruning, orthogonal feature modulation, and modality-adaptive role steering to resolve modality-role interference in multimodal RPAs.
-
Context Learning for Multi-Agent Discussion
M2CL trains per-agent context generators with a self-adaptive mechanism to maintain coherence and reduce output discrepancies in multi-LLM discussions, yielding 20-50% gains on reasoning, embodied, and mobile control tasks.
-
Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings
LLMs beat humans and supervised models at next speaker prediction in meetings using only text, while multimodal LLMs improve on addressee and turn-change tasks but remain below human performance.