Pith. sign in

REVIEW 5 cited by

Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.17234 v2 pith:7K2VUCQT submitted 2023-09-29 cs.CL cs.CYcs.LG

classification cs.CLcs.CYcs.LG
keywords negotiationagentsgamesllmsmodelsmulti-agentbenchmarkcommunication
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There is an growing interest in using Large Language Models (LLMs) in multi-agent systems to tackle interactive real-world tasks that require effective collaboration and assessing complex situations. Yet, we still have a limited understanding of LLMs' communication and decision-making abilities in multi-agent setups. The fundamental task of negotiation spans many key features of communication, such as cooperation, competition, and manipulation potentials. Thus, we propose using scorable negotiation to evaluate LLMs. We create a testbed of complex multi-agent, multi-issue, and semantically rich negotiation games. To reach an agreement, agents must have strong arithmetic, inference, exploration, and planning capabilities while integrating them in a dynamic and multi-turn setup. We propose multiple metrics to rigorously quantify agents' performance and alignment with the assigned role. We provide procedures to create new games and increase games' difficulty to have an evolving benchmark. Importantly, we evaluate critical safety aspects such as the interaction dynamics between agents influenced by greedy and adversarial players. Our benchmark is highly challenging; GPT-3.5 and small models mostly fail, and GPT-4 and SoTA large models (e.g., Llama-3 70b) still underperform.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Too Human to Model:The Uncanny Valley of LLMs in Social Simulation -- When Generative Language Agents Misalign with Modelling Principles

    cs.CY 2025-07 conditional novelty 7.0 of 10

    A position paper contends that LLM agents, despite their human-like talk, are often too rich in detail to serve as scientific models, and proposes conditions where they still excel.

  2. The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.

  3. Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLM-driven agents in a simulated MMO economy reproduce role specialization and price responses to supply and demand, though the price result is partly shaped by what the AI is told.

  4. Finding Common Ground: Using Large Language Models to Detect Agreement in Multi-Agent Decision Conferences

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLM agents can run a simulated decision conference, and a dedicated agreement-detection agent helps the debate cover topics that match a real expert workshop.

  5. Tackling One Health Risks: How Large Language Models are leveraged for Risk Negotiation and Consensus-building

    cs.MA 2025-09 conditional novelty 4.0 of 10

    A new human-in-the-loop workflow combines LLM multi-agent negotiation with the Ehling-Schulz risk negotiation framework and is demonstrated in two role-played One Health scenarios.

Pith tools