Pith. sign in

REVIEW 11 cited by

Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.17234 v2 pith:7K2VUCQT submitted 2023-09-29 cs.CL cs.CYcs.LG

classification cs.CLcs.CYcs.LG
keywords negotiationagentsgamesllmsmodelsmulti-agentbenchmarkcommunication
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

There is an growing interest in using Large Language Models (LLMs) in multi-agent systems to tackle interactive real-world tasks that require effective collaboration and assessing complex situations. Yet, we still have a limited understanding of LLMs' communication and decision-making abilities in multi-agent setups. The fundamental task of negotiation spans many key features of communication, such as cooperation, competition, and manipulation potentials. Thus, we propose using scorable negotiation to evaluate LLMs. We create a testbed of complex multi-agent, multi-issue, and semantically rich negotiation games. To reach an agreement, agents must have strong arithmetic, inference, exploration, and planning capabilities while integrating them in a dynamic and multi-turn setup. We propose multiple metrics to rigorously quantify agents' performance and alignment with the assigned role. We provide procedures to create new games and increase games' difficulty to have an evolving benchmark. Importantly, we evaluate critical safety aspects such as the interaction dynamics between agents influenced by greedy and adversarial players. Our benchmark is highly challenging; GPT-3.5 and small models mostly fail, and GPT-4 and SoTA large models (e.g., Llama-3 70b) still underperform.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Too Human to Model:The Uncanny Valley of LLMs in Social Simulation -- When Generative Language Agents Misalign with Modelling Principles

    cs.CY 2025-07 conditional novelty 7.0 of 10

    A position paper contends that LLM agents, despite their human-like talk, are often too rich in detail to serve as scientific models, and proposes conditions where they still excel.

  2. The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.

  3. Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLM-driven agents in a simulated MMO economy reproduce role specialization and price responses to supply and demand, though the price result is partly shaped by what the AI is told.

  4. The Power of Stories: Narrative Priming Shapes How LLM Agents Collaborate and Compete

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Narrative priming with shared cooperative stories increases LLM-agent contributions in a repeated public goods game, while different or self-interested prompts reduce cooperation.

  5. Generative Adversarial Reviews: When LLMs Become the Critic

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new LLM-agent framework, GAR, generates peer reviews from a graph representation of manuscripts and predicts conference acceptance decisions, reportedly matching or exceeding human reviewer performance.

  6. PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making

    cs.AI 2024-11 reject novelty 6.0 of 10

    An LLM can generate executable world-model components that, combined with MCTS, outperform LLM-as-policy in GOPS and match it in Taboo, though the evaluation under-supports the partial-observability claim.

  7. Finding Common Ground: Using Large Language Models to Detect Agreement in Multi-Agent Decision Conferences

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLM agents can run a simulated decision conference, and a dedicated agreement-detection agent helps the debate cover topics that match a real expert workshop.

  8. LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language Models

    cs.MM 2025-02 conditional novelty 5.0 of 10

    LLMER uses LLM-generated JSON data instead of code to create interactive XR worlds, cutting token use and task completion time in a small user study.

  9. Tackling One Health Risks: How Large Language Models are leveraged for Risk Negotiation and Consensus-building

    cs.MA 2025-09 conditional novelty 4.0 of 10

    A new human-in-the-loop workflow combines LLM multi-agent negotiation with the Ehling-Schulz risk negotiation framework and is demonstrated in two role-played One Health scenarios.

  10. From Divergence to Consensus: Evaluating the Role of Large Language Models in Facilitating Agreement through Adaptive Strategies

    cs.HC 2025-02 conditional novelty 4.0 of 10

    In a 75-session pilot with two Greek students per session, ChatGPT 4.0 produced consensus proposals with higher average cosine similarity to initial participant opinions and required fewer iterations than Mistral Larg...

  11. A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios

    cs.CL 2024-12 conditional novelty 3.0 of 10

    LLM-based game-playing agents are surveyed across choice-focused and communication-focused games, with a comparative performance table and future directions.

Pith tools