REVIEW 11 cited by
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
There is an growing interest in using Large Language Models (LLMs) in multi-agent systems to tackle interactive real-world tasks that require effective collaboration and assessing complex situations. Yet, we still have a limited understanding of LLMs' communication and decision-making abilities in multi-agent setups. The fundamental task of negotiation spans many key features of communication, such as cooperation, competition, and manipulation potentials. Thus, we propose using scorable negotiation to evaluate LLMs. We create a testbed of complex multi-agent, multi-issue, and semantically rich negotiation games. To reach an agreement, agents must have strong arithmetic, inference, exploration, and planning capabilities while integrating them in a dynamic and multi-turn setup. We propose multiple metrics to rigorously quantify agents' performance and alignment with the assigned role. We provide procedures to create new games and increase games' difficulty to have an evolving benchmark. Importantly, we evaluate critical safety aspects such as the interaction dynamics between agents influenced by greedy and adversarial players. Our benchmark is highly challenging; GPT-3.5 and small models mostly fail, and GPT-4 and SoTA large models (e.g., Llama-3 70b) still underperform.
Forward citations
Cited by 11 Pith papers
-
Too Human to Model:The Uncanny Valley of LLMs in Social Simulation -- When Generative Language Agents Misalign with Modelling Principles
A position paper contends that LLM agents, despite their human-like talk, are often too rich in detail to serve as scientific models, and proposes conditions where they still excel.
-
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.
-
Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling
LLM-driven agents in a simulated MMO economy reproduce role specialization and price responses to supply and demand, though the price result is partly shaped by what the AI is told.
-
The Power of Stories: Narrative Priming Shapes How LLM Agents Collaborate and Compete
Narrative priming with shared cooperative stories increases LLM-agent contributions in a repeated public goods game, while different or self-interested prompts reduce cooperation.
-
Generative Adversarial Reviews: When LLMs Become the Critic
A new LLM-agent framework, GAR, generates peer reviews from a graph representation of manuscripts and predicts conference acceptance decisions, reportedly matching or exceeding human reviewer performance.
-
PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making
An LLM can generate executable world-model components that, combined with MCTS, outperform LLM-as-policy in GOPS and match it in Taboo, though the evaluation under-supports the partial-observability claim.
-
Finding Common Ground: Using Large Language Models to Detect Agreement in Multi-Agent Decision Conferences
LLM agents can run a simulated decision conference, and a dedicated agreement-detection agent helps the debate cover topics that match a real expert workshop.
-
LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language Models
LLMER uses LLM-generated JSON data instead of code to create interactive XR worlds, cutting token use and task completion time in a small user study.
-
Tackling One Health Risks: How Large Language Models are leveraged for Risk Negotiation and Consensus-building
A new human-in-the-loop workflow combines LLM multi-agent negotiation with the Ehling-Schulz risk negotiation framework and is demonstrated in two role-played One Health scenarios.
-
From Divergence to Consensus: Evaluating the Role of Large Language Models in Facilitating Agreement through Adaptive Strategies
In a 75-session pilot with two Greek students per session, ChatGPT 4.0 produced consensus proposals with higher average cosine similarity to initial participant opinions and required fewer iterations than Mistral Larg...
-
A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios
LLM-based game-playing agents are surveyed across choice-focused and communication-focused games, with a comparative performance table and future directions.
Discussion (0). Continue with ORCID to comment.