Pith. sign in

REVIEW 5 cited by

Welfare Diplomacy: Benchmarking Language Model Cooperation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08901 v1 pith:XY7BIFJG submitted 2023-10-13 cs.MA cs.AIcs.CL

classification cs.MAcs.AIcs.CL
keywords diplomacywelfarecapabilitiescooperativebenchmarksexperimentslanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The growing capabilities and increasingly widespread deployment of AI systems necessitate robust benchmarks for measuring their cooperative capabilities. Unfortunately, most multi-agent benchmarks are either zero-sum or purely cooperative, providing limited opportunities for such measurements. We introduce a general-sum variant of the zero-sum board game Diplomacy -- called Welfare Diplomacy -- in which players must balance investing in military conquest and domestic welfare. We argue that Welfare Diplomacy facilitates both a clearer assessment of and stronger training incentives for cooperative capabilities. Our contributions are: (1) proposing the Welfare Diplomacy rules and implementing them via an open-source Diplomacy engine; (2) constructing baseline agents using zero-shot prompted language models; and (3) conducting experiments where we find that baselines using state-of-the-art models attain high social welfare but are exploitable. Our work aims to promote societal safety by aiding researchers in developing and assessing multi-agent AI systems. Code to evaluate Welfare Diplomacy and reproduce our experiments is available at https://github.com/mukobi/welfare-diplomacy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A survey proposing adaptability as a three-part taxonomy (learning, policy, scenario-driven) for organizing and evaluating MARL under changing conditions.

  3. Preventing Rogue Agents Improves Multi-Agent Collaboration

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A monitor trained on agent uncertainty statistics triggers rollbacks that improve multi-agent LLM collaboration across three test environments.

  4. Information Bargaining: Bilateral Commitment in Bayesian Persuasion

    cs.GT 2025-06 reject novelty 4.0 of 10

    Bayesian persuasion is restated as a two-sided bargaining game, but the proof reduces to a relabeling and the empirical validation is circular.

  5. A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A narrative review of behavioral and representational Theory of Mind in LLMs, with a taxonomy of safety risks and mitigation directions.

Pith tools