Pith. sign in

REVIEW 3 cited by

Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.05492 v1 pith:GU2NAXJR submitted 2022-10-11 cs.GT cs.AIcs.LGcs.MA

classification cs.GTcs.AIcs.LGcs.MA
keywords humanlearningalgorithmdiplomacygameinvolvingmodelno-press
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

No-press Diplomacy is a complex strategy game involving both cooperation and competition that has served as a benchmark for multi-agent AI research. While self-play reinforcement learning has resulted in numerous successes in purely adversarial games like chess, Go, and poker, self-play alone is insufficient for achieving optimal performance in domains involving cooperation with humans. We address this shortcoming by first introducing a planning algorithm we call DiL-piKL that regularizes a reward-maximizing policy toward a human imitation-learned policy. We prove that this is a no-regret learning algorithm under a modified utility function. We then show that DiL-piKL can be extended into a self-play reinforcement learning algorithm we call RL-DiL-piKL that provides a model of human play while simultaneously training an agent that responds well to this human model. We used RL-DiL-piKL to train an agent we name Diplodocus. In a 200-game no-press Diplomacy tournament involving 62 human participants spanning skill levels from beginner to expert, two Diplodocus agents both achieved a higher average score than all other participants who played more than two games, and ranked first and third according to an Elo ratings model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Solving Zero-Sum Convex Markov Games

    cs.GT 2025-06 conditional novelty 7.0 of 10

    Independent policy-gradient algorithms provably compute approximate Nash equilibria in two-player zero-sum convex Markov games.

  2. SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.

  3. Online Competitive Information Gathering for Partially Observable Trajectory Games

    cs.GT 2025-06 reject novelty 5.0 of 10

    A particle-based stochastic-gradient planner for finite-history partially observable trajectory games is shown to produce active information-gathering behavior in continuous pursuit-evasion and warehouse-pickup simula...

Pith tools