Global PSRO minimizes Population Exploitability to expand restricted strategy sets, yielding lower exploitability and Nash approximations with fewer iterations than prior PSRO variants across tested zero-sum games.
hub
Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
14 Pith papers cite this work, alongside 145 external citations. Polarity classification is still indexing.
abstract
Many real-world applications can be described as large-scale games of imperfect information. To deal with these challenging domains, prior work has focused on computing Nash equilibria in a handcrafted abstraction of the domain. In this paper we introduce the first scalable end-to-end approach to learning approximate Nash equilibria without prior domain knowledge. Our method combines fictitious self-play with deep reinforcement learning. When applied to Leduc poker, Neural Fictitious Self-Play (NFSP) approached a Nash equilibrium, whereas common reinforcement learning methods diverged. In Limit Texas Holdem, a poker game of real-world scale, NFSP learnt a strategy that approached the performance of state-of-the-art, superhuman algorithms based on significant domain expertise.
hub tools
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
PPO in a new competitive game fails due to five implementation bugs and then competitive overfitting where self-play stays near 50% but generalization drops to 21.6%; mixing 20% random opponents restores generalization to 77.1%.
OpenAI Five achieved superhuman performance in Dota 2 by defeating the world champions using scaled self-play reinforcement learning.
Self-play RL in a takeover auction model shows optimal due diligence is modest and finite, decreasing with cost and competition, while simple general methods outperform specialized ones in large intractable games.
Self-play RL with a vision transformer policy, powered by a 10,000x faster JAX simulator, produces an agent that ranks #1 on the Generals.io leaderboard and wins 199-70 against top humans.
GARIP uses a running-average reference to minimize peak lag among causal convex averages, proves local last-iterate convergence at constant anchor strength, and matches R-NaD robustness on matrix, Coin, Connect Four, and Othello games while being easier to tune.
Multi-agent RL with self-play trains quadrotors that beat a human champion at 22 m/s races while halving collisions versus single-agent methods and generalizing zero-shot to human opponents.
PopuLoRA shows that co-evolving populations of LoRA adapters through cross-evaluated self-play can outperform compute-matched single-agent baselines on multiple code and math reasoning benchmarks.
A controlled study using a fixed Gin Rummy expert as a yardstick isolates which lightweight RL training choices help (TRPO, knock-first reward, curriculum, warm-start, best-checkpoint) and which fail (imitation, dense rewards, embeddings, LLM opponents), finding the performance ceiling is set by the
FootsiesGym is an open-source, vectorized fighting-game benchmark for two-player zero-sum imperfect-information RL that isolates non-transitive neutral-game dynamics while remaining tractable on standard hardware.
StratFormer uses a two-phase curriculum with dual-turn tokens and bucket-rate features to model and exploit opponents in Leduc Hold'em, gaining +0.106 BB/hand on average over GTO while keeping near-equilibrium safety.
Randomly masking a proposer's output vocabulary during training and generation sustains curriculum diversity and improves solver accuracy by +4.4 points at 8B in LLM co-evolution.
Periodically re-centering the KL-regularizer on the current policy in self-play yields a policy-gradient algorithm that, in its exact form, provably converges to a Nash equilibrium.
PPO with moderate entropy regularization and current-policy self-play outperforms Monte Carlo Q, SARSA, and Q-learning in a controlled self-play framework for the imperfect-information game Big 2.
citing papers explorer
-
Global Policy-Space Response Oracles for Two-Player Zero-Sum Games
Global PSRO minimizes Population Exploitability to expand restricted strategy sets, yielding lower exploitability and Nash approximations with fewer iterations than prior PSRO variants across tested zero-sum games.
-
Territory Paint Wars: Diagnosing and Mitigating Failure Modes in Competitive Multi-Agent PPO
PPO in a new competitive game fails due to five implementation bugs and then competitive overfitting where self-play stays near 50% but generalization drops to 21.6%; mixing 20% random opponents restores generalization to 77.1%.
-
Dota 2 with Large Scale Deep Reinforcement Learning
OpenAI Five achieved superhuman performance in Dota 2 by defeating the world champions using scaled self-play reinforcement learning.
-
How Much Due Diligence Before You Bid? Learning in Intractable Takeover Auctions
Self-play RL in a takeover auction model shows optimal due diligence is modest and finite, decreasing with cost and competition, while simple general methods outperform specialized ones in large intractable games.
-
Superhuman AI for Generals.io Using Self-Play Reinforcement Learning
Self-play RL with a vision transformer policy, powered by a 10,000x faster JAX simulator, produces an agent that ranks #1 on the Generals.io leaderboard and wins 199-70 against top humans.
-
GARIP: A Running-Average Moving Reference for Last-Iterate Self-Play in Two-Player Zero-Sum Games
GARIP uses a running-average reference to minimize peak lag among causal convex averages, proves local last-iterate convergence at constant anchor strength, and matches R-NaD robustness on matrix, Coin, Connect Four, and Othello games while being easier to tune.
-
Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
Multi-agent RL with self-play trains quadrotors that beat a human champion at 22 m/s races while halving collisions versus single-agent methods and generalizing zero-shot to human opponents.
-
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play
PopuLoRA shows that co-evolving populations of LoRA adapters through cross-evaluated self-play can outperform compute-matched single-agent baselines on multiple code and math reasoning benchmarks.
-
A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong
A controlled study using a fixed Gin Rummy expert as a yardstick isolates which lightweight RL training choices help (TRPO, knock-first reward, curriculum, warm-start, best-checkpoint) and which fail (imitation, dense rewards, embeddings, LLM opponents), finding the performance ceiling is set by the
-
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games
FootsiesGym is an open-source, vectorized fighting-game benchmark for two-player zero-sum imperfect-information RL that isolates non-transitive neutral-game dynamics while remaining tractable on standard hardware.
-
StratFormer: Adaptive Opponent Modeling and Exploitation in Imperfect-Information Games
StratFormer uses a two-phase curriculum with dual-turn tokens and bucket-rate features to model and exploit opponents in Leduc Hold'em, gaining +0.106 BB/hand on average over GTO while keeping near-equilibrium safety.
-
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
Randomly masking a proposer's output vocabulary during training and generation sustains curriculum diversity and improves solver accuracy by +4.4 points at 8B in LLM co-evolution.
-
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria
Periodically re-centering the KL-regularizer on the current policy in self-play yields a policy-gradient algorithm that, in its exact form, provably converges to a Nash equilibrium.
-
Self-Play Reinforcement Learning under Imperfect Information in Big 2
PPO with moderate entropy regularization and current-policy self-play outperforms Monte Carlo Q, SARSA, and Q-learning in a controlled self-play framework for the imperfect-information game Big 2.