Pith. sign in

hub

Deep Reinforcement Learning from Self-Play in Imperfect-Information Games

14 Pith papers cite this work, alongside 145 external citations. Polarity classification is still indexing.

14 Pith papers citing it
145 external citations · Pith
abstract

Many real-world applications can be described as large-scale games of imperfect information. To deal with these challenging domains, prior work has focused on computing Nash equilibria in a handcrafted abstraction of the domain. In this paper we introduce the first scalable end-to-end approach to learning approximate Nash equilibria without prior domain knowledge. Our method combines fictitious self-play with deep reinforcement learning. When applied to Leduc poker, Neural Fictitious Self-Play (NFSP) approached a Nash equilibrium, whereas common reinforcement learning methods diverged. In Limit Texas Holdem, a poker game of real-world scale, NFSP learnt a strategy that approached the performance of state-of-the-art, superhuman algorithms based on significant domain expertise.

hub tools

citation-role summary

background 1

citation-polarity summary

roles

background 1

polarities

background 1

representative citing papers

Global Policy-Space Response Oracles for Two-Player Zero-Sum Games

cs.AI · 2026-05-27 · unverdicted · novelty 7.0

Global PSRO minimizes Population Exploitability to expand restricted strategy sets, yielding lower exploitability and Nash approximations with fewer iterations than prior PSRO variants across tested zero-sum games.

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play

cs.AI · 2026-05-16 · unverdicted · novelty 6.0

PopuLoRA shows that co-evolving populations of LoRA adapters through cross-evaluated self-play can outperform compute-matched single-agent baselines on multiple code and math reasoning benchmarks.

A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

cs.LG · 2026-07-07 · accept · novelty 5.0

A controlled study using a fixed Gin Rummy expert as a yardstick isolates which lightweight RL training choices help (TRPO, knock-first reward, curriculum, warm-start, best-checkpoint) and which fail (imitation, dense rewards, embeddings, LLM opponents), finding the performance ceiling is set by the

citing papers explorer

Showing 14 of 14 citing papers.