Pith. sign in

REVIEW 6 cited by

DREAM: Deep Regret minimization with Advantage baselines and Model-free learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.10410 v2 pith:HLQLOD42 submitted 2020-06-18 cs.LG cs.GTstat.ML

classification cs.LGcs.GTstat.ML
keywords dreamgamesalgorithmsdeeplearningalgorithmequilibriummodel-free
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce DREAM, a deep reinforcement learning algorithm that finds optimal strategies in imperfect-information games with multiple agents. Formally, DREAM converges to a Nash Equilibrium in two-player zero-sum games and to an extensive-form coarse correlated equilibrium in all other games. Our primary innovation is an effective algorithm that, in contrast to other regret-based deep learning algorithms, does not require access to a perfect simulator of the game to achieve good performance. We show that DREAM empirically achieves state-of-the-art performance among model-free algorithms in popular benchmark games, and is even competitive with algorithms that do use a perfect simulator.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Correlated Chance Sampling for Monte Carlo Counterfactual Regret Minimization

    cs.GT 2026-07 conditional novelty 6.0 of 10

    Persistent randomized Weyl streams at each chance node cut MCCFR exploitability 19–34% on Kuhn and Leduc poker with local O(log N/N) frequency guarantees and no new hyperparameters.

  2. NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria

    cs.LG 2025-10 unverdicted novelty 6.0 of 10

    NashPG is a policy-gradient method with iteratively refined regularization that guarantees monotonic convergence to Nash equilibria in two-player zero-sum extensive-form games and scales to large benchmarks.

  3. Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria

    cs.MA 2026-06 unverdicted novelty 5.0 of 10

    Phi-Actor-Critic is a new method that steers multi-agent reinforcement learning toward Pareto-efficient correlated equilibria using regret minimization and Lagrangian selection.

  4. Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    DAGS initializes policy-gradient self-play from human-derived intermediate states to reduce exploitability in challenging imperfect-information games, with a multi-task flag fix for resulting bias and new benchmark en...

  5. AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    AlphaExploitem adds a hierarchical transformer encoder and a diverse pool of exploitable opponents to AlphaHoldem, enabling exploitation of suboptimal poker play while preserving performance against Nash-equilibrium o...

  6. NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria

    cs.LG 2025-10 conditional novelty 4.0 of 10

    Periodically re-centering the KL-regularizer on the current policy in self-play yields a policy-gradient algorithm that, in its exact form, provably converges to a Nash equilibrium.

Pith tools