Pith. sign in

REVIEW 2 cited by

DREAM: Deep Regret minimization with Advantage baselines and Model-free learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.10410 v2 pith:HLQLOD42 submitted 2020-06-18 cs.LG cs.GTstat.ML

classification cs.LGcs.GTstat.ML
keywords dreamgamesalgorithmsdeeplearningalgorithmequilibriummodel-free
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce DREAM, a deep reinforcement learning algorithm that finds optimal strategies in imperfect-information games with multiple agents. Formally, DREAM converges to a Nash Equilibrium in two-player zero-sum games and to an extensive-form coarse correlated equilibrium in all other games. Our primary innovation is an effective algorithm that, in contrast to other regret-based deep learning algorithms, does not require access to a perfect simulator of the game to achieve good performance. We show that DREAM empirically achieves state-of-the-art performance among model-free algorithms in popular benchmark games, and is even competitive with algorithms that do use a perfect simulator.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Correlated Chance Sampling for Monte Carlo Counterfactual Regret Minimization

    cs.GT 2026-07 conditional novelty 6.0 of 10

    Persistent randomized Weyl streams at each chance node cut MCCFR exploitability 19–34% on Kuhn and Leduc poker with local O(log N/N) frequency guarantees and no new hyperparameters.

  2. NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria

    cs.LG 2025-10 unverdicted novelty 6.0 of 10

    Periodically re-centering the KL-regularizer on the current policy in self-play yields a policy-gradient algorithm that, in its exact form, provably converges to a Nash equilibrium.

Pith tools