Pith. sign in

REVIEW 4 major objections 6 minor 65 references

Experience-replay Innovative Dynamics

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read ERID is a stateless multi-agent learning algorithm whose experience-replay updates, with a tunable revision protocol, prove convergence to the trajectories of BNN, Smith, and Smith-replicator pairwise dynamics in the limit of vanishing…

desk verdict Novel and readable, but Theorem 1 is false as stated—the boundary counterexample holds—and the proof's key expectation step is invalid. read the letter →

arxiv 2501.12199 v2 pith:EELWG5VN submitted 2025-01-21 cs.LG cs.GTcs.MA

classification cs.LGcs.GTcs.MA
keywords experiencereplaymulti-agentreinforcementlearningevolutionarygametheoryinnovativedynamicsBNNSmithrevisionprotocolsNashequilibrium
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces ERID, a stateless multi-agent reinforcement learning algorithm that uses an experience-replay buffer to average rewards and a tunable revision protocol to update policies. Its central claim is that when the buffer size grows and the learning rate shrinks together, the policy trajectories of ERID converge to the trajectories of BNN, Smith, and Smith-replicator-based pairwise dynamics, three 'innovative' evolutionary dynamics. This matters because replicator-dynamics-based MARL is unstable or slow in null-stable games like zero-sum games, whereas these innovative dynamics converge there; ERID inherits those guarantees. The paper verifies the match in matching pennies, biased rock-paper-scissors, and a nonstationary rock-paper-scissors setting.

What carries the argument

The load-bearing mechanism is experience replay interpreted as a reward-smoothing average. At each step the algorithm keeps a buffer of the last $K$ action–reward pairs, computes per-action averages $\bar{r}_i$ and the global average $\bar{r}$, then updates the policy with a discrete revision-protocol equation, $\pi_i(t+1) \leftarrow \pi_i(t) + \alpha(\sum_j \pi_j(t)\eta_{ji} - \pi_i(t)\sum_j \eta_{ij})$. Each target dynamics is selected by a protocol factor $\eta_{ij}$ built from the reward averages: $[\bar{r}_j - \bar{r}]_+$ for BNN, $[\bar{r}_j - \bar{r}_i]_+$ for Smith, and the constrained version (12) for Smith-replicator pairwise dynamics. The proof works by showing that as $\alpha K\to 0$ and $K\to\infty$, the buffer averages converge to current expected payoffs, so the discrete update becomes the corresponding ODE.

What would settle it

Run ERID with the BNN protocol on a zero-sum game over a grid of $K$ and $\alpha$ values with $\alpha K$ small and compare the empirical policy trajectory to the BNN ODE trajectory; if the distance does not shrink to zero, or if a direct calculation exhibits a case with random $|I_i|$ where $E(\sum_{j\in I_i} b_j / |I_i|) \neq E(\sum_{j\in I_i} b_j)/E(|I_i|)$, the central convergence claim is refuted.

Watch

Extended reading notes

Core claim

The central claim is that a single experience-replay learning rule can reproduce the trajectories of three 'innovative' evolutionary dynamics. For the BNN protocol factor $\eta_{ij} = [\bar{r}_j - \bar{r}]_+$, the update in equation (6) becomes, in the limit $\alpha K \to 0$ with $K \to \infty$, the BNN differential equation (4)–(5); Theorems 2 and 3 state the same trajectory convergence for Smith dynamics and for Smith-replicator-based pairwise dynamics. The result is a bridge: MARL algorithms built on ERID inherit the convergence properties of these dynamics in stable and null-stable games, a guarantee previously available mainly for replicator dynamics. The empirical section shows the stochastic ERID trajectories tracking the deterministic dynamics in matching pennies and biased RPS, and tracking the shifting Nash equilibrium in a nonstationary RPS game.

Load-bearing premise

The proof assumes that the average reward observed for an action in the replay buffer equals the ratio of expected sums and that policies stay essentially constant over the whole buffer window, so buffer averages behave like current expected payoffs; if those assumptions fail, the claimed trajectory match is not established.

Editorial extensions

If this is right

  • In zero-sum games, ERID with the BNN or Smith protocol factor can converge to the Nash equilibrium, because the underlying innovative dynamics do; this is exactly where replicator-dynamics learners only orbit or need time averaging.
  • In nonstationary environments, ERID's NashConv keeps dropping after payoff changes, while time-averaged replicator-based learners lag because their accumulated averages retain old equilibrium bias.
  • Swapping the revision protocol in the update rule selects a different evolutionary dynamics, so the same algorithm covers BNN, Smith, and constrained pairwise dynamics without changing the replay mechanism.
  • The matching-pennies and biased-RPS experiments show the stochastic ERID trajectories track the deterministic dynamics trajectories closely enough to inherit their qualitative convergence behavior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same replay-averaging construction could be applied to other revision protocols, such as projection or logit dynamics, yielding a whole family of MARL algorithms indexed by protocol factor; the paper does not explore these.
  • A finite-sample analysis of the buffer averages would reveal how $K$ and $\alpha$ must be coupled in practice; the theorem only states the joint limit $\alpha K\to 0$, leaving the rate of convergence unspecified.
  • Because ERID is stateless and uses only aggregate rewards, it could be combined with function approximation by applying the update to a parameterized policy in expectation, though the paper does not address deep RL.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces Experience-replay Innovative Dynamics (ERID), a stateless multi-agent reinforcement learning algorithm that uses a replay buffer and a tunable revision-protocol factor to approximate three evolutionary dynamics: Brown-von Neumann-Nash (BNN), Smith, and Smith-replicator-based pairwise dynamics. The central theoretical claim is that, under the joint limit αK→0 and K→∞, the policy trajectories of ERID converge to the trajectories of these dynamics (Theorems 1–3). The paper also reports experiments in matching pennies, biased Rock-Paper-Scissors, and a nonstationary Rock-Paper-Scissors game, comparing ERID with cross learning and claiming that ERID adapts better to environmental changes.

Significance. If the theoretical claims were established, ERID would provide a general mechanism for transferring the convergence guarantees of innovative dynamics to MARL, going beyond the replicator dynamics that dominate the current EGT-MARL literature. The algorithmic idea of using experience replay to implement non-linear revision protocols is original, and the empirical comparisons, especially the nonstationary RPS experiment, are suggestive and align with the qualitative story. The paper also makes a useful conceptual point about time-averaged replicator dynamics being slow to adapt. However, the central theorems are not proven, and Theorem 1 is false as stated; these issues currently prevent the paper from delivering on its main promise.

major comments (4)
  1. [Theorem 1, Section 3.3] Theorem 1 is false as stated because it does not restrict initial policies to the interior of the simplex. If an action i has probability zero at some time, then by Eq. (2) the set I_i is empty and r̄_i is set to 0; since action i is never selected, I_i remains empty forever and r̄_i stays 0. In the two-player normal-form game with payoff matrix [[2,2],[0,0]] for both players and initial policy (0,1), the BNN dynamics (4) has a positive derivative for the first action at the boundary and converges to (1,0), whereas ERID update (6) keeps the first action at probability 0 for all time because all positive-part terms are zero. The proof explicitly assumes the initial probability of action i is positive immediately before Eq. (7), so the proof does not cover the statement as written.
  2. [Equation (7), Section 3.3] The first equality in Eq. (7) is mathematically invalid. The expected average reward of action i is the expectation of the ratio of the sum of rewards in I_i to the size of I_i, and this is not equal to the ratio of the expectation of the sum to the expectation of the size, because |I_i| is a random variable depending on the sequence of sampled actions. The proof therefore does not justify the subsequent replacement of buffer averages by current expected payoffs, even when all probabilities are positive. A correct argument would need a proper stochastic-approximation or two-time-scale analysis, which is not supplied.
  3. [Limit αK→0 and K→∞, Section 3.3] The proof asserts that under αK→0 and K→∞ the policies remain effectively constant over the buffer window and then invokes the law of large numbers, but this stochastic-averaging step is not made rigorous. The bound on the policy difference over j steps only controls the drift for a fixed j; summing up to K requires a uniform argument with condition αK→0, and the interaction between the two limits, K growing for averaging and αK shrinking for policy constancy, is never formalized. In addition, the passage from the discrete update to the ODE (9) via the parameter θ is only asserted; no error bounds or compactness argument is given.
  4. [Theorems 2 and 3, Sections 3.4 and 3.5] Theorems 2 and 3 are stated without proofs, with the remark that they are similar to Theorem 1. Since the proof of Theorem 1 is invalid, these results are unsupported. Moreover, the boundary failure identified in Theorem 1 applies equally to the Smith update (11) and the Smith-replicator pairwise update (13), because r̄_i is zero for any never-sampled action, preventing the policy from leaving the boundary. The manuscript needs either complete proofs or a clear restriction to interior initial conditions and non-degenerate exploration for all three theorems.
minor comments (6)
  1. [Section 3 heading] The heading contains a typo: 'main contibution' should be 'main contribution'.
  2. [Algorithm 1] Algorithm 1 lists the input learning rate as theta, while the text and update rules consistently use alpha; please unify the notation.
  3. [Section 1.1 and reference [19]] The text refers to 'Hennis et al.', but the cited work is by Hennes et al.; please correct the name.
  4. [Section 3.5, reference [9]] The Smith-replicator-based pairwise dynamics is attributed to reference [9], which is a paper on urban drainage systems; this reference does not appear to define the dynamics in Eq. (12). Please cite the correct source for these dynamics.
  5. [Equation (7) notation] The superscript notation in Eq. (7) is confusing: r̄ with superscript [1] is used for the player-1 empirical average, but the expectation E is not formally defined with respect to the buffer randomness conditional on the current policy; please clarify the probability space and conditioning.
  6. [Figures 2 and 3] The figures compare simulated dynamics and ERID trajectories only visually; adding a quantitative distance metric, such as average Euclidean distance over time, would make the claimed agreement more precise.

Circularity Check

0 steps flagged · score 2.0 of 10

ERID's BNN/Smith convergence is largely designed into the protocol factors, but the stochastic-averaging step is independent content; no fitted-parameter or self-citation circularity.

full rationale

The central convergence theorems are not circular in the fitted-prediction sense. Equation (3) is a generic discrete revision-protocol update; the BNN, Smith, and Smith-replicator variants are obtained by substituting protocol factors eta_ij = [rbar_j - rbar]+, eta_ij = [rbar_j - rbar_i]+, and the constrained form in Eq. (12). Thus the continuous-time limit in Eq. (9) is, by design, the BNN ODE in Eq. (4): the algorithm is intentionally a discretized version of the target dynamics. That design choice is the paper's stated contribution ('by appropriately adjusting the revision protocols, the behavior of our algorithm mirrors the trajectories'), and it is not a hidden fitted parameter. The independent mathematical content lies in Eq. (7), which attempts to replace buffer averages by current expected payoffs under K->infinity and alpha K -> 0, and in the discrete-to-continuous limit. That content is not derived from, and does not assume, the theorem conclusion; it is a stochastic-approximation argument. Its flaws -- the E(sum b_j)/E(|I_i|) interchange, the boundary case rbar_i = 0 for unvisited actions, and the omitted proofs of Theorems 2 and 3 (Sections 3.4 and 3.5) -- are correctness and completeness concerns, not circularity. There are no load-bearing self-citations: references to the authors' earlier work are background or motivational. No fitted parameter is later reported as a prediction. Accordingly, no circular step rises to the enumerated patterns; the only mild issue is that the convergence is largely designed into the update rule, which suppresses the circularity score only slightly above zero.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The central claim rests on an ideal limiting regime and two unproved interchange steps; no fitted parameters beyond the tuning of α and K are used, and those are not fitted to data.

free parameters (2)
  • Learning rate α = 1e-5 in experiments; in theory α→0
    The convergence theorem requires αK→0, and experiments use α=1e-5. The choice affects the speed and accuracy of the limiting dynamics.
  • Buffer size K = 1000 in experiments; in theory K→∞
    The convergence theorem requires K→∞; experiments use K=1000. Finite buffers introduce bias relative to the ideal dynamics.
assumptions (5)
  • domain assumption The reward function R is bounded and lies in [R_min, R_max] for finite normal-form games.
    Used in the proof of Theorem 1 to bound δ; holds for finite payoff matrices.
  • ad hoc to paper E(r̄_i) = E(Σ b_j)/E(|I_i|), a ratio of expectations, holds for the replay buffer average.
    Equation (7) uses this equality, but it is not generally true for random denominators; no justification is given.
  • ad hoc to paper Under αK→0 and K→∞, policies remain effectively constant over the buffer window, so buffer averages equal current expected payoffs.
    This is the key idealization that makes the algorithm reduce to the intended ODE; it is not validated for finite K.
  • domain assumption The policy update rule (3) keeps policies on the probability simplex.
    The algorithm assumes π remains a valid mixed strategy; no proof of nonnegativity and normalization invariance is provided.
  • standard math The discrete update with step α becomes the continuous-time ODE as α→0.
    Standard stochastic approximation argument, as in Börgers and Sarin [11].

how reviews work

0 comments
Cite this review

Pith. "Pith review of Experience-replay Innovative Dynamics." pith.science (2026). https://pith.science/paper/EELWG5VN

@misc{pith2026250112199,
  author       = {Pith},
  title        = {Pith review of: Experience-replay Innovative Dynamics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EELWG5VN}},
  note         = {Machine review of arXiv:2501.12199}
}
read the original abstract

Despite its groundbreaking success, multi-agent reinforcement learning (MARL) still suffers from instability and nonstationarity. Replicator dynamics, the most well-known model from evolutionary game theory (EGT), provide a theoretical framework for the convergence of the trajectories to Nash equilibria and, as a result, have been used to ensure formal guarantees for MARL algorithms in stable game settings. However, they exhibit the opposite behavior in other settings, which poses the problem of finding alternatives to ensure convergence. In contrast, innovative dynamics, such as the Brown-von Neumann-Nash (BNN) or Smith, result in periodic trajectories with the potential to approximate Nash equilibria. Yet, no MARL algorithms based on these dynamics have been proposed. In response to this challenge, we develop a novel experience replay-based MARL algorithm that incorporates revision protocols as tunable hyperparameters. We demonstrate, by appropriately adjusting the revision protocols, that the behavior of our algorithm mirrors the trajectories resulting from these dynamics. Importantly, our contribution provides a framework capable of extending the theoretical guarantees of MARL algorithms beyond replicator dynamics. Finally, we corroborate our theoretical findings with empirical results.

Figures

Figures reproduced from arXiv: 2501.12199 by the authors.

Figure 1
Figure 1. Policy NashConv of BNN, Replicator and Hedge algorithms in nonstationary RPS, with the game phases every 3000 iterations separated by vertical red lines. basic rules of the modified RPS game are identical to the traditional RPS, where the winner of each round receives a payoff of +1, while the loser incurs a payoff of -1. If both players choose the same action, the outcome is a tie and, therefore, both players recei… view at source ↗
Figure 2
Figure 2. Innovative dynamics (left) vs policy trajectories [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Innovative dynamics (left) vs policy trajectories [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Policy NashConv and relative NashConv of ERID with BNN, ERID with Smith, Cross learning in nonstationary RPS. away from the Nash equilibrium. This is due to the strategy being on a periodic orbit near the boundary of the simplex at the time of the shift. As the Nash eq…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

65 extracted references · 57 canonical work pages

  1. [1]

    Sherief Abdallah and Victor Lesser. 2008. A multiagent reinforcement learning algorithm with non-linear dynamics. Journal of Artificial Intelligence Research 33 (2008), 521–549

  2. [2]

    Jeffrey L Adler and Victor J Blue. 2002. A cooperative multi-agent transporta- tion management and route guidance system. Transportation Research Part C: Emerging Technologies 10, 5-6 (2002), 433–454

  3. [3]

    Aurell and B

    A. Aurell and B. Djehiche. 2018. Mean-field type modeling of nonlocal crowd aversion in pedestrian crowd dynamics.SIAM Journal on Control and Optimization 56, 1 (2018), 434–455

  4. [4]

    Aurell and B

    A. Aurell and B. Djehiche. 2020. Behavior near walls in the mean-field approach to crowd dynamics. SIAM J. Appl. Math. 80, 3 (2020), 1153–1174

  5. [5]

    Wolfram Barfuss, Jonathan F Donges, and Jürgen Kurths. 2019. Deterministic limit of temporal difference reinforcement learning for stochastic games.Physical Review E 99, 4 (2019), 043305

  6. [6]

    Barreiro-Gomez

    J. Barreiro-Gomez. 2022. Stochastic differential games for crowd evacuation problems: A paradox. Automatica 140, 2022 (2022), 110271

  7. [7]

    Barreiro-Gomez, I

    J. Barreiro-Gomez, I. Mas, J. I. Giribet, P. Moreno, C. Ocampo-Martínez, R. Sánchez- Pe na, and N. Quijano. 2021. Distributed data-driven UAV formation control via evolutionary games: Experimental results. Journal of the Franklin Institute 358, 10 (2021), 5334–5352

  8. [8]

    Julian Barreiro-Gomez and Nader Masmoudi. 2023. Differential games for crowd dynamics and applications. Mathematical Models and Methods in Applied Sciences 33, 13 (2023), 2703–2742

Show all 65 references
  1. [9]

    Julian Barreiro-Gomez, Germán Obando, Gerardo Riaño-Briceño, Nicanor Qui- jano, and Carlos Ocampo-Martínez. 2015. Decentralized control for urban drainage systems via population dynamics: Bogotá case study. In 2015 Euro- pean Control Conference (ECC) . IEEE, 2426–2431

  2. [10]

    Barreiro-Gomez and H

    J. Barreiro-Gomez and H. Tembine. 2018. Constrained evolutionary games by using a mixture of imitation dynamics. Automatica 97 (2018), 254–262

  3. [11]

    Tilman Börgers and Rajiv Sarin. 1997. Learning through reinforcement and replicator dynamics. Journal of economic theory 77, 1 (1997), 1–14

  4. [12]

    1950.Solutions of games by differential equations

    George W Brown and John Von Neumann. 1950.Solutions of games by differential equations. Rand Corporation

  5. [13]

    J. A. Carrillo, S. Martin, and M. Wolfram. 2016. An improved version of the Hughes model for pedestrian flow. Mathematical Models and Methods in Applied Sciences 26, 4 (2016), 671–697

  6. [14]

    Jorge Cortes, Sonia Martinez, Timur Karatas, and Francesco Bullo. 2004. Coverage control for mobile sensing networks.IEEE Transactions on robotics and Automation 20, 2 (2004), 243–255

  7. [15]

    John G Cross. 1973. A stochastic learning model of economic behavior. The quarterly journal of economics 87, 2 (1973), 239–266

  8. [16]

    Drew Fudenberg. 1991. Game theory. MIT press

  9. [17]

    Luís García, Julian Barreiro-Gomez, Eduardo Escobar, Duván Téllez, Nicanor Quijano, and Carlos Ocampo-Martínez. 2015. Modeling and real-time control of urban drainage systems: A review. Advances in Water Resources 85 (2015), 120–132

  10. [18]

    Daniel Hennes, Michael Kaisers, and Karl Tuyls. 2010. RESQ-learning in stochastic games. In Adaptive and Learning Agents Workshop at AAMAS . Citeseer, 8

  11. [19]

    Daniel Hennes, Dustin Morrill, Shayegan Omidshafiei, Rémi Munos, Julien Pero- lat, Marc Lanctot, Audrunas Gruslys, Jean-Baptiste Lespiau, Paavo Parmas, Edgar Duéñez-Guzmán, et al. 2020. Neural replicator dynamics: Multiagent learning via hedging policy gradients. In Proceeding...

  12. [20]

    Daniel Hennes, Karl Tuyls, and Matthias Rauterberg. 2009. State-coupled replica- tor dynamics.. In AAMAS (2). 789–796

  13. [21]

    Josef Hofbauer. 2011. Deterministic evolutionary game dynamics. (2011)

  14. [22]

    Josef Hofbauer and William H Sandholm. 2009. Stable games and their dynamics. Journal of Economic theory 144, 4 (2009), 1665–1693

  15. [23]

    Josef Hofbauer, Sylvain Sorin, and Yannick Viossat. 2009. Time average replicator and best-reply dynamics. Mathematics of Operations Research 34, 2 (2009), 263– 269

  16. [24]

    R. L. Hughes. 2002. A continuum theory for the flow of pedestrians.Transportation Research Part B 36, 2002 (2002), 507–535

  17. [25]

    R. L. Hughes. 2003. The flow of human crowds. Annual Review of Fluid Mechanics 18, 1 (2003), 169–182

  18. [26]

    Michael Kaisers, Daan Bloembergen, and Karl Tuyls. 2012. A common gradient in multi-agent reinforcement learning. In AAMAS. 1393–1394

  19. [27]

    Michael Kaisers and Karl Tuyls. 2010. Frequency adjusted multi-agent Q-learning. In Proceedings of the 9th International Conference on Autonomous Agents and Multiagent Systems: volume 1-Volume 1. 309–316

  20. [28]

    Ardeshir Kianercy and Aram Galstyan. 2012. Dynamics of Boltzmann Q learning in two-player two-action games. Physical Review E 85, 4 (2012), 041145

  21. [29]

    Tomas Klos, Gerrit Jan Van Ahee, and Karl Tuyls. 2010. Evolutionary dynamics of regret minimization. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 82–96

  22. [30]

    Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel. 2017. A unified game- theoretic approach to multiagent reinforcement learning. Advances in neural information processing systems 30 (2017)

  23. [31]

    Sascha Lange, Thomas Gabel, and Martin Riedmiller. 2012. Batch reinforcement learning. In Reinforcement learning: State-of-the-art. Springer, 45–73

  24. [32]

    Jae Won Lee, Jonghun Park, O Jangmin, Jongwoo Lee, and Euyseok Hong. 2007. A multiagent approach to𝑞-learning for daily stock trading. IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans 37, 6 (2007), 864–877

  25. [33]

    Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 (2015)

  26. [34]

    Jason R Marden. 2012. State based potential games. Automatica 48, 12 (2012), 3075–3088

  27. [35]

    Jason R Marden, Shalom D Ruben, and Lucy Y Pao. 2013. A model-free approach to wind farm control using game theoretic methods. IEEE Transactions on Control Systems Technology 21, 4 (2013), 1207–1214

  28. [36]

    Nuno C Martins, Jair Certório, and Matthew S Hankins. 2024. Counterclockwise Dissipativity, Potential Games and Evolutionary Nash Equilibrium Learning. arXiv preprint arXiv:2408.00647 (2024)

  29. [37]

    Diego Marti Mason, Leonardo Stella, and Dario Bauso. 2020. Evolutionary game dynamics for crowd behavior in emergency evacuations. In 2020 59th IEEE Con- ference on Decision and Control (CDC) . IEEE, 1672–1677

  30. [38]

    Panayotis Mertikopoulos and William H Sandholm. 2016. Learning in games via reinforcement and regularization. Mathematics of Operations Research 41, 4 (2016), 1297–1324

  31. [39]

    Panayotis Mertikopoulos and William H Sandholm. 2018. Riemannian game dynamics. Journal of Economic Theory 177 (2018), 315–364

  32. [40]

    Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015. Human-level control through deep reinforcement learning. Nature 518, 7540 (2015), 529–533

  33. [41]

    Julien Perolat, Bart De Vylder, Daniel Hennes, Eugene Tarassov, Florian Strub, Vincent de Boer, Paul Muller, Jerome T Connor, Neil Burch, Thomas Anthony, et al

  34. [42]

    Julien Perolat, Remi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro Ortega, Neil Burch, Thomas Anthony, David Balduzzi, Bart De Vylder, et al. 2021. From Poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regular...

  35. [43]

    Ramirez-Llanos and N

    E. Ramirez-Llanos and N. Quijano. 2010. A population dynamics approach for the water distribution problem. Internat. J. Control 83 (2010), 1947–1964. Issue 9

  36. [44]

    William H Sandholm. 2010. Population games and evolutionary dynamics . MIT press

  37. [45]

    William H Sandholm. 2015. Population games and deterministic evolutionary dynamics. In Handbook of game theory with economic applications . Vol. 4. Elsevier, 703–778

  38. [46]

    William H Sandholm, Emin Dokumacı, and Ratul Lahkar. 2008. The projection dynamic and the replicator dynamic. Games and Economic Behavior 64, 2 (2008), 666–683

  39. [47]

    David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. 2017. Mastering the game of go without human knowledge. nature 550, 7676 (2017), 354–359

  40. [48]

    J Maynard Smith. 1974. The theory of games and the evolution of animal conflicts. Journal of theoretical biology 47, 1 (1974), 209–221

  41. [49]

    J Maynard Smith and George R Price. 1973. The logic of animal conflict. Nature 246, 5427 (1973), 15–18

  42. [50]

    Michael J Smith. 1984. The stability of a dynamic model of traffic assignment—an application of a method of Lyapunov.Transportation science 18, 3 (1984), 245–252

  43. [51]

    Leonardo Stella, Wouter Baar, and Dario Bauso. 2022. Lower network degrees promote cooperation in the prisoner’s dilemma with environmental feedback. IEEE Control Systems Letters 6 (2022), 2725–2730

  44. [52]

    Leonardo Stella and Dario Bauso. 2019. Bio-inspired evolutionary dynamics on complex networks under uncertain cross-inhibitory signals. Automatica 100 (2019), 61–66

  45. [53]

    Leonardo Stella and Dario Bauso. 2023. The impact of irrational behaviors in the optional prisoner’s dilemma with game-environment feedback. International Journal of Robust and Nonlinear Control 33, 9 (2023), 5145–5158

  46. [54]

    Andrew R Tilman, Joshua B Plotkin, and Erol Akçay. 2020. Evolutionary games with environmental feedbacks. Nature communications 11, 1 (2020), 915

  47. [55]

    Karl Tuyls, Dries Heytens, Ann Nowe, and Bernard Manderick. 2003. Extended replicator dynamics as a key to reinforcement learning in multi-agent systems. In Machine Learning: ECML 2003: 14th European Conference on Machine Learning, Cavtat-Dubrovnik, Croatia, September 22-26, 2...

  48. [56]

    Karl Tuyls, Katja Verbeeck, and Tom Lenaerts. 2003. A selection-mutation model for q-learning in multi-agent systems. In Proceedings of the second international joint conference on Autonomous agents and multiagent systems . 693–700

  49. [57]

    Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, An- drew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al. 2019. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature 575, 7782 (2019), 350–354

  50. [58]

    Yannick Viossat and Andriy Zapechelnyuk. 2013. No-regret dynamics and ficti- tious play. Journal of Economic Theory 148, 2 (2013), 825–842

  51. [59]

    Peter Vrancx, Karl Tuyls, Ronald L Westra, and Ann Nowé. 2008. Switching dynamics of multi-agent learning. AAMAS (1) 2008 (2008), 307–313

  52. [60]

    Jörgen W Weibull. 1997. Evolutionary game theory . MIT press

  53. [61]

    Joshua S Weitz, Ceyhun Eksin, Keith Paarporn, Sam P Brown, and William C Ratcliff. 2016. An oscillating tragedy of the commons in replicator dynamics with game-environment feedback. Proceedings of the National Academy of Sciences 113, 47 (2016), E7518–E7525

  54. [62]

    Yaodong Yang and Jun Wang. 2020. An overview of multi-agent reinforcement learning from game theoretical perspective. arXiv preprint arXiv:2011.00583 (2020)

  55. [63]

    Kaiqing Zhang, Zhuoran Yang, and Tamer Başar. 2021. Multi-agent reinforce- ment learning: A selective overview of theories and algorithms. Handbook of reinforcement learning and control (2021), 321–384

  56. [64]

    Tuo Zhang, Harsh Gupta, Kumar Suprabhat, and Leonardo Stella. 2023. A multi- agent reinforcement learning approach to promote cooperation in evolutionary games on networks with environmental feedback. In 2023 62nd IEEE Conference on Decision and Control (CDC) . IEEE, 2196–2201

  57. [2022]

    Science 378, 6623 (2022), 990–996

    Mastering the game of Stratego with model-free multiagent reinforcement learning. Science 378, 6623 (2022), 990–996

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.