REVIEW 4 major objections 6 minor 65 references
Experience-replay Innovative Dynamics
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read ERID is a stateless multi-agent learning algorithm whose experience-replay updates, with a tunable revision protocol, prove convergence to the trajectories of BNN, Smith, and Smith-replicator pairwise dynamics in the limit of vanishing…
desk verdict Novel and readable, but Theorem 1 is false as stated—the boundary counterexample holds—and the proof's key expectation step is invalid. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is experience replay interpreted as a reward-smoothing average. At each step the algorithm keeps a buffer of the last $K$ action–reward pairs, computes per-action averages $\bar{r}_i$ and the global average $\bar{r}$, then updates the policy with a discrete revision-protocol equation, $\pi_i(t+1) \leftarrow \pi_i(t) + \alpha(\sum_j \pi_j(t)\eta_{ji} - \pi_i(t)\sum_j \eta_{ij})$. Each target dynamics is selected by a protocol factor $\eta_{ij}$ built from the reward averages: $[\bar{r}_j - \bar{r}]_+$ for BNN, $[\bar{r}_j - \bar{r}_i]_+$ for Smith, and the constrained version (12) for Smith-replicator pairwise dynamics. The proof works by showing that as $\alpha K\to 0$ and $K\to\infty$, the buffer averages converge to current expected payoffs, so the discrete update becomes the corresponding ODE.
What would settle it
Run ERID with the BNN protocol on a zero-sum game over a grid of $K$ and $\alpha$ values with $\alpha K$ small and compare the empirical policy trajectory to the BNN ODE trajectory; if the distance does not shrink to zero, or if a direct calculation exhibits a case with random $|I_i|$ where $E(\sum_{j\in I_i} b_j / |I_i|) \neq E(\sum_{j\in I_i} b_j)/E(|I_i|)$, the central convergence claim is refuted.
Extended reading notes
Core claim
The central claim is that a single experience-replay learning rule can reproduce the trajectories of three 'innovative' evolutionary dynamics. For the BNN protocol factor $\eta_{ij} = [\bar{r}_j - \bar{r}]_+$, the update in equation (6) becomes, in the limit $\alpha K \to 0$ with $K \to \infty$, the BNN differential equation (4)–(5); Theorems 2 and 3 state the same trajectory convergence for Smith dynamics and for Smith-replicator-based pairwise dynamics. The result is a bridge: MARL algorithms built on ERID inherit the convergence properties of these dynamics in stable and null-stable games, a guarantee previously available mainly for replicator dynamics. The empirical section shows the stochastic ERID trajectories tracking the deterministic dynamics in matching pennies and biased RPS, and tracking the shifting Nash equilibrium in a nonstationary RPS game.
Load-bearing premise
The proof assumes that the average reward observed for an action in the replay buffer equals the ratio of expected sums and that policies stay essentially constant over the whole buffer window, so buffer averages behave like current expected payoffs; if those assumptions fail, the claimed trajectory match is not established.
Editorial extensions
If this is right
- In zero-sum games, ERID with the BNN or Smith protocol factor can converge to the Nash equilibrium, because the underlying innovative dynamics do; this is exactly where replicator-dynamics learners only orbit or need time averaging.
- In nonstationary environments, ERID's NashConv keeps dropping after payoff changes, while time-averaged replicator-based learners lag because their accumulated averages retain old equilibrium bias.
- Swapping the revision protocol in the update rule selects a different evolutionary dynamics, so the same algorithm covers BNN, Smith, and constrained pairwise dynamics without changing the replay mechanism.
- The matching-pennies and biased-RPS experiments show the stochastic ERID trajectories track the deterministic dynamics trajectories closely enough to inherit their qualitative convergence behavior.
Reading between the lines
- The same replay-averaging construction could be applied to other revision protocols, such as projection or logit dynamics, yielding a whole family of MARL algorithms indexed by protocol factor; the paper does not explore these.
- A finite-sample analysis of the buffer averages would reveal how $K$ and $\alpha$ must be coupled in practice; the theorem only states the joint limit $\alpha K\to 0$, leaving the rate of convergence unspecified.
- Because ERID is stateless and uses only aggregate rewards, it could be combined with function approximation by applying the update to a parameterized policy in expectation, though the paper does not address deep RL.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Experience-replay Innovative Dynamics (ERID), a stateless multi-agent reinforcement learning algorithm that uses a replay buffer and a tunable revision-protocol factor to approximate three evolutionary dynamics: Brown-von Neumann-Nash (BNN), Smith, and Smith-replicator-based pairwise dynamics. The central theoretical claim is that, under the joint limit αK→0 and K→∞, the policy trajectories of ERID converge to the trajectories of these dynamics (Theorems 1–3). The paper also reports experiments in matching pennies, biased Rock-Paper-Scissors, and a nonstationary Rock-Paper-Scissors game, comparing ERID with cross learning and claiming that ERID adapts better to environmental changes.
Significance. If the theoretical claims were established, ERID would provide a general mechanism for transferring the convergence guarantees of innovative dynamics to MARL, going beyond the replicator dynamics that dominate the current EGT-MARL literature. The algorithmic idea of using experience replay to implement non-linear revision protocols is original, and the empirical comparisons, especially the nonstationary RPS experiment, are suggestive and align with the qualitative story. The paper also makes a useful conceptual point about time-averaged replicator dynamics being slow to adapt. However, the central theorems are not proven, and Theorem 1 is false as stated; these issues currently prevent the paper from delivering on its main promise.
major comments (4)
- [Theorem 1, Section 3.3] Theorem 1 is false as stated because it does not restrict initial policies to the interior of the simplex. If an action i has probability zero at some time, then by Eq. (2) the set I_i is empty and r̄_i is set to 0; since action i is never selected, I_i remains empty forever and r̄_i stays 0. In the two-player normal-form game with payoff matrix [[2,2],[0,0]] for both players and initial policy (0,1), the BNN dynamics (4) has a positive derivative for the first action at the boundary and converges to (1,0), whereas ERID update (6) keeps the first action at probability 0 for all time because all positive-part terms are zero. The proof explicitly assumes the initial probability of action i is positive immediately before Eq. (7), so the proof does not cover the statement as written.
- [Equation (7), Section 3.3] The first equality in Eq. (7) is mathematically invalid. The expected average reward of action i is the expectation of the ratio of the sum of rewards in I_i to the size of I_i, and this is not equal to the ratio of the expectation of the sum to the expectation of the size, because |I_i| is a random variable depending on the sequence of sampled actions. The proof therefore does not justify the subsequent replacement of buffer averages by current expected payoffs, even when all probabilities are positive. A correct argument would need a proper stochastic-approximation or two-time-scale analysis, which is not supplied.
- [Limit αK→0 and K→∞, Section 3.3] The proof asserts that under αK→0 and K→∞ the policies remain effectively constant over the buffer window and then invokes the law of large numbers, but this stochastic-averaging step is not made rigorous. The bound on the policy difference over j steps only controls the drift for a fixed j; summing up to K requires a uniform argument with condition αK→0, and the interaction between the two limits, K growing for averaging and αK shrinking for policy constancy, is never formalized. In addition, the passage from the discrete update to the ODE (9) via the parameter θ is only asserted; no error bounds or compactness argument is given.
- [Theorems 2 and 3, Sections 3.4 and 3.5] Theorems 2 and 3 are stated without proofs, with the remark that they are similar to Theorem 1. Since the proof of Theorem 1 is invalid, these results are unsupported. Moreover, the boundary failure identified in Theorem 1 applies equally to the Smith update (11) and the Smith-replicator pairwise update (13), because r̄_i is zero for any never-sampled action, preventing the policy from leaving the boundary. The manuscript needs either complete proofs or a clear restriction to interior initial conditions and non-degenerate exploration for all three theorems.
minor comments (6)
- [Section 3 heading] The heading contains a typo: 'main contibution' should be 'main contribution'.
- [Algorithm 1] Algorithm 1 lists the input learning rate as theta, while the text and update rules consistently use alpha; please unify the notation.
- [Section 1.1 and reference [19]] The text refers to 'Hennis et al.', but the cited work is by Hennes et al.; please correct the name.
- [Section 3.5, reference [9]] The Smith-replicator-based pairwise dynamics is attributed to reference [9], which is a paper on urban drainage systems; this reference does not appear to define the dynamics in Eq. (12). Please cite the correct source for these dynamics.
- [Equation (7) notation] The superscript notation in Eq. (7) is confusing: r̄ with superscript [1] is used for the player-1 empirical average, but the expectation E is not formally defined with respect to the buffer randomness conditional on the current policy; please clarify the probability space and conditioning.
- [Figures 2 and 3] The figures compare simulated dynamics and ERID trajectories only visually; adding a quantitative distance metric, such as average Euclidean distance over time, would make the claimed agreement more precise.
Circularity Check
ERID's BNN/Smith convergence is largely designed into the protocol factors, but the stochastic-averaging step is independent content; no fitted-parameter or self-citation circularity.
full rationale
The central convergence theorems are not circular in the fitted-prediction sense. Equation (3) is a generic discrete revision-protocol update; the BNN, Smith, and Smith-replicator variants are obtained by substituting protocol factors eta_ij = [rbar_j - rbar]+, eta_ij = [rbar_j - rbar_i]+, and the constrained form in Eq. (12). Thus the continuous-time limit in Eq. (9) is, by design, the BNN ODE in Eq. (4): the algorithm is intentionally a discretized version of the target dynamics. That design choice is the paper's stated contribution ('by appropriately adjusting the revision protocols, the behavior of our algorithm mirrors the trajectories'), and it is not a hidden fitted parameter. The independent mathematical content lies in Eq. (7), which attempts to replace buffer averages by current expected payoffs under K->infinity and alpha K -> 0, and in the discrete-to-continuous limit. That content is not derived from, and does not assume, the theorem conclusion; it is a stochastic-approximation argument. Its flaws -- the E(sum b_j)/E(|I_i|) interchange, the boundary case rbar_i = 0 for unvisited actions, and the omitted proofs of Theorems 2 and 3 (Sections 3.4 and 3.5) -- are correctness and completeness concerns, not circularity. There are no load-bearing self-citations: references to the authors' earlier work are background or motivational. No fitted parameter is later reported as a prediction. Accordingly, no circular step rises to the enumerated patterns; the only mild issue is that the convergence is largely designed into the update rule, which suppresses the circularity score only slightly above zero.
Assumptions & free parameters
free parameters (2)
- Learning rate α =
1e-5 in experiments; in theory α→0
- Buffer size K =
1000 in experiments; in theory K→∞
assumptions (5)
- domain assumption The reward function R is bounded and lies in [R_min, R_max] for finite normal-form games.
- ad hoc to paper E(r̄_i) = E(Σ b_j)/E(|I_i|), a ratio of expectations, holds for the replay buffer average.
- ad hoc to paper Under αK→0 and K→∞, policies remain effectively constant over the buffer window, so buffer averages equal current expected payoffs.
- domain assumption The policy update rule (3) keeps policies on the probability simplex.
- standard math The discrete update with step α becomes the continuous-time ODE as α→0.
Cite this review
Pith. "Pith review of Experience-replay Innovative Dynamics." pith.science (2026). https://pith.science/paper/EELWG5VN
@misc{pith2026250112199,
author = {Pith},
title = {Pith review of: Experience-replay Innovative Dynamics},
year = {2026},
howpublished = {\url{https://pith.science/paper/EELWG5VN}},
note = {Machine review of arXiv:2501.12199}
}
read the original abstract
Despite its groundbreaking success, multi-agent reinforcement learning (MARL) still suffers from instability and nonstationarity. Replicator dynamics, the most well-known model from evolutionary game theory (EGT), provide a theoretical framework for the convergence of the trajectories to Nash equilibria and, as a result, have been used to ensure formal guarantees for MARL algorithms in stable game settings. However, they exhibit the opposite behavior in other settings, which poses the problem of finding alternatives to ensure convergence. In contrast, innovative dynamics, such as the Brown-von Neumann-Nash (BNN) or Smith, result in periodic trajectories with the potential to approximate Nash equilibria. Yet, no MARL algorithms based on these dynamics have been proposed. In response to this challenge, we develop a novel experience replay-based MARL algorithm that incorporates revision protocols as tunable hyperparameters. We demonstrate, by appropriately adjusting the revision protocols, that the behavior of our algorithm mirrors the trajectories resulting from these dynamics. Importantly, our contribution provides a framework capable of extending the theoretical guarantees of MARL algorithms beyond replicator dynamics. Finally, we corroborate our theoretical findings with empirical results.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Sherief Abdallah and Victor Lesser. 2008. A multiagent reinforcement learning algorithm with non-linear dynamics. Journal of Artificial Intelligence Research 33 (2008), 521–549
work page 2008
-
[2]
Jeffrey L Adler and Victor J Blue. 2002. A cooperative multi-agent transporta- tion management and route guidance system. Transportation Research Part C: Emerging Technologies 10, 5-6 (2002), 433–454
work page 2002
-
[3]
A. Aurell and B. Djehiche. 2018. Mean-field type modeling of nonlocal crowd aversion in pedestrian crowd dynamics.SIAM Journal on Control and Optimization 56, 1 (2018), 434–455
work page 2018
-
[4]
A. Aurell and B. Djehiche. 2020. Behavior near walls in the mean-field approach to crowd dynamics. SIAM J. Appl. Math. 80, 3 (2020), 1153–1174
work page 2020
-
[5]
Wolfram Barfuss, Jonathan F Donges, and Jürgen Kurths. 2019. Deterministic limit of temporal difference reinforcement learning for stochastic games.Physical Review E 99, 4 (2019), 043305
work page 2019
-
[6]
J. Barreiro-Gomez. 2022. Stochastic differential games for crowd evacuation problems: A paradox. Automatica 140, 2022 (2022), 110271
work page 2022
-
[7]
J. Barreiro-Gomez, I. Mas, J. I. Giribet, P. Moreno, C. Ocampo-Martínez, R. Sánchez- Pe na, and N. Quijano. 2021. Distributed data-driven UAV formation control via evolutionary games: Experimental results. Journal of the Franklin Institute 358, 10 (2021), 5334–5352
work page 2021
-
[8]
Julian Barreiro-Gomez and Nader Masmoudi. 2023. Differential games for crowd dynamics and applications. Mathematical Models and Methods in Applied Sciences 33, 13 (2023), 2703–2742
work page 2023
Show all 65 references
-
[9]
Julian Barreiro-Gomez, Germán Obando, Gerardo Riaño-Briceño, Nicanor Qui- jano, and Carlos Ocampo-Martínez. 2015. Decentralized control for urban drainage systems via population dynamics: Bogotá case study. In 2015 Euro- pean Control Conference (ECC) . IEEE, 2426–2431
2015
-
[10]
Barreiro-Gomez and H
J. Barreiro-Gomez and H. Tembine. 2018. Constrained evolutionary games by using a mixture of imitation dynamics. Automatica 97 (2018), 254–262
2018
-
[11]
Tilman Börgers and Rajiv Sarin. 1997. Learning through reinforcement and replicator dynamics. Journal of economic theory 77, 1 (1997), 1–14
1997
-
[12]
1950.Solutions of games by differential equations
George W Brown and John Von Neumann. 1950.Solutions of games by differential equations. Rand Corporation
1950
-
[13]
J. A. Carrillo, S. Martin, and M. Wolfram. 2016. An improved version of the Hughes model for pedestrian flow. Mathematical Models and Methods in Applied Sciences 26, 4 (2016), 671–697
2016
-
[14]
Jorge Cortes, Sonia Martinez, Timur Karatas, and Francesco Bullo. 2004. Coverage control for mobile sensing networks.IEEE Transactions on robotics and Automation 20, 2 (2004), 243–255
2004
-
[15]
John G Cross. 1973. A stochastic learning model of economic behavior. The quarterly journal of economics 87, 2 (1973), 239–266
1973
-
[16]
Drew Fudenberg. 1991. Game theory. MIT press
1991
-
[17]
Luís García, Julian Barreiro-Gomez, Eduardo Escobar, Duván Téllez, Nicanor Quijano, and Carlos Ocampo-Martínez. 2015. Modeling and real-time control of urban drainage systems: A review. Advances in Water Resources 85 (2015), 120–132
2015
-
[18]
Daniel Hennes, Michael Kaisers, and Karl Tuyls. 2010. RESQ-learning in stochastic games. In Adaptive and Learning Agents Workshop at AAMAS . Citeseer, 8
2010
-
[19]
Daniel Hennes, Dustin Morrill, Shayegan Omidshafiei, Rémi Munos, Julien Pero- lat, Marc Lanctot, Audrunas Gruslys, Jean-Baptiste Lespiau, Paavo Parmas, Edgar Duéñez-Guzmán, et al. 2020. Neural replicator dynamics: Multiagent learning via hedging policy gradients. In Proceeding...
2020
-
[20]
Daniel Hennes, Karl Tuyls, and Matthias Rauterberg. 2009. State-coupled replica- tor dynamics.. In AAMAS (2). 789–796
2009
-
[21]
Josef Hofbauer. 2011. Deterministic evolutionary game dynamics. (2011)
2011
-
[22]
Josef Hofbauer and William H Sandholm. 2009. Stable games and their dynamics. Journal of Economic theory 144, 4 (2009), 1665–1693
2009
-
[23]
Josef Hofbauer, Sylvain Sorin, and Yannick Viossat. 2009. Time average replicator and best-reply dynamics. Mathematics of Operations Research 34, 2 (2009), 263– 269
2009
-
[24]
R. L. Hughes. 2002. A continuum theory for the flow of pedestrians.Transportation Research Part B 36, 2002 (2002), 507–535
2002
-
[25]
R. L. Hughes. 2003. The flow of human crowds. Annual Review of Fluid Mechanics 18, 1 (2003), 169–182
2003
-
[26]
Michael Kaisers, Daan Bloembergen, and Karl Tuyls. 2012. A common gradient in multi-agent reinforcement learning. In AAMAS. 1393–1394
2012
-
[27]
Michael Kaisers and Karl Tuyls. 2010. Frequency adjusted multi-agent Q-learning. In Proceedings of the 9th International Conference on Autonomous Agents and Multiagent Systems: volume 1-Volume 1. 309–316
2010
-
[28]
Ardeshir Kianercy and Aram Galstyan. 2012. Dynamics of Boltzmann Q learning in two-player two-action games. Physical Review E 85, 4 (2012), 041145
2012
-
[29]
Tomas Klos, Gerrit Jan Van Ahee, and Karl Tuyls. 2010. Evolutionary dynamics of regret minimization. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 82–96
2010
-
[30]
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel. 2017. A unified game- theoretic approach to multiagent reinforcement learning. Advances in neural information processing systems 30 (2017)
2017
-
[31]
Sascha Lange, Thomas Gabel, and Martin Riedmiller. 2012. Batch reinforcement learning. In Reinforcement learning: State-of-the-art. Springer, 45–73
2012
-
[32]
Jae Won Lee, Jonghun Park, O Jangmin, Jongwoo Lee, and Euyseok Hong. 2007. A multiagent approach to𝑞-learning for daily stock trading. IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans 37, 6 (2007), 864–877
2007
-
[33]
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 (2015)
2015 arXiv
-
[34]
Jason R Marden. 2012. State based potential games. Automatica 48, 12 (2012), 3075–3088
2012
-
[35]
Jason R Marden, Shalom D Ruben, and Lucy Y Pao. 2013. A model-free approach to wind farm control using game theoretic methods. IEEE Transactions on Control Systems Technology 21, 4 (2013), 1207–1214
2013
-
[36]
Nuno C Martins, Jair Certório, and Matthew S Hankins. 2024. Counterclockwise Dissipativity, Potential Games and Evolutionary Nash Equilibrium Learning. arXiv preprint arXiv:2408.00647 (2024)
2024 arXiv
-
[37]
Diego Marti Mason, Leonardo Stella, and Dario Bauso. 2020. Evolutionary game dynamics for crowd behavior in emergency evacuations. In 2020 59th IEEE Con- ference on Decision and Control (CDC) . IEEE, 1672–1677
2020
-
[38]
Panayotis Mertikopoulos and William H Sandholm. 2016. Learning in games via reinforcement and regularization. Mathematics of Operations Research 41, 4 (2016), 1297–1324
2016
-
[39]
Panayotis Mertikopoulos and William H Sandholm. 2018. Riemannian game dynamics. Journal of Economic Theory 177 (2018), 315–364
2018
-
[40]
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015. Human-level control through deep reinforcement learning. Nature 518, 7540 (2015), 529–533
2015
-
[41]
Julien Perolat, Bart De Vylder, Daniel Hennes, Eugene Tarassov, Florian Strub, Vincent de Boer, Paul Muller, Jerome T Connor, Neil Burch, Thomas Anthony, et al
-
[42]
Julien Perolat, Remi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro Ortega, Neil Burch, Thomas Anthony, David Balduzzi, Bart De Vylder, et al. 2021. From Poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regular...
2021
-
[43]
Ramirez-Llanos and N
E. Ramirez-Llanos and N. Quijano. 2010. A population dynamics approach for the water distribution problem. Internat. J. Control 83 (2010), 1947–1964. Issue 9
2010
-
[44]
William H Sandholm. 2010. Population games and evolutionary dynamics . MIT press
2010
-
[45]
William H Sandholm. 2015. Population games and deterministic evolutionary dynamics. In Handbook of game theory with economic applications . Vol. 4. Elsevier, 703–778
2015
-
[46]
William H Sandholm, Emin Dokumacı, and Ratul Lahkar. 2008. The projection dynamic and the replicator dynamic. Games and Economic Behavior 64, 2 (2008), 666–683
2008
-
[47]
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. 2017. Mastering the game of go without human knowledge. nature 550, 7676 (2017), 354–359
2017
-
[48]
J Maynard Smith. 1974. The theory of games and the evolution of animal conflicts. Journal of theoretical biology 47, 1 (1974), 209–221
1974
-
[49]
J Maynard Smith and George R Price. 1973. The logic of animal conflict. Nature 246, 5427 (1973), 15–18
1973
-
[50]
Michael J Smith. 1984. The stability of a dynamic model of traffic assignment—an application of a method of Lyapunov.Transportation science 18, 3 (1984), 245–252
1984
-
[51]
Leonardo Stella, Wouter Baar, and Dario Bauso. 2022. Lower network degrees promote cooperation in the prisoner’s dilemma with environmental feedback. IEEE Control Systems Letters 6 (2022), 2725–2730
2022
-
[52]
Leonardo Stella and Dario Bauso. 2019. Bio-inspired evolutionary dynamics on complex networks under uncertain cross-inhibitory signals. Automatica 100 (2019), 61–66
2019
-
[53]
Leonardo Stella and Dario Bauso. 2023. The impact of irrational behaviors in the optional prisoner’s dilemma with game-environment feedback. International Journal of Robust and Nonlinear Control 33, 9 (2023), 5145–5158
2023
-
[54]
Andrew R Tilman, Joshua B Plotkin, and Erol Akçay. 2020. Evolutionary games with environmental feedbacks. Nature communications 11, 1 (2020), 915
2020
-
[55]
Karl Tuyls, Dries Heytens, Ann Nowe, and Bernard Manderick. 2003. Extended replicator dynamics as a key to reinforcement learning in multi-agent systems. In Machine Learning: ECML 2003: 14th European Conference on Machine Learning, Cavtat-Dubrovnik, Croatia, September 22-26, 2...
2003
-
[56]
Karl Tuyls, Katja Verbeeck, and Tom Lenaerts. 2003. A selection-mutation model for q-learning in multi-agent systems. In Proceedings of the second international joint conference on Autonomous agents and multiagent systems . 693–700
2003
-
[57]
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, An- drew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al. 2019. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature 575, 7782 (2019), 350–354
2019
-
[58]
Yannick Viossat and Andriy Zapechelnyuk. 2013. No-regret dynamics and ficti- tious play. Journal of Economic Theory 148, 2 (2013), 825–842
2013
-
[59]
Peter Vrancx, Karl Tuyls, Ronald L Westra, and Ann Nowé. 2008. Switching dynamics of multi-agent learning. AAMAS (1) 2008 (2008), 307–313
2008
-
[60]
Jörgen W Weibull. 1997. Evolutionary game theory . MIT press
1997
-
[61]
Joshua S Weitz, Ceyhun Eksin, Keith Paarporn, Sam P Brown, and William C Ratcliff. 2016. An oscillating tragedy of the commons in replicator dynamics with game-environment feedback. Proceedings of the National Academy of Sciences 113, 47 (2016), E7518–E7525
2016
-
[62]
Yaodong Yang and Jun Wang. 2020. An overview of multi-agent reinforcement learning from game theoretical perspective. arXiv preprint arXiv:2011.00583 (2020)
2020 arXiv
-
[63]
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar. 2021. Multi-agent reinforce- ment learning: A selective overview of theories and algorithms. Handbook of reinforcement learning and control (2021), 321–384
2021
-
[64]
Tuo Zhang, Harsh Gupta, Kumar Suprabhat, and Leonardo Stella. 2023. A multi- agent reinforcement learning approach to promote cooperation in evolutionary games on networks with environmental feedback. In 2023 62nd IEEE Conference on Decision and Control (CDC) . IEEE, 2196–2201
2023
-
[2022]
Science 378, 6623 (2022), 990–996
Mastering the game of Stratego with model-free multiagent reinforcement learning. Science 378, 6623 (2022), 990–996
2022
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.