Pith. sign in

REVIEW 3 minor 158 references

The Game Changer Problem: Controlling Equilibria with Discrete Rewards

T0 review · 0 major / 3 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read An external designer can force any chosen pure action profile to become the unique Nash equilibrium by picking all rewards from a fixed finite set.

desk verdict The paper frames equilibrium steering as selecting discrete rewards from a finite set to force a target pure profile, with simple feasibility rules and DP for two-player zero-sum and general-sum games. read the letter →

arxiv 2606.29012 v1 pith:MSVR7MZ2 submitted 2026-06-27 cs.GT

classification cs.GT
keywords gamechangerproblemequilibriumcontroldiscreterewardsNashzero-sumgamesgeneral-sumdynamicprogrammingrewardredesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper defines the game changer problem in which a designer alters every entry of a game's reward matrix, but must draw each value from a predetermined finite set, with the goal of making one specific pure action profile the sole equilibrium. For two-player zero-sum games and for general-sum games the authors supply direct feasibility tests that decide whether such a redesign is possible. Because rewards are restricted to discrete values rather than arbitrary reals, the resulting programs admit exact optimal solutions that can be computed by dynamic programming, in contrast to earlier continuous formulations that rely on linear programming.

What carries the argument

The finite discrete reward set together with the feasibility conditions that certify when a target pure profile can be made the unique equilibrium.

What would settle it

A concrete two-player game and finite reward set together with a target profile for which the stated feasibility test returns yes, yet no assignment from the set actually makes that profile the unique equilibrium.

Watch

Extended reading notes

Core claim

In the game changer problem an external designer modifies a game's reward matrix so that a target pure action profile becomes the unique equilibrium, subject to every matrix entry belonging to a given finite set. Simple feasibility characterizations are given for two-player zero-sum games and for general-sum games. The discrete reward structure yields exact optimality and supports efficient dynamic programming algorithms, supplying a sharper alternative to prior continuous reward redesign formulations based on linear programming.

Load-bearing premise

The designer may assign any value from the finite reward set to any matrix entry without further limits on which entries can be changed or on the size of the action spaces.

Editorial extensions

If this is right

  • Feasibility reduces to simple checks rather than solving an optimization problem.
  • Dynamic programming yields exact optimal reward assignments for both zero-sum and general-sum cases.
  • The discrete formulation avoids the approximation gaps that appear when continuous linear programs are discretized after the fact.
  • Any finite reward set that satisfies the characterizations immediately produces a working redesign without further search.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same discrete-reward idea might be tested on games with three or more players to see whether analogous feasibility conditions survive.
  • If the finite set is interpreted as a menu of practical payoff levels, the characterizations give a direct way to decide whether a desired outcome can be engineered with those levels.
  • Dynamic programming routines could be run on small action spaces to produce explicit redesign tables for concrete games.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. The paper introduces the game changer problem, in which an external designer modifies entries of a game's reward matrix—each drawn from a finite discrete set—to render a designated pure action profile the unique Nash equilibrium. For two-player zero-sum and general-sum games the authors supply feasibility characterizations; the discrete structure is further shown to admit exact optimality results together with efficient dynamic-programming algorithms, positioned as a sharper alternative to prior continuous-reward redesign formulations that rely on linear programming.

Significance. If the characterizations and DP optimality claims hold, the work supplies a precise, computationally tractable method for equilibrium control under the realistic constraint that rewards must belong to a finite set. The explicit feasibility conditions and the exact (rather than approximate) optimality results constitute a clear technical contribution relative to continuous LP baselines.

minor comments (3)
  1. [§2] §2, Definition 1: the finite reward set R is introduced without an accompanying small-scale example that illustrates how the designer’s choice of entries interacts with the target profile; adding one would clarify the subsequent feasibility statements.
  2. [§4.2] §4.2, Algorithm 1: the dynamic-programming recurrence is stated at a high level; the base case and the precise state representation (action-profile versus payoff-vector) should be written explicitly to allow direct implementation.
  3. [Table 1] Table 1: the reported running times compare the DP approach only against a generic LP solver; a column showing the size of the action spaces used in the experiments would help readers assess scalability claims.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the careful summary of our work and the recommendation of minor revision. No specific major comments were provided in the report, so we have no individual points to address at this time.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

The paper introduces the game changer problem with discrete rewards from a finite set as an explicit modeling choice and derives feasibility characterizations plus DP algorithms directly from that structure. These are presented as theoretical results (feasibility conditions for zero-sum and general-sum cases) without any reduction to fitted parameters, self-definitional loops, or load-bearing self-citations. The discrete structure is an input assumption, not an output derived from the claimed characterizations. No equations or steps in the abstract or described claims exhibit the enumerated circular patterns; the work is self-contained against external benchmarks.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract-only review; no free parameters, axioms, or invented entities are stated or derivable from the given text.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Game Changer Problem: Controlling Equilibria with Discrete Rewards." pith.science (2026). https://pith.science/paper/MSVR7MZ2

@misc{pith2026260629012,
  author       = {Pith},
  title        = {Pith review of: The Game Changer Problem: Controlling Equilibria with Discrete Rewards},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MSVR7MZ2}},
  note         = {Machine review of arXiv:2606.29012}
}
read the original abstract

We introduce the game changer problem, where an external designer modifies a game's reward matrix to make a target pure action profile the unique equilibrium, subject to the constraint that all entries of the reward matrix come from a finite set. We give simple feasibility characterizations for two-player zero-sum games and general-sum games, and the discrete reward structure yields exact optimality and enables efficient dynamic programming algorithms, providing a sharper alternative to prior continuous reward redesign formulations based on linear programming.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

158 extracted references · 158 canonical work pages

  1. [1]

    Apprenticeship learning via inverse reinforcement learning , author =

  2. [2]

    Proceedings of The 8th International Conference on Autonomous Agents and Multiagent Systems-Volume 2 , pages =

    Multiagent reinforcement learning: algorithm converging to nash equilibrium in general-sum discounted stochastic games , author =. Proceedings of The 8th International Conference on Autonomous Agents and Multiagent Systems-Volume 2 , pages =

  3. [3]

    Proceedings of the 9th International Conference on Autonomous Agents and Multiagent Systems: volume 1-Volume 1 , pages =

    Internal implementation , author =. Proceedings of the 9th International Conference on Autonomous Agents and Multiagent Systems: volume 1-Volume 1 , pages =

  4. [4]

    Journal of the Operational Research Society , publisher =

    On the uniqueness of solutions to linear programs , author =. Journal of the Operational Research Society , publisher =

  5. [5]

    SIAM journal on computing , publisher =

    The nonstochastic multiarmed bandit problem , author =. SIAM journal on computing , publisher =

  6. [6]

    , author =

    Multi-Agent Reinforcement Learning in Common Interest and Fixed Sum Stochastic Games: An Experimental Study. , author =

  7. [7]

    Defense against reward poisoning attacks in reinforcement learning , author =

  8. [8]

    Admissible Policy Teaching through Reward Design , author =

Show all 158 references
  1. [9]

    International Conference on Machine Learning and Data Mining in Pattern Recognition , pages =

    Vulnerability of deep reinforcement learning to policy induction attacks , author =. International Conference on Machine Learning and Data Mining in Pattern Recognition , pages =

  2. [10]

    Dota 2 with large scale deep reinforcement learning , author =

  3. [11]

    International Conference on Artificial Intelligence and Statistics , pages =

    Stochastic linear bandits robust to adversarial attacks , author =. International Conference on Artificial Intelligence and Statistics , pages =

  4. [12]

    ICML , pages =

    Convergence problems of general-sum multiagent reinforcement learning , author =. ICML , pages =

  5. [13]

    International joint conference on artificial intelligence , volume = 17, pages =

    Rational and convergent learning in stochastic games , author =. International joint conference on artificial intelligence , volume = 17, pages =

  6. [14]

    Advances in Neural Information Processing Systems , volume = 34, pages =

    Offline rl without off-policy evaluation , author =. Advances in Neural Information Processing Systems , volume = 34, pages =

  7. [15]

    , author =

    Libratus: The Superhuman AI for No-Limit Poker. , author =. IJCAI , pages =

  8. [16]

    Science , publisher =

    Superhuman AI for multiplayer poker , author =. Science , publisher =

  9. [17]

    Foundations and Trends

    Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems , author =. Foundations and Trends

  10. [18]

    , year = 2021, note =

    Clancey, William J. , year = 2021, note =

  11. [19]

    0803.4058 , archiveprefix =

    Crime and punishment in scientific research , author =. 0803.4058 , archiveprefix =

  12. [20]

    Pluto: The 'Other' Red Planet , author =

  13. [21]

    , year = 1979, address =

    Clancey, William J. , year = 1979, address =

  14. [22]

    , year = 1983, booktitle =

    Clancey, William J. , year = 1983, booktitle =

  15. [23]

    , year = 1984, booktitle =

    Clancey, William J. , year = 1984, booktitle =

  16. [24]

    When is Offline Two-Player Zero-Sum Markov Game Solvable? , author =

  17. [25]

    When is Offline Two-Player Zero-Sum

    Cui, Qiwen and Du, Simon S , year = 2022, journal =. When is Offline Two-Player Zero-Sum

  18. [26]

    Linear programming and extensions , author =

  19. [27]

    Autonomous Robots , publisher =

    A taxonomy for multi-agent robotics , author =. Autonomous Robots , publisher =

  20. [28]

    Blackboard Systems , year = 1986, publisher =

  21. [29]

    Volunteer's dilemma ---

  22. [30]

    Revisiting some common practices in cooperative multi-agent reinforcement learning , author =

  23. [31]

    Adversarial Attacks on Linear Contextual Bandits , author =

  24. [32]

    Adversarial policies: Attacking deep reinforcement learning , author =

  25. [33]

    SIAM Review , publisher =

    f-Finger Morra , author =. SIAM Review , publisher =

  26. [34]

    2017 IEEE international conference on robotics and automation (ICRA) , pages =

    Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates , author =. 2017 IEEE international conference on robotics and automation (ICRA) , pages =

  27. [35]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume = 34, pages =

    Robust stochastic bandit algorithms under probabilistic unbounded adversarial attack , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume = 34, pages =

  28. [36]

    International Conference on Machine Learning , pages =

    Adversarial policy learning in two-player competitive games , author =. International Conference on Machine Learning , pages =

  29. [37]

    Econometrica , publisher =

    A simple adaptive procedure leading to correlated equilibrium , author =. Econometrica , publisher =

  30. [38]

    International Journal of Man-Machine Studies , volume = 20, number = 1, pages =

    Strategic explanations for a diagnostic consultation system , author =. International Journal of Man-Machine Studies , volume = 20, number = 1, pages =. doi:https://doi.org/10.1016/S0020-7373(84)80003-6 , issn =

  31. [39]

    and Rennels, Glenn R

    Hasling, Diane Warner and Clancey, William J. and Rennels, Glenn R. and Test, Thomas , year = 1983, journal =

  32. [40]

    International conference on machine learning , pages =

    Fictitious self-play in extensive-form games , author =. International conference on machine learning , pages =

  33. [41]

    International Conference on Autonomous Agents and Multiagent Systems , pages =

    Towards a fast detection of opponents in repeated stochastic games , author =. International Conference on Autonomous Agents and Multiagent Systems , pages =

  34. [42]

    International Journal of Game Theory , publisher =

    Uniqueness of equilibrium points in bimatrix games , author =. International Journal of Game Theory , publisher =

  35. [43]

    Journal of machine learning research , volume = 4, number =

    Nash Q-learning for general-sum stochastic games , author =. Journal of machine learning research , volume = 4, number =

  36. [44]

    Adversarial attacks on neural network policies , author =

  37. [45]

    International Conference on Decision and Game Theory for Security , pages =

    Deceptive reinforcement learning under adversarial manipulations on cost signals , author =. International Conference on Decision and Game Theory for Security , pages =

  38. [46]

    Proceedings of the 36th International Conference on Machine Learning , publisher =

    Multi-Agent Adversarial Inverse Reinforcement Learning , author =. Proceedings of the 36th International Conference on Machine Learning , publisher =

  39. [47]

    Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 1 , location =

    Value-Based Policy Teaching with Active Indirect Elicitation , author =. Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 1 , location =

  40. [48]

    Science , publisher =

    Human-level performance in 3D multiplayer games with population-based reinforcement learning , author =. Science , publisher =

  41. [49]

    Offline decentralized multi-agent reinforcement learning , author =

  42. [50]

    International Conference on Machine Learning , pages =

    Is pessimism provably efficient for offline rl? , author =. International Conference on Machine Learning , pages =

  43. [51]

    Advances in Neural Information Processing Systems , booktitle =

    Adversarial attacks on stochastic bandits , author =. Advances in Neural Information Processing Systems , booktitle =

  44. [52]

    2020 57th ACM/IEEE Design Automation Conference (DAC) , pages =

    TrojDRL: evaluation of backdoor attacks on deep reinforcement learning , author =. 2020 57th ACM/IEEE Design Automation Conference (DAC) , pages =

  45. [53]

    The International Journal of Robotics Research , publisher =

    Reinforcement learning in robotics: A survey , author =. The International Journal of Robotics Research , publisher =

  46. [54]

    Delving into adversarial attacks on deep policies , author =

  47. [55]

    Journal of Economic Dynamics and Control , publisher =

    Learning competitive pricing strategies by multi-agent reinforcement learning , author =. Journal of Economic Dynamics and Control , publisher =

  48. [56]

    International Conference on Database and Expert Systems Applications , pages =

    A multi-agent Q-learning framework for optimizing stock trading systems , author =. International Conference on Database and Expert Systems Applications , pages =

  49. [57]

    IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans , publisher =

    A multiagent approach to q -learning for daily stock trading , author =. IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans , publisher =

  50. [58]

    Multi-agent reinforcement learning in sequential social dilemmas , author =

  51. [59]

    IEEE Transactions on Games , publisher =

    Multiagent inverse reinforcement learning for two-person zero-sum games , author =. IEEE Transactions on Games , publisher =

  52. [60]

    Tactics of adversarial attack on deep reinforcement learning agents , author =

  53. [61]

    Journal of Artificial Intelligence Research , volume = 66, pages =

    Multi-agent inverse reinforcement learning for certain general-sum stochastic games , author =. Journal of Artificial Intelligence Research , volume = 66, pages =

  54. [62]

    Machine learning proceedings 1994 , publisher =

    Markov games as a framework for multi-agent reinforcement learning , author =. Machine learning proceedings 1994 , publisher =

  55. [63]

    International Conference on Machine Learning , pages =

    Data poisoning attacks on stochastic bandits , author =. International Conference on Machine Learning , pages =

  56. [64]

    , author =

    Value Function Transfer for Deep Multi-Agent Reinforcement Learning Based on N-Step Returns. , author =. IJCAI , pages =

  57. [65]

    ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages =

    Action-manipulation attacks on stochastic bandits , author =. ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages =

  58. [66]

    Provably Efficient Black-Box Action Poisoning Attacks Against Reinforcement Learning , author =

  59. [67]

    Algorithms in multi-agent systems: a holistic perspective from reinforcement learning and game theory , author =

  60. [68]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume = 35, pages =

    Stochastic Graphical Bandits with Adversarial Corruptions , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume = 35, pages =

  61. [69]

    Conference on Learning Theory , pages =

    Corruption-robust exploration in episodic reinforcement learning , author =. Conference on Learning Theory , pages =

  62. [70]

    International Conference on Decision and Game Theory for Security , pages =

    Data poisoning attacks in contextual bandits , author =. International Conference on Decision and Game Theory for Security , pages =

  63. [71]

    Advances in Neural Information Processing Systems , booktitle =

    Policy poisoning in batch reinforcement learning and control , author =. Advances in Neural Information Processing Systems , booktitle =

  64. [72]

    Game Redesign in No-regret Game Playing , author =

  65. [73]

    Markov games of incomplete information for multi-agent reinforcement learning , author =

  66. [74]

    Uniqueness of solution in linear programming , author =

  67. [75]

    Dynamic economic emissions dispatch optimisation using multi-agent reinforcement learning , author =

  68. [76]

    Observable actions , author =

    Markov perfect equilibrium: I. Observable actions , author =. Journal of Economic Theory , publisher =

  69. [77]

    Offline pre-trained multi-agent decision transformer: One big sequence model conquers all starcraftii tasks , author =

  70. [78]

    Naval Research Logistics Quarterly , publisher =

    Constructing bimatrix games with special properties , author =. Naval Research Logistics Quarterly , publisher =

  71. [79]

    Journal of Dynamics & Games , publisher =

    On the uniqueness of Nash equilibrium in strategic-form games , author =. Journal of Dynamics & Games , publisher =

  72. [80]

    Adversarial Bandits with Corruptions: Regret Lower Bound and No-regret Algorithm , author =

  73. [81]

    Journal of Artificial Intelligence Research , publisher =

    Multi-agent Inverse Reinforcement Learning for Certain General-sum Stochastic Games , author =. Journal of Artificial Intelligence Research , publisher =. doi:10.1613/jair.1.11541 , url =

  74. [82]

    Flatland-RL: Multi-agent reinforcement learning on trains , author =

  75. [83]

    Proceedings of the 4th ACM conference on Electronic Commerce , pages =

    k-Implementation , author =. Proceedings of the 4th ACM conference on Electronic Commerce , pages =

  76. [84]

    Journal of Artificial Intelligence Research , volume = 21, pages =

    k-Implementation , author =. Journal of Artificial Intelligence Research , volume = 21, pages =

  77. [85]

    Proceedings of the national academy of sciences , publisher =

    Equilibrium points in n-person games , author =. Proceedings of the national academy of sciences , publisher =

  78. [86]

    Workshop on Learning in Distributed Artificial Intelligence Systems, Workshop on Learning, Interaction, and Organization in Multiagent Environments , pages =

    A modular approach to multi-agent reinforcement learning , author =. Workshop on Learning in Distributed Artificial Intelligence Systems, Workshop on Learning, Interaction, and Organization in Multiagent Environments , pages =

  79. [87]

    An introduction to game theory , author =

  80. [88]

    International Conference on Machine Learning , pages =

    Plan better amid conservatism: Offline multi-agent reinforcement learning with actor rectification , author =. International Conference on Machine Learning , pages =

  81. [89]

    Robust deep reinforcement learning with adversarial attacks , author =

  82. [90]

    A study of gradient descent schemes for general-sum stochastic games , author =

  83. [91]

    Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems , pages =

    Two-timescale algorithms for learning Nash equilibria in general-sum stochastic games , author =. Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems , pages =

  84. [92]

    Uniqueness of

    Quintas, Luis G , year = 1988, publisher =. Uniqueness of

  85. [93]

    Science , publisher =

    New Ways to Make Microcircuits Smaller , author =. Science , publisher =. doi:10.1126/science.208.4447.1019 , issn =

  86. [94]

    , year = 1980, journal =

    Robinson, Arthur L. , year = 1980, journal =

  87. [95]

    Rice, James , year = 1986, number =

  88. [96]

    Handbook of game theory with economic applications , publisher =

    Zero-sum two-person games , author =. Handbook of game theory with economic applications , publisher =

  89. [97]

    arXiv preprint arXiv:2003.12909 , booktitle =

    Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning , author =. arXiv preprint arXiv:2003.12909 , booktitle =

  90. [98]

    Journal of Machine Learning Research , volume = 22, number = 210, pages =

    Policy teaching in reinforcement learning via environment poisoning attacks , author =. Journal of Machine Learning Research , volume = 22, number = 210, pages =

  91. [99]

    Reward poisoning in reinforcement learning: Attacks against unknown learners in unknown environments , author =

  92. [100]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume = 36, pages =

    Saving stochastic bandits from poisoning attacks via limited data verification , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume = 36, pages =

  93. [101]

    Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence,

    Understanding the Limits of Poisoning Attacks in Episodic Reinforcement Learning , author =. Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence,. doi:10.24963/ijcai.2022/471 , url =

  94. [102]

    2012 IEEE International Conference on Systems, Man, and Cybernetics (SMC) , pages =

    Inverse reinforcement learning for decentralized non-cooperative multiagent systems , author =. 2012 IEEE International Conference on Systems, Man, and Cybernetics (SMC) , pages =

  95. [103]

    Autonomous Robots , publisher =

    Reinforcement learning for robot soccer , author =. Autonomous Robots , publisher =

  96. [104]

    doi:10.48550/ARXIV.1702.02284 , url =

    Adversarial Attacks on Neural Network Policies , author =. doi:10.48550/ARXIV.1702.02284 , url =

  97. [105]

    doi:10.48550/ARXIV.1705.06452 , url =

    Delving into adversarial attacks on deep policies , author =. doi:10.48550/ARXIV.1705.06452 , url =

  98. [106]

    Safe, multi-agent, reinforcement learning for autonomous driving , author =

  99. [107]

    Proceedings of the national academy of sciences , publisher =

    Stochastic games , author =. Proceedings of the national academy of sciences , publisher =

  100. [108]

    nature , publisher =

    Mastering the game of Go with deep neural networks and tree search , author =. nature , publisher =

  101. [109]

    nature , publisher =

    Mastering the game of go without human knowledge , author =. nature , publisher =

  102. [110]

    Introduction to multi-armed bandits , author =

  103. [111]

    Proceedings of the first international joint conference on Autonomous agents and multiagent systems: Part 1 , pages =

    A multiagent reinforcement learning algorithm using extended optimal response , author =. Proceedings of the first international joint conference on Autonomous agents and multiagent systems: Part 1 , pages =

  104. [112]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume = 34, pages =

    Stealthy and efficient adversarial attacks against deep reinforcement learning , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume = 34, pages =

  105. [113]

    Vulnerability-aware poisoning mechanism for online rl with unknown dynamics , author =

  106. [114]

    Who is the strongest enemy? towards optimal and efficient evasion attacks in deep rl , author =

  107. [115]

    Revue d'analyse num

    On the uniqueness of the optimal solution in linear programming , author =. Revue d'analyse num

  108. [116]

    Congress of the Italian Association for Artificial Intelligence , pages =

    Hypotheses about typical general human strategic behavior in a concrete case , author =. Congress of the Italian Association for Artificial Intelligence , pages =

  109. [117]

    cooperative agents , author =

    Multi-agent reinforcement learning: Independent vs. cooperative agents , author =. Proceedings of the tenth international conference on machine learning , pages =

  110. [118]

    Advances in Neural Information Processing Systems , volume = 34, pages =

    Pettingzoo: Gym for multi-agent reinforcement learning , author =. Advances in Neural Information Processing Systems , volume = 34, pages =

  111. [119]

    Game Transformations That Preserve

    Emanuel Tewolde , year = 2023, journal =. Game Transformations That Preserve

  112. [120]

    Nature , publisher =

    Grandmaster level in StarCraft II using multi-agent reinforcement learning , author =. Nature , publisher =

  113. [121]

    International Conference on Interactive Collaborative Robotics , pages =

    Multi-agent robotic systems in collaborative robotics , author =. International Conference on Interactive Collaborative Robotics , pages =

  114. [122]

    International Conference on Machine Learning , pages =

    Competitive multi-agent inverse reinforcement learning with sub-optimal demonstrations , author =. International Conference on Machine Learning , pages =

  115. [123]

    Optimism in reinforcement learning with generalized linear function approximation , author =

  116. [124]

    Backdoorl: Backdoor attack against competitive reinforcement learning , author =

  117. [125]

    International Conference on Algorithmic Learning Theory , pages =

    A model selection approach for corruption robust reinforcement learning , author =. International Conference on Algorithmic Learning Theory , pages =

  118. [126]

    2006 , publisher=

    Numerical Optimization , author=. 2006 , publisher=

  119. [127]

    Crop: Certifying robust policies for reinforcement learning through functional smoothing , author =

  120. [128]

    COPA: Certifying Robust Policies for Offline Reinforcement Learning against Poisoning Attacks , author =

  121. [129]

    Reward Poisoning Attacks on Offline Multi-Agent Reinforcement Learning , author =

  122. [130]

    On Faking a

    Wu, Young and McMahan, Jeremy and Zhu, Xiaojin and Xie, Qiaomin , year = 2023, journal =. On Faking a

  123. [131]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume = 37, pages =

    Reward poisoning attacks on offline multi-agent reinforcement learning , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume = 37, pages =

  124. [132]

    arXiv preprint arXiv:2311.00582 , year=

    Minimally Modifying a Markov Game to Achieve Any Nash Equilibrium and Value , author=. arXiv preprint arXiv:2311.00582 , year=

  125. [133]

    Conference on learning theory , pages =

    Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium , author =. Conference on learning theory , pages =

  126. [134]

    CAAI Transactions on Intelligence Technology , publisher =

    Multi-robot path planning based on a deep reinforcement learning DQN algorithm , author =. CAAI Transactions on Intelligence Technology , publisher =

  127. [135]

    , author =

    Value-Based Policy Teaching with Active Indirect Elicitation. , author =. AAAI , volume = 8, pages =

  128. [136]

    Proceedings of the 10th ACM conference on Electronic commerce , pages =

    Policy teaching through reward function learning , author =. Proceedings of the 10th ACM conference on Electronic commerce , pages =

  129. [137]

    arXiv preprint arXiv:2003.12613 , booktitle =

    Adaptive reward-poisoning attacks against reinforcement learning , author =. arXiv preprint arXiv:2003.12613 , booktitle =

  130. [138]

    Advances in Neural Information Processing Systems , volume = 33, pages =

    Robust deep reinforcement learning against adversarial perturbations on state observations , author =. Advances in Neural Information Processing Systems , volume = 33, pages =

  131. [139]

    Corruption-robust offline reinforcement learning , author =

  132. [140]

    Handbook of Reinforcement Learning and Control , publisher =

    Multi-agent reinforcement learning: A selective overview of theories and algorithms , author =. Handbook of Reinforcement Learning and Control , publisher =

  133. [141]

    arXiv preprint arXiv:2101.08452 , booktitle =

    Robust policy gradient against strong data corruption , author =. arXiv preprint arXiv:2101.08452 , booktitle =

  134. [142]

    The ai economist: Improving equality and productivity with ai-driven tax policies , author =

  135. [143]

    Pessimistic minimax value iteration: Provably efficient equilibrium learning from offline datasets , author =

  136. [144]

    Near Optimal Adversarial Attack on UCB Bandits , author =

  137. [145]

    Proceedings of the twenty-ninth annual ACM-SIAM symposium on discrete algorithms , pages=

    Cycles in adversarial regularized learning , author=. Proceedings of the twenty-ninth annual ACM-SIAM symposium on discrete algorithms , pages=. 2018 , organization=

  138. [146]

    Bimatrix games with unique equilibrium points , author=

  139. [147]

    and Shapley, LS , title=

    Bohnenblust, HF and Karlin, S. and Shapley, LS , title=. Contributions to the Theory of Games , number=. 1950 , publisher=

  140. [148]

    International Conference on Machine Learning , pages=

    From chaos to order: Symmetry and conservation laws in game dynamics , author=. International Conference on Machine Learning , pages=. 2020 , organization=

  141. [149]

    1997 , publisher=

    Introduction to linear optimization , author=. 1997 , publisher=

  142. [150]

    European Journal of Operational Research , volume=

    Sensitivity analysis in linear programming: just be careful! , author=. European Journal of Operational Research , volume=. 1997 , publisher=

  143. [151]

    Scientific Reports , volume=

    Precision game engineering through reshaping strategic payoffs , author=. Scientific Reports , volume=. 2024 , publisher=

  144. [152]

    Handbook of social Choice and Welfare , volume=

    Implementation theory , author=. Handbook of social Choice and Welfare , volume=. 2002 , publisher=

  145. [153]

    Econometrica: Journal of the Econometric Society , pages=

    Nash implementation: a full characterization , author=. Econometrica: Journal of the Econometric Society , pages=. 1990 , publisher=

  146. [154]

    arXiv preprint arXiv:2402.09695 , year=

    Reward poisoning attack against offline reinforcement learning , author=. arXiv preprint arXiv:2402.09695 , year=

  147. [155]

    Proceedings of the AAAI conference on artificial intelligence , volume=

    Optimal attack and defense for reinforcement learning , author=. Proceedings of the AAAI conference on artificial intelligence , volume=

  148. [156]

    Advances in Neural Information Processing Systems , volume=

    Efficient adversarial attacks on online multi-agent reinforcement learning , author=. Advances in Neural Information Processing Systems , volume=

  149. [157]

    arXiv preprint arXiv:2503.03676 , year=

    Optimally Installing Strict Equilibria , author=. arXiv preprint arXiv:2503.03676 , year=

  150. [158]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    Data poisoning to fake a nash equilibria for markov games , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.