Pith. sign in

Paper Citation Record · LEDGER

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria

As of 10 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2510.18183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.18183 v3

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:58:33.585535Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T05:35:27.427721Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-21T05:39:40.969302Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1088073e-d548-4f06-9931-19bf43e7120b · outbound

This paper cites Oliehoek.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Oliehoek

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:28.629538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:28.629538Z digest=sha256:86929ea5c6806485e77d8da17b14abb0a07093ece6483c00cca55bd40bb0dc4d

Observation 5feadff1-0f2c-4d4b-b486-682ca9f3edf2 · outbound

This paper cites JAX: composable transfor- mations of Python+NumPy programs, 2018.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria JAX: composable transfor- mations of Python+NumPy programs, 2018

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:28.767199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:28.767199Z digest=sha256:bd8072b053ac495b7f552ffa2227c4919ff3d73d3026a4eae089d57be93c0a00

Observation afa7b3a7-1e18-46e9-84b2-db7d8811c84e · outbound

This paper cites Iterative solution of games by fictitious play.Act.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Iterative solution of games by fictitious play.Act

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:28.896870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:28.896870Z digest=sha256:15be460a052e3e9d46589fb0eb332133110686a99a98ee068747e461cbe702bc

Observation 41a573df-1c21-47ed-9bfa-b31e9d07894e · outbound

This paper cites Superhuman ai for heads-up no-limit poker: Libratus beats top profession- als.Science, 359(6374):418–424, 2018.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Superhuman ai for heads-up no-limit poker: Libratus beats top profession- als.Science, 359(6374):418–424, 2018

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:29.020148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:29.020148Z digest=sha256:0991e6ac3ef985658570d1bdfcf9ea869c01e4de2718e738a3b19bfa6e946854

Observation e47adac6-4aeb-4ff0-9528-65e7599374b8 · outbound

This paper cites Superhuman ai for multiplayer poker.Science, 365(6456):885–890, 2019.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Superhuman ai for multiplayer poker.Science, 365(6456):885–890, 2019

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:29.131292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:29.131292Z digest=sha256:deca952eaf0a5c1d6f378940339c21c0d92e8d1482638ae5178b6e0fb82ef0ea

Observation 53985372-219d-419a-b951-5e50b110b483 · outbound

This paper cites Deep counterfactual regret minimization.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Deep counterfactual regret minimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:29.250877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:29.250877Z digest=sha256:8fa73af150895c2f94ef6944fdc295e1b40ee6d04217712fcd0c1d4d218c4841

Observation 9c660a0f-b007-44a7-8269-8148aac7de85 · outbound

This paper cites A comprehensive survey of multiagent reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 38(2): 156–172, 2008.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A comprehensive survey of multiagent reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 38(2): 156–172, 2008

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:29.365021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:29.365021Z digest=sha256:7227d553ef9be44577f351091543c09ebe113909d134b5a70dc64d24d7d08361

Observation ce753de8-fec7-4ee7-bc74-7becb206f8b9 · outbound

This paper cites Fast policy extragradient methods for competitive games with entropy regularization.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Fast policy extragradient methods for competitive games with entropy regularization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:29.496706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:29.496706Z digest=sha256:2cd7ac80f1e077706d573cc9129de378e45a242ff9c8499f3004b16a699205de

Observation d78ab0f0-543d-4421-ae16-e9282366fa44 · outbound

This paper cites Last-iterate convergence: Zero-sum games and constrained min- max optimization.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Last-iterate convergence: Zero-sum games and constrained min- max optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:29.658080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:29.658080Z digest=sha256:151e4a35089db931a7d64de2ddd3508e6a0f073b0f4d51c62fe9cab0714d152c

Observation 31d7292d-e306-4c18-9ebf-315c6a6270e5 · outbound

This paper cites Training gans with optimism.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Training gans with optimism

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:29.815653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:29.815653Z digest=sha256:aeab78d07d5ea2f4e920744290abe7ebe325186265920dc3ce4eb3fd893e8b14

Observation ff6dd9e0-7c05-4dac-b29c-6515361a597d · outbound

This paper cites Springer, 2003.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Springer, 2003

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:29.952558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:29.952558Z digest=sha256:e964903f3177f8acb226fa66bbfd65314d4c2471055de1f4785dbb22dcc62243

Observation 60ef99f7-ca8a-4ed4-b836-bb3902e2d54c · outbound

This paper cites Solving for best responses in extensive-form games using reinforcement learning methods.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Solving for best responses in extensive-form games using reinforcement learning methods

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:30.076837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:30.076837Z digest=sha256:d98bbb7081d1bd290131ee3f1ebdd55aad2a91246fb109b89ca69109bef5d4c5

Observation 4975bc87-7f7f-4e1d-a9fe-a659391fe349 · outbound

This paper cites Deep Reinforcement Learning from Self-Play in Imperfect-Information Games.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Deep Reinforcement Learning from Self-Play in Imperfect-Information Games

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:30.196277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:30.196277Z digest=sha256:e177b68543cbda9e5cc19ef11a69977bea0c4e4825bbef39a9aacbc34ad39cca

Observation 859c70e3-709f-462b-bbd5-7b3fc6db3baf · outbound

This paper cites Learning in perturbed asymmetric games.Games Econ.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Learning in perturbed asymmetric games.Games Econ

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:30.366974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:30.366974Z digest=sha256:b9ce814ac00c4b944d2474d4583fa600160388c472bb3f377d45f437e07cf0a6

Observation ec4c67cc-b910-4747-878e-d0b27864dccc · outbound

This paper cites Adaptive learning in continuous games: Optimal regret bounds and convergence to nash equilibrium.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Adaptive learning in continuous games: Optimal regret bounds and convergence to nash equilibrium

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:30.507732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:30.507732Z digest=sha256:c624a1fac633eaa40f7183cb710644173e36bcabf0950a99c11743e4272bd95e

Observation 4c26b87a-0a5d-4630-959a-c45f8b8e989b · outbound

This paper cites Extensive games and the problem of information.Contributions to the Theory of Games, 2(28): 193–216, 1953.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Extensive games and the problem of information.Contributions to the Theory of Games, 2(28): 193–216, 1953

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:30.623234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:30.623234Z digest=sha256:1c8db32fc899c85c43ba33d454d25b5dc85e582eec495950a64bfb2a6b247122

Observation 0bc8ca25-92d9-4bd8-aa23-5ce5a533b33d · outbound

This paper cites A unified game-theoretic approach to multiagent reinforcement learning.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A unified game-theoretic approach to multiagent reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:30.740485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:30.740485Z digest=sha256:30774641ee1b8157842f9d3ead00c635f2805f6331b2f491c12227c9e7a3fc84

Observation 11abbae1-60db-444a-9c05-9e5f8785eb4a · outbound

This paper cites A class of gap functions for variational inequalities.Mathematical Programming, 64(1):53–79, 1994.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A class of gap functions for variational inequalities.Mathematical Programming, 64(1):53–79, 1994

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:31.095675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:31.095675Z digest=sha256:b2d094c51e99841a0ffd70ba8853d994909099c4e36e5829348b881be57f45b9

Observation df1daf7d-e64e-4663-b3b4-3e275a278838 · outbound

This paper cites Last-iterate convergence in extensive-form games.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Last-iterate convergence in extensive-form games

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:31.247211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:31.247211Z digest=sha256:c53db6e6f253e16d665a376bc869cd9ff645b9cc21b040541972907f9db98af5

Observation 38948cf2-175e-4488-8087-3d45ada886f8 · outbound

This paper cites Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:31.407152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:31.407152Z digest=sha256:c23c8e3aa1d607e2f54f10d8e03b93752f583b5047373c4e9b3a465915974cfe

Observation 8ed87ff8-aa5a-49a6-b185-c31f55e8f2bf · outbound

This paper cites A survey of nash equilibrium strategy solving based on cfr.Archives of Computational Methods in Engineering, 28(4), 2021.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A survey of nash equilibrium strategy solving based on cfr.Archives of Computational Methods in Engineering, 28(4), 2021

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:31.542497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:31.542497Z digest=sha256:e9712c40e2d51e878ac8cc5333cb6f4280ab0b1f668a418e941ec4281d99f303

Observation c8aff3e5-001b-458d-ae7a-f16c007ace6e · outbound

This paper cites Ozdaglar, Tiancheng Yu, and Kaiqing Zhang.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Ozdaglar, Tiancheng Yu, and Kaiqing Zhang

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:31.642192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:31.642192Z digest=sha256:2080ed6babb48f51e883091881b5f3589780db14b40287303953ab9751e0c4d0

Observation bff831ab-f6e3-44af-9079-190a12f4c9c9 · outbound

This paper cites Lanier, Roy Fox, and Pierre Baldi.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Lanier, Roy Fox, and Pierre Baldi

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:31.795089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:31.795089Z digest=sha256:75865a8e079f044f7c316b3a8cc7e04f4e97d36d6cc3ae7537d19802267f6ada

Observation d6df8c92-e7a7-45c8-a867-62924aee2d43 · outbound

This paper cites ESCHER: eschewing impor- tance sampling in games by computing a history value function to estimate regret.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria ESCHER: eschewing impor- tance sampling in games by computing a history value function to estimate regret

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:31.910301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:31.910301Z digest=sha256:f66a8f69978bc7fd5f4a15b0f31ce82cc847986d8cbeb1155e74b13e408df585

Observation aa134598-acf0-4880-9fce-e5827b7c4f5b · outbound

This paper cites Wang, Pierre Baldi, Tuomas Sandholm, and Roy Fox.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Wang, Pierre Baldi, Tuomas Sandholm, and Roy Fox

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:32.047562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:32.047562Z digest=sha256:9c9ac79dd359c45c220e74445bc720226c468fa79cd7725007329a33c644975d

Observation fbb710e8-16e1-42ab-9edd-21c7129bc98d · outbound

This paper cites McKelvey and Thomas R.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria McKelvey and Thomas R

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:32.194779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:32.194779Z digest=sha256:bb20bcf6d13d5938d2b8aa7a4c372881131f0d80fb64109c621d6ed0699ab4b6

Observation 96b8ebf1-617d-4cd4-8ebe-bfd287d45af3 · outbound

This paper cites Ortega, Neil Burch, Thomas W.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Ortega, Neil Burch, Thomas W

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:32.332116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:32.332116Z digest=sha256:707d3c46238f79d84492ce67b03fcfcd51ace719683c238822e301dd7f6e6ffa

Observation b215b5f6-1445-4ce5-82e8-87c9200e7c63 · outbound

This paper cites Reevaluating Policy Gradient Methods for Imperfect-Information Games.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Reevaluating Policy Gradient Methods for Imperfect-Information Games

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:32.453363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:32.453363Z digest=sha256:ba11c7ffcd81ad6960c58c5d38b0e79ab6d56db67cc63d80b30e2ffc6934f802

Observation 48308ba7-0f88-4847-861b-4ace5d1a8776 · outbound

This paper cites Proximal Policy Optimization Algorithms.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:32.616577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:32.616577Z digest=sha256:e179639d14ffcb1770698f315f919e6c62a34ca07e61485c4402334ee961b94e

Observation 2bdaee0d-e23a-4ec4-9d6b-4f6675424cca · outbound

This paper cites If multi-agent learning is the answer, what is the question? Artificial intelligence, 171(7):365–377, 2007.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria If multi-agent learning is the answer, what is the question? Artificial intelligence, 171(7):365–377, 2007

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:32.782011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:32.782011Z digest=sha256:57083695cbc45539250301b28b65b233a362096ffd06410e9fc622102379c7d5

Observation 1386e903-1dae-4db7-a854-17a227d72b91 · outbound

This paper cites an unresolved cited work.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:32.896649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:32.896649Z digest=sha256:1f469bb9eef059598a407bb71817c157090f8a1f922c47a6efca78c58b64a142

Observation 7834897e-f1c0-4fdd-bce2-35353c5d097f · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362 (6419):1140–1144, 2018.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362 (6419):1140–1144, 2018

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:33.066122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:33.066122Z digest=sha256:d5848f51e5099926fce70731d8d312a469be22223696b5ac623d45f25e615a52

Observation c3a4f671-454f-4744-a567-844d8c41fbf4 · outbound

This paper cites Zico Kolter, Nicolas Loizou, Marc Lanctot, Ioannis Mitliagkas, Noam Brown, and Christian Kroer.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Zico Kolter, Nicolas Loizou, Marc Lanctot, Ioannis Mitliagkas, Noam Brown, and Christian Kroer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:33.183853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:33.183853Z digest=sha256:03b3c2af31b92bb9dd007429ee2a92e63f777649a854e6d348769b071a1376a9

Observation 47174499-4607-4911-b25f-e7df15edd7d8 · outbound

This paper cites DREAM: Deep Regret minimization with Advantage baselines and Model-free learning.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria DREAM: Deep Regret minimization with Advantage baselines and Model-free learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:33.334458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:33.334458Z digest=sha256:31d977aa40d9687b9d0bd053b1b06a4365db67b352f0a9956e0dbfd92964ddf8

Observation a44373cb-2979-4fab-a3df-2ddd9ad8e83c · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:33.510525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:33.510525Z digest=sha256:377b4fb44a72baa5035547b62891d6c0a5092532b0fbb8e35f658df9d0940f60

Observation 7f3c8b1f-34cd-47b2-9efa-4e0937f01be4 · outbound

This paper cites Bowling, and Carmelo Piccione.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Bowling, and Carmelo Piccione

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:33.585535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:33.585535Z digest=sha256:9705842152ee22bbfad5a3a9de17cfdbe00574da372d3b6b2289956ffc05e781

Observation ff5553eb-4cf0-4138-917c-f9d3bc75472a · outbound

This paper cites OpenSpiel: A Framework for Reinforcement Learning in Games.

NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria OpenSpiel: A Framework for Reinforcement Learning in Games

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T08:58:30.975937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:58:30.975937Z digest=sha256:2f4de0000a68b516cc2eefa327e4b740085fab7ebd17a019856756b76bb0268c

Pith citing papers

Observation 69a90323-31b8-4649-a4df-2ea4fed6bc62 · inbound

Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX cites this paper.

Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:39:40.970877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T05:35:27.427721Z digest=sha256:774906c8a09bc10149f54bde4ff09541afd7817114f86a793c8a82925c6a9e7c