Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:58:33.585535Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2510.18183.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:58:33.585535Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T05:35:27.427721Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-21T05:39:40.969302Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1088073e-d548-4f06-9931-19bf43e7120b · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Oliehoek
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5feadff1-0f2c-4d4b-b486-682ca9f3edf2 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria JAX: composable transfor- mations of Python+NumPy programs, 2018
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa7b3a7-1e18-46e9-84b2-db7d8811c84e · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Iterative solution of games by fictitious play.Act
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a573df-1c21-47ed-9bfa-b31e9d07894e · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Superhuman ai for heads-up no-limit poker: Libratus beats top profession- als.Science, 359(6374):418–424, 2018
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e47adac6-4aeb-4ff0-9528-65e7599374b8 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Superhuman ai for multiplayer poker.Science, 365(6456):885–890, 2019
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53985372-219d-419a-b951-5e50b110b483 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Deep counterfactual regret minimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c660a0f-b007-44a7-8269-8148aac7de85 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A comprehensive survey of multiagent reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 38(2): 156–172, 2008
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce753de8-fec7-4ee7-bc74-7becb206f8b9 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Fast policy extragradient methods for competitive games with entropy regularization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d78ab0f0-543d-4421-ae16-e9282366fa44 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Last-iterate convergence: Zero-sum games and constrained min- max optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d7292d-e306-4c18-9ebf-315c6a6270e5 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Training gans with optimism
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6dd9e0-7c05-4dac-b29c-6515361a597d · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Springer, 2003
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60ef99f7-ca8a-4ed4-b836-bb3902e2d54c · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Solving for best responses in extensive-form games using reinforcement learning methods
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4975bc87-7f7f-4e1d-a9fe-a659391fe349 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 859c70e3-709f-462b-bbd5-7b3fc6db3baf · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Learning in perturbed asymmetric games.Games Econ
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec4c67cc-b910-4747-878e-d0b27864dccc · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Adaptive learning in continuous games: Optimal regret bounds and convergence to nash equilibrium
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c26b87a-0a5d-4630-959a-c45f8b8e989b · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Extensive games and the problem of information.Contributions to the Theory of Games, 2(28): 193–216, 1953
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bc8ca25-92d9-4bd8-aa23-5ce5a533b33d · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A unified game-theoretic approach to multiagent reinforcement learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11abbae1-60db-444a-9c05-9e5f8785eb4a · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A class of gap functions for variational inequalities.Mathematical Programming, 64(1):53–79, 1994
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df1daf7d-e64e-4663-b3b4-3e275a278838 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Last-iterate convergence in extensive-form games
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38948cf2-175e-4488-8087-3d45ada886f8 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed87ff8-aa5a-49a6-b185-c31f55e8f2bf · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A survey of nash equilibrium strategy solving based on cfr.Archives of Computational Methods in Engineering, 28(4), 2021
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8aff3e5-001b-458d-ae7a-f16c007ace6e · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Ozdaglar, Tiancheng Yu, and Kaiqing Zhang
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bff831ab-f6e3-44af-9079-190a12f4c9c9 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Lanier, Roy Fox, and Pierre Baldi
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6df8c92-e7a7-45c8-a867-62924aee2d43 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria ESCHER: eschewing impor- tance sampling in games by computing a history value function to estimate regret
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa134598-acf0-4880-9fce-e5827b7c4f5b · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Wang, Pierre Baldi, Tuomas Sandholm, and Roy Fox
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbb710e8-16e1-42ab-9edd-21c7129bc98d · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria McKelvey and Thomas R
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b8ebf1-617d-4cd4-8ebe-bfd287d45af3 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Ortega, Neil Burch, Thomas W
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b215b5f6-1445-4ce5-82e8-87c9200e7c63 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Reevaluating Policy Gradient Methods for Imperfect-Information Games
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48308ba7-0f88-4847-861b-4ace5d1a8776 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Proximal Policy Optimization Algorithms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bdaee0d-e23a-4ec4-9d6b-4f6675424cca · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria If multi-agent learning is the answer, what is the question? Artificial intelligence, 171(7):365–377, 2007
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1386e903-1dae-4db7-a854-17a227d72b91 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7834897e-f1c0-4fdd-bce2-35353c5d097f · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362 (6419):1140–1144, 2018
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3a4f671-454f-4744-a567-844d8c41fbf4 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Zico Kolter, Nicolas Loizou, Marc Lanctot, Ioannis Mitliagkas, Noam Brown, and Christian Kroer
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47174499-4607-4911-b25f-e7df15edd7d8 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria DREAM: Deep Regret minimization with Advantage baselines and Model-free learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44373cb-2979-4fab-a3df-2ddd9ad8e83c · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f3c8b1f-34cd-47b2-9efa-4e0937f01be4 · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Bowling, and Carmelo Piccione
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5553eb-4cf0-4138-917c-f9d3bc75472a · outbound
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria OpenSpiel: A Framework for Reinforcement Learning in Games
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69a90323-31b8-4649-a4df-2ea4fed6bc62 · inbound
Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.