Pith. sign in

Paper Citation Record · LEDGER

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

As of 6 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2606.23995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.23995 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T08:32:51.215581Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T03:17:09.608451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T03:24:28.854619Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 984d085a-1262-499d-a632-3a30937199c4 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Dota 2 with Large Scale Deep Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:49:45.627271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:e767d997c86a3932afa3e749f2844583c970fedde0433310519db6fbd27e58c8

Observation f62fbe3c-7e1d-480e-a2e2-28a8f5d9c5c3 · outbound

This paper cites an unresolved cited work.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:73958de0f5632ec737cfccafc4328a18cc9297b479a9faece332030c55121828

Observation 3ee2af48-9dc9-4e27-8063-3e7d41dbf3b4 · outbound

This paper cites Superhuman AI for heads-up no-limit poker: Libratus beats top professionals.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Superhuman AI for heads-up no-limit poker: Libratus beats top professionals

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:9e4ee19df5b0abffdd56bfba214cf1e3a3c99d0cb003a298df75cad56eacbd63

Observation 5598f9b3-fb33-498a-ba01-cffa5466558b · outbound

This paper cites Combining deep reinforce- ment learning and search for imperfect-information games.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Combining deep reinforce- ment learning and search for imperfect-information games

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:e1416dab0fc1b04d8f403ff046d7d35f92f841a3550ac1180f75072aa43f852d

Observation 1a539a98-6405-4cb0-b436-c66614d458f8 · outbound

This paper cites Enhancing robustness in multi-agent reinforcement learn- ing via temporal consistency regularization: A self-distillation framework.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Enhancing robustness in multi-agent reinforcement learn- ing via temporal consistency regularization: A self-distillation framework

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:ff0ea6d86a2c545b28c4fc42ac489624433c3ef714439e6ad22b0c1e6448aa88

Observation 3b2d4df2-29be-4222-84ef-03cbd18daf6c · outbound

This paper cites V ortices instead of equilibria in minmax opti- mization: Chaos and butterfly effects of online learning in zero-sum games.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games V ortices instead of equilibria in minmax opti- mization: Chaos and butterfly effects of online learning in zero-sum games

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:cf1db2632663ef52d385106a6f2ee79ce02f98110c5e0d959dc978878702fb4c

Observation 20b54c9c-b277-48dd-812e-3c0f54a8c452 · outbound

This paper cites Deep reinforcement learning from self-play in imperfect- information games, 2016.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Deep reinforcement learning from self-play in imperfect- information games, 2016

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:4bfc9f5e77aa536139669f51e4ff1a7880574f7698ca8aca82bb7faa5521fc9b

Observation ea49ad59-abab-4821-8123-40b377d80bf7 · outbound

This paper cites Neural replicator dynamics: Multiagent learning via hedging policy gradients.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Neural replicator dynamics: Multiagent learning via hedging policy gradients

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:83537d132b86f4970ca3731ea095da2cd599066ab5c90f0210cc7409188de221

Observation 1ce71ab4-aa94-4e53-b469-062069b599a4 · outbound

This paper cites Averaging Weights Leads to Wider Optima and Better Generalization.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Averaging Weights Leads to Wider Optima and Better Generalization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:49:45.624333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:7a042d44b203bcc6cf1ee75bfea0b8d8f441a9ac32abab44b931cdd5a1511093

Observation 301de6d5-1f19-4935-bd61-2a427ca04bc8 · outbound

This paper cites A unified game-theoretic approach to multiagent reinforcement learning.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games A unified game-theoretic approach to multiagent reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:9a3580ddc07d8517b4d59fdc738f590ebd65695d40106e248c204075b7949145

Observation 6d3373b7-8fe3-4760-8118-942df6d1bcf8 · outbound

This paper cites OpenSpiel: A Framework for Reinforcement Learning in Games.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games OpenSpiel: A Framework for Reinforcement Learning in Games

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:45.645442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:79cbede4ba426504256732582cd1eb2eb87a12f922d0d93c2d83f6f9413e6f11

Observation ec34513a-70b3-4db5-a6bd-aa510e8f7f44 · outbound

This paper cites Data-augmented game starts for accelerating self-play exploration in imperfect information games.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Data-augmented game starts for accelerating self-play exploration in imperfect information games

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:b45f9a4ce8d9c5b04fa7b5b6c96d02537df8985e35c0be1376ff87d40b578f08

Observation 461870a2-4fd7-48ce-9bed-c2f10d417315 · outbound

This paper cites Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise Networks.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise Networks

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:45.644943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:20c868d8684f3c21da0581218224d317d255eb2939a1d9c70e96aaad28d27a90

Observation 237ec56f-758d-4d71-8213-81ffd9e49069 · outbound

This paper cites Continuous control with deep reinforcement learning, September 15 2020.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Continuous control with deep reinforcement learning, September 15 2020

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:1662d60409050ce6f48650c99577e7bdb4da5750d231753c464adee4682f3431

Observation daed2e5f-c9a5-4c4c-bc67-e291dd74f5b6 · outbound

This paper cites NeuPL: Neural population learning.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games NeuPL: Neural population learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:393c7a16210ba2003483a777cf92610eb6785c96dc9f9e09f7ccdb799ad08fc1

Observation 1020c128-b3b9-4a23-a854-e486c4c0d4b9 · outbound

This paper cites Pipeline PSRO: A scalable approach for finding approximate Nash equilibria in large games.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Pipeline PSRO: A scalable approach for finding approximate Nash equilibria in large games

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:c4b0dde7f6f36e1f9530b50f02ded3d7cf447abb5226cf765d2add3bd4d6560f

Observation 5327bcdd-9dc5-49a0-91d4-f7256feb7804 · outbound

This paper cites Wang, Pierre Baldi, Tuomas Sandholm, and Roy Fox.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Wang, Pierre Baldi, Tuomas Sandholm, and Roy Fox

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:59b1c26b0039d39d33559803490adf48519d3bc8bc53ac9712bfe2d26f4562f2

Observation e105f52a-d209-41dc-a3e2-9e8d386505bc · outbound

This paper cites Wang, Pierre Baldi, and Roy Fox.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Wang, Pierre Baldi, and Roy Fox

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:67f5b8070b944a0f9fed00dcc67fe09fe83be996fece2db7de0a559bcb566bd0

Observation e19039d1-30dc-48c4-998b-18e544ecd2ec · outbound

This paper cites Escher: Eschewing importance sampling in games by computing a history value function to estimate regret.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Escher: Eschewing importance sampling in games by computing a history value function to estimate regret

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:5f9ecb24603cde0d222c667be502d3382262836ed859f9852ac56f527215e249

Observation ae1f20ff-fce4-4c6b-9adb-6ddb1c058e79 · outbound

This paper cites Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:45.639309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:09deb537b14323bcca51dbcb34afb5ea11986f4a39c75b1a6af5233512ce4717

Observation 538b7c0c-8c2c-46c6-80f6-01a288e4c52c · outbound

This paper cites Connor, Neil Burch, Thomas Anthony, Stephen McAleer, Romuald Elie, Sarah H.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Connor, Neil Burch, Thomas Anthony, Stephen McAleer, Romuald Elie, Sarah H

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:a2fe0e7fc9005ae20ee8673ee8f98a4f9cfc8e10e583a2d26c8bff82f742c4a3

Observation 35c0e572-c422-4a72-b107-562b57def5ca · outbound

This paper cites WARP: On the Benefits of Weight Averaged Rewarded Policies.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:45.642601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:dd749e339d5702eed1d6c7fe9d651e41d7d270ef333452ff0785cfc328d4b3dc

Observation 94b7a3e7-37f6-4453-b80a-fbf73b5fba81 · outbound

This paper cites Zico Kolter, Amy Zhang, Gabriele Farina, Eugene Vinitsky, and Samuel Sokota.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Zico Kolter, Amy Zhang, Gabriele Farina, Eugene Vinitsky, and Samuel Sokota

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:8fcb533355b986489b00538d31f7e125c30d7343594e32cbd9f2a8218cdb83d4

Observation a2fcb582-38ac-4c9e-a4e4-4c0d53ff6c7a · outbound

This paper cites Proximal policy optimization algorithms, 2017.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Proximal policy optimization algorithms, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:29df6ddfcb016ba20e8592866687fad11d5f4c6d5c843b33c32e500516015644

Observation 42a77ce0-23c9-450e-b7ac-3ce761cc85a0 · outbound

This paper cites A unified approach to reinforcement learning, quantal response equilibria, and two-player zero-sum games.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games A unified approach to reinforcement learning, quantal response equilibria, and two-player zero-sum games

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:cf61907aeb73e55a9c3a46837234784b17223ef78485c6a89d39cdd6b332af49

Observation 5f3d1628-2832-473a-911c-195487ac3fcf · outbound

This paper cites Superhuman AI for Stratego using self-play reinforcement learning and test-time search.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Superhuman AI for Stratego using self-play reinforcement learning and test-time search

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:45.634952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:876cbbea43eacffa7b281c01339a41f4f53f01f02ac8858e6cbb0237ec355a71

Observation ad0fd4fe-fb5a-4bec-9fca-66e0b3fa2095 · outbound

This paper cites DREAM: Deep regret minimization with advantage baselines and model-free learning, 2020.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games DREAM: Deep regret minimization with advantage baselines and model-free learning, 2020

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:b49a276bec503cfb58271082db9ab218443deda408f089146867c5cb50a339ba

Observation ef0cf67c-5846-468a-bd2b-0f068121c1bc · outbound

This paper cites Czarnecki, et al.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Czarnecki, et al

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:04109a2a52a0b59612070e8ce53d3f9c5bd0915b36e3577d2710a9a446d8c9f8

Observation 621f8a96-a157-4332-8427-d555b045817e · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:bdfbe103b1f7f9624ea98ab4a2763f68833a8622696c08ef60ac3de6d333922a

Observation d253e9da-bcdc-43d5-b1bb-5f4dbe6b08e9 · outbound

This paper cites arXiv preprint arXiv:2602.04417 , year=.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games arXiv preprint arXiv:2602.04417 , year=

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:45.636250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:4a9ee7d9c0b9afbc5a480d8e7f4770750319d4db56ec7fc3cb5f185b90083ccc

Observation fb321d39-4358-454c-92c7-8e25ac807575 · outbound

This paper cites an unresolved cited work.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:ddf3579ab9fbeb6e5b6f22b8263db9ecf367ab93f617e2060189cf2c1afe7e8c

Observation 8a76422c-4821-41ba-9cdc-58687d905a5a · outbound

This paper cites model soups.

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games model soups

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T08:32:51.215581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:32:51.215581Z digest=sha256:eb35bc0e2d4d5e05097709d8239d5344ba37340251084cca65696a411b46c272

Pith citing papers

Observation 568a338a-ce4f-4bd7-b9da-fb4b4d52ed5d · inbound

FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games cites this paper.

FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T03:24:28.855837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-08T03:17:09.608451Z digest=sha256:d209876c6b9f0b18559d34ad4508737802f789bf094d229eaf093e288a656f4c