Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T08:32:51.215581Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2606.23995.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T08:32:51.215581Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-08T03:17:09.608451Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T03:24:28.854619Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 984d085a-1262-499d-a632-3a30937199c4 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Dota 2 with Large Scale Deep Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f62fbe3c-7e1d-480e-a2e2-28a8f5d9c5c3 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee2af48-9dc9-4e27-8063-3e7d41dbf3b4 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5598f9b3-fb33-498a-ba01-cffa5466558b · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Combining deep reinforce- ment learning and search for imperfect-information games
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a539a98-6405-4cb0-b436-c66614d458f8 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Enhancing robustness in multi-agent reinforcement learn- ing via temporal consistency regularization: A self-distillation framework
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2d4df2-29be-4222-84ef-03cbd18daf6c · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games V ortices instead of equilibria in minmax opti- mization: Chaos and butterfly effects of online learning in zero-sum games
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b54c9c-b277-48dd-812e-3c0f54a8c452 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Deep reinforcement learning from self-play in imperfect- information games, 2016
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea49ad59-abab-4821-8123-40b377d80bf7 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Neural replicator dynamics: Multiagent learning via hedging policy gradients
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce71ab4-aa94-4e53-b469-062069b599a4 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Averaging Weights Leads to Wider Optima and Better Generalization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 301de6d5-1f19-4935-bd61-2a427ca04bc8 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games A unified game-theoretic approach to multiagent reinforcement learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d3373b7-8fe3-4760-8118-942df6d1bcf8 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games OpenSpiel: A Framework for Reinforcement Learning in Games
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec34513a-70b3-4db5-a6bd-aa510e8f7f44 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Data-augmented game starts for accelerating self-play exploration in imperfect information games
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 461870a2-4fd7-48ce-9bed-c2f10d417315 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise Networks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 237ec56f-758d-4d71-8213-81ffd9e49069 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Continuous control with deep reinforcement learning, September 15 2020
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daed2e5f-c9a5-4c4c-bc67-e291dd74f5b6 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games NeuPL: Neural population learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1020c128-b3b9-4a23-a854-e486c4c0d4b9 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Pipeline PSRO: A scalable approach for finding approximate Nash equilibria in large games
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5327bcdd-9dc5-49a0-91d4-f7256feb7804 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Wang, Pierre Baldi, Tuomas Sandholm, and Roy Fox
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e105f52a-d209-41dc-a3e2-9e8d386505bc · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Wang, Pierre Baldi, and Roy Fox
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e19039d1-30dc-48c4-998b-18e544ecd2ec · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Escher: Eschewing importance sampling in games by computing a history value function to estimate regret
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae1f20ff-fce4-4c6b-9adb-6ddb1c058e79 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 538b7c0c-8c2c-46c6-80f6-01a288e4c52c · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Connor, Neil Burch, Thomas Anthony, Stephen McAleer, Romuald Elie, Sarah H
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35c0e572-c422-4a72-b107-562b57def5ca · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94b7a3e7-37f6-4453-b80a-fbf73b5fba81 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Zico Kolter, Amy Zhang, Gabriele Farina, Eugene Vinitsky, and Samuel Sokota
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2fcb582-38ac-4c9e-a4e4-4c0d53ff6c7a · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Proximal policy optimization algorithms, 2017
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42a77ce0-23c9-450e-b7ac-3ce761cc85a0 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games A unified approach to reinforcement learning, quantal response equilibria, and two-player zero-sum games
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3d1628-2832-473a-911c-195487ac3fcf · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Superhuman AI for Stratego using self-play reinforcement learning and test-time search
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ad0fd4fe-fb5a-4bec-9fca-66e0b3fa2095 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games DREAM: Deep regret minimization with advantage baselines and model-free learning, 2020
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0cf67c-5846-468a-bd2b-0f068121c1bc · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Czarnecki, et al
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 621f8a96-a157-4332-8427-d555b045817e · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d253e9da-bcdc-43d5-b1bb-5f4dbe6b08e9 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games arXiv preprint arXiv:2602.04417 , year=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fb321d39-4358-454c-92c7-8e25ac807575 · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a76422c-4821-41ba-9cdc-58687d905a5a · outbound
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games model soups
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 568a338a-ce4f-4bd7-b9da-fb4b4d52ed5d · inbound
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.