Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:04.168498Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.13862.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:04.168498Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 86109e89-1663-4981-a269-60bc7099eba2 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent single-task RL
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7264f06-876d-4bf3-ac89-6f1a90f7d4ea · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf13b061-6f02-4338-947f-3b3082204d4b · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7014e099-0c24-42b9-a2c3-688f78ce2d9a · outbound
StaQ it! Growing neural networks for Policy Mirror Descent MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13004cb4-2553-452c-87f0-0eb8c1412822 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8bb50f7b-3cd9-4073-b96f-4cbeae8614ef · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c052211d-e640-432d-8c14-0d3c16cdc075 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent For PQN, we use the CleanRL implementation (Huang et al., 2022)
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2977c183-1317-43ac-b637-925c2a76c054 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 604c68c7-509d-41e0-a758-99730384f708 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent In the MinAtar environ- ments α is linearly annealed from 1 to 0 over the course of learning.∗Humanoid-v4 uses a hidden layer size of256
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e27c280b-70f5-4b1e-bb4a-d2578ef42c6b · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81bef3d0-d1a3-4e91-8a59-3480fe621da4 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Classic and MinAtar hyperparameters are based on the original paper (Gallici et al., 2025), while MuJoCo hyperparameters were found by hyperparameter tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b0f0d1f-d8c6-4ac8-bb09-b376ac207a24 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent doi: 10.1016/S0167-6377(02)00231-6
Reference 2003
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89fea3c6-8f12-4a8f-99a4-1d118f9ff6c4 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Mirror Descent Policy Optimization
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a11d5ff7-9cf5-4c62-91ea-fcb73de79e63 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent We can see in Fig
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82fdb66c-f27d-4497-abf3-de8584454647 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb783b25-fc5b-4a52-b22e-b3a4fd72315a · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Homotopic Policy Mirror Descent: Policy Convergence, Implicit Regularization, and Improved Sample Complexity
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 726edaa6-805d-4d80-9f87-9537bbeb5c61 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f1bc58c-a35c-4fa4-a565-7dc90dcc5f57 · outbound
StaQ it! Growing neural networks for Policy Mirror Descent Maxmin Q-learning: Controlling the Estimation Bias of Q-learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7c1b16f-eca0-4fbb-9424-7b73cac187dd · outbound
StaQ it! Growing neural networks for Policy Mirror Descent van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.