Pith. sign in

Paper Citation Record · LEDGER

StaQ it! Growing neural networks for Policy Mirror Descent

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.13862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13862 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:04.168498Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86109e89-1663-4981-a269-60bc7099eba2 · outbound

This paper cites single-task RL.

StaQ it! Growing neural networks for Policy Mirror Descent single-task RL

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.503389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:03.294608Z digest=sha256:069bfed05e4c54a5e6f7e9192c49a8802d4e349c2d514a24e3d092ac4636cd28

Observation f7264f06-876d-4bf3-ac89-6f1a90f7d4ea · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:07.288593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:02.894245Z digest=sha256:40559fbc32a93152b1c9fd64a38c83caec3596695e5bfc75039bf4f3ae2aca1c

Observation bf13b061-6f02-4338-947f-3b3082204d4b · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:06.730051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:03.161264Z digest=sha256:a94805dad53c9d1f5b200541a737a72e7c61eff053d7c70dc71a954528d8e816

Observation 7014e099-0c24-42b9-a2c3-688f78ce2d9a · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

StaQ it! Growing neural networks for Policy Mirror Descent MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.491366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.491366Z digest=sha256:b36b802629cfc927f8b5adfaac0b561842b9746088b0c75e46cf389961d1749d

Observation 13004cb4-2553-452c-87f0-0eb8c1412822 · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:07.545576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:02.732844Z digest=sha256:54e4d16ed23373bb16408701df5d8089ce68053b5b312551d920ad26338dca1f

Observation 8bb50f7b-3cd9-4073-b96f-4cbeae8614ef · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:06.224679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:03.429724Z digest=sha256:acd3066c8d0d69d9479153434dc11f3adddcb63acbc50304a7649d439c7982ac

Observation c052211d-e640-432d-8c14-0d3c16cdc075 · outbound

This paper cites For PQN, we use the CleanRL implementation (Huang et al., 2022).

StaQ it! Growing neural networks for Policy Mirror Descent For PQN, we use the CleanRL implementation (Huang et al., 2022)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.917222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:03.537326Z digest=sha256:3eb559785b3db5a8e511d6fa32c4f4a304f1ebe1f19ad2b9d8527d2cb78f51f2

Observation 2977c183-1317-43ac-b637-925c2a76c054 · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:05.694862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:03.697384Z digest=sha256:81b733890dcbf06a87a89bfe56ad6e60c2fda64fd26e4696d29359d4e3ed3c98

Observation 604c68c7-509d-41e0-a758-99730384f708 · outbound

This paper cites In the MinAtar environ- ments α is linearly annealed from 1 to 0 over the course of learning.∗Humanoid-v4 uses a hidden layer size of256.

StaQ it! Growing neural networks for Policy Mirror Descent In the MinAtar environ- ments α is linearly annealed from 1 to 0 over the course of learning.∗Humanoid-v4 uses a hidden layer size of256

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.424094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:03.814872Z digest=sha256:7ea8c89e252213c5028b9dbac37e300142eb35c42e5e2aa19002ddaa361d13da

Observation e27c280b-70f5-4b1e-bb4a-d2578ef42c6b · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 18

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:34:05.131796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:04.036926Z digest=sha256:3fa70c68c96ca78cb8249f516fb15ed99406957cd13e3229036a93be7494d100

Observation 81bef3d0-d1a3-4e91-8a59-3480fe621da4 · outbound

This paper cites Classic and MinAtar hyperparameters are based on the original paper (Gallici et al., 2025), while MuJoCo hyperparameters were found by hyperparameter tuning.

StaQ it! Growing neural networks for Policy Mirror Descent Classic and MinAtar hyperparameters are based on the original paper (Gallici et al., 2025), while MuJoCo hyperparameters were found by hyperparameter tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:04.937775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:04.168498Z digest=sha256:02222b7ccff31b4d347ad1622b581460bb434a9efce44c1557fd11fd05774bb4

Observation 0b0f0d1f-d8c6-4ac8-bb09-b376ac207a24 · outbound

This paper cites doi: 10.1016/S0167-6377(02)00231-6.

StaQ it! Growing neural networks for Policy Mirror Descent doi: 10.1016/S0167-6377(02)00231-6

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.737396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.737396Z digest=sha256:e788e6aa9fb2e887d8c9ccd4c49f777d3c98a1306865d51d3ada601148ab597a

Observation 89fea3c6-8f12-4a8f-99a4-1d118f9ff6c4 · outbound

This paper cites Mirror Descent Policy Optimization.

StaQ it! Growing neural networks for Policy Mirror Descent Mirror Descent Policy Optimization

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.274823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.274823Z digest=sha256:bdaa56624a9e1ce994c5c04ee490d36550031242f1c9d0d170199ae479cd45bd

Observation a11d5ff7-9cf5-4c62-91ea-fcb73de79e63 · outbound

This paper cites We can see in Fig.

StaQ it! Growing neural networks for Policy Mirror Descent We can see in Fig

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:07.019219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:34:03.027368Z digest=sha256:30d918e5a0ef2a7c9dbf72dfd351a5cdb326b2cb342ed24ef484e756518585ce

Observation 82fdb66c-f27d-4497-abf3-de8584454647 · outbound

This paper cites Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies.

StaQ it! Growing neural networks for Policy Mirror Descent Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.598175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.598175Z digest=sha256:6a9248823c329ce295eeea47b7fec5e705a1fa9529ff67e1a7f37b8e6d77c6ea

Observation eb783b25-fc5b-4a52-b22e-b3a4fd72315a · outbound

This paper cites Homotopic Policy Mirror Descent: Policy Convergence, Implicit Regularization, and Improved Sample Complexity.

StaQ it! Growing neural networks for Policy Mirror Descent Homotopic Policy Mirror Descent: Policy Convergence, Implicit Regularization, and Improved Sample Complexity

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.124836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.124836Z digest=sha256:0c552fa9721a5d879f23915f77e4ba77f9d7d8ce6f56fa5b51c195ddfbd750c2

Observation 726edaa6-805d-4d80-9f87-9537bbeb5c61 · outbound

This paper cites Randomized Ensembled Double Q-Learning: Learning Fast Without a Model.

StaQ it! Growing neural networks for Policy Mirror Descent Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.863748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.863748Z digest=sha256:84d013ad99d595321c36b20dfdb691a62022fdab9090a5d4c1a6f0cd0b4de595

Observation 6f1bc58c-a35c-4fa4-a565-7dc90dcc5f57 · outbound

This paper cites Maxmin Q-learning: Controlling the Estimation Bias of Q-learning.

StaQ it! Growing neural networks for Policy Mirror Descent Maxmin Q-learning: Controlling the Estimation Bias of Q-learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.965749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.965749Z digest=sha256:4243cd701d68601045cf9d25acc3863fc060b4a5903133bc0a7983b7693437d2

Observation b7c1b16f-eca0-4fbb-9424-7b73cac187dd · outbound

This paper cites van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J.

StaQ it! Growing neural networks for Policy Mirror Descent van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.354750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.354750Z digest=sha256:c1273217bb124a590473d9d910324bfe1a67d093f8c6c25676e930e2bbcd1078

Pith citing papers

No inbound Pith citation observations are available.