Pith. sign in

Paper Citation Record · LEDGER

Training Larger Networks for Deep Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2102.07920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.07920 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:38:32.950750Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:13:58.812654Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f09bf7b5-5e32-4408-bdd7-9f6b7ea461d6 · inbound

Improving Vision-Language-Action Model with Online Reinforcement Learning cites this paper.

Improving Vision-Language-Action Model with Online Reinforcement Learning Training Larger Networks for Deep Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T11:39:56.097738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:39:56.097738Z digest=sha256:0bc0aea41f592d863868690e71c758717e4a316474942b0a6b27707c965998b4

Observation 8ead2cf2-0c21-4434-b27b-eabe73a84501 · inbound

Exploring the robustness of TractOracle methods in RL-based tractography cites this paper.

Exploring the robustness of TractOracle methods in RL-based tractography Training Larger Networks for Deep Reinforcement Learning

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:13:58.870852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:13:58.268289Z digest=sha256:59b57883fbc5fa545d3672132df07b25f7cf5c3f5c29b36cb94868b4ec1073ed

Observation bb4c30b8-c5e7-4989-9635-bf16f2ea5a52 · inbound

Population-aware Online Mirror Descent for Mean-Field Games with Common Noise by Deep Reinforcement Learning cites this paper.

Population-aware Online Mirror Descent for Mean-Field Games with Common Noise by Deep Reinforcement Learning Training Larger Networks for Deep Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:38:32.950750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:38:32.950750Z digest=sha256:99fc30d3a13a0cb8d290de4e4fc5624ae680f48041d4be7d9fffb4337e3d0176

Observation 4b20b069-19e4-4ba2-8f6e-040ab9f682c0 · inbound

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments cites this paper.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Training Larger Networks for Deep Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.108589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.108589Z digest=sha256:07e2f1e0b1a8758cff64557bb8f154018fdd6e8cf292f07ea97533bd00915a5b

Observation 82132760-d177-4b33-99fa-520321209624 · inbound

Learning Reach-Avoid Task with Reinforcement Learning: Vectorized Simulation and Benchmark cites this paper.

Learning Reach-Avoid Task with Reinforcement Learning: Vectorized Simulation and Benchmark Training Larger Networks for Deep Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:54:36.714784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:54:36.714784Z digest=sha256:222ddb969f1fd3e60b1e00b86a925f2582adb71a3c54a33326a622189661b4ee