Pith. sign in

Paper Citation Record · LEDGER

Leveraging Procedural Generation to Benchmark Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:1912.01588.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.01588 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:59:40.090855Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

171
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 253eddd0-5bf8-4c84-9f40-fa56f9e8ab76 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:31.413770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:1b8560cacb8cb90692a4575572f5c5b7bf06323bd38d3c88a3eb3305b9dfc6c6

Observation 104aae9c-0c71-449a-9de8-c9c97c00416d · inbound

Proximal Policy Distillation cites this paper.

Proximal Policy Distillation Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T23:13:36.980164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T23:09:42.421333Z digest=sha256:7b90324962b083d7792cf2198960fcc8da04cb27a1d50d63bd4388f568e20658

Observation b42d011d-d382-4b29-ae7a-c14f333f4633 · inbound

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models cites this paper.

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:40.090855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:59:40.090855Z digest=sha256:da1c5fdc22385a6eca150fffee5608cdfae823d70d1394589c179f124aae41bb

Observation 2d7ab01c-84ce-482a-990a-51ff1f76c9ce · inbound

Dyn-O: Building Structured World Models with Object-Centric Representations cites this paper.

Dyn-O: Building Structured World Models with Object-Centric Representations Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:02.762041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:02.762041Z digest=sha256:b3006db7e11f32ded6e6342e9486759bad4f622d2a6dfdc381d6a8724b09b6e8

Observation 6d28784e-5ce8-4fe5-8a07-7470b597d00c · inbound

Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines cites this paper.

Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:42:07.614826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T06:38:33.969927Z digest=sha256:62895672454b1d2da9eadd4f7f0409ee0f745dc3b924b4403223e0b3fb1860cd

Observation be0497ab-3a38-4fa9-a49f-2eeb5dbd12c0 · inbound

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics cites this paper.

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:38:46.903783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:38:46.903783Z digest=sha256:d4e7eee956a8ee8661456464cbafca77f1a3d586b54ffdbdbe971ae775ed7c24

Observation e1284f15-6fcf-4e2e-89e0-e9ca4c5edb42 · inbound

Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains cites this paper.

Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-06T10:34:35.473241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:34:35.473241Z digest=sha256:6ab3155a841747701d40e1daa21572777ca97b5823415d8ddfaf306b0e3fab73

Observation 41bf3f85-9e3b-4c37-bee0-20063b01e320 · inbound

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning cites this paper.

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T05:21:35.522064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:21:35.522064Z digest=sha256:42a79483e0885e6fc22549167e9d42195211b9010cae211c403edd6629fc218f

Observation ff30be29-828b-4205-8356-f5a39941ec16 · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:20.671717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:256d672d6e68553ce762073ceaa02bf4d69c04b56c1b60ee12a78d7f5b324094

Observation 70afcc4e-51ae-4410-9c78-348c1d0cdaef · inbound

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization cites this paper.

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:36:29.552457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T09:55:00.402411Z digest=sha256:b6956672fc98ec2d7cef48272e0edf192cfe6f75dda74fd5e4d38689d96d99da

Observation d063ccc5-8c68-409d-b95c-a9aaa4d03dc4 · inbound

Reinforcement Learning Foundation Models Should Already Be A Thing cites this paper.

Reinforcement Learning Foundation Models Should Already Be A Thing Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:05.044993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T21:50:00.974590Z digest=sha256:53451ad41273fbbf44db200957ba736cd19ef075b4f591a2091fdb70b0e3adf9

Observation d81c660d-93d4-4273-92f3-c62ff223fc21 · inbound

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training cites this paper.

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:48.168776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:48.168776Z digest=sha256:1ee2c32822d4fb28b201e1af3e9303524ef49da2337e6cd848a0fa4f015e4baa