Pith. sign in

Paper Citation Record · LEDGER

Leveraging Procedural Generation to Benchmark Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:1912.01588.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.01588 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:29:20.075172Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

171
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 253eddd0-5bf8-4c84-9f40-fa56f9e8ab76 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:31.413770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:a3feb3f2f3a47b17029cc6308fbe4d9338bd0be3f81620e2a43bc566e7cb3ae0

Observation 104aae9c-0c71-449a-9de8-c9c97c00416d · inbound

Proximal Policy Distillation cites this paper.

Proximal Policy Distillation Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T23:13:36.980164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T23:09:42.421333Z digest=sha256:6209c17e67dfa57ff2468c8c3027d854fc8c9f8af31270345a52c3272456ae03

Observation 7cd0f961-0f24-48d1-acc7-6ebb1642d356 · inbound

Bilinear Convolution Decomposition for Causal RL Interpretability cites this paper.

Bilinear Convolution Decomposition for Causal RL Interpretability Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:54:49.062845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:54:49.062845Z digest=sha256:35fb5316471488f3ec9003e7ae1c3994cc1e34bd8a10cfbe6982938dd59f353e

Observation e545fb5e-91db-4ff5-9bc7-e808cb260031 · inbound

MineStudio: A Streamlined Package for Minecraft AI Agent Development cites this paper.

MineStudio: A Streamlined Package for Minecraft AI Agent Development Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:58.231308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:51:58.231308Z digest=sha256:74a16e2abff3058007ce9d8f538d5f09643bca3098c9e8fe4d5924b2d1238ab7

Observation bc175a12-9662-49e2-8ae2-d02c5c225af5 · inbound

Episodic Novelty Through Temporal Distance cites this paper.

Episodic Novelty Through Temporal Distance Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:58.788678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:25:58.788678Z digest=sha256:faa1877baf3ad8b13fab862e4fab36e3ad49d8623471a234247aa6a4e9835c4c

Observation f34a3016-e5e6-45cf-b792-ae3d899f5a37 · inbound

Cross-environment Cooperation Enables Zero-shot Multi-agent Coordination cites this paper.

Cross-environment Cooperation Enables Zero-shot Multi-agent Coordination Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:29:20.075172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:29:20.075172Z digest=sha256:bb813fc8cb065f85b66239982bb41434a60ca0a2f721e9535080782548d4901f

Observation b42d011d-d382-4b29-ae7a-c14f333f4633 · inbound

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models cites this paper.

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:40.090855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:59:40.090855Z digest=sha256:3e19a0ab758ff59ed252e38ba6f3ea7561e7dbffcacb5bb6256fe5aa7a0f0686

Observation 2d7ab01c-84ce-482a-990a-51ff1f76c9ce · inbound

Dyn-O: Building Structured World Models with Object-Centric Representations cites this paper.

Dyn-O: Building Structured World Models with Object-Centric Representations Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:02.762041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:02.762041Z digest=sha256:b3006db7e11f32ded6e6342e9486759bad4f622d2a6dfdc381d6a8724b09b6e8

Observation 6d28784e-5ce8-4fe5-8a07-7470b597d00c · inbound

Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines cites this paper.

Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:42:07.614826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T06:38:33.969927Z digest=sha256:9476f6b6e03eeca9094638bbb1f552072e3083d8f45ef47f5f685e13c0e0eec3

Observation be0497ab-3a38-4fa9-a49f-2eeb5dbd12c0 · inbound

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics cites this paper.

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:38:46.903783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:38:46.903783Z digest=sha256:96ef7b0c65e7afa761568d6b580862e9f2a699fc29b6cbedee9f072507288d98

Observation e1284f15-6fcf-4e2e-89e0-e9ca4c5edb42 · inbound

Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains cites this paper.

Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-06T10:34:35.473241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:34:35.473241Z digest=sha256:24361959c70074602bd65dbdc968a20a2ce4a2494ec608aded9cb37e714a0cf3

Observation 41bf3f85-9e3b-4c37-bee0-20063b01e320 · inbound

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning cites this paper.

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T05:21:35.522064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:21:35.522064Z digest=sha256:631e176dd027d366c67203e642fe1636c88442e416ac51a38c3fbbb9583f5ca5

Observation ff30be29-828b-4205-8356-f5a39941ec16 · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:20.671717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:1a671512e04b192a316397e1cec8970e53f66e975a2d99ce773a11be85bd840b

Observation 70afcc4e-51ae-4410-9c78-348c1d0cdaef · inbound

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization cites this paper.

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:36:29.552457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T09:55:00.402411Z digest=sha256:572bb607f688d4e0d9976896e26f6d87d98af4dfd6b59f2b849324587e67408b

Observation d063ccc5-8c68-409d-b95c-a9aaa4d03dc4 · inbound

Reinforcement Learning Foundation Models Should Already Be A Thing cites this paper.

Reinforcement Learning Foundation Models Should Already Be A Thing Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:05.044993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T21:50:00.974590Z digest=sha256:ffcbd7053873f7a83edea9f48b4def8cdab222d691a01e336c2038f601e17076

Observation d81c660d-93d4-4273-92f3-c62ff223fc21 · inbound

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training cites this paper.

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:48.168776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:48.168776Z digest=sha256:ff23609ccd3e0cd73e50466ec81b3289dc32bd3df73084ac179985e2d0db30cd

Observation 945fd4b1-9e8b-4285-80f0-423b5f106b91 · inbound

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks cites this paper.

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T10:11:39.627419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T10:11:39.627419Z digest=sha256:bd5fc4b3d179bc05ab5f533fa8ab06ca3a73e52e83d766189b01b77ccefcb61f

Observation d2b40c51-68c2-42fa-a5d8-a57fd9f57dd0 · inbound

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks cites this paper.

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-14T04:45:49.460695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:45:49.460695Z digest=sha256:afd2d953baa67849106a49ecf36203e290dc0680e4d5d672b97cea0a53859239