Pith. sign in

Paper Citation Record · LEDGER

Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2306.09884.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.09884 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:54:39.273506Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9099ace3-0b92-4b5a-88e7-8cdc8116c452 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:31.315387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:e3bfa2730e79c8b45704a09fca6622f6da6fd39a5da5b7e1ece2ccb9b0c8f009

Observation c2b8b9d1-3d81-43d4-8c9f-0bc4ced60bf3 · inbound

Gymnasium: A Standard Interface for Reinforcement Learning Environments cites this paper.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.581158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:5ad08dce93600b2890138da87a41586fadc60f9b3a8765f37a44132f10ff795d

Observation 4f501582-ae8c-409c-abc9-2efc2ca501d9 · inbound

Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX cites this paper.

Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:39.273506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:54:39.273506Z digest=sha256:1b5224904a6ae239360089e1cf237a83cf0c39aae7da0b2e5162ce9c712403c5

Observation 5888563a-2747-4fc3-b190-2c4515fb867d · inbound

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments cites this paper.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.286849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.286849Z digest=sha256:34e8f68b1c904708e3769ab6a7cfe03eb71bf3b74b0af7b4f4a452d7a4fe92fe

Observation 5055a177-3d43-467d-a838-dad7540087ec · inbound

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation cites this paper.

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:12:55.256675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:12:07.920183Z digest=sha256:0e7eb564b4f1980d103e34605c407f8891133078b844346345803eccc7817135

Observation f54aa429-faff-4f7e-8b84-7923555950ee · inbound

Finding the Time to Think: Learning Planning Budgets in Real-Time RL cites this paper.

Finding the Time to Think: Learning Planning Budgets in Real-Time RL Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:51.943255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:08:19.504454Z digest=sha256:fe541f5ef20f019802c4432bd9d9828428db7ab9c0ec834104da765380878ab4

Observation c26d9e7c-5aa4-4022-b140-395749ef39fb · inbound

Finding the Time to Think: Learning Planning Budgets in Real-Time RL cites this paper.

Finding the Time to Think: Learning Planning Budgets in Real-Time RL Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:34:34.919692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T09:26:31.944405Z digest=sha256:09f2dc42eddf3c0b97c907a9e0acaf78445aef2a7ab2ba93d9217d6014620598

Observation 7c1fb0cb-613b-40d1-9b18-05945d554523 · inbound

Position: RL Researchers Need to Distinguish Between Solving Simulators and Using Simulators as a Proxy cites this paper.

Position: RL Researchers Need to Distinguish Between Solving Simulators and Using Simulators as a Proxy Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:25:48.447624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T01:28:44.074862Z digest=sha256:b9f133f508df6db18cb9e1f6541e3f9418921a6659a692a839f7fa3e09349155

Observation 9de7e73d-3082-4698-90f8-6dd10cb56ec4 · inbound

PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation cites this paper.

PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:15:39.082889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T21:09:25.719945Z digest=sha256:65698fec8c5ee17ffd22d142adc602a0122f0ed6585f4d80be824b4693ecb6bd