Pith. sign in

Paper Citation Record · LEDGER

Gymnasium: A Standard Interface for Reinforcement Learning Environments

As of 4 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 100 inbound Pith citation observations for arXiv:2407.17032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.17032 v4

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T17:29:49.186565Z

measured 134 of 134 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 100 of 148 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:50:39.504717Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact16
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch16

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4cd2fe2d-4ae0-4f3b-bcb7-695472a7e997 · outbound

This paper cites Hindsight Experience Replay.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Hindsight Experience Replay

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.674922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:b709a60df9dd6609d8aafd77748953aaa14c782233de24c27eb162691b2d2765

Observation 7a2ab3fb-be5c-48af-a7a1-597576227b25 · outbound

This paper cites DeepMind Lab2D.

Gymnasium: A Standard Interface for Reinforcement Learning Environments DeepMind Lab2D

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.543978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:9cebab5c1bb13488104e634b8c54af1dd6da5323d1a499cee62dff127ee849d5

Observation 47ba3158-f8f6-4f10-99d9-13016bb3d87d · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Dota 2 with Large Scale Deep Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T22:18:16.018301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:f3a3ebd675880f96239a3e480389ad1159eb25ddbe39484649c90ac035729d7d

Observation c2b8b9d1-3d81-43d4-8c9f-0bc4ced60bf3 · outbound

This paper cites Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.581158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:18f4d0f849204e8c028ed661c51e9840e9aa8f56f61fa6fcb2909bc799d348f0

Observation b1ccc781-d854-4b8a-8e83-e8c95781704c · outbound

This paper cites 10 Yann Bouteiller, Edouard GEZE, GobeX, Stefan Kuhn, and pius.

Gymnasium: A Standard Interface for Reinforcement Learning Environments 10 Yann Bouteiller, Edouard GEZE, GobeX, Stefan Kuhn, and pius

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.350796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:15997738b0f41dbcc49dee7e7fbc785c3018ce1c94d8d89a89507cb14045098a

Observation 2140442b-99ea-4ae9-a6a9-444fdf8c3e9b · outbound

This paper cites James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang.

Gymnasium: A Standard Interface for Reinforcement Learning Environments James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang

Reference 6

Resolution
verified exact
doi, observed 2026-05-11T17:29:49.294340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:01cc8044ba18d22a0ee023e9ac21435f9d4eaebf9096211cbc434df200ede60e

Observation 82dfe8da-00d1-4e51-af99-91d2ef66bc45 · outbound

This paper cites OpenAI Gym.

Gymnasium: A Standard Interface for Reinforcement Learning Environments OpenAI Gym

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:49:39.087655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:4d510f604c953af5777ec5cc23e2ad703a427ed3fdf5d12acb23ebe87797c89d

Observation 29c766d6-d04f-443d-9164-287c68df9a9d · outbound

This paper cites Dopamine: A Research Framework for Deep Reinforcement Learning.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Dopamine: A Research Framework for Deep Reinforcement Learning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.633591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:f76ca7bf1f4cf11eb56077cc0204a840ee3782db99b16b0dd2acc99cd403ff3e

Observation b9d3c1e1-02ec-42ad-b175-483b43ce798d · outbound

This paper cites ICU-Sepsis: A Benchmark MDP Built from Real Medical Data.

Gymnasium: A Standard Interface for Reinforcement Learning Environments ICU-Sepsis: A Benchmark MDP Built from Real Medical Data

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.646206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:e2082d6c60daeaf6a6fd3c9f107fe8bd0dad5dce9b6aa0f71c5c8fca337caf1d

Observation 43309982-9de4-4ee2-b6e5-20e8161f5ba2 · outbound

This paper cites Accelerating Reinforcement Learning through GPU Atari Emulation.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Accelerating Reinforcement Learning through GPU Atari Emulation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.361174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:eface17ee144454f3bde1475e03823409b7d473debb6b404e689aa050b273dc1

Observation 45f8087c-9b2e-4864-869b-69f171fa9bac · outbound

This paper cites OCAtari: Object-Centric Atari 2600 Reinforcement Learning Environments.

Gymnasium: A Standard Interface for Reinforcement Learning Environments OCAtari: Object-Centric Atari 2600 Reinforcement Learning Environments

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.375358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:c2caf1f32b713b79f00fdd44c8d2509bdb871c0bafdf3462922837e6b5dde1d4

Observation e505c5a5-281e-4d75-9c32-a743585f3e73 · outbound

This paper cites Alegre, Ann Nowé, Ana L.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Alegre, Ann Nowé, Ana L

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T17:29:49.690741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:f90c51cc233033b167e4d62ead25216a61200c467f8d3a0d2f7d501151e49e58

Observation 6a9d7e02-f680-4442-9920-1cce1b8bb4b7 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:48:10.852058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:f85b02251577a74938b0576bead25dde76dcfcafaff09412ec0b753cae47d2a2

Observation 5a4769fb-9a4c-44e1-934a-12ae09581d4b · outbound

This paper cites Acme: A Research Framework for Distributed Reinforcement Learning.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Acme: A Research Framework for Distributed Reinforcement Learning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.403651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:c22eab763b1db5bce9b31715bd506e25e389504686155c7945135f191074bc33

Observation 3986d04a-32dc-415c-8fcb-d5f74d588f80 · outbound

This paper cites Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.414324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:96ef8adb4a530f6085990d5fc4cdd3bf40d104880836e7325d38c8f930416d4b

Observation a56f97f1-baa0-4e74-b9cb-987d8424917e · outbound

This paper cites Littman, and Anthony R.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Littman, and Anthony R

Reference 19

Resolution
metadata mismatch
doi, observed 2026-05-11T17:29:49.269427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:0ac2f6f60daf10b36e5f863d5d19efb922014d2a1356faf1556148074d7d5d8e

Observation 8dfd740d-e136-4f10-a72a-2dcf61f958e5 · outbound

This paper cites GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS.

Gymnasium: A Standard Interface for Reinforcement Learning Environments GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.426309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:78ee319f82ad03fcecab87c82c77ad1885eee76a685474a2d4aa473c412ed582

Observation 28b4214c-4e7e-482a-84d7-be39b08f8ab4 · outbound

This paper cites ViZDoom: A Doom-based AI Research Platform for Visual Reinforcement Learning.

Gymnasium: A Standard Interface for Reinforcement Learning Environments ViZDoom: A Doom-based AI Research Platform for Visual Reinforcement Learning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.441443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:81ddc8b22d75527721c12e90c8e5a6437cc519e64600603a454780ef2d35a16b

Observation 9f0b7086-9d97-4fd3-9433-66dfd4d0b058 · outbound

This paper cites Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.458516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:075654c79022e8fab0ed2a5036de82b8ab1a2f7fc1b9eb303c1036345ed8ee3c

Observation 1d64acc1-f01a-4972-8322-7a90de2d153e · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Playing Atari with Deep Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:29:49.467867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:ab22cff1ef6cfcac81dd8d7572f2120279cb5d38d33d82fded9d27f1aeeeacc6

Observation e8131224-c63d-4fd7-a1fc-9185ed1cd0b9 · outbound

This paper cites Human-level control through deep reinforcement learning.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Human-level control through deep reinforcement learning

Reference 24

Resolution
metadata mismatch
doi, observed 2026-05-11T17:29:49.280781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:605a88a831c07093cf0e2ce48224b83d73eac6ff168826e06c8251ffa3c97740

Observation 951569ed-fd37-43a5-82d5-0f64b2f2ceb3 · outbound

This paper cites Behaviour Suite for Reinforcement Learning.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Behaviour Suite for Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.477736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:d8ddc4732c3ec4afcc528edbe877223456dd0fa4cde5b39635d55536ae77c11d

Observation 05960e65-0149-4b05-a8f7-30612bab9e4e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Proximal Policy Optimization Algorithms

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:29:49.484894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:2358aa2ebdba292ce501db16128bf7d9488f87653760bc6f7d3986af4c4db93e

Observation ba409965-1b10-484d-936f-75b5a6402f34 · outbound

This paper cites Mask Atari for Deep Reinforcement Learning as POMDP Benchmarks.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Mask Atari for Deep Reinforcement Learning as POMDP Benchmarks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.493298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:891e37e0c73054c34239f31e3dcb8ed9be8a133fb4a3e847b33400edd99828f7

Observation 7f1aee20-8479-4610-9cc4-a576405ad418 · outbound

This paper cites PufferLib: Making Reinforcement Learning Libraries and Environments Play Nice.

Gymnasium: A Standard Interface for Reinforcement Learning Environments PufferLib: Making Reinforcement Learning Libraries and Environments Play Nice

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.500000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:5959f68340401061dbf94ec0feb75d92ac622957e74f70cb779dd27985fe8f20

Observation 40010bf7-6710-44fa-8473-23467e9d535f · outbound

This paper cites PyFlyt -- UAV Simulation Environments for Reinforcement Learning Research.

Gymnasium: A Standard Interface for Reinforcement Learning Environments PyFlyt -- UAV Simulation Environments for Reinforcement Learning Research

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.510529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:051534e3e1491e71ff9609951c642462e7b281f483b3c51f436acd8130a7bc51

Observation aaf865b1-736a-4343-adaf-663e70cd6bfc · outbound

This paper cites Multiplayer Support for the Arcade Learning Environment.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Multiplayer Support for the Arcade Learning Environment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.520253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:7834bf2042e340d381994694a01931ebc3097b89dc6e7cffeb090ca0d78f656b

Observation b6698461-fc85-434a-867e-7f4fd4d33335 · outbound

This paper cites Mujoco: A physics en- gine for model-based control, in: 2012 IEEE/RSJ International Con- ference on Intelligent Robots and Systems, IEEE.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Mujoco: A physics en- gine for model-based control, in: 2012 IEEE/RSJ International Con- ference on Intelligent Robots and Systems, IEEE

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.318904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:231ccb3bfe8b30c5076000e62468f5f84b5c0409bb38a8a78fe8ca1f14bc9608

Observation 5d71b5d0-159d-40d0-a35c-4a37b0eccf94 · outbound

This paper cites 2020 , issn =.

Gymnasium: A Standard Interface for Reinforcement Learning Environments 2020 , issn =

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.338382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:a5219fff511a41cfbfd7015ebbc2da645b9f81e48e9ad5ede6cc983a67b85c87

Observation 22b1d74f-e0e4-4154-9188-acb09ad8e412 · outbound

This paper cites Tianshou: a Highly Modularized Deep Reinforcement Learning Library.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Tianshou: a Highly Modularized Deep Reinforcement Learning Library

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.526610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:dc29a99dbba9210937f73e3526187073815c2eb6af63a34c98df596047e528db

Observation 69376ba7-5857-4018-81b4-35b60e4d6ff4 · outbound

This paper cites Kenny Young and Tian Tian.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Kenny Young and Tian Tian

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T17:29:49.681594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:4fb4dc9d096e7038625110ad90926ab6c8aed037b4715560df03ceec6805cd05

Observation b918eef3-a7da-4d4f-b861-6ee4a0c7baba · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

Gymnasium: A Standard Interface for Reinforcement Learning Environments MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.654592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:40b0c8013b7e78ee9382c93388804c5df871a816c49a39d4b0da6704809c37b7

Observation e9ac05e1-c4d7-4699-b34f-433c7a1c65d8 · outbound

This paper cites Siyuan Zhang and Nan Jiang.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Siyuan Zhang and Nan Jiang

Reference 37

Resolution
metadata mismatch
doi, observed 2026-05-11T17:29:49.254957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:381d287c10cae0ef3eeecf652e9fb54cf6999ec3644b57b912b1dade0d50660c

Observation dcf5e911-491b-4013-b9ad-76fa0aa02644 · outbound

This paper cites Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.533695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:1b3da3e8213ecba88be11c55a8d0e248e3e5816657092b8468af8571b4bbf2c9

Pith citing papers

Observation 1ad05e3f-1abb-400b-8ee8-3e6c7d5bc467 · inbound

Fast State Stabilization using Deep Reinforcement Learning for Measurement-based Quantum Feedback Control cites this paper.

Fast State Stabilization using Deep Reinforcement Learning for Measurement-based Quantum Feedback Control Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:53:29.693750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T21:51:19.092922Z digest=sha256:b28833e78f933c0d2e8986c126503478f1d341583d0e898483ccb83c1d9da5a3

Observation e3984839-5ff8-4ae8-a313-c5bbae50d858 · inbound

Simultaneous Multi-die Floorplanning and Technology Assignment cites this paper.

Simultaneous Multi-die Floorplanning and Technology Assignment Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-23T03:12:27.823806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T03:12:22.952242Z digest=sha256:7518e47797c17cacc14aaa211a1cc0b38a34ea07d4f1ea8aae344094d9989a04

Observation dc687816-0c83-40b4-9119-be571fae50fa · inbound

Scalable Multi-Task Learning through Spiking Neural Networks with Adaptive Task-Switching Policy for Intelligent Autonomous Agents cites this paper.

Scalable Multi-Task Learning through Spiking Neural Networks with Adaptive Task-Switching Policy for Intelligent Autonomous Agents Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:47:01.207783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T19:45:07.172201Z digest=sha256:6b5c046333d770dcb6b77bbcae7b24e560002fa1907b4b19fe37d40392362370

Observation 76ee0218-c645-4b50-a96e-061bf73025c9 · inbound

Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cram\'er Surrogate cites this paper.

Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cram\'er Surrogate Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T16:41:47.655473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T16:36:48.916557Z digest=sha256:77fd14e509695f04b64db20da78dab9484eca3d9eaa793d1a935c5eb3ebf56eb

Observation 26fcbf04-c1e7-4db9-9f69-92a9ccdc7c54 · inbound

Leveraging Analytic Gradients in Provably Safe Reinforcement Learning cites this paper.

Leveraging Analytic Gradients in Provably Safe Reinforcement Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:07:15.364046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T11:04:23.284562Z digest=sha256:2b3f88ba6c082772540ba6323239b1f625e6cefe19e88281978b5a210cf78d61

Observation f45010df-8e02-4ec1-adec-3a1a9117f319 · inbound

DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty cites this paper.

DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:12:14.637710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:09:51.531488Z digest=sha256:c029add5759337b812b2e1babfb194aa0d68087638b313e5fb16033c1b453233

Observation ddd202c7-ae37-4a30-8447-3b99b380b73b · inbound

A Survey of Continual Reinforcement Learning cites this paper.

A Survey of Continual Reinforcement Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:27:11.186560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T08:27:03.376909Z digest=sha256:6c20c5e2b8437d68d554a3d9366e2305a72f6abd68f9a9bc160fe45dccac79a7

Observation a59e7710-90a6-467e-bbb0-bb781937730f · inbound

Deep Double Q-learning cites this paper.

Deep Double Q-learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:24:27.686506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-22T00:23:41.373753Z digest=sha256:97c601cc92e923896dd334362ea346dbd26f474ac04a05b1bc176b4da220993e

Observation f7f68170-0374-4a9d-acd1-92ff3a8dc84a · inbound

Adaptive Ensemble Aggregation for Actor-Critics cites this paper.

Adaptive Ensemble Aggregation for Actor-Critics Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:16:59.343268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T02:14:26.752215Z digest=sha256:27284dc9019c54e8b54eb81ecc9e06026621e4773b5b14b6acb8506165f61ce2

Observation 5080fb25-fd1f-480b-8f92-a5523977057c · inbound

Multimodal Remote Inference cites this paper.

Multimodal Remote Inference Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-19T00:31:55.722803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T00:28:59.722036Z digest=sha256:fd6d0871bd0c512c0ab87af120f64243c6a71dabcb66b5eaae5dbcfb12273f60

Observation 702ff302-e68b-4952-a337-6e711e05ce63 · inbound

A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions cites this paper.

A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 134

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:41:53.816999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T22:37:32.388931Z digest=sha256:89a257f1b8476b8b57c4e2d3aec7003f786fa4aba09aa36ad6b388bb1027c976

Observation 07600789-4410-451c-9e67-1c81d76bf389 · inbound

Frictional Q-Learning cites this paper.

Frictional Q-Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:51:30.330206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-18T14:47:49.452254Z digest=sha256:ae68831e5296a8fab547bfe4fb075e2b2733f77cf8be4771d9e7534cb1999a39

Observation 6887066f-0c03-4bd1-97f9-612a5687b85b · inbound

Activation Function Design Sustains Plasticity in Continual Learning cites this paper.

Activation Function Design Sustains Plasticity in Continual Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:01:23.654432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T13:00:27.749673Z digest=sha256:e71302cbdfe037888c597eb4978af42177aa3afb5ae312d2b8420661c51e7641

Observation b4df839b-8ef7-49df-8dae-7dab23f409c4 · inbound

Sample-Efficient and Smooth Cross-Entropy Method Model Predictive Control Using Deterministic Samples cites this paper.

Sample-Efficient and Smooth Cross-Entropy Method Model Predictive Control Using Deterministic Samples Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T09:41:11.868668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T09:41:09.249720Z digest=sha256:40352b568697580020892dacf20c699db8d3425f6a033806c223fec0ae63ed8f

Observation 0e918685-e03b-490d-8bf8-5b012e6cfbc8 · inbound

Transformer-Guided Deep Reinforcement Learning for Optimal Takeoff Trajectory Design of an eVTOL Drone cites this paper.

Transformer-Guided Deep Reinforcement Learning for Optimal Takeoff Trajectory Design of an eVTOL Drone Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.603806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:22:14.574161Z digest=sha256:c50c8ee838dd369ff4e0a4d6acc620f4ad7ed1b9c2e1991c140f86f96f0fca14

Observation 2424d9ec-c41b-4199-a6d3-96a67b382a48 · inbound

Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance cites this paper.

Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:01:31.980297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T04:59:49.126016Z digest=sha256:10dfd6bddd274dc0c9105e0e4f09d3310fbdf50a12dad9258e8a16cb4c902f01

Observation 6fa6eba9-beb3-414e-8c84-3f9f0a826e2d · inbound

Bench-Push: Benchmarking Pushing-based Navigation and Manipulation Tasks for Mobile Robots cites this paper.

Bench-Push: Benchmarking Pushing-based Navigation and Manipulation Tasks for Mobile Robots Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T16:50:39.504717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:50:39.504717Z digest=sha256:feb41b882cafc5a50783847b55b5b77fec6f57c4107131c86facec92ab4e7149

Observation e3c5a018-ffd2-4a77-aa52-d87da4bbf2b6 · inbound

Neural CDEs as Correctors for Learned Time Series Models cites this paper.

Neural CDEs as Correctors for Learned Time Series Models Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:28:41.114510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T23:23:44.658505Z digest=sha256:8ccdc6c29f7efa359c40e7326aea5b0afc9f9f88728dcbd1b3c5ae1e0523a9d4

Observation d9f81cab-8180-4215-9a49-47f63f87206c · inbound

The HydroGym Reinforcement Learning Platform for Fluid Dynamics cites this paper.

The HydroGym Reinforcement Learning Platform for Fluid Dynamics Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T15:16:58.894784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:16:58.894784Z digest=sha256:70675775e8bd441b9cd93dafc9f5cbeb7126dad96bc17df504cc8218b8e59b1f

Observation be4c433f-3587-481d-9bfb-28318e12f44e · inbound

About Time: Model-free Reinforcement Learning with Timed Reward Machines cites this paper.

About Time: Model-free Reinforcement Learning with Timed Reward Machines Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:01:16.070561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T21:00:52.571711Z digest=sha256:ada7a5128388acbccf0cae81118aeece0f82f2d6ac8a34e27c43a438e835e308

Observation e6f94848-e6ac-44c4-b7c4-2a52a9098c75 · inbound

Quantum Nonlinearity for Optical Neural Computing cites this paper.

Quantum Nonlinearity for Optical Neural Computing Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T12:51:11.368695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:51:11.368695Z digest=sha256:cab6939621b60ee754938e4ca8335aad75afa05fe30f5aec71408bc7e865c857

Observation be9ed1c7-6ae6-4714-acfd-bb9e1c5669a7 · inbound

Plug-and-Play Benchmarking of Reinforcement Learning Algorithms for Large-Scale Flow Control cites this paper.

Plug-and-Play Benchmarking of Reinforcement Learning Algorithms for Large-Scale Flow Control Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T09:03:32.549413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:03:32.549413Z digest=sha256:b0bcb1251b9fbf7db3cd69452a6d19b36903329b1dee012af080f420aa18d85e

Observation a6654283-b7bc-4612-bdc2-65b07539f193 · inbound

GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning cites this paper.

GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-03T07:16:24.426964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:16:24.426964Z digest=sha256:4582c4d472e04c998b9eb252e0da0ec1eba29a32b2685b3c1361a1469febd911

Observation 6ecb2fed-6451-453d-9547-b6a274ba7a3f · inbound

Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods cites this paper.

Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:45.267364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:45.267364Z digest=sha256:7c31d31ea8f68b365236931da12ebb6529e49d41ea17d0c20f2076e8af9f0c02

Observation aeb34403-e2bb-4364-9ba7-e810f91b3b1f · inbound

Agile Reinforcement Learning through Separable Neural Architecture and Applications cites this paper.

Agile Reinforcement Learning through Separable Neural Architecture and Applications Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T06:13:01.613750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:13:01.613750Z digest=sha256:3e7daa6edc78c71bfaf2f541c4d0f8b7f68b8571cd108ac825300ef234a7507d

Observation 938737e7-a3c8-49b3-8879-d20a0ff525ce · inbound

SUSD: Structured Unsupervised Skill Discovery through State Factorization cites this paper.

SUSD: Structured Unsupervised Skill Discovery through State Factorization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-03T05:40:30.125476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:40:30.125476Z digest=sha256:1119bb010fc3dc5f56dd46bd734db075d7a0bd6e0ffb6077b70d4f2de331fd29

Observation 46381aa2-bfae-4a83-83a6-21bda85e38d4 · inbound

The hidden risks of temporal resampling in clinical reinforcement learning cites this paper.

The hidden risks of temporal resampling in clinical reinforcement learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:07:29.623534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T07:05:16.379996Z digest=sha256:4ca754e64926799bef562e13671de8c77f722d7eccb0201ff3b63ecdeecd5e96

Observation 8ce43951-02aa-4c32-9da1-fd317ea5175d · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.882418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.882418Z digest=sha256:aa69b7bcce4b2d96ac5563097b4f5437684d0478434d1eca7bce7a90624b87ef

Observation d07f17c4-bf89-49c1-82cd-606914e9b0c2 · inbound

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents cites this paper.

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 106

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:51:25.902030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T10:49:35.846590Z digest=sha256:97c14f47e36b1b78ed46a87754598e13cb78f67606003320997bf99439ec098f

Observation 68e8a048-bded-4632-a9cf-b15a1136354f · inbound

RLGT: A reinforcement learning framework for extremal graph theory cites this paper.

RLGT: A reinforcement learning framework for extremal graph theory Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:21:38.151961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T21:21:07.914983Z digest=sha256:393ebf70eac8deaed9cda5f1d0fec1b1cf751f51bdb96572546ce1b68b2b334d

Observation 412567f6-b7de-4094-85a3-813aadda6717 · inbound

Asymptotically Optimal Sequential Testing with Markovian Data cites this paper.

Asymptotically Optimal Sequential Testing with Markovian Data Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T22:20:34.201622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:20:34.201622Z digest=sha256:597a9498f1ed9e810ea88fb4f49b367b7d6adf51261c9379b7806dd1b09ade77

Observation e04756ea-2138-4fbc-9f8b-6d4201bd6a28 · inbound

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation cites this paper.

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T20:06:58.253783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:06:58.253783Z digest=sha256:f3989120836d6472f0491fb509808185b607060a3c7e435d5ca094c065c38ef9

Observation b95d71bc-a1f9-49af-906e-eeb623321e46 · inbound

HY-WU (Part I): An Extensible Functional Neural Memory Framework and An Instantiation in Text-Guided Image Editing cites this paper.

HY-WU (Part I): An Extensible Functional Neural Memory Framework and An Instantiation in Text-Guided Image Editing Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-15T13:24:24.041663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:24:24.041663Z digest=sha256:ebf9a299ffea0bfde42ab1de39ff809a6d8ddd0d8e01fc1dcc8e590b83c9f99d

Observation 1f82c125-154d-45af-aa97-ec4c8389fe68 · inbound

A Survey of Reinforcement Learning For Economics cites this paper.

A Survey of Reinforcement Learning For Economics Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T18:34:54.957047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:34:54.957047Z digest=sha256:2121ce2515accd4315c261a18620ab7f3da0e56d59bf29f50a77f14ed1f5b9fd

Observation 0983aeb4-6f21-4a5f-bc14-1ddbf0418525 · inbound

Automatic Generation of High-Performance RL Environments cites this paper.

Automatic Generation of High-Performance RL Environments Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:05:01.478494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T11:04:11.718672Z digest=sha256:e00cc668bd1cc656b68a3ca46ab238e0bc13b1554b458127a8ecb70ae6745e66

Observation 0d12ab06-5198-4034-8ee9-93d26aa0eb85 · inbound

WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems cites this paper.

WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:44:09.358367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T11:40:28.002606Z digest=sha256:1cc83c5136e3f21f5506f0def52476b8f3307c30ba951cb5fe351bc648db7bff

Observation 0b8a735f-c51c-41e5-a642-9a0b52630854 · inbound

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels cites this paper.

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:09:22.556119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T04:09:22.328844Z digest=sha256:be699630de290fdaa493642a46a9fa6d2ee822f8e2d128f2604c4913292e1675

Observation eb33c335-3a86-411a-93e6-c0eea64e0500 · inbound

Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts cites this paper.

Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-15T11:52:28.028968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:52:28.028968Z digest=sha256:2e942d61a16e0f2989dfd7a2729849f8eb78e3abf2ef2d7c4a64e2c8741bf6d4

Observation 2da56bd9-ac95-45da-9a9c-979ccd5fb88c · inbound

Temporal Logic Control of Nonlinear Stochastic Systems with Online Performance Optimization cites this paper.

Temporal Logic Control of Nonlinear Stochastic Systems with Online Performance Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:48:18.951183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:47:55.704455Z digest=sha256:0052a1b1c99a1beba4f0d689e01977f2dff37830855c3ffa2439c53c0fe92929

Observation 82c0e00a-6532-41c3-b085-c0ed6ff1c9fc · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:04:56.512544Z digest=sha256:424b0f32924271bf419b018791531dde2c66ecfe7cbf3912fadacdbc8bdbc067

Observation ac930a4a-9087-4b3b-8972-4a2775899f8e · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:12:41.330190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T17:08:31.770889Z digest=sha256:bb53de05294263e5e92fb58e9da910bd65199464388784a2dc6b3e1bcea2142f

Observation 35d27494-b206-40b0-96cb-47564bfeb305 · inbound

Gym-Anything: Turn any Software into an Agent Environment cites this paper.

Gym-Anything: Turn any Software into an Agent Environment Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T19:34:42.621666Z digest=sha256:b2295b2c792de117adab35626b683e8076f3d67cf592b20b0e1841559763a126

Observation f5c93907-05d9-4019-bc1f-a31a7a913153 · inbound

RAMP: Hybrid DRL for Online Learning of Numeric Action Models cites this paper.

RAMP: Hybrid DRL for Online Learning of Numeric Action Models Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T17:06:22.252782Z digest=sha256:828f3d49bdaf2a4724bc72be62f2ed7e7ffd35f9cda9e4bb6db1e62f5cf56376

Observation b0edf386-0068-491b-a43a-595b228b0729 · inbound

SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning cites this paper.

SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T17:50:56.101903Z digest=sha256:842d9c0089a3c96bec9a9e2978913c808ee1944056d618a8f5614442a2c3fcd9

Observation c71f7138-7a7a-4d13-9d0b-2260efe82d3f · inbound

GPU-Accelerated Continuous-Time Successive Convexification for Contact-Implicit Legged Locomotion cites this paper.

GPU-Accelerated Continuous-Time Successive Convexification for Contact-Implicit Legged Locomotion Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:49:33.010827Z digest=sha256:9634ba61a3c706bc24c7db6aeea4da7fd7a885c7edff9b88f4ec7dda223bbe40

Observation e3c9d446-9494-4470-9654-b63c66a024a7 · inbound

[COMP25] The Automated Negotiating Agents Competition (ANAC) 2025 Challenges and Results cites this paper.

[COMP25] The Automated Negotiating Agents Competition (ANAC) 2025 Challenges and Results Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T12:13:21.122831Z digest=sha256:45ad51eecdeceb20bde26662effc89ee96939aaad848d957b18f58a72bed7474

Observation 4f003452-8588-4760-9c6d-b6d6b6097cda · inbound

Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning cites this paper.

Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T19:46:39.624903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:46:39.624903Z digest=sha256:732e9888b7fa2d1dab666ea4fe70f26950226dd4d2af492cbfb7013c8800c46d

Observation f60a74c0-6484-4a8a-8930-4ad7b94ce600 · inbound

Efficient Federated RLHF via Zeroth-Order Policy Optimization cites this paper.

Efficient Federated RLHF via Zeroth-Order Policy Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T04:48:15.394329Z digest=sha256:22be5207c1312403e68be146324081c61d3c380c6172944a4f9f229e2f9ead55

Observation 67f64fa7-21cb-408d-bbad-219f4bdd540e · inbound

Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation cites this paper.

Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T02:32:58.434291Z digest=sha256:52c4ebbaf0de3500fb568bbbfdc2a968f9f34d89212e0fdc95339583d6030dae

Observation c24c3f55-33f8-4999-b328-8d0c7dd47223 · inbound

RL-ABC: Reinforcement Learning for Accelerator Beamline Control cites this paper.

RL-ABC: Reinforcement Learning for Accelerator Beamline Control Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T03:18:53.581919Z digest=sha256:a433e786b4eed28136a066db27298a5f86d9da8391bedf8f84fc4b96b6890db1

Observation c9ae02b6-77aa-452e-8942-b0f8eb253b20 · inbound

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps cites this paper.

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:23:56.777888Z digest=sha256:b7419b8440f5d07c8a33291b625b3ef00e0e83021934aa5408e95501f9f2ddaa

Observation 843bb37b-6d6e-4a03-9df5-8633bf4045d2 · inbound

Benefits of Low-Cost Bio-Inspiration in the Age of Overparametrization cites this paper.

Benefits of Low-Cost Bio-Inspiration in the Age of Overparametrization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:21:37.686664Z digest=sha256:3adb58ef04f94af37036f54c3b150e3cb6131d71610de79930922e0ff61c72fb

Observation 846606e3-da43-46a0-9da3-4d1f94647727 · inbound

Replay-buffer engineering for noise-robust quantum circuit optimization cites this paper.

Replay-buffer engineering for noise-robust quantum circuit optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T22:11:29.151910Z digest=sha256:a487b55b03682b9e892bdcd46ecb6e7356ad01c36e1c3f68af0510d9f1825ae7

Observation eb4729b3-9039-45f5-ae73-17271a39e760 · inbound

SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning cites this paper.

SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:51:30.614872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:01:31.928056Z digest=sha256:81c0565437286fb10b1caf9f844ae47892c9f7b44eac59904241d5f332a56340

Observation aee757c0-49d9-46de-98a6-b48d0997e3b4 · inbound

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning cites this paper.

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-12T00:16:20.265668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T15:39:17.505374Z digest=sha256:6463d4328475e4373bd35a9aa794a64e8280ca19e73d56211c59d5a52480017d

Observation 6cef9daf-55b2-4021-848d-24eb401a2ca5 · inbound

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations cites this paper.

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:36:26.187420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T10:27:55.337253Z digest=sha256:797447865e1b5069d282db22c6af69634a55adeaed7f98d6ed0e5f7367f515b3

Observation 1a18a7ab-c157-464b-8c4b-646c84dd8d1c · inbound

Your Loss is My Gain: Low Stake Attacks on Liquid Staking Pools cites this paper.

Your Loss is My Gain: Low Stake Attacks on Liquid Staking Pools Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T17:57:37.260492Z digest=sha256:c30eb408c3b46faa289e6c3328c7715f5d08874514ae371dc6c642987d15f6f5

Observation 2180fb82-a11c-4be6-b42b-be9544e69170 · inbound

Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs cites this paper.

Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-09T15:12:54.575483Z digest=sha256:17b0b59a9416dc48b9c8d7fd5110d1798ef45a3f5ba30cbfaa7dce88709777bc

Observation 5045a22c-35ab-447d-9b57-45b74d31c37b · inbound

A Multi-View Media Profiling Suite: Resources, Evaluation, and Analysis cites this paper.

A Multi-View Media Profiling Suite: Resources, Evaluation, and Analysis Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-09T15:22:45.263415Z digest=sha256:e3a0a1b641512c0f5e4a8fd57c5fbfbb17be692cdb4b33fa0fc5b10486008255

Observation c36c3f00-1402-40bc-8a79-fd8900a8d0af · inbound

PACE: Parameter Change for Unsupervised Environment Design cites this paper.

PACE: Parameter Change for Unsupervised Environment Design Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T14:21:34.002785Z digest=sha256:ac95ada82329bc20db715f183b66e6bb834c884df67fad2403330f65377d06fc

Observation e74b8414-3dea-40a7-a85c-03f7683dced1 · inbound

Perturb and Correct: Post-Hoc Ensembles using Affine Redundancy cites this paper.

Perturb and Correct: Post-Hoc Ensembles using Affine Redundancy Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-09T14:27:18.519853Z digest=sha256:cc256362eb5c331c500215321d324ee0ef697920763b757fa7bf9359b010daf3

Observation b12f7b7a-b0d2-478c-be31-5fe0b5064eb7 · inbound

Stable GFlowNets with Probabilistic Guarantees cites this paper.

Stable GFlowNets with Probabilistic Guarantees Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T19:22:03.198766Z digest=sha256:dda4b8935462db03b9dd8e893d4e9b83a611de77f8f391f8f22f6f1e05cd0316

Observation 1e8a6a7f-572f-4583-9f9a-37aa4dd0cb5b · inbound

Zero-Shot, Safe and Time-Efficient UAV Navigation via Potential-Based Reward Shaping, Control Lyapunov and Barrier Functions cites this paper.

Zero-Shot, Safe and Time-Efficient UAV Navigation via Potential-Based Reward Shaping, Control Lyapunov and Barrier Functions Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T17:01:36.202798Z digest=sha256:242f428521dea3ea0282c318e67789e473e677eae2a614d1155baf6474a8b079

Observation 207a6697-804f-4d84-951d-7d65cd4d25e1 · inbound

Training Non-Differentiable Networks via Optimal Transport cites this paper.

Training Non-Differentiable Networks via Optimal Transport Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T15:37:42.167420Z digest=sha256:29a2b7bc57661b5e3939ef6d53a865103e6bf14084a399e550d8ab24bb9312e1

Observation 3fd50936-1b7d-423e-9448-324f3a2867ae · inbound

Bridging the Gap Between Average and Discounted TD Learning cites this paper.

Bridging the Gap Between Average and Discounted TD Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T19:17:19.065418Z digest=sha256:8d8ff73cb8f595763562dab7c6b7f071dee9655716005f937776683656d5ea21

Observation 158da49e-1e54-4dae-bdf4-f70888606603 · inbound

ANO: A Principled Approach to Robust Policy Optimization cites this paper.

ANO: A Principled Approach to Robust Policy Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T19:34:12.002351Z digest=sha256:870c5ff683dd0f4069174a982aa13b07658d1241970f3004a60837c05323ae34

Observation 5b13e0b7-7c9e-44a6-83b8-8998fc5256c2 · inbound

Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters cites this paper.

Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T19:29:51.400740Z digest=sha256:a7d3f0592d8c080e5ffa5dfdac353201cc6819c66f686a37be635937275c376f

Observation 162b1949-5a21-4f25-87a3-c75fae374c83 · inbound

Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters cites this paper.

Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T00:25:09.191771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-01T00:16:02.179335Z digest=sha256:c1f7ba27d5f3c93cd0debb36e2a76c234016e8d8e03690e801a4edd1cf993c82

Observation 67c23977-be50-4503-ba65-2dd429e54f54 · inbound

Extending Differential Temporal Difference Methods for Episodic Problems cites this paper.

Extending Differential Temporal Difference Methods for Episodic Problems Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T18:19:57.472765Z digest=sha256:d70c18928ee70f02a75bde8191c21467df196d7a0d0666053cc0ef7aff61350c

Observation 8d3081a2-af6b-42e8-b0f6-04f455eb6e29 · inbound

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior cites this paper.

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 160

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T17:47:09.591001Z digest=sha256:ca488b6c94857e3c07d2174160d29e37831fd74d1ec6a9eb87669bfcb72b2ff0

Observation 95470c55-0608-4c4b-8887-a4a54a6580d5 · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:08.425241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T14:54:30.895137Z digest=sha256:73d068c09fc1d712ad7ccae15c628d888162d28b173733a77369824429979f9b

Observation acd0fd55-2ae8-4abe-a073-f7b53325d2ee · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:19:56.733509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T09:15:11.280343Z digest=sha256:c89d562c61771a2e271f7df29b183065ba788479d84ceffb8a6aec0e9c6003e4

Observation 336dcce8-05a3-40ab-803c-2688c86a6353 · inbound

Causal Reinforcement Learning for Complex Card Games: A Magic The Gathering Benchmark cites this paper.

Causal Reinforcement Learning for Complex Card Games: A Magic The Gathering Benchmark Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:46:09.927990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T13:59:01.188044Z digest=sha256:5108872054422c120c1c9cf8100f906fc056b24ff3ea8bec1c3a1c001206df23

Observation a78fd08b-e7c9-46aa-9f4c-5d22555b8d53 · inbound

AdaGamma: State-Dependent Discounting for Temporal Adaptation in Reinforcement Learning cites this paper.

AdaGamma: State-Dependent Discounting for Temporal Adaptation in Reinforcement Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:51:06.882767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T13:47:18.720944Z digest=sha256:9ee6c30fb697802665c664d49e39bfcb83799ad53b1133a6cb33309bc8e86e93

Observation fa132447-0738-46ef-af65-b86162cde18b · inbound

Beyond the Independence Assumption: Finite-Sample Guarantees for Deep Q-Learning under $\tau$-Mixing cites this paper.

Beyond the Independence Assumption: Finite-Sample Guarantees for Deep Q-Learning under $\tau$-Mixing Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T21:36:13.901669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T05:01:53.381535Z digest=sha256:10564e65f89c66fda6fbc851d85cda400208838dd1e9a5cb35b384a4d5503719

Observation e7486c67-4604-407b-be41-54b9499ba9b0 · inbound

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents cites this paper.

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:29:49.696310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-11T01:22:32.713175Z digest=sha256:d2474d8cfcfb01db9cd150c32c9727d59f5e3004e300030cffd836d0dd1489e6

Observation a62cd7d9-b02e-4c33-9fba-1f821390babc · inbound

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents cites this paper.

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:52:57.917077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-14T20:51:26.063471Z digest=sha256:19beb0705bc98107611f8e6e27f255f6521ccb576645205ca809d1c3c832066b

Observation ec253524-8436-41c9-a6dd-a2256f406f7f · inbound

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias cites this paper.

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:41:23.994400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T00:51:58.315109Z digest=sha256:40a4f7a061dda35a7e3e64bd9b1490c6ae2b423a5d0f4661366400d460f5c967

Observation 32d6f508-ee97-4b25-990f-7081447cdc9b · inbound

Learning When to Stop: Selective Imitation Learning Under Arbitrary Dynamics Shift cites this paper.

Learning When to Stop: Selective Imitation Learning Under Arbitrary Dynamics Shift Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T07:21:25.340301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T03:27:41.845716Z digest=sha256:eab815faba84d1e20ac2a1e304d16ede71dd86fa7794d4067b930038ce4a0cee

Observation 77669a96-ff19-4e39-8ecc-1cc8cf0ad116 · inbound

Learning When to Stop: Selective Imitation Learning Under Arbitrary Dynamics Shift cites this paper.

Learning When to Stop: Selective Imitation Learning Under Arbitrary Dynamics Shift Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:09:07.577044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-20T22:04:57.245024Z digest=sha256:722d127ef97a3736c542c70dafa3b3f3288811c527db4bd7de9fae891209cc56

Observation e046d3de-47df-49ed-a95b-636afa8a1255 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 163

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:18.997445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:3cac48baedf1d6ef9a5a16d597265ba749ce9698ae8d3443bfb79c794c648454

Observation 6ace471f-92f0-40f3-be52-3920914b9610 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 163

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:47:27.046006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:37c87a2f67e19809514cf3241a015c6a6da068b25d93211a9bcb7e5df5218d0a

Observation 6ab60d41-7c0f-4098-90bb-ad112c1ef884 · inbound

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning cites this paper.

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:24.938712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:28:30.722268Z digest=sha256:f28faf17774315bfaf6851975c7b29c44479122b7424c37b82fececf80dedbd9

Observation 86cd9959-e117-4691-9a7a-f24d9fef8575 · inbound

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency cites this paper.

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:52:08.791092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:50:20.040271Z digest=sha256:c3ca91c0cabe35d400edab5508e5f9c0ede32d1b0e1178481f7c920d2eae1a39

Observation 6c8800c4-36c0-480a-920a-9039effada32 · inbound

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs cites this paper.

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:27:29.077799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T07:27:12.302109Z digest=sha256:7f8a1fe3997ccad16cba12a3daabce414d70ef5c586a6377f0e372e355fab3c9

Observation c0607a29-0cae-4f0b-88b5-5e8e094bac8d · inbound

Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance cites this paper.

Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:42:57.299765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:42:06.571221Z digest=sha256:2fd100c22c74eb5a153a7268c68428790beb5051ea8b81ceaded5df624ffe70a

Observation 7f60a866-ded5-4b7c-ae9c-3ad0fe4700ac · inbound

Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models cites this paper.

Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:25:04.202338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T21:15:15.155369Z digest=sha256:8228a45ebcd865c1b0feba2c6c2676091a5497b530385050357828d750ce1be4

Observation c6832d17-d591-4e2c-9597-18ad1de8311e · inbound

Chrono-Gymnasium: An Open-Source, Gymnasium-Compatible Distributed Simulation Framework cites this paper.

Chrono-Gymnasium: An Open-Source, Gymnasium-Compatible Distributed Simulation Framework Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T20:45:03.897568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T20:41:31.816389Z digest=sha256:9f0f0ae6da3b84b01d5c471613f0cb49978f8ae764b7baf4aeffb8552d105897

Observation 236bb79c-c07c-40fe-b1b0-2de68588f6d4 · inbound

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning cites this paper.

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:27:41.312843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T17:26:04.762451Z digest=sha256:610e697b5edc7cbbcf7c388c619db03f6ba6b1b6ba36b29bcd510a858e73c138

Observation f5214580-3491-490f-a012-b1ecb8bd8cc5 · inbound

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning cites this paper.

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:15:04.597378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T21:08:47.084947Z digest=sha256:95f0d5e3ebbb7f88c6a7e86fa2793eb8912a75b04fd06578e81d0e31317fb8df

Observation 6f7ee100-2056-4600-a74e-5b8bebabeb3b · inbound

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents cites this paper.

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:28:14.503778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T11:24:48.558423Z digest=sha256:d4de406ae96bb0d569c27a76816ccdca10f319baf08398a1a2ef1646468db0dc

Observation ec72341c-efe2-43c7-b204-e6b7fd44e9e6 · inbound

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind cites this paper.

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:09:46.256886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T07:09:37.399954Z digest=sha256:d3c719e6bac19f845c3620f540fa7ccd4c529bcf2e05d3c8508d40721db55947

Observation 434d659b-ceb6-4c72-aad6-a97ac7ce6634 · inbound

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale cites this paper.

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:59:45.569781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T06:56:27.532299Z digest=sha256:dfba8ba7dfa199b600f8468db4dd2d297693e9f2b3fc6d4688232172bd4c06c1

Observation 0d9ceeef-ef5f-4ede-a589-277ab7314472 · inbound

stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation cites this paper.

stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:20.238336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T08:57:19.834179Z digest=sha256:44c885bbe9717399acf8077c33b73ab507f50e5c5118ad92885fb81a4f98c541

Observation 2fb51365-d222-4e43-98f5-7ba26eb10f64 · inbound

Score-Based One-step MeanFlow Policy Optimization cites this paper.

Score-Based One-step MeanFlow Policy Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.029853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:7145b36e6163bf7b4506e468b50e131e6d93432007e4d517ac3fd42e52a47dd2

Observation f47238c6-5d8b-4993-80bc-02e0e25620f5 · inbound

Quantum Frog: Emergent Cooperation and Difficulty Scaling in a Quantized-Time Cooperative Game cites this paper.

Quantum Frog: Emergent Cooperation and Difficulty Scaling in a Quantized-Time Cooperative Game Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-05T06:00:43.858829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-05T05:56:17.409844Z digest=sha256:6bb898d7db3031b839e59352788399a1b7e30bcf529fde6eec03cf597d29a7f0

Observation 0dfdb80a-7112-49be-92c5-ba8fc8deea52 · inbound

Distilling Game Code World Model Generation into Lightweight Large Language Models cites this paper.

Distilling Game Code World Model Generation into Lightweight Large Language Models Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:04:44.683226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T13:58:37.956333Z digest=sha256:32c714b166cb69f5991b5078bf41517149d26fcf22feabccd4d4cebbc88241ae

Observation 85933b69-edae-42c1-a1db-25f63d3a3a95 · inbound

Cost-Aware Adaptive Conformal Inference for Runtime Assurance in Dynamic Environments cites this paper.

Cost-Aware Adaptive Conformal Inference for Runtime Assurance in Dynamic Environments Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:24:40.135912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T13:19:24.185337Z digest=sha256:0a872319d6742509a129193fd65249f3f36cbe0e1120ea5bccfd8e00ffcda917

Observation 636da051-2c89-4e39-b7b0-f02b3547825b · inbound

When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability cites this paper.

When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:35:48.581702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T01:20:04.292547Z digest=sha256:10b0e03fd4839c58a703a0de49a052f4a9b8beb86ac906383b6d0fbf09e910c8

Observation 99af8fe6-e639-42de-a804-1f26e58a7b5d · inbound

Adversarial Dual On-Policy Distillation from Expressive Teacher cites this paper.

Adversarial Dual On-Policy Distillation from Expressive Teacher Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T19:03:51.574458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T18:55:23.777484Z digest=sha256:6732b874dec0b1423ef40a5aae04ef12c68deb8bebffa11ce041a683b73af388