Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:36.790121Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.15040.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:36.790121Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:05:07.831643Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T05:05:07.890229Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c6106fea-1de2-43bd-88c1-f064a50f103d · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task What matters in on-policy reinforcement learning? a large-scale empirical study, 2020
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f99f2355-1bc8-4c2c-af04-2fad1097b93c · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959058a0-368f-4807-9fe7-87f25a6ef6e0 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0af9d966-68b8-4da4-9841-1f4664e5c7dc · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Kovalev, and Aleksandr I
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4302e45-ad40-4aaa-bf2f-d0656927aaa6 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks, 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79cc64d2-2fff-4366-90b6-451c7b92622b · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Empirical evaluation of gated recurrent neural networks on sequence modeling, 2014
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad80d92-8636-4b54-a4c6-8c27093eac20 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Le, and Ruslan Salakhutdinov
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f3b796-ca9d-475f-b969-3f42124e8064 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac5bf077-82ac-44d3-b52d-d5b77df93f78 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Addressing function approximation error in actor-critic methods
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cb30f82-13a3-476f-8bef-461584aee8f8 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Mamba: Linear-time sequence modeling with selective state spaces, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79fc7973-4af8-420d-9a92-4a68c75c8cb0 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Safe multi-agent reinforcement learning for multi-robot control.Artificial Intelligence, 319:103905, 2023
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83232638-4bd6-41a4-84fc-00c1b75aece6 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Safe and balanced: A framework for constrained multi-objective reinforcement learning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4aa36a4-8f82-44c3-9459-8e32a412ba21 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba726a9-a79f-4f66-936a-e9f0be26ca87 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Long short-term memory.Neural computation, 9(8):1735–1780, 1997
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e2b146-19c2-4334-954b-2e5eb1e22cff · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task The 37 implementation details of proximal policy optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 996a4878-2a4d-4e5f-a1c8-5791e6ba4936 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0013807-108f-4a30-ba82-095d37b8792d · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Decision mamba: Reinforcement learning via hybrid selective sequence modeling, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ba1c3c0-f64a-49f6-a824-1acc9db8bb54 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Du, and Huazhe Xu
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbd87fdd-e57a-44b2-bedc-0657196f08e6 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Efficient recurrent off-policy RL requires a context-encoder-specific learning rate
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a1c6f3f-f8ee-4dd5-b4c4-48f014ae176c · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d29e2f8f-3b87-4eb5-a936-b175e1663822 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Popgym: Benchmarking partially observable reinforcement learning, 2023
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2b295c3-ef45-4214-90ae-3fa2b1792084 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task When do transformers shine in rl? decoupling memory from credit assignment, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b568af3-cf56-4d9a-8073-3086ca7d744a · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Decision mamba: Reinforcement learning via sequence modeling with selective state spaces, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6e0cc61-88c0-4881-b376-71a9166ddc2d · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Francis Song, Jack W
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 896e1418-ee3c-4b2f-98cd-e6c2d2b0e9d9 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Transformerxl as episodic memory in proximal policy optimization.Github Repository, 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 806735a1-f767-4017-89ef-f929645697fd · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Memory gym: Towards endless tasks to benchmark memory capabilities of agents, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 094d1ed2-1d67-4e4e-b3d8-ac6d31fa2498 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Proximal policy optimization algorithms, 2017
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca03cca5-66c7-480e-b7cb-686f2195a7ca · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Mas- tering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1e8e2f-396c-41eb-adba-8127e43dcef3 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Mujoco: A physics engine for model-based control
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7625643-bee5-4928-9685-38c8dfbd0af1 · outbound
RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9610706-0245-42e6-82d7-53dc9175e2f6 · inbound
Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning RLBenchNet: The Right Network for the Right Reinforcement Learning Task
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.