Pith. sign in

Paper Citation Record · LEDGER

RLBenchNet: The Right Network for the Right Reinforcement Learning Task

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.15040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15040 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:36.790121Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:05:07.831643Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:05:07.890229Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6106fea-1de2-43bd-88c1-f064a50f103d · outbound

This paper cites What matters in on-policy reinforcement learning? a large-scale empirical study, 2020.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task What matters in on-policy reinforcement learning? a large-scale empirical study, 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:41.938153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:33.206676Z digest=sha256:54ea45bbe3117c705bee6c6bbf93c8dd5c606e92ccbe03401750388b3f73757f

Observation f99f2355-1bc8-4c2c-af04-2fad1097b93c · outbound

This paper cites an unresolved cited work.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:33.282105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:33.282105Z digest=sha256:09963008065e2bc497889434d53d65f65d0ba9fae107a09c0a6e23b7c6b1f70b

Observation 959058a0-368f-4807-9fe7-87f25a6ef6e0 · outbound

This paper cites an unresolved cited work.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:41.670666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:33.436858Z digest=sha256:80598d3b53bbc33e6398661dbf8f745054dda4a2331c02f5e546ab869c66e910

Observation 0af9d966-68b8-4da4-9841-1f4664e5c7dc · outbound

This paper cites Kovalev, and Aleksandr I.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Kovalev, and Aleksandr I

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:41.438864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:33.551580Z digest=sha256:0caefe8ed56191860bf1e7da69cf9746cf5c6bbb63ead3f660bdbc4049c8c96a

Observation f4302e45-ad40-4aaa-bf2f-d0656927aaa6 · outbound

This paper cites Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:41.211628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:33.668828Z digest=sha256:943a05af335ca186eec025eaffe2982227a21780e96c74f37facbd4c48031d56

Observation 79cc64d2-2fff-4366-90b6-451c7b92622b · outbound

This paper cites Empirical evaluation of gated recurrent neural networks on sequence modeling, 2014.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Empirical evaluation of gated recurrent neural networks on sequence modeling, 2014

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:33.788457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:33.788457Z digest=sha256:146f43a5b519865e7805b16c1a226da93c1d5bc7d77f2e687275bc16d92b1f20

Observation cad80d92-8636-4b54-a4c6-8c27093eac20 · outbound

This paper cites Le, and Ruslan Salakhutdinov.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Le, and Ruslan Salakhutdinov

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:33.909597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:33.909597Z digest=sha256:be1dc501d4e1699b1f0baa5ce9b82d239292dc2fe427c21ebb4c9b5ae392d08d

Observation 25f3b796-ca9d-475f-b969-3f42124e8064 · outbound

This paper cites Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:40.997531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:34.079669Z digest=sha256:d3292acff13d0a87beebc462cd07fd6a5e7505d4119d04b24cee6973a24da74b

Observation ac5bf077-82ac-44d3-b52d-d5b77df93f78 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Addressing function approximation error in actor-critic methods

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:40.719291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:34.209058Z digest=sha256:6a2d5ee6939f58f93828a007b95ba4025dd0f3ddf4f3530e42fa81ea75d4f029

Observation 3cb30f82-13a3-476f-8bef-461584aee8f8 · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Mamba: Linear-time sequence modeling with selective state spaces, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.339686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.339686Z digest=sha256:e5efe14b04d4cc463218a47c568478317d52e4338d269a2776f4ac7553d3929c

Observation 79fc7973-4af8-420d-9a92-4a68c75c8cb0 · outbound

This paper cites Safe multi-agent reinforcement learning for multi-robot control.Artificial Intelligence, 319:103905, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Safe multi-agent reinforcement learning for multi-robot control.Artificial Intelligence, 319:103905, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.462727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.462727Z digest=sha256:1e0a9950eef51eabac750f19dc61930accfd838ac9a858b4ccba296ea510d06a

Observation 83232638-4bd6-41a4-84fc-00c1b75aece6 · outbound

This paper cites Safe and balanced: A framework for constrained multi-objective reinforcement learning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Safe and balanced: A framework for constrained multi-objective reinforcement learning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:40.474550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:34.598603Z digest=sha256:6e4b890a24e798a2130fdb1c7051582f6a4965d37b8d209aa6ec558f638aac00

Observation e4aa36a4-8f82-44c3-9459-8e32a412ba21 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.737787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.737787Z digest=sha256:018b6d1fd2bdc83d09efc6e30c37d85886a4a57cc24d4cb582c0fcb9f1b4b85f

Observation bba726a9-a79f-4f66-936a-e9f0be26ca87 · outbound

This paper cites Long short-term memory.Neural computation, 9(8):1735–1780, 1997.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Long short-term memory.Neural computation, 9(8):1735–1780, 1997

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.828624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.828624Z digest=sha256:730a358e3148b2d734d9bc19c85e7959fb47760277c2c0d3b0b43be13ff8b9c2

Observation d3e2b146-19c2-4334-954b-2e5eb1e22cff · outbound

This paper cites The 37 implementation details of proximal policy optimization.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task The 37 implementation details of proximal policy optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.979587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.979587Z digest=sha256:f88f1ae2fbf3d979c8633ba7657f50a2d94018407135f17348d937a91a3a3d99

Observation 996a4878-2a4d-4e5f-a1c8-5791e6ba4936 · outbound

This paper cites an unresolved cited work.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:39.966752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:35.265411Z digest=sha256:d13ae5d3968f206e9c58c6eba47c860e1e679ffa99b8e2dc31a2e3bf39a5880d

Observation c0013807-108f-4a30-ba82-095d37b8792d · outbound

This paper cites Decision mamba: Reinforcement learning via hybrid selective sequence modeling, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Decision mamba: Reinforcement learning via hybrid selective sequence modeling, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:39.702547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:35.380323Z digest=sha256:a804aecd91e35d335ae561c8a5f5accc16a6ad0a140e6e0482e29a9558410f85

Observation 3ba1c3c0-f64a-49f6-a824-1acc9db8bb54 · outbound

This paper cites Du, and Huazhe Xu.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Du, and Huazhe Xu

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:39.471916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:35.513888Z digest=sha256:0ccf75cadd4d98da4eb0aceee683199243727e378c73c64a733bd9f37f5aeae5

Observation bbd87fdd-e57a-44b2-bedc-0657196f08e6 · outbound

This paper cites Efficient recurrent off-policy RL requires a context-encoder-specific learning rate.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Efficient recurrent off-policy RL requires a context-encoder-specific learning rate

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:39.170120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:35.656320Z digest=sha256:9e60c71b7adc4ccabce54616dab92bd64f122d7c983384099eea78d375a0ce12

Observation 2a1c6f3f-f8ee-4dd5-b4c4-48f014ae176c · outbound

This paper cites Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl, 2025.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:38.809913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:35.846182Z digest=sha256:71f3bb744ce1e07e067fad6257b4cc7ac8ed895117b4b281e0e52c7c00d8a60b

Observation d29e2f8f-3b87-4eb5-a936-b175e1663822 · outbound

This paper cites Popgym: Benchmarking partially observable reinforcement learning, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Popgym: Benchmarking partially observable reinforcement learning, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:38.519287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:35.950671Z digest=sha256:95f61913fd64cfcb1091de1a14d83fca0561d32b51c2f11b8c7c19e9f7690a74

Observation b2b295c3-ef45-4214-90ae-3fa2b1792084 · outbound

This paper cites When do transformers shine in rl? decoupling memory from credit assignment, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task When do transformers shine in rl? decoupling memory from credit assignment, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:38.227986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:36.065512Z digest=sha256:c5278ae77c246050a83a57bfbe6b63d52187a1f2362486f8bc928ebfd4dfe5e4

Observation 0b568af3-cf56-4d9a-8073-3086ca7d744a · outbound

This paper cites Decision mamba: Reinforcement learning via sequence modeling with selective state spaces, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Decision mamba: Reinforcement learning via sequence modeling with selective state spaces, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:37.933547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:36.196985Z digest=sha256:d5bb085f9b96cdcf2644b5a333f21218f6a3fe818fdce46c374ea32465726cd9

Observation a6e0cc61-88c0-4881-b376-71a9166ddc2d · outbound

This paper cites Francis Song, Jack W.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Francis Song, Jack W

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:37.694539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:36.332756Z digest=sha256:e98582c19f1e1eeec79ab1f2c276002f082e3699df52134da4fe08b899ec031f

Observation 896e1418-ee3c-4b2f-98cd-e6c2d2b0e9d9 · outbound

This paper cites Transformerxl as episodic memory in proximal policy optimization.Github Repository, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Transformerxl as episodic memory in proximal policy optimization.Github Repository, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:37.480182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:36.456294Z digest=sha256:c65cd03800d0333e9c5952e91f75911017cd76f1a0e502be63bd5b5487f9fae4

Observation 806735a1-f767-4017-89ef-f929645697fd · outbound

This paper cites Memory gym: Towards endless tasks to benchmark memory capabilities of agents, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Memory gym: Towards endless tasks to benchmark memory capabilities of agents, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:37.293317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:36.550031Z digest=sha256:b20390f326d414803a30c70ace07e7ddfa74c9f6db8e4901bb3f8032ff0d7a1a

Observation 094d1ed2-1d67-4e4e-b3d8-ac6d31fa2498 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Proximal policy optimization algorithms, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:36.638559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:36.638559Z digest=sha256:65150ace36c77e5e6007880ad844110d96ac594e94993e8dfd80e85cef99c35f

Observation ca03cca5-66c7-480e-b7cb-686f2195a7ca · outbound

This paper cites Mas- tering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Mas- tering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:36.723806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:36.723806Z digest=sha256:1fadc65bad377a8c4b0493931fcb3d0fbe1d4899849e3cede801e94e755a670e

Observation ed1e8e2f-396c-41eb-adba-8127e43dcef3 · outbound

This paper cites Mujoco: A physics engine for model-based control.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Mujoco: A physics engine for model-based control

Reference 29

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:28:37.050856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:36.790121Z digest=sha256:a844f84fd80289353caf693b3b059926705115005cf57433401e9f977957b837

Observation b7625643-bee5-4928-9685-38c8dfbd0af1 · outbound

This paper cites an unresolved cited work.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:40.268447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:28:35.157057Z digest=sha256:07bc9d98b770e109ed6bd576aa7e063a9f00188480cf0c1ea1c396829c5001b2

Pith citing papers

Observation d9610706-0245-42e6-82d7-53dc9175e2f6 · inbound

Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning cites this paper.

Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning RLBenchNet: The Right Network for the Right Reinforcement Learning Task

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T05:05:07.908039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T05:05:07.831643Z digest=sha256:5a678f10c8ff9e117e094c1563b23d3d76ac92aef3473f3cdc62d23007e07267