Pith. sign in

Paper Citation Record · LEDGER

RLBenchNet: The Right Network for the Right Reinforcement Learning Task

As of 7 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.15040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15040 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:36.790121Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:05:07.831643Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:05:07.890229Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6106fea-1de2-43bd-88c1-f064a50f103d · outbound

This paper cites What matters in on-policy reinforcement learning? a large-scale empirical study, 2020.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task What matters in on-policy reinforcement learning? a large-scale empirical study, 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:41.938153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:33.206676Z digest=sha256:4da036f5ea950023c8fea8c194c502c04dbf5ecff26c6499abda610c4db2dc63

Observation f99f2355-1bc8-4c2c-af04-2fad1097b93c · outbound

This paper cites an unresolved cited work.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:33.282105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:33.282105Z digest=sha256:09963008065e2bc497889434d53d65f65d0ba9fae107a09c0a6e23b7c6b1f70b

Observation 959058a0-368f-4807-9fe7-87f25a6ef6e0 · outbound

This paper cites an unresolved cited work.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:41.670666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:33.436858Z digest=sha256:b925b144f24a752e1699c61c721cdd3e65847950cc51af37dafd4915b2f5f525

Observation 0af9d966-68b8-4da4-9841-1f4664e5c7dc · outbound

This paper cites Kovalev, and Aleksandr I.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Kovalev, and Aleksandr I

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:41.438864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:33.551580Z digest=sha256:d6ce7754c51fa5000a4563a77c00f1b780d1eb197f83cdee85a89969efaa4a31

Observation f4302e45-ad40-4aaa-bf2f-d0656927aaa6 · outbound

This paper cites Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:41.211628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:33.668828Z digest=sha256:07a017a40c821c000c8399a88129aad7d53c4e2a72396d2969576aaa78f5845c

Observation 79cc64d2-2fff-4366-90b6-451c7b92622b · outbound

This paper cites Empirical evaluation of gated recurrent neural networks on sequence modeling, 2014.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Empirical evaluation of gated recurrent neural networks on sequence modeling, 2014

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:33.788457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:33.788457Z digest=sha256:146f43a5b519865e7805b16c1a226da93c1d5bc7d77f2e687275bc16d92b1f20

Observation cad80d92-8636-4b54-a4c6-8c27093eac20 · outbound

This paper cites Le, and Ruslan Salakhutdinov.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Le, and Ruslan Salakhutdinov

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:33.909597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:33.909597Z digest=sha256:be1dc501d4e1699b1f0baa5ce9b82d239292dc2fe427c21ebb4c9b5ae392d08d

Observation 25f3b796-ca9d-475f-b969-3f42124e8064 · outbound

This paper cites Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:40.997531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:34.079669Z digest=sha256:ed45ff74d247ff66a6349a11090a491892a98e04eb45085fd57f03a6c791c010

Observation ac5bf077-82ac-44d3-b52d-d5b77df93f78 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Addressing function approximation error in actor-critic methods

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:40.719291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:34.209058Z digest=sha256:33c188f13ca3104775cd9eb9be33bca12f4b3f2b3789ee28afe4380c826c0865

Observation 3cb30f82-13a3-476f-8bef-461584aee8f8 · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Mamba: Linear-time sequence modeling with selective state spaces, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.339686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.339686Z digest=sha256:e5efe14b04d4cc463218a47c568478317d52e4338d269a2776f4ac7553d3929c

Observation 79fc7973-4af8-420d-9a92-4a68c75c8cb0 · outbound

This paper cites Safe multi-agent reinforcement learning for multi-robot control.Artificial Intelligence, 319:103905, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Safe multi-agent reinforcement learning for multi-robot control.Artificial Intelligence, 319:103905, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.462727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.462727Z digest=sha256:1e0a9950eef51eabac750f19dc61930accfd838ac9a858b4ccba296ea510d06a

Observation 83232638-4bd6-41a4-84fc-00c1b75aece6 · outbound

This paper cites Safe and balanced: A framework for constrained multi-objective reinforcement learning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Safe and balanced: A framework for constrained multi-objective reinforcement learning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:40.474550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:34.598603Z digest=sha256:6c494b00ef72333502ce2d11de0a457937e2126c3863671dbd6d1f34f0b92c2d

Observation e4aa36a4-8f82-44c3-9459-8e32a412ba21 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.737787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.737787Z digest=sha256:018b6d1fd2bdc83d09efc6e30c37d85886a4a57cc24d4cb582c0fcb9f1b4b85f

Observation bba726a9-a79f-4f66-936a-e9f0be26ca87 · outbound

This paper cites Long short-term memory.Neural computation, 9(8):1735–1780, 1997.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Long short-term memory.Neural computation, 9(8):1735–1780, 1997

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.828624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.828624Z digest=sha256:730a358e3148b2d734d9bc19c85e7959fb47760277c2c0d3b0b43be13ff8b9c2

Observation d3e2b146-19c2-4334-954b-2e5eb1e22cff · outbound

This paper cites The 37 implementation details of proximal policy optimization.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task The 37 implementation details of proximal policy optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:34.979587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:34.979587Z digest=sha256:f88f1ae2fbf3d979c8633ba7657f50a2d94018407135f17348d937a91a3a3d99

Observation 996a4878-2a4d-4e5f-a1c8-5791e6ba4936 · outbound

This paper cites an unresolved cited work.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:39.966752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:35.265411Z digest=sha256:2d2625634d9de493aa4e4f92b5e909be97b029b8c31cdbee72ffed6a95ccfcb5

Observation c0013807-108f-4a30-ba82-095d37b8792d · outbound

This paper cites Decision mamba: Reinforcement learning via hybrid selective sequence modeling, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Decision mamba: Reinforcement learning via hybrid selective sequence modeling, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:39.702547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:35.380323Z digest=sha256:b11b3902671f7e027a4d16b6747c4806f8c6f633406131e44e9fa3a694a84f95

Observation 3ba1c3c0-f64a-49f6-a824-1acc9db8bb54 · outbound

This paper cites Du, and Huazhe Xu.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Du, and Huazhe Xu

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:39.471916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:35.513888Z digest=sha256:8cd37903528f50b2e0f2868577a21163c1e3c6e056d06982d2db1a6b82372570

Observation bbd87fdd-e57a-44b2-bedc-0657196f08e6 · outbound

This paper cites Efficient recurrent off-policy RL requires a context-encoder-specific learning rate.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Efficient recurrent off-policy RL requires a context-encoder-specific learning rate

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:39.170120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:35.656320Z digest=sha256:f687e2617107abaf565376fa2ee241f342dbaae13aec7f166ed3353b73ed827c

Observation 2a1c6f3f-f8ee-4dd5-b4c4-48f014ae176c · outbound

This paper cites Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl, 2025.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:38.809913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:35.846182Z digest=sha256:ec653a6f85aeae9e9070d14757cc740c3b686a2adc13de6d582d12241ee553b4

Observation d29e2f8f-3b87-4eb5-a936-b175e1663822 · outbound

This paper cites Popgym: Benchmarking partially observable reinforcement learning, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Popgym: Benchmarking partially observable reinforcement learning, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:38.519287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:35.950671Z digest=sha256:6f5b7baf141cbf8fdd9632a3c7d7cb9337c47fbf8ea7438006b465d79567b8fa

Observation b2b295c3-ef45-4214-90ae-3fa2b1792084 · outbound

This paper cites When do transformers shine in rl? decoupling memory from credit assignment, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task When do transformers shine in rl? decoupling memory from credit assignment, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:38.227986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:36.065512Z digest=sha256:17f4b20c2f49d1bd7e27c0d0b520742f0dcd6fd934e6a189e7fc9045c2e4079a

Observation 0b568af3-cf56-4d9a-8073-3086ca7d744a · outbound

This paper cites Decision mamba: Reinforcement learning via sequence modeling with selective state spaces, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Decision mamba: Reinforcement learning via sequence modeling with selective state spaces, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:37.933547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:36.196985Z digest=sha256:93e52abde6ee62d19f7589ceae331b114f6e9a9086eb313c5747e6b18a600899

Observation a6e0cc61-88c0-4881-b376-71a9166ddc2d · outbound

This paper cites Francis Song, Jack W.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Francis Song, Jack W

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:37.694539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:36.332756Z digest=sha256:0077f7752f5e2ab298c6d6b44bb875e3bebad13ab5936d1e382bcc07b211bdf8

Observation 896e1418-ee3c-4b2f-98cd-e6c2d2b0e9d9 · outbound

This paper cites Transformerxl as episodic memory in proximal policy optimization.Github Repository, 2023.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Transformerxl as episodic memory in proximal policy optimization.Github Repository, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:37.480182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:36.456294Z digest=sha256:1e07d8d83659f24d39ec1aa4e4f36326052e546658d940038c14ee5a92196d51

Observation 806735a1-f767-4017-89ef-f929645697fd · outbound

This paper cites Memory gym: Towards endless tasks to benchmark memory capabilities of agents, 2024.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Memory gym: Towards endless tasks to benchmark memory capabilities of agents, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:37.293317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:36.550031Z digest=sha256:f2a9e708e3668df747f445bd7c1c5b47182b5834bb3c3a91585131c26a09cdf0

Observation 094d1ed2-1d67-4e4e-b3d8-ac6d31fa2498 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Proximal policy optimization algorithms, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:36.638559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:36.638559Z digest=sha256:65150ace36c77e5e6007880ad844110d96ac594e94993e8dfd80e85cef99c35f

Observation ca03cca5-66c7-480e-b7cb-686f2195a7ca · outbound

This paper cites Mas- tering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Mas- tering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:36.723806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:36.723806Z digest=sha256:1fadc65bad377a8c4b0493931fcb3d0fbe1d4899849e3cede801e94e755a670e

Observation ed1e8e2f-396c-41eb-adba-8127e43dcef3 · outbound

This paper cites Mujoco: A physics engine for model-based control.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Mujoco: A physics engine for model-based control

Reference 29

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:28:37.050856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:36.790121Z digest=sha256:ddfd4c24516edb25b71ecb3c1c4d4051b7268bc522f901e1a79755a989bf33d8

Observation b7625643-bee5-4928-9685-38c8dfbd0af1 · outbound

This paper cites an unresolved cited work.

RLBenchNet: The Right Network for the Right Reinforcement Learning Task Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:40.268447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:28:35.157057Z digest=sha256:00ad6217b373e6ad6d11eb721d00c1aec780a876359181ac57dc68d0da61e9f7

Pith citing papers

Observation d9610706-0245-42e6-82d7-53dc9175e2f6 · inbound

Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning cites this paper.

Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning RLBenchNet: The Right Network for the Right Reinforcement Learning Task

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T05:05:07.908039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:05:07.831643Z digest=sha256:5ae7c7ef2b15e82b0a76d8e746123ebec8e33554efbd97162627c797edab8557