Pith. sign in

Paper Citation Record · LEDGER

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method

As of 22 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2501.18093.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18093 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:16.642210Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact4
  • verified fuzzy11
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fca48c31-ca0f-4934-a54c-733cd16d1874 · outbound

This paper cites Retroactive and graded prioritization of memory by reward.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Retroactive and graded prioritization of memory by reward

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.940322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.545411Z digest=sha256:b03e300c06c27172b88bacf8ade4e71611e957f009bd8f3d6894598e38cee775

Observation 48c6e192-95b4-4a0f-b6ed-1bcd102d91c7 · outbound

This paper cites Prioritized Sequence Experience Replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Prioritized Sequence Experience Replay

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:46:16.776638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.550779Z digest=sha256:d36123d3fd23592d50f59256442bb9cb43e92d27a7f493eec5c3051c23cdac10

Observation 3811aabf-5dfe-456a-9459-b2a9cc30b594 · outbound

This paper cites OpenAI Gym.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method OpenAI Gym

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.555177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.555177Z digest=sha256:287081c7023527ab6fd7fe81cb78a87d14eeba41a03bc02f0081c30087b3ac83

Observation 5dd4dd80-7c33-4c25-99ed-cda645d4ada6 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Addressing function approximation error in actor-critic methods

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.559700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.559700Z digest=sha256:a8c6b39f79a4760cc367d8278da57ca64784dda2011c1457bba3592ca94325b1

Observation 768f7d28-d0d2-47e5-aebb-a65bb76b4821 · outbound

This paper cites An equivalence between loss functions and non-uniform sampling in experience replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method An equivalence between loss functions and non-uniform sampling in experience replay

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.922804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.563926Z digest=sha256:3884d1c1a423aaa6372915e098da36d859255012ec4b2a29206b4249d872e6f1

Observation ec438277-4960-48ce-bca8-f220d4e6c5e8 · outbound

This paper cites World Models.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method World Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.567840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.567840Z digest=sha256:0976cbe999f022a5926307c50b501dfed647535ff65b5d367bee074e84bb333c

Observation e3e764fd-596c-48db-bcb0-60326d916eed · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.911071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.572472Z digest=sha256:88f013f971713a78eae69e2d08d37419a18b35dd2889f72206d77ac2b08d0477

Observation c850a82a-9523-4ad5-a77b-e9da44762105 · outbound

This paper cites Deep reinforcement learning that matters.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Deep reinforcement learning that matters

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.576292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.576292Z digest=sha256:970052b8f001321aa63449295923fc6f180d60a180f8cd60419c1746ed5e03e1

Observation c4192335-3331-4262-92c4-c28f272b389a · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Rainbow: Combining improvements in deep reinforcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.579886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.579886Z digest=sha256:0903ab1076e79bff379d041b40dea21069878ebf8cf69f1755a2f3168039cf4e

Observation a14d55f6-644f-471e-b34e-4cd0758c2cae · outbound

This paper cites How to train your robot with deep reinforcement learning: lessons we have learned.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method How to train your robot with deep reinforcement learning: lessons we have learned

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.886457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.583477Z digest=sha256:206a1e053c31745d29da9ca8365961c638dd904e02b0b82e1516f5451f30e1e5

Observation fcc293e7-eebf-40cf-98e4-bded471f54b2 · outbound

This paper cites Model-free and model-based reinforcement learning, the intersection of learning and planning.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Model-free and model-based reinforcement learning, the intersection of learning and planning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.875019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.587123Z digest=sha256:377def00cf7a357db1c2a1d619fc41265cedb75c181f687c18caac477a1adac8

Observation 336cf85c-be76-4f60-8eb8-1945baa95748 · outbound

This paper cites Dopamine, updated: reward prediction error and beyond.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Dopamine, updated: reward prediction error and beyond

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.863802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.590790Z digest=sha256:93137156b951bbe0b4880fb82b27aadd01709bc95ccdc76d3ba9bc651056b70b

Observation 99401917-524a-46aa-ac0b-99b3ab018dac · outbound

This paper cites Self-improving reactive agents based on reinforcement learning, planning and teaching.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Self-improving reactive agents based on reinforcement learning, planning and teaching

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.853092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.594290Z digest=sha256:342bcf379d2671334975eabf7a7aa90c9653f5ac4b70dffd9c37041593e6473c

Observation 704908b2-4150-4858-84b0-743bc87943ba · outbound

This paper cites Human-level control through deep reinforcement learning.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Human-level control through deep reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.597703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.597703Z digest=sha256:b714349505bbca3993730d8ca941236b360dee2f879aae7eae0b996592de4c49

Observation dcf946c6-347f-4fe7-9658-b92ddd7c673c · outbound

This paper cites Model-augmented prioritized experience replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Model-augmented prioritized experience replay

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.835866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.601132Z digest=sha256:39f0e652173ba603c23f4904ca9f66fec6b6370b116f531c77be9ea74cde765f

Observation dcd3e038-42c6-49ce-a9f3-4eac8812cc6c · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Curiosity-driven exploration by self-supervised prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.604656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.604656Z digest=sha256:738db815b86271c90adc96946984b01c57090440f3ad3e5227b7d8054cae0041

Observation e3d14543-ee11-418d-8717-fb49a14a6c94 · outbound

This paper cites Building and breaking the chain: A model of reward prediction error integration and segmentation of memory.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Building and breaking the chain: A model of reward prediction error integration and segmentation of memory

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.819053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.607985Z digest=sha256:d255029fc22728fa11e5ed3482d3eda6b16ec3af31b96982c3aaede81ecad333

Observation 3ae85502-4af5-49f4-b26b-210b6591a7a8 · outbound

This paper cites Actor Prioritized Experience Replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Actor Prioritized Experience Replay

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:46:16.740208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.611328Z digest=sha256:4a61ecd4ce61fec6e2932f39d2cfdc9b0ef8d14186cb243e3630d73299eeee03

Observation 99068ee9-2732-44ff-a3fd-e1611fb5cf2d · outbound

This paper cites Prioritized Experience Replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Prioritized Experience Replay

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.615256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.615256Z digest=sha256:66e32f6743b1ab382e62b7fd1608ed38fc16899f8bca161b7139a8b87458a3ca

Observation e981b5a1-4960-479d-9263-597206e77da2 · outbound

This paper cites Reward prediction error.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Reward prediction error

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.807147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.619442Z digest=sha256:1d5c1a7ff5f64c3abd67a330457f0ef20badc1a6bc750a39b930c0eebf127afd

Observation a3c8aa65-9bf4-4b97-bbb7-e00418ce551c · outbound

This paper cites Loss is its own Reward: Self-Supervision for Reinforcement Learning.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Loss is its own Reward: Self-Supervision for Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.622999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.622999Z digest=sha256:231d87edf5b250c40144828a600d90209e418f235cf44b3fb9dcfcb9f40fa3e9

Observation 46c97f99-e63c-481d-90ae-0e58e21e326a · outbound

This paper cites Reward Prediction Error as an Exploration Objective in Deep RL.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Reward Prediction Error as an Exploration Objective in Deep RL

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:46:16.705585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.627084Z digest=sha256:19d64881857ff3fdf53cdec3b6305e10f21db3d90157ac6a0a450cdf9c348c71

Observation 5e3cd7c3-1e91-4a00-bab0-26d78690fdf5 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Mujoco: A physics engine for model-based control

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.630766Z digest=sha256:b808701293ef9cec9f47b17c9211c5058186dddc954b9c505f6e6a08e9223a11

Observation ad5cd5ab-ffaf-46ff-848a-3fd911c0a52e · outbound

This paper cites Experience Replay Optimization.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Experience Replay Optimization

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:46:16.689312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.634365Z digest=sha256:d03dd4934512bb0f1e1dd03d09b89be29cc6fd65601a3192fb9c7623686331a8

Observation 6a105918-acd3-4fbf-aed2-097036a68bd8 · outbound

This paper cites Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.638153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.638153Z digest=sha256:8585f2c79c51557a4f89292e0f9a092adf351f70ef82e4a7102b3f0660e3c642

Observation d4181f3f-583b-4401-8f20-67187c000994 · outbound

This paper cites write newline.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method write newline

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.642210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.642210Z digest=sha256:f839590b9bc4bfc17f5aa86438583dfe9c52a983d31a4024937f33fc0c4b7e3b

Pith citing papers

No inbound Pith citation observations are available.