Pith. sign in

Paper Citation Record · LEDGER

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method

As of 22 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2501.18093.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18093 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:16.642210Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact4
  • verified fuzzy11
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fca48c31-ca0f-4934-a54c-733cd16d1874 · outbound

This paper cites Retroactive and graded prioritization of memory by reward.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Retroactive and graded prioritization of memory by reward

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.940322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.545411Z digest=sha256:da13f1efc2ceae2cc90590a7a6d32d5efd054dcf3b2767a576097a12820632ae

Observation 48c6e192-95b4-4a0f-b6ed-1bcd102d91c7 · outbound

This paper cites Prioritized Sequence Experience Replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Prioritized Sequence Experience Replay

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:46:16.776638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.550779Z digest=sha256:569b17b04f89b66730b5dd28f3d3e66779576468dbff2f24ef50caedd86a82fa

Observation 3811aabf-5dfe-456a-9459-b2a9cc30b594 · outbound

This paper cites OpenAI Gym.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method OpenAI Gym

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.555177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.555177Z digest=sha256:287081c7023527ab6fd7fe81cb78a87d14eeba41a03bc02f0081c30087b3ac83

Observation 5dd4dd80-7c33-4c25-99ed-cda645d4ada6 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Addressing function approximation error in actor-critic methods

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.559700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.559700Z digest=sha256:a8c6b39f79a4760cc367d8278da57ca64784dda2011c1457bba3592ca94325b1

Observation 768f7d28-d0d2-47e5-aebb-a65bb76b4821 · outbound

This paper cites An equivalence between loss functions and non-uniform sampling in experience replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method An equivalence between loss functions and non-uniform sampling in experience replay

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.922804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.563926Z digest=sha256:f3c6bfb2b19fe678a75dd624b21808c8e7e0d1108d5836dd8e90605281949d8a

Observation ec438277-4960-48ce-bca8-f220d4e6c5e8 · outbound

This paper cites World Models.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method World Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.567840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.567840Z digest=sha256:0976cbe999f022a5926307c50b501dfed647535ff65b5d367bee074e84bb333c

Observation e3e764fd-596c-48db-bcb0-60326d916eed · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.911071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.572472Z digest=sha256:e84cc2cb5e9b21e8be6dbfdc9c51210500facae7b4c5d170e4161449d10990a1

Observation c850a82a-9523-4ad5-a77b-e9da44762105 · outbound

This paper cites Deep reinforcement learning that matters.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Deep reinforcement learning that matters

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.576292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.576292Z digest=sha256:970052b8f001321aa63449295923fc6f180d60a180f8cd60419c1746ed5e03e1

Observation c4192335-3331-4262-92c4-c28f272b389a · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Rainbow: Combining improvements in deep reinforcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.579886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.579886Z digest=sha256:0903ab1076e79bff379d041b40dea21069878ebf8cf69f1755a2f3168039cf4e

Observation a14d55f6-644f-471e-b34e-4cd0758c2cae · outbound

This paper cites How to train your robot with deep reinforcement learning: lessons we have learned.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method How to train your robot with deep reinforcement learning: lessons we have learned

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.886457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.583477Z digest=sha256:98be602e06f1879eb2306b91ee4ec3415dce9e926db86a6bb11486e9b3e6c9be

Observation fcc293e7-eebf-40cf-98e4-bded471f54b2 · outbound

This paper cites Model-free and model-based reinforcement learning, the intersection of learning and planning.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Model-free and model-based reinforcement learning, the intersection of learning and planning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.875019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.587123Z digest=sha256:3b517599c7a2f532b1938cd182bfc02d87935f29928b7258a9bcdf08ff32df4a

Observation 336cf85c-be76-4f60-8eb8-1945baa95748 · outbound

This paper cites Dopamine, updated: reward prediction error and beyond.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Dopamine, updated: reward prediction error and beyond

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.863802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.590790Z digest=sha256:d8cdeee563f064767af2d0340a5153fc6e3f84e456e3231636888f9c0a76de47

Observation 99401917-524a-46aa-ac0b-99b3ab018dac · outbound

This paper cites Self-improving reactive agents based on reinforcement learning, planning and teaching.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Self-improving reactive agents based on reinforcement learning, planning and teaching

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.853092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.594290Z digest=sha256:0c08246192a6c6a2c1b0136f57d3ac901fad420a9f7d98d2e9b4187c00f234e7

Observation 704908b2-4150-4858-84b0-743bc87943ba · outbound

This paper cites Human-level control through deep reinforcement learning.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Human-level control through deep reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.597703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.597703Z digest=sha256:b714349505bbca3993730d8ca941236b360dee2f879aae7eae0b996592de4c49

Observation dcf946c6-347f-4fe7-9658-b92ddd7c673c · outbound

This paper cites Model-augmented prioritized experience replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Model-augmented prioritized experience replay

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.835866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.601132Z digest=sha256:5634b423cf3df190ed228e7a6c47e6dafb10ae678ea356625ef9922e0d8fa5a8

Observation dcd3e038-42c6-49ce-a9f3-4eac8812cc6c · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Curiosity-driven exploration by self-supervised prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.604656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.604656Z digest=sha256:738db815b86271c90adc96946984b01c57090440f3ad3e5227b7d8054cae0041

Observation e3d14543-ee11-418d-8717-fb49a14a6c94 · outbound

This paper cites Building and breaking the chain: A model of reward prediction error integration and segmentation of memory.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Building and breaking the chain: A model of reward prediction error integration and segmentation of memory

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.819053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.607985Z digest=sha256:15c5e3e741f02b45e1d397597176a273bbd48f4b840650a6e183a43c6a27f79a

Observation 3ae85502-4af5-49f4-b26b-210b6591a7a8 · outbound

This paper cites Actor Prioritized Experience Replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Actor Prioritized Experience Replay

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:46:16.740208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.611328Z digest=sha256:037971b95279c0cce213c9ade1fcb2aa77fe8ed7f1d0574816876963cc2d6792

Observation 99068ee9-2732-44ff-a3fd-e1611fb5cf2d · outbound

This paper cites Prioritized Experience Replay.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Prioritized Experience Replay

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.615256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.615256Z digest=sha256:66e32f6743b1ab382e62b7fd1608ed38fc16899f8bca161b7139a8b87458a3ca

Observation e981b5a1-4960-479d-9263-597206e77da2 · outbound

This paper cites Reward prediction error.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Reward prediction error

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.807147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.619442Z digest=sha256:e390e2eef0f41a04fbaf640be129fc0fc6bf02ff10a3b66ad88011ee5384d8d1

Observation a3c8aa65-9bf4-4b97-bbb7-e00418ce551c · outbound

This paper cites Loss is its own Reward: Self-Supervision for Reinforcement Learning.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Loss is its own Reward: Self-Supervision for Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.622999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.622999Z digest=sha256:231d87edf5b250c40144828a600d90209e418f235cf44b3fb9dcfcb9f40fa3e9

Observation 46c97f99-e63c-481d-90ae-0e58e21e326a · outbound

This paper cites Reward Prediction Error as an Exploration Objective in Deep RL.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Reward Prediction Error as an Exploration Objective in Deep RL

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:46:16.705585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.627084Z digest=sha256:c7f569ab555bcc5ceb7531837e6a45aadf18b1623fe6ad7c9503c045d973204d

Observation 5e3cd7c3-1e91-4a00-bab0-26d78690fdf5 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Mujoco: A physics engine for model-based control

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:16.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.630766Z digest=sha256:2cecd3b323149d4e2e5359796cea455a4eef3754a880ff826626ce20a2df8334

Observation ad5cd5ab-ffaf-46ff-848a-3fd911c0a52e · outbound

This paper cites Experience Replay Optimization.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Experience Replay Optimization

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:46:16.689312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-10T00:46:16.634365Z digest=sha256:bbb93fc35d38706d919fde639b97a8c2178318a26fa83d9583fda7d16ac7bc1c

Observation 6a105918-acd3-4fbf-aed2-097036a68bd8 · outbound

This paper cites Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.638153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.638153Z digest=sha256:8585f2c79c51557a4f89292e0f9a092adf351f70ef82e4a7102b3f0660e3c642

Observation d4181f3f-583b-4401-8f20-67187c000994 · outbound

This paper cites write newline.

Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method write newline

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:16.642210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:16.642210Z digest=sha256:f839590b9bc4bfc17f5aa86438583dfe9c52a983d31a4024937f33fc0c4b7e3b

Pith citing papers

No inbound Pith citation observations are available.