Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:16.642210Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2501.18093.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:16.642210Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fca48c31-ca0f-4934-a54c-733cd16d1874 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Retroactive and graded prioritization of memory by reward
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 48c6e192-95b4-4a0f-b6ed-1bcd102d91c7 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Prioritized Sequence Experience Replay
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3811aabf-5dfe-456a-9459-b2a9cc30b594 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method OpenAI Gym
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dd4dd80-7c33-4c25-99ed-cda645d4ada6 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Addressing function approximation error in actor-critic methods
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768f7d28-d0d2-47e5-aebb-a65bb76b4821 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method An equivalence between loss functions and non-uniform sampling in experience replay
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ec438277-4960-48ce-bca8-f220d4e6c5e8 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method World Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3e764fd-596c-48db-bcb0-60326d916eed · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c850a82a-9523-4ad5-a77b-e9da44762105 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Deep reinforcement learning that matters
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4192335-3331-4262-92c4-c28f272b389a · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Rainbow: Combining improvements in deep reinforcement learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a14d55f6-644f-471e-b34e-4cd0758c2cae · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method How to train your robot with deep reinforcement learning: lessons we have learned
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fcc293e7-eebf-40cf-98e4-bded471f54b2 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Model-free and model-based reinforcement learning, the intersection of learning and planning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 336cf85c-be76-4f60-8eb8-1945baa95748 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Dopamine, updated: reward prediction error and beyond
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 99401917-524a-46aa-ac0b-99b3ab018dac · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Self-improving reactive agents based on reinforcement learning, planning and teaching
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 704908b2-4150-4858-84b0-743bc87943ba · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Human-level control through deep reinforcement learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcf946c6-347f-4fe7-9658-b92ddd7c673c · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Model-augmented prioritized experience replay
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dcd3e038-42c6-49ce-a9f3-4eac8812cc6c · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Curiosity-driven exploration by self-supervised prediction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d14543-ee11-418d-8717-fb49a14a6c94 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Building and breaking the chain: A model of reward prediction error integration and segmentation of memory
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3ae85502-4af5-49f4-b26b-210b6591a7a8 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Actor Prioritized Experience Replay
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 99068ee9-2732-44ff-a3fd-e1611fb5cf2d · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Prioritized Experience Replay
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e981b5a1-4960-479d-9263-597206e77da2 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Reward prediction error
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a3c8aa65-9bf4-4b97-bbb7-e00418ce551c · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Loss is its own Reward: Self-Supervision for Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c97f99-e63c-481d-90ae-0e58e21e326a · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Reward Prediction Error as an Exploration Objective in Deep RL
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5e3cd7c3-1e91-4a00-bab0-26d78690fdf5 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Mujoco: A physics engine for model-based control
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ad5cd5ab-ffaf-46ff-848a-3fd911c0a52e · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Experience Replay Optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6a105918-acd3-4fbf-aed2-097036a68bd8 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4181f3f-583b-4401-8f20-67187c000994 · outbound
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method write newline
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.