Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:48.526180Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2506.05968.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:48.526180Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 74e6c653-6a1b-4973-b4a1-9294f7e4089e · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a08e0af-1b8e-48e7-b92f-3e42aa666699 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning S., Courville, A., and Bellemare, M
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45079820-dde8-4551-8f97-e5219b05fda6 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b45f3e0e-0cd1-46ca-9229-5a4b4b04e45e · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1013eb96-90e9-44e8-9bf9-e923536a9d37 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning G., and Courville, A
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5e87e7c-7cf3-4f7a-acd6-c4e0c3a5e466 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Addressing function approximation error in actor-critic methods
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 615f40d1-1469-45de-88b2-87b8ae539466 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Extreme q-learning: Maxent RL without entropy
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f268b2a8-9a3c-48e5-ac1e-0112a52da0c5 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Reinforcement learning with deep energy-based policies
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ce08810-5d86-4fcd-880d-ab9041f99c4f · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de6e7585-9838-4414-8067-1fdabd457bb0 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning and Montana, G
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6880da0-a7c5-4aa6-86ec-43e00e352a7f · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Seizing serendipity: exploiting the value of past success in off-policy actor-critic
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56b0703c-6b37-4983-9bb8-31292def33a7 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Scalable deep reinforcement learning for vision-based robotic manipulation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a42b7053-01e2-453a-b0b1-4cd20fc06f7c · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Offline reinforcement learning with implicit q-learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52bab990-8632-44ea-8975-839eadde2fac · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Stabilizing off-policy q-learning via bootstrapping error reduction
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b0fc27d-a44b-4f6f-9c8d-a40993c67cb9 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Conservative q-learning for offline reinforcement learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02c3b8f7-5121-48ec-b9d4-4ba63752c3a9 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Maxmin q-learning: Controlling the estimation bias of q-learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 881e225e-6388-435d-a79a-5ca3fe2a34eb · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Continuous control with deep reinforcement learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6facc739-eb3e-4637-98f3-83ab11f51b1f · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Playing atari with deep reinforcement learning, 2013
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5d58713-2671-421c-a068-32ac12f65beb · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning A., Veness, J., Bellemare, M
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b359830d-c8cd-45a7-9fd2-ceabd3fa7bc2 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Curriculum dropout
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d6005e5-7c19-40d1-addb-4b852fb2fe78 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Stabilizing extreme q-learning by maclaurin expansion
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03340be4-19ca-4bad-aa9b-5a9dcd729d3f · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning and Niranjan, M
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab148013-765b-4f15-8fca-d6660c3ab6d3 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Trust region policy optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f6d060f-9686-4adb-a57a-768e76528ca4 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Proximal policy optimization algorithms, 2017
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c87f85e4-8b3d-4c0a-9f5b-cc00c061ec5c · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Solving continuous control via q-learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9f5f068-c203-4565-aeb8-405dca379fed · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Growing Q -networks: S olving continuous control tasks with adaptive control resolution
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 924d39d6-b358-436e-a5da-3c2b6a0cbbef · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Dual RL : Unification and new methods for reinforcement and imitation learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65d2a1ec-1373-4076-86ee-2f8c9a73238c · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Learning to predict by the method of temporal differences
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fa07a46-5611-4109-8c48-654dcdd26bf6 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac71ebbe-adae-4ae9-8bb1-1263ef7bbdc6 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Deepmind control suite, 2018
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50bf035f-377f-4013-907d-d15fe03e0952 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Action branching architectures for deep reinforcement learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 423310dc-48f6-40a0-a343-1e8d9b011de3 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning A deep hierarchical approach to lifelong learning in minecraft
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3a9e91-1812-4068-9ed5-f05dd61910fb · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning and Schwartz, A
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b510924-3e4f-40b8-9538-aa1f4b053952 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning dm\_control: Software and tasks for continuous control
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a28cff0-59c5-41b6-999e-9d67139228da · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43eeef58-3b37-41e5-9679-cb3fedc0bc5e · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0aa191d-f2e4-40fa-9ff7-402ae11b50b8 · outbound
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.