Pith. sign in

Paper Citation Record · LEDGER

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2506.05968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05968 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:48.526180Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 74e6c653-6a1b-4973-b4a1-9294f7e4089e · outbound

This paper cites write newline.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:45.614896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:45.614896Z digest=sha256:37aa4b9eceec589084116349c9f640ad2dac73419fc5134e32a9279d666a58f3

Observation 5a08e0af-1b8e-48e7-b92f-3e42aa666699 · outbound

This paper cites S., Courville, A., and Bellemare, M.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning S., Courville, A., and Bellemare, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:53.429166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:45.701112Z digest=sha256:e952d9141d7a31d9f05251b2037eca12947059fc928213d40fb44162b9d6744d

Observation 45079820-dde8-4551-8f97-e5219b05fda6 · outbound

This paper cites P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:53.179191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:45.756172Z digest=sha256:4f70aa81e360a359485a21d7649834acd9d8795679ba5f37810f1137ad3cae66

Observation b45f3e0e-0cd1-46ca-9229-5a4b4b04e45e · outbound

This paper cites an unresolved cited work.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:52.930988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:45.842821Z digest=sha256:e8d9e2df6ef14e4168efbd95c6cee1a431142bf953c24c45daf0c74205c1d917

Observation 1013eb96-90e9-44e8-9bf9-e923536a9d37 · outbound

This paper cites G., and Courville, A.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning G., and Courville, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:52.768549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:45.953313Z digest=sha256:3f9644e9ed50681946d492c044a65001775a6a0c13546d1432e294ba5398587c

Observation c5e87e7c-7cf3-4f7a-acd6-c4e0c3a5e466 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Addressing function approximation error in actor-critic methods

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:52.533275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.036985Z digest=sha256:5ba3d75e4358ffc967775ea631a5502043a52909bf48bfc93db19f30c8665bf0

Observation 615f40d1-1469-45de-88b2-87b8ae539466 · outbound

This paper cites Extreme q-learning: Maxent RL without entropy.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Extreme q-learning: Maxent RL without entropy

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:52.217667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.160570Z digest=sha256:dd859ed9c88ca26a80d6f31ebff4d1777273c4431598838b5ba0bfb3d091add5

Observation f268b2a8-9a3c-48e5-ac1e-0112a52da0c5 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Reinforcement learning with deep energy-based policies

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:51.890421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.253819Z digest=sha256:ba9bf1a97f59cbbfc33321e9d0a8350d7679cc60051a0b213806285cdb5b933f

Observation 5ce08810-5d86-4fcd-880d-ab9041f99c4f · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:51.632867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.332706Z digest=sha256:f51ab32af495387bcff72a6d8ba967d5601c1846a73591094af1479812972ec0

Observation de6e7585-9838-4414-8067-1fdabd457bb0 · outbound

This paper cites and Montana, G.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning and Montana, G

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:51.407864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.423189Z digest=sha256:ccae9dee573aec79bb438e120a1b019538fd5dbd4aa39c9e912aab8b97ef1a3f

Observation c6880da0-a7c5-4aa6-86ec-43e00e352a7f · outbound

This paper cites Seizing serendipity: exploiting the value of past success in off-policy actor-critic.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Seizing serendipity: exploiting the value of past success in off-policy actor-critic

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:51.117765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.499898Z digest=sha256:b394bb185d67dacf9de02a8491a9c35fd470cd0abbd2815b4e7b961531205317

Observation 56b0703c-6b37-4983-9bb8-31292def33a7 · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:50.905539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.565608Z digest=sha256:ed784cfdc43cdb145834de1e92346828789bab6b6adb4bc1200a45fd0130785c

Observation a42b7053-01e2-453a-b0b1-4cd20fc06f7c · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Offline reinforcement learning with implicit q-learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:46.634948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:46.634948Z digest=sha256:75f46806638ce8c38db2b64c640701f13717bbada10a637a1c3d14f76b056d0e

Observation 52bab990-8632-44ea-8975-839eadde2fac · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:50.694062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.698273Z digest=sha256:20c1400aa94fa454f71c479b8c429cb825728934047659952488b173c37c3633

Observation 3b0fc27d-a44b-4f6f-9c8d-a40993c67cb9 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Conservative q-learning for offline reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:50.480294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.770953Z digest=sha256:9193b4ab88b5f686637582f1f3705d6770a8a70e58bcc6e0b54de790c7138b98

Observation 02c3b8f7-5121-48ec-b9d4-4ba63752c3a9 · outbound

This paper cites Maxmin q-learning: Controlling the estimation bias of q-learning.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Maxmin q-learning: Controlling the estimation bias of q-learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:50.243561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:46.831857Z digest=sha256:2ba17d4b37b41a9f408364a607a25af8fdc98a4e642d40f998a3c3520c534a3c

Observation 881e225e-6388-435d-a79a-5ca3fe2a34eb · outbound

This paper cites Continuous control with deep reinforcement learning.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Continuous control with deep reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:46.934739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:46.934739Z digest=sha256:0fbc4a5e58fabb5341d24eec1dbf659d1609e26977bf3bdba6f09f366c39cc36

Observation 6facc739-eb3e-4637-98f3-83ab11f51b1f · outbound

This paper cites Playing atari with deep reinforcement learning, 2013.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Playing atari with deep reinforcement learning, 2013

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:47.029223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:47.029223Z digest=sha256:ccfc02d198fea1f5e629b149ffa138fbcddbe0d49809ac41a9a5a2106aea440c

Observation b5d58713-2671-421c-a068-32ac12f65beb · outbound

This paper cites A., Veness, J., Bellemare, M.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning A., Veness, J., Bellemare, M

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:47.140492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:47.140492Z digest=sha256:76e574dd6341c7b8f9e99d27b6154d9aa4e17ed6653146a7d66d6cb9e1f65589

Observation b359830d-c8cd-45a7-9fd2-ceabd3fa7bc2 · outbound

This paper cites Curriculum dropout.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Curriculum dropout

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:50.064348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:47.193512Z digest=sha256:cacd68129c67fae3143568ff9f134d08158a1e2c36c8ae75d70a7377e1650ae7

Observation 3d6005e5-7c19-40d1-addb-4b852fb2fe78 · outbound

This paper cites Stabilizing extreme q-learning by maclaurin expansion.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Stabilizing extreme q-learning by maclaurin expansion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:49.970652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:47.263505Z digest=sha256:ba47bcf253ab7242b5d27b925844977759b386c1c1b7be9fd6c5c7af80af9647

Observation 03340be4-19ca-4bad-aa9b-5a9dcd729d3f · outbound

This paper cites and Niranjan, M.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning and Niranjan, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:49.866604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:47.342525Z digest=sha256:0cb0e066b8b55c9b0082cc862e23533a2382986c666abac7aaef5eca38c9de08

Observation ab148013-765b-4f15-8fca-d6660c3ab6d3 · outbound

This paper cites Trust region policy optimization.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Trust region policy optimization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:49.776975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:47.409151Z digest=sha256:0193ea467a17e9aed10c3ed0da80f87235cd58aa6a495f83e3954bd54bab8598

Observation 7f6d060f-9686-4adb-a57a-768e76528ca4 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Proximal policy optimization algorithms, 2017

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:47.463138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:47.463138Z digest=sha256:3694cfba8fc192c2535926b5686b7163fe9650c4eb43659f68da6fc2ae81300f

Observation c87f85e4-8b3d-4c0a-9f5b-cc00c061ec5c · outbound

This paper cites Solving continuous control via q-learning.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Solving continuous control via q-learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:49.667511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:47.567098Z digest=sha256:73f4b7da927cb700410db8fedd92e9342f64eae17229e9299d23aef936b3a22d

Observation e9f5f068-c203-4565-aeb8-405dca379fed · outbound

This paper cites Growing Q -networks: S olving continuous control tasks with adaptive control resolution.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Growing Q -networks: S olving continuous control tasks with adaptive control resolution

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:49.559832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:47.665258Z digest=sha256:661c8382855b492a67b03c19abfb6f3fe21a6e4ff7637552587ddde1818e9b0c

Observation 924d39d6-b358-436e-a5da-3c2b6a0cbbef · outbound

This paper cites Dual RL : Unification and new methods for reinforcement and imitation learning.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Dual RL : Unification and new methods for reinforcement and imitation learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:49.478106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:47.739882Z digest=sha256:cdaea114826209b500525dbb518f441a70f4cc0815f7c7f65fff7de96ff2e239

Observation 65d2a1ec-1373-4076-86ee-2f8c9a73238c · outbound

This paper cites Learning to predict by the method of temporal differences.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Learning to predict by the method of temporal differences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:47.818822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:47.818822Z digest=sha256:efe251a4c76baf99da7e29afb6dbaf6ab63838b0651967cf481f4489472c4ee2

Observation 1fa07a46-5611-4109-8c48-654dcdd26bf6 · outbound

This paper cites an unresolved cited work.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:47.874160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:47.874160Z digest=sha256:665d867fb3d8767fb4f81d5ca5423e07dfeb0affb854e6ae19dc1d0c6958f538

Observation ac71ebbe-adae-4ae9-8bb1-1263ef7bbdc6 · outbound

This paper cites Deepmind control suite, 2018.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Deepmind control suite, 2018

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:49.324942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:47.949840Z digest=sha256:d33d698f6c5519398c7653325b227b81754f8e706a65882cd562fdef35e7da0d

Observation 50bf035f-377f-4013-907d-d15fe03e0952 · outbound

This paper cites Action branching architectures for deep reinforcement learning.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Action branching architectures for deep reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:49.202948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:48.022560Z digest=sha256:86eaa329d8a78f5dae230604d95723e06b4e4d8cb7ef53ce474322905aa91c5a

Observation 423310dc-48f6-40a0-a343-1e8d9b011de3 · outbound

This paper cites A deep hierarchical approach to lifelong learning in minecraft.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning A deep hierarchical approach to lifelong learning in minecraft

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:48.122721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:48.122721Z digest=sha256:7b53b146069db27564941c9998843a321348e3d848d129ef13d3f70575b8a414

Observation 0e3a9e91-1812-4068-9ed5-f05dd61910fb · outbound

This paper cites and Schwartz, A.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning and Schwartz, A

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:49.108975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:48.202504Z digest=sha256:a626e605fcf1df9afce372474cb617bac6257b1981ea148d42fc08da2cac8654

Observation 7b510924-3e4f-40b8-9538-aa1f4b053952 · outbound

This paper cites dm\_control: Software and tasks for continuous control.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning dm\_control: Software and tasks for continuous control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:48.288319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:48.288319Z digest=sha256:b9f518500f674cba6bf42316e8286e065b3c882b9e0659201fd69519c5350c30

Observation 7a28cff0-59c5-41b6-999e-9d67139228da · outbound

This paper cites an unresolved cited work.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:48.975376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:48.364793Z digest=sha256:ff26ac04e66213ce07b81e4506dad6c98146a16abc64443f1ea68f709af6a7fd

Observation 43eeef58-3b37-41e5-9679-cb3fedc0bc5e · outbound

This paper cites an unresolved cited work.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:48.449376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:48.449376Z digest=sha256:3e6811167ae7963a8cb67a45107c4b955bb6879106c8596cf6d9aa9d6f8c22e3

Observation a0aa191d-f2e4-40fa-9ff7-402ae11b50b8 · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:48.813577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:19:48.526180Z digest=sha256:b3ee475cd2084a6bae3c1dc59889bc3b2a68a29dd6d3d677ee7631604f286840

Pith citing papers

No inbound Pith citation observations are available.