Pith. sign in

Paper Citation Record · LEDGER

Continual Reinforcement Learning by Planning with Online World Models

As of 21 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 3 inbound Pith citation observations for arXiv:2507.09177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09177 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:10:19.203780Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T19:24:48.899301Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:26.615794Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bcef066c-562c-4fe1-ba23-6bce07c6075a · outbound

This paper cites write newline.

Continual Reinforcement Learning by Planning with Online World Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:15.146099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:10:15.146099Z digest=sha256:bea7cf54489d60947fd1c04e5e59c8a12781d307b8c7bc1ca878415702cc7a12

Observation a1428363-22f9-41ee-83ee-3a81a56f7c4b · outbound

This paper cites an unresolved cited work.

Continual Reinforcement Learning by Planning with Online World Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:11:19.973952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.170221Z digest=sha256:e81ee04ddea119c7601c7b0cbb950bce1633bf590c3bccff0828765e92b013ae

Observation 3bef8d88-881f-4838-90c0-b3b628cda6a7 · outbound

This paper cites P., and Singh, S.

Continual Reinforcement Learning by Planning with Online World Models P., and Singh, S

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:19.808082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.203364Z digest=sha256:526bc369f3b91e84992fa2ca41062405e5403287cbfcb27269ee9a68eed65694

Observation 222e613d-4b87-4f7a-b59f-0a43e2adcbf6 · outbound

This paper cites Gradient based sample selection for online continual learning.

Continual Reinforcement Learning by Planning with Online World Models Gradient based sample selection for online continual learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:19.663322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.245228Z digest=sha256:5da09b4fd08368c75acdb79d6a0de72515d5a77d20b3f73f5686685c1fde98fa

Observation f52bf769-aa2d-482c-a0bd-033c5c328056 · outbound

This paper cites Selfless sequential learning.

Continual Reinforcement Learning by Planning with Online World Models Selfless sequential learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:19.517144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.287952Z digest=sha256:0bc4240bbbcb8cd919acdbc6c2dcf7875fc23328f2a7de4181167e7efbb1759d

Observation b403af44-f0e8-4aba-9ef5-e21207d7242d · outbound

This paper cites G., Naddaf, Y., Veness, J., and Bowling, M.

Continual Reinforcement Learning by Planning with Online World Models G., Naddaf, Y., Veness, J., and Bowling, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:19.420054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.322815Z digest=sha256:997d8d3e253f42dbd21c8758901ad35b70923284abfdd73fe3ae81b99905bb4a

Observation 7654f282-6daa-4a29-91e2-68d2b5fbcb0a · outbound

This paper cites A markovian decision process.

Continual Reinforcement Learning by Planning with Online World Models A markovian decision process

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:19.223912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.362639Z digest=sha256:195a2e2794b397e5d0dfe826d084cf3f0bc570ae16cec4837cfee57e8ffef7dc

Observation a01e71f1-fe0c-45d4-b2da-790ffd11e411 · outbound

This paper cites Class-incremental continual learning into the extended der-verse.

Continual Reinforcement Learning by Planning with Online World Models Class-incremental continual learning into the extended der-verse

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:19.107268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.401256Z digest=sha256:64c426dba2c41ea43dada31e58832b07b3af6fae5dcbba23bb8c5d2387cf6ef7

Observation 24baf1fd-cd45-4741-a602-7956ec43e93a · outbound

This paper cites Efficient lifelong learning with A-GEM.

Continual Reinforcement Learning by Planning with Online World Models Efficient lifelong learning with A-GEM

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:18.941144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.438261Z digest=sha256:18a13e8a0f199320c85e577bb2c12a3afd1781be19871ed75e12b186d385f312

Observation b9e7680b-8b1b-4a3e-80a5-e88a5cc8ec24 · outbound

This paper cites Continual learning with tiny episodic memories.

Continual Reinforcement Learning by Planning with Online World Models Continual learning with tiny episodic memories

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:18.833051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.475347Z digest=sha256:09c2e196ed075815fcef86d98791545eade3b70f5db80b0e139e11ff63d1aa4e

Observation 03b9ab8f-27f1-427e-8697-3abb3aa88aa7 · outbound

This paper cites Deep reinforcement learning in a handful of trials using probabilistic dynamics models.

Continual Reinforcement Learning by Planning with Online World Models Deep reinforcement learning in a handful of trials using probabilistic dynamics models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:15.513110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:10:15.513110Z digest=sha256:e89e41e019a0106ca62f657897c62f928fa0da8631a9be9484b36431545fe15b

Observation a86586b8-8ace-452e-a33f-ca8228b5ea28 · outbound

This paper cites Leveraging procedural generation to benchmark reinforcement learning.

Continual Reinforcement Learning by Planning with Online World Models Leveraging procedural generation to benchmark reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:18.739423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.571933Z digest=sha256:b881c6a9cf7f4864641c74e5681e6eff3720295ec1c9958364e6c49e390eefd6

Observation c8488fd5-6762-45da-8876-c217ef5c9211 · outbound

This paper cites P., Mannor, S., and Rubinstein, R.

Continual Reinforcement Learning by Planning with Online World Models P., Mannor, S., and Rubinstein, R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:18.641746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.605091Z digest=sha256:70f44c6a505f393e2f4c9aba64b4f63769d14f466496280c91d0b1d7c4a9d343

Observation bba66180-0264-4d0a-853b-d3b57dd244d4 · outbound

This paper cites Orthogonal gradient descent for continual learning.

Continual Reinforcement Learning by Planning with Online World Models Orthogonal gradient descent for continual learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:18.533855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.646313Z digest=sha256:f50413ae79819e1e7bab158dae7ab059a43d08dfac922835cba7d6a10f6fc34d

Observation 4806d2a2-973d-4563-b189-dec4109d1ea3 · outbound

This paper cites E., Prett, D.

Continual Reinforcement Learning by Planning with Online World Models E., Prett, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:18.440369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.774071Z digest=sha256:8388cce0ef33b33137572257fdf496947d9e4f4fcf5fd9339136b42fc15db114

Observation 735f65e6-ddd4-446b-8ada-4a9df7c2d61e · outbound

This paper cites an unresolved cited work.

Continual Reinforcement Learning by Planning with Online World Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:11:18.295080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.828706Z digest=sha256:1a62f8a5ca21d1939443f4112155decc2330a17726ff896b910d547909ea0a9f

Observation 87b97f19-ff11-477c-bdb8-802b38c5f9bc · outbound

This paper cites Building a subspace of policies for scalable continual learning.

Continual Reinforcement Learning by Planning with Online World Models Building a subspace of policies for scalable continual learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:18.221087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.884057Z digest=sha256:934ae1a06442428fde707c0ca14a6a244f87f49f9fad8379b30e8a03558f9f91

Observation 7dd5b270-61f1-4a4e-ac3b-cbff40db3651 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Continual Reinforcement Learning by Planning with Online World Models Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:18.077729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.940751Z digest=sha256:6ae27ba976ab44e5f92a51b005fcbd8ab42bf30d5050c954503507174037ac72

Observation 2c32fd81-cccc-402b-a755-74a42ac606f3 · outbound

This paper cites K., et al.

Continual Reinforcement Learning by Planning with Online World Models K., et al

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:17.934115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:15.997879Z digest=sha256:32b4f671686b815d889990f77885265f20172f21cfd1c41c1b007e4600d49543

Observation 7432bbca-23d6-4717-845d-783972b524e0 · outbound

This paper cites Continual model-based reinforcement learning with hypernetworks.

Continual Reinforcement Learning by Planning with Online World Models Continual model-based reinforcement learning with hypernetworks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:17.739091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.074901Z digest=sha256:6c398365b72745e9457c6243ae409ae261dd5476ccb71ce639a88695ee5a8897

Observation 1231d891-ef7e-40fc-ba0f-6b350745ef2d · outbound

This paper cites A theory of universal artificial intelligence based on algorithmic complexity.

Continual Reinforcement Learning by Planning with Online World Models A theory of universal artificial intelligence based on algorithmic complexity

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:17.626556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.168793Z digest=sha256:7658f9400e9d639b478d9a9feaa99fe526da47b8dbffaf905b6e8ca09c20fbae

Observation d0b8e610-57df-4736-9afd-ba62473f840c · outbound

This paper cites and Cosgun, A.

Continual Reinforcement Learning by Planning with Online World Models and Cosgun, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:17.518041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.231032Z digest=sha256:77d7c989d5421437f0853ccbd8c34f3e61c47af3cc53a8cd417326c212891c5b

Observation d9d65a93-6dde-4dc3-9a1a-b8862b131f64 · outbound

This paper cites an unresolved cited work.

Continual Reinforcement Learning by Planning with Online World Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:11:17.394837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.306480Z digest=sha256:a0f4181cde3e81361b230ebf1c8251e5b861cf281fd269da919a82ffd44daf57

Observation fb3d6e27-a890-42f3-8b73-d412c66b0039 · outbound

This paper cites R., Hwang, S.

Continual Reinforcement Learning by Planning with Online World Models R., Hwang, S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:17.252092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.382778Z digest=sha256:32916c478578a1c8eaeb600b01baeadf63b099c53f4ba71ae6bdf98571bac562

Observation af2c2f29-c797-4ed1-8bf8-0865b9f59b3d · outbound

This paper cites J., Zohren, S., and Roberts, S.

Continual Reinforcement Learning by Planning with Online World Models J., Zohren, S., and Roberts, S

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:17.162487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.450734Z digest=sha256:94e9c8bafa28ad372b0a486431c6f9e0b94b58259ba9ab7b0574949e3d9c1f58

Observation aa13dbb3-1635-41ec-a7bc-caeb50cfb20b · outbound

This paper cites an unresolved cited work.

Continual Reinforcement Learning by Planning with Online World Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:11:17.129523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.525313Z digest=sha256:d058fa484bc9f20ecfb41630e0a38ce89cdfa3a1fbab9f3b750454afd788e067

Observation 76988404-6cc3-4711-bfd3-8f78e750082a · outbound

This paper cites Towards continual reinforcement learning: A review and perspectives.

Continual Reinforcement Learning by Planning with Online World Models Towards continual reinforcement learning: A review and perspectives

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:16.999266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.635145Z digest=sha256:c766241c6cf3d5f67b5ea09ead107d10c630b64f591cadb729432fe34b6df4cd

Observation 7d2667ea-0b44-435a-8f5e-51689b3a4949 · outbound

This paper cites an unresolved cited work.

Continual Reinforcement Learning by Planning with Online World Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:16.731665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:10:16.731665Z digest=sha256:c7eff46bb6c455f87a166d8bef090bbd903a5e5bba2ec347f456c46c5e9f4c42

Observation 87b20be5-795c-4899-8e7d-7adad73193e9 · outbound

This paper cites A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al.

Continual Reinforcement Learning by Planning with Online World Models A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:16.906941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.781120Z digest=sha256:cbbccd84ffbfed0d0627d91f8b94692589fafb9f61cfaad47e55820d471e6a54

Observation dd1b8742-3c0b-4850-ac32-dc19d36c1a78 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Continual Reinforcement Learning by Planning with Online World Models AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:16.832020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:10:16.832020Z digest=sha256:758a1835d1cbef19f1d6200ed3e7917db4d422ddac408f6bad2f9c8350f2ae2d

Observation edd5aaa1-ab79-4980-92a4-8921934cbc10 · outbound

This paper cites u ttler, H., Nardelli, N., Miller, A., Raileanu, R., Selvatici, M., Grefenstette, E., and Rockt \.

Continual Reinforcement Learning by Planning with Online World Models u ttler, H., Nardelli, N., Miller, A., Raileanu, R., Selvatici, M., Grefenstette, E., and Rockt \

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:16.801382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.889860Z digest=sha256:1cb5c36ff98909c7c0bf8ae04773869502bb2a7dcb48e91ef46d5f2c5b085bf0

Observation e45eaaae-6a69-490a-a0f2-98975964f710 · outbound

This paper cites S., and Lin, M.

Continual Reinforcement Learning by Planning with Online World Models S., and Lin, M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:16.658919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:16.964447Z digest=sha256:f3c5e0e500a04189650279f3630aa436fb760d6e83e6bc2d571f391572d2e2bd

Observation 51524395-bfbd-4be1-ae87-baac32b95709 · outbound

This paper cites and Lazebnik, S.

Continual Reinforcement Learning by Planning with Online World Models and Lazebnik, S

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:16.471864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.095740Z digest=sha256:1ae0875aaf590574b643b532b0fffa7937a5143e9a1836dfb9e3b173fa25813b

Observation 9ce9e3f1-959b-482b-a229-0fe5e5776b5a · outbound

This paper cites Deep online learning via meta-learning: Continual adaptation for model-based rl.

Continual Reinforcement Learning by Planning with Online World Models Deep online learning via meta-learning: Continual adaptation for model-based rl

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:16.353061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.212292Z digest=sha256:d78d369370301c24085e0e92209ea1047da8c411b9cb1aecce7c10ce31c93b9f

Observation d2104027-d851-4333-ba15-e566b568854d · outbound

This paper cites R., De Schutter, B., Wiering, M.

Continual Reinforcement Learning by Planning with Online World Models R., De Schutter, B., Wiering, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:16.259082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.298275Z digest=sha256:c5fd05f3b757f9ed977146aae15f51585ec59a3ce22d33c7dbcd6dcb657e993a

Observation d0299dec-4fab-402a-b3e2-d175b8d33653 · outbound

This paper cites and Vidal, R.

Continual Reinforcement Learning by Planning with Online World Models and Vidal, R

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:16.145695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.373685Z digest=sha256:eb4c0ebb7cccedba25ae2709dd63db02809315a8a0cde50a93d8688b8b516f29

Observation 3884e2c7-1e0c-4e97-b4c0-114dd928cec7 · outbound

This paper cites V., and Vidal, R.

Continual Reinforcement Learning by Planning with Online World Models V., and Vidal, R

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.968542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.440877Z digest=sha256:843aa4accc6831cd4b87a0c477e9dd94ac61aadf8ee6642953550934177e1e33

Observation 477a1953-439a-4b86-8e8d-a78d2bbf0d10 · outbound

This paper cites O., and Calandra, R.

Continual Reinforcement Learning by Planning with Online World Models O., and Calandra, R

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.857791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.529574Z digest=sha256:e67fa173a43fc6d536b04796b6bec2385a91df4bc96c0a9b66b12f0bfd611981

Observation 70a7bfb6-2d6a-4799-8535-c0e3911b44bb · outbound

This paper cites Sample-efficient cross-entropy method for real-time planning.

Continual Reinforcement Learning by Planning with Online World Models Sample-efficient cross-entropy method for real-time planning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.756384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.625496Z digest=sha256:081f4695a75587dcbeb4a97437fbd40a1a9a224eed9cd1d5897fdf10742ff071

Observation ea16f5d9-736c-4f97-8771-f2d6264ad600 · outbound

This paper cites Cora: Benchmarks, baselines, and metrics as a platform for continual reinforcement learning agents.

Continual Reinforcement Learning by Planning with Online World Models Cora: Benchmarks, baselines, and metrics as a platform for continual reinforcement learning agents

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.616354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.698008Z digest=sha256:d9f6cc7abc2e1c71d5dd027a3b33624c4680bdbc65fae9a6512ccd130dd1d40a

Observation d4f7d7f4-8338-42c5-a0cb-1f810958a38e · outbound

This paper cites Learning to learn without forgetting by maximizing transfer and minimizing interference.

Continual Reinforcement Learning by Planning with Online World Models Learning to learn without forgetting by maximizing transfer and minimizing interference

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.553428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.754884Z digest=sha256:ac8e2738a387c99622678fd683f401a397f6deec884c1cac6c65c54f38f1d3e6

Observation f232caa8-7826-467d-98a3-01518d114fce · outbound

This paper cites Experience replay for continual learning.

Continual Reinforcement Learning by Planning with Online World Models Experience replay for continual learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.527004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.807528Z digest=sha256:14a46870e4db23d9d480675a8cbc402367c9a45f115d594ca656f66446a78379

Observation 87c44d83-294b-4227-9b40-e659a84e8c8f · outbound

This paper cites The cross-entropy method for combinatorial and continuous optimization.

Continual Reinforcement Learning by Planning with Online World Models The cross-entropy method for combinatorial and continuous optimization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.435368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.916005Z digest=sha256:27efe47597ee8c89e743720d7cbaefcd666c83aa7752ed3ede82acf6bb8f5d1c

Observation adf74431-fe53-4ee6-a69b-d0fa59e045d3 · outbound

This paper cites Curious exploration via structured world models yields zero-shot object manipulation.

Continual Reinforcement Learning by Planning with Online World Models Curious exploration via structured world models yields zero-shot object manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.338379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:17.986840Z digest=sha256:ff9afd526c725656104adc8eabc8e4bb8c7ce4f004f002b29a95069c4d19b8c3

Observation 6b726db6-5975-45df-b231-b5696bf2a301 · outbound

This paper cites M., Grabska-Barwinska, A., Teh, Y.

Continual Reinforcement Learning by Planning with Online World Models M., Grabska-Barwinska, A., Teh, Y

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.258193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.059835Z digest=sha256:35342f5a65e95f88cde3098d05329ad83454dda45e09f001e8b6abe2df77f071

Observation c1cbf20f-c36a-441b-a36b-5366a0c145d2 · outbound

This paper cites an unresolved cited work.

Continual Reinforcement Learning by Planning with Online World Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:11:15.123043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.127946Z digest=sha256:faab9dccfb90443d1ee09808fff90d474fad108b4ae0e1e6ba00a5d84cab577a

Observation 36674e8b-0cdd-4ef2-a85c-3c37a992cdf2 · outbound

This paper cites Autonomous reinforcement learning: Formalism and benchmarking.

Continual Reinforcement Learning by Planning with Online World Models Autonomous reinforcement learning: Formalism and benchmarking

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:15.029920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.218857Z digest=sha256:46297f0b4d43356aa7abe9485f7d080226a9d48c53ab71a7b639c1381120487e

Observation 187b17b3-fbb3-47e8-b78c-bc6a6075ade2 · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Continual Reinforcement Learning by Planning with Online World Models Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:14.929917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.273145Z digest=sha256:59feab281d2a22d753d64f62a8b103812662414d7b39cf84d9c9e95b3c1b53c2

Observation ae1e6d2b-ad26-4669-811a-b8dbcf469a76 · outbound

This paper cites an unresolved cited work.

Continual Reinforcement Learning by Planning with Online World Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:11:14.902729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.318361Z digest=sha256:708b273715cd062600021b132a91806b93a78b9105b00174b649b6477bbecc73

Observation 2289b0f7-af4c-40f8-8a28-f763a7cb7df3 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Continual Reinforcement Learning by Planning with Online World Models Mujoco: A physics engine for model-based control

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:14.838708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.380419Z digest=sha256:d6e047c052fb9c6a75ee1680ecdd6f01a3ae0d771719b8313ae1ec45654c96d8

Observation 2bc98a74-7c0c-4211-bf50-332b3052b698 · outbound

This paper cites an unresolved cited work.

Continual Reinforcement Learning by Planning with Online World Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:11:14.755578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.454588Z digest=sha256:7e155d0a513a3ff9bfbd88b21235fadc94252744ba76c79275d58799430223bc

Observation 9465a9fc-214e-4397-b9f4-cf242043823e · outbound

This paper cites and Ba, J.

Continual Reinforcement Learning by Planning with Online World Models and Ba, J

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:14.610854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.530369Z digest=sha256:acc46db242c5dbc778a413b3762cc667ae9a3ac442834607a574351a0a4212fa

Observation 2ba237eb-ed53-4787-ba6c-59981f511b95 · outbound

This paper cites Model Predictive Path Integral Control using Covariance Variable Importance Sampling.

Continual Reinforcement Learning by Planning with Online World Models Model Predictive Path Integral Control using Covariance Variable Importance Sampling

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:18.638329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:10:18.638329Z digest=sha256:3a9071d016319ce7a83877eae306267e5f057ccd9fe551e06a12addcd9770258

Observation 238598ba-ef15-4136-8ce9-e21085f140ff · outbound

This paper cites Continual world: A robotic benchmark for continual reinforcement learning.

Continual Reinforcement Learning by Planning with Online World Models Continual world: A robotic benchmark for continual reinforcement learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:14.518922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.750710Z digest=sha256:2c2edeac4fb7b2985affe5b4be5fcc2026fc71ee5b87f1f701cb700da6c6ae4d

Observation f0b98d82-4964-4612-90d2-2583f87170d5 · outbound

This paper cites Continual task allocation in meta-policy network via sparse prompting.

Continual Reinforcement Learning by Planning with Online World Models Continual task allocation in meta-policy network via sparse prompting

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:14.418368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.806077Z digest=sha256:335d860389188934e3d3effb411f8053ffc7bf9bc57faa7da2c91b61364d30e7

Observation 438b62fe-351d-46cf-bf4e-e9796a7295cc · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

Continual Reinforcement Learning by Planning with Online World Models Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:14.267972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.908717Z digest=sha256:9316ce4dabd5e9eb3850654a7dae96f46b67ed239e7604b301a7cbaeccf3ed32

Observation c36e89c5-84d1-4553-94a0-59853dfb233f · outbound

This paper cites Continual learning through synaptic intelligence.

Continual Reinforcement Learning by Planning with Online World Models Continual learning through synaptic intelligence

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:19.891697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:18.974652Z digest=sha256:ac440f84b20b861e330eb72566a8146f19ac3924163da42ebafcbe8c565577f7

Observation 61f55a7f-cb68-44fa-b2b9-84baa3f27b02 · outbound

This paper cites The Schur complement and its applications, volume 4.

Continual Reinforcement Learning by Planning with Online World Models The Schur complement and its applications, volume 4

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:19.723179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:19.057358Z digest=sha256:09f70606ff2219bdd29cf49dc65fea45c167b8c5c1234b202636a49f31e793d3

Observation 29fabc1a-1ff6-4b94-9699-6590b2f89abe · outbound

This paper cites ACIL : Analytic class-incremental learning with absolute memorization and privacy protection.

Continual Reinforcement Learning by Planning with Online World Models ACIL : Analytic class-incremental learning with absolute memorization and privacy protection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:19.581119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:19.108937Z digest=sha256:021a01dd1461f0258674963031f7c2ffa3f214dd4348cbd7725515ea7cedd7be

Observation 572e2b75-52be-4b2c-86cd-ae47ff919a37 · outbound

This paper cites GKEAL : Gaussian kernel embedded analytic learning for few-shot class incremental task.

Continual Reinforcement Learning by Planning with Online World Models GKEAL : Gaussian kernel embedded analytic learning for few-shot class incremental task

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:19.459214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:19.154747Z digest=sha256:ff12f6e3e631209835f3a8567b10b90c8e4adce3b206dc43e6e371a07adc1ab4

Observation 69969800-dead-4048-a9d5-4e087eb8d5b2 · outbound

This paper cites DS-AL : A dual-stream analytic learning for exemplar-free class-incremental learning.

Continual Reinforcement Learning by Planning with Online World Models DS-AL : A dual-stream analytic learning for exemplar-free class-incremental learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:19.377619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:10:19.203780Z digest=sha256:b2b5a82493b32cfe58e4cad4c4c44561c50de614b25ee317489f841ff1c6a97a

Pith citing papers

Observation f477ed61-5763-4485-a849-ab24ba3c43b7 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Continual Reinforcement Learning by Planning with Online World Models

Reference 186

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.172861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:6dfe92a0f0cc5a56ad09204e3f3b03b542ba9dadd6c3a4766026ce8e014ede91

Observation 289902f6-1611-4218-b269-0f3be319bdd3 · inbound

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions cites this paper.

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions Continual Reinforcement Learning by Planning with Online World Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:26.617162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T10:51:30.546134Z digest=sha256:37e52231fba731b64c9d34c0fbee1a97dcd1a6d6685ede618a9dee69ca7845c4

Observation b8bb22d9-ae2e-47c6-847a-11af09512cfd · inbound

Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling cites this paper.

Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling Continual Reinforcement Learning by Planning with Online World Models

Reference 154

Resolution
unresolved
no resolver link, observed 2026-07-11T19:24:48.899301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T19:24:48.899301Z digest=sha256:f07ca91350c360a588b8a581b132a78df1e82547114bdfc17bc33fce1b1f8342