Pith. sign in

Paper Citation Record · LEDGER

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 2 inbound Pith citation observations for arXiv:2506.06261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06261 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:03:49.520373Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:27:31.026539Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:09:14.926116Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy25
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8092b47-592e-4a82-b66a-cd0fce5fc801 · outbound

This paper cites write newline.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:44.662883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:44.662883Z digest=sha256:ec1046fe1cf0bde467eab4430abb5cf5e7567719a9e7411d5d2bc6807329db29

Observation c1e2bd2b-cdcf-4cf5-8a4c-2568b5e59295 · outbound

This paper cites T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:58.817378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:44.734527Z digest=sha256:5d4dce62fb554ed921acf05f6a30a3c376ef135f07c3e05c0d7b648ebcd0c426

Observation 22a2a20c-39f7-4aec-ad69-8d7dc93f94d7 · outbound

This paper cites Deep Reinforcement Learning at the Edge of the Statistical Precipice.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Deep Reinforcement Learning at the Edge of the Statistical Precipice

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:44.808251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:44.808251Z digest=sha256:f6f1dee062de06e26c5cd9655bd5fd7cb95e37de67d88c9e7197aad65b23a92e

Observation 05899d7e-9df4-4707-b93a-b9b427f6ea1e · outbound

This paper cites an unresolved cited work.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:58.554367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:44.892153Z digest=sha256:870396b914c27a6d7fdd7c67b8cb14528d864f98a08b1aea5a4595ff7e5df01f

Observation f78c6806-41a6-42ce-bed7-7574043fae77 · outbound

This paper cites and Dulac-Arnold, G.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens and Dulac-Arnold, G

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:58.335194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:44.989810Z digest=sha256:5789fecceac6d43fbf6cac669dc2fe94da8677cee8f1757c957919339c6523d6

Observation 46e4276a-0f8e-48b6-a9fa-e2bc0d3300df · outbound

This paper cites Experiment tracking with weights and biases, 2020.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Experiment tracking with weights and biases, 2020

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:58.058084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.042792Z digest=sha256:ee617a25e85590cf31b6dfc9cc3de87f7090af766cb2cbe35ac399f9e8dca36b

Observation 9f0003b4-1283-495a-a48d-a56f1b787393 · outbound

This paper cites T., Wenjie, S., and Ye, J.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens T., Wenjie, S., and Ye, J

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:57.739705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.123090Z digest=sha256:d865e7b30f0dcc1875ff417832668296e1407432c2fbb3a43832adcfe2eea730

Observation ee9dc7b7-43f8-491c-b5c0-6781847d8ddc · outbound

This paper cites Deep reinforcement learning in a handful of trials using probabilistic dynamics models.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Deep reinforcement learning in a handful of trials using probabilistic dynamics models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:57.533127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.201119Z digest=sha256:1e0bca615e78979b9bb4c7a91ca47db3d381f506c20acee9ac7a27af8cb52fb8

Observation 9428f096-c1a4-48d4-8f2e-2aa7816e0afc · outbound

This paper cites S., Abbeel, P., Levine, S., and Finn, C.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens S., Abbeel, P., Levine, S., and Finn, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:57.285004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.285290Z digest=sha256:4de2fe851aeb543d2fa79ca60e57e36f0ffd79ee56be82b526a0203dfe769fae

Observation ca17b460-a715-47aa-8594-b91f28163e4d · outbound

This paper cites Offline meta reinforcement learning -- identifiability challenges and effective data collection strategies.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline meta reinforcement learning -- identifiability challenges and effective data collection strategies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:56.969870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.365659Z digest=sha256:51c89a464c4f9b65dcbdee2ff321142d7bf2b433d6677c2cc11742a1fb79a0fc

Observation ab536021-8957-42d9-8dcb-f46a45381c51 · outbound

This paper cites an unresolved cited work.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:56.712428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.411352Z digest=sha256:7bf15c4fa527e88a37761e89faa501db9d6b34c0ed26773e18070217f505d748

Observation 4c5c8bac-27bf-41dc-8dc2-08557a1ec706 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:45.544896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:45.544896Z digest=sha256:41764487121bbb0af9268a4147418f5d745ab148b4c92cc3155ce793eb29d82e

Observation b76c7758-8e57-4861-9361-d4dd75684348 · outbound

This paper cites and Gu, S.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens and Gu, S

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:45.611725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:45.611725Z digest=sha256:52374c576c27570295034f67dccbb6eb146a5f9a69d7a3761d032a5ee5a2bfc4

Observation b7b5d722-cb99-4da3-b011-02228b4b32f4 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Off-policy deep reinforcement learning without exploration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:56.465987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.696943Z digest=sha256:dba9740e5ab7f32b7fdd10511b6d5aae41623167ec03b2e576740fbdc02fd4a4

Observation c639d093-f020-46bd-81ab-9e58b7ccd681 · outbound

This paper cites Bayesian reinforcement learning: A survey.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Bayesian reinforcement learning: A survey

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:56.146775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.765536Z digest=sha256:8726d9d8a51a50abc0330b5ea35ef00b89dda60d51a2327a36a8a4f65893bf7b

Observation f6fef5e4-79ef-4c5d-ba64-f1a2869dbb01 · outbound

This paper cites P., and Levine, S.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens P., and Levine, S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:55.861945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.831667Z digest=sha256:1207879adfcdb613c6c31b6ab55c018b4f5b66d1537ab50806c110e875daf5b1

Observation 27827645-388a-40b6-9c8a-672b7ecd6df1 · outbound

This paper cites Offline RL policies should be trained to be adaptive.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline RL policies should be trained to be adaptive

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:55.550935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.937699Z digest=sha256:377fd0afcee0d2326856a9ab9c92c3eed69788240367e78190bd6dbd23c604dd

Observation 321a86f3-abf5-44f5-8209-44aef3c51753 · outbound

This paper cites Efficient bayes-adaptive reinforcement learning using sample-based search.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Efficient bayes-adaptive reinforcement learning using sample-based search

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:03:50.770223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.012409Z digest=sha256:a04f1265ff4bc67aa1b15b2a7b4fe6ccaf532e2349a051643b9a4081a42a9827

Observation 705bd01e-1c79-4bfe-a40d-afbf8fe81c58 · outbound

This paper cites When to trust your model: Model-based policy optimization.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens When to trust your model: Model-based policy optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:55.308363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.096762Z digest=sha256:d15be2103f5b38d153f814dd6e0d09811631cb8fc8f1faa7b8814b2223c2d758

Observation 42ad2618-6304-40fe-b34f-6df7def337bc · outbound

This paper cites Planning with diffusion for flexible behavior synthesis.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Planning with diffusion for flexible behavior synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:54.959745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.183550Z digest=sha256:0dd0e224e78bb862b09616436082b495ae639cded1fc5ea6846f5675bb319105

Observation 0ab800cf-6a1f-4238-bbc1-dfc6983d4a9c · outbound

This paper cites Is pessimism provably efficient for offline rl? In Meila, M.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Is pessimism provably efficient for offline rl? In Meila, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:54.657314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.260650Z digest=sha256:43da883f9bd057cf7f69f42e594dc2dce300ed0f3192cc22b447a73925e11679

Observation ca289220-e219-4c8d-96c5-209ad2b4d40e · outbound

This paper cites P., Littman, M.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens P., Littman, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:54.324634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.347403Z digest=sha256:cc936d098adc86933974bb6739b22af3e4d9557033a486657d3a097d114bc07d

Observation 0ba19711-98f1-4007-9ad5-5c42e5c10026 · outbound

This paper cites Morel : Model-based offline reinforcement learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Morel : Model-based offline reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:53.981841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.452733Z digest=sha256:e4e27f413d30367d53dd687c0809cda24c427f9a371341fc17337f37a788854a

Observation dd60bff0-8a58-4589-850d-21e8f415fd4d · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline reinforcement learning with implicit q-learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:46.549158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:46.549158Z digest=sha256:2f9425006bbac2bcaddafad9e5ea4c19a0bcb426b7aa27120aebd78ea0873d96

Observation 55e25cb3-2daf-442c-bced-130b23261ac5 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:53.649910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.664728Z digest=sha256:e54a205e1b258c3fb95c4e58cb440a0f95c1cf32c571e45af96176f69dd8a04a

Observation a226a1e9-50ab-4673-abec-69396cd17915 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Conservative q-learning for offline reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:53.293634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.791841Z digest=sha256:56b983eecd8e4ca033de478e14c9621333be2f4ae758df189a224e8f22e68260

Observation b3e38fe3-479d-4d2b-97b8-d3cc76116f40 · outbound

This paper cites Reinforcement learning and control as probabilistic inference: Tutorial and review, 2018.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Reinforcement learning and control as probabilistic inference: Tutorial and review, 2018

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.991334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.990592Z digest=sha256:f2daa8948472a8dca35fc7e8af5f292139e1205c7edc9bd23dd02f92355aad14

Observation d0d22f6d-3053-4300-8c3b-b4cbd78b069d · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:47.167284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:47.167284Z digest=sha256:ceac52bdcd87a2ed646f44f3c150a463997d55a878db03daeecb90943e6060c7

Observation e95f3e36-e84f-46ae-b63d-338187ba32a5 · outbound

This paper cites Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:47.372549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:47.372549Z digest=sha256:db74895c2a662f6d19837b002b05a63f77a686600e0e6ba45f9e013780366350

Observation e17cfde7-9275-4d2c-bf81-faf758631575 · outbound

This paper cites Revisiting design choices in offline model based reinforcement learning, 2021.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Revisiting design choices in offline model based reinforcement learning, 2021

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.717159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:47.506550Z digest=sha256:59b044a7c27ebd46f81e0e4816cc45a83fc18beb6b5118da47f24dc9b520bc4b

Observation 8de4fcca-fc38-43a3-a230-aa0ca047e3ea · outbound

This paper cites Deep Dynamics Models for Learning Dexterous Manipulation.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Deep Dynamics Models for Learning Dexterous Manipulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:47.644501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:47.644501Z digest=sha256:36f278af0959b9eed361622250cfb57d49eb4eae4e42f8b64fcc7684f3a15261

Observation 3425e546-0fed-405c-9f01-b5c0be6bdda4 · outbound

This paper cites and Taniguchi, T.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens and Taniguchi, T

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.311287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:47.808954Z digest=sha256:6a7fb2625b63d4dec803335f4c0c76a5393d948342256a63724e896650b68a9f

Observation 76999567-9b58-4024-a5a9-a64c44821feb · outbound

This paper cites Probabilistic planning with sequential monte carlo methods.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Probabilistic planning with sequential monte carlo methods

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.002407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.015339Z digest=sha256:34b182cf8e9a517ac759547614574d6a05cbb29ff468a456077eb3c607f9e179

Observation 3dcb864d-2bad-4af4-8797-ee40806379e7 · outbound

This paper cites RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:48.178163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:48.178163Z digest=sha256:9a10678405054d7d294f4a87734018a44761c9e278da3d391ca01f464eb38f6d

Observation 6b04e500-2cc4-4f42-87c8-27ab9d7706b4 · outbound

This paper cites Learning off-policy with online planning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Learning off-policy with online planning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:51.673944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.345210Z digest=sha256:8ed4d234d62502bd5dfd029f35a9da46ccc59c784c8ab2b0d2d427b1e6d3b211

Observation 300b965f-9130-4fc9-bd52-e230fd1425a8 · outbound

This paper cites Improved sampling-importance resampling and reduced bias importance sampling.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Improved sampling-importance resampling and reduced bias importance sampling

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:03:50.402148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.477888Z digest=sha256:2c5d4db4e85470d0270e2f2ab14aaf8e8359c221efe6a158387ba939f4736b85

Observation 2d6978e4-bf2e-413e-9d8d-5540e253e3ac · outbound

This paper cites an unresolved cited work.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:48.624377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:48.624377Z digest=sha256:5925421d9b21252ffb32aa65c47ec22930bf7fdfb2ce7513ea67cd93b242db3e

Observation 85faaf2e-b21c-4330-83d4-9a7482d1a1f0 · outbound

This paper cites Model predictive path integral control using covariance variable importance sampling, 2015.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Model predictive path integral control using covariance variable importance sampling, 2015

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:51.303992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.752678Z digest=sha256:69085775e4233ad2b252f9423783a4cb59c9a4e95c9a64b1b31608dd4e7c1699

Observation 6c80e0e6-da96-4a3c-8da9-f2437ea0e9a8 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Behavior Regularized Offline Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:48.860136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:48.860136Z digest=sha256:c94adde4f2043edd7f3ff0d603ea56b1017e445f49e726c7ad5eaa7429eebbc3

Observation 7f0ae814-39de-4658-a455-60ce2bad1014 · outbound

This paper cites Mopo: Model-based offline policy optimization.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Mopo: Model-based offline policy optimization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:50.976072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:49.021233Z digest=sha256:d97d766334bfe3d462dd132da369ed660af102f324eaa4def6c26608ab9cd3d1

Observation 40591d21-29e5-46d1-b52f-9a758d043e95 · outbound

This paper cites COMBO: Conservative Offline Model-Based Policy Optimization.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens COMBO: Conservative Offline Model-Based Policy Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:49.226355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:49.226355Z digest=sha256:378cfc62283d89ec21a3a8d221c4d533bef735dc0cd2603929119844e755bf13

Observation 377ac002-9335-4d18-923f-0e1bace05914 · outbound

This paper cites Model-based offline planning with trajectory pruning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Model-based offline planning with trajectory pruning

Reference 42

Resolution
verified exact
doi, observed 2026-08-07T06:03:49.820614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:49.380210Z digest=sha256:dc7799c28b1f7fc9168c0150b61df14607f88983041249d68fecf09baa6cde85

Observation c4e594ae-0268-48e9-92cc-2d4f6d18b387 · outbound

This paper cites VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:49.520373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:49.520373Z digest=sha256:f1d5366c03a9e56bbef39e989b194c0b422431d56d3957dcb9215ec85d5b9157

Pith citing papers

Observation 26f71b9d-ad91-4084-834d-b3d50183cc9f · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.907040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:e4ad4b3b4d9ba16918f702b6c67cf0e8ea38cc5b8715759153a7604cc1def18e

Observation c1530e55-3dc4-43ea-861d-91ed9c3c525d · inbound

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning cites this paper.

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:09:14.928534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:27:31.026539Z digest=sha256:ba4edf4b43cca48d34f9cb335321230bb4d4d995472d96a867044e6ed6548a6e