Pith. sign in

Paper Citation Record · LEDGER

Fully Offline Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 2 inbound Pith citation observations for arXiv:2505.22442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22442 v3

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:16:05.159203Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:25:39.829506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T01:48:51.121586Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact11
  • verified fuzzy41
  • unresolved28
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f4fa7dd-2576-4833-918f-082dad491144 · outbound

This paper cites Aitchison.

Fully Offline Reinforcement Learning Aitchison

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:09.424246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:57.315944Z digest=sha256:4288c87d0a35d0300cff361ca86045750dd6b320f610b313a61ff6d675c6b7aa

Observation aa23fa33-78ea-43be-ad09-ec04a2ec318a · outbound

This paper cites Alaa and Mihaela van der Schaar.

Fully Offline Reinforcement Learning Alaa and Mihaela van der Schaar

Reference 2

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:16:09.081121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:57.356190Z digest=sha256:1e2a3e76482b460ba23bc0d447096d68d5bf51fe82b88a6116aa0ae16ffd01cb

Observation faa3004a-74fb-466a-b59d-28e8b50b6332 · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Fully Offline Reinforcement Learning Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.413976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.413976Z digest=sha256:e4c2c90a93630f1e098d7b1fbf5f545f9851e63ce1ad718494a907e8fe3dc77e

Observation 3ad69933-f655-4d69-8dee-5f498d7ff47c · outbound

This paper cites Asymptotically minimax bayes predictive densities.The Annals of Statistics, 34(6):2921–2938, 2006.

Fully Offline Reinforcement Learning Asymptotically minimax bayes predictive densities.The Annals of Statistics, 34(6):2921–2938, 2006

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.474090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.474090Z digest=sha256:3d6541d4b2b1c240c873f89560700657d55d524489fba50a15f8889ef7a37fe0

Observation 02430707-3cef-41eb-acea-41eedfcd62cc · outbound

This paper cites Augmented world models facilitate zero-shot dynamics generalization from a single offline environment.

Fully Offline Reinforcement Learning Augmented world models facilitate zero-shot dynamics generalization from a single offline environment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.536060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.536060Z digest=sha256:53279ddb823255a91ae555050935f008eec703bcd8c4629f02edf3b96b96018a

Observation b48111ec-9df1-470f-b256-4122791e1c06 · outbound

This paper cites Information-theoretic characterization of bayes performance and the choice of priors in parametric and nonparametric problems.

Fully Offline Reinforcement Learning Information-theoretic characterization of bayes performance and the choice of priors in parametric and nonparametric problems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.255216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:57.590263Z digest=sha256:a34062c0098fcb7c0042d447236ca7a53a63f6448766901583c855b6d8cbe445

Observation ab150b87-375f-4aca-b853-765cb9f73e2f · outbound

This paper cites Barron.The Exponential Convergence of Posterior Probabilities with Implications for Bayes Estimators of Density Functions.

Fully Offline Reinforcement Learning Barron.The Exponential Convergence of Posterior Probabilities with Implications for Bayes Estimators of Density Functions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.989180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:57.757860Z digest=sha256:31cdc65797bfd7a7c34ba49546958ee8825703f8d43023aa5f55c74880352346

Observation d16ea67a-4b7d-4dcf-83cb-b64d4ddd1675 · outbound

This paper cites Bass.Real Analysis for Graduate Students, chapter 21.

Fully Offline Reinforcement Learning Bass.Real Analysis for Graduate Students, chapter 21

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.820274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:57.822068Z digest=sha256:38498773d00483385f1ced6658436097319a3f73e5324e92b14559c3f2e6b831

Observation 14fff60b-933b-4eff-8483-c0f9dc237f61 · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

Fully Offline Reinforcement Learning A Tutorial on Meta-Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.942175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.942175Z digest=sha256:ba84e251aa96ab08db4fa3cf51d2f2fa6077e062c6feb877d06444cf5bfb5bf0

Observation c275b9b1-b9bd-4ad5-a6de-e2c89d65fe86 · outbound

This paper cites A problem in the sequential design of experiments.Sankhy ¯a: The Indian Journal of Statistics (1933-1960), 16(3/4):221–229, 1956.

Fully Offline Reinforcement Learning A problem in the sequential design of experiments.Sankhy ¯a: The Indian Journal of Statistics (1933-1960), 16(3/4):221–229, 1956

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:08.501815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.001783Z digest=sha256:013c1c61c1d0ebca41186c116bc1d5fa3c135689caf6c00c350f215abfb07731

Observation 53dbe3be-d34e-4b8f-b8c4-eeae6d108a02 · outbound

This paper cites Dynamic programming and stochastic control processes.Information and Control, 1(3):228–239, 1958.

Fully Offline Reinforcement Learning Dynamic programming and stochastic control processes.Information and Control, 1(3):228–239, 1958

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:15:58.052512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:58.052512Z digest=sha256:160fc5f25c001d4c1e31650e2ac576985e3e9142d6e4225bce6b1b4081c75771

Observation ad5273c6-cfb0-4297-8b5b-d36395eea30e · outbound

This paper cites Foster, and Daniel M.

Fully Offline Reinforcement Learning Foster, and Daniel M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.427736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.116894Z digest=sha256:a79cf15d852385673b971d50aee96617ab4ed4d494822b34fd3e8a987dbac3e3

Observation 3f3dade5-de7b-468d-8deb-ae49eabbfc43 · outbound

This paper cites On the foundations of statistical inference.Journal of the American Statistical Association, 57(298):269–306, 1962.

Fully Offline Reinforcement Learning On the foundations of statistical inference.Journal of the American Statistical Association, 57(298):269–306, 1962

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:58.171525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:58.171525Z digest=sha256:fa4bfdf6f479ea089b252dd3e0b65d3f25488c03a0ca3e1f4859b2515a861933

Observation 50ccbea6-642b-4c41-abfb-6ccd7d75f62c · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:20.279450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.236104Z digest=sha256:9588deae6b4996f9bcc2e68a6bfdb1957008ce681366107120d8b775fe85f97e

Observation de0a4ee4-4b43-46b8-9c14-1c86432ce338 · outbound

This paper cites Bayes adaptive monte carlo tree search for of- fline model-based reinforcement learning, 2024.

Fully Offline Reinforcement Learning Bayes adaptive monte carlo tree search for of- fline model-based reinforcement learning, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.135149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.297640Z digest=sha256:f77b7f2ca7e897906297292a26de389d8a4adbb81a414352ef69b81ebb6e080b

Observation cfe50663-3e99-4113-815f-551afb1c29ee · outbound

This paper cites Conser- vative uncertainty estimation by fitting prior networks.

Fully Offline Reinforcement Learning Conser- vative uncertainty estimation by fitting prior networks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.983352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.372662Z digest=sha256:6f3f9eb3fec3cc108db4efe2ad8ed3bf409bfd066bb2236fbb3cc320b56a3aa7

Observation 016295ea-6bda-4420-99f2-2c9d7ae85950 · outbound

This paper cites Clarke and A.R.

Fully Offline Reinforcement Learning Clarke and A.R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.743980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.448968Z digest=sha256:35d3454ea444183beab4bc67b2774c1bf7ec30c6b0682fcbdd270a1da231c0cc

Observation 4ba4c40f-40e6-47f5-a44a-23eb99b56a5f · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:07.959162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.500423Z digest=sha256:d771199315013d7715acfc6feeda1a7ab411a4ac64ce921fca57381ded94f4ee

Observation 54158c56-5f57-431f-be73-7a763611224d · outbound

This paper cites Observation of a markov process through a noisy channel.PhD Thesis, 1962.

Fully Offline Reinforcement Learning Observation of a markov process through a noisy channel.PhD Thesis, 1962

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.551859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.563857Z digest=sha256:0f34780ad6a304a52b00e8586322b7c50aab3321e5a06f4711f00ea57f0d6810

Observation d5475ef1-f8d4-4265-9802-3e75ba2d6176 · outbound

This paper cites Fast reinforcement learning via slow reinforcement learning.

Fully Offline Reinforcement Learning Fast reinforcement learning via slow reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.341391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.618468Z digest=sha256:a19ed97bb27ff5b068f2f85bc1f751f2d8201fb49c04e8795fc19c483227f4ad

Observation fd95ba64-4cef-4f53-a399-142fcde22945 · outbound

This paper cites PhD thesis, 2002.

Fully Offline Reinforcement Learning PhD thesis, 2002

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.101212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.684957Z digest=sha256:9baf0aa27a57412b47f5b4083d0f7bada0b21a983e1ad045745d370960dc1b8a

Observation b34d3302-8720-46c0-88bf-8404861c1c57 · outbound

This paper cites Bayesian exploration networks.

Fully Offline Reinforcement Learning Bayesian exploration networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.931137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.801124Z digest=sha256:68de96d766583271d3c84ce7694a31460403d3449f2a8e7eef661b03a6b296ef

Observation af8f42f6-eebf-41e4-a8ed-29b0c9113c05 · outbound

This paper cites D4rl: Datasets for deep data-driven reinforcement learning, 2020.

Fully Offline Reinforcement Learning D4rl: Datasets for deep data-driven reinforcement learning, 2020

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.703635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.874936Z digest=sha256:492421229617c7ad68e5226a4233d53c37ef8a2f27ee39345a9bca6793dae476

Observation 37eca544-1166-4b74-b6d2-90890599c766 · outbound

This paper cites A minimalist approach to offline reinforcement learn- ing.

Fully Offline Reinforcement Learning A minimalist approach to offline reinforcement learn- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.497264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:58.962879Z digest=sha256:0e8f47950bec0d493b1f52067036e203018a97bb3c0b7bb63109635c1583197d

Observation b819815a-87ac-40bc-8d62-8771f3c71ef1 · outbound

This paper cites A new proof of the likelihood principle.The British Journal for the Philosophy of Science, 66(3):475–503, 2015.

Fully Offline Reinforcement Learning A new proof of the likelihood principle.The British Journal for the Philosophy of Science, 66(3):475–503, 2015

Reference 25

Resolution
verified exact
doi, observed 2026-08-07T13:16:06.339819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:59.057545Z digest=sha256:301b2b53e635069ba91514ef9fd59716fb65ef0a9dc0aba2487f7e6630e48895

Observation ef74a39f-c140-461e-9299-238a9cacd29d · outbound

This paper cites Efficient bayes-adaptive reinforcement learn- ing using sample-based search.

Fully Offline Reinforcement Learning Efficient bayes-adaptive reinforcement learn- ing using sample-based search

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.234691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:59.158171Z digest=sha256:32625d31526e2d9d8471c72d175c26b9f95698fb1a1cf087a4e481c71a3d0a54

Observation 66c454f4-a472-49ce-9f75-be6bfcf5830a · outbound

This paper cites Scalable and efficient bayes-adaptive reinforcement learning based on monte-carlo tree search.Journal of Artificial Intelligence Research, 48:841– 883, 10 2013.

Fully Offline Reinforcement Learning Scalable and efficient bayes-adaptive reinforcement learning based on monte-carlo tree search.Journal of Artificial Intelligence Research, 48:841– 883, 10 2013

Reference 27

Resolution
verified exact
doi, observed 2026-08-07T13:16:06.134303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:59.266350Z digest=sha256:ed5b9c2eb97484b11d169fdc28f5b344e53c93a59469d06d779ff3e9541a3bdc

Observation 5ae9635e-e761-4966-9f5c-475929cbdaa3 · outbound

This paper cites Bayes-adaptive simulation-based search with value function approximation.

Fully Offline Reinforcement Learning Bayes-adaptive simulation-based search with value function approximation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.029970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:59.378247Z digest=sha256:8bbab09f43b6d342095c595a0386b602f525f70e0e21e791d1cee93a31459deb

Observation 0bf909f1-3bb1-4868-9454-17edfad43c83 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:17.800520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:59.456794Z digest=sha256:58f930cdf32af4a55233a1bd69df17e284aa12c10b8a2e38a036a25aacc5752c

Observation 9d4a2e4d-9dec-438c-8374-ef6516ec5cf1 · outbound

This paper cites A Clean Slate for Offline Reinforcement Learning.

Fully Offline Reinforcement Learning A Clean Slate for Offline Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:59.584911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:59.584911Z digest=sha256:c5ad76453fee4b3814b720c188ac50fe2c5f7e7529485112f22dffb752d4a7b9

Observation cea3e1a3-04be-4a0e-b586-9593314cd513 · outbound

This paper cites Relu to the rescue: Improve your on-policy actor-critic with positive advantages.

Fully Offline Reinforcement Learning Relu to the rescue: Improve your on-policy actor-critic with positive advantages

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.425672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:59.840841Z digest=sha256:66a691bae8a5b4165721756eeb063098f7c98d6a65ce3104816006345d40461e

Observation 3d7c401b-dbee-484d-97e6-317c68901b68 · outbound

This paper cites Littman, and Anthony R.

Fully Offline Reinforcement Learning Littman, and Anthony R

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.259390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:59.939273Z digest=sha256:86eb8f6d97111f23c506551ac9ce41135f2a8f13f5a25936c03bad60787f05fe

Observation c8b21251-6035-4794-b409-1b41bf81643a · outbound

This paper cites The validity of posterior expansions based on laplace’s method.Bayesian and Likelihood Methods in Statistics and Economics, pages 473–488, 1990.

Fully Offline Reinforcement Learning The validity of posterior expansions based on laplace’s method.Bayesian and Likelihood Methods in Statistics and Economics, pages 473–488, 1990

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.095540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:00.030618Z digest=sha256:2bd4c2273258a85cda9708bb2ec48359c0449f8482a91d4dbdc07f269303cc2d

Observation 5f93167a-8333-4b84-9aa6-02e28e1791d0 · outbound

This paper cites Morel: Model-based offline reinforcement learning.

Fully Offline Reinforcement Learning Morel: Model-based offline reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.938047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:00.149839Z digest=sha256:57dada42f72bcaa251e0266b3998671178bf409cf4695bbd19dfb8347dc28d0d

Observation b1048332-956b-40ae-97f3-948b0f060334 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Fully Offline Reinforcement Learning Adam: A Method for Stochastic Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:00.249033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:00.249033Z digest=sha256:a29c7ae718cb2a9aa6117b907f7ad43863c53c696852df1b90fa90b97c5c5d88

Observation 07d239ca-ddc3-4b44-bd0a-7fa83874c750 · outbound

This paper cites Kleijn and A.W.

Fully Offline Reinforcement Learning Kleijn and A.W

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:00.328567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:00.328567Z digest=sha256:b35e59eda34c3abd56a3e858d66ad60979fdea7959d104d8442b3fc763e9b99e

Observation 60336ccb-c3d8-410e-b70c-1f30de21c1e6 · outbound

This paper cites On asymptotic properties of predictive distributions.Biometrika, 83(2):299–313, 06 1996.

Fully Offline Reinforcement Learning On asymptotic properties of predictive distributions.Biometrika, 83(2):299–313, 06 1996

Reference 37

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.898060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:00.414859Z digest=sha256:aad20dca7cb80583b68729501b3cfc3d2625c087a75901ef7bb4dcf56cbb20dc

Observation 5a7d85c1-7902-4a0e-b581-c368ca353447 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Fully Offline Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:00.532185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:00.532185Z digest=sha256:86d6197eec9c28ed503c4dc3f5cc73caf80fd9a9eaaf92399372793b50633993

Observation 30ecba1b-69b7-4d9b-858a-2f4e7484e894 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:16.756648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:00.631981Z digest=sha256:f6bae83611b506306a7b3e3b5897e7bedd6bf3865d5a94cb5f4665afe1b26dd3

Observation ffe6b090-67c1-4b10-80c7-98f457a65e28 · outbound

This paper cites Conserva- tive q-learning for offline reinforcement learning.

Fully Offline Reinforcement Learning Conserva- tive q-learning for offline reinforcement learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.535991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:00.720547Z digest=sha256:ad118b9fb1c2ff2e61a822d87f1d6028d70bc970ed716699a858100c61a61f3a

Observation ca65e117-c96a-4482-946a-1488ab71b6ce · outbound

This paper cites Springer Berlin Heidelberg, Berlin, Heidelberg, 2012.

Fully Offline Reinforcement Learning Springer Berlin Heidelberg, Berlin, Heidelberg, 2012

Reference 41

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.713211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:00.803761Z digest=sha256:45caf5e6a9ded1ecc873977cfb12995d643d9f391270d51db76b5e89ba9a0362

Observation 59003146-ce29-4e1a-b6de-9a5eeb746557 · outbound

This paper cites On some asymptotic properties of maximum likelihood estimates and related bayes’ estimates.

Fully Offline Reinforcement Learning On some asymptotic properties of maximum likelihood estimates and related bayes’ estimates

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.367331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:00.913603Z digest=sha256:9c45fcd5c5f22602abb93b9fa99aad39587eb9bfdcf56a481bba333207195ee7

Observation c3b51491-c4f4-492e-853b-98679cf6c568 · outbound

This paper cites Efficient backprop.

Fully Offline Reinforcement Learning Efficient backprop

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.174298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:01.009737Z digest=sha256:8e53c16c20b9e98d498e884a72cb865f6cd26d5d37e6d3267bfa13e7c9cb2bec

Observation f588b94a-fc8f-49ea-8df6-56381841e2ab · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Fully Offline Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:01.099532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:01.099532Z digest=sha256:98e3849b1f2b34b9961f2feb46b193ddf54e941ba75b01dc0acf70babd87a7ee

Observation 763cabbd-dca8-4f36-9bd2-80503a8a572c · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.993087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:01.205003Z digest=sha256:2997468e3ed8e9cbd4249b89628b55baaf1c3d934031fc4a2169796fa837e406

Observation 7f2bc730-51c6-4fd0-bd3b-dfa8cec553e6 · outbound

This paper cites Discovered policy optimisation.Advances in Neural Information Processing Systems, 35:16455–16468, 2022.

Fully Offline Reinforcement Learning Discovered policy optimisation.Advances in Neural Information Processing Systems, 35:16455–16468, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:15.846851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:01.369676Z digest=sha256:28c6165a3cfd6a94ec853cd95b18be633cff7b1d01b545221e8f5fd41eda7770

Observation c8edee02-d5e4-4aca-8ff9-e8d4973c8c16 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.734238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:01.541553Z digest=sha256:54af6e3f694ae7f79c8cf99ecfbd6d31eec93786ccbd12554086af025d96ca3b

Observation 0efa23b4-2053-4f0e-bc2e-4046f7604aa4 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.552146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:01.622269Z digest=sha256:3db7d5fc26bf56017967404f96447473442fd1f22a599e04da9651ddad668d3a

Observation 109b40e2-9355-43a8-919c-180a8a979deb · outbound

This paper cites Reinforcement learning: An overview, 2024.

Fully Offline Reinforcement Learning Reinforcement learning: An overview, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:01.713071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:01.713071Z digest=sha256:6a6c58ea7e3796947f87a99c837c7863deacdb19aa3587748a8726be18e57d9f

Observation faa5c574-0182-4877-94e7-37959d84b22d · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.280859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:01.806401Z digest=sha256:9dc407153f7638cab5ebdd6b39ef4e84aea8b4008865fbbba168bdab82b5500e

Observation 5ad12ac9-c509-4e25-b824-4f25a4667b3f · outbound

This paper cites Randomized prior functions for deep reinforcement learning.

Fully Offline Reinforcement Learning Randomized prior functions for deep reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:15.019714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:01.878513Z digest=sha256:bf09ab3a7f2f7f3041ddc5b5fd62877f0ba138176902c8e697f0fb1ba63c7486

Observation 3f69e0b2-8fea-4d05-92fa-65b796d5717b · outbound

This paper cites Hyperparameter Selection for Offline Reinforcement Learning.

Fully Offline Reinforcement Learning Hyperparameter Selection for Offline Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.204750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.204750Z digest=sha256:d0058761f6b90a8bf2ee0eb566ce64dd2b99625ce29f9499ac5d7461083685ab

Observation 3cdb724d-205c-4577-bfcf-229b3a585e22 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:14.501296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:02.297041Z digest=sha256:e02a4c1617255ace529011e4fda3501f29f036e38353b6f2f8ac3f94274bc661

Observation 01135a97-f1f5-4186-8595-cbb430cf666a · outbound

This paper cites Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Programming.

Fully Offline Reinforcement Learning Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Programming

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:14.254455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:02.466920Z digest=sha256:dd42ff4afc6ed0c51c7fa14e6aba59c45a43e7757a912349ddeaf426028e5453

Observation d3db7484-5096-47a1-8f1d-99d7ad8a2b2e · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:13.976552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:02.570369Z digest=sha256:ef0402d63ec055520fa705596c86d51b5042a68c80b2445abc92108271906150

Observation a6c1a408-8b9c-43d6-aae5-a7adf262494e · outbound

This paper cites Roberts and Jeffrey S.

Fully Offline Reinforcement Learning Roberts and Jeffrey S

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.646018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.646018Z digest=sha256:b20081cc2ae0147a76178f6ada8f22e282bc6498499e10f80ad3e849f94eab25

Observation 6548a700-b2b2-4587-8ecf-bf30c365eb12 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Fully Offline Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.763290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.763290Z digest=sha256:a34212ffa48e94f422ff0e355e9ea8ec1e1dfac8966b8f4918ef39596dac8400

Observation c9630eac-592d-458e-962e-7de77b2b2c11 · outbound

This paper cites The edge-of-reach problem in offline model-based reinforcement learning.

Fully Offline Reinforcement Learning The edge-of-reach problem in offline model-based reinforcement learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.702817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:02.940355Z digest=sha256:4804b61e85a547a94958aa79fba97d873f2cd366f51e15843ec7169cc3e1945d

Observation a4892ae3-3add-4651-956c-4ebe72a9e2f7 · outbound

This paper cites Smallwood and Edward J.

Fully Offline Reinforcement Learning Smallwood and Edward J

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.474924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:03.098532Z digest=sha256:0a91c6a3fb9b4fa6c7941e31ee88bf7ae7dac6731bc5c06b7d6bf27d973919ae

Observation 3dbaf9de-ba74-42ea-a427-5573b308c687 · outbound

This paper cites A Strong Baseline for Batch Imitation Learning.

Fully Offline Reinforcement Learning A Strong Baseline for Batch Imitation Learning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:16:07.261939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:03.231306Z digest=sha256:8379b4c0911035f9e203f7a3a02c454c210ef16deb8cbf106e1f6515a61bc562

Observation 47b7cc56-f8b3-4513-a944-01b1baf5d6da · outbound

This paper cites Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Scholkopf, and Gert R.

Fully Offline Reinforcement Learning Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Scholkopf, and Gert R

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.245354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:03.346225Z digest=sha256:d71aff457499c0fc1a474149c7bad9eb88af093f968bd080773f04afc5f49faf

Observation 9d4ce3f3-a4e3-43ae-9e00-b4997b3e5697 · outbound

This paper cites Model-Bellman inconsistency for model-based offline reinforcement learning.

Fully Offline Reinforcement Learning Model-Bellman inconsistency for model-based offline reinforcement learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.080447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:03.495666Z digest=sha256:12330f03251743a18e3fde6de98b3394976b3c177d9344df0909276ea124fc20

Observation 6f3323e5-478d-4efc-99b4-5fca93c0204d · outbound

This paper cites Sutton and Andrew G.

Fully Offline Reinforcement Learning Sutton and Andrew G

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:12.860042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:03.611536Z digest=sha256:3ffcaf601a10529da0fb3cd63032d7dcb4e04e6fcba50324f016b3bfda1dcc12

Observation 00410bc4-e85a-44a8-81b3-59f65070875b · outbound

This paper cites Algorithms for Reinforcement Learning.Synthesis Lectures on Artificial Intelligence and Machine Learning, 4(1):1–103, 2010.

Fully Offline Reinforcement Learning Algorithms for Reinforcement Learning.Synthesis Lectures on Artificial Intelligence and Machine Learning, 4(1):1–103, 2010

Reference 64

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.532252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:03.712384Z digest=sha256:aaa473ce0dd9e1b916ad800fb05506831a27e633913f2442dc85428474a6abe6

Observation 148591fc-2b94-4a6b-95e3-f4149d0dd418 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning.

Fully Offline Reinforcement Learning Revisiting the minimalist approach to offline reinforcement learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:12.677528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:03.845273Z digest=sha256:345c9bcb563ece6ada5960844b378d351f08a47902a7fe38fa3016cf061a9e8a

Observation 3f1d74b1-20f4-4f2d-b848-9d4d2ddcc294 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:12.382724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:03.993853Z digest=sha256:6793f23cf8fab2d678730e2f56e37cf4c0c0189a33b2bfac11179a7e282151a7

Observation 2c8874e2-5f69-48de-8f6f-a726461790cb · outbound

This paper cites Kass, and Joseph B.

Fully Offline Reinforcement Learning Kass, and Joseph B

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:12.102869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:04.077806Z digest=sha256:c0dcc78232c6d208b7e37795cbe95cacf922ba6324c1dfca19a1d30b16c1cf1f

Observation 52eacba7-dd3b-435e-9019-4de3411263b5 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 68

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.370696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:04.164651Z digest=sha256:cb0c894773924ac425b22f3b09ebc3a82d1f25d77235476abfe47e927a8c8fd8

Observation 200ba08e-31ef-46ab-9499-8a18109ff2ee · outbound

This paper cites Information rates of nonparametric gaussian process methods.J.

Fully Offline Reinforcement Learning Information rates of nonparametric gaussian process methods.J

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:11.719273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:04.247779Z digest=sha256:30684fac5bcdf97501f9d3f0f2ecdeb14177845b52d80ffff05a6273cf0c5cc9

Observation a5ff4e8f-5141-4868-a201-40cd77bbf3f9 · outbound

This paper cites No more pesky hyperparameters: Offline hyperparameter tuning for RL.Transactions on Machine Learning Research, 2022.

Fully Offline Reinforcement Learning No more pesky hyperparameters: Offline hyperparameter tuning for RL.Transactions on Machine Learning Research, 2022

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:11.312056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:04.382384Z digest=sha256:baadf69385eb1e95b85a5b5769be2c38078c2279573454c7dffd0e4e1f062390

Observation 6db256f6-ee66-49ff-94bc-74333915930d · outbound

This paper cites Foster, and Sham M.

Fully Offline Reinforcement Learning Foster, and Sham M

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:10.928204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:04.503412Z digest=sha256:f8da6845c5790d48f58a809a3af6ebec2a46e639dde198728d2d64cf5b80b219

Observation 377f2112-f27e-47e5-b138-44b928cd80b0 · outbound

This paper cites Differential-space.Journal of Mathematics and Physics, 2(1-4):131–174, 1923.

Fully Offline Reinforcement Learning Differential-space.Journal of Mathematics and Physics, 2(1-4):131–174, 1923

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:04.598127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:04.598127Z digest=sha256:4e2f3c40c3eced4b30a09df815c33ddb883fe2b2379223a0bceb20243c6bccfe

Observation d1f9866b-e8e5-4644-b4b1-5d10a190d644 · outbound

This paper cites Information-theoretic determination of minimax rates of convergence.The Annals of Statistics, 27(5):1564 – 1599, 1999.

Fully Offline Reinforcement Learning Information-theoretic determination of minimax rates of convergence.The Annals of Statistics, 27(5):1564 – 1599, 1999

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:04.682823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:04.682823Z digest=sha256:bd5f52d48b2d3f1631b5b7849d3b3743f6fe069943a90be7b5fbbabd2a6228f3

Observation 32d5bb05-0908-42bb-9fad-fc95a8b99a71 · outbound

This paper cites Mopo: Model-based offline policy optimization.

Fully Offline Reinforcement Learning Mopo: Model-based offline policy optimization

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:10.584652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:04.761857Z digest=sha256:52014da0720677e7290d490b0a1fbf9523565fedb2baba7288e2ffb76d0d0d79

Observation c848b3c1-b773-4246-9597-50e98808cd6f · outbound

This paper cites Combo: Conservative offline model-based policy optimization.

Fully Offline Reinforcement Learning Combo: Conservative offline model-based policy optimization

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:10.295053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:04.859926Z digest=sha256:26cbeab70396b4d7658a2ac5324f24abee571cb06fef9cd0503416fa3140a843

Observation d95966f2-0215-4863-a16c-934f6f31a520 · outbound

This paper cites On the importance of hyperparameter optimization for model-based reinforcement learning.

Fully Offline Reinforcement Learning On the importance of hyperparameter optimization for model-based reinforcement learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:09.953911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:04.932417Z digest=sha256:5ec50e3cb60f7cf6d6045e78370c836cde8393a15716dc72b196a0c6e1b7421e

Observation 01f6eb3b-010c-4d61-ab64-ee67d6026ba2 · outbound

This paper cites Varibad: A very good method for bayes-adaptive deep rl via meta- learning.

Fully Offline Reinforcement Learning Varibad: A very good method for bayes-adaptive deep rl via meta- learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:09.714445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:05.017057Z digest=sha256:43521b52bb8935a3199994f901222638cd42870bb4456423b410be36d4f09e67

Observation e4a4b1c3-7783-4bf3-86a2-7dba28d76125 · outbound

This paper cites N−1X i=0 1 N D 2 log(2π) + 1 2 D−1X d=0 logσ 2 θd (xi) + (yid −µ θd (xi))2 σ2 θd (xi) !!# , =Ei∼UN.

Fully Offline Reinforcement Learning N−1X i=0 1 N D 2 log(2π) + 1 2 D−1X d=0 logσ 2 θd (xi) + (yid −µ θd (xi))2 σ2 θd (xi) !!# , =Ei∼UN

Reference 78

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:06.856195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:05.082852Z digest=sha256:c15bcd9f46dcec5ac845c447119e15ac447ad3e1e4212335d40e2882ea32d621

Observation 584f4e18-5a1d-4826-9888-eb69dfcfea54 · outbound

This paper cites Eθ∼PΘ(DN ).

Fully Offline Reinforcement Learning Eθ∼PΘ(DN )

Reference 83

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:16:06.605584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:05.159203Z digest=sha256:407c8b6bbda5850a8451ab424a75f9073c52c48eabd669ef2c3943af825a9eb4

Observation d995e4c3-886c-44f7-9c7a-9ca3fbf81918 · outbound

This paper cites doi: 10.1093/oso/9780198504856.003.0002.

Fully Offline Reinforcement Learning doi: 10.1093/oso/9780198504856.003.0002

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.660083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.660083Z digest=sha256:0a45f8db76b3ba544d38dc436de64d58514ecedadb2681e45b8ddefbe18351c0

Observation aa4e3f11-d59b-4bd4-a552-fd3592fb27fd · outbound

This paper cites URL https://books.google.co.uk/books?id= s6mVlgEACAAJ.

Fully Offline Reinforcement Learning URL https://books.google.co.uk/books?id= s6mVlgEACAAJ

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.621497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:57.884088Z digest=sha256:44db94736731563bd6b9bdace34f24ae6703c705a3b84cfa2836fe0505d03bac

Observation 1e434d6f-04c0-42fb-b1dc-edaa8c4dbd63 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:17.599212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:15:59.744897Z digest=sha256:7b76cce14d51d7cfb8d0403fbf6bce17023390ec5c5204861dbe1c58da0a9590

Observation 8c85133c-fff4-4179-b7a6-273632e7103f · outbound

This paper cites URL http://papers.nips.cc/paper/8080- randomized-prior-functions-for-deep-reinforcement-learning.pdf.

Fully Offline Reinforcement Learning URL http://papers.nips.cc/paper/8080- randomized-prior-functions-for-deep-reinforcement-learning.pdf

Reference 8629

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:14.762978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:16:02.029997Z digest=sha256:85cba5ca88f2507ec8853597c37a5cf6118c1321c2b835b6b4b6a0da36a8def6

Pith citing papers

Observation 7a38f31e-998b-44d0-a06c-e8f6a48dbd1e · inbound

Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism cites this paper.

Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism Fully Offline Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-17T00:20:38.764461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T01:46:16.923948Z digest=sha256:c2e3d3bb8a0b1c1d5fc6dc9623d2ac69c234288afb71827f9b403bf23d2b8823

Observation 00b3f2c1-259b-40d1-9384-9e49144431e9 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Fully Offline Reinforcement Learning

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:39.829506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:39.829506Z digest=sha256:8fdefa3eeb707efaf0702707ec0500c1b987c2a981a477b95845a7236cb5621e