Pith. sign in

Paper Citation Record · LEDGER

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling

As of 10 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2502.06491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06491 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:22:57.334192Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 026cb979-7686-4b92-8880-2ee95fe78810 · outbound

This paper cites an unresolved cited work.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:22:57.743081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.203760Z digest=sha256:e330b7e603ceca12df68f698c678772c95da2fbe0968b8e41b8c10289ee39e1f

Observation 55d23ecb-21c4-4d8d-989e-ef35a4696425 · outbound

This paper cites De- cision transformer: Reinforcement learning via sequence modeling.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling De- cision transformer: Reinforcement learning via sequence modeling

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.705723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.216414Z digest=sha256:0ca70230d35ebe723995e6f9d68fe28c4c2714582ad824f3595701637a382610

Observation 4269d9f1-b62e-46d1-bdca-f7fb7b3a5dd8 · outbound

This paper cites UMBRELLA: Uncertainty-Aware Model-Based Offline Reinforcement Learning Leveraging Planning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling UMBRELLA: Uncertainty-Aware Model-Based Offline Reinforcement Learning Leveraging Planning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:57.223190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:22:57.223190Z digest=sha256:f6bbb70222f1a1315c74979de1fc2aa84fe1d7490f1a29bb939264185aea2649

Observation 70be4e63-6abf-45d9-ab32-ecd9e84a69c1 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:57.226625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:22:57.226625Z digest=sha256:f8e8d63146b7506bdbf2d39029603471bffca211ed00857caabaf5f3927e2eb9

Observation 29e74946-14b7-4ce1-95f6-9cd394fd2d09 · outbound

This paper cites Bridging the data gap between training and inference for unsupervised neural machine translation.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Bridging the data gap between training and inference for unsupervised neural machine translation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.670597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.235914Z digest=sha256:124eb7e8d23fa48069e0fa0973d399647b002375044816f6256bc5196be3e270

Observation 18f60e09-4984-470e-b489-ca6c1d27ce53 · outbound

This paper cites Offline reinforcement learning as one big se- quence modeling problem.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Offline reinforcement learning as one big se- quence modeling problem

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.645403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.245241Z digest=sha256:b1e0070859c6f55f99b3ec2a5df8509d6cc0139594ef1b7abc06e19a651fd870

Observation 6280b86b-f4ff-4442-9386-8508074da6a1 · outbound

This paper cites CTRL: A Conditional Transformer Language Model for Controllable Generation.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling CTRL: A Conditional Transformer Language Model for Controllable Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:57.248534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:22:57.248534Z digest=sha256:a6a5fed44d8d303f6b0f1ab518252a690fa035973f07787aaf9f70f11780b355

Observation 7f3a3ed9-5e91-4949-9333-ac83f829dc08 · outbound

This paper cites Morel: Model-based offline reinforcement learning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Morel: Model-based offline reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.636829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.252753Z digest=sha256:5b7b49a56e6f9334571aac2e7a6e146fe79edb6e99ca97d90019f8eb9e1f278e

Observation 5c4afab4-e449-4a73-9b09-458d5cbbe853 · outbound

This paper cites Offline reinforcement learning with im- plicit q-learning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Offline reinforcement learning with im- plicit q-learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.627969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.256445Z digest=sha256:65b5802419b84ab049c110601e448516d6d21baad95ff9e620cd9d22a642cc1b

Observation 41586ed2-9973-4f06-a537-f562519308fd · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Conservative q-learning for offline reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.619426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.259885Z digest=sha256:da80abe815ea8944e57a8e3882b41dd5f5f08b9fc8af120aeef41620d37ed65d

Observation 9a951ff5-7ddd-471d-a61b-333113962602 · outbound

This paper cites Multi-game decision transformers.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Multi-game decision transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.611076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.263329Z digest=sha256:8ddc95dab0fc8a466bcfc6ee6da2584e3bcd721e83d743e5715682075b7519f3

Observation 717ee278-1cab-4ba7-832c-ac210c4d99b9 · outbound

This paper cites Distribution- conditioned adversarial variational autoencoder for valid instrumental variable generation.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Distribution- conditioned adversarial variational autoencoder for valid instrumental variable generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.601910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.266857Z digest=sha256:39a3cc8f1e4b06e2ad5d47d10a01b66f0223fede2bae6a24d4d3a284adef1f86

Observation 92cd01e8-920c-40c4-be8c-f1d1247e33d2 · outbound

This paper cites Diffstitch: Boosting of- fline reinforcement learning with diffusion-based trajec- tory stitching.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Diffstitch: Boosting of- fline reinforcement learning with diffusion-based trajec- tory stitching

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.592871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.270618Z digest=sha256:0e79bc7e0c01b8e74ffd984869040058668313997f84cd726834f95ca06da896

Observation c7f7a13b-1aec-4776-9d28-dcef19612457 · outbound

This paper cites Ball, Yee Whye Teh, and Jack Parker-Holder.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Ball, Yee Whye Teh, and Jack Parker-Holder

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.583625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.273848Z digest=sha256:040a2a60d53947a803322a12fd51b07f75916fe2fe14181bc107996f0134a19d

Observation fe2073c9-2ccb-4673-8f77-654dfba6d71c · outbound

This paper cites Luis, Alessandro G.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Luis, Alessandro G

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.575212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.277371Z digest=sha256:520cfe49cc73da8c5bb92d06f57fd35d12006598440b1f186226b0dee1a6adf4

Observation 1950a005-210e-4aeb-96cf-7037685de849 · outbound

This paper cites Double check your state before trusting it: Confidence- aware bidirectional offline model-based imagination.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Double check your state before trusting it: Confidence- aware bidirectional offline model-based imagination

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.566186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.280821Z digest=sha256:ce957cc4a9e9290240f667f423551a1bb385622433501866b2e191a0dc63db4d

Observation 61e7ecd6-f73c-4b00-9986-25798328309e · outbound

This paper cites Reining generalization in offline re- inforcement learning via representation distinction.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Reining generalization in offline re- inforcement learning via representation distinction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.556304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.284177Z digest=sha256:8f9eb91c140b53c55e0fee5e546854c492e0ea3c97872a2943c13d027ee09238

Observation 8c61a2e2-8a3a-498f-9eed-d19572ee22c4 · outbound

This paper cites Offline Imitation Learning with Model-based Reverse Augmentation.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Offline Imitation Learning with Model-based Reverse Augmentation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:22:57.369429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.287578Z digest=sha256:3400dea26381bcdc09b6a43ad31fd00455d360feccd51d6e6570c46d661a5835

Observation 01e72fb7-2537-4b69-8fa2-00159d6e4547 · outbound

This paper cites Model-bellman in- consistency for model-based offline reinforcement learn- ing.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Model-bellman in- consistency for model-based offline reinforcement learn- ing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.547316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.291356Z digest=sha256:2cb49df4b25cc4f318f5c1893750318a19344f015b9eb5a6d5baa37619f495a4

Observation c2d81fe2-a9c1-4f6d-8151-5914df624f47 · outbound

This paper cites an unresolved cited work.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:22:57.538430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.294834Z digest=sha256:439d21283cae569d2c6c950da4d36adc443dea0ea63ece5301b127b3804d4561

Observation 15d7fb2a-a021-40a5-b218-a061e4b2056b · outbound

This paper cites Self-correcting models for model-based reinforcement learning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Self-correcting models for model-based reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.529437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.298237Z digest=sha256:54980ba33a468f562dca4855459f6247ca80dc4fcdc41b1bcc3b1dd7e9e5285c

Observation b78e3e17-89ba-4638-b1eb-edf7f8c15c02 · outbound

This paper cites Visualizing data using t-sne.Journal of machine learning research, 9(11),.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Visualizing data using t-sne.Journal of machine learning research, 9(11),

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:57.304653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:22:57.304653Z digest=sha256:a0c571370647c5c0891607b8504861db0e5e80a2b2a36e7de3ec0da95939774f

Observation 56d541da-ebea-4604-b467-bd85d01fe907 · outbound

This paper cites Offline reinforcement learning with reverse model-based imagination.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Offline reinforcement learning with reverse model-based imagination

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.495528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.310939Z digest=sha256:bb6161727a1e9f3bb7be4a0ab1157e5394aee59281158543775a0c9717215a9d

Observation c4cb93a7-0fb8-41f1-81f3-58e809ab790d · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional se- quence modelling in offline RL.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Q-learning decision transformer: Leveraging dynamic programming for conditional se- quence modelling in offline RL

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.484638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.313868Z digest=sha256:67f4671909a8b36b342352e4dc4b144b95d294c32149129d137df290842f7af4

Observation d188fd78-b098-42f4-8718-2de07b8da280 · outbound

This paper cites Pareto policy pool for model-based offline reinforcement learning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Pareto policy pool for model-based offline reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.475757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.316920Z digest=sha256:ce8cd4d25e38f25d710fc4694c76cce917dc002ec10aa5489fcdf879d425726b

Observation cde9e3bd-eef7-43ed-8821-d20a5a4b7949 · outbound

This paper cites Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.464744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.320106Z digest=sha256:6afc3aa4934282463bc32754fb3a0ce86eca2e3c26d401b769d26043a7661c85

Observation 0a3006d4-6613-4f3f-91c4-5dfd042a72a8 · outbound

This paper cites Combo: Conservative offline model-based policy opti- mization.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Combo: Conservative offline model-based policy opti- mization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.453045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.322865Z digest=sha256:f14c6d22fec888e933d47cefa0c93edb1dbca1ad3f9eb8fb6b70513e54f88005

Observation 4ec02951-1bfa-467f-8cb0-8740c0533f86 · outbound

This paper cites Model-based offline planning with trajectory prun- ing.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Model-based offline planning with trajectory prun- ing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.442149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.325731Z digest=sha256:0906521b77684b95bcd1a982de8384b02116eff9725062a165f061aba7505c65

Observation c8406b5e-fd84-4b11-90de-47f9c3e0f528 · outbound

This paper cites Uncertainty-driven trajectory truncation for data augmen- tation in offline reinforcement learning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Uncertainty-driven trajectory truncation for data augmen- tation in offline reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.430685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.328642Z digest=sha256:016fa8cf6e6773ffa6b7df6b9a8192c9b75da2ae6e956401245376f3394c9ffc

Observation c1b29786-5bdb-4317-9ae6-90c9dca0a147 · outbound

This paper cites Conditional vari- ational autoencoder for sign language translation with cross-modal alignment.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Conditional vari- ational autoencoder for sign language translation with cross-modal alignment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.419360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.331555Z digest=sha256:481588f7baffd9edab4f5047682958fce3444754ee45465a62c806fb092ca6fc

Observation 13d9d2d0-9298-4b86-92c8-54bdfe8e946b · outbound

This paper cites Is model ensemble necessary? model- based RL via a single model with lipschitz regularized value function.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Is model ensemble necessary? model- based RL via a single model with lipschitz regularized value function

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.407879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.334192Z digest=sha256:a4ed19db7dbdbc1a87d6749e26f6d09d833a952baca71d77b07e81b8b05029cd

Observation 556ad3d2-8dce-4ba1-b319-0d167b3e30f2 · outbound

This paper cites Jamieson.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Jamieson

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.505853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.307781Z digest=sha256:50f81e12cf6737ec08325d5a0060161bc0d6a5f852be3bdee98a69437c567f17

Observation a880dac2-8a39-4295-ad3c-93c285a6fd89 · outbound

This paper cites When to trust your model: Model-based policy optimization.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling When to trust your model: Model-based policy optimization

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.653888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.242035Z digest=sha256:9b5904b20f939136d3f3783f8ff84d601d49789b42fcf0209d9e2c2d14e5f27d

Observation fc4a1364-2801-4845-9c84-56e87ae577b8 · outbound

This paper cites Behavioral cloning from observation.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Behavioral cloning from observation

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.520050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.301527Z digest=sha256:1ff47293d0758741ed9252fdb4b92ee94ff6090398f44590ffe87fed3d7bda8f

Observation 8c6bcd7a-3cab-481b-92bc-ee2992815790 · outbound

This paper cites Waypoint trans- former: Reinforcement learning via supervised learning with intermediate targets.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Waypoint trans- former: Reinforcement learning via supervised learning with intermediate targets

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.733466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.207356Z digest=sha256:bbd8d6904fde2decd2fee4718109a901817928931ea4dede9ed8c203f114bdc2

Observation 166a5bd9-fec1-4769-959d-2c9967232283 · outbound

This paper cites ACT: empowering decision transformer with dynamic program- ming via advantage conditioning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling ACT: empowering decision transformer with dynamic program- ming via advantage conditioning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.679007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.232925Z digest=sha256:8fbd587e24fcf0e34c71666d8ba7a843345a596e7eb59a6a1763b4c72eb035d9

Observation 96854808-e76e-48ee-a68b-f5274a2c64e7 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Off-policy deep reinforcement learning without exploration

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.687089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.230025Z digest=sha256:4a241954720f8edfe2aff0e55f979dd4514e9fdd5b2465a0d6369f181e49189a

Observation 31eb5e0d-8161-47fc-a12a-57f08cbf02b9 · outbound

This paper cites Lapo: Latent-variable advantage-weighted policy optimization for offline rein- forcement learning.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Lapo: Latent-variable advantage-weighted policy optimization for offline rein- forcement learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.695616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.219766Z digest=sha256:d785d4d4ce72630014da92374db2dead8ba5aa1a5d56105f22fbb6f09adbb856

Observation f3934870-e60d-4ca6-bc41-691e8e65e092 · outbound

This paper cites an unresolved cited work.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:22:57.662108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.239136Z digest=sha256:04f6eca766cbbade48ac54c0b60f8444f3278492d50029043983d407bc60b1bc

Observation 56ec4d35-2a57-4f8c-8796-f22eeffb4fc8 · outbound

This paper cites Data-efficient task generalization via probabilistic model-based meta reinforcement learn- ing.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Data-efficient task generalization via probabilistic model-based meta reinforcement learn- ing

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.724076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.210404Z digest=sha256:14293e5e00698bcb4689002566f3cf6c3b69a9b9ed26c83abf5acbb66009ffd2

Observation 4b3ba3be-fe11-49f6-9117-8c0278e0b5c8 · outbound

This paper cites Blanchet, Miao Lu, Tong Zhang, and Han Zhong.

Model-Based Offline Reinforcement Learning with Reliability-Guaranteed Sequence Modeling Blanchet, Miao Lu, Tong Zhang, and Han Zhong

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:22:57.715280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:22:57.213411Z digest=sha256:de1bea32a686678957627f6e923d3e7997be2950f1c604314251deb42ab8ab60

Pith citing papers

No inbound Pith citation observations are available.