Pith. sign in

Paper Citation Record · LEDGER

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies

As of 13 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2509.05735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.05735 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:09:29.269978Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee287cfc-e534-4381-acbc-8d2791ef3b3f · outbound

This paper cites Combating the Compounding-Error Problem with a Multi-step Model.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Combating the Compounding-Error Problem with a Multi-step Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.205646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.205646Z digest=sha256:c8c667f1bd38f2f29eeb7748938776455855cf1ba8aedbf5406cea4949cb96c1

Observation 32071725-ddd9-4c2f-b5c2-facf1b812fed · outbound

This paper cites Below, the training of the world model M includes training all components in Eq.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Below, the training of the world model M includes training all components in Eq

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.425275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.261773Z digest=sha256:468b5ff78d13d8c2c0ad3637a21f4166d823b4d9e508b8a6f22e7790f04415ab

Observation 86bb3937-28c7-4ce8-90fb-f09a9540103e · outbound

This paper cites Regularized Behavior Value Estimation.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Regularized Behavior Value Estimation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.215472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.215472Z digest=sha256:2aced9bd142af5e94786465f29a1eba0933bfd83f6f16dfe2dd9db1a9dc51a8b

Observation 783a355f-10d1-4167-9e5b-b7b0dd5e819c · outbound

This paper cites A Survey on Offline Model-Based Reinforcement Learning.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies A Survey on Offline Model-Based Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:09:29.343895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.225492Z digest=sha256:3c44cd5a1072ffefb2a4d0b15d07fa95120750020e0535c18f9e3cf7ec9bcc1b

Observation aeccf0ce-10ea-4c16-809e-db11a1d5e6ce · outbound

This paper cites Discor: Corrective feedback in reinforcement learning via distribution correction.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Discor: Corrective feedback in reinforcement learning via distribution correction

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.454915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.234724Z digest=sha256:7115479270930f4da9ba05a15d8b24015fd2cdb0805e3c4b700cf765e0510b2e

Observation 49a3ba7d-8906-4510-b54e-b22e15165640 · outbound

This paper cites Andrew Bagnell.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Andrew Bagnell

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.445196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.238447Z digest=sha256:b2c1e1f82afa86651fd7a02394d659100298e65c45d5cb9a4fe71d88637dfe49

Observation d18d4844-7d4d-422d-8037-d70e1072919e · outbound

This paper cites The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:09:29.324068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.241251Z digest=sha256:0de55bdcc5d83efc947481dd03656c4b1362fffec4d9572d79d73b316a69f6b6

Observation 372920bb-6144-4d2a-9de6-51c63bbeb966 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Understanding the performance gap between online and offline alignment algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.245182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.245182Z digest=sha256:248148f84e0f6e07eea4aafa889cbd2248d68a6e23fc8404d3ef11749f7f1d41

Observation ea4bdc3f-4d47-4feb-b7de-06859784cd58 · outbound

This paper cites Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.248222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.248222Z digest=sha256:a7a680cd745c57d0ab8008f89ff85fa84f884f49882ebbd0877005b5d5049622

Observation 05cd2f78-30c3-45ae-a746-040560acf259 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.251166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.251166Z digest=sha256:24642efe8217b9f3534289257dd5daf7355c38f762154d2093500ff2791dbfe3

Observation 8ba0897d-c032-412c-8d51-8790df24a9bc · outbound

This paper cites A Implementation Details A.1 Runtime Overview Our experiments comprised approximately 2000 runs, totaling 20000 GPU hours.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies A Implementation Details A.1 Runtime Overview Our experiments comprised approximately 2000 runs, totaling 20000 GPU hours

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.436594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.258053Z digest=sha256:1e82d4170f348b804ec3939bcb486d490ca8f6767c124d395b06e33e26b0c628

Observation ab6cea53-7737-43db-94cd-71264f022b31 · outbound

This paper cites -same”, while the different model initialization is marked with “-diff.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies -same”, while the different model initialization is marked with “-diff

Reference 19

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T05:09:29.415889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.265767Z digest=sha256:a15e0dbb2d40d64aac62c8dba82b20fe3d3b50853cfe55b2886fed20079a285d

Observation a571720c-4a66-4583-aae5-55809fd18630 · outbound

This paper cites an unresolved cited work.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:09:29.406556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.269978Z digest=sha256:670b37c8a2426968cdcbc5554172549fee497199fb199397e475d34881716c12

Observation 45390c4c-59ae-464e-8f66-d06cee0d787c · outbound

This paper cites Mastering Diverse Domains through World Models.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Mastering Diverse Domains through World Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.222266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.222266Z digest=sha256:950388a85485fc4d9bf808ea8eaff4da230eba7ddc3c27708041e373542e44a4

Observation 5422bfd4-e554-4c47-847b-13464b51c3e6 · outbound

This paper cites Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.228691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.228691Z digest=sha256:b309510833ec2d0237dbfd9c26c75dd1779fc9877d1e6a77967873e99304c20d

Observation 970f977b-940e-4d93-b523-185692526cbc · outbound

This paper cites Exploring generalization and adaptability of offline reinforcement learning for robot manipulation.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Exploring generalization and adaptability of offline reinforcement learning for robot manipulation

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.464306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.231579Z digest=sha256:00e048da712ffbe45711987e3c9c6aca96947e8763bfcce7dc9fb8a1d6ae5b5f

Observation 912e2045-4ae9-4daa-b764-dd6721502493 · outbound

This paper cites Boosting Offline Reinforcement Learning via Data Rebalancing.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Boosting Offline Reinforcement Learning via Data Rebalancing

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.254979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.254979Z digest=sha256:f1361253b8f440b4c144630c040fe637feb64cc4c286388f8e1b4d8c96cd9376

Observation e885a2ad-9441-46a0-af61-63eaf5aa4185 · outbound

This paper cites Learning Latent Dynamics for Planning from Pixels.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Learning Latent Dynamics for Planning from Pixels

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.219130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.219130Z digest=sha256:b2e029a00c79d1c7d6e2340790dd7dcebe05a9c7dd9d6fb9b77956bed577369a

Observation ac127ea4-8e9e-4825-89da-61919d9a4b69 · outbound

This paper cites Knowledge Transfer from Teachers to Learners in Growing-Batch Reinforcement Learning.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Knowledge Transfer from Teachers to Learners in Growing-Batch Reinforcement Learning

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T05:09:29.377909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.212419Z digest=sha256:18d3e48236fac5c4308de7a23580cf159b4760043904f026174f1ed8e5dfeeb6

Observation a1289718-f32d-456a-b445-3dc26cc19514 · outbound

This paper cites Behavioral Priors and Dynamics Models: Improving Performance and Domain Transfer in Offline RL.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Behavioral Priors and Dynamics Models: Improving Performance and Domain Transfer in Offline RL

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:09:29.389938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:09:29.209188Z digest=sha256:f799e6be553e48d76b8c06d8d437792f1560708171859cfded630bdb49f7ff33

Pith citing papers

No inbound Pith citation observations are available.