Pith. sign in

Paper Citation Record · LEDGER

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2412.14312.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14312 v3

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:24:23.487246Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31202fdf-95ac-4a0c-b951-4be5873e32a9 · outbound

This paper cites write newline.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.250361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.250361Z digest=sha256:108e1d9527169c65ccda72d5071e6074e35d8f173151f4324ea25daf0f6b6f4d

Observation c049cbc4-fc97-48da-ab2b-68620653c8dc · outbound

This paper cites S., Courville, A., and Bellemare, M.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning S., Courville, A., and Bellemare, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.619058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.256377Z digest=sha256:b35abbaa641bda3cf8cfef4e3f1897fb327796e49faadd0857f7a3c4c398f530

Observation 25a816c6-d86a-413d-bd71-f995b566de4b · outbound

This paper cites an unresolved cited work.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.260388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.260388Z digest=sha256:0f7f1f9ba95956043d569af0116dc8b800ba95b7736db23bb5939c70d9512089

Observation 055b6c35-a945-4e4d-9708-176af80a38ac · outbound

This paper cites an unresolved cited work.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:24:24.607353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.264359Z digest=sha256:bd1ab908ffe0b34a9cb27b6b4885a6d703a7a14d33284e3c4291631dc8584718

Observation 9349635a-e5ad-4c5a-9bf3-dd863c8a9f19 · outbound

This paper cites and Schaal, S.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning and Schaal, S

Reference 5

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T12:24:24.172511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.267954Z digest=sha256:8f69553f2a3e7fa998602ff03fe7e4eb9e382cdeaec32197177993053f4a1809

Observation e66e9819-d457-4d60-9163-c67407f2a904 · outbound

This paper cites Layer Normalization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Layer Normalization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.272015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.272015Z digest=sha256:af092756892b42635a4b6a8f2d87813d8a5feac445494ff5280dfe152a095645

Observation 51c7b76e-ca68-490e-b910-5003ed4348b9 · outbound

This paper cites J., Smith, L., Kostrikov, I., and Levine, S.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning J., Smith, L., Kostrikov, I., and Levine, S

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.595513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.277055Z digest=sha256:725b30c5620200d7471308c7a206333b619fe437d3b930649442021079890bc7

Observation 8e6b2bd3-6170-401e-9d1d-84e0dd096dd2 · outbound

This paper cites M., Gebru, T., McMillan-Major, A., and Shmitchell, S.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning M., Gebru, T., McMillan-Major, A., and Shmitchell, S

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.281329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.281329Z digest=sha256:6dd0b84d7951e5f83eba352b4c0f3a2c12d73e0d3182484a72530b275194ab1f

Observation d4c6f1ec-2034-42cf-94b6-8f7a803c6a4d · outbound

This paper cites J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., Vander P las, J., Wanderman- M ilne, S., and Zhang, Q.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., Vander P las, J., Wanderman- M ilne, S., and Zhang, Q

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.285322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.285322Z digest=sha256:8652b78ef858ededcb18853f39f7f7a4eba65a5cbb1ca5282566d7d1df8f4fe4

Observation 16858469-0400-4ebf-8be7-8afb840ce37e · outbound

This paper cites OpenAI Gym.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning OpenAI Gym

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.289322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.289322Z digest=sha256:fd53579599adb842783ffe0862916fbfcd8ac05ea8a196083b0015ff6bf068ae

Observation 08c422c3-3c26-4ac7-b35c-adc7a8159ae9 · outbound

This paper cites Sample-efficient reinforcement learning with stochastic ensemble value expansion.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Sample-efficient reinforcement learning with stochastic ensemble value expansion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.577716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.293432Z digest=sha256:dda109ca605912b2a929e45720e6c41b752a126ccff13ff473d172efb6b21ce6

Observation bad5d73f-ab3c-4f1b-afa3-1a048313de29 · outbound

This paper cites G., and Silver, D.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning G., and Silver, D

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.565813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.297288Z digest=sha256:865807a14b7056f4df011de74f79da8fcc0f709c2173654f0dd75c1818d3bace

Observation 8e21eb4d-2ddf-4f97-824c-4df356c8e3db · outbound

This paper cites an unresolved cited work.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:24:24.555082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.301171Z digest=sha256:20ef2731f290af287b7a2db104f87ab61586b006ecf453a3c37c385afb2b8b03

Observation 61d81e54-88b5-4d08-97bf-9c7c951a6fb0 · outbound

This paper cites Dyna-style model-based reinforcement learning with model-free policy optimization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Dyna-style model-based reinforcement learning with model-free policy optimization

Reference 14

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T12:24:23.966181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.305689Z digest=sha256:86ca2d45947d1642aeb2633b7c8ce3aa5c7732f0f854760312996d7c6f3335d8

Observation da59c95c-2f3c-4fd8-a635-6538bbacadd6 · outbound

This paper cites G., and Courville, A.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning G., and Courville, A

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.544249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.309610Z digest=sha256:1f2bef71d76cc883d6685d3f520d2d3ca23463eb6949b4bfaeabc54a2628b401

Observation 7da63c09-e45b-4e8c-8a6a-eb698e74db61 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Addressing function approximation error in actor-critic methods

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.533847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.313519Z digest=sha256:3cf14448fbcde24713563411f9c34e7f8062ab6b9720b3cfa61d1a2d268b701c

Observation 336df748-4204-4ae1-bc18-c3ac8aef2926 · outbound

This paper cites Simplifying model-based rl: learning representations, latent-space models, and policies with one objective.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Simplifying model-based rl: learning representations, latent-space models, and policies with one objective

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.522697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.318786Z digest=sha256:1aeec4a183a636b5936148d143bdb01e569f032b2d5af57daf04603b4c60f95b

Observation fd95f24f-db71-43bd-9b9d-9432b389fd2e · outbound

This paper cites Continuous deep q-learning with model-based acceleration.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Continuous deep q-learning with model-based acceleration

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.511123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.322562Z digest=sha256:7f8ae7e9f852f7b678a7c16f22726d76813c70529cd3ff360c5ea15e54f22f6a

Observation 4bac0b49-657c-4837-b044-de808bd7dbe0 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.326340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.326340Z digest=sha256:e0ef21a6729e9acb2355cad3875a5b866aa0ccb556fb2a9ab6a6b293e57e3862

Observation e37e8501-1df5-4d9a-a3a4-c7fc6834b4e4 · outbound

This paper cites Dream to control: learning behaviors by latent imagination.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Dream to control: learning behaviors by latent imagination

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.499061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.330692Z digest=sha256:cada05f64074a68123ee5733fe65a4514f7959101126b85bb114d6b0447077bb

Observation 512fc7d4-90e2-401b-aebb-c20323a4d466 · outbound

This paper cites Mastering diverse control tasks through world models.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mastering diverse control tasks through world models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.334574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.334574Z digest=sha256:bff19b61a38034f7638c48d9dd9205a5dd3c6f02ed8f75ef89d3e9baa4c9cfae

Observation c12c7111-984c-4c29-af52-84db5adc6da5 · outbound

This paper cites A., Hosny, A., Khodakarami, F., Waldron, L., Wang, B., McIntosh, C., Goldenberg, A., Kundaje, A., Greene, C.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Hosny, A., Khodakarami, F., Waldron, L., Wang, B., McIntosh, C., Goldenberg, A., Kundaje, A., Greene, C

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.338375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.338375Z digest=sha256:5a74c414899dabe6c5cae2b3af2c05493fbf656ae145aa5fc8abe6beac412a8a

Observation 321aa030-88b4-46d9-bdca-040ebfc6199c · outbound

This paper cites A., Su, H., and Wang, X.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Su, H., and Wang, X

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.487487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.342417Z digest=sha256:8ec3e8d27138781a872e9a3a0eaa63faec946dfab8e7521b6522fcedb038faca

Observation 2dd2e89b-2a79-47a7-b756-7dc1bdb7a3d2 · outbound

This paper cites The Effect of Planning Shape on Dyna-style Planning in High-dimensional State Spaces.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning The Effect of Planning Shape on Dyna-style Planning in High-dimensional State Spaces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.346757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.346757Z digest=sha256:9e156cbc57af3d9dafeb3308f0f4d3f75b2092a0615dab3e2771f5f5ca48d7ea

Observation 605a0198-5307-4163-a747-bed6628e1f3a · outbound

This paper cites A., Gilitschenski, I., Farahmand, A.-m., and Eaton, E.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Gilitschenski, I., Farahmand, A.-m., and Eaton, E

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.476030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.351317Z digest=sha256:54017854f0e3e523481dc3e5adb16b09a89f6be81a1235c0b9eabbe2e3830865

Observation ccc7b9e6-e63e-4bfc-adfc-89856c4b815c · outbound

This paper cites Mbpo: Model-based policy optimization, 2019.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mbpo: Model-based policy optimization, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.464607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.355133Z digest=sha256:f4acbb7376babbba37f819eda9e9d0a1026c848d1a30cf15c76138b9bf9062f1

Observation 63f6b51a-bba8-4421-b1fd-42e6614f862b · outbound

This paper cites When to trust your model: Model-based policy optimization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning When to trust your model: Model-based policy optimization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.453713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.359923Z digest=sha256:0fecdbdfd174901df51b3809ec4e20cdb0e7e8c9e14dbda92becb95e0beaf2cc

Observation ac8118e3-0368-4a1b-809d-78fa799e1ecb · outbound

This paper cites Position: Benchmarking is Limited in Reinforcement Learning Research.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Position: Benchmarking is Limited in Reinforcement Learning Research

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.363794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.363794Z digest=sha256:59959483ab62e860b4b21a5b5170e6b2f5d23b87579a2ae773ba628320c895ce

Observation 9d7a26f6-2b14-483c-9b88-671519d2ac0c · outbound

This paper cites and Boedecker, J.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning and Boedecker, J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.442103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.367848Z digest=sha256:c13c8dfa34c483894d042ee3d719839c3fe3e451e637628fc6aa9c8b301fce4a

Observation bdd8d717-bdc0-4711-b137-8a2d50dc0b28 · outbound

This paper cites Bidirectional model-based policy optimization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Bidirectional model-based policy optimization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.430788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.371890Z digest=sha256:68ec7ac0459af84563e456114219a54f3c4ef08627ad3db269eb9eaa51474b1a

Observation 2660e7cc-109b-449e-b5c2-593a61c5f2e3 · outbound

This paper cites On effective scheduling ofmModel-based reinforcement Learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning On effective scheduling ofmModel-based reinforcement Learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.419372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.375682Z digest=sha256:764d340cdefa21991613dfd3dd8c2e0529bc96fc68c444ba7123854a79c9e809

Observation e3f6eed4-39d0-47c4-90ee-3da3668beaec · outbound

This paper cites Reinforcement learning with augmented data.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Reinforcement learning with augmented data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.407272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.379359Z digest=sha256:047910cb643583d14e6c3a7e69ec18b844f74d635bca485afac224d10b717913

Observation a6f8c3df-dd88-4e57-90b5-bb1a4c2414e4 · outbound

This paper cites When to Trust Your Data: Enhancing Dyna-Style Model-Based Reinforcement Learning With Data Filter.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning When to Trust Your Data: Enhancing Dyna-Style Model-Based Reinforcement Learning With Data Filter

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:24:23.863440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.383233Z digest=sha256:f2e9aec8e12aa1e5862307a07880dc3e62e56411a02d2cf006d43ac7ec9f8a3f

Observation 85de09b5-5783-474f-bc32-e06e10de20cb · outbound

This paper cites Continuous control with deep reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Continuous control with deep reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.387283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.387283Z digest=sha256:c9ea353763635998f7acc01052cd101c7040782759ab945fe0bdfdf467b73864

Observation 0a038221-1b2c-49f6-9cb4-8b1479d39ce5 · outbound

This paper cites [Re] When to trust your model: model-based policy optimization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning [Re] When to trust your model: model-based policy optimization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.395511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.391118Z digest=sha256:e1676527376068d24efb170faea010c03d2084f48e8cf43677bb6dcf39aad2af

Observation d1a74e53-dcdb-4138-8021-f8c90d856980 · outbound

This paper cites A survey on model-based reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A survey on model-based reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.383969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.394649Z digest=sha256:ec8daf2f8a536dfa28ef8e278bffb88ecf0a3da8151d7bc38f2d6d0a8c4baf70

Observation 6fc80cd3-2bb6-4c6f-880e-06e5dfe909ed · outbound

This paper cites Understanding and preventing capacity loss in reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Understanding and preventing capacity loss in reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.372657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.398437Z digest=sha256:01ed8f97bfdc0d42cc256e5ffa1c7629ca2551aca5add2a46d533a47e8ff4233

Observation eaee252a-918c-4170-825c-80f2810cc0cd · outbound

This paper cites Overestimation, overfitting, and plasticity in actor-critic: the Bitter lesson of Reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Overestimation, overfitting, and plasticity in actor-critic: the Bitter lesson of Reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.360750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.402220Z digest=sha256:8019c31ccba0dc6f9afec0574d5bfd9872cf07d44d7728a664d4b8e2c225070b

Observation f545d2d9-bcec-49dd-88b1-fa07be1ba06d · outbound

This paper cites The primacy bias in deep reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning The primacy bias in deep reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.349395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.406259Z digest=sha256:64a08e66bb8cdcbeb702d6cc459d49c5c8d8f45b19f5f7c1f4230cc5f623f46c

Observation 0023f1d2-c5bb-49b3-ab31-091baf614736 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.409998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.409998Z digest=sha256:a4d7120b1c53566035ae7117f7957c279fde64b50760cd52e282b9ed3d3792ab

Observation 99d35366-c4ad-4da8-b331-b1f76aa7d831 · outbound

This paper cites Mind the model, not the agent: the primacy bias in model-based RL.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mind the model, not the agent: the primacy bias in model-based RL

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.338389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.417742Z digest=sha256:1b6131ffb8e5e799d90872c07d50f19e6e6a1864e4323504bd9a80f5f47691d5

Observation fa064f53-cf58-4fff-bc70-c5500eb04cbf · outbound

This paper cites EPOpt : learning robust neural network policies using model ensembles.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning EPOpt : learning robust neural network policies using model ensembles

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.326369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.421858Z digest=sha256:18f96e62fe4badf1bebda3549a94b08cb874a9142c3370461f49d80c488d77b1

Observation adb063b5-5af3-4573-9ec2-4a7938a5d6ca · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.425323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.425323Z digest=sha256:a99fc889ff4b3b9053af3445fb710e574065bc8e7cab6e013e38fb61dea6a340

Observation fe9bb814-f39b-4457-865b-6a4c15d7f29d · outbound

This paper cites Learning off-policy with online planning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Learning off-policy with online planning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.314096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.429324Z digest=sha256:6f68423dc7486395ee648b750a346f5157dea89d230c93c371f210d9386c6d37

Observation 8f83b69a-7944-4425-b67d-49585f03fe94 · outbound

This paper cites CURL : contrastive unsupervised representations for reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning CURL : contrastive unsupervised representations for reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.302236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.433345Z digest=sha256:b31cd5db0afed6227f931040c8a1c709cb27fd6ec6b2a9d083d7ff9865b07533

Observation f60c9680-34a0-490c-ade9-38be67f12759 · outbound

This paper cites an unresolved cited work.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.436853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.436853Z digest=sha256:8c0beb0c070d2fe28f3ac1f629af4bc4595f2727625268d3699e7d2792c66374

Observation 2e9fbeeb-a76a-4e36-8b9f-337035181481 · outbound

This paper cites Model regularization for stable sample rollouts.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Model regularization for stable sample rollouts

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.289386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.440340Z digest=sha256:4ae452ee1d54310453b7ee38d72c7db59f4a52fa65aca21dec7009a9557032f1

Observation 99c30342-9c85-4a5a-99d2-522a309a5e44 · outbound

This paper cites dm\_control: Software and Tasks for Continuous Control.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning dm\_control: Software and Tasks for Continuous Control

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.443886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.443886Z digest=sha256:5024d353139236bb12e31d0d81a5b9696078af028168427cb7ddbb659fd32cd5

Observation c0572604-5990-4223-94c0-095580640edf · outbound

This paper cites and Schwartz, A.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning and Schwartz, A

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.277247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.448321Z digest=sha256:0cfcbffe33698738e8aa74bcece8eca68ec744fff0a30f4355fc820ebfb7343f

Observation 769a9005-b6cf-4e3e-a9ba-d59524ddbce9 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mujoco: A physics engine for model-based control

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.265820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.452784Z digest=sha256:90385d10ccbbc819af45adde1c587a86b0b6e0ba3c9ec7dbc49dc5ac0e689c9c

Observation 7b65fa95-5ba6-47ac-a51c-0a46879514b0 · outbound

This paper cites A., Hussing, M., and Eaton, E.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Hussing, M., and Eaton, E

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.253446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.456565Z digest=sha256:1963f980d3c379d98afbcbeb0cde6b84e6d8fbf607cb66f0690a506a5f86e0f6

Observation 499c58fd-5976-4aa2-988e-f06982d918c4 · outbound

This paper cites Model-based policy optimization under approximate Bayesian inference.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Model-based policy optimization under approximate Bayesian inference

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.241358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.460825Z digest=sha256:a78ccaace0efbd9b296b17e9d657aae2a6a756b08d51cb05b69af2b84f7b6210

Observation 95a2e144-6071-440b-8710-2cce7bea435b · outbound

This paper cites Benchmarking Model-Based Reinforcement Learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Benchmarking Model-Based Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.465176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.465176Z digest=sha256:814688e8851f96f13b41b70b16e125dd3f9081420395dfc991089f98959fa97a

Observation d4175e3b-3d47-45d3-ba63-93dfa091b947 · outbound

This paper cites Live in the moment: learning dynamics model adapted to evolving policy.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Live in the moment: learning dynamics model adapted to evolving policy

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.227657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.469704Z digest=sha256:b9b9c151177c80bf92bb1d68594fb2924c975ca078f045043971ee1233d7b164

Observation 5156ca77-6341-43a8-a8cf-659949f2563d · outbound

This paper cites Accelerated policy learning with parallel differentiable simulation.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Accelerated policy learning with parallel differentiable simulation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.214542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.473557Z digest=sha256:42894ea3fb92001bd73fd35265e436751039551e1f5614b1b222b768628ebbf8

Observation a9466429-95f2-4e9c-a0c2-86a3eab5737e · outbound

This paper cites Mastering visual continuous control: improved data-augmented reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mastering visual continuous control: improved data-augmented reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.201085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.477697Z digest=sha256:12a6986d0cb0b2324ad322e12fda8e2bbd80ecbcc3f5215eafd30bfcf224ffa6

Observation d2fa609a-25a9-4c43-bc43-a34530cc16f4 · outbound

This paper cites Is model ensemble necessary? Model-based RL via a single model with Lipschitz regularized value function.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Is model ensemble necessary? Model-based RL via a single model with Lipschitz regularized value function

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.185347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.487246Z digest=sha256:a196f42314918b7c6e727c18967dede369b383ee8804a1cbf38843fd41c740a5

Pith citing papers

No inbound Pith citation observations are available.