Pith. sign in

Paper Citation Record · LEDGER

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2412.14312.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14312 v3

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:24:23.487246Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31202fdf-95ac-4a0c-b951-4be5873e32a9 · outbound

This paper cites write newline.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.250361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.250361Z digest=sha256:108e1d9527169c65ccda72d5071e6074e35d8f173151f4324ea25daf0f6b6f4d

Observation c049cbc4-fc97-48da-ab2b-68620653c8dc · outbound

This paper cites S., Courville, A., and Bellemare, M.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning S., Courville, A., and Bellemare, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.619058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.256377Z digest=sha256:233cbeda1a3e5fed75c8dd5490268abcf8ba3151f55048be62181143d53e6fbd

Observation 25a816c6-d86a-413d-bd71-f995b566de4b · outbound

This paper cites an unresolved cited work.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.260388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.260388Z digest=sha256:0f7f1f9ba95956043d569af0116dc8b800ba95b7736db23bb5939c70d9512089

Observation 055b6c35-a945-4e4d-9708-176af80a38ac · outbound

This paper cites an unresolved cited work.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:24:24.607353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.264359Z digest=sha256:18816b95f1ddc42401af0f97b1ee800e10223d497712a4f846ccd24732c65b66

Observation 9349635a-e5ad-4c5a-9bf3-dd863c8a9f19 · outbound

This paper cites and Schaal, S.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning and Schaal, S

Reference 5

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T12:24:24.172511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.267954Z digest=sha256:4f090302adec5d21a841a4c696277c34ace7ece3731726b6e217c6c2b58bf994

Observation e66e9819-d457-4d60-9163-c67407f2a904 · outbound

This paper cites Layer Normalization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Layer Normalization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.272015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.272015Z digest=sha256:af092756892b42635a4b6a8f2d87813d8a5feac445494ff5280dfe152a095645

Observation 51c7b76e-ca68-490e-b910-5003ed4348b9 · outbound

This paper cites J., Smith, L., Kostrikov, I., and Levine, S.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning J., Smith, L., Kostrikov, I., and Levine, S

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.595513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.277055Z digest=sha256:fce658826bfb1b03c6045dd609005f81f83e54807a351835b47f6e56a94e8e41

Observation 8e6b2bd3-6170-401e-9d1d-84e0dd096dd2 · outbound

This paper cites M., Gebru, T., McMillan-Major, A., and Shmitchell, S.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning M., Gebru, T., McMillan-Major, A., and Shmitchell, S

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.281329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.281329Z digest=sha256:6dd0b84d7951e5f83eba352b4c0f3a2c12d73e0d3182484a72530b275194ab1f

Observation d4c6f1ec-2034-42cf-94b6-8f7a803c6a4d · outbound

This paper cites J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., Vander P las, J., Wanderman- M ilne, S., and Zhang, Q.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., Vander P las, J., Wanderman- M ilne, S., and Zhang, Q

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.285322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.285322Z digest=sha256:8652b78ef858ededcb18853f39f7f7a4eba65a5cbb1ca5282566d7d1df8f4fe4

Observation 16858469-0400-4ebf-8be7-8afb840ce37e · outbound

This paper cites OpenAI Gym.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning OpenAI Gym

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.289322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.289322Z digest=sha256:fd53579599adb842783ffe0862916fbfcd8ac05ea8a196083b0015ff6bf068ae

Observation 08c422c3-3c26-4ac7-b35c-adc7a8159ae9 · outbound

This paper cites Sample-efficient reinforcement learning with stochastic ensemble value expansion.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Sample-efficient reinforcement learning with stochastic ensemble value expansion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.577716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.293432Z digest=sha256:a08428ecc042b73355bc2c9248ef5cad0a7ee7547e5557b569d79fd5625686a9

Observation bad5d73f-ab3c-4f1b-afa3-1a048313de29 · outbound

This paper cites G., and Silver, D.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning G., and Silver, D

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.565813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.297288Z digest=sha256:4eb87082734a91d52a12324e838da9b4fda5ccb8239147d9c0f241e5c09bca58

Observation 8e21eb4d-2ddf-4f97-824c-4df356c8e3db · outbound

This paper cites an unresolved cited work.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:24:24.555082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.301171Z digest=sha256:34247d1a90e3df8368ff8b8d54c5737b59f24c493b52979fe5a79bbea79afc6e

Observation 61d81e54-88b5-4d08-97bf-9c7c951a6fb0 · outbound

This paper cites Dyna-style model-based reinforcement learning with model-free policy optimization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Dyna-style model-based reinforcement learning with model-free policy optimization

Reference 14

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T12:24:23.966181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.305689Z digest=sha256:8a8552cf75fe70203e33fa6f66cee1599d5c7b19f09d00fc3e1b705513f8d8f2

Observation da59c95c-2f3c-4fd8-a635-6538bbacadd6 · outbound

This paper cites G., and Courville, A.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning G., and Courville, A

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.544249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.309610Z digest=sha256:f8ebe6a2cb73e9e53074b672f85d27ea55e20e29e9427029a055491bc91ac0e2

Observation 7da63c09-e45b-4e8c-8a6a-eb698e74db61 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Addressing function approximation error in actor-critic methods

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.533847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.313519Z digest=sha256:3672aeef6ca207dde111cfc764659aea304053a89dfa14ea10ec60de249071e6

Observation 336df748-4204-4ae1-bc18-c3ac8aef2926 · outbound

This paper cites Simplifying model-based rl: learning representations, latent-space models, and policies with one objective.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Simplifying model-based rl: learning representations, latent-space models, and policies with one objective

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.522697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.318786Z digest=sha256:ec60d85d901f0fefaca027cb487e1fceab2ea1973aedecf9c4635211370de2d7

Observation fd95f24f-db71-43bd-9b9d-9432b389fd2e · outbound

This paper cites Continuous deep q-learning with model-based acceleration.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Continuous deep q-learning with model-based acceleration

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.511123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.322562Z digest=sha256:000f1660f66260a0bb264fe958e1dbcf7610c2b1b4a8f74b371801d4de457f94

Observation 4bac0b49-657c-4837-b044-de808bd7dbe0 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.326340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.326340Z digest=sha256:e0ef21a6729e9acb2355cad3875a5b866aa0ccb556fb2a9ab6a6b293e57e3862

Observation e37e8501-1df5-4d9a-a3a4-c7fc6834b4e4 · outbound

This paper cites Dream to control: learning behaviors by latent imagination.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Dream to control: learning behaviors by latent imagination

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.499061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.330692Z digest=sha256:f822b517501d3fc8c97b07bd2aefbe918c6bddf820de2c84198fa52e816c60b2

Observation 512fc7d4-90e2-401b-aebb-c20323a4d466 · outbound

This paper cites Mastering diverse control tasks through world models.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mastering diverse control tasks through world models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.334574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.334574Z digest=sha256:bff19b61a38034f7638c48d9dd9205a5dd3c6f02ed8f75ef89d3e9baa4c9cfae

Observation c12c7111-984c-4c29-af52-84db5adc6da5 · outbound

This paper cites A., Hosny, A., Khodakarami, F., Waldron, L., Wang, B., McIntosh, C., Goldenberg, A., Kundaje, A., Greene, C.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Hosny, A., Khodakarami, F., Waldron, L., Wang, B., McIntosh, C., Goldenberg, A., Kundaje, A., Greene, C

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.338375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.338375Z digest=sha256:5a74c414899dabe6c5cae2b3af2c05493fbf656ae145aa5fc8abe6beac412a8a

Observation 321aa030-88b4-46d9-bdca-040ebfc6199c · outbound

This paper cites A., Su, H., and Wang, X.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Su, H., and Wang, X

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.487487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.342417Z digest=sha256:3c7b30afaba083a267163029ba702ae20dae0de54df63424d25c7fa12d6993a3

Observation 2dd2e89b-2a79-47a7-b756-7dc1bdb7a3d2 · outbound

This paper cites The Effect of Planning Shape on Dyna-style Planning in High-dimensional State Spaces.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning The Effect of Planning Shape on Dyna-style Planning in High-dimensional State Spaces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.346757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.346757Z digest=sha256:9e156cbc57af3d9dafeb3308f0f4d3f75b2092a0615dab3e2771f5f5ca48d7ea

Observation 605a0198-5307-4163-a747-bed6628e1f3a · outbound

This paper cites A., Gilitschenski, I., Farahmand, A.-m., and Eaton, E.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Gilitschenski, I., Farahmand, A.-m., and Eaton, E

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.476030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.351317Z digest=sha256:b500fbf1014ef993e9226ef0485cb38097c719ca076dda822f97ca75179e242f

Observation ccc7b9e6-e63e-4bfc-adfc-89856c4b815c · outbound

This paper cites Mbpo: Model-based policy optimization, 2019.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mbpo: Model-based policy optimization, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.464607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.355133Z digest=sha256:6297d296572d8ba72c3f50eea4891dd2e2dbf2ab2a94288b48ea9b9b75753097

Observation 63f6b51a-bba8-4421-b1fd-42e6614f862b · outbound

This paper cites When to trust your model: Model-based policy optimization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning When to trust your model: Model-based policy optimization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.453713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.359923Z digest=sha256:0d098f9d41ac58b6f097db7c27eddfaaa5c33148ecf184f00573a622a8eb16f6

Observation ac8118e3-0368-4a1b-809d-78fa799e1ecb · outbound

This paper cites Position: Benchmarking is Limited in Reinforcement Learning Research.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Position: Benchmarking is Limited in Reinforcement Learning Research

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.363794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.363794Z digest=sha256:59959483ab62e860b4b21a5b5170e6b2f5d23b87579a2ae773ba628320c895ce

Observation 9d7a26f6-2b14-483c-9b88-671519d2ac0c · outbound

This paper cites and Boedecker, J.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning and Boedecker, J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.442103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.367848Z digest=sha256:5194dccfb399e69ebaadda4e25209ad4a3c863aa83bce42a353b7420a5885d93

Observation bdd8d717-bdc0-4711-b137-8a2d50dc0b28 · outbound

This paper cites Bidirectional model-based policy optimization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Bidirectional model-based policy optimization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.430788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.371890Z digest=sha256:ba0c1f190cf0e7a67b5847baf83754ec31f7278a2cc133195e5bc6298a60312e

Observation 2660e7cc-109b-449e-b5c2-593a61c5f2e3 · outbound

This paper cites On effective scheduling ofmModel-based reinforcement Learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning On effective scheduling ofmModel-based reinforcement Learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.419372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.375682Z digest=sha256:7d063e1821ee5ab8aa806e236ef1157b15c98e257c34c8f123c074d030b6fc09

Observation e3f6eed4-39d0-47c4-90ee-3da3668beaec · outbound

This paper cites Reinforcement learning with augmented data.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Reinforcement learning with augmented data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.407272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.379359Z digest=sha256:c611dd13b80d403a5c8705eab174b2106463ad513d96e373f05442e569375c05

Observation a6f8c3df-dd88-4e57-90b5-bb1a4c2414e4 · outbound

This paper cites When to Trust Your Data: Enhancing Dyna-Style Model-Based Reinforcement Learning With Data Filter.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning When to Trust Your Data: Enhancing Dyna-Style Model-Based Reinforcement Learning With Data Filter

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:24:23.863440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.383233Z digest=sha256:81cb030fcb4155ed710bd27547a4491c8ececc750d0ab37f4a2cb52e7aaf250f

Observation 85de09b5-5783-474f-bc32-e06e10de20cb · outbound

This paper cites Continuous control with deep reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Continuous control with deep reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.387283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.387283Z digest=sha256:c9ea353763635998f7acc01052cd101c7040782759ab945fe0bdfdf467b73864

Observation 0a038221-1b2c-49f6-9cb4-8b1479d39ce5 · outbound

This paper cites [Re] When to trust your model: model-based policy optimization.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning [Re] When to trust your model: model-based policy optimization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.395511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.391118Z digest=sha256:43eef1a0d006b34eaed0c2e19c65b66364aea82782105248fbb0ff5c49a77b0b

Observation d1a74e53-dcdb-4138-8021-f8c90d856980 · outbound

This paper cites A survey on model-based reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A survey on model-based reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.383969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.394649Z digest=sha256:9d8c585be8a1b741fe2dcd0314bc997247aea67b33c1a7a875e27f0a1b348945

Observation 6fc80cd3-2bb6-4c6f-880e-06e5dfe909ed · outbound

This paper cites Understanding and preventing capacity loss in reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Understanding and preventing capacity loss in reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.372657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.398437Z digest=sha256:646513b2658e32076d6e874a2098e58851dd2aae50551a3010e7307de6bbb28e

Observation eaee252a-918c-4170-825c-80f2810cc0cd · outbound

This paper cites Overestimation, overfitting, and plasticity in actor-critic: the Bitter lesson of Reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Overestimation, overfitting, and plasticity in actor-critic: the Bitter lesson of Reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.360750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.402220Z digest=sha256:25a36727067a2c5ab5a1b9231c61aaf0e103e02d86eda96e6fc0e7c00607d81a

Observation f545d2d9-bcec-49dd-88b1-fa07be1ba06d · outbound

This paper cites The primacy bias in deep reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning The primacy bias in deep reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.349395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.406259Z digest=sha256:39b5e6641285dfb17f313d3e89ab4c8cd6399d88a4522fb6bc1a331d7b2c7c4c

Observation 0023f1d2-c5bb-49b3-ab31-091baf614736 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.409998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.409998Z digest=sha256:a4d7120b1c53566035ae7117f7957c279fde64b50760cd52e282b9ed3d3792ab

Observation 99d35366-c4ad-4da8-b331-b1f76aa7d831 · outbound

This paper cites Mind the model, not the agent: the primacy bias in model-based RL.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mind the model, not the agent: the primacy bias in model-based RL

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.338389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.417742Z digest=sha256:c10570ded60a615a4cfd2ab5fefa997e76a22858650606e3337e9c29abe1887a

Observation fa064f53-cf58-4fff-bc70-c5500eb04cbf · outbound

This paper cites EPOpt : learning robust neural network policies using model ensembles.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning EPOpt : learning robust neural network policies using model ensembles

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.326369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.421858Z digest=sha256:a898f34d3926360fe6b933ac72cb7d03fa85dc8a9666b72de8d10e5ed46bca03

Observation adb063b5-5af3-4573-9ec2-4a7938a5d6ca · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.425323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.425323Z digest=sha256:a99fc889ff4b3b9053af3445fb710e574065bc8e7cab6e013e38fb61dea6a340

Observation fe9bb814-f39b-4457-865b-6a4c15d7f29d · outbound

This paper cites Learning off-policy with online planning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Learning off-policy with online planning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.314096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.429324Z digest=sha256:7b74ea512e62934ef05e1c1170aea8ee554477ae3f171d8ea291332956dfa20e

Observation 8f83b69a-7944-4425-b67d-49585f03fe94 · outbound

This paper cites CURL : contrastive unsupervised representations for reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning CURL : contrastive unsupervised representations for reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.302236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.433345Z digest=sha256:a3ddd74526abfd09daca526b5a0315212a9ac3a666e6fb705cca33ea94cd7f28

Observation f60c9680-34a0-490c-ade9-38be67f12759 · outbound

This paper cites an unresolved cited work.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.436853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.436853Z digest=sha256:8c0beb0c070d2fe28f3ac1f629af4bc4595f2727625268d3699e7d2792c66374

Observation 2e9fbeeb-a76a-4e36-8b9f-337035181481 · outbound

This paper cites Model regularization for stable sample rollouts.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Model regularization for stable sample rollouts

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.289386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.440340Z digest=sha256:a83c013d8f9d38afc8102560a2ff73f26b04125a0c319341d5fb0e055010d89d

Observation 99c30342-9c85-4a5a-99d2-522a309a5e44 · outbound

This paper cites dm\_control: Software and Tasks for Continuous Control.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning dm\_control: Software and Tasks for Continuous Control

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.443886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.443886Z digest=sha256:5024d353139236bb12e31d0d81a5b9696078af028168427cb7ddbb659fd32cd5

Observation c0572604-5990-4223-94c0-095580640edf · outbound

This paper cites and Schwartz, A.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning and Schwartz, A

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.277247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.448321Z digest=sha256:4eaa81d567df5cbbcce54b5c77892f1639b403201d8e01d1c127eb212be6298f

Observation 769a9005-b6cf-4e3e-a9ba-d59524ddbce9 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mujoco: A physics engine for model-based control

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.265820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.452784Z digest=sha256:543b9ec1d9bba284b124ef1a17edb1f8aa1c069aa29e2caffa10de0aee56085b

Observation 7b65fa95-5ba6-47ac-a51c-0a46879514b0 · outbound

This paper cites A., Hussing, M., and Eaton, E.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Hussing, M., and Eaton, E

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.253446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.456565Z digest=sha256:d17e0084ae8e344263f37f70947a25f7811e3c6713d8b9599f60daea76f4eb43

Observation 499c58fd-5976-4aa2-988e-f06982d918c4 · outbound

This paper cites Model-based policy optimization under approximate Bayesian inference.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Model-based policy optimization under approximate Bayesian inference

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.241358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.460825Z digest=sha256:989620eeb6f3570b36af0a4d0b80f479f8f6a12bd0d6bb4038a2d3660753004d

Observation 95a2e144-6071-440b-8710-2cce7bea435b · outbound

This paper cites Benchmarking Model-Based Reinforcement Learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Benchmarking Model-Based Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:24:23.465176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:24:23.465176Z digest=sha256:814688e8851f96f13b41b70b16e125dd3f9081420395dfc991089f98959fa97a

Observation d4175e3b-3d47-45d3-ba63-93dfa091b947 · outbound

This paper cites Live in the moment: learning dynamics model adapted to evolving policy.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Live in the moment: learning dynamics model adapted to evolving policy

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.227657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.469704Z digest=sha256:5534efe2b21f57487bfb37528ef812e16a51eed77f82449cb4e1321c1ee21f5e

Observation 5156ca77-6341-43a8-a8cf-659949f2563d · outbound

This paper cites Accelerated policy learning with parallel differentiable simulation.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Accelerated policy learning with parallel differentiable simulation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.214542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.473557Z digest=sha256:40118927849ebb3e363debd25d30c7f27e1bff744c97bac70e624bdb9a5d756f

Observation a9466429-95f2-4e9c-a0c2-86a3eab5737e · outbound

This paper cites Mastering visual continuous control: improved data-augmented reinforcement learning.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mastering visual continuous control: improved data-augmented reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.201085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.477697Z digest=sha256:79f6c7c1185afb8e528e033c68f07dc258893c4ed0cd9936bb4fa0fdc1160eb2

Observation d2fa609a-25a9-4c43-bc43-a34530cc16f4 · outbound

This paper cites Is model ensemble necessary? Model-based RL via a single model with Lipschitz regularized value function.

Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Is model ensemble necessary? Model-based RL via a single model with Lipschitz regularized value function

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:24:24.185347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T12:24:23.487246Z digest=sha256:9562c1074a6d9aee60bb41790f0e6505a379fd6cfe7fb62ae912e839e99b0011

Pith citing papers

No inbound Pith citation observations are available.