Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:24:23.487246Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2412.14312.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:24:23.487246Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 31202fdf-95ac-4a0c-b951-4be5873e32a9 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c049cbc4-fc97-48da-ab2b-68620653c8dc · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning S., Courville, A., and Bellemare, M
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 25a816c6-d86a-413d-bd71-f995b566de4b · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 055b6c35-a945-4e4d-9708-176af80a38ac · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9349635a-e5ad-4c5a-9bf3-dd863c8a9f19 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning and Schaal, S
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e66e9819-d457-4d60-9163-c67407f2a904 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Layer Normalization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51c7b76e-ca68-490e-b910-5003ed4348b9 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning J., Smith, L., Kostrikov, I., and Levine, S
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e6b2bd3-6170-401e-9d1d-84e0dd096dd2 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning M., Gebru, T., McMillan-Major, A., and Shmitchell, S
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c6f1ec-2034-42cf-94b6-8f7a803c6a4d · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., Vander P las, J., Wanderman- M ilne, S., and Zhang, Q
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16858469-0400-4ebf-8be7-8afb840ce37e · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning OpenAI Gym
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c422c3-3c26-4ac7-b35c-adc7a8159ae9 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Sample-efficient reinforcement learning with stochastic ensemble value expansion
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bad5d73f-ab3c-4f1b-afa3-1a048313de29 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning G., and Silver, D
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e21eb4d-2ddf-4f97-824c-4df356c8e3db · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 61d81e54-88b5-4d08-97bf-9c7c951a6fb0 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Dyna-style model-based reinforcement learning with model-free policy optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation da59c95c-2f3c-4fd8-a635-6538bbacadd6 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning G., and Courville, A
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7da63c09-e45b-4e8c-8a6a-eb698e74db61 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Addressing function approximation error in actor-critic methods
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 336df748-4204-4ae1-bc18-c3ac8aef2926 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Simplifying model-based rl: learning representations, latent-space models, and policies with one objective
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd95f24f-db71-43bd-9b9d-9432b389fd2e · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Continuous deep q-learning with model-based acceleration
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4bac0b49-657c-4837-b044-de808bd7dbe0 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Soft Actor-Critic Algorithms and Applications
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e37e8501-1df5-4d9a-a3a4-c7fc6834b4e4 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Dream to control: learning behaviors by latent imagination
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 512fc7d4-90e2-401b-aebb-c20323a4d466 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mastering diverse control tasks through world models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c12c7111-984c-4c29-af52-84db5adc6da5 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Hosny, A., Khodakarami, F., Waldron, L., Wang, B., McIntosh, C., Goldenberg, A., Kundaje, A., Greene, C
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321aa030-88b4-46d9-bdca-040ebfc6199c · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Su, H., and Wang, X
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2dd2e89b-2a79-47a7-b756-7dc1bdb7a3d2 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning The Effect of Planning Shape on Dyna-style Planning in High-dimensional State Spaces
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605a0198-5307-4163-a747-bed6628e1f3a · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Gilitschenski, I., Farahmand, A.-m., and Eaton, E
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ccc7b9e6-e63e-4bfc-adfc-89856c4b815c · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mbpo: Model-based policy optimization, 2019
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 63f6b51a-bba8-4421-b1fd-42e6614f862b · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning When to trust your model: Model-based policy optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ac8118e3-0368-4a1b-809d-78fa799e1ecb · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Position: Benchmarking is Limited in Reinforcement Learning Research
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d7a26f6-2b14-483c-9b88-671519d2ac0c · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning and Boedecker, J
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bdd8d717-bdc0-4711-b137-8a2d50dc0b28 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Bidirectional model-based policy optimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2660e7cc-109b-449e-b5c2-593a61c5f2e3 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning On effective scheduling ofmModel-based reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e3f6eed4-39d0-47c4-90ee-3da3668beaec · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Reinforcement learning with augmented data
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a6f8c3df-dd88-4e57-90b5-bb1a4c2414e4 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning When to Trust Your Data: Enhancing Dyna-Style Model-Based Reinforcement Learning With Data Filter
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 85de09b5-5783-474f-bc32-e06e10de20cb · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Continuous control with deep reinforcement learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a038221-1b2c-49f6-9cb4-8b1479d39ce5 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning [Re] When to trust your model: model-based policy optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d1a74e53-dcdb-4138-8021-f8c90d856980 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A survey on model-based reinforcement learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6fc80cd3-2bb6-4c6f-880e-06e5dfe909ed · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Understanding and preventing capacity loss in reinforcement learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eaee252a-918c-4170-825c-80f2810cc0cd · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Overestimation, overfitting, and plasticity in actor-critic: the Bitter lesson of Reinforcement learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f545d2d9-bcec-49dd-88b1-fa07be1ba06d · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning The primacy bias in deep reinforcement learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0023f1d2-c5bb-49b3-ab31-091baf614736 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d35366-c4ad-4da8-b331-b1f76aa7d831 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mind the model, not the agent: the primacy bias in model-based RL
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fa064f53-cf58-4fff-bc70-c5500eb04cbf · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning EPOpt : learning robust neural network policies using model ensembles
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation adb063b5-5af3-4573-9ec2-4a7938a5d6ca · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe9bb814-f39b-4457-865b-6a4c15d7f29d · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Learning off-policy with online planning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8f83b69a-7944-4425-b67d-49585f03fe94 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning CURL : contrastive unsupervised representations for reinforcement learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f60c9680-34a0-490c-ade9-38be67f12759 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9fbeeb-a76a-4e36-8b9f-337035181481 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Model regularization for stable sample rollouts
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 99c30342-9c85-4a5a-99d2-522a309a5e44 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning dm\_control: Software and Tasks for Continuous Control
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0572604-5990-4223-94c0-095580640edf · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning and Schwartz, A
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 769a9005-b6cf-4e3e-a9ba-d59524ddbce9 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mujoco: A physics engine for model-based control
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7b65fa95-5ba6-47ac-a51c-0a46879514b0 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning A., Hussing, M., and Eaton, E
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 499c58fd-5976-4aa2-988e-f06982d918c4 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Model-based policy optimization under approximate Bayesian inference
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 95a2e144-6071-440b-8710-2cce7bea435b · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Benchmarking Model-Based Reinforcement Learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4175e3b-3d47-45d3-ba63-93dfa091b947 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Live in the moment: learning dynamics model adapted to evolving policy
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5156ca77-6341-43a8-a8cf-659949f2563d · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Accelerated policy learning with parallel differentiable simulation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9466429-95f2-4e9c-a0c2-86a3eab5737e · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Mastering visual continuous control: improved data-augmented reinforcement learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d2fa609a-25a9-4c43-bc43-a34530cc16f4 · outbound
Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning Is model ensemble necessary? Model-based RL via a single model with Lipschitz regularized value function
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.