Pith. sign in

Paper Citation Record · LEDGER

Reinforcement learning with world model

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:1908.11494.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.11494 v4

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:18:25.102620Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c30b289a-d30d-4f23-90a4-f5fd0b49f44a · outbound

This paper cites Addressing Function Approximation Error in Actor-Critic Methods.

Reinforcement learning with world model Addressing Function Approximation Error in Actor-Critic Methods

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:24.814500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:24.814500Z digest=sha256:36a427b6bb4d6442367c333a1a56718e22289506144b009ccc654aede0f4a91a

Observation d3e4474c-660f-4f0b-85ea-146e0a1186e2 · outbound

This paper cites Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic.

Reinforcement learning with world model Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:24.835876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:24.835876Z digest=sha256:a11eadf8c9935921979efc5a3855963ca09155187367247d1dcd97e14afd1ff2

Observation 37e15af6-b280-4c48-be43-25f9573e870a · outbound

This paper cites an unresolved cited work.

Reinforcement learning with world model Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:18:25.612418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T10:18:24.882753Z digest=sha256:ea67e19ccb9fee7808c4bfa2470ca0d26b9954ef9be92374fb0cc95807c6d9fb

Observation 23a28fc0-1fc6-4a87-9258-82a4d2bd732f · outbound

This paper cites World Models.

Reinforcement learning with world model World Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:24.896081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:24.896081Z digest=sha256:9bcafa4a2fc1ddc0fb4a589727f5b9b5c488a5262ac00815e1bc7e46fd4e3e59

Observation d7f11a55-fe57-4f24-9b4c-7ce59a8e92fd · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Reinforcement learning with world model Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:24.900666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:24.900666Z digest=sha256:c461dfb7116aa71d401b14fcb706f051caffadc4b1995d0fcfdc48e699d31638

Observation e0741495-2176-46c1-8476-ce40822600bf · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Reinforcement learning with world model Soft Actor-Critic Algorithms and Applications

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:24.908511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:24.908511Z digest=sha256:fe9a0253263ec111d2b3a950371c1b65727636a0202b83cfce4981046d74e7ae

Observation f92161fa-1217-4924-bfd6-bd973296a385 · outbound

This paper cites Learning Latent Dynamics for Planning from Pixels.

Reinforcement learning with world model Learning Latent Dynamics for Planning from Pixels

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:24.925231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:24.925231Z digest=sha256:3beddbcd28bd8dd464b8cde1136d56925d7cc371fcb74aa2bee8b58672345c44

Observation 400f5975-43d4-4c70-a986-0fa70c7adccd · outbound

This paper cites and Stone, P.

Reinforcement learning with world model and Stone, P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:18:25.550726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T10:18:24.929993Z digest=sha256:ba7a40cf03686f5e76d8742315c6a1177b4e67c724b6f25190877fb5b46c764c

Observation 5e2bdabb-1414-4b60-b7d0-3b6b0dcb1eed · outbound

This paper cites an unresolved cited work.

Reinforcement learning with world model Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:18:25.535609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T10:18:24.934087Z digest=sha256:71c690b8f176ec8d8c1008df63a3b2f508b06bac2ca2270e073dd44beb864f7b

Observation 8d8b4e75-6295-4be8-a79b-7c8ccbb5dde1 · outbound

This paper cites an unresolved cited work.

Reinforcement learning with world model Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:18:25.519535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T10:18:24.938452Z digest=sha256:9ba5c9e0ec9b6256b41d1b87e09168969a6ed8dd3e5723f9da3ecaf4ce345fd4

Observation 6be48e78-51bf-4234-b2ea-83d42ce6f826 · outbound

This paper cites an unresolved cited work.

Reinforcement learning with world model Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:18:25.504441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T10:18:24.942112Z digest=sha256:beaf1de4a924b606c3b4db09108af571f56e44b05d9c0dc534feb85dd1003521

Observation 7f808845-d084-4410-bb73-801edff98fac · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Reinforcement learning with world model Adam: A Method for Stochastic Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:24.945914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:24.945914Z digest=sha256:300606fea1196616213c88d2e56d2f8ddc8c5d808ac1f82698995c8db15342b8

Observation dcd3d62a-93b8-4da9-a63e-3559a09c515c · outbound

This paper cites Model-Ensemble Trust-Region Policy Optimization.

Reinforcement learning with world model Model-Ensemble Trust-Region Policy Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:24.950016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:24.950016Z digest=sha256:62d668f4472ecf5318058a7e9bca8fe40892bd5772602624689738e382c23949

Observation 5af25c26-5953-430a-96cc-44a20346927a · outbound

This paper cites Continuous control with deep reinforcement learning.

Reinforcement learning with world model Continuous control with deep reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:24.993438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:24.993438Z digest=sha256:74c884072d539a6deb0d3fca0c2bb81a7a97ee08b48d2d47c2dea3869496179b

Observation 30f91050-9507-46ec-8060-743c4eeca76e · outbound

This paper cites an unresolved cited work.

Reinforcement learning with world model Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:18:25.485324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T10:18:25.017884Z digest=sha256:886feaf23368af90e12e9d7b64df92764ed861258b58af19cba3eb305e4ba986

Observation 54828f07-5e14-408f-913b-84ea4fc2d222 · outbound

This paper cites Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees.

Reinforcement learning with world model Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:25.033668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:25.033668Z digest=sha256:59fccf31fccac5b261581f24e94a47dc33dc0a379699a7127142bf15d2f8624e

Observation 0e1a90cd-f5f9-417c-8d19-bb0964f8bb04 · outbound

This paper cites A., Veness, J., Bellemare, M.

Reinforcement learning with world model A., Veness, J., Bellemare, M

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:18:25.468770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T10:18:25.076410Z digest=sha256:7d50083a9d38c2b6f1122c12ee6c60fc6b82188f3a6ede4e1cc16dceb369f68a

Observation 14b9fd5d-a4e2-40bc-88f3-8ba3642272af · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement learning with world model Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:25.087028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:25.087028Z digest=sha256:75db073973997ec958cc5efa9090c04612d0bd77da8e9f9afd8f0669615e2de2

Observation 4396dc90-809d-464f-a702-f66f82188335 · outbound

This paper cites R., Yang, C., McGreavy, C., and Li, Z.

Reinforcement learning with world model R., Yang, C., McGreavy, C., and Li, Z

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:18:25.452871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T10:18:25.094998Z digest=sha256:2eb296e479539b7efa1a0ed68f766354dd53e09e56405aec2d923d30039af327

Observation 4c6054ca-8ea3-400c-a716-afae2bed813d · outbound

This paper cites an unresolved cited work.

Reinforcement learning with world model Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:25.098719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:18:25.098719Z digest=sha256:6280710cff38560b663e9a343fb5c8040c59422c53c11e838c0dc34696f8af24

Observation cc5025e0-db4b-4a08-9b1d-e20cacb75f12 · outbound

This paper cites an unresolved cited work.

Reinforcement learning with world model Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:18:25.322421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T10:18:25.102620Z digest=sha256:d7d463b0946778b68b74149421e7abbbf64498029e61abc27370bdf82cbc0262

Pith citing papers

No inbound Pith citation observations are available.