Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:1802.01561.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:54:49.069229Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T14:09:52.701191Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 01cb16bc-ee80-47b5-8119-b8eecee0f75e · inbound
Shaping Belief States with Generative Environment Models for RL IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8d9172cd-bbb9-4d08-acab-e7bd88f61f5e · inbound
Growing Action Spaces IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e6691caa-35d9-48f6-a089-98eb4e6a0373 · inbound
Learning Safe Unlabeled Multi-Robot Planning with Motion Constraints IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 93d6e0d7-947c-457a-a553-f5a87b26683d · inbound
Dream to Control: Learning Behaviors by Latent Imagination IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c641b33c-8e7f-405f-acf7-910c8374e3f7 · inbound
Dota 2 with Large Scale Deep Reinforcement Learning IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 01b67e85-c8f9-4bd1-bfe4-702ab7aaa0ce · inbound
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 295
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6620b507-68df-41b9-95ea-d3089bcdcfce · inbound
Mastering Diverse Domains through World Models IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 30932dcd-0afc-4cf7-b23d-1156aedd7cc6 · inbound
Bilinear Convolution Decomposition for Causal RL Interpretability IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e9b4180-c2af-4ab1-9420-7e91d9c3ec09 · inbound
Federated Learning of Dynamic Bayesian Network via Continuous Optimization from Time Series Data IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6467b71b-4615-43d1-9bfb-192d6363263d · inbound
Divergence-Augmented Policy Optimization IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d0c68a-fe9d-415c-a8ea-b0196c1bf0c6 · inbound
Generative AI for Autonomous Driving: A Review IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 257
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a90ddc6-33be-42a8-bfe5-3425ed2a759d · inbound
Rethinking Agent Design: From Top-Down Workflows to Bottom-Up Skill Evolution IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac34978-7ca8-41c5-85b1-5523b877bc77 · inbound
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d7b1a2-8640-4d15-aa2a-d20528885b22 · inbound
M3PO: Massively Multi-Task Model-Based Policy Optimization IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e162b0-6260-4029-9169-179e10bf90f1 · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c400599a-47be-4c03-b5fc-8696aad2b151 · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f0b3a030-5ac9-4eed-9226-7b1e879ead9e · inbound
Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 67e90f2d-fd7c-4ba9-8993-fba9c2b4e2b9 · inbound
Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 795112a6-2e44-45b6-a3e3-63ce5fd9bae1 · inbound
Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e784af15-4b08-43b3-987b-f5840a9852cf · inbound
Bridging Performance and Generalization in Reinforcement Learning for Agile Flight IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5f4ac08a-1039-44ae-9c64-b296258edbab · inbound
Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0374c552-206a-49c1-b20c-2af14c8ba37e · inbound
Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.