Pith. sign in

Paper Citation Record · LEDGER

IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:1802.01561.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1802.01561 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:54:49.069229Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T14:09:52.701191Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 01cb16bc-ee80-47b5-8119-b8eecee0f75e · inbound

Shaping Belief States with Generative Environment Models for RL cites this paper.

Shaping Belief States with Generative Environment Models for RL IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-25T18:57:06.758886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T18:56:19.459447Z digest=sha256:6c4ec1955049989eff938e8053b0591c73102914b0b72911c7f99e886daf35cb

Observation 8d9172cd-bbb9-4d08-acab-e7bd88f61f5e · inbound

Growing Action Spaces cites this paper.

Growing Action Spaces IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-25T13:25:52.358606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T13:24:59.682609Z digest=sha256:e3eeaf911460f5fee2652745dc8f1d994bf688d74de8af345db527d19f6d6c97

Observation e6691caa-35d9-48f6-a089-98eb4e6a0373 · inbound

Learning Safe Unlabeled Multi-Robot Planning with Motion Constraints cites this paper.

Learning Safe Unlabeled Multi-Robot Planning with Motion Constraints IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-24T23:15:03.078689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T23:14:06.926188Z digest=sha256:11bee9a12b7150b16ac85051bbaf0c9580b950bba70a52747d21e6f6a97faf03

Observation 93d6e0d7-947c-457a-a553-f5a87b26683d · inbound

Dream to Control: Learning Behaviors by Latent Imagination cites this paper.

Dream to Control: Learning Behaviors by Latent Imagination IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:16:36.481403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T01:16:36.399272Z digest=sha256:b4db683b3f73ea2f8bc8402cf7bc63ef141dcf3822ee43fc54b5f7d505c3ca59

Observation c641b33c-8e7f-405f-acf7-910c8374e3f7 · inbound

Dota 2 with Large Scale Deep Reinforcement Learning cites this paper.

Dota 2 with Large Scale Deep Reinforcement Learning IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T22:18:15.786172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T22:18:15.591361Z digest=sha256:c3f5444574dd4af2a9ec42445f29c37ffdfab47e0057b9a90d7f4885320dc0ee

Observation 01b67e85-c8f9-4bd1-bfe4-702ab7aaa0ce · inbound

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems cites this paper.

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 295

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:33:21.664988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T11:33:20.892688Z digest=sha256:7297e0e477e1fa065c89b6e1a359184e27fc8e5bcefa1a6202ecdc8aeb3110b4

Observation 6620b507-68df-41b9-95ea-d3089bcdcfce · inbound

Mastering Diverse Domains through World Models cites this paper.

Mastering Diverse Domains through World Models IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 66

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T09:08:21.921972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T09:08:21.677362Z digest=sha256:22d9be2ac48862c5c459c9393e74e42f4bfa31620044c5c4c03002b1fd7b5e7c

Observation 30932dcd-0afc-4cf7-b23d-1156aedd7cc6 · inbound

Bilinear Convolution Decomposition for Causal RL Interpretability cites this paper.

Bilinear Convolution Decomposition for Causal RL Interpretability IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:54:49.069229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:54:49.069229Z digest=sha256:5861c28f1f3ade0969a0624fdad64cb49f65f2def81527e679dc656ee8dcc474

Observation 8e9b4180-c2af-4ab1-9420-7e91d9c3ec09 · inbound

Federated Learning of Dynamic Bayesian Network via Continuous Optimization from Time Series Data cites this paper.

Federated Learning of Dynamic Bayesian Network via Continuous Optimization from Time Series Data IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:42.955351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:42.955351Z digest=sha256:17f7e20f586d974d42fd9e07fee7e2476ec30e21e8effb3ff7d037b1b0d8f3e0

Observation 6467b71b-4615-43d1-9bfb-192d6363263d · inbound

Divergence-Augmented Policy Optimization cites this paper.

Divergence-Augmented Policy Optimization IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.896067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.896067Z digest=sha256:521efef1ea06548111bdc0f783774daaadb74e78f13b663247c8af384d4d4953

Observation 03d0c68a-fe9d-415c-a8ea-b0196c1bf0c6 · inbound

Generative AI for Autonomous Driving: A Review cites this paper.

Generative AI for Autonomous Driving: A Review IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 257

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:07.347431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:07.347431Z digest=sha256:400b9086a89367d2f7440a32a3c1cf76d60a03fda3f259e3f6cacf6878cb515c

Observation 0a90ddc6-33be-42a8-bfe5-3425ed2a759d · inbound

Rethinking Agent Design: From Top-Down Workflows to Bottom-Up Skill Evolution cites this paper.

Rethinking Agent Design: From Top-Down Workflows to Bottom-Up Skill Evolution IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:31.462520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:31.462520Z digest=sha256:f3ea11358755fcc61d4c1bcbbcb2f6606ff37f4cd2b5ff28bfbffad3a05e57dd

Observation fac34978-7ca8-41c5-85b1-5523b877bc77 · inbound

Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN cites this paper.

Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:27.819005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:27.819005Z digest=sha256:658bf27c19e8ff1e66fcf972c0904109b3969967b3163f5de2ad0b81078c240f

Observation 24d7b1a2-8640-4d15-aa2a-d20528885b22 · inbound

M3PO: Massively Multi-Task Model-Based Policy Optimization cites this paper.

M3PO: Massively Multi-Task Model-Based Policy Optimization IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:50.021754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:50.021754Z digest=sha256:f3813433de43e583db7b40a3f8fdaeb84b82e13bd032abfe364b1e41d78444a2

Observation a8e162b0-6260-4029-9169-179e10bf90f1 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:42.435634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:c29d16e456c2b3cfb89c234d0e775b4ffc7929e8b31f6015e383d9cdc586c7c9

Observation c400599a-47be-4c03-b5fc-8696aad2b151 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.721995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:701d5560fe9cbc869eabee5c29ba39eedb055bc4a7a7eb3d162261741dc12aac

Observation f0b3a030-5ac9-4eed-9226-7b1e879ead9e · inbound

Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL cites this paper.

Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:05.312851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T16:58:39.520620Z digest=sha256:3d3783d28af449b2c15bcf31cf7ef54a8fc1c83b10f168aa06f260c8bae913b5

Observation 67e90f2d-fd7c-4ba9-8993-fba9c2b4e2b9 · inbound

Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL cites this paper.

Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.134398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T23:41:30.860309Z digest=sha256:a13340f0d449db1be26b13039bde4eec9d4053771c8ad8d55d6655b32e3bc0f1

Observation 795112a6-2e44-45b6-a3e3-63ce5fd9bae1 · inbound

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models cites this paper.

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:07:26.508407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T19:10:53.882876Z digest=sha256:4aa60edc6331834e471a4556aa44d0a999018be4f21c179698adb02bd437dc39

Observation e784af15-4b08-43b3-987b-f5840a9852cf · inbound

Bridging Performance and Generalization in Reinforcement Learning for Agile Flight cites this paper.

Bridging Performance and Generalization in Reinforcement Learning for Agile Flight IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-04T14:09:52.702576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T04:33:20.981811Z digest=sha256:25a6ff64bd2ed2a0315af319748115f13c175674520dc41a5c1c8d6ea0465017

Observation 5f4ac08a-1039-44ae-9c64-b296258edbab · inbound

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF cites this paper.

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T18:55:58.783845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T01:23:07.441218Z digest=sha256:d1d01d6b0f50c372b45dc8ebe404fce522a0ab5007c0bf87420ac27fcc165f39

Observation 0374c552-206a-49c1-b20c-2af14c8ba37e · inbound

Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning cites this paper.

Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T07:56:35.435467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:56:35.435467Z digest=sha256:9738a54ccb38c4d4c28c1bd0dabf8683d15e8021daa1e8cbd5a1c2aa999da55c