Pith. sign in

Paper Citation Record · LEDGER

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference

As of 12 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 4 inbound Pith citation observations for arXiv:2412.14355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14355 v1

Coverage vector

measured 100 of 112 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:27:15.062636Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:07:41.420435Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:29:51.963553Z

Reference resolution

100 of 112 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved67
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42676e21-90fe-446a-b0fe-7caeef5aa0e3 · outbound

This paper cites Context-Specific Representation Abstraction for Deep Option Learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Context-Specific Representation Abstraction for Deep Option Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.639773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.639773Z digest=sha256:8f34f12d44657b2550f9dd251ca5705500255f588bffc9ab332568caaf49b8a9

Observation 2d40583f-5efd-4973-89c4-2797ff8d7f54 · outbound

This paper cites Blind decision making: Reinforcement learning with delayed observations.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Blind decision making: Reinforcement learning with delayed observations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.646038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.646038Z digest=sha256:f87933eeb1b48a36ce3f2d801163c6169a9476d0c5e70362fd2cf796643d4cb1

Observation c1e272b4-c239-47b7-a543-dac16b596e68 · outbound

This paper cites Efficient Black-Box Planning Using Macro-Actions with Focused Effects.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Efficient Black-Box Planning Using Macro-Actions with Focused Effects

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.650868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.650868Z digest=sha256:49b36ed1343eb1efe29e8b6e9de47f58e503c93db92bf2f47e843416df80207e

Observation 71a818d4-11a9-4253-90f1-49fdedb815e8 · outbound

This paper cites The option-critic architecture.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference The option-critic architecture

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.656009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.656009Z digest=sha256:e225062987e55b29fb866404064a5f738faf7b13cd608344fa6461830689ee36

Observation cb1d5fbe-ffbb-4c57-8db9-278b9434b4f8 · outbound

This paper cites Reinforcement learning and its relationship to supervised learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Reinforcement learning and its relationship to supervised learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.660540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.660540Z digest=sha256:e71c0dc003a99c726b2241917b2e639927b854fb86cf97b8ca65abc3976526e4

Observation 2f49b150-b874-4263-94a8-1a251ed295f8 · outbound

This paper cites Continual Learning with Self-Organizing Maps.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Continual Learning with Self-Organizing Maps

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.665152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.665152Z digest=sha256:21870d83887744145d965dbba7094ad941b54bdde93d22e4e24fdfb55b15e8df

Observation 44fa17d6-3c07-4c5b-a62d-c1df64154080 · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference The arcade learning environment: An evaluation platform for general agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.669925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.669925Z digest=sha256:00550fe875972476b0f01fff6dc6f44f31ae92f099492613156c8a267cbd10ee

Observation 194ed486-1d2b-4673-856e-269f0e3f2f01 · outbound

This paper cites Dynamic programming and optimal control.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Dynamic programming and optimal control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.673979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.673979Z digest=sha256:35fe1750112efc3a133c1bffa96517a34ffd5892dde4e7af757881a3b9d4b321

Observation 169d906c-f333-4347-81e7-2436fe8d4bd7 · outbound

This paper cites Reinforcement learning with random delays.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Reinforcement learning with random delays

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.677966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.677966Z digest=sha256:b866cea1998195663c07fbf16c9a6b84de7ab7e56af24a44f6460fd575b97323

Observation 9ca0c860-d6c6-4694-91d8-ab7882b8eee4 · outbound

This paper cites OpenAI Gym.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference OpenAI Gym

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.682133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.682133Z digest=sha256:31088f0628b2c6a9bdc547f91697139d19042573ad6adee80b44a905371bbb99

Observation 94faf705-8ed9-472e-8412-03dfeaa07b3e · outbound

This paper cites Recursive routing networks: Learning to compose modules for language understanding.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Recursive routing networks: Learning to compose modules for language understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.686543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.686543Z digest=sha256:4098e20cd5240008e668c4d4367ae9242dae0d93671ce04545425bdd3b2ab03e

Observation 7fa4cb4a-60de-4f4e-aa8f-93d6caa884e9 · outbound

This paper cites Automatically composing representation transformations as a means for generalization.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Automatically composing representation transformations as a means for generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.690534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.690534Z digest=sha256:1675499ad107ca7241f62d79366b3dd59545cd627c652f9344b8dd0e665f28d9

Observation e13ebc03-b3e9-4807-9a19-daf62942e8f0 · outbound

This paper cites Leveraging procedural generation to benchmark reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Leveraging procedural generation to benchmark reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.694381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.694381Z digest=sha256:473b36022c99cdf2813bb043293ae82fae2547619f8f528f44a8d66267f805ad

Observation 4d3f0895-871a-4fdf-b5bc-65912eb48c3f · outbound

This paper cites Learning how to Interact with a Complex Interface using Hierarchical Reinforcement Learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Learning how to Interact with a Complex Interface using Hierarchical Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.698765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.698765Z digest=sha256:9d8a1ae0c680696dfe7cceb10f8f4fee980ab5ec08ffb9f4aa047902ceb67a4a

Observation c5bb3772-46e2-4d74-a958-0bbc51ab575e · outbound

This paper cites PepCVAE: Semi-Supervised Targeted Design of Antimicrobial Peptide Sequences.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference PepCVAE: Semi-Supervised Targeted Design of Antimicrobial Peptide Sequences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.703211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.703211Z digest=sha256:b4fb5655a4f4101c44cd9493fe262e4559cb913085f638caeaadc43bfb36f800

Observation a5ab8989-f0ae-474b-b522-f621be23ba0f · outbound

This paper cites Acting in delayed environments with non- stationary markov policies.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Acting in delayed environments with non- stationary markov policies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.707857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.707857Z digest=sha256:8ed521067e52493a307735524a3c0d749d38cc46a41a0f845a7ece28fbb38b4c

Observation 13629a57-e29b-42a8-82e6-c075bb028faa · outbound

This paper cites Acting in Delayed Environments with Non-Stationary Markov Policies.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Acting in Delayed Environments with Non-Stationary Markov Policies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.712015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.712015Z digest=sha256:fed6fde8c0bf00f64371df0778504c63de3a3639a7b7a8027358c7a080c006f8

Observation 9a191e21-725c-4ed0-84e0-d800acdd5bb8 · outbound

This paper cites Contextual Moral Value Alignment Through Context-Based Aggregation.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Contextual Moral Value Alignment Through Context-Based Aggregation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.716776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.716776Z digest=sha256:a5389129374379d7afd7962c55c249e7138f2b88d9137e85915d09e630faba11

Observation 349b4bff-0547-47b2-b915-dff83405b97e · outbound

This paper cites Challenges of Real-World Reinforcement Learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Challenges of Real-World Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.721266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.721266Z digest=sha256:7c54e2f26741bbad052b1724894804fcb94bdead355f0d9f3980dee630899c17

Observation de3df47f-eebc-44f7-91e5-a27c6e7581bf · outbound

This paper cites Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.725578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.725578Z digest=sha256:51e866dd3b44cbc8d1dc76c07a82c6da4337308432ea82a44a8c63a19f29427a

Observation 5964f396-5f02-4494-bdd0-58bd445fade2 · outbound

This paper cites Reducing the cost of cycle-time tuning for real-world policy optimization.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Reducing the cost of cycle-time tuning for real-world policy optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.729690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.729690Z digest=sha256:1986cca8e991674991183928b667c87977da17cb92782a841c761b126476f8c4

Observation bd75baed-9b4d-4d18-bd29-8cdae67d2fd8 · outbound

This paper cites Learning with opponent-learning awareness.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Learning with opponent-learning awareness

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.733660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.733660Z digest=sha256:cb1b71dfaac4eb3ebe6a621d2247fba5ef5676a80d3716acfb98b88e1ae75acd

Observation ae080094-a2fb-47da-9889-ea18ad9ee9f4 · outbound

This paper cites Dice: The infinitely differentiable monte carlo estimator.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Dice: The infinitely differentiable monte carlo estimator

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.737671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.737671Z digest=sha256:dd4ef5d7ee4bc5a30ecccf1f0caf00255824d36cccbd7f7ab8a5f657773c6004

Observation 44b7e040-700e-49f0-b4d5-552a23d47862 · outbound

This paper cites Harris, K.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Harris, K

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.741631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.741631Z digest=sha256:dc47938995dd618b3cfc83213b756e41301a90646792e7b85cb21287bd448440

Observation 32d673fd-24fd-4d52-a0d1-eefa674a9b93 · outbound

This paper cites Alexandria: Extensible framework for rapid exploration of social media.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Alexandria: Extensible framework for rapid exploration of social media

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.745545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.745545Z digest=sha256:bb65d89da876aa6c5adf7762385feba9c75703af7cbf86076e3fede614deaa17

Observation 31f3d686-b03c-4046-b10b-e6d213177db5 · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Rainbow: Combining improvements in deep reinforcement learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.750142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.750142Z digest=sha256:4f902e8433d74509ff4404c5846f0358a85889b50c432bed348a37d5f195ca27

Observation e05a2e96-ad17-4189-9bbb-1a27e0d81f16 · outbound

This paper cites Texplore: real-time sample-efficient reinforcement learning for robots.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Texplore: real-time sample-efficient reinforcement learning for robots

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.754482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.754482Z digest=sha256:eb8407068cdf051cc85aef243eea41867b67ee07d70037d57268864398fa994d

Observation df57cb1f-6ff2-4986-8d33-09c9f3979804 · outbound

This paper cites A Real-Time Model-Based Reinforcement Learning Architecture for Robot Control.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference A Real-Time Model-Based Reinforcement Learning Architecture for Robot Control

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.758876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.758876Z digest=sha256:7ea87c8af25e9b1c5460a6141df987c8764975860411bfdf86092334520eb005

Observation ddfbd057-651a-47f1-8fd8-5b95952d4372 · outbound

This paper cites Rtmba: A real-time model-based reinforce- ment learning architecture for robot control.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Rtmba: A real-time model-based reinforce- ment learning architecture for robot control

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.763424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.763424Z digest=sha256:83d8894b763b04a230fc591b1e0bfd71ca0bb047d83dc8f9e149f2ef0391168b

Observation 7442c6fd-60dc-48e6-bb1b-14d483ec7761 · outbound

This paper cites Why can a machine beat mario but not pokemon?, 2018.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Why can a machine beat mario but not pokemon?, 2018

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.767749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.767749Z digest=sha256:60540c5ea5f001557e2bb25248625ecd2e077640b8aec0b85637527c5f5d7dfd

Observation 5c27494a-714e-4e5c-9bf3-44bb2e3c7c28 · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Near-optimal regret bounds for reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.771958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.771958Z digest=sha256:bd6c8a772fd27f3740b3aa7e74423f7a036ba8204b9fc1d1e2618a0aaf676769

Observation 26401b76-537f-4e1e-918b-704f6422f873 · outbound

This paper cites Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.776263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.776263Z digest=sha256:836ea07bec272f4e2504b4c38b801623d85ea6e65b6f5f1db96ba176a593b72a

Observation 77b41dbe-7544-4752-ba91-7b270144b7ad · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.780488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.780488Z digest=sha256:6e4c9ef94ce09dee95866517d64d1d77be9536a93fa20964948e4a4f1e2133b3

Observation 3ce58390-5d87-474e-be42-7def0c668028 · outbound

This paper cites Scaling Laws for Neural Language Models.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Scaling Laws for Neural Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.784837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.784837Z digest=sha256:1915c06087f74af59e63279d896e3b8ee637b3d6aec86650531aa1ffb99ba1ef

Observation dbe37a15-1d70-4757-a357-4888eb7480a7 · outbound

This paper cites Reinforcement Learning from Delayed Observations via World Models.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Reinforcement Learning from Delayed Observations via World Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.789255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.789255Z digest=sha256:bd91368470e61ce2568a1f9755d8e67361028d210a1217cd1e898b5a6f16f168

Observation 83d15995-a8ea-4f7e-862d-cb6cada52e74 · outbound

This paper cites Dynamic decision frequency with continuous options.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Dynamic decision frequency with continuous options

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.793283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.793283Z digest=sha256:506ea8e2340819e0e5e3a57784f531f6db7038ed1f8b49f05762cb0ce93da2e7

Observation b108ebf0-0d82-4fef-84d9-30830ab0c2f4 · outbound

This paper cites Markov decision processes with delays and asynchronous cost collection.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Markov decision processes with delays and asynchronous cost collection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.797034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.797034Z digest=sha256:3bbd2507134d7095d4a5ba655b59c00b76faec6116374ec59cbb6472a8989555

Observation 087bd5a7-b645-4b38-9430-cf987ff8456f · outbound

This paper cites Domain scoping for subject matter experts.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Domain scoping for subject matter experts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.800965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.800965Z digest=sha256:96038d0a25dd295df7e41c6746264287a92dc6153c6a7366ac97b7cac587d9d1

Observation 75a007d3-514e-428a-be07-d87fc97bcc6f · outbound

This paper cites Towards Continual Reinforcement Learning: A Review and Perspectives.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Towards Continual Reinforcement Learning: A Review and Perspectives

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.804991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.804991Z digest=sha256:4551ea029140f034e4f0cd1481cb89b4e432bf2a8b9a79359713b425d8599b80

Observation f89c5bb1-9424-4a4e-9165-ed4f06ed33b1 · outbound

This paper cites Learning Hierarchical Teaching Policies for Cooperative Agents.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Learning Hierarchical Teaching Policies for Cooperative Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.809102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.809102Z digest=sha256:b31863e84d1696c4609bfd334c1fd5375fbba9b94a9d093ac430d8954fbc15e5

Observation 7e908fff-b058-45d9-8091-f9e192ae7dd0 · outbound

This paper cites Hetero- geneous knowledge transfer via hierarchical teaching in cooperative multiagent reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Hetero- geneous knowledge transfer via hierarchical teaching in cooperative multiagent reinforcement learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.813510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.813510Z digest=sha256:3a3258acebc579b5e24165e090a7ee8fe1bd63b4f4c7f8cc27abbf88468864f5

Observation fb8f1726-d4f6-4b65-be0a-8faa5e205f36 · outbound

This paper cites A policy gradient algorithm for learning to learn in multiagent reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference A policy gradient algorithm for learning to learn in multiagent reinforcement learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.817747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.817747Z digest=sha256:1a9505f9f7e291fc2296e8e054230eebd4fd56291c6c6b656b38993915c9e28a

Observation c91f14d2-e8b4-45e4-beee-7d5e60eefea8 · outbound

This paper cites Influencing long-term behavior in multiagent reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Influencing long-term behavior in multiagent reinforcement learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.821739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.821739Z digest=sha256:18c0bfe0642c22f792d7043ea03dba5e7f17ec636ba5d5a7dcd717edab7a13bc

Observation f8edd0f0-d037-4b0a-ba5f-578cc6f92c37 · outbound

This paper cites Game-Theoretical Perspectives on Active Equilibria: A Preferred Solution Concept over Nash Equilibria.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Game-Theoretical Perspectives on Active Equilibria: A Preferred Solution Concept over Nash Equilibria

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.825834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.825834Z digest=sha256:dc27d7f622238306cdc05030f5fb9bf4a3ac02c618baa27341a665cbb23962cb

Observation f3f95e65-dbf7-4896-9c0b-d4c1037673de · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Overcoming catastrophic forgetting in neural networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.830395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.830395Z digest=sha256:29ed61aa25c40144a1f7c811bb3505c50911bc2a787065c18904448173369db2

Observation b8724588-9917-4033-8c29-c6bb131b118b · outbound

This paper cites A Study of Compositional Generalization in Neural Models.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference A Study of Compositional Generalization in Neural Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.834450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.834450Z digest=sha256:4f35402ea2a324327c6b1678b07235297fe6afa66f66853f1a3cb196774783f1

Observation 7294d8a5-f384-4ed1-9829-6b7ede3f8c5e · outbound

This paper cites Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized Teaming.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized Teaming

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:27:15.374748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.838998Z digest=sha256:1a7bc430e668d367d7a716349aaec127972492b1137488a1d40e357eb0de8760

Observation 9a29402c-a35f-4eff-bb1a-96ad8a4442db · outbound

This paper cites Asynchronous coagent networks.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Asynchronous coagent networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.843219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.843219Z digest=sha256:fcba9284de164267d8c934723370c3b9bc067252c69ffedb56ac0959f94320cc

Observation eccb639e-cd4c-40f5-8cdb-49714ff3c24c · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Conservative q-learning for offline reinforcement learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.847013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.847013Z digest=sha256:05ff00717a618c3c9a28a36f5bc1e03cd59e21efb18ef847b1386465ac2738d0

Observation 83da1e7c-a131-4ac2-bb3c-432c4a5d0382 · outbound

This paper cites Slow learners are fast.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Slow learners are fast

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.280996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.851260Z digest=sha256:634e048213ddae7662b31b6da6c6c26832da3d6217e5f774c56cdf185f3569b1

Observation e0c1c3ec-0f47-4c65-8a31-424bfcf47cc8 · outbound

This paper cites Emergent Multi-Agent Communication in the Deep Learning Era.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Emergent Multi-Agent Communication in the Deep Learning Era

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.855190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.855190Z digest=sha256:d8e968ef96b812f77927a1505d8eb721e1151e07b3b5a168a9ccc5b916a91c75

Observation fe7b2f65-c7c4-49fa-82f0-ccd26a83eed5 · outbound

This paper cites Towards a unified theory of state abstraction for mdps.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Towards a unified theory of state abstraction for mdps

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.267152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.859334Z digest=sha256:104c30ca2009f8b7a2d17c68b8cd83919d243b69d4f5a25e189b5c25bff904f4

Observation 9cfd5235-f7a9-4961-8853-88dbcd0b95be · outbound

This paper cites Learning without forgetting.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Learning without forgetting

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.863327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.863327Z digest=sha256:7f37d834f217d55094377e73e5eb8b446843c0f490da671fa71a71de8573e6ea

Observation 1582e272-c6d7-4a8f-abc4-230e35e0d5be · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Multi-agent actor-critic for mixed cooperative-competitive environments

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.867203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.867203Z digest=sha256:55dc4fe512d0ddf7e87be7ce897c4cca351522843e16454267fca12e6f4da116

Observation 83d5a837-bfa6-42b4-afb4-64e1c7438229 · outbound

This paper cites Consolidation via Policy Information Regularization in Deep RL for Multi-Agent Games.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Consolidation via Policy Information Regularization in Deep RL for Multi-Agent Games

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.871555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.871555Z digest=sha256:c30313506a10aa2ee58a5f2a2385ea62f010ca969f37a4ec456b6f6759c1d994

Observation 94d8b9ad-6844-4c6f-9212-bff1a9880467 · outbound

This paper cites Deep RL With Information Constrained Policies: Generalization in Continuous Control.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Deep RL With Information Constrained Policies: Generalization in Continuous Control

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.876277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.876277Z digest=sha256:35a31e25f91fa4084b52ed6e03c635427c3a942488b47e1005d9a641b7a2705f

Observation be58ace0-fd77-46cc-b5f4-d596d4c7c5f5 · outbound

This paper cites Rl generalization in a theory of mind game through a sleep metaphor (student abstract).

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Rl generalization in a theory of mind game through a sleep metaphor (student abstract)

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.234368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.880741Z digest=sha256:25a19264c13c0d9dfdcf5e701b06ac94bb330f5f0ee0f03e00971fe394634fca

Observation dd9c3ceb-76a4-4441-8896-0a1dd9f94a43 · outbound

This paper cites Capacity-limited decentralized actor-critic for multi-agent games.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Capacity-limited decentralized actor-critic for multi-agent games

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.218789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.885308Z digest=sha256:7fd045be9568dcaca4a1afb44464fb00f7921839467f404edc1f4dc59688aa41

Observation ba88cb19-9bf8-4ee1-bb5c-4eb1ccd6af90 · outbound

This paper cites Learning in Factored Domains with Information-Constrained Visual Representations.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Learning in Factored Domains with Information-Constrained Visual Representations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.889666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.889666Z digest=sha256:f9b19b449593f976a5e1c12033ac2a314101cbc70b9fee712d2b84d561be4e59

Observation a854065a-249f-4f60-9250-b92a158ff0a5 · outbound

This paper cites Summarizing societies: Agent abstraction in multi-agent reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Summarizing societies: Agent abstraction in multi-agent reinforcement learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.203318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.894379Z digest=sha256:29da4dfb4e25533f5b10e945e10db54a7d1204d394dd071cff9fbc4150d41568

Observation 6003dc43-48a3-468a-a42c-c6d92c295c64 · outbound

This paper cites Human-level control through deep reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Human-level control through deep reinforcement learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.898919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.898919Z digest=sha256:e067f2072c3e062dc36736a5d88d771b5abe29ee472fdb63788420266d134080

Observation 92b42fe9-2d7c-4243-a275-2cb75521561a · outbound

This paper cites Asynchronous methods for deep rein- forcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Asynchronous methods for deep rein- forcement learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.179908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.903285Z digest=sha256:fe992680aa0f67512dbf4f3a4eec5a56f960aa88121fcc18a2d3747443145203

Observation 2142e7b2-56ac-4c58-b245-0f0f934c1885 · outbound

This paper cites Revisiting state augmentation methods for reinforcement learning with stochastic delays.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Revisiting state augmentation methods for reinforcement learning with stochastic delays

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.166795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.911492Z digest=sha256:7771c98ef62029022114ce1cf8f7e9a252ac99d8281be66bebabec1a157ac9dc

Observation e40cf72f-80c2-46b5-9dc1-fa4dfa0c973f · outbound

This paper cites Gotta Learn Fast: A New Benchmark for Generalization in RL.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Gotta Learn Fast: A New Benchmark for Generalization in RL

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.915508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.915508Z digest=sha256:2fcc6463394927a909aa13cdd9c56dd9d7147d5f625dae8fd7204ec86b561c17

Observation 964a3708-080b-4a5f-a9a7-fc1e0e337a40 · outbound

This paper cites Sequoia: A Software Framework to Unify Continual Learning Research.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Sequoia: A Software Framework to Unify Continual Learning Research

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.920225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.920225Z digest=sha256:bc47ddaf1b5bed1cfe1a1e4c76f6cc11b3dae3479bbac7ffba64a00168e50180

Observation 365efc24-4f5c-449d-a27c-1caf5bfdf147 · outbound

This paper cites Learning to teach in cooperative multiagent reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Learning to teach in cooperative multiagent reinforcement learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.153483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.924602Z digest=sha256:951ecdc2c42ed1598a975ed13fecd8ef3b46a84b3e7568223a20f0629b5d6fd7

Observation e3e178e2-fdde-4391-b71d-ec5b3c533ade · outbound

This paper cites Regret bounds for reinforcement learning via markov chain concentration.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Regret bounds for reinforcement learning via markov chain concentration

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.139550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.928720Z digest=sha256:1b31621ae45afb466ab6e4cd693208736ca17c7783a2b75880949ffb14a4f030

Observation bc55efdd-e18d-40fb-939d-a23595e4aaa8 · outbound

This paper cites Comvas: Contextual moral values alignment system.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Comvas: Contextual moral values alignment system

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.125714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.932722Z digest=sha256:e88366d4e915d349cac9b61f21e0ad32bf9f8ccc24a08501d227d93db70e0d60

Observation 59619b19-2fa8-47ec-a564-554197c1acd5 · outbound

This paper cites Pytorch: An imperative style, high- performance deep learning library.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Pytorch: An imperative style, high- performance deep learning library

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.112292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.936808Z digest=sha256:be8a3c530b5ebada7899c396b2270d700731b9f37df394621620c8344213d81f

Observation 485f0c2d-b6fc-4abc-8224-c807355768e4 · outbound

This paper cites Markov decision processes.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Markov decision processes

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.098346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.940874Z digest=sha256:3cf88d8eb16b24ac06baa28d2024dbe6c696249964f0e24fe51d5a30aff8528b

Observation ef4834a4-75a3-47de-baaa-43a96a653fe1 · outbound

This paper cites Machine theory of mind.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Machine theory of mind

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.085394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.945027Z digest=sha256:25d7544a99bf9a7e009170e815beb51c54ea24f0e484785b430aa7eba1a60319

Observation 00ffd810-928e-42fc-9160-d50b7f54943a · outbound

This paper cites Real-time reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Real-time reinforcement learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.071913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.948899Z digest=sha256:8156d615ff28243c837cbd6b828f4301d4dcc65cdcfaf1e1cbb40092b9a62b80

Observation 1e513550-13ed-4c87-b0c2-2f7096b92e2f · outbound

This paper cites Distributed computing in social media analytics.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Distributed computing in social media analytics

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.058272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.952925Z digest=sha256:24ad3f68a69166427421996374f28552b2203709ca7c2b71dec4c869289a042a

Observation 3f30e355-d817-4ef3-84d7-3090c89a9d9c · outbound

This paper cites A deep learning and knowledge transfer based architecture for social media user characteristic determination.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference A deep learning and knowledge transfer based architecture for social media user characteristic determination

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.044082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.956822Z digest=sha256:e25d67fe256e355df4a96d5f32ec0e97ee041a350e81620dc6c32bd6e4350286

Observation 5b830176-b8e6-4551-ad81-f577447ab746 · outbound

This paper cites Correcting forecasts with multifactor neural attention.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Correcting forecasts with multifactor neural attention

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.029248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.960748Z digest=sha256:e43c1e02e4afd6f07e78e76584bd049850cb8554e7039d76b36e922c5d409da8

Observation 61a8096b-df49-4e8e-b248-36f0d219153e · outbound

This paper cites Generative knowledge distillation for general purpose function compression.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Generative knowledge distillation for general purpose function compression

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.015456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.964907Z digest=sha256:72f23d5d71592f4c4960016ed58deee0e0f5730820922320301d5bed6d3f6a84

Observation a8f8927e-63b9-4476-9df8-0f23708c9c8e · outbound

This paper cites Representation Stability as a Regularizer for Improved Text Analytics Transfer Learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Representation Stability as a Regularizer for Improved Text Analytics Transfer Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.968725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.968725Z digest=sha256:5f283cdfb6297b3eda441c8db4aa1201dbf14ac316264ac3924296e3c11b71a7

Observation e1a0638b-e413-4d87-8941-44a12c08e55b · outbound

This paper cites Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:14.973031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:14.973031Z digest=sha256:21bab0b3eaa2f7ee1b6a37b3e67da5db7df9db46c7333bae4b830628de1b5113

Observation 95d8d66a-5b04-40f3-8c63-3f20f3521bda · outbound

This paper cites Learning abstract options.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Learning abstract options

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:16.001919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.977621Z digest=sha256:9698b19902d5e7f932c340de3706cbdae9fb797eb95cdf60b02abfdcc3a4753a

Observation 4c844ca2-6768-4694-a8a8-e9fc22383800 · outbound

This paper cites Scalable recollections for continual lifelong learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Scalable recollections for continual lifelong learning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.987843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.981462Z digest=sha256:e8195cfd8b9ce7cbeae960aebf4c3241cd8e869a3977e66f31c541a9f2505078

Observation 3d37eb38-8477-48fb-a9ba-f9312bedacf3 · outbound

This paper cites On the role of weight sharing during deep option learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference On the role of weight sharing during deep option learning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.973742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.985584Z digest=sha256:7a2ebd67f49c222c8c4f79a08be8d3c838cc3f11721e17a6ddad51831337e27a

Observation 0738d394-5aa3-4e8e-9362-4cfc7fa0009c · outbound

This paper cites Continual learning in environments with polynomial mixing times.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Continual learning in environments with polynomial mixing times

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.954100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.989450Z digest=sha256:e5f1f0141b767a66e89f07fff280d32e39456ff40e9f2be3414ec36b87a1c1e2

Observation 35cb7252-1ffa-4c2b-85ff-a4fa89b20e2c · outbound

This paper cites Balancing context length and mixing times for reinforcement learning at scale.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Balancing context length and mixing times for reinforcement learning at scale

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.940664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.993355Z digest=sha256:5d273c3e0fe25c2226fe9ca9cca97e08d239dc8187e45dd28686bbb67208b08a

Observation 45d37930-939e-46b0-8809-2d22b0d76147 · outbound

This paper cites Realtime reinforcement learning: Towards rapid asynchronous deployment of large models.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Realtime reinforcement learning: Towards rapid asynchronous deployment of large models

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.926945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:14.997330Z digest=sha256:81c83535da432977d21b4d2e96a310477597d4857fe2f4d1762a65ad802292b2

Observation 14ae1d6f-3604-4a1e-8a74-3eff9f09b904 · outbound

This paper cites Routing networks: Adaptive selection of non-linear functions for multi-task learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Routing networks: Adaptive selection of non-linear functions for multi-task learning

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.912708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:15.001208Z digest=sha256:0b5f7c20230cd35c4418129d59128c326afae02cc84fbbca436ee2bf4cb24253

Observation 8c2b864e-8597-455f-a980-70684bf74bfb · outbound

This paper cites Dispatched routing networks.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Dispatched routing networks

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.897118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:15.005734Z digest=sha256:e8773d20335b62a301c1700f6bd05baab564c6a8b44c21af5b04b06dee08dbd7

Observation f9594823-3c5d-4c40-9b8e-d448449301d9 · outbound

This paper cites Routing Networks and the Challenges of Modular and Compositional Computation.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Routing Networks and the Challenges of Modular and Compositional Computation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:15.010184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:15.010184Z digest=sha256:54fe45b2c594a002dff421732ee50d315e1abb676024340422a02af4c47f5f50

Observation 8f4188cf-bed3-4482-a278-a647c36e1b64 · outbound

This paper cites Control delay in reinforcement learning for real-time dynamic systems: A memoryless approach.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Control delay in reinforcement learning for real-time dynamic systems: A memoryless approach

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.882923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:15.014901Z digest=sha256:dc5339067b3dcb3576f349b887f2fbc1b0ad12971370fd3b466a12d8e5a9db82

Observation 3f84a9fa-23e7-4b1a-bf67-d66d1b12ef4f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Proximal Policy Optimization Algorithms

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:15.018952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:15.018952Z digest=sha256:4fc7f6c24e1aba93f2cc381746806e0c46a791e95e8cedcdb43d0e6aa2ceb9a0

Observation 55e8fdd2-de7d-425c-b16e-e8b46d8855e4 · outbound

This paper cites Bigger, better, faster: Human-level atari with human-level efficiency.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Bigger, better, faster: Human-level atari with human-level efficiency

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.868449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:15.023116Z digest=sha256:d048953e2a56c0e05ac32f1fe399202d9b3a3e8a23404677a7dd8da438a5d283

Observation b7ca989e-4e7b-4154-9904-c9d6755c90e7 · outbound

This paper cites Pokémon red/blue/rng manipulation faq, 2020.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Pokémon red/blue/rng manipulation faq, 2020

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.855244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:15.027245Z digest=sha256:00847c4ef83649b41f9a9d15e179af31872b8290ea1f42892f22ab9542e0ae5d

Observation 5604b39a-e6e6-4edd-ae16-9714f802bc66 · outbound

This paper cites Pokemon red/blue leaderboard, 2024.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Pokemon red/blue leaderboard, 2024

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.841314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:15.031036Z digest=sha256:1ff6918f6caffa763755f2604ea5c3fb61fc6704a92eef8bbac6ab846eaca89d

Observation 34eefcfc-d87e-45bb-b237-007c867ccbbd · outbound

This paper cites Reinforcement learning: An introduction.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Reinforcement learning: An introduction

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:15.034810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:15.034810Z digest=sha256:93e7e22f116f0e9926cb8f74ef6ee23a362835e149e695ccbf63a9b9d3a86e62

Observation 75eb3a9b-b0f5-4a05-996a-bf558d451f03 · outbound

This paper cites Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:15.038803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:15.038803Z digest=sha256:8af6deff1a243f16127ee5401aaa05bb43aebdd77527b0d10672c4c38373cdc7

Observation 5e576380-2621-4e92-8b46-1f5e3252e11b · outbound

This paper cites Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.808522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:15.042627Z digest=sha256:63a831aa6a555422bd684ab6c204c9d9d3ead843203d2862f5177353023dd5bd

Observation 5647d77d-8fcb-4be3-9ed2-b50b3c8d46ed · outbound

This paper cites A Deep Dive into the Trade-Offs of Parameter-Efficient Preference Alignment Techniques.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference A Deep Dive into the Trade-Offs of Parameter-Efficient Preference Alignment Techniques

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:15.046470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:15.046470Z digest=sha256:9a1bcede3170cb1d497c25b7d16e954d777c96ffd80a379fd868a8b56cd41255

Observation 0bc06504-d1ea-4892-811f-4bd2bfe5d1cf · outbound

This paper cites Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:15.050862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:15.050862Z digest=sha256:09808e13fad4031305426fc387bb0c69a0e3aff88ac19cd031dc1dba7e61a525

Observation d5b35cfe-96bd-4be9-b3c7-92771cc55c1a · outbound

This paper cites an unresolved cited work.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:27:15.794586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:15.054937Z digest=sha256:d1597025d6be24677574472f6dce65e032fedebeaf023245e6cfeb0ddfb9ddc4

Observation 3db06a8e-4974-486a-be49-7b6f79c7b686 · outbound

This paper cites Scalable approaches for a theory of many minds.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Scalable approaches for a theory of many minds

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:27:15.781218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:27:15.058713Z digest=sha256:e87912f073bf117e0daad502b6f770026ee88c0f032b8193eaf2b8d5cd0ddeca

Observation 80348ff6-5ca7-4d79-86b4-96438e8153d7 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T12:27:15.062636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:27:15.062636Z digest=sha256:739cff1851a7564ad5381cb3cbb9270db94e504e0e738d65e180f7f672284272

Pith citing papers

Observation de804997-082a-43ea-92b9-a3953a2a84a2 · inbound

Position: Theory of Mind Benchmarks are Broken for Large Language Models cites this paper.

Position: Theory of Mind Benchmarks are Broken for Large Language Models Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T00:07:41.420435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:07:41.420435Z digest=sha256:6b8e3640fe0f60ccc0b8fd1ff7d31fde3226146677f73b02887c6830fed23c18

Observation 7a2b46ba-fbf4-420c-bf80-95161b1c80c2 · inbound

RT-HCP: Dealing with Inference Delays and Sample Efficiency to Learn Directly on Robotic Platforms cites this paper.

RT-HCP: Dealing with Inference Delays and Sample Efficiency to Learn Directly on Robotic Platforms Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:17:52.993979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:17:52.993979Z digest=sha256:a77a6d6634724a84d83c6cc95271bcb73774d6938b5b5514bc17b3daa6d1c21f

Observation 6e89106f-9536-49e9-9b41-6d31c9ab3a81 · inbound

Finding the Time to Think: Learning Planning Budgets in Real-Time RL cites this paper.

Finding the Time to Think: Learning Planning Budgets in Real-Time RL Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:29:51.964800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T05:08:19.504454Z digest=sha256:c7cf26597120701aa69147c9ac99de19d986bdfc794038a75d4c503aec5f6bc3

Observation 20e97d7f-8c81-448b-9e81-6fec7a9ca9ef · inbound

Finding the Time to Think: Learning Planning Budgets in Real-Time RL cites this paper.

Finding the Time to Think: Learning Planning Budgets in Real-Time RL Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:34:34.896004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T09:26:31.944405Z digest=sha256:2ddda9c3c7b07df430e990c22f65b98912931ff28c84ae74b960e0c56f563290