Pith. sign in

Paper Citation Record · LEDGER

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies

As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2508.03194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.03194 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:10.038427Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact6
  • verified fuzzy6
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f31cc6b1-dd95-41d7-a058-ac0fba07e79b · outbound

This paper cites Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:01.594756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:01.594756Z digest=sha256:a748c87049b5fd1a25f65e349a7ec4d1bbe693ef913022cfaa58f4857c55f5a9

Observation 92cfd7de-2a3c-40c0-a7fb-d1657f3d4f4a · outbound

This paper cites DeepSeek-V3 Technical Report.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.097893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.097893Z digest=sha256:03b0152ae7c45ba9805501ef4823b443fe8dfe75e03b0eec66c7a1fffb151e58

Observation 296bd663-b3a3-44d7-bc7d-4bd91fe21d45 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.167095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.167095Z digest=sha256:e25dae83e381e88fd267eac9b1599299aba9dea5224c4b637b8aea8d6fdd66c1

Observation 35bd2519-08ce-4023-b351-bbb96f5dc2d8 · outbound

This paper cites Girshick.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Girshick

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:39:13.713021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:02.300702Z digest=sha256:cbb50f1971e97fe78a42dfccf89f2dc64e61f55074d55f95fd903f229df86b56

Observation 96e5cc89-4cfe-4d15-ab9a-c4abf82b6a4d · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Laws for Autoregressive Generative Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.375698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.375698Z digest=sha256:5fd522592fb8dbce1abb43cbd41d69480dffd09053aa63ef1a8bcd0529247f36

Observation 5d04cc46-6e85-4cf5-9072-1bfbf717c71d · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Laws for Neural Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.650942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.650942Z digest=sha256:5d65358a387beb44dac3f5588eb46118a363b73c825fbbe40cc3b89cd16a8a7d

Observation 47ec75fa-639c-468d-930a-43b7db5d5488 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Kimi K2: Open Agentic Intelligence

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.714676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.714676Z digest=sha256:f7186f8e140e4e62e9db31e9b705a23974a767afc9b6a1c16412285ef654f95c

Observation 6bf65cda-86bb-4f74-8cc7-f8bb7f90a255 · outbound

This paper cites SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.901710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.901710Z digest=sha256:a9f9563c96f287be3fddd7b81158858d00597aef988cf153cf6411ecde766006

Observation c276368b-ee5a-4e92-a097-6a6b93eaab6a · outbound

This paper cites Hyperspherical Normalization for Scalable Deep Reinforcement Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Hyperspherical Normalization for Scalable Deep Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.965845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.965845Z digest=sha256:d1d974e02a049d299b1ea0477c9ca2b22502845fbadac3fcd0d61813f130a7fd

Observation 8aeb3b80-7dba-4d80-92f9-104d6f66f7a0 · outbound

This paper cites HIPODE: Enhancing Offline Reinforcement Learning with High-Quality Synthetic Data from a Policy-Decoupled Approach.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies HIPODE: Enhancing Offline Reinforcement Learning with High-Quality Synthetic Data from a Policy-Decoupled Approach

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:39:12.697025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:03.194469Z digest=sha256:f0d76586b7a611d10fff08b5777dbef929e8aa189f995ceeefa338c5ae487dbf

Observation 56a88053-8484-4f20-8e2e-d2afacf86fab · outbound

This paper cites Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:03.382391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:03.382391Z digest=sha256:9455679becd6775415613ff1103df76147f2e7bafcc24dff52b888b4f012f910

Observation 07612d32-42f0-4bb3-aa7b-7ae7f7c8a48a · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Encouraging divergent thinking in large language models through multi-agent debate

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:39:13.524958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:03.493412Z digest=sha256:fda62ab0106eb19b0c518f49046aeaabb5aff2d61816d41795a0a7a8f8851546

Observation ae716863-0aee-4800-8dcb-9cc490fe3269 · outbound

This paper cites Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:03.647835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:03.647835Z digest=sha256:faaaea574bdfe3fc466c8172997c81cb709239d359c97a9e4a03173a27816534

Observation e62efe48-f42d-494a-8f2f-e426e9358a9c · outbound

This paper cites Continuous control with deep reinforcement learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Continuous control with deep reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:03.814369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:03.814369Z digest=sha256:4ddc99ead7cc14a509649efba45211b2c187d0e5724d8590529d7441b132609f

Observation 74b92fd2-80fe-457e-8dfc-f7d363cc0b3b · outbound

This paper cites Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:04.024925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:04.024925Z digest=sha256:dd887c03fd80b3aedd37017cde7e88d1d1e694870031f871b3309cdd68f493ac

Observation fc48908a-06b0-4ce2-b747-1ddc66b74c69 · outbound

This paper cites Evolutionary Action Selection for Gradient-based Policy Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Evolutionary Action Selection for Gradient-based Policy Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:39:12.504235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:04.163212Z digest=sha256:a80035c72d1ae2f0588788740db6823c3e8e2583473b1e7918edc4993952cc42

Observation 51f80b10-ba54-4325-9ee9-7544de71c491 · outbound

This paper cites Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:04.390690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:04.390690Z digest=sha256:105bba351114b903ef65fd062f69479c4c41182b254931353100c0eb89ea426e

Observation 92c82929-f926-4f15-9a69-7230b4fb38e3 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies SmolVLM: Redefining small and efficient multimodal models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:04.547064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:04.547064Z digest=sha256:9ab61ef4aa3c9c83e0888a3397dfe9cb9f16cf181829ef473e66331eec563a5c

Observation b41a1983-3b77-479e-b2b4-547f72c1f91f · outbound

This paper cites The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:39:12.305554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:04.763706Z digest=sha256:6945ea32ccb745c33ef0930b7d2081e85f55b206807fbdb3425827bf3adaec46

Observation da9c01ce-fac7-4106-a460-3eab38e783ba · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Playing Atari with Deep Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:04.975317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:04.975317Z digest=sha256:3e4a2ddc44f213ce95bc5f09bea1fca4b2426c86da6c0db36a7e38c8ccd7a3d2

Observation 4ac0e78e-37c2-4042-a3bb-e7e062983975 · outbound

This paper cites s1: Simple test-time scaling.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies s1: Simple test-time scaling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:05.355060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:05.355060Z digest=sha256:8332367b23bd34f81d1013c2ba5f472f5c1d2f9537ba34a66cd7c8b3c14acbb0

Observation bc77ead2-1f30-42fd-9916-5280d93a49ee · outbound

This paper cites Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:05.516882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:05.516882Z digest=sha256:3ffa7dcbcb463d50c6feefe967e76677e309610ed22d700e5f2dac5a2ccb8aac

Observation d69bd5f6-b29c-4465-8e1a-1022b874b711 · outbound

This paper cites Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:05.661675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:05.661675Z digest=sha256:4f4abaa9090fd3e9c05f24bf537764f8b3d9fb0bc9399b2da8c449bd2a92ddd1

Observation 4a4c2d2b-a15e-4637-9c60-fc9a674349fd · outbound

This paper cites GPT-4 Technical Report.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies GPT-4 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:05.858336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:05.858336Z digest=sha256:64b64764a8fa95c9296964f3029adaffa25fc10aa7c5d5ca336342f93a0224e6

Observation 47bb1096-30d7-43a5-a96a-02dbf597181b · outbound

This paper cites Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:06.014406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:06.014406Z digest=sha256:6b10e71d6cac3c2f58e8c71b0432a5a97c896f8e257358cdacdc0ebf66f10815

Observation 91a7e734-9132-4d03-924d-422473fb67a7 · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Generative Agents: Interactive Simulacra of Human Behavior

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:06.216173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:06.216173Z digest=sha256:343bbb3102d557029d6603c99dc27f24da4c3cb3d3d4bbeef04a8f7e2f4d4c3b

Observation c74ba78b-7db9-4f46-8b3c-08b0b3c1f5da · outbound

This paper cites Horizon reduction makes rl scalable.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Horizon reduction makes rl scalable

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:06.374492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:06.374492Z digest=sha256:8d5ca44a702be20d7e643007c1bb9f490b377b5636eac6b3e901544472c2b794

Observation 4322a865-7636-4446-89a9-6b6f44e3eac0 · outbound

This paper cites Qwen3 Technical Report.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Qwen3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:06.606659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:06.606659Z digest=sha256:f9f6049962ca76da8560522753ab7b291a7ad138bc4dd277c98714882d03913e

Observation e7d746db-8c02-4bce-8d81-72aab31f19d0 · outbound

This paper cites Value-Based Deep RL Scales Predictably.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Value-Based Deep RL Scales Predictably

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:06.797394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:06.797394Z digest=sha256:9be95dc60d0502db1ec0968a9603a6e4d43326fdcce40a7a38726c6f8e4f0988

Observation 0977ce65-9c9f-4280-b9d4-0b48ef20c4eb · outbound

This paper cites Proximal Policy Optimization Algorithms.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Proximal Policy Optimization Algorithms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:06.960332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:06.960332Z digest=sha256:c4c4f4fe2153918ffaf5d0b22c867245dc008182b5b405e783ad46b3c06b14a4

Observation 2866452a-e27c-4565-b71c-001936c8ccbe · outbound

This paper cites FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:07.111116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:07.111116Z digest=sha256:c0d2677cbb6bb7471fc2c4c981110a1db8a04d6673d7616e8b99717761610aeb

Observation 14655a44-f98d-4c09-9a60-a5429f3718e3 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:07.312483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:07.312483Z digest=sha256:67ba046a3ae8b1ab694b0135723076319e6003387a71cb7eef4937004007866b

Observation 8e775b0e-68c3-4b0e-923c-35b5b6e8b0b1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:07.470093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:07.470093Z digest=sha256:442b623ff75102c492ceffd19b8b37da28dc50ad5cf349e385e0f3c9285d4424

Observation 04bfc398-424b-4cec-927b-236058396c63 · outbound

This paper cites Q-Learning for Continuous Actions with Cross-Entropy Guided Policies.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Q-Learning for Continuous Actions with Cross-Entropy Guided Policies

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T04:39:11.817965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:07.677891Z digest=sha256:833da871ea20b846ac0f2cac86bf55e3edab74a695c5616cc9b78566af9a5aeb

Observation 7ede08de-cc52-470f-b1a8-d678163e58ef · outbound

This paper cites Accelerated Methods for Deep Reinforcement Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Accelerated Methods for Deep Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:07.841303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:07.841303Z digest=sha256:d564fab292d2417671d105c98ece25e9935740a972c49605b1610f0b73e06d86

Observation 5a627790-67da-4c89-ad59-60d9bc95771b · outbound

This paper cites DeepMind Control Suite.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies DeepMind Control Suite

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:07.992274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:07.992274Z digest=sha256:151635e2f56d75824f7dc1ff5079ae89eddf4139659c289cbc1ad6268bf1d7d6

Observation 7141075c-d58a-45f0-8b00-941519ddc798 · outbound

This paper cites MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:39:11.383388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:08.410191Z digest=sha256:8361d59d2c6e0b9743465a2c8f93aeb1b0743f8ed296dd079a891e20703459ad

Observation 84e66027-628c-4006-9e36-350330d06cc5 · outbound

This paper cites 1000 layer networks for self-supervised rl: Scaling depth can enable new goal-reaching capabilities.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies 1000 layer networks for self-supervised rl: Scaling depth can enable new goal-reaching capabilities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:08.569861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:08.569861Z digest=sha256:075088a220b4692f342dc7c2eb13e8e4a78b54e7cddf4471fbf557df9634b591

Observation 337e92b0-c427-4c5a-89f0-58d39828bdf0 · outbound

This paper cites Prioritized Generative Replay.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Prioritized Generative Replay

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:08.778688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:08.778688Z digest=sha256:84b5a61033eaf2c230541a24f3582dcc1aa19a508bbf15e1dcf0210258e368c9

Observation 0b721e05-0aca-4a3c-acb4-07f2f438a046 · outbound

This paper cites Aggressive Q-Learning with Ensembles: Achieving Both High Sample Efficiency and High Asymptotic Performance.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Aggressive Q-Learning with Ensembles: Achieving Both High Sample Efficiency and High Asymptotic Performance

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:39:11.084778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:08.949726Z digest=sha256:b3a4243e0efc1b1f57b02ee2fe7d41283a65ffb513fab18bddcf1280d296060b

Observation cabf089c-1e07-48ef-ae7f-f1337a470e7f · outbound

This paper cites Higher Replay Ratio Empowers Sample-Efficient Multi-Agent Reinforcement Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Higher Replay Ratio Empowers Sample-Efficient Multi-Agent Reinforcement Learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:39:13.115603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:09.151122Z digest=sha256:c6ef24b90254123b6553f5d4e5e34db0e4339010b9e25eaa76a85fae7b12b95b

Observation 31bb396d-a529-48b7-9603-00011f03b0a5 · outbound

This paper cites ISBN 979-8-3503-5067-8.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies ISBN 979-8-3503-5067-8

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:09.364581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:09.364581Z digest=sha256:db910f4a37e070f169b07ae28252167493fb9b4efab0ed4afd90dcd75def023e

Observation 876de244-d1a8-43a5-949e-3d1b876a59b0 · outbound

This paper cites Towards Applicable Reinforcement Learning: Improving the Generalization and Sample Efficiency with Policy Ensemble.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Towards Applicable Reinforcement Learning: Improving the Generalization and Sample Efficiency with Policy Ensemble

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T04:39:10.681491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:09.528457Z digest=sha256:6eb838903defb809b0ded384ec6011c72db8e252acc2e5646e0c86b8f43f8d94

Observation 7d462103-b0b5-4e41-af02-4e275f73f6ba · outbound

This paper cites Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:09.721695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:09.721695Z digest=sha256:4a793459937de9ab9d8f6350e5842742d3fe9841a99a90250ff29f31bd597ada

Observation 4b44c69d-357d-472a-9f0c-5501184607d3 · outbound

This paper cites doi: 10.3233/FAIA230609.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies doi: 10.3233/FAIA230609

Reference 58

Resolution
verified exact
doi, observed 2026-08-06T04:39:10.335341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:09.891263Z digest=sha256:24c293308f6e42a3ab543c7a6d283edea38bf1997deca5dba2c6ce45a4bec4a3

Observation 862ee5f6-ec79-47a6-9c0c-48aa3c87c5a4 · outbound

This paper cites Towards A Unified Policy Abstraction Theory and Representation Learning Approach in Markov Decision Processes.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Towards A Unified Policy Abstraction Theory and Representation Learning Approach in Markov Decision Processes

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:10.038427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:10.038427Z digest=sha256:7e74b33bc7e32197314f07b8689cdee81a93786560a70728aa9b163dd86bf713

Observation f8ff49ef-a2ff-469f-a26a-7272292de7a0 · outbound

This paper cites URLB: Unsupervised Reinforcement Learning Benchmark.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies URLB: Unsupervised Reinforcement Learning Benchmark

Reference 1989

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.796315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.796315Z digest=sha256:b835b1ded2ac7def7dac5a13405ea08547d2c701ce35e55ad2d8d9336786bcd9

Observation 7187d104-04a6-4c4b-a531-0a2095337020 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Representation Learning with Contrastive Predictive Coding

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:08.202891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:08.202891Z digest=sha256:03edbe3c806834615711ad79cd5929d6032ed22e7a5ef5cca864ed7b0f6789cf

Observation a134f84a-fae4-43cf-b484-559d0e84c0b2 · outbound

This paper cites Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:39:13.306588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:05.143346Z digest=sha256:2f62824962368045817d1e49059d5c45b2985efe84dfe950240b7bae506e59d6

Observation 6a5219e7-61fd-45cc-965a-1298248ee48f · outbound

This paper cites MyoSuite -- A contact-rich simulation suite for musculoskeletal motor control.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies MyoSuite -- A contact-rich simulation suite for musculoskeletal motor control

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:01.907116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:01.907116Z digest=sha256:983e9a3d5d01e61046e37ec752b2eb364d83c568a559a8533794ec280ee13518

Observation 60cc0364-9d78-44e9-bb00-880a11a28b55 · outbound

This paper cites Scaling laws for single-agent reinforcement learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling laws for single-agent reinforcement learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.461192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.461192Z digest=sha256:5be4636396e5c441c58b1b1f484dd35a63f40d413a64df4a67964b6ff43fd760

Observation 0a5445b3-83e8-45f8-bdab-b98429e1a043 · outbound

This paper cites Simplifying Deep Temporal Difference Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Simplifying Deep Temporal Difference Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.239780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.239780Z digest=sha256:48d8ea84bdd10c99171434ed99860f1f18a241590ac4f0c32ad2d0736bb1e183

Observation f1ea7b3f-1edb-49f9-93a2-677f9a47c633 · outbound

This paper cites UCB Exploration via Q-Ensembles.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies UCB Exploration via Q-Ensembles

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:01.982238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:01.982238Z digest=sha256:8dfae6cc6b6dd1b10f046c1e89af0240dd72c595d0503e814d1d1798bada6ada

Observation 115464bc-a918-4615-9f44-a6b1b78e8fdc · outbound

This paper cites OpenAI Gym.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies OpenAI Gym

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:01.830561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:01.830561Z digest=sha256:c1269f19688242c83a50041672775cc991d85648f9b362aec0fd2964d038850e

Observation 41b4c15e-3f08-4f6d-a758-9d2b65e1d10a · outbound

This paper cites Phasic policy gradient.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Phasic policy gradient

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:39:13.844864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:02.050002Z digest=sha256:432a9de0942e5187d27be34afea91f368c281975e91e500448b6f2263b377601

Observation b7addc73-5c8a-44e7-ac44-dd8f8b35bfbe · outbound

This paper cites Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:03.060592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:03.060592Z digest=sha256:7a232054825f0a386db37d2da792291228a6d3a4a6201242779fc9070cb463a0

Observation 4cb6c6bc-60cc-4f75-89d2-424822a76ed4 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Dota 2 with Large Scale Deep Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:01.745617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:01.745617Z digest=sha256:a9a2449a6dbff7d4e102a3d1351466cd6e1bff68c336ca9a76c0bcea671bde41

Observation decb3e77-a2a3-4f19-9afd-b62762b65cb1 · outbound

This paper cites Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.562616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.562616Z digest=sha256:f2f487874f19ef722771968eaa66e15c858cca6e2398cd94cd2136858df27174

Observation 469c9732-ca93-42e9-ba3f-7db283b79c39 · outbound

This paper cites Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:39:14.006027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T04:39:01.673860Z digest=sha256:08c198b897edb225249555fd077c26f47c0652458bf97b6ec18aa00811ddc4a9

Pith citing papers

No inbound Pith citation observations are available.