Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:10.038427Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2508.03194.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:10.038427Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f31cc6b1-dd95-41d7-a058-ac0fba07e79b · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92cfd7de-2a3c-40c0-a7fb-d1657f3d4f4a · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies DeepSeek-V3 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 296bd663-b3a3-44d7-bc7d-4bd91fe21d45 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35bd2519-08ce-4023-b351-bbb96f5dc2d8 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Girshick
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96e5cc89-4cfe-4d15-ab9a-c4abf82b6a4d · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Laws for Autoregressive Generative Modeling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d04cc46-6e85-4cf5-9072-1bfbf717c71d · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Laws for Neural Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47ec75fa-639c-468d-930a-43b7db5d5488 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Kimi K2: Open Agentic Intelligence
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf65cda-86bb-4f74-8cc7-f8bb7f90a255 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c276368b-ee5a-4e92-a097-6a6b93eaab6a · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Hyperspherical Normalization for Scalable Deep Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aeb3b80-7dba-4d80-92f9-104d6f66f7a0 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies HIPODE: Enhancing Offline Reinforcement Learning with High-Quality Synthetic Data from a Policy-Decoupled Approach
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56a88053-8484-4f20-8e2e-d2afacf86fab · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07612d32-42f0-4bb3-aa7b-7ae7f7c8a48a · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Encouraging divergent thinking in large language models through multi-agent debate
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae716863-0aee-4800-8dcb-9cc490fe3269 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62efe48-f42d-494a-8f2f-e426e9358a9c · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Continuous control with deep reinforcement learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b92fd2-80fe-457e-8dfc-f7d363cc0b3b · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc48908a-06b0-4ce2-b747-1ddc66b74c69 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Evolutionary Action Selection for Gradient-based Policy Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51f80b10-ba54-4325-9ee9-7544de71c491 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c82929-f926-4f15-9a69-7230b4fb38e3 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies SmolVLM: Redefining small and efficient multimodal models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b41a1983-3b77-479e-b2b4-547f72c1f91f · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da9c01ce-fac7-4106-a460-3eab38e783ba · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Playing Atari with Deep Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ac0e78e-37c2-4042-a3bb-e7e062983975 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies s1: Simple test-time scaling
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc77ead2-1f30-42fd-9916-5280d93a49ee · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d69bd5f6-b29c-4465-8e1a-1022b874b711 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4c2d2b-a15e-4637-9c60-fc9a674349fd · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies GPT-4 Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47bb1096-30d7-43a5-a96a-02dbf597181b · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a7e734-9132-4d03-924d-422473fb67a7 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Generative Agents: Interactive Simulacra of Human Behavior
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c74ba78b-7db9-4f46-8b3c-08b0b3c1f5da · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Horizon reduction makes rl scalable
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4322a865-7636-4446-89a9-6b6f44e3eac0 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Qwen3 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d746db-8c02-4bce-8d81-72aab31f19d0 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Value-Based Deep RL Scales Predictably
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0977ce65-9c9f-4280-b9d4-0b48ef20c4eb · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Proximal Policy Optimization Algorithms
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2866452a-e27c-4565-b71c-001936c8ccbe · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14655a44-f98d-4c09-9a60-a5429f3718e3 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e775b0e-68c3-4b0e-923c-35b5b6e8b0b1 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04bfc398-424b-4cec-927b-236058396c63 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Q-Learning for Continuous Actions with Cross-Entropy Guided Policies
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ede08de-cc52-470f-b1a8-d678163e58ef · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Accelerated Methods for Deep Reinforcement Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a627790-67da-4c89-ad59-60d9bc95771b · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies DeepMind Control Suite
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7141075c-d58a-45f0-8b00-941519ddc798 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84e66027-628c-4006-9e36-350330d06cc5 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies 1000 layer networks for self-supervised rl: Scaling depth can enable new goal-reaching capabilities
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337e92b0-c427-4c5a-89f0-58d39828bdf0 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Prioritized Generative Replay
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b721e05-0aca-4a3c-acb4-07f2f438a046 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Aggressive Q-Learning with Ensembles: Achieving Both High Sample Efficiency and High Asymptotic Performance
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cabf089c-1e07-48ef-ae7f-f1337a470e7f · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Higher Replay Ratio Empowers Sample-Efficient Multi-Agent Reinforcement Learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31bb396d-a529-48b7-9603-00011f03b0a5 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies ISBN 979-8-3503-5067-8
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 876de244-d1a8-43a5-949e-3d1b876a59b0 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Towards Applicable Reinforcement Learning: Improving the Generalization and Sample Efficiency with Policy Ensemble
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d462103-b0b5-4e41-af02-4e275f73f6ba · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b44c69d-357d-472a-9f0c-5501184607d3 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies doi: 10.3233/FAIA230609
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 862ee5f6-ec79-47a6-9c0c-48aa3c87c5a4 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Towards A Unified Policy Abstraction Theory and Representation Learning Approach in Markov Decision Processes
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ff49ef-a2ff-469f-a26a-7272292de7a0 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies URLB: Unsupervised Reinforcement Learning Benchmark
Reference 1989
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7187d104-04a6-4c4b-a531-0a2095337020 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Representation Learning with Contrastive Predictive Coding
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a134f84a-fae4-43cf-b484-559d0e84c0b2 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a5219e7-61fd-45cc-965a-1298248ee48f · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies MyoSuite -- A contact-rich simulation suite for musculoskeletal motor control
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60cc0364-9d78-44e9-bb00-880a11a28b55 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling laws for single-agent reinforcement learning
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5445b3-83e8-45f8-bdab-b98429e1a043 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Simplifying Deep Temporal Difference Learning
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ea7b3f-1edb-49f9-93a2-677f9a47c633 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies UCB Exploration via Q-Ensembles
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 115464bc-a918-4615-9f44-a6b1b78e8fdc · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies OpenAI Gym
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b4c15e-3f08-4f6d-a758-9d2b65e1d10a · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Phasic policy gradient
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7addc73-5c8a-44e7-ac44-dd8f8b35bfbe · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb6c6bc-60cc-4f75-89d2-424822a76ed4 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Dota 2 with Large Scale Deep Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation decb3e77-a2a3-4f19-9afd-b62762b65cb1 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 469c9732-ca93-42e9-ba3f-7db283b79c39 · outbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.