Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T18:45:23.213117Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 2 inbound Pith citation observations for arXiv:2603.06009.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T18:45:23.213117Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:56:51.940356Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T08:51:25.700420Z
69 of 69 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 324d468a-0ef7-41ab-b8a8-d109b3222946 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ea3926-2720-442f-ae65-02428ba04317 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf2f041-1fb4-4e78-9f94-50969c0b27b6 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Unifying count-based exploration and intrinsic motivation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28fb3107-307e-41a2-8469-f286488a6e10 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Nonlinear programming
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1df8c173-9bf6-4d1f-ac95-936ba5d5a96f · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Staggered environment resets improve massively parallel on-policy reinforcement learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb988d13-6dab-4f1c-b529-8a7da5239187 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Towards deeper deep reinforcement learning with spectral normalization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5888563a-2747-4fc3-b190-2c4515fb867d · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef4fe975-aabb-4c6f-92e3-5ef0f4d70094 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Mixtures of experts unlock parameter scaling for deep RL
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7eca6dc-852f-436d-8d4b-ad834820aeea · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Two-timescale networks for nonlinear value function approximation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1301120-9d07-4b2a-9c3e-3a6924f17330 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Bayen, Stuart Russell, Andrew Critch, and Sergey Levine
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2954065a-4892-47da-8cff-e2e127b8d2b8 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Revisiting lars for large batch training generalization of neural networks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25f19ba-1351-4903-a07a-df1185db8a76 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments First return, then explore
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43518c25-4e3d-4773-8eda-17dedb3c887c · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84acd0c9-d1fa-439f-ada0-722524b2f624 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f8730c-77fb-48c0-ac05-e358dc8b2efb · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c7485fc-90ed-4c3c-8089-40ae59723b91 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8e47bd5-a90c-43b4-8454-cfca8ed370d0 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Learning rates as a function of batch size: A random matrix theory approach to neural network training
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6d2f4c-d71e-477b-934b-848029d12e20 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Batch size-invariance for policy optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e6dcbc-c46c-4c9a-bfac-ee5c337af806 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Scaling laws for single-agent reinforcement learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb6e709-9532-4297-ae33-d5c542ea6ebf · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Position: Open-endedness is essential for artificial superhuman intelligence
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd5cc85-b604-4cb7-af86-976b18d92667 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments A Closer Look at Deep Policy Gradients
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf2602e-2935-463b-bfca-1da5cd997345 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Prioritized level replay
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ea2447c-fc09-4fd3-ab26-5eb21f0958eb · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Adam: A Method for Stochastic Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993ae644-4fa5-4a2f-bea7-cbd992480a14 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments One weird trick for parallelizing convolutional neural networks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ef049f-6854-45cd-a5fb-b9f56bc40c82 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments u ttler, Nantas Nardelli, Alexander Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rockt \
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44a1b1d-885f-44a0-80f2-e25c2b5b5dd8 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments gymnax : A JAX -based reinforcement learning environment library, 2022
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 838de2d9-fb23-4d12-b0b4-99a48bdbe22a · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Wurman, Jaegul Choo, Peter Stone, and Takuma Seno
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f9d073a-5ea2-4f7f-a7d9-3f3f246c324d · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Hyperspherical Normalization for Scalable Deep Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f861d90e-6353-42d7-87e6-c02021e58cae · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Linear and nonlinear programming, volume 2
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b46b9e-831e-4f74-8524-fb4bc35f4966 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Understanding and preventing capacity loss in reinforcement learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe001b47-43e7-46af-9320-0ad2ea4bede0 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Disentangling the causes of plasticity loss in neural networks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe17d574-42d3-444e-ac1e-727b272e7504 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Isaac gym: High performance GPU based physics simulation for robot learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a551a44c-467b-461f-9d42-f1d455fc40b9 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments On the sdes and scaling rules for adaptive gradient algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c67d8891-9c09-4bc8-b702-32f708d5aaba · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Craftax: A lightning-fast benchmark for open-ended reinforcement learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f068f8c5-a260-429a-a3cb-6ec93fe5ea20 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de666f82-b53d-4177-8517-a5745efb7be0 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments An Empirical Model of Large-Batch Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cbd09d7-fade-4eab-b894-4cf64e724026 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Multi-task reinforcement learning enables parameter scaling
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ae9d672-186c-45b1-99fe-59c41a664d58 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Bigger, regularized, optimistic: scaling for compute and sample efficient continuous control
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f8fc35-d3ca-401d-a589-16f3f8687429 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments The primacy bias in deep reinforcement learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8744c93-dc41-4b3f-bdbd-866b779031bf · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments XL and-minigrid: Scalable meta-reinforcement learning environments in JAX
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb04fbd9-728a-441e-80c6-3c6487dd0c41 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 25a4f697-d4df-4743-8780-96d65cf35f4a · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Numerical optimization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b20b069-19e4-4ba2-8f6e-040ab9f682c0 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Training Larger Networks for Deep Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc16fb24-25c5-4cc7-be89-c9b5748c58fc · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Evolving curricula with regret-based environment design
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7f2e7d-8748-499a-af22-96db5cc90e98 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Some methods of speeding up the convergence of iteration methods
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b763ee1-377f-43f5-a79f-e43d3f2b647e · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments A stochastic approximation method
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec434f3f-33f7-441b-a989-de607ede3ee9 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments JaxMARL: Multi-Agent RL Environments and Algorithms in JAX
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ba84456-3169-4b8a-8839-d9cba76eae96 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments No regrets: Investigating and improving regret approximations for curriculum discovery
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5377ea3c-39fd-4430-a539-6e3b7fb450d0 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Value-based deep RL scales predictably
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8a4a067-209c-474a-ba56-eca104e763b0 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b01029d-8b70-44eb-a41b-3577843ee218 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Proximal Policy Optimization Algorithms
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16f3a565-b1f6-40be-8766-d6f0c9ed186c · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Bigger, better, faster: Human-level atari with human-level efficiency
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1224cad9-23c2-4029-8ef8-44a43e3a4e2c · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments SAPG: Split and Aggregate Policy Gradients
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc960eb-0465-489d-9429-06cd16a32162 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Smith, Pieter-Jan Kindermans, and Quoc V
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92398630-9fb0-470f-8327-f3ce3b78e661 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Why open-endedness matters
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ded6a708-c11a-47c7-a5ac-091b9e5bdee0 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Characterization and Mitigation of Training Instabilities in Microscaling Formats
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276c54a0-5753-4591-b5aa-47e461ed3223 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments On Bonus-Based Exploration Methods in the Arcade Learning Environment
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 058f5053-c8a0-4149-a692-54718f1bd4a8 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Beyond the Boundaries of Proximal Policy Optimization
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 438a996f-6afd-4a5e-9be5-15d374122f90 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Human-Timescale Adaptation in an Open-Ended Task Space
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fb0f036-004b-491f-abaa-551f1b531280 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Open-Ended Learning Leads to Generally Capable Agents
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68765e25-3cf7-4553-8bc6-45ea749fe06e · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Efficient exploration in reinforcement learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f0c469-6455-43a9-ab8b-b8f11ad12f5b · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Stoix: Distributed Single-Agent Reinforcement Learning End-to-End in JAX , April 2024
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1eab8c5-2e04-4362-a662-209a46724942 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments 1000 layer networks for self-supervised RL : Scaling depth can enable new goal-reaching capabilities
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a12b9ae3-56ba-47c9-97dd-068273d774d4 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Truly proximal policy optimization
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58b54141-451e-44eb-83bd-2b6fda085b36 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments A finite-time analysis of two time-scale actor-critic methods
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615725cc-be96-4575-8c2d-cecc3ccf8d5a · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Kahrs, Carlo Sferrazza, Yuval Tassa, and Pieter Abbeel
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 719932f3-f12e-4724-8a26-baa72efd50e6 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Fast two-time-scale stochastic gradient method with applications in reinforcement learning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cccd674e-dd4a-47b1-9455-cc35bd91bd20 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments A two-time-scale stochastic optimization framework with applications in control and reinforcement learning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97d2c54e-5a2d-460d-983a-cfc813566cd2 · outbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Stabilizing reinforcement learning with llms: Formulation and practices
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 783a5005-aee6-441d-b65a-5f2230404964 · inbound
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8e32634-548b-434a-9051-726edda6527f · inbound
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.