Pith. sign in

Paper Citation Record · LEDGER

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments

As of 10 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 2 inbound Pith citation observations for arXiv:2603.06009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.06009 v2

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T18:45:23.213117Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:56:51.940356Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:51:25.700420Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 324d468a-0ef7-41ab-b8a8-d109b3222946 · outbound

This paper cites write newline.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.203273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.203273Z digest=sha256:328a340731668fd1c1951846bab711584c0d6b1fcaab8586b66dc19e1e9a2cd6

Observation 16ea3926-2720-442f-ae65-02428ba04317 · outbound

This paper cites What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.216901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.216901Z digest=sha256:47811463cf9b7535df809c2e634fed255bf0e0c08bdd00c0900694b78e788200

Observation faf2f041-1fb4-4e78-9f94-50969c0b27b6 · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Unifying count-based exploration and intrinsic motivation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.229450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.229450Z digest=sha256:7b59f89747e4bf667fe590162f9405444c0176c14f5c46f7f48f6080f51934df

Observation 28fb3107-307e-41a2-8469-f286488a6e10 · outbound

This paper cites Nonlinear programming.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Nonlinear programming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.241718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.241718Z digest=sha256:2bd4b90031ea8501bde0ced01ba6e79f60091b6a30d47b7387aeb9dedc61139e

Observation 1df8c173-9bf6-4d1f-ac95-936ba5d5a96f · outbound

This paper cites Staggered environment resets improve massively parallel on-policy reinforcement learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Staggered environment resets improve massively parallel on-policy reinforcement learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.253010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.253010Z digest=sha256:ed43d491cfd0bda21246ecf0f530464241b05aab464699d12767797c9cfbf81a

Observation eb988d13-6dab-4f1c-b529-8a7da5239187 · outbound

This paper cites Towards deeper deep reinforcement learning with spectral normalization.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Towards deeper deep reinforcement learning with spectral normalization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.267228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.267228Z digest=sha256:c943732643d0b23bf1835f73bedfe4bc07acdc57fedbdd368b75f64718fd9dbb

Observation 5888563a-2747-4fc3-b190-2c4515fb867d · outbound

This paper cites Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.286849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.286849Z digest=sha256:857275ad00c5e1aed4151e9ba10c57e15233d2bd3f74d600be6c262b0f9d5057

Observation ef4fe975-aabb-4c6f-92e3-5ef0f4d70094 · outbound

This paper cites Mixtures of experts unlock parameter scaling for deep RL.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Mixtures of experts unlock parameter scaling for deep RL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.303337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.303337Z digest=sha256:7a7941c5f6555fe4ead530db0083dc737b36a566b1e33985ca56994eb44eebf1

Observation f7eca6dc-852f-436d-8d4b-ad834820aeea · outbound

This paper cites Two-timescale networks for nonlinear value function approximation.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Two-timescale networks for nonlinear value function approximation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.319817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.319817Z digest=sha256:47ecf9e576923f88e37b6ffbd9b04f2e1f53197726a5c8f62c0f86b7d3da6386

Observation b1301120-9d07-4b2a-9c3e-3a6924f17330 · outbound

This paper cites Bayen, Stuart Russell, Andrew Critch, and Sergey Levine.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Bayen, Stuart Russell, Andrew Critch, and Sergey Levine

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.332979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.332979Z digest=sha256:27ab30d68f1e946e35da2ab3ae56309b95b74a2deab2c4c2603a3241673329e6

Observation 2954065a-4892-47da-8cff-e2e127b8d2b8 · outbound

This paper cites Revisiting lars for large batch training generalization of neural networks.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Revisiting lars for large batch training generalization of neural networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.346469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.346469Z digest=sha256:253c1cf9793fb4a3c9a7326b57cffa241070d827bc2a0492e556a9d44737bde7

Observation e25f19ba-1351-4903-a07a-df1185db8a76 · outbound

This paper cites First return, then explore.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments First return, then explore

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.411579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.411579Z digest=sha256:030accf5bf96d07754eabf81e8804dc51081f5fabe0c188fa0cbfc5ca0fed7a3

Observation 43518c25-4e3d-4773-8eda-17dedb3c887c · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.537597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.537597Z digest=sha256:b8d24d2b161df448d6deec2f596132419f6c66f7e3b130bf7a9ec60c656ee2b3

Observation 84acd0c9-d1fa-439f-ada0-722524b2f624 · outbound

This paper cites Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.648056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.648056Z digest=sha256:f068144c3eb1448082a2b20da85b95cd678219266ae08211ca4e0549e12938cf

Observation c0f8730c-77fb-48c0-ac05-e358dc8b2efb · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.779259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.779259Z digest=sha256:7fd98292256729ed223515a50de894141ef047c1bf27a02bd9c45f5644821b0e

Observation 7c7485fc-90ed-4c3c-8089-40ae59723b91 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.901597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.901597Z digest=sha256:28338d9024d73cd2741ef8f2e01580ac2c5d98e3e6efe101c61cf430811fad1c

Observation f8e47bd5-a90c-43b4-8454-cfca8ed370d0 · outbound

This paper cites Learning rates as a function of batch size: A random matrix theory approach to neural network training.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Learning rates as a function of batch size: A random matrix theory approach to neural network training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.987553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.987553Z digest=sha256:ea560f958748376bf2f11e5fb324f0dfd682bf788ed81368739590ddcfe6d339

Observation 5d6d2f4c-d71e-477b-934b-848029d12e20 · outbound

This paper cites Batch size-invariance for policy optimization.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Batch size-invariance for policy optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:21.086389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:21.086389Z digest=sha256:4b140658f920d6fe12699d0707c92b6fe3754d59a62740d80a043a96230f3337

Observation 23e6dcbc-c46c-4c9a-bfac-ee5c337af806 · outbound

This paper cites Scaling laws for single-agent reinforcement learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Scaling laws for single-agent reinforcement learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:21.203626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:21.203626Z digest=sha256:d04185db73903a533675b952dfe1a15747053965bf94045c7e5cea64773560df

Observation 1fb6e709-9532-4297-ae33-d5c542ea6ebf · outbound

This paper cites Position: Open-endedness is essential for artificial superhuman intelligence.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Position: Open-endedness is essential for artificial superhuman intelligence

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:21.320970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:21.320970Z digest=sha256:4695996f69fd0b5a14d91d0aae6d35451e37249f11832f0d0e19683a25aa334d

Observation 7dd5cc85-b604-4cb7-af86-976b18d92667 · outbound

This paper cites A Closer Look at Deep Policy Gradients.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments A Closer Look at Deep Policy Gradients

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:21.412418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:21.412418Z digest=sha256:3cd5c40d45d2ac26bf6b8e617768f8cd5716c3dc26e29a83e7e5b29f70b92fd3

Observation 6cf2602e-2935-463b-bfca-1da5cd997345 · outbound

This paper cites Prioritized level replay.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Prioritized level replay

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:21.492814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:21.492814Z digest=sha256:064cfae8cb72b0ff22bd6d4843f59e2c197ec31a332b524fa3cd8b3f466efa5c

Observation 6ea2447c-fc09-4fd3-ab26-5eb21f0958eb · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Adam: A Method for Stochastic Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:21.623678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:21.623678Z digest=sha256:108b92e56f5d90e861e3958208378414780967c068e0c4c523b745edc2778a89

Observation 993ae644-4fa5-4a2f-bea7-cbd992480a14 · outbound

This paper cites One weird trick for parallelizing convolutional neural networks.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments One weird trick for parallelizing convolutional neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:21.702560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:21.702560Z digest=sha256:5390783bea1b073d11ea00a3403f6a4b6f2cef644a24c8bbb677ce0509eb3c81

Observation 92ef049f-6854-45cd-a5fb-b9f56bc40c82 · outbound

This paper cites u ttler, Nantas Nardelli, Alexander Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rockt \.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments u ttler, Nantas Nardelli, Alexander Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rockt \

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:21.850672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:21.850672Z digest=sha256:8527faac6f5cf6d98c1162cceb97b6190acbd001421909256ade70d24604ae17

Observation d44a1b1d-885f-44a0-80f2-e25c2b5b5dd8 · outbound

This paper cites gymnax : A JAX -based reinforcement learning environment library, 2022.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments gymnax : A JAX -based reinforcement learning environment library, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:21.931314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:21.931314Z digest=sha256:d244fb64e95d2080b1899aaa09711d1a606f8a41343c2ec2e3a42131d2b67cdc

Observation 838de2d9-fb23-4d12-b0b4-99a48bdbe22a · outbound

This paper cites Wurman, Jaegul Choo, Peter Stone, and Takuma Seno.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Wurman, Jaegul Choo, Peter Stone, and Takuma Seno

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.011868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.011868Z digest=sha256:5d26dc2c5e4f43018256191367888f8cb5709b485f004e5f2dd994159ef4f878

Observation 4f9d073a-5ea2-4f7f-a7d9-3f3f246c324d · outbound

This paper cites Hyperspherical Normalization for Scalable Deep Reinforcement Learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Hyperspherical Normalization for Scalable Deep Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.126026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.126026Z digest=sha256:808b6ec440cf65331feaac894439feced0dc54bd346d8fba4f1fbf3cd7fcae96

Observation f861d90e-6353-42d7-87e6-c02021e58cae · outbound

This paper cites Linear and nonlinear programming, volume 2.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Linear and nonlinear programming, volume 2

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.251346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.251346Z digest=sha256:962ddb2c5a9a9d5e97dd930385fd7179bfd95c93d95cfba2d8483b15fdc479eb

Observation a2b46b9e-831e-4f74-8524-fb4bc35f4966 · outbound

This paper cites Understanding and preventing capacity loss in reinforcement learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Understanding and preventing capacity loss in reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.373975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.373975Z digest=sha256:ab94e01f18ffe704a1adb37385997a4c0bbb2c8852d0074b139b133d3de9ade0

Observation fe001b47-43e7-46af-9320-0ad2ea4bede0 · outbound

This paper cites Disentangling the causes of plasticity loss in neural networks.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Disentangling the causes of plasticity loss in neural networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.423849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.423849Z digest=sha256:5d622bd4a05c05d5e96fe31a1b96375ee046c9e2f60da366cc536b8af0cef3ff

Observation fe17d574-42d3-444e-ac1e-727b272e7504 · outbound

This paper cites Isaac gym: High performance GPU based physics simulation for robot learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Isaac gym: High performance GPU based physics simulation for robot learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.489621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.489621Z digest=sha256:6fb689550242d89b01afd5f1571719b26124b3da711a0ffee44248267fbda852

Observation a551a44c-467b-461f-9d42-f1d455fc40b9 · outbound

This paper cites On the sdes and scaling rules for adaptive gradient algorithms.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments On the sdes and scaling rules for adaptive gradient algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.574471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.574471Z digest=sha256:a40e373d32a96791ab251266fb333d5684e4e48aaf91f9c9e8a4dfb9f003ec35

Observation c67d8891-9c09-4bc8-b702-32f708d5aaba · outbound

This paper cites Craftax: A lightning-fast benchmark for open-ended reinforcement learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Craftax: A lightning-fast benchmark for open-ended reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.677138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.677138Z digest=sha256:81a60971afbbecf63f305c36f86a332d5ffe0346b4b6ace02068d282ca672271

Observation f068f8c5-a260-429a-a3cb-6ec93fe5ea20 · outbound

This paper cites Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.786428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.786428Z digest=sha256:f8c83bb6dfe8ca7a7ec2e5239d3e3aa17543bd669696d9de63052af194c96a48

Observation de666f82-b53d-4177-8517-a5745efb7be0 · outbound

This paper cites An Empirical Model of Large-Batch Training.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments An Empirical Model of Large-Batch Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.899915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.899915Z digest=sha256:3410b691c6fe8cf96f321f6d1f8bb34019d0da741323e1f0dc66fbd79ca24011

Observation 8cbd09d7-fade-4eab-b894-4cf64e724026 · outbound

This paper cites Multi-task reinforcement learning enables parameter scaling.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Multi-task reinforcement learning enables parameter scaling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:22.948817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:22.948817Z digest=sha256:13aeec53843995b9268f64e19792531c79302b9ff0954b8310bf700e6552b9c2

Observation 7ae9d672-186c-45b1-99fe-59c41a664d58 · outbound

This paper cites Bigger, regularized, optimistic: scaling for compute and sample efficient continuous control.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Bigger, regularized, optimistic: scaling for compute and sample efficient continuous control

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.005083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.005083Z digest=sha256:3178637e8fdc97043ef4f94bde78ba84c60d88dab900058a4ae63cb0dff8310d

Observation 01f8fc35-d3ca-401d-a589-16f3f8687429 · outbound

This paper cites The primacy bias in deep reinforcement learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments The primacy bias in deep reinforcement learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.062351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.062351Z digest=sha256:26093f706ee67340f9eadcc233c0ae2331010554b8ba1490a048309ae2c06f8b

Observation c8744c93-dc41-4b3f-bdbd-866b779031bf · outbound

This paper cites XL and-minigrid: Scalable meta-reinforcement learning environments in JAX.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments XL and-minigrid: Scalable meta-reinforcement learning environments in JAX

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.094899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.094899Z digest=sha256:42056811a1bc9a0ae58af5ccd96f8d8067d6cf1a728e68d5f6811ef820eb3caa

Observation cb04fbd9-728a-441e-80c6-3c6487dd0c41 · outbound

This paper cites XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement Learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement Learning

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-02T18:48:31.995581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-02T18:45:23.099039Z digest=sha256:e465f2e0b8eeab2aeb2dafb9536031a75ca9057b5a16b08a03698d92b4f0b374

Observation 25a4f697-d4df-4743-8780-96d65cf35f4a · outbound

This paper cites Numerical optimization.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Numerical optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.103513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.103513Z digest=sha256:f64bc7c5f3d865c1cb44fbbb7b152a54097a4392cdee864d63ed17e91261abe5

Observation 4b20b069-19e4-4ba2-8f6e-040ab9f682c0 · outbound

This paper cites Training Larger Networks for Deep Reinforcement Learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Training Larger Networks for Deep Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.108589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.108589Z digest=sha256:cf8f8d0920b2ca37b26ece9f365f2f4423b8a1ee731adc31aa2bd2fd5fb15c0c

Observation bc16fb24-25c5-4cc7-be89-c9b5748c58fc · outbound

This paper cites Evolving curricula with regret-based environment design.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Evolving curricula with regret-based environment design

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.113006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.113006Z digest=sha256:38acb8dcb5e87b5326be35c58905c761fb2225a61a02d560ba6f98a6d47d14fa

Observation 5d7f2e7d-8748-499a-af22-96db5cc90e98 · outbound

This paper cites Some methods of speeding up the convergence of iteration methods.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Some methods of speeding up the convergence of iteration methods

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.117078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.117078Z digest=sha256:d2caa22ab24f3335b39df923268ca16321ca151638884698360febf2ec9444ce

Observation 2b763ee1-377f-43f5-a79f-e43d3f2b647e · outbound

This paper cites A stochastic approximation method.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments A stochastic approximation method

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.121237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.121237Z digest=sha256:c132c50aa0dcae561828f3a7ebee70b2747058343f95e448751b9a22bd0602fa

Observation ec434f3f-33f7-441b-a989-de607ede3ee9 · outbound

This paper cites JaxMARL: Multi-Agent RL Environments and Algorithms in JAX.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments JaxMARL: Multi-Agent RL Environments and Algorithms in JAX

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.125183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.125183Z digest=sha256:231e186e403927600b586515f1b33de8206b809d6cb71558ff98264e4a65a1c1

Observation 0ba84456-3169-4b8a-8839-d9cba76eae96 · outbound

This paper cites No regrets: Investigating and improving regret approximations for curriculum discovery.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments No regrets: Investigating and improving regret approximations for curriculum discovery

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.129183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.129183Z digest=sha256:b501386d627427766de0075254a843dc0bef2da443e722d3d3767c63fe0b1d7e

Observation 5377ea3c-39fd-4430-a539-6e3b7fb450d0 · outbound

This paper cites Value-based deep RL scales predictably.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Value-based deep RL scales predictably

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.133253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.133253Z digest=sha256:0dd1fb5beeec949c61b55db88602670d76fd967e987efa98fa4cd95b13a6efd1

Observation c8a4a067-209c-474a-ba56-eca104e763b0 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.137181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.137181Z digest=sha256:abe2a01909e1345751db9b2ba34deae799641c26166ac69a6106ea46dff9db7a

Observation 1b01029d-8b70-44eb-a41b-3577843ee218 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Proximal Policy Optimization Algorithms

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.141968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.141968Z digest=sha256:b54e7fbd77cd5f43d4f0bea48b294366b37ec04d23ef64d5540a70072f86837b

Observation 16f3a565-b1f6-40be-8766-d6f0c9ed186c · outbound

This paper cites Bigger, better, faster: Human-level atari with human-level efficiency.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Bigger, better, faster: Human-level atari with human-level efficiency

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.145966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.145966Z digest=sha256:a609647901ae0674c65d6011ed25033d9b7023162a9b7e5884150477c86d76f7

Observation 1224cad9-23c2-4029-8ef8-44a43e3a4e2c · outbound

This paper cites SAPG: Split and Aggregate Policy Gradients.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments SAPG: Split and Aggregate Policy Gradients

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.149868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.149868Z digest=sha256:9927164928920e4df123a8c39598ffced55040f1e1187778ebf3a82730bff1de

Observation abc960eb-0465-489d-9429-06cd16a32162 · outbound

This paper cites Smith, Pieter-Jan Kindermans, and Quoc V.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Smith, Pieter-Jan Kindermans, and Quoc V

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.153984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.153984Z digest=sha256:bae30125865d5c313782e6568f405a1f87eac8d1ec144a7db8088786db8faff9

Observation 92398630-9fb0-470f-8327-f3ce3b78e661 · outbound

This paper cites Why open-endedness matters.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Why open-endedness matters

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.158003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.158003Z digest=sha256:760853b10df2aab725f9194a9c31591310cc727bbd36290bbc289a47e37090b1

Observation ded6a708-c11a-47c7-a5ac-091b9e5bdee0 · outbound

This paper cites Characterization and Mitigation of Training Instabilities in Microscaling Formats.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Characterization and Mitigation of Training Instabilities in Microscaling Formats

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.162011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.162011Z digest=sha256:0c7ee1c2bd7a5e5260ef07893a701cfd1794f45a94e48b7289e609f0b990428f

Observation 276c54a0-5753-4591-b5aa-47e461ed3223 · outbound

This paper cites On Bonus-Based Exploration Methods in the Arcade Learning Environment.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments On Bonus-Based Exploration Methods in the Arcade Learning Environment

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.166185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.166185Z digest=sha256:0eeff9da1559fd7378d0de77c7f1479c34cc6e573b53b27ca88d3035f7066ab9

Observation 058f5053-c8a0-4149-a692-54718f1bd4a8 · outbound

This paper cites Beyond the Boundaries of Proximal Policy Optimization.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Beyond the Boundaries of Proximal Policy Optimization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.170247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.170247Z digest=sha256:e676a876258fc8730ffb897b72e735d45cbc0df18bb9338f83fb64a97ebd01e5

Observation 438a996f-6afd-4a5e-9be5-15d374122f90 · outbound

This paper cites Human-Timescale Adaptation in an Open-Ended Task Space.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Human-Timescale Adaptation in an Open-Ended Task Space

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.174550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.174550Z digest=sha256:1953a1910e72f649ec353dfef55f2a5e4efa5453cf0318d408c8ded02e370145

Observation 4fb0f036-004b-491f-abaa-551f1b531280 · outbound

This paper cites Open-Ended Learning Leads to Generally Capable Agents.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Open-Ended Learning Leads to Generally Capable Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.178594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.178594Z digest=sha256:df537e8c04be73f46a57698841558dc995e396790c26f36c77b7d5fc2173bcab

Observation 68765e25-3cf7-4553-8bc6-45ea749fe06e · outbound

This paper cites Efficient exploration in reinforcement learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Efficient exploration in reinforcement learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.182591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.182591Z digest=sha256:eeb1254d2779873e3a056b0fdb200d3fe51f4930bbe7ba2391f02937b3372186

Observation 06f0c469-6455-43a9-ab8b-b8f11ad12f5b · outbound

This paper cites Stoix: Distributed Single-Agent Reinforcement Learning End-to-End in JAX , April 2024.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Stoix: Distributed Single-Agent Reinforcement Learning End-to-End in JAX , April 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.186289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.186289Z digest=sha256:99de91c9d6967c663bb7c53590f991752f1a5a91b6d928997adc3a85d97c8cb5

Observation d1eab8c5-2e04-4362-a662-209a46724942 · outbound

This paper cites 1000 layer networks for self-supervised RL : Scaling depth can enable new goal-reaching capabilities.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments 1000 layer networks for self-supervised RL : Scaling depth can enable new goal-reaching capabilities

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.189894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.189894Z digest=sha256:9060293dc05e300483b9e886b44973d299b02f98444dbcdcaa4544f75d08188c

Observation a12b9ae3-56ba-47c9-97dd-068273d774d4 · outbound

This paper cites Truly proximal policy optimization.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Truly proximal policy optimization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.193619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.193619Z digest=sha256:bd4cad8240dcd5791668eb006a99863b265c3b4b518dfce28499c8f78f854250

Observation 58b54141-451e-44eb-83bd-2b6fda085b36 · outbound

This paper cites A finite-time analysis of two time-scale actor-critic methods.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments A finite-time analysis of two time-scale actor-critic methods

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.197554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.197554Z digest=sha256:a7c370ddd4952112d665ac8cfbb9bbd07e03bcb9a5658bf7e71b2ffc7870aff0

Observation 615725cc-be96-4575-8c2d-cecc3ccf8d5a · outbound

This paper cites Kahrs, Carlo Sferrazza, Yuval Tassa, and Pieter Abbeel.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Kahrs, Carlo Sferrazza, Yuval Tassa, and Pieter Abbeel

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.201163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.201163Z digest=sha256:f33563eb6e8eac15d9deb9d8c767537d2a933f79fdc39466169d321d11e59f8c

Observation 719932f3-f12e-4724-8a26-baa72efd50e6 · outbound

This paper cites Fast two-time-scale stochastic gradient method with applications in reinforcement learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Fast two-time-scale stochastic gradient method with applications in reinforcement learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.205161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.205161Z digest=sha256:3d26d2e144e836eab63b8a04bba1d438fdb5bc235f1bca21eeb3efc84aa0aefc

Observation cccd674e-dd4a-47b1-9455-cc35bd91bd20 · outbound

This paper cites A two-time-scale stochastic optimization framework with applications in control and reinforcement learning.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments A two-time-scale stochastic optimization framework with applications in control and reinforcement learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.209039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.209039Z digest=sha256:6185b86c5b78d8ed21d479326d26d13f7dac91d95deb200829f5af4642a40608

Observation 97d2c54e-5a2d-460d-983a-cfc813566cd2 · outbound

This paper cites Stabilizing reinforcement learning with llms: Formulation and practices.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Stabilizing reinforcement learning with llms: Formulation and practices

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:23.213117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:23.213117Z digest=sha256:aac77f523c1399ea22c329b0287c645fa63ee65826117cd0d12868ea640d2bca

Pith citing papers

Observation 783a5005-aee6-441d-b65a-5f2230404964 · inbound

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control cites this paper.

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-21T01:19:54.005423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T13:34:06.461850Z digest=sha256:56bfcb8f9c3965119d0ac8d268017411399c62f2e17713c980834d4436fbb1f3

Observation b8e32634-548b-434a-9051-726edda6527f · inbound

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control cites this paper.

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-21T01:19:54.005423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:56:51.940356Z digest=sha256:bc9f27ac3a0a720236ab22e4f58c1fb3bb72deb94be1e6cc8fb49535a043fc1d