Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 44 inbound Pith citation observations for arXiv:2510.13786.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:25:49.318391Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ae038977-eb78-49da-8e44-bfeb7f50ad5d · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 28668106-2e17-437f-9bed-27f75d82d55f · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Cwm: An open-weights llm for research on code generation with world models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3134f12a-286c-4f99-add0-6531030943b1 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd699a5d-3ef4-4bbd-9808-05ad9436106a · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 99ddcc83-aa44-45f4-8bdc-3c5d3c0a65e1 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Measuring Mathematical Problem Solving With the MATH Dataset
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4f6f3687-96ac-4848-bbd8-7f190b9b76ab · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Training Compute-Optimal Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6eaf8934-e196-45dd-91e2-c8af25ef7b38 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a999b556-f4d4-42ed-b4f6-329a44108c6c · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Scaling Laws for Neural Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8bf06b28-cf5e-4ad1-b1c1-798304e1d709 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Kimi K2: Open Agentic Intelligence
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 150fcd04-9d9e-427d-9dba-c88f8f72ac38 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 38ae454e-4f7b-403d-a66c-cc2cee262063 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Quantifying Variance in Evaluation Benchmarks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16e52cfc-216f-4d74-b354-52098de5ed16 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b26de1a-636b-4273-b24e-264f667f04ff · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Scaling Data-Constrained Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dce19e5b-ab95-4c40-b085-3e52dadf25c3 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs OpenAI o1 System Card
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 75ebfafe-8d71-4a7a-90c1-c041c162b77d · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs How predictable is language model benchmark performance?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 33c67b5a-9d66-4a3c-abb8-35214a292016 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Resolving Discrepancies in Compute-Optimal Scaling of Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9720cbe-1dcf-4070-8683-8df9cd1a62cd · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Observational Scaling Laws and the Predictability of Language Model Performance
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1591812d-8566-4314-bc4e-a05bd89db3d4 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Proximal Policy Optimization Algorithms
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9356063b-adf3-4de8-9c53-d56af148e932 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a3e4bc57-bbd7-42e9-a1b9-7b1ca3204f87 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0f043711-7720-401a-a758-a72b39b21b18 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9590be79-55de-442f-a855-e92239a3bc6d · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Qwen3 Technical Report
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2498844c-269a-4133-ac73-4c622c4676f8 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ac9027f1-458d-4505-a4d9-19450ba8c1d2 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2216c8df-3317-471e-98f8-c4ba2496518a · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b6ce8d30-d08c-425b-bdea-ea3187e8cd7e · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0eb73f6f-96c1-401d-bf78-f5b1608761cf · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0cb31d43-77b5-44dc-aea0-286e1738bf4d · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs (2017) for LLM fine-tuning with verifiable rewards
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f435c24c-0f8c-484b-a314-662c4c3ba130 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b75671b6-7cc2-4069-87b3-ae90f9f15f8c · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs The lowerϵ is to avoid gradient clipping (epsilon underflow) (Wortsman et al., 2023)
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e3afc2a3-101b-4fad-ace2-16d0932efb15 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs We use a custom code execution environment for coding problems involving unit tests and desired outputs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9677fa14-0eda-4d7d-88f6-1027e68e1c17 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs We consider two approaches: (a)interruptions, used in works like GLM-4.1V (GLM-V Team et al., 2025), and Qwen3 (Yang et al
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c173bf54-65e5-4a42-909e-fce8989db88f · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Okay, time is up. Let me stop thinking and formulate a final answer</think>
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 96997020-d3c7-48e3-ba2e-29fe91c5f2a9 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1632bf95-b2b1-4e9d-adbb-c9724059a112 · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bbf1ffff-e25f-4d55-acec-73cf9da2baac · outbound
The Art of Scaling Reinforcement Learning Compute for LLMs Specifically DAPO drops 0-variance prompts and samples more prompts until the batch is full
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 890744d3-a472-471a-80ab-4a0d270be6ee · inbound
Reinforcement Learning from Human Feedback The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7132b63c-3286-4b58-8493-a9a9e8218e2f · inbound
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332ccb78-ae5a-485b-a8ad-083a7a05303a · inbound
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e3886273-f2e8-45b3-9ed8-cf179c0f6d12 · inbound
Toward Training Superintelligent Software Agents through Self-Play SWE-RL The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8f6329e0-b46f-458e-b5d8-3be52158cd30 · inbound
Toward Training Superintelligent Software Agents through Self-Play SWE-RL The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60fdfe69-e7a5-42ec-929f-df86bb10a6d7 · inbound
StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ceacded-530a-40f6-a185-993217b6f506 · inbound
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d3b94f34-1b5e-429e-9d82-6df92ea03926 · inbound
Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98c45028-e813-4e69-853c-a91652bc950b · inbound
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20f917b-130e-4f41-a5f5-0e635846e0e1 · inbound
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8591c9cf-73fa-4f71-aae3-0acc77639aa9 · inbound
Continued AI Scaling Requires Repeated Efficiency Doublings The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 80bf37f4-c32e-4562-b049-4a83d4b7a9be · inbound
Target Policy Optimization The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4cd2668a-9b7d-46d4-98e4-0089d5428ba9 · inbound
Beyond Distribution Sharpening: The Importance of Task Rewards The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c957c6d0-7941-4db5-98ca-246120b096c0 · inbound
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d659a415-43c3-4a8b-ab55-4ee4580f6cba · inbound
Scaling Self-Play with Self-Guidance The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2e311385-4a51-4600-840a-99f31adea9d6 · inbound
Cost-Aware Learning The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bc70a71-71d9-4439-9d9f-5b13c6f717c5 · inbound
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a63f6d2-3f06-4cc1-88fc-82de48b8aa5d · inbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 283135aa-3c85-418a-9a72-0b50b15b7fae · inbound
ZAYA1-8B Technical Report The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 158
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da05758b-252e-4c53-a5a5-1d4486cbbafc · inbound
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4ae07ba-b1e9-4406-b13c-ed8be5e58f04 · inbound
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd73d658-bfb2-4e24-a3b5-dfeae5a15333 · inbound
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8cf8928e-2eaa-4253-bed9-5eb90f8e9d28 · inbound
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f4d0cb26-0180-4e15-be8c-a0b6b6dbee06 · inbound
KL for a KL: On-Policy Distillation with Control Variate Baseline The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9e8e7edf-f295-47ff-9b1f-1ba7ae188728 · inbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cccee1d1-89b2-44d5-b2cd-786401296878 · inbound
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2ad4c347-8285-43e9-945c-61ee19bac263 · inbound
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 48bbc7fc-73a8-4fbd-a577-d5bd6be6699a · inbound
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 336770af-d9bc-4cc6-9709-f44452ac097b · inbound
Learning, Fast and Slow: Towards LLMs That Adapt Continually The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ce70be2-cb62-4b06-ae07-e233b13cdf09 · inbound
Learning, Fast and Slow: Towards LLMs That Adapt Continually The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0d6d125-b30f-4223-8bd9-f922b29aa54d · inbound
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e7e3fe6c-d9f9-44af-8dfd-0893782dd9f5 · inbound
Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d40eddb5-23b1-42d0-9853-690827aeed2a · inbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fbc048f1-6ec6-4991-813c-a93f9d112039 · inbound
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3b985535-95dd-49a0-a977-75a043a7f9a9 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 90fe2acd-3c6f-4d6f-a600-5727a08e177a · inbound
Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1513e3a3-426d-4284-a48d-ea0393642148 · inbound
ZONOS2 Technical Report The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 183
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ea59d1be-41d0-4456-b515-19d2e8cd1779 · inbound
ZONOS2 Technical Report The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 183
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16474127-7b58-4d61-a363-5702f9e930e4 · inbound
Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b886bcb1-417d-4c26-8913-2c0937be3a9f · inbound
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af318e1f-883f-4e64-93af-cb0701df1549 · inbound
Mach-Mind-4-Flash Technical Report The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5290bcba-6c1a-4b26-8172-e73dffe05623 · inbound
Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84ad023-1d0f-4acb-8a1b-f1c422cb1437 · inbound
Understanding Reasoning from Pretraining to Post-Training The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e72b1c8-2290-4794-a5a2-a98a76ad0bd4 · inbound
ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 168
Source-reported events for the cited work
Unavailable: canonical work link unavailable.