Pith. sign in

Paper Citation Record · LEDGER

The Art of Scaling Reinforcement Learning Compute for LLMs

As of 4 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 44 inbound Pith citation observations for arXiv:2510.13786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.13786 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:25:49.318391Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact12
  • verified fuzzy6
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ae038977-eb78-49da-8e44-bfeb7f50ad5d · outbound

This paper cites an unresolved cited work.

The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:29:14.085252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:870449644f11e404b6a05a5ff7de5e22cc7095a72aca1d34f8aeba108b6bf6eb

Observation 28668106-2e17-437f-9bed-27f75d82d55f · outbound

This paper cites Cwm: An open-weights llm for research on code generation with world models.

The Art of Scaling Reinforcement Learning Compute for LLMs Cwm: An open-weights llm for research on code generation with world models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:13.976324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:b271ddd77c82208aea274b9795711c5dcbfa8acb5367268be544ae3e2c01bd48

Observation 3134f12a-286c-4f99-add0-6531030943b1 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

The Art of Scaling Reinforcement Learning Compute for LLMs The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:13.979906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:5910974ac8c8b9c8a3755a4a7b429d571f7d72d622b160dc1b3278d7b07bbbb4

Observation bd699a5d-3ef4-4bbd-9808-05ad9436106a · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

The Art of Scaling Reinforcement Learning Compute for LLMs GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:29:13.983934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:3394eba208f61a08cdbac55fa198eebec903b3f5d30fda4d14b0d1de24ac1762

Observation 99ddcc83-aa44-45f4-8bdc-3c5d3c0a65e1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

The Art of Scaling Reinforcement Learning Compute for LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:29:13.972253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:39436405c23e55ab06beb847702b090a150c0f4a32ed1f7568ef219ce3a2598c

Observation 4f6f3687-96ac-4848-bbd8-7f190b9b76ab · outbound

This paper cites Training Compute-Optimal Large Language Models.

The Art of Scaling Reinforcement Learning Compute for LLMs Training Compute-Optimal Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:13.987105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:97f5c7eb2b11c3b8365d109320f450f5853130765d05e187bc091136fe4a0334

Observation 6eaf8934-e196-45dd-91e2-c8af25ef7b38 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

The Art of Scaling Reinforcement Learning Compute for LLMs REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:13.990510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:d0680fec0ba7d9fc7bed3ca64f7183696a41aa08f63893c47a6ab37253b2f284

Observation a999b556-f4d4-42ed-b4f6-329a44108c6c · outbound

This paper cites Scaling Laws for Neural Language Models.

The Art of Scaling Reinforcement Learning Compute for LLMs Scaling Laws for Neural Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:13.993355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:461867ebe8968d5fb39281608ba19d556db545a44ebb278934c87d3e1ae41f7d

Observation 8bf06b28-cf5e-4ad1-b1c1-798304e1d709 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

The Art of Scaling Reinforcement Learning Compute for LLMs Kimi K2: Open Agentic Intelligence

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:13.996692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:7d3b65608b2e1bc8a435d84c8e2e18a2a11894286643c435d41787d05e63b064

Observation 150fcd04-9d9e-427d-9dba-c88f8f72ac38 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

The Art of Scaling Reinforcement Learning Compute for LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:c4dbc9081c0a85723a6975271e4f7d5f47303c2985e5eb03f7c11f62485e5241

Observation 38ae454e-4f7b-403d-a66c-cc2cee262063 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

The Art of Scaling Reinforcement Learning Compute for LLMs Quantifying Variance in Evaluation Benchmarks

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.003804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:6bdea95d28cc11e550a1802900895fafbd2f4ec97f8c5d942a4e5416cb1ff60d

Observation 16e52cfc-216f-4d74-b354-52098de5ed16 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

The Art of Scaling Reinforcement Learning Compute for LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:29:14.006669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:a0f3d8610addba82c6865d591c63042e5dee86e1c7c7ccbb46e99df0a9515cb3

Observation 7b26de1a-636b-4273-b24e-264f667f04ff · outbound

This paper cites Scaling Data-Constrained Language Models.

The Art of Scaling Reinforcement Learning Compute for LLMs Scaling Data-Constrained Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:35:21.865142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:547ad32d8478130c85780a639a286ce43ef165d959cadbde39a05f90bc386ce0

Observation dce19e5b-ab95-4c40-b085-3e52dadf25c3 · outbound

This paper cites OpenAI o1 System Card.

The Art of Scaling Reinforcement Learning Compute for LLMs OpenAI o1 System Card

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:14.013407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:141987a2e7047bbba9f606290cc1a29349242c5097634886b7ad57f3c8bfbc93

Observation 75ebfafe-8d71-4a7a-90c1-c041c162b77d · outbound

This paper cites How predictable is language model benchmark performance?.

The Art of Scaling Reinforcement Learning Compute for LLMs How predictable is language model benchmark performance?

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.017077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:5cd1567516d484e86e83f678e1cea26d186f05da7ed890f3e44e7f477f8855f6

Observation 33c67b5a-9d66-4a3c-abb8-35214a292016 · outbound

This paper cites Resolving Discrepancies in Compute-Optimal Scaling of Language Models.

The Art of Scaling Reinforcement Learning Compute for LLMs Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.020476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:3f73bd7ac528126604d1a43d4ec0597d8bda0de3829a42d30faa9ae233fffbee

Observation b9720cbe-1dcf-4070-8683-8df9cd1a62cd · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

The Art of Scaling Reinforcement Learning Compute for LLMs Observational Scaling Laws and the Predictability of Language Model Performance

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.024807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:73e92f15e61ccfeabca2cca9009b6eb2421510207d8585fbfb45e60f49fb977b

Observation 1591812d-8566-4314-bc4e-a05bd89db3d4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Art of Scaling Reinforcement Learning Compute for LLMs Proximal Policy Optimization Algorithms

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:29:14.028011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:e5f6aa0c516b11f6e78bb1f7d3c8b2d9994c0df2d98e4b31e163091169f901fa

Observation 9356063b-adf3-4de8-9c53-d56af148e932 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

The Art of Scaling Reinforcement Learning Compute for LLMs Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T01:42:19.263401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:3bc182296049eb4043598edc5eedb246e93b3407bfd643eab2abd2d55fb9e971

Observation a3e4bc57-bbd7-42e9-a1b9-7b1ca3204f87 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

The Art of Scaling Reinforcement Learning Compute for LLMs Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:14.034600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:3849c2dea74132568e61566229fd86a969e3f22b08521d65f4924788866e61d9

Observation 0f043711-7720-401a-a758-a72b39b21b18 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

The Art of Scaling Reinforcement Learning Compute for LLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.038199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:25ae77db6c5c3c8d82d9e48360da968768a343948c521a115f89cb8273a4ddcd

Observation 9590be79-55de-442f-a855-e92239a3bc6d · outbound

This paper cites Qwen3 Technical Report.

The Art of Scaling Reinforcement Learning Compute for LLMs Qwen3 Technical Report

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:29:14.042708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:848ac63da6addaa9e431487c96e0176c9cc6f57b1fb1ac2d3b7c142b95e8f6d7

Observation 2498844c-269a-4133-ac73-4c622c4676f8 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

The Art of Scaling Reinforcement Learning Compute for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:29:14.046510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:3753d704770673df99895f909492f520c4a440930aadb238bb9004953f7271df

Observation ac9027f1-458d-4505-a4d9-19450ba8c1d2 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

The Art of Scaling Reinforcement Learning Compute for LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.050525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:9166f3897d47bb104c5744b9a7d5f7b5dbbdae14cc912b14eeb42c23dc00ae47

Observation 2216c8df-3317-471e-98f8-c4ba2496518a · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

The Art of Scaling Reinforcement Learning Compute for LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:14.058755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:cff0ed50720772344ac69fcf1e8fbb8d9d393806609ffe6ad1bce7fe5ccc681b

Observation b6ce8d30-d08c-425b-bdea-ea3187e8cd7e · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

The Art of Scaling Reinforcement Learning Compute for LLMs Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.055173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:ece5e6cff088232885fffdb6a6afe1165034b3cb8305b3ed2cff2ad47e7559db

Observation 0eb73f6f-96c1-401d-bf78-f5b1608761cf · outbound

This paper cites an unresolved cited work.

The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:29:14.063973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:adb0e69b07d322ebf1910596474580a7f9e339f79039ae86d3672b0884578dd3

Observation 0cb31d43-77b5-44dc-aea0-286e1738bf4d · outbound

This paper cites (2017) for LLM fine-tuning with verifiable rewards.

The Art of Scaling Reinforcement Learning Compute for LLMs (2017) for LLM fine-tuning with verifiable rewards

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:29:14.066378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:bd4a3cdedf47a471a86d2246d0d9f6449f6c0325acb1f2120a1565b2a3a07882

Observation f435c24c-0f8c-484b-a314-662c4c3ba130 · outbound

This paper cites an unresolved cited work.

The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:29:14.068606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:c3c284da16cc483ac3478ae2966fc892f4ba2cc31ca18cddf5097b2a04059f5f

Observation b75671b6-7cc2-4069-87b3-ae90f9f15f8c · outbound

This paper cites The lowerϵ is to avoid gradient clipping (epsilon underflow) (Wortsman et al., 2023).

The Art of Scaling Reinforcement Learning Compute for LLMs The lowerϵ is to avoid gradient clipping (epsilon underflow) (Wortsman et al., 2023)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:29:14.071043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:2952c69673ad0eaa25583de6d32fdfb8f4d441b3fed81bf919d1e420947b371f

Observation e3afc2a3-101b-4fad-ace2-16d0932efb15 · outbound

This paper cites We use a custom code execution environment for coding problems involving unit tests and desired outputs.

The Art of Scaling Reinforcement Learning Compute for LLMs We use a custom code execution environment for coding problems involving unit tests and desired outputs

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:29:14.073989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:fda0923bee2dfa5e68dad9bd2a08cdcb5b073a0572c46e3d1680d9277b6a1a88

Observation 9677fa14-0eda-4d7d-88f6-1027e68e1c17 · outbound

This paper cites We consider two approaches: (a)interruptions, used in works like GLM-4.1V (GLM-V Team et al., 2025), and Qwen3 (Yang et al.

The Art of Scaling Reinforcement Learning Compute for LLMs We consider two approaches: (a)interruptions, used in works like GLM-4.1V (GLM-V Team et al., 2025), and Qwen3 (Yang et al

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:29:14.076619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:ff852a80533ee2a14c6dd14805ff5a2b6bac234da76c37011661579e0ce60c1c

Observation c173bf54-65e5-4a42-909e-fce8989db88f · outbound

This paper cites Okay, time is up. Let me stop thinking and formulate a final answer</think>.

The Art of Scaling Reinforcement Learning Compute for LLMs Okay, time is up. Let me stop thinking and formulate a final answer</think>

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:29:14.078959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:a233bda587bef3b453bdcf263edbc9ca23d323b0e5e0256d640b7ab9b5d8377d

Observation 96997020-d3c7-48e3-ba2e-29fe91c5f2a9 · outbound

This paper cites an unresolved cited work.

The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:29:14.081042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:c2b7c7a3355fcb16de523c2b706f19705a8124225291dfb267b8a54e4d0cee05

Observation 1632bf95-b2b1-4e9d-adbb-c9724059a112 · outbound

This paper cites an unresolved cited work.

The Art of Scaling Reinforcement Learning Compute for LLMs Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:29:14.083158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:2e9fd12fe15790cb235e3bfa819e70e49ed07f43dc8d6cbb4b56f53e4027a233

Observation bbf1ffff-e25f-4d55-acec-73cf9da2baac · outbound

This paper cites Specifically DAPO drops 0-variance prompts and samples more prompts until the batch is full.

The Art of Scaling Reinforcement Learning Compute for LLMs Specifically DAPO drops 0-variance prompts and samples more prompts until the batch is full

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:29:14.061537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:3aef693a8945a4c15456861da3e0760286c376cddf22af6b64d525919b98e9e1

Pith citing papers

Observation 890744d3-a472-471a-80ab-4a0d270be6ee · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.153805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:d9178b4acc768da507a2aff173d1bf4e95957851015822724090e9e8a011eb25

Observation 7132b63c-3286-4b58-8493-a9a9e8218e2f · inbound

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments cites this paper.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.327269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.327269Z digest=sha256:e76a6ac84baa0d814ea99415eefe1c2678f20a91b1d9193016930a2ccd327e7b

Observation 332ccb78-ae5a-485b-a8ad-083a7a05303a · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T22:43:37.855854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:205b7725979233d27f3ba6226aa82d815ec9783e5b78d178b6a650fda18f21f2

Observation e3886273-f2e8-45b3-9ed8-cf179c0f6d12 · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:10:20.191635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T16:07:48.570995Z digest=sha256:73ad6a86320cffd44b3ccc1cf5631e5dee2a5b8657b5e148171f3df047698b6e

Observation 8f6329e0-b46f-458e-b5d8-3be52158cd30 · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T15:02:11.739581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:02:11.739581Z digest=sha256:8c294ff97978c21f38403b943e655098d11fc17bf657ead0809eaa2432366be0

Observation 60fdfe69-e7a5-42ec-929f-df86bb10a6d7 · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:49.318391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:49.318391Z digest=sha256:8f3a4c1c2ee140beded35533431038d68a967d25b166616b8215430bc006a521

Observation 5ceacded-530a-40f6-a185-993217b6f506 · inbound

Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL cites this paper.

Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:40:12.439577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:37:20.114152Z digest=sha256:b7649bbe6f56355d4428ed8f7d72ace1bbd399ad5b1bcd7c7169a230fc61b207

Observation d3b94f34-1b5e-429e-9d82-6df92ea03926 · inbound

Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design cites this paper.

Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:44:11.524200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:42:03.896311Z digest=sha256:8168f0818304cb51d18cb74bc2e85c5a1a3d7563402754465ab79431411bde47

Observation 98c45028-e813-4e69-853c-a91652bc950b · inbound

F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare cites this paper.

F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:55:00.850928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:55:00.850928Z digest=sha256:b0e595719f84e4f5b2c1632034049026c013969d226a06eb121e3ba87c7a5989

Observation a20f917b-130e-4f41-a5f5-0e635846e0e1 · inbound

Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning cites this paper.

Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T21:38:56.393217Z digest=sha256:7d40309ad9f6cae5f41922976cc65761d166707944b10680323021dbf4395f67

Observation 8591c9cf-73fa-4f71-aae3-0acc77639aa9 · inbound

Continued AI Scaling Requires Repeated Efficiency Doublings cites this paper.

Continued AI Scaling Requires Repeated Efficiency Doublings The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:05.116491Z digest=sha256:8539511d97d76e2d47709b46a75399ac06944d9ab95b4bf90d298158811891d3

Observation 80bf37f4-c32e-4562-b049-4a83d4b7a9be · inbound

Target Policy Optimization cites this paper.

Target Policy Optimization The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:21:37.744591Z digest=sha256:3afb0c51503b275deb620c5a833fefc1315c257242a3f9f0cc9161a240f3a24b

Observation 4cd2668a-9b7d-46d4-98e4-0089d5428ba9 · inbound

Beyond Distribution Sharpening: The Importance of Task Rewards cites this paper.

Beyond Distribution Sharpening: The Importance of Task Rewards The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T08:07:14.691463Z digest=sha256:f7d915e9ba9a256c66081e676e386c8df883f4b0bd65814ab5ec7274cc4620be

Observation c957c6d0-7941-4db5-98ca-246120b096c0 · inbound

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes cites this paper.

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:18.468593Z digest=sha256:5f101a984c895c2e02d221ba6a9ea221eb575ed3abf710059e34649400775672

Observation d659a415-43c3-4a8b-ab55-4ee4580f6cba · inbound

Scaling Self-Play with Self-Guidance cites this paper.

Scaling Self-Play with Self-Guidance The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T01:31:06.090698Z digest=sha256:bf21dc2ab13f3c61dab69b464d4b867cfb58ce252d860d8518b36fe2a04bf88d

Observation 2e311385-4a51-4600-840a-99f31adea9d6 · inbound

Cost-Aware Learning cites this paper.

Cost-Aware Learning The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:11:01.131590Z digest=sha256:5b7a747c633aca0d3f9b486322e6033b7ab2270798c642ae61e72785910aacee

Observation 9bc70a71-71d9-4439-9d9f-5b13c6f717c5 · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:a400d53f47ae00f89241e034c1c1062afe2a43c82df44433fe26f75a1c23c8fb

Observation 5a63f6d2-3f06-4cc1-88fc-82de48b8aa5d · inbound

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO cites this paper.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:9552799c6e79575278e22cf1b21bf5d436acfc1c6112992d99e10ba889808990

Observation 283135aa-3c85-418a-9a72-0b50b15b7fae · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 158

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:ff3addeda196ebc50067c7c5fedb5b061b8be77aafbfaa74416bf9ded1f63615

Observation da05758b-252e-4c53-a5a5-1d4486cbbafc · inbound

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key cites this paper.

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T09:35:47.501360Z digest=sha256:127e6f4ebac45ce377c10d9f85b95b049844dcd3b64db629ec629763565f8896

Observation e4ae07ba-b1e9-4406-b13c-ed8be5e58f04 · inbound

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key cites this paper.

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T03:16:59.195706Z digest=sha256:b1d7962687d3cd275eb36c9b8039069c393bceef6853d14b34b7622a35051210

Observation bd73d658-bfb2-4e24-a3b5-dfeae5a15333 · inbound

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key cites this paper.

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:39:10.411305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T22:36:13.781114Z digest=sha256:de786da1907dc8616e9f07b8fc839cdf761a529d4cc18fb8b7d43da86d18f3e9

Observation 8cf8928e-2eaa-4253-bed9-5eb90f8e9d28 · inbound

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents cites this paper.

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:25:44.578007Z digest=sha256:6a9260a4ef92d919131ba64fb10dce11e8952276574123e20ca0208a96ad4a78

Observation f4d0cb26-0180-4e15-be8c-a0b6b6dbee06 · inbound

KL for a KL: On-Policy Distillation with Control Variate Baseline cites this paper.

KL for a KL: On-Policy Distillation with Control Variate Baseline The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T03:31:07.462474Z digest=sha256:25868896d29ada5cb212c1a388ad2c2b68f249eca11cf49a2c3f78db21bf983f

Observation 9e8e7edf-f295-47ff-9b1f-1ba7ae188728 · inbound

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models cites this paper.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:53774db5e697d3ac9e1b0a67d29af58ded0cabbe21aa033670005f3f3eed2ba8

Observation cccee1d1-89b2-44d5-b2cd-786401296878 · inbound

TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment cites this paper.

TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:56:09.363823Z digest=sha256:8c5443317b7fc8d6cbc8d76c03edcbafc1b0d703d95e989f3bdc4537a93d988c

Observation 2ad4c347-8285-43e9-945c-61ee19bac263 · inbound

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation cites this paper.

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T07:01:46.325498Z digest=sha256:51045b7168be4c4cedd8a83ebff903162e109b1b80e26811153edd7936537b3c

Observation 48bbc7fc-73a8-4fbd-a577-d5bd6be6699a · inbound

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation cites this paper.

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T21:07:43.168136Z digest=sha256:01634b51047383925708ba7e905b78e4446a067f2c6baf825723c45b6965aaf5

Observation 336770af-d9bc-4cc6-9709-f44452ac097b · inbound

Learning, Fast and Slow: Towards LLMs That Adapt Continually cites this paper.

Learning, Fast and Slow: Towards LLMs That Adapt Continually The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:00:31.452781Z digest=sha256:21762b500afeb5962fc1036de96b99de4f601e22bada0c888288e967618f3294

Observation 8ce70be2-cb62-4b06-ae07-e233b13cdf09 · inbound

Learning, Fast and Slow: Towards LLMs That Adapt Continually cites this paper.

Learning, Fast and Slow: Towards LLMs That Adapt Continually The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T05:19:05.368681Z digest=sha256:2524115801736dc4ae6470534013db1910730cd9e7a6da743f4eda39bee1efe8

Observation e0d6d125-b30f-4223-8bd9-f922b29aa54d · inbound

Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design cites this paper.

Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T18:58:53.857181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T18:58:11.197587Z digest=sha256:c17e108655614874d20c07029f7aebd297412be40910b169bff73575dcbce7c1

Observation e7e3fe6c-d9f9-44af-8dfd-0893782dd9f5 · inbound

Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing cites this paper.

Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:06:17.099267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T15:40:55.535914Z digest=sha256:63022b1199a2c73cd24f2f24ab96a5329d6a69a67398019b561cf3f0ce90cb35

Observation d40eddb5-23b1-42d0-9853-690827aeed2a · inbound

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training cites this paper.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:46:28.994858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:8f320a582495b32d8acdd9096558450478d78252ba5eab3150e97334be9a35a3

Observation fbc048f1-6ec6-4991-813c-a93f9d112039 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 88

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:47:59.843785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:00cb08a6fe09ad1b68b2724a0dc61e9417ad9cf48af5611b4739b8a40790d6d8

Observation 3b985535-95dd-49a0-a977-75a043a7f9a9 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:56.161137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:0bce5d1f09b45a8fd03068475b353dfdde8d930e2ac6e938c0863cf84dbbbffa

Observation 90fe2acd-3c6f-4d6f-a600-5727a08e177a · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 146

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T08:57:48.095122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:ed0be67952f402e377ede61f064a80881c772ef2a49ce7e9e994a7d807425a0f

Observation 1513e3a3-426d-4284-a48d-ea0393642148 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 183

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T18:40:03.282572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:a9a4a75bf6c47ec370ed7731b8e9a757460a09d72680494fb10926d5ae4e4fea

Observation ea59d1be-41d0-4456-b515-19d2e8cd1779 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 183

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T18:15:58.953101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:62719761019ae5d332cb278c7ecac8267315d633598b96a57fb832153413c898

Observation 16474127-7b58-4d61-a363-5702f9e930e4 · inbound

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL cites this paper.

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:57.677335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-03T20:59:57.539909Z digest=sha256:5783b93fc92b57b633e9f6ef4f788782ffc4e98a2c8ac85e68969d3a5f8893b9

Observation b886bcb1-417d-4c26-8913-2c0937be3a9f · inbound

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments cites this paper.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:63a90ef420a4ecbdb90e0de47f33e79182a1c91e5f078bb13c0927e65955df38

Observation af318e1f-883f-4e64-93af-cb0701df1549 · inbound

Mach-Mind-4-Flash Technical Report cites this paper.

Mach-Mind-4-Flash Technical Report The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T03:29:34.486347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:29:34.486347Z digest=sha256:6d58bcc3c8798007ef2739d1d9acff53016701a74bb64d2176239b6b7b37cbd7

Observation 5290bcba-6c1a-4b26-8172-e73dffe05623 · inbound

Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback cites this paper.

Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:22:33.420002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:22:33.420002Z digest=sha256:39cba0140c1dc77010b9c1572a63105147f9ad9e8649ddac5acc49a9e79f1e84

Observation b84ad023-1d0f-4acb-8a1b-f1c422cb1437 · inbound

Understanding Reasoning from Pretraining to Post-Training cites this paper.

Understanding Reasoning from Pretraining to Post-Training The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T21:26:32.726200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:26:32.726200Z digest=sha256:cb2169e1045826e7f45adb803efb7c707c53d8d5f20bd38a417d58f237515cdb

Observation 6e72b1c8-2290-4794-a5a2-a98a76ad0bd4 · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:05.363334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:05.363334Z digest=sha256:848b84a397eb135a8210ea93ac2909977f4f3508d3816c58f15db47813246e87