Pith. sign in

Paper Citation Record · LEDGER

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training

As of 6 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2605.26606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26606 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T19:47:09.817243Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact17
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7a46cd7-7851-4120-bc6c-79ba15bf77a2 · outbound

This paper cites arXiv preprint arXiv:2503.18929 , year=.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training arXiv preprint arXiv:2503.18929 , year=

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.801831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:2ee6b829fb14dc0038a0d305832e79b37af725e3e07cf7f5feb189f88859cf81

Observation bc1ff799-a1ca-44df-85c5-e36db16bccdd · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T19:47:09.817243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:010967ac439a1c3164f55d850dfe69710a20cd00a4a7a7d9a2e50d7e229b08e8

Observation 8bf88d01-7253-4077-8366-29b2d1bb9843 · outbound

This paper cites arXiv preprint arXiv:2510.01135 , year=.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training arXiv preprint arXiv:2510.01135 , year=

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.804948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:fced91720779701906c45b11c1f75e55febba581f13899c9b2452a98f0b48263

Observation ea838eab-8e97-476f-8d80-a2736669e7f2 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.807818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:67db37058a7babb4c519c9aa76ef9227c1f94d68346a5e59757b1095b47dee13

Observation b8f67a05-ae1b-49c7-9cb0-56028ac47fa2 · outbound

This paper cites Vcrl: Variance-based curriculum reinforcement learning for large language models.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Vcrl: Variance-based curriculum reinforcement learning for large language models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.798410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:f3735f2c412dce66023fe83067090bdbf8611fe47b9e7130f7ff8b61d7102c04

Observation d31b275c-48fa-449e-9618-c02a80b45c53 · outbound

This paper cites Tinker, 2025.https://thinkingmachines.ai/tinker/.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Tinker, 2025.https://thinkingmachines.ai/tinker/

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T19:47:09.817243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:fc7ea890f9eb76ce51a4908f8ff47b69a342dc5fefaf54a8daf22b97dffd2eac

Observation 49ee130d-b97a-481a-bbc6-aad4c82d7f60 · outbound

This paper cites Knapsack rl: Unlocking exploration of llms via optimizing budget allocation.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Knapsack rl: Unlocking exploration of llms via optimizing budget allocation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.792753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:03c1553e61cf930abf5ef66fd3fb8431923937c7ab0b33cb7b58dc17f1697bc0

Observation 2a7091f0-4665-48a8-985e-2e89b0a6701c · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.748045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:046fad4980238186c79b91ac521d4cda33a1ed11e82340b4c6f122fa7f1bd5d5

Observation 4c01d6d6-43f7-46fb-98a3-f07d1c7e050e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Understanding R1-Zero-Like Training: A Critical Perspective

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.750872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:d8eb709403723c02fec1e76a0f1dc2fc85fd06e09e8032bd997dc9edbcf801d2

Observation fca16e96-44c4-4d55-97a6-5abccc197354 · outbound

This paper cites Yun Qu, Qi Wang, Yixiu Mao, Vincent Tao Hu, Bj¨orn Ommer, and Xiangyang Ji.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Yun Qu, Qi Wang, Yixiu Mao, Vincent Tao Hu, Bj¨orn Ommer, and Xiangyang Ji

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.754059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:da023bd6aa53d70e3e5ff0ba70b0d5a14e580d6f37d910e99f3f3d9ed63c06ee

Observation 1627ce8b-ed7d-4c44-85cb-af20c9fc0f63 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Proximal Policy Optimization Algorithms

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.773627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:1e0a1fb6e6dd057bbf4085d0a4c4b824d49260f829bb4880865d1ca3af695467

Observation 46bd31e0-233b-4137-a939-e71ad415c689 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.804385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:281053b8c29f77bcc69cb29d814b6281be9bae8191aebffd460ca0128b26834d

Observation cbab1d9f-87b4-4cd9-8e91-b32ecddde4a7 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training HybridFlow: A Flexible and Efficient RLHF Framework

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.810454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:e53c44ceff647decaebd8ce9c2977c5f9fe43235a72a1fcc7ffc7af6c406bdef

Observation a4ec028c-7bf4-477f-9a6d-f5b3532c8d46 · outbound

This paper cites Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.807528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:0d0d3e95dfe9690b43184b20e05e37455908324c308e0ea8f2e5d7bfff5ba0b7

Observation 3ca7aa93-bd09-4601-9abd-9e0159acd6aa · outbound

This paper cites LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.792604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:398b5f364db52c52367b0958371cfd19727bb0bed17f06ea2713af5c2489e5cf

Observation 2daae097-2b59-42e0-a985-0b65c6a84215 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.795093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:740774dc6770f89c3d54c9c67e44616156687b3b10dca5889cbe87463ec799a0

Observation f2b379f6-5468-45d0-a1de-67dbfc86a075 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.795243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:f2833a9e248cb5b0e733363520ea7632673252f14b86762ad954229fc0245b3f

Observation 729a0b42-3aba-4068-b985-aed5c8184b19 · outbound

This paper cites Qwen3 Technical Report.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Qwen3 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.811407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:23684cbfb3af0829adb62c95db3ca63660a4b6b77d156a9961f4878ba13a543f

Observation 36f0d166-0b90-4c03-9c82-f3d3b7ca186f · outbound

This paper cites DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.785979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:1fdb1da89552dc46637a27608719152da5c97333821610278eaa6bfe8ab8b16f

Observation dc0b3142-4e43-48af-a3ee-e1c664c391b6 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.782296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:25f3eb7914bd048bb8bcce58a004a769628f81d96ccd69d723c6d5f4c3ab5259

Observation 37914b51-0026-4c78-83a7-9a0a1fb415f4 · outbound

This paper cites Speed-rl: Faster training of reasoning models via online curriculum learning.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Speed-rl: Faster training of reasoning models via online curriculum learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.801652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:a19c09f6cf2e1a732def358475f0cde090b5e99783d5699d399c4c83141df43e

Observation 942d9897-902c-48e5-9eb4-31e12c4ae2cd · outbound

This paper cites Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.782288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:d61bd8cbbb74ce3de8af45299e13b774577ce22fac0bc7e8b1f7758f947b6d5e

Observation 0db1c8d8-bc8f-4af4-aaee-27cb384a58c2 · outbound

This paper cites Rollouts Trained.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Rollouts Trained

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-06-29T19:47:09.817243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:64007f0e3437ea4cb9f628aca0ba589932539f41c1d4d1662cf25071169e549a

Pith citing papers

No inbound Pith citation observations are available.