Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T19:47:09.817243Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2605.26606.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T19:47:09.817243Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c7a46cd7-7851-4120-bc6c-79ba15bf77a2 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training arXiv preprint arXiv:2503.18929 , year=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc1ff799-a1ca-44df-85c5-e36db16bccdd · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bf88d01-7253-4077-8366-29b2d1bb9843 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training arXiv preprint arXiv:2510.01135 , year=
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea838eab-8e97-476f-8d80-a2736669e7f2 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8f67a05-ae1b-49c7-9cb0-56028ac47fa2 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Vcrl: Variance-based curriculum reinforcement learning for large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d31b275c-48fa-449e-9618-c02a80b45c53 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Tinker, 2025.https://thinkingmachines.ai/tinker/
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ee130d-b97a-481a-bbc6-aad4c82d7f60 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Knapsack rl: Unlocking exploration of llms via optimizing budget allocation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2a7091f0-4665-48a8-985e-2e89b0a6701c · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4c01d6d6-43f7-46fb-98a3-f07d1c7e050e · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Understanding R1-Zero-Like Training: A Critical Perspective
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fca16e96-44c4-4d55-97a6-5abccc197354 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Yun Qu, Qi Wang, Yixiu Mao, Vincent Tao Hu, Bj¨orn Ommer, and Xiangyang Ji
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1627ce8b-ed7d-4c44-85cb-af20c9fc0f63 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Proximal Policy Optimization Algorithms
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46bd31e0-233b-4137-a939-e71ad415c689 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cbab1d9f-87b4-4cd9-8e91-b32ecddde4a7 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training HybridFlow: A Flexible and Efficient RLHF Framework
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a4ec028c-7bf4-477f-9a6d-f5b3532c8d46 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ca7aa93-bd09-4601-9abd-9e0159acd6aa · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2daae097-2b59-42e0-a985-0b65c6a84215 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2b379f6-5468-45d0-a1de-67dbfc86a075 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 729a0b42-3aba-4068-b985-aed5c8184b19 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Qwen3 Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 36f0d166-0b90-4c03-9c82-f3d3b7ca186f · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dc0b3142-4e43-48af-a3ee-e1c664c391b6 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 37914b51-0026-4c78-83a7-9a0a1fb415f4 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Speed-rl: Faster training of reasoning models via online curriculum learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 942d9897-902c-48e5-9eb4-31e12c4ae2cd · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0db1c8d8-bc8f-4af4-aaee-27cb384a58c2 · outbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Rollouts Trained
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.