Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T01:09:27.543625Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2605.07114.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T01:09:27.543625Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f1be6d6c-2aaf-413f-a5da-4baa2a475be2 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Step guided reasoning: Improving mathematical reasoning using guidance generation and step reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 41fad8c2-676d-4432-a421-b9e43f16a3cb · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Evaluating Large Language Models Trained on Code
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af22db31-4f35-4b14-a9e4-1b318f3c9147 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Differential smooth- ing mitigates sharpening and improves llm reasoning.arXiv preprint arXiv:2511.19942
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 83043b97-c0d9-48c3-ab21-c59f1780777c · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Rewarding the unlikely: Lifting grpo beyond distribution sharpening
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 933a320e-f56f-4b02-b16d-e2c3b08f44bb · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Measuring Mathematical Problem Solving With the MATH Dataset
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 706c466b-8b7c-4da7-9c75-f47773ce47a5 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Understanding R1-Zero-Like Training: A Critical Perspective
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d25db324-9220-4eec-85a0-b096ff28838f · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Adaptive rollout allocation for online reinforcement learning with verifiable rewards
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9742e471-1cd7-4ea7-8607-6d6c3920cad0 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60aa04d6-5746-4ab5-88d9-46e5d1f1cab0 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR https://arxiv.org/abs/2410.03131
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b394ff8a-c1c7-4159-bef6-1ebce99cdd98 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 336fd8d0-5125-42c1-b43a-7399cb6f9eb8 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 68092739-f75c-4792-8f60-b7d3fd897957 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e8e6f52-3640-4926-bd8d-e9c90a41bfa5 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f2752ca1-08e0-4c46-950b-80e8ee3d8af5 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR The invisible leash: Why rlvr may or may not escape its origin
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53a922e5-6fa6-435c-8a4e-6deed89a33f8 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Qwen2.5 Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eea7058f-6ac4-4a5c-be4a-ec370f40a36a · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb832be7-f953-45a4-b813-6f382984b891 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf135d09-7837-4246-b4b7-4bd9ba93b27a · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9d6e023-75e0-4971-8103-52dfbd9ceff7 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d49cfb48-64b8-4180-9a8e-4e09621d3627 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 372bcb32-3471-4e8e-8a13-27986db56ee2 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR All runs use the TRL implementation of GRPO with vLLM rollout in colocate mode
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c2bf8eb-68a4-4030-8ded-903593f7f9dd · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 184449ce-eb5a-4cff-826a-0c58d63715e4 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5aad2fe9-7091-4b50-9c86-2303e012a8d6 · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c0cb0aea-e525-4d72-b7c3-a37e5017adba · outbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR The fixedBeta(1, 1)prior used by HORA (solid blue) is compared against the five learned-prior variants described above
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.