Pith. sign in

Paper Citation Record · LEDGER

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2605.07114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07114 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:09:27.543625Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact17
  • verified fuzzy4
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1be6d6c-2aaf-413f-a5da-4baa2a475be2 · outbound

This paper cites Step guided reasoning: Improving mathematical reasoning using guidance generation and step reasoning.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Step guided reasoning: Improving mathematical reasoning using guidance generation and step reasoning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:06:55.750928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:5f39d043cd5c594479f9f080461e32280ff62854602d482504f4d0e46e92c04b

Observation 41fad8c2-676d-4432-a421-b9e43f16a3cb · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Evaluating Large Language Models Trained on Code

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:40:59.548091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:29d806377be3bf92814104a929e4b47d3a61eaaf218482adb61fad7e5dabee43

Observation af22db31-4f35-4b14-a9e4-1b318f3c9147 · outbound

This paper cites Differential smooth- ing mitigates sharpening and improves llm reasoning.arXiv preprint arXiv:2511.19942.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Differential smooth- ing mitigates sharpening and improves llm reasoning.arXiv preprint arXiv:2511.19942

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:40:59.556925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:ba31ee0d4ba0e454bee784fe0d66c745373b1174799d086dc2efd7f70dd04237

Observation 83043b97-c0d9-48c3-ab21-c59f1780777c · outbound

This paper cites Rewarding the unlikely: Lifting grpo beyond distribution sharpening.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Rewarding the unlikely: Lifting grpo beyond distribution sharpening

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:06:55.741460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:00647e835baa9e722887f783c9545db7090ef7b75a05eee0277c153e0299fad2

Observation 933a320e-f56f-4b02-b16d-e2c3b08f44bb · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:40:59.539550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:a811c603db064ad6a2f7d58c90e1921b63c9894fb8e8fc509ebf66f0dc411cf4

Observation 706c466b-8b7c-4da7-9c75-f47773ce47a5 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Understanding R1-Zero-Like Training: A Critical Perspective

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:40:59.574754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:5e70cb6631a9325eb3c25c25984b502695595498adcf5374705beef4c39fd6e7

Observation d25db324-9220-4eec-85a0-b096ff28838f · outbound

This paper cites Adaptive rollout allocation for online reinforcement learning with verifiable rewards.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Adaptive rollout allocation for online reinforcement learning with verifiable rewards

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:40:59.592552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:772dfaa5ee83672e1b45a18c5df4963a681aa029728a7c7ef62b8d934e852ea1

Observation 9742e471-1cd7-4ea7-8607-6d6c3920cad0 · outbound

This paper cites HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-01T02:03:29.943232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:c65c3d80bc57ea883e062194bc70762a1f0cbd54ed01c7b10151188467132516

Observation 60aa04d6-5746-4ab5-88d9-46e5d1f1cab0 · outbound

This paper cites https://arxiv.org/abs/2410.03131.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR https://arxiv.org/abs/2410.03131

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:40:59.568271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:a6660edb328d599a453263935ea793c5b6e70721cd7eeb1537b89ebb333d098b

Observation b394ff8a-c1c7-4159-bef6-1ebce99cdd98 · outbound

This paper cites Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:33.173797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:108b90e17788120cc6c018bec50a3505e6fce921a0f83487a0d20ea0fb1db44d

Observation 336fd8d0-5125-42c1-b43a-7399cb6f9eb8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:40:59.524414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:ecbfdc16bbc1ce1ee6f09e68d257fbb906bbcb0e793d64045b68f89136532a59

Observation 68092739-f75c-4792-8f60-b7d3fd897957 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:12:09.004283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:4faca461a5b48fdf3a96885005e323e3c20f906dfe1cb94dcc0482603a8a005d

Observation 3e8e6f52-3640-4926-bd8d-e9c90a41bfa5 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:8dcc3dae64b2440f16cb035040e295dc8a11e8a13dbacd4fb7999a8e383ae440

Observation f2752ca1-08e0-4c46-950b-80e8ee3d8af5 · outbound

This paper cites The invisible leash: Why rlvr may or may not escape its origin.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR The invisible leash: Why rlvr may or may not escape its origin

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:40:59.426365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:a6e5da7805bf7e4ec3694bd1fe6dacca78db345b1d3472aa872db5926f86a901

Observation 53a922e5-6fa6-435c-8a4e-6deed89a33f8 · outbound

This paper cites Qwen2.5 Technical Report.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Qwen2.5 Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:40:59.515986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:a1f097b3727be233b0fad6fc12621376a856543011e14e67d25b73bbc1df9e83

Observation eea7058f-6ac4-4a5c-be4a-ec370f40a36a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:40:59.532268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:4652f672bd50aba2dfa101a825def7245cbda8e1f7a10ad6018efa9c6ea1b3d3

Observation cb832be7-f953-45a4-b813-6f382984b891 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:40:59.500399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:0d6a8ca35996cfb9ee6dd540760237eddd9053415e693023948878818b9b16c4

Observation cf135d09-7837-4246-b4b7-4bd9ba93b27a · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:21:05.971096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:ee065bd566626f77f2da7056a85b4ffc410bee7fd26573db6d27294c06b86e06

Observation b9d6e023-75e0-4971-8103-52dfbd9ceff7 · outbound

This paper cites Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:40:59.493685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:74af99c8caa5e4ed1ec52cbddd6a6d4048c41b58dda09fd6887c8e952fe2d650

Observation d49cfb48-64b8-4180-9a8e-4e09621d3627 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:40:59.442026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:102cd8a0fcb95c7acb5d7a76e64f13031bcfa7e88c12b433a3ad6ea4fafa0e94

Observation 372bcb32-3471-4e8e-8a13-27986db56ee2 · outbound

This paper cites All runs use the TRL implementation of GRPO with vLLM rollout in colocate mode.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR All runs use the TRL implementation of GRPO with vLLM rollout in colocate mode

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:06:55.733072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:5eb313347a69ea0f6b5629da036753a43722ed0be8f5686ef2272671f1931b9d

Observation 6c2bf8eb-68a4-4030-8ded-903593f7f9dd · outbound

This paper cites an unresolved cited work.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-16T00:06:55.754100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:4a9bc288145883dd08d5a0bdeb2d2054f7b1902f22db448a42604506ac51ee2b

Observation 184449ce-eb5a-4cff-826a-0c58d63715e4 · outbound

This paper cites an unresolved cited work.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-16T00:06:55.747722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:c2b3796bc4dcfd772e01e17b3c9bb498e173a36c3445264f2cecea541f160d68

Observation 5aad2fe9-7091-4b50-9c86-2303e012a8d6 · outbound

This paper cites an unresolved cited work.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-16T00:06:55.744782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:0b35e6941cb5a275165c5e9560a7abeb530e31bd0836e8ccfdc21d5b91e41c58

Observation c0cb0aea-e525-4d72-b7c3-a37e5017adba · outbound

This paper cites The fixedBeta(1, 1)prior used by HORA (solid blue) is compared against the five learned-prior variants described above.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR The fixedBeta(1, 1)prior used by HORA (solid blue) is compared against the five learned-prior variants described above

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:06:55.738172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:f55bc870b003c96487e036335b567c0e937c03bc843e455fcc768479957bca2a

Pith citing papers

No inbound Pith citation observations are available.