Pith. sign in

Paper Citation Record · LEDGER

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 28 inbound Pith citation observations for arXiv:2506.02177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02177 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:33:58.345391Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:22:11.829438Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.849796Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9831c0c-05e7-4a31-a01c-0fbc8f226c97 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.422853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.422853Z digest=sha256:7d512f3f1c79b8c12882627e5e1c8cf4bc9e2b7a78eeae62bc9bcb5cfe064565

Observation 0d8a338a-89e1-47db-ba73-d11448e44cda · outbound

This paper cites Training.Our method is implemented based on verl (Sheng et al.,.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Training.Our method is implemented based on verl (Sheng et al.,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.133936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:57.817375Z digest=sha256:3ef474ce4a8e874cc87779232650f99fc8c02b6b63ca2849176f8a1a55830781

Observation 23c8d49a-973f-49ee-95cc-884995a9bb2d · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.037842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.037842Z digest=sha256:77c902591dfc02a9b320586b2f36b495ae1085fce873d19f37b24cb7d778f271

Observation f9c3030d-ccd9-4bcc-bafd-d5082f7d4ebe · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.160410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.160410Z digest=sha256:ddd54e7fce1b91117e6e8c0716bb493986d9a306fd415e60fbec09eb3087c9d9

Observation 1513ae04-df18-4551-b3b0-5004701717db · outbound

This paper cites Data-efficient finetuning using cross-task nearest neighbors.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Data-efficient finetuning using cross-task nearest neighbors

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:01.140721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:54.446718Z digest=sha256:ee4ebbd884e960fd065db0ef2cf53b91b38cdbb367fea9b66f6cae1ef46915a8

Observation 9b82e5ce-5627-4432-acbb-de0cc0225403 · outbound

This paper cites Large-Scale Data Selection for Instruction Tuning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Large-Scale Data Selection for Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.560575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.560575Z digest=sha256:311f441d8933b64b20ba702e72e5e6f80ecf4d2186e219cf9823554380c63b2c

Observation aedd8de2-cdae-4679-b01c-fc73610148f5 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.702600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.702600Z digest=sha256:77da900f2c9445334200432f80999b6380e5fdc035d25b8080fc4cf2a3a0196e

Observation 04d656e2-c62d-4612-92f6-7b1e770b01bd · outbound

This paper cites Let's Verify Step by Step.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Let's Verify Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.000411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.000411Z digest=sha256:c918822adca15ee15f706e33cf40b24df0e26acfe963005a7a32c0950456345a

Observation c30c30e7-a2a5-40f5-aa80-9eade33d7799 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Understanding R1-Zero-Like Training: A Critical Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.145743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.145743Z digest=sha256:cdc276752886302d44d4d031a0dc89030ba375841b6562434b6510ce21b014b8

Observation 0f96e0ce-7823-441d-bc85-09fc21a62170 · outbound

This paper cites Enabling Weak LLMs to Judge Response Reliability via Meta Ranking.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Enabling Weak LLMs to Judge Response Reliability via Meta Ranking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.280113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.280113Z digest=sha256:e800830bc43abf14d830ab3bb489509c25e4a92806501290ab22bbcba0c566df

Observation 281fdffb-7daf-4088-a06d-b603cad1e1f9 · outbound

This paper cites s1: Simple test-time scaling.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts s1: Simple test-time scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.416178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.416178Z digest=sha256:5b86949ba9f04ade132b9feeb66933bd3f07ee6837ebbbefb16ae3c2d8ad25cc

Observation 6d8283ee-dc53-42bc-bf0b-7e40a6c8fb88 · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.521985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.521985Z digest=sha256:92dd080afb1bbf2686033be4c9dede9ab9facb30ec8c5b2b4b15d42afccfa51c

Observation 8325e58a-b228-4603-8f90-fa9d1df8c754 · outbound

This paper cites OpenAI o1 System Card.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts OpenAI o1 System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.643161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.643161Z digest=sha256:dfc2151e97a3042b60c712149390e267f932a8b181c51652b3fb596c80c2eabe

Observation 58fd25b9-d2f3-4476-98de-d4e90c16c08c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.813166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.813166Z digest=sha256:dc15ab3e46e8138f59c2bbe4f83d5f3b6704018bfe14a3c9ff0a6eee00f8732f

Observation 10a3f3e9-9bb3-4dc5-aaa8-8ef6b83ed13e · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts HybridFlow: A Flexible and Efficient RLHF Framework

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.929448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.929448Z digest=sha256:1eef6b86c58a70400a9b786deff154efae4bf7a9e201442e7e01ca889ecd9a65

Observation 8bb4b8c2-2e6a-4fec-8a11-2d49ab3cd961 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.042626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.042626Z digest=sha256:59ce64b3e49c1a3806ec16ff109c5fb18ea94d7977b1ce008d9b1746a3a5b56e

Observation b7f103ad-20e3-4ede-b2fb-bba1a62b64ea · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.370913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.370913Z digest=sha256:be0fbc8875704db7150c09a14205223a51ae13d42395dd8f3fe74388f9d57d94

Observation f678aa86-9893-4cd6-8d94-f0914b3a7036 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.534466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.534466Z digest=sha256:dd4900b471b5e54e76291551f21b843b9dbfd44a3bcfe0cf588bddf10824b842

Observation 87883627-e99d-41f5-b282-cc9822319cae · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.659506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.659506Z digest=sha256:51d07b268b28e1f8900c18eff2223625fb47786a97a88f401bd5d186fda849c8

Observation 40fb4ab4-d45f-4e03-adc9-c5a196ab3651 · outbound

This paper cites LIMO: Less is More for Reasoning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMO: Less is More for Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.783335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.783335Z digest=sha256:09231d9604ec9a373952b207361fe4ffaa3aa14e8105fd1bcd5a474e186c5b1f

Observation 209e468b-62bb-4f5a-b75b-ea1eb0702bd0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.907899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.907899Z digest=sha256:8528b1e168094daf5d79f2c3f4027f3f8cb7c773dc49ebe5a889485948afe890

Observation 1871f2f7-c535-4ef7-9dcc-a2ad22c02ed6 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.031251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.031251Z digest=sha256:34d089110f3b9546cedaf2195ec2d66b0a3299905469ac800e24a980edb56cac

Observation 0550c118-947f-459e-bfbe-4a6c265c01e0 · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.154091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.154091Z digest=sha256:28b203b69a81865c4bfdbd4048f0a456557beb35afcda8a942896c0e301c32bd

Observation af96af09-a3de-4845-a58d-785ada5276b4 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.305024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.305024Z digest=sha256:4c7b5f60b27e62b5ab8a3f0b3648290130a4695438b1d648ee0cf0f209498af9

Observation 39da3963-895b-4fa0-9746-a8b686108307 · outbound

This paper cites Coverage-centric coreset selection for high pruning rates.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Coverage-centric coreset selection for high pruning rates

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.880737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:57.447663Z digest=sha256:2040a54fd2d835457a3fa95f35dfb49fa7e8a432959abecde1eb2f42cf6c836d

Observation b4b7d512-6da2-4e00-a08b-ae0adce931fb · outbound

This paper cites LLNL-affiliated authors were supported under Contract DE-AC52-07NA27344 and supported by the LLNL-LDRD Program under Project Nos.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LLNL-affiliated authors were supported under Contract DE-AC52-07NA27344 and supported by the LLNL-LDRD Program under Project Nos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.655878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:57.550465Z digest=sha256:bcdd3716d9ca58397860d5a24fbb8906ef159dad80c22e3700bc36a6dc0a54e1

Observation 9d4a53b0-6123-428b-ad5b-49517ad0fe3a · outbound

This paper cites We find that training on DAPO alone can degrade performance on LaTeX-based benchmarks, so we augment it with MATH to preserve formatting diversity and improve generalization.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We find that training on DAPO alone can degrade performance on LaTeX-based benchmarks, so we augment it with MATH to preserve formatting diversity and improve generalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.370119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:57.673181Z digest=sha256:8d6a130b661a2efea32774e035d71f8f376bce4442993fffc50ee10e4c4be7e3

Observation 97e561c9-9c20-4a84-abea-f46fe94f1050 · outbound

This paper cites We use 4xH100 for Qwen2.5-Math-1.5B training and 8xH100 for Qwen2.5-Math-7B and DeepSeek-R1-Distill-Qwen-1.5B.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We use 4xH100 for Qwen2.5-Math-1.5B training and 8xH100 for Qwen2.5-Math-7B and DeepSeek-R1-Distill-Qwen-1.5B

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.873088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:57.964485Z digest=sha256:d036885d9e3b39a91c9bccacdd2b4eb2a1e0cc5f605bd0fce58b26eacaee12ac

Observation 7eafab92-670f-4b36-bd34-28395fb1569c · outbound

This paper cites an unresolved cited work.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:33:59.357792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:58.212664Z digest=sha256:6183a4ac1970d8330c3a9e367600be5bff0d32448fccf968c62c84c907ea3176

Observation 64209bf2-bcdd-4765-acb2-35b1ea20ce8e · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.510451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.510451Z digest=sha256:fc477bd4df7b7f5b82ec3df043c936635b1a9909b750333c84d59967ca3fd9ae

Observation 8efec914-804d-416f-bbd1-df6fdc0ad840 · outbound

This paper cites We use β1 = 0.9, β2 = 0.999, and apply a weight decay of 0.01.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We use β1 = 0.9, β2 = 0.999, and apply a weight decay of 0.01

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.602692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:58.080472Z digest=sha256:4ea7d6c22ac3cd1a51a1be868c496d3efa782724c8c2e057c0904496b250552e

Observation 2a05e5e0-9c29-4444-89dd-426e927ce2b7 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.204875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.204875Z digest=sha256:55af67df1452b6557f936d9b4e6265c0d592c2a169f7a31ae42b8d1e3cd22302

Observation 078a963e-511f-47f7-90a5-90629883b038 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.297876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.297876Z digest=sha256:3bceb888366eb4b6f74faf778be6339c44fb80bc3e5d6f11c55b2e1cf2793689

Observation 5c790190-4f4c-4b4f-a4a2-43d7f28d5c46 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMR: Less is More for RL Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.876617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.876617Z digest=sha256:8aaf668abd5efe25a7511b21710a073a5c4e14c0326b2e17fe614e6a24bdbbcc

Observation 5182579f-4527-4fea-820f-fafc04ade777 · outbound

This paper cites Active Preference Optimization for Sample Efficient RLHF.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Active Preference Optimization for Sample Efficient RLHF

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.628034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.628034Z digest=sha256:212e32230ae235426617593c0d246884646ca8136a04b0f86e82f10186e5a2f4

Observation de08b487-e30d-48b1-90f0-2e748fca46e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.786117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.786117Z digest=sha256:07dacfe1980d567b1e80bf3a376a968de1c7dc2b21db704d2e503cadf320357d

Observation d5eb21f6-e440-4230-a022-5f7f5ebe07d1 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.https://github.com/huggingface/open-r1.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Open r1: A fully open reproduction of deepseek-r1, January 2025.https://github.com/huggingface/open-r1

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.916913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.916913Z digest=sha256:d4afb38d4ec496c3e2b1069a3f76a658e29358fd5044b67b23185d35feb25848

Observation 9d935f1b-5438-49e7-a02b-8e5183d840bb · outbound

This paper cites We evaluate all models with temperature = 1 and repeat the test set 4 times for evaluation stability, i.e.,pass@1(avg@4), for all benchmarks.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We evaluate all models with temperature = 1 and repeat the test set 4 times for evaluation stability, i.e.,pass@1(avg@4), for all benchmarks

Reference 8196

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.131928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:58.345391Z digest=sha256:fa960531f98edd5d36f5c75c55dc5e2e20392d3f33e6a44877c511e0d9c60bde

Pith citing papers

Observation b4d40336-1471-4269-b24f-77da40690d0f · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.638499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:5e3b4b1cae79170634699232c9d6b52aaa3faaead28ea4f7082f1770ff3f35b0

Observation db8baea6-a880-4384-a29e-9fd1060a96ae · inbound

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO cites this paper.

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:02.195880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:15:04.225059Z digest=sha256:bf9b781caeefe0586a7914aa1959a41e8d9db5677647ae39a2c90b0331eb38ff

Observation 332989d2-bc58-4d9e-b58b-55c974b8e2be · inbound

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models cites this paper.

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:06:52.840498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T07:05:45.081955Z digest=sha256:d8c759ab377967e73a95159bb9161f8ddbe7cd070d5091f4b370b6db26d0c196

Observation 3644689d-62f7-4cea-8a5c-5ec10e9c720c · inbound

Cost-Aware Learning cites this paper.

Cost-Aware Learning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:36:30.855968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T05:11:01.131590Z digest=sha256:6a0190262d005a202eaba34b619ef69e55566c4edf063bc66f1cf93b3cc2ad91

Observation 920e3085-d9e2-44c6-a90f-4243e10ddbe6 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 184

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.463309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:a0de8903d3a033dc074bb29c2bb56b134b8e9d5ec3ed9451ca11c11f0cff612b

Observation cb3b60b6-8cee-4536-9fa1-7a060ba2fb04 · inbound

Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL cites this paper.

Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:09.059414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T15:31:23.161009Z digest=sha256:d6b3db365efea8619ab79dbe85c936a31186c94ffaaf8619beda34cc0dc6af3b

Observation 1bd91141-483c-4119-97cc-f75940b57616 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:15.988767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:543243e4a509ae53b88737c532857e4163e29e4cfa84b7ba94af05386c80206b

Observation 99e4c427-1350-4aaa-bef3-db35fa1bc34b · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.200705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:9832e79b247311ba08cfba155e615032a1581e28930a2cb59f85bae3b22978f0

Observation e8dfdb1f-37e0-4527-abcc-0816e3b30fcd · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.733368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:c9f9b8d2b454c25e8e2bd6cfb77218327b28ff01db2ecd03da26df2e8803502d

Observation 56fe42d0-5d57-4e03-8d3e-2fa1432befd8 · inbound

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States cites this paper.

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:54.882249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:08:28.113371Z digest=sha256:865519a54f5729d31fd74e1d65af66ee9da29176c82568c831d979b466cfb993

Observation 1bb2620f-65b2-4965-a974-ffd22c58c6b1 · inbound

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States cites this paper.

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:45.835310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:01:25.989133Z digest=sha256:025b0b111900342f1e17524e1d3c0cc7f580e76bdf3e67c4ae9222fd2ab44d04

Observation f86cb607-eed8-4ff2-a252-14fb94b01f17 · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 43

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T07:46:28.647091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:b4c6a1b341ebc936d385d4582e54b8053e6f060eea63cf1aeb7178031ecb72e7

Observation 12169cac-4b43-40bd-ae4d-62099273376f · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.327965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:37260d40a9dfd92908c826aee1fdf564a9f67ee70cf2956a0915af89dc94e02b

Observation 117727e2-7d6b-4893-9a70-1b6ea6b10b7a · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:22:06.806520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:754ee420032100221d34043926b7e594ed57d9e46a1a791738bcc6e4c5b485b2

Observation 39a0d34e-3e63-4bc3-b2c7-ecb5096dfb5c · inbound

AIS: Adaptive Importance Sampling for Quantized RL cites this paper.

AIS: Adaptive Importance Sampling for Quantized RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:14:52.595524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T03:13:14.384567Z digest=sha256:298a5fffa32d760f5e4e1ee48c5b8d8fc6cece1f569cf2a9004d1129f02482a1

Observation 2bd94e0d-39a0-4977-a38e-91ea45ade716 · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:39:00.348750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:34891668a2f6538c2324acff4cb76f901a914cd47667e35452be40ea3d565a42

Observation 72467af6-8d03-4f29-8f8f-1424b2767578 · inbound

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs cites this paper.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:43.788110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:c4ca407e17b035f0a7d29866960ec3dc0833369f3e4c044cf0b2b78ad6107f42

Observation 942d9897-902c-48e5-9eb4-31e12c4ae2cd · inbound

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training cites this paper.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.782288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:fabed9ff7d8582938ba1e553bdf0004e3b480a5bdce53c21d91dd61401f8d24a

Observation 24f670bc-ea5d-4b7e-a474-bd4cdf87ed2c · inbound

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training cites this paper.

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:43:15.361248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T08:38:52.411671Z digest=sha256:4c6fa0ee418a49fc3f7c4786632db2d71e75041f48c51c1bbcd6775a9341f285

Observation ba73a442-9db2-4e22-ab0a-1a24cbdae46a · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 70

Resolution
malformed identifier
arxiv_id, observed 2026-07-01T21:16:13.345991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:74edea5a754c83e22efd948b0b5926761d42b657c36e0aff554bb1071b9eb3b6

Observation 4c242cd4-ec7c-43f0-901d-3edb7e28c52e · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:56.403920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:dcf162c1fbec8a21209bc8d229931e194302de8e21516a9b5c26f48759e7462a

Observation 9fe1ca1e-f844-438f-8466-60f0a59ecddc · inbound

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO cites this paper.

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:08.917602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:54:21.740892Z digest=sha256:dce4155ebba8ec3bf54e09e9150b08648fe76b2f3e6b7f426f6aab590ed91684

Observation 342d495b-9320-40ba-adb9-58a2db117d97 · inbound

CATPO: Critique-Augmented Tree Policy Optimization cites this paper.

CATPO: Critique-Augmented Tree Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:47:28.013303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T19:30:07.971424Z digest=sha256:09a92b4afc2da480136417a9329c4b2bfe85a29dc61990f88a27186f533e6ad9

Observation 125b38d1-05be-4ade-bf2d-03ff4f6c5ced · inbound

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization cites this paper.

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:37.386022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T14:02:05.833651Z digest=sha256:39c383d92df52a3793b92bccce041298b57208db2dff75098fe0affce2c13054

Observation 1e5be0ca-5cbb-4df1-9a9e-4cb4ed65cca2 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.851250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:210112addbe98b8c2e7cf5bdd41f3f9e39ce0ee91f1fd15a47d38be14c5044dc

Observation 3ea09dec-71d9-4b2d-9af3-1544d52c4ef9 · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.908981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.908981Z digest=sha256:6754b10967363fb80519f1f95ecd5df295e7edcd7842509d61f044769eb00890

Observation 90c868fa-b460-46cc-b1cf-3190f1bd4b4e · inbound

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning cites this paper.

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-31T01:31:13.934392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:31:13.934392Z digest=sha256:da4ad5893601fbe25cae14ee6b9700e26a9230b67958a9d262e8413bee9f79a6

Observation 79f0a914-cc9b-46b6-bf4d-f8132541aa86 · inbound

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance cites this paper.

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T00:22:11.829438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:22:11.829438Z digest=sha256:1becf9dc74ac2f9bce0a6d8390a06ad3b6bf82d4174b05af8d9e1215d8356c90