Pith. sign in

Paper Citation Record · LEDGER

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

As of 20 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 30 inbound Pith citation observations for arXiv:2506.02177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02177 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:33:58.345391Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:19:46.719051Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.849796Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9831c0c-05e7-4a31-a01c-0fbc8f226c97 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.422853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.422853Z digest=sha256:d9e84493315f9afe251ef91bc43b80a656854529bdc8f98d36f4cda494a0464f

Observation 0d8a338a-89e1-47db-ba73-d11448e44cda · outbound

This paper cites Training.Our method is implemented based on verl (Sheng et al.,.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Training.Our method is implemented based on verl (Sheng et al.,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.133936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:57.817375Z digest=sha256:87af756bb352efe075d978f9d4313d42ddfcb53159eaf4b8c5c1be8d5e79be71

Observation 23c8d49a-973f-49ee-95cc-884995a9bb2d · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.037842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.037842Z digest=sha256:a0a89f364d91a59e0ddaa04ef699ca60e1bf32748f7f24a8af7fa11f4e0b3b13

Observation f9c3030d-ccd9-4bcc-bafd-d5082f7d4ebe · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.160410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.160410Z digest=sha256:3fb3083cfc3ed3783547c151337b232d1cc8247f7d7ec0c3a524492f0b289f76

Observation 1513ae04-df18-4551-b3b0-5004701717db · outbound

This paper cites Data-efficient finetuning using cross-task nearest neighbors.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Data-efficient finetuning using cross-task nearest neighbors

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:01.140721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:54.446718Z digest=sha256:3b740a9d88e10f68de4c706200f6b64ce9defb2e07bf1bdbd25971dfc6396224

Observation 9b82e5ce-5627-4432-acbb-de0cc0225403 · outbound

This paper cites Large-Scale Data Selection for Instruction Tuning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Large-Scale Data Selection for Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.560575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.560575Z digest=sha256:711c12975e42ab6b5e2f399967517ca5ea244048ac5597c959aba3aaa1ac8504

Observation aedd8de2-cdae-4679-b01c-fc73610148f5 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.702600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.702600Z digest=sha256:bf98eccb54cacfe2d59fc0f60109c4ba224b8f353f7bfe954cdd6fcb63b1bb0d

Observation 04d656e2-c62d-4612-92f6-7b1e770b01bd · outbound

This paper cites Let's Verify Step by Step.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Let's Verify Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.000411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.000411Z digest=sha256:1702633dfad0254a53e96a546e6d5526ca1284227c60604761f2a64a0d9e7b5c

Observation c30c30e7-a2a5-40f5-aa80-9eade33d7799 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Understanding R1-Zero-Like Training: A Critical Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.145743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.145743Z digest=sha256:4143fbbd8d774e2d4db7171c8b80b16b744d39bfaa58e1099e55d7cd937de0f9

Observation 0f96e0ce-7823-441d-bc85-09fc21a62170 · outbound

This paper cites Enabling Weak LLMs to Judge Response Reliability via Meta Ranking.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Enabling Weak LLMs to Judge Response Reliability via Meta Ranking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.280113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.280113Z digest=sha256:1ec03357508badf6da6d59053ba16f1e4c1de6fdff5bfb894a03d111fe0e05ee

Observation 281fdffb-7daf-4088-a06d-b603cad1e1f9 · outbound

This paper cites s1: Simple test-time scaling.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts s1: Simple test-time scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.416178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.416178Z digest=sha256:f6d53776027d1fa8837ffb6e5d81cddeb53674c215101a598b1f5d6a1ce42479

Observation 6d8283ee-dc53-42bc-bf0b-7e40a6c8fb88 · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.521985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.521985Z digest=sha256:8615b91aaf2f2dccc22422881ac0d9d930546e9a7b30de20380a0314021ba2f1

Observation 8325e58a-b228-4603-8f90-fa9d1df8c754 · outbound

This paper cites OpenAI o1 System Card.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts OpenAI o1 System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.643161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.643161Z digest=sha256:02edb60e145197b7444a5e81786bf28d43886046bcb658bb720aef0d3c73745f

Observation 58fd25b9-d2f3-4476-98de-d4e90c16c08c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.813166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.813166Z digest=sha256:f5a25dc912d8e34f1c2e516f30dbafb8b05ce2814df12b8b16ddbabe0ee19a9c

Observation 10a3f3e9-9bb3-4dc5-aaa8-8ef6b83ed13e · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts HybridFlow: A Flexible and Efficient RLHF Framework

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.929448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.929448Z digest=sha256:c5d9bf6f43b2c804823f6b608651914ddad06049b6f6e7eb282bcf152107b8b5

Observation 8bb4b8c2-2e6a-4fec-8a11-2d49ab3cd961 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.042626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.042626Z digest=sha256:fc6c31d8f797f89f637f91a8e123d42ed16762251956842482d0e34f6e8e21e0

Observation b7f103ad-20e3-4ede-b2fb-bba1a62b64ea · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.370913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.370913Z digest=sha256:eea82a03360aba92d590b303e224770ea7c290faa0ec02481cbf0bccf1f44553

Observation f678aa86-9893-4cd6-8d94-f0914b3a7036 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.534466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.534466Z digest=sha256:cf720fcc3cfcaadf712aeaa6d476c7f3e3fb4c756e8e634ca4e889e0b5c43762

Observation 87883627-e99d-41f5-b282-cc9822319cae · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.659506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.659506Z digest=sha256:f4dbb670cfd85e0dd453943be7948d16e85f15a2c10c913cb4b9491c99cc0c40

Observation 40fb4ab4-d45f-4e03-adc9-c5a196ab3651 · outbound

This paper cites LIMO: Less is More for Reasoning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMO: Less is More for Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.783335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.783335Z digest=sha256:bfd6c94aa31d4ffd44c1a0f0116ec8bfbc52b172819eb5346c3ed46c380460d1

Observation 209e468b-62bb-4f5a-b75b-ea1eb0702bd0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.907899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.907899Z digest=sha256:d925729ff69591c89fcf483582e1f692fc240031c69cf1d1ca38f48832e2d4fb

Observation 1871f2f7-c535-4ef7-9dcc-a2ad22c02ed6 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.031251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.031251Z digest=sha256:a6defa0706ee0c50c6bbf9aabdf0f16d9c423c9897486530f7d7d82f60e52eb7

Observation 0550c118-947f-459e-bfbe-4a6c265c01e0 · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.154091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.154091Z digest=sha256:9df55abace969f6be9cf165b7d67042a654f73facde43dfaff07e9f39ba726ce

Observation af96af09-a3de-4845-a58d-785ada5276b4 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.305024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.305024Z digest=sha256:46dc9e6482cb2de025ee94a1cae65975391d243040a7ed0cadb75c161082b927

Observation 39da3963-895b-4fa0-9746-a8b686108307 · outbound

This paper cites Coverage-centric coreset selection for high pruning rates.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Coverage-centric coreset selection for high pruning rates

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.880737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:57.447663Z digest=sha256:5fda7104db0ea18df21b31b1d491aa7ecce7f22f02ba750aba690aaaded3bf98

Observation b4b7d512-6da2-4e00-a08b-ae0adce931fb · outbound

This paper cites LLNL-affiliated authors were supported under Contract DE-AC52-07NA27344 and supported by the LLNL-LDRD Program under Project Nos.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LLNL-affiliated authors were supported under Contract DE-AC52-07NA27344 and supported by the LLNL-LDRD Program under Project Nos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.655878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:57.550465Z digest=sha256:5fc5262f857a915036bd85b22d678f70835f020168080309aaded824dbfd9118

Observation 9d4a53b0-6123-428b-ad5b-49517ad0fe3a · outbound

This paper cites We find that training on DAPO alone can degrade performance on LaTeX-based benchmarks, so we augment it with MATH to preserve formatting diversity and improve generalization.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We find that training on DAPO alone can degrade performance on LaTeX-based benchmarks, so we augment it with MATH to preserve formatting diversity and improve generalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.370119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:57.673181Z digest=sha256:4c8fa7c77c8550cc943f39babadbf929c39db96fdb443e095947823e2c04a6b9

Observation 97e561c9-9c20-4a84-abea-f46fe94f1050 · outbound

This paper cites We use 4xH100 for Qwen2.5-Math-1.5B training and 8xH100 for Qwen2.5-Math-7B and DeepSeek-R1-Distill-Qwen-1.5B.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We use 4xH100 for Qwen2.5-Math-1.5B training and 8xH100 for Qwen2.5-Math-7B and DeepSeek-R1-Distill-Qwen-1.5B

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.873088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:57.964485Z digest=sha256:7138e246c178de8f96398f95f70802e59a21b0be9219ee95d901aec305cc1220

Observation 7eafab92-670f-4b36-bd34-28395fb1569c · outbound

This paper cites an unresolved cited work.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:33:59.357792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:58.212664Z digest=sha256:b7fb629cc32698271f3e5336f2561e50bd39e505319746ed78f5e1811ad3adc4

Observation 64209bf2-bcdd-4765-acb2-35b1ea20ce8e · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.510451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.510451Z digest=sha256:f9b274fedfe50d648c341743fad78c05d02fe57c51874ae6bc8e040e4ccc7a8c

Observation 8efec914-804d-416f-bbd1-df6fdc0ad840 · outbound

This paper cites We use β1 = 0.9, β2 = 0.999, and apply a weight decay of 0.01.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We use β1 = 0.9, β2 = 0.999, and apply a weight decay of 0.01

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.602692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:58.080472Z digest=sha256:2e8e4b6b6059893082bc30585c61ed0469504052c50484d45c26a259a553bb06

Observation 2a05e5e0-9c29-4444-89dd-426e927ce2b7 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.204875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.204875Z digest=sha256:8b98f85f7849491ff0dd64172fe6cfc07e82fff15e2f8efd02bf1fe3c8fbd17b

Observation 078a963e-511f-47f7-90a5-90629883b038 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.297876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.297876Z digest=sha256:0370ddfd417644fe7e7216a69a4534c85a5674a9230d0ef858a9a1d4f889f782

Observation 5c790190-4f4c-4b4f-a4a2-43d7f28d5c46 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMR: Less is More for RL Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.876617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.876617Z digest=sha256:36e7b0edeac82ad1956f523af71a0e16053e431aee3a6b9d98456efffce0e12c

Observation 5182579f-4527-4fea-820f-fafc04ade777 · outbound

This paper cites Active Preference Optimization for Sample Efficient RLHF.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Active Preference Optimization for Sample Efficient RLHF

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.628034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.628034Z digest=sha256:4045f989e515d267c506b20add4b9efa533ec6f808bbbf5ced98b6f36860fb1d

Observation de08b487-e30d-48b1-90f0-2e748fca46e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.786117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.786117Z digest=sha256:ef5788740a241e3c8529eda9d2d3d0ff753ece55215446d9865d7824d07f9208

Observation d5eb21f6-e440-4230-a022-5f7f5ebe07d1 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.https://github.com/huggingface/open-r1.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Open r1: A fully open reproduction of deepseek-r1, January 2025.https://github.com/huggingface/open-r1

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.916913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.916913Z digest=sha256:91364397dec0fbe223d8db7ef7b44f9ff94655cfc948723d03eb1ef31b95abd3

Observation 9d935f1b-5438-49e7-a02b-8e5183d840bb · outbound

This paper cites We evaluate all models with temperature = 1 and repeat the test set 4 times for evaluation stability, i.e.,pass@1(avg@4), for all benchmarks.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We evaluate all models with temperature = 1 and repeat the test set 4 times for evaluation stability, i.e.,pass@1(avg@4), for all benchmarks

Reference 8196

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.131928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:58.345391Z digest=sha256:777cb660ffce619370f4366161cef3b5f1afc300c8320981cfb381294d357d3f

Pith citing papers

Observation d084dc3c-910a-4cbc-9535-8d8786694e82 · inbound

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance cites this paper.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.719051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.719051Z digest=sha256:ac8d04a12993dc2fab7dc7489e3391bdeeac884c896cf30f2d46634529e52ace

Observation b4d40336-1471-4269-b24f-77da40690d0f · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.638499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c3863cb206fbd427b88f926b740ad2ce6ba4064575dbdeae3eaedc9d23dea3f9

Observation db8baea6-a880-4384-a29e-9fd1060a96ae · inbound

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO cites this paper.

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:02.195880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:15:04.225059Z digest=sha256:48e8d147ce0903704b122d5fbb1cb5d98854cd0136ffbd3703869093ae2375af

Observation 332989d2-bc58-4d9e-b58b-55c974b8e2be · inbound

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models cites this paper.

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:06:52.840498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T07:05:45.081955Z digest=sha256:45bcd8b39ec592c3453f18854325026014a020aadbbb53f5c43f90d215b38b95

Observation 3644689d-62f7-4cea-8a5c-5ec10e9c720c · inbound

Cost-Aware Learning cites this paper.

Cost-Aware Learning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:36:30.855968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T05:11:01.131590Z digest=sha256:33a7190c43809511ec9817e7fc27fd9fd53792d0af3d0ca7c10ea27c15006b39

Observation 920e3085-d9e2-44c6-a90f-4243e10ddbe6 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 184

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.463309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:7d417f0223c5db9a558dfb5498c062cd35ed6fdb4de4f4558bb2ad7ff00131d8

Observation cb3b60b6-8cee-4536-9fa1-7a060ba2fb04 · inbound

Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL cites this paper.

Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:09.059414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T15:31:23.161009Z digest=sha256:fe6f9d3aabc3c3607be505699e318660bfaac2a1b768088e6ece3ac9ef86ff40

Observation 1bd91141-483c-4119-97cc-f75940b57616 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:15.988767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:f206fa9868dfae9ad472d3c6a8ce7a6f0a4a214dca9966305117cbc6cd0e3e36

Observation 99e4c427-1350-4aaa-bef3-db35fa1bc34b · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.200705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:7f038951df58083076709e385f357ceb974d0f0cfeb0b0794e1621aad56c6961

Observation e8dfdb1f-37e0-4527-abcc-0816e3b30fcd · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.733368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:5bc743b3275bd26308d767f7f9b0eeebf70450da60fb930c28ed6b42e9dc3296

Observation 56fe42d0-5d57-4e03-8d3e-2fa1432befd8 · inbound

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States cites this paper.

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:54.882249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T02:08:28.113371Z digest=sha256:fa1b11b64caccecaf4217f094950a36cfe6bab77b8c12e09d20bd8f8bc10a72b

Observation 1bb2620f-65b2-4965-a974-ffd22c58c6b1 · inbound

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States cites this paper.

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:45.835310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:01:25.989133Z digest=sha256:b74c135ab3b15a0012f4f2c112a0dd2df7cceb7de3a15927fb22303560edce18

Observation f86cb607-eed8-4ff2-a252-14fb94b01f17 · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 43

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T07:46:28.647091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:f28b897c7abfcc5678a80066d284113495ed0580bfd24574e956987b3ed969f9

Observation 12169cac-4b43-40bd-ae4d-62099273376f · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.327965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:f3c448e300824fbcf4298ad085a10746d71d8f34b04549a2f1cc9daf778a5ed3

Observation 117727e2-7d6b-4893-9a70-1b6ea6b10b7a · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:22:06.806520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:93d0de9b6b9242dcb745f18fce9b4cca53a10dfa6b841ca2d034064b7fcbf4af

Observation 39a0d34e-3e63-4bc3-b2c7-ecb5096dfb5c · inbound

AIS: Adaptive Importance Sampling for Quantized RL cites this paper.

AIS: Adaptive Importance Sampling for Quantized RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:14:52.595524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T03:13:14.384567Z digest=sha256:6f357b474ab9ccb9f8a5558b314fa747934f49313ed82bdab3cc4ad017c779cd

Observation 2bd94e0d-39a0-4977-a38e-91ea45ade716 · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:39:00.348750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:a628cd417b84fa959dcdb006cda11457479d7c4b1a5fb69113d6b8ee207aab10

Observation 72467af6-8d03-4f29-8f8f-1424b2767578 · inbound

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs cites this paper.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:43.788110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:2bac0a444659f68f97f03fc46eae6bc77c2ef440235079d9b87a9ddf82056a59

Observation 942d9897-902c-48e5-9eb4-31e12c4ae2cd · inbound

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training cites this paper.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.782288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:8a59b9828d0c8b868a486ff54b1e8c8044e52d92f8354935456df3f9c7a18a2a

Observation 24f670bc-ea5d-4b7e-a474-bd4cdf87ed2c · inbound

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training cites this paper.

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:43:15.361248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T08:38:52.411671Z digest=sha256:5c0def5cc575a1b175b315a129dcc96258b5e0eadae2b052bdc2c8bb7c081203

Observation ba73a442-9db2-4e22-ab0a-1a24cbdae46a · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 70

Resolution
malformed identifier
arxiv_id, observed 2026-07-01T21:16:13.345991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:45efb86d5feec626f346402334567b78f2471cf7eab9729091871f5099c2a8bd

Observation 4c242cd4-ec7c-43f0-901d-3edb7e28c52e · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:56.403920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:9eb083e0cad4e5baa348f638f0c81a448b22bd285c5315ba62efac7df96b7098

Observation 9fe1ca1e-f844-438f-8466-60f0a59ecddc · inbound

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO cites this paper.

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:08.917602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T22:54:21.740892Z digest=sha256:71e6ccb48784a9b18c904037b428267eba6f224e84d66ba113d422fac453e65b

Observation 342d495b-9320-40ba-adb9-58a2db117d97 · inbound

CATPO: Critique-Augmented Tree Policy Optimization cites this paper.

CATPO: Critique-Augmented Tree Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:47:28.013303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T19:30:07.971424Z digest=sha256:0fabed50e173ba37f89fb8785357497eb038e7ae2b28e22081bffbf6b5f22e71

Observation 125b38d1-05be-4ade-bf2d-03ff4f6c5ced · inbound

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization cites this paper.

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:37.386022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T14:02:05.833651Z digest=sha256:b90bb4ba1d8f8215d81a7bcab97d7f503875932f18ba9042a6c2b8cb36e9a540

Observation 1e5be0ca-5cbb-4df1-9a9e-4cb4ed65cca2 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.851250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:b802b5c76bdeaed4cc63a2a9ade9bea2059e06e3f2befbb68280434e08284095

Observation 3ea09dec-71d9-4b2d-9af3-1544d52c4ef9 · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.908981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.908981Z digest=sha256:9d5fb5e38928c0e238af050d44add086d48fc95d4c6399f38fd6c3eccceefc87

Observation 90c868fa-b460-46cc-b1cf-3190f1bd4b4e · inbound

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning cites this paper.

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-31T01:31:13.934392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:31:13.934392Z digest=sha256:2db500a6728465ac52539feab466d4c4025333e4f12c8a492d01619988785427

Observation 79f0a914-cc9b-46b6-bf4d-f8132541aa86 · inbound

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance cites this paper.

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T00:22:11.829438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:22:11.829438Z digest=sha256:a5e3c1ce54d6c6002e61f661ffa684ff234f9e8842b748b600699e2920fe0b08

Observation 2ac7fa79-d18f-4e95-83b3-d54f63cf1327 · inbound

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR cites this paper.

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:22:55.603592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:22:55.603592Z digest=sha256:3e1c4e23dff465303804484347a5ebea6347a10c6e9fa1e9ef6e08062c1ec29f