Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:33:58.345391Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 28 inbound Pith citation observations for arXiv:2506.02177.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:33:58.345391Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T00:22:11.829438Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T09:47:59.849796Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d9831c0c-05e7-4a31-a01c-0fbc8f226c97 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d8a338a-89e1-47db-ba73-d11448e44cda · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Training.Our method is implemented based on verl (Sheng et al.,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 23c8d49a-973f-49ee-95cc-884995a9bb2d · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts On Designing Effective RL Reward at Training Time for LLM Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9c3030d-ccd9-4bcc-bafd-d5082f7d4ebe · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1513ae04-df18-4551-b3b0-5004701717db · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Data-efficient finetuning using cross-task nearest neighbors
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b82e5ce-5627-4432-acbb-de0cc0225403 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Large-Scale Data Selection for Instruction Tuning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aedd8de2-cdae-4679-b01c-fc73610148f5 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d656e2-c62d-4612-92f6-7b1e770b01bd · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Let's Verify Step by Step
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c30c30e7-a2a5-40f5-aa80-9eade33d7799 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Understanding R1-Zero-Like Training: A Critical Perspective
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f96e0ce-7823-441d-bc85-09fc21a62170 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 281fdffb-7daf-4088-a06d-b603cad1e1f9 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts s1: Simple test-time scaling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d8283ee-dc53-42bc-bf0b-7e40a6c8fb88 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8325e58a-b228-4603-8f90-fa9d1df8c754 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts OpenAI o1 System Card
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fd25b9-d2f3-4476-98de-d4e90c16c08c · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10a3f3e9-9bb3-4dc5-aaa8-8ef6b83ed13e · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts HybridFlow: A Flexible and Efficient RLHF Framework
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bb4b8c2-2e6a-4fec-8a11-2d49ab3cd961 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f103ad-20e3-4ede-b2fb-bba1a62b64ea · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f678aa86-9893-4cd6-8d94-f0914b3a7036 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87883627-e99d-41f5-b282-cc9822319cae · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40fb4ab4-d45f-4e03-adc9-c5a196ab3651 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMO: Less is More for Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 209e468b-62bb-4f5a-b75b-ea1eb0702bd0 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1871f2f7-c535-4ef7-9dcc-a2ad22c02ed6 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0550c118-947f-459e-bfbe-4a6c265c01e0 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af96af09-a3de-4845-a58d-785ada5276b4 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39da3963-895b-4fa0-9746-a8b686108307 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Coverage-centric coreset selection for high pruning rates
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b4b7d512-6da2-4e00-a08b-ae0adce931fb · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LLNL-affiliated authors were supported under Contract DE-AC52-07NA27344 and supported by the LLNL-LDRD Program under Project Nos
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d4a53b0-6123-428b-ad5b-49517ad0fe3a · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We find that training on DAPO alone can degrade performance on LaTeX-based benchmarks, so we augment it with MATH to preserve formatting diversity and improve generalization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97e561c9-9c20-4a84-abea-f46fe94f1050 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We use 4xH100 for Qwen2.5-Math-1.5B training and 8xH100 for Qwen2.5-Math-7B and DeepSeek-R1-Distill-Qwen-1.5B
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7eafab92-670f-4b36-bd34-28395fb1569c · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64209bf2-bcdd-4765-acb2-35b1ea20ce8e · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8efec914-804d-416f-bbd1-df6fdc0ad840 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We use β1 = 0.9, β2 = 0.999, and apply a weight decay of 0.01
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a05e5e0-9c29-4444-89dd-426e927ce2b7 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 078a963e-511f-47f7-90a5-90629883b038 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c790190-4f4c-4b4f-a4a2-43d7f28d5c46 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMR: Less is More for RL Scaling
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5182579f-4527-4fea-820f-fafc04ade777 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Active Preference Optimization for Sample Efficient RLHF
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de08b487-e30d-48b1-90f0-2e748fca46e8 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5eb21f6-e440-4230-a022-5f7f5ebe07d1 · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Open r1: A fully open reproduction of deepseek-r1, January 2025.https://github.com/huggingface/open-r1
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d935f1b-5438-49e7-a02b-8e5183d840bb · outbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We evaluate all models with temperature = 1 and repeat the test set 4 times for evaluation stability, i.e.,pass@1(avg@4), for all benchmarks
Reference 8196
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b4d40336-1471-4269-b24f-77da40690d0f · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db8baea6-a880-4384-a29e-9fd1060a96ae · inbound
LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 332989d2-bc58-4d9e-b58b-55c974b8e2be · inbound
MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3644689d-62f7-4cea-8a5c-5ec10e9c720c · inbound
Cost-Aware Learning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 920e3085-d9e2-44c6-a90f-4243e10ddbe6 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 184
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb3b60b6-8cee-4536-9fa1-7a060ba2fb04 · inbound
Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1bd91141-483c-4119-97cc-f75940b57616 · inbound
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99e4c427-1350-4aaa-bef3-db35fa1bc34b · inbound
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e8dfdb1f-37e0-4527-abcc-0816e3b30fcd · inbound
Gradient Extrapolation-Based Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56fe42d0-5d57-4e03-8d3e-2fa1432befd8 · inbound
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1bb2620f-65b2-4965-a974-ffd22c58c6b1 · inbound
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f86cb607-eed8-4ff2-a252-14fb94b01f17 · inbound
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12169cac-4b43-40bd-ae4d-62099273376f · inbound
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 117727e2-7d6b-4893-9a70-1b6ea6b10b7a · inbound
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39a0d34e-3e63-4bc3-b2c7-ecb5096dfb5c · inbound
AIS: Adaptive Importance Sampling for Quantized RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2bd94e0d-39a0-4977-a38e-91ea45ade716 · inbound
Learning from Language Feedback via Variational Policy Distillation Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72467af6-8d03-4f29-8f8f-1424b2767578 · inbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 942d9897-902c-48e5-9eb4-31e12c4ae2cd · inbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 24f670bc-ea5d-4b7e-a474-bd4cdf87ed2c · inbound
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba73a442-9db2-4e22-ab0a-1a24cbdae46a · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c242cd4-ec7c-43f0-901d-3edb7e28c52e · inbound
On Advantage Estimates for Max@K Policy Gradients Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9fe1ca1e-f844-438f-8466-60f0a59ecddc · inbound
AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 342d495b-9320-40ba-adb9-58a2db117d97 · inbound
CATPO: Critique-Augmented Tree Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 125b38d1-05be-4ade-bf2d-03ff4f6c5ced · inbound
N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e5be0ca-5cbb-4df1-9a9e-4cb4ed65cca2 · inbound
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ea09dec-71d9-4b2d-9af3-1544d52c4ef9 · inbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c868fa-b460-46cc-b1cf-3190f1bd4b4e · inbound
Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f0a914-cc9b-46b6-bf4d-f8132541aa86 · inbound
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.