Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T01:15:49.658412Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 10 inbound Pith citation observations for arXiv:2508.00222.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T01:15:49.658412Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:34:06.659805Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T22:27:26.448121Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7b22eef4-03c0-4df3-93c4-657d5cecb0a0 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ae68cf3-9cc2-46be-b62b-fa959f8e26cb · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b96672c0-0bed-4193-9ae3-4c8ab85aac4f · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Evaluating Large Language Models Trained on Code
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f0cd3ad-c5a0-4f7f-b891-ee41de0ec916 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c02c08fd-8d9a-4359-bab6-f00a2eeacb8f · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d341d843-23ac-4668-85cc-2a4cd35da58b · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Process Reinforcement through Implicit Rewards
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9588cdb-293f-4ccf-bcee-e4773f677bf4 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cccdf65-b86f-4eaa-8584-d24c5534aadb · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Teaching Large Language Models to Reason with Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a05552eb-1069-4963-ae17-3dd4728d2867 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d23aded9-6e25-426d-8887-42be8eb3e84e · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e437c7e8-593e-4319-abf4-4ed102200f61 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90be6e53-f651-415a-b8a1-e2b32ae7e8e1 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization SuperRL: Reinforcement Learning with Supervision to Boost Language Model Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22272cb6-e500-487e-8f2a-9bfa08c6b452 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization John Wiley & Sons
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fab2b4f4-8975-4504-98b3-c6df7177afda · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Proximal Policy Optimization Algorithms
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3614caa9-715a-41bb-9e26-b7ea0abf370f · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c62b401a-4ddc-4b1c-8099-2e529cba08b3 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization HybridFlow: A Flexible and Efficient RLHF Framework
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25803ded-892f-4eb0-8dfd-08c9acb5dbfa · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e1e3c82-2fa3-4b9b-8617-47f0d56dc2e1 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7132f0ea-f4be-46f5-b269-58b3540dbda3 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Learning to Reason under Off-Policy Guidance
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 022e8f73-c2b8-4372-957f-d47abeb89bf0 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 179d49d9-5143-4989-a9c2-f99e8d816395 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1690ad0a-2a23-4c81-9284-9c8fccd9f666 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 724ed7d1-bae9-44e3-b0da-9ea529315ae7 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization First, we dissect the bias and variance issues inherent to standard Importance Sampling (IS) when using data from a single behavior policy
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3195e19-88f0-4e10-a4bc-89576fb6853d · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Both theχ 2-divergence and the more commonly known KL-divergence (DKL(πθ∥πω)) are measures of dissimilarity between distributions (both are instances of f-divergences)
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7d5442c-d63d-492c-86aa-29264ba58b21 · outbound
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1de1f09c-09e0-4355-983e-fcdd4ee855d9 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5dc6cc1-442a-4442-89f9-ab6ee5d81ee2 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization For our approach, one of the model-generated rollouts is replaced with a correct reasoning trajectory from the training dataset
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44f5a834-1b23-4e05-842f-ab92b39fa250 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 732bfdc5-a802-4494-8539-6bbe1f4e0228 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization During evaluation, we set the sampling temperature to 0.6 and report the average pass@1 score over 5 runs by default
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 609a9226-2fbf-49ab-a86f-b939df71e029 · outbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization The second category consists of four straightforward baselines: 1)SFT, supervised fine-tuning using external reasoning trajectory data
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f544093f-ec86-4b34-af8a-850645821afe · inbound
Beyond the Sampled Token: Preserving Candidate Support in RLVR RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da083899-7760-4574-bf2b-59638c3afc2a · inbound
Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f852c4b-d9b7-4d86-b2d2-d3e785e29c46 · inbound
Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 895f2d4e-3db9-4fed-a88e-6bfe5a51724f · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b2c95e0-46d3-4174-9cd4-f5bdfe9d76c1 · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54793fe2-ed12-4d07-9310-e806fa66d86f · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 988a5c69-d1bf-4da7-a7e1-f9ce65807ce5 · inbound
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 026b34fc-4a97-43da-bb62-a9a4b3d59496 · inbound
Rollout-Level Advantage-Prioritized Experience Replay for GRPO RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ee16055-4ec2-4e07-a709-aaf498fc08b5 · inbound
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 637b06e1-7920-48cd-8901-029268f58fa7 · inbound
Preference Tuning as Spectral Update Reorganization RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.