Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2505.13026.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:42:32.507410Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T22:27:26.462219Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 43b5554e-0bf3-4687-b7f5-aed6cf3c50d0 · inbound
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ae68cf3-9cc2-46be-b62b-fa959f8e26cb · inbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 994317d1-da8b-437a-8e2f-ff8730d29f96 · inbound
EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d64f13b-c276-4196-aa2a-14bdc9bec298 · inbound
CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96c74d77-4f08-43bf-b901-c92e189d6bbc · inbound
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b9a0335-f9cf-4929-8a2d-e56052ccac1c · inbound
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1134a14-495b-415f-8224-ae2b5b3a2183 · inbound
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation caf1d41f-2ff2-403c-ae32-e47e3c03e4d7 · inbound
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0028337c-5869-44ff-a2b4-d846267ce88b · inbound
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78f0c023-f494-4c08-a0cd-acbfa6d5ac05 · inbound
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81be6abc-94e2-4b28-8412-c180b2b289e9 · inbound
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.