Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2506.02208.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:15:50.891014Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-07T12:33:45.190652Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 8d152d08-91fd-4ac5-b84e-58206f25b39b · inbound
Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5971ec2e-4c87-4006-8d9b-3bd1882168b5 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 198
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b486ff-92f8-4c8a-8e3f-8cc7583ff877 · inbound
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20561600-97a2-4b49-a18c-94857ce0b58f · inbound
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79670b7c-b18d-4eb9-968e-3dbbe5ac9c10 · inbound
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cb03e82-9d72-416d-8c1f-d69e288d04b6 · inbound
Characterizing Model-Native Skills KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf4c15fe-1558-42a2-9fac-0b2e4c0513e2 · inbound
Structured Role-Aware Policy Optimization for Multimodal Reasoning KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e791c93a-73ea-4480-8d79-6d9ac748c9e9 · inbound
Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f34ad15-5089-42fe-93f8-744e0e90f97d · inbound
AIPO: Learning to Reason from Active Interaction KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63c379f9-c1f8-456d-86c1-137068c3cfcc · inbound
AIPO: Learning to Reason from Active Interaction KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71e57212-c1da-43f2-8692-a5d3608b8faf · inbound
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d23ae8c3-3675-4eba-bd0e-96a9b678e0d9 · inbound
Multi-Rollout On-Policy Distillation via Peer Successes and Failures KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9246a7d8-c7b7-4c3b-bca6-6d190305d287 · inbound
Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3a484d2-b832-4b9e-baa0-f76a8cebdf82 · inbound
One-Way Policy Optimization for Self-Evolving LLMs KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94c30858-349e-4749-b2b8-ac3df692453f · inbound
Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3988326-fb10-4f49-996b-d6c8a77ad7b1 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b01bd5d6-69f5-4f3b-b464-8a3f70d817e0 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64d41ed5-9b4c-4adb-a93e-dcf3a9891769 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation defd7e1c-7a86-4920-a21b-04d34c45bcf9 · inbound
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f2baf39-57d6-467d-a5f9-05b0373855a4 · inbound
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8857167-3609-4823-94d3-4c04bdb123df · inbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.