Pith. sign in

Paper Citation Record · LEDGER

KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2506.02208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02208 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:15:50.891014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:33:45.190652Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8d152d08-91fd-4ac5-b84e-58206f25b39b · inbound

Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance cites this paper.

Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T19:02:43.873939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:02:43.873939Z digest=sha256:bd17a20600255b45890f381cadcf9571458a8b7fad1c245049bc20f1d62f02f3

Observation 5971ec2e-4c87-4006-8d9b-3bd1882168b5 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 198

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:44.633196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:44.633196Z digest=sha256:4fcf2fcd4d5d0fd71ff569fd94ccfb562f4bd37787cca4a75c2d579a3413cc9e

Observation f1b486ff-92f8-4c8a-8e3f-8cc7583ff877 · inbound

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation cites this paper.

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T07:57:26.211367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:57:26.211367Z digest=sha256:0aec64df807209d1a6852b8416e156c399d053adca9b7b17f9a559c4fd491393

Observation 20561600-97a2-4b49-a18c-94857ce0b58f · inbound

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents cites this paper.

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:31:00.859941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T15:28:07.981488Z digest=sha256:fda89eeff694fe06fb1aa1b79a62372d273be4bf89907362144b5ffa3d7df461

Observation 79670b7c-b18d-4eb9-968e-3dbbe5ac9c10 · inbound

SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting cites this paper.

SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T22:24:51.874340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:24:51.874340Z digest=sha256:3d84ceaf572bd3bd3599827a8291a6b76b9347d6dd7b3bbfb9984ab897a0f44e

Observation 3cb03e82-9d72-416d-8c1f-d69e288d04b6 · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:11:20.617234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:3d06b8ac659783b864dd82d237290ae7ad9ea33f6eb96ecdd0d06c9774137394

Observation bf4c15fe-1558-42a2-9fac-0b2e4c0513e2 · inbound

Structured Role-Aware Policy Optimization for Multimodal Reasoning cites this paper.

Structured Role-Aware Policy Optimization for Multimodal Reasoning KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:58.237485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:24:20.286120Z digest=sha256:793de6d699c6f4f274f89135054aa28d5ab39c72983d85222dea6617f06a60ca

Observation e791c93a-73ea-4480-8d79-6d9ac748c9e9 · inbound

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair cites this paper.

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:59.234734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:21:43.166304Z digest=sha256:0c34e1efda40bff805766cdb3a286f1ea7e7fd9bf83445c1d1debc9775f59ee7

Observation 2f34ad15-5089-42fe-93f8-744e0e90f97d · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.269454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:afcf5c52f31be0afe3e79ce2b6c5b5dc69940d3d7e8a2101875b003661be6579

Observation 63c379f9-c1f8-456d-86c1-137068c3cfcc · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:07:42.300161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:5aaa0eed5d257bf493be8249a09c5c4193337a5c2e2f1ddfd9074b9414e3b50b

Observation 71e57212-c1da-43f2-8692-a5d3608b8faf · inbound

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization cites this paper.

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:26.466360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T00:59:44.364491Z digest=sha256:698273a222a774084712a2d7f72b967abb062f0227d5b1520371fe4635c679f2

Observation d23ae8c3-3675-4eba-bd0e-96a9b678e0d9 · inbound

Multi-Rollout On-Policy Distillation via Peer Successes and Failures cites this paper.

Multi-Rollout On-Policy Distillation via Peer Successes and Failures KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.839839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T21:29:36.803832Z digest=sha256:f6bb9b65b2ccac5f419e0735534234e190ac236e8a1ad4d69944e45503bfed35

Observation 9246a7d8-c7b7-4c3b-bca6-6d190305d287 · inbound

Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence cites this paper.

Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:12:54.742130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:12:32.638453Z digest=sha256:8e7670355aacba40fefc815af7995f1109fe5e372b04a0a647931ff286a3a669

Observation b3a484d2-b832-4b9e-baa0-f76a8cebdf82 · inbound

One-Way Policy Optimization for Self-Evolving LLMs cites this paper.

One-Way Policy Optimization for Self-Evolving LLMs KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:01:15.509689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:f64b964300a6edac69bb83ed384dba428d92d879c0cf8d4c71ec0364225dc8f1

Observation 94c30858-349e-4749-b2b8-ac3df692453f · inbound

Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models cites this paper.

Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:26:14.307856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:04:19.035841Z digest=sha256:6ef3a67fe104433d7fe4bc4507e73d35aa7b1b4ebd9a6be3ceebe3f338c16fbe

Observation e3988326-fb10-4f49-996b-d6c8a77ad7b1 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.061982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:a6b5257896ca5410749f41292a9b2f1eac2a537719de41755011a4ec3e26cc78

Observation b01bd5d6-69f5-4f3b-b464-8a3f70d817e0 · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:33:45.192381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-07T12:31:42.224094Z digest=sha256:1e1f37076a3884cf56d4abb212f49e60bef9733fc86bbe847ed3026043b3cca7

Observation 64d41ed5-9b4c-4adb-a93e-dcf3a9891769 · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:5dc9bf05acd2a7666a317f91167690f53ec2b87fbbdc769603f83ef9c3c41126

Observation defd7e1c-7a86-4920-a21b-04d34c45bcf9 · inbound

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation cites this paper.

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T11:42:23.970188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:42:23.970188Z digest=sha256:9d62a46f37203d311eb3fa52d0bded3d2c62e702f2645227726800e691a9fee7

Observation 9f2baf39-57d6-467d-a5f9-05b0373855a4 · inbound

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance cites this paper.

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T00:22:11.804686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:22:11.804686Z digest=sha256:fbb05321ae1fc9e20ba9bbaeb7288a8d146c1a53f8123504f1c556ea0be15a74

Observation f8857167-3609-4823-94d3-4c04bdb123df · inbound

Agentic Reinforcement Learning with Self-Distilled Reward Shaping cites this paper.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.891014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.891014Z digest=sha256:00867221b3afd68cd354c3e2be1dc3934f6e58cad676b9919f4fcfa0e39b768d