Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T16:22:46.913172Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 7 inbound Pith citation observations for arXiv:2605.05040.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T16:22:46.913172Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T13:40:41.757803Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T16:28:38.434907Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4474d615-5a35-4844-b49f-9ae84066fe9b · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 630c912b-ad0c-47d1-9028-18a1be8b467e · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Jiamu Bai, Xin Yu, Meilong Xu, Weitao Lu, Xin Pan, Kiwan Maeng, Daniel Kifer, Jian Wang, and Yu Wang
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8faf494b-48c4-4a43-bd64-d82676d7a9d2 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c0428810-736e-46d0-9b3c-106fd2d52cab · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Hdpo: Hybrid distillation policy optimization via privileged self-distillation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8d5f5526-0bc8-43d3-be39-ac30142f40cc · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d9515901-543c-4e25-9510-56747c002831 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization OpenThoughts: Data Recipes for Reasoning Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee7fa4a1-f0ab-488a-9d7f-3c473b5bd133 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization PFedDST: Personalized Federated Learning with Decentralized Selection Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 567fd72d-d9e7-49f2-a8a3-6d495536ef3d · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Reinforcement Learning via Self-Distillation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9d5a2af6-6dcb-4107-91c3-39c25dac3069 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation df82730e-31b0-4bb9-8542-16c0a6a29f35 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Unifying group-relative and self-distillation policy optimization via sample routing
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d9ae9a9-a5d8-4a79-bb35-f2cc6941ae39 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization On-policy distillation.Thinking Machines Lab: Con- nectionism
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 01af755b-8ef8-4f4e-b71c-8b107f072748 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization and Ravikumar, Pradeep and Wainwright, Martin J
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c8adf3e0-6aa8-4421-ba5c-f644bfdc0e22 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 495db4f8-69f6-4cf1-ac3a-3ae3a45effac · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 35a7df0e-6aea-4477-9302-d7e7eda67e84 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 39d5b8f9-faab-425f-989a-5b24af4e0c76 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Self-Distillation Enables Continual Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7e0ba19f-1c61-46b8-bfe5-33b9a18640ab · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization A Survey of On-Policy Distillation for Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 74db5023-b36d-4051-905b-025d7df7ca94 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8c0199b6-0411-417c-994c-25e0da9c78da · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Self-Distilled RLVR
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8057ce1c-bdcd-408b-8aff-517f74b16fba · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f6e75022-a335-4d15-a081-e89bed8931fd · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Embarrassingly Simple Self-Distillation Improves Code Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 90e88614-6bed-449f-8b19-c2a86e878f3a · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec45e4a4-4f48-4097-955a-686854b0ab13 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae0a0656-b9cc-4b25-b2b6-9541c4759d6a · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Appendix E develops the technical details behind the statistical analysis in the main text
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 547ffe04-b74c-4323-8039-c3461eec9345 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e6476b21-681f-466c-8ac3-196244b87de2 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization [2026a] whenever applicable so that the comparison against prior baselines isolates the effect of the proposed PBSD objective as cleanly as possible
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 517b923f-e1e3-4f87-8ba7-675bb2a4e767 · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Tool-use data.For the additional tool-use study, we follow the setup in Shenfeld et al
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c11ee6f9-3076-42b1-bd70-682e600ffc2c · outbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Base (Student)
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dc822383-633f-4f61-a170-3acb28b804b1 · inbound
A Brief Overview: On-Policy Self-Distillation In Large Language Models Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6159c1c3-988f-4682-a15d-d069dcea0db3 · inbound
A Brief Overview: On-Policy Self-Distillation In Large Language Models Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e63f3d9a-5097-4b76-87f0-b60c9afd40c1 · inbound
RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c1e27acb-7add-46b5-b717-b3c636aedb14 · inbound
DemoPSD: Disagreement-Modulated Policy Self-Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 685b8c5c-e0e9-4cd6-a1bb-a24693448c5e · inbound
DemoPSD: Disagreement-Modulated Policy Self-Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9efda571-6255-4f55-ae5f-ead04887d3fa · inbound
DemoPSD: Disagreement-Modulated Policy Self-Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed59648d-76be-4a61-9bd2-e75318123df3 · inbound
Contrastive On-Policy Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.