Pith. sign in

Paper Citation Record · LEDGER

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

As of 4 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 7 inbound Pith citation observations for arXiv:2605.05040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05040 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T16:22:46.913172Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:40:41.757803Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T16:28:38.434907Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact20
  • verified fuzzy4
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4474d615-5a35-4844-b49f-9ae84066fe9b · outbound

This paper cites On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:16:10.878876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:94930bbc90c09d0c0968888dd7454627edc75fc61485ff8a306fad2cd302309b

Observation 630c912b-ad0c-47d1-9028-18a1be8b467e · outbound

This paper cites Jiamu Bai, Xin Yu, Meilong Xu, Weitao Lu, Xin Pan, Kiwan Maeng, Daniel Kifer, Jian Wang, and Yu Wang.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Jiamu Bai, Xin Yu, Meilong Xu, Weitao Lu, Xin Pan, Kiwan Maeng, Daniel Kifer, Jian Wang, and Yu Wang

Reference 2

Resolution
verified exact
doi, observed 2026-05-08T20:24:09.051021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:07ac01126a1b7a388f365875db9b77591417302eeeab8219e83c4587b8ef886d

Observation 8faf494b-48c4-4a43-bd64-d82676d7a9d2 · outbound

This paper cites OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:43:12.126898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:759a5d28dca8b97da91e544549f471f60b9bd3d001b790996a80c9944968b60a

Observation c0428810-736e-46d0-9b3c-106fd2d52cab · outbound

This paper cites Hdpo: Hybrid distillation policy optimization via privileged self-distillation.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Hdpo: Hybrid distillation policy optimization via privileged self-distillation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:10.854456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:5e68c3cc8ab986466819e916204360afcfec72da5cf56ca693f9de59789459df

Observation 8d5f5526-0bc8-43d3-be39-ac30142f40cc · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:16:10.948921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:2a6eb427e8bd15fd65c7a029778703bb8a32f1bac06156fa6bdb0f5ca01cd939

Observation d9515901-543c-4e25-9510-56747c002831 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization OpenThoughts: Data Recipes for Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:57:51.597669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:04ae30b746c3fa3aaa3886cb98794eb334b019dea6f15a32548ade1a6db269b0

Observation ee7fa4a1-f0ab-488a-9d7f-3c473b5bd133 · outbound

This paper cites PFedDST: Personalized Federated Learning with Decentralized Selection Training.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization PFedDST: Personalized Federated Learning with Decentralized Selection Training

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:16:10.973582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:c8247a08beabad48e241c7a7e2574dcdd54bb118b52453fd4a9821bc7cb26915

Observation 567fd72d-d9e7-49f2-a8a3-6d495536ef3d · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Reinforcement Learning via Self-Distillation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:29:18.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:fb7baffd82089d2decdde601172c251a45419cb15b5c737cac2a9d59632da911

Observation 9d5a2af6-6dcb-4107-91c3-39c25dac3069 · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:16:10.921197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:10627c70f071ad1f576eee98fb34a9ebc2b19c6a629d706f64dda454f234f6a5

Observation df82730e-31b0-4bb9-8542-16c0a6a29f35 · outbound

This paper cites Unifying group-relative and self-distillation policy optimization via sample routing.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Unifying group-relative and self-distillation policy optimization via sample routing

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:10.828081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:2f8526e0438bd3d8b5fa99405addb1b0b364d4709c611708c9873c9a626779ef

Observation 2d9ae9a9-a5d8-4a79-bb35-f2cc6941ae39 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Con- nectionism.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization On-policy distillation.Thinking Machines Lab: Con- nectionism

Reference 11

Resolution
verified exact
doi, observed 2026-05-08T20:24:09.044640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:d1db253b509269db7af653160e6b6b41be1556dee227e4e6e41710e0f2200844

Observation 01af755b-8ef8-4f4e-b71c-8b107f072748 · outbound

This paper cites and Ravikumar, Pradeep and Wainwright, Martin J.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization and Ravikumar, Pradeep and Wainwright, Martin J

Reference 12

Resolution
metadata mismatch
doi, observed 2026-05-08T20:24:09.055047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:adb865e5894ef7617a02aba2245b0b1f6b35202fcb5ea9eeea32e0f5a3fa8930

Observation c8adf3e0-6aa8-4421-ba5c-f644bfdc0e22 · outbound

This paper cites Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:10.962351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:cf1b6ecfcef9106babdda2cfa6e70192e6138d5a99e9503946e5c28a380fb533

Observation 495db4f8-69f6-4cf1-ac3a-3ae3a45effac · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:16:10.996011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:e7f2598d1094c65a97ff5e49f1b33fc2e60d51476c1a1d17ed8901274c3b62d6

Observation 35a7df0e-6aea-4477-9302-d7e7eda67e84 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:16:10.835642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:267c994ae07dee306735f872e2de0ec0aa2223172b7cfa2351449039191a2121

Observation 39d5b8f9-faab-425f-989a-5b24af4e0c76 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Self-Distillation Enables Continual Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:52.723793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:10fa3d22602950652c0ef97d50d27fee1e59f9b1135eeebb2aa389ea93d68f5e

Observation 7e0ba19f-1c61-46b8-bfe5-33b9a18640ab · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization A Survey of On-Policy Distillation for Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:47:19.238105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:c4523ae4734060476f7030068bcd6af5a9b2c7c33dcac116086cde864c1ccf24

Observation 74db5023-b36d-4051-905b-025d7df7ca94 · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:16:10.847225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:a3174475b4f4fb14e997825e09540c23f7fde5c598263383d033b8e7a45766c9

Observation 8c0199b6-0411-417c-994c-25e0da9c78da · outbound

This paper cites Self-Distilled RLVR.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Self-Distilled RLVR

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:16:10.782925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:5aed558543918c23ecaaf8b7bb716f8443235106a79d469cbd5e4aa64d2e0c35

Observation 8057ce1c-bdcd-408b-8aff-517f74b16fba · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:16:10.817694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:0f835b54f60f6f8a9b1c45bb4335f9492089880e109f45f57f8dae52dd6eea19

Observation f6e75022-a335-4d15-a081-e89bed8931fd · outbound

This paper cites Embarrassingly Simple Self-Distillation Improves Code Generation.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Embarrassingly Simple Self-Distillation Improves Code Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-26T01:15:18.478768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:24c62b1bef5625550afb0d376098992775789e2df86189d41e25a0df1bede9bf

Observation 90e88614-6bed-449f-8b19-c2a86e878f3a · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:54:31.187112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:a084373bb34b6e31b29472ed544389ebf584c9b4fdebf36e056871e4e914f958

Observation ec45e4a4-4f48-4097-955a-686854b0ab13 · outbound

This paper cites Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:10.933497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:0e937b0ed6e2a798e27e5044f216bb647acefa2e74acaae63373f2c155f9f81f

Observation ae0a0656-b9cc-4b25-b2b6-9541c4759d6a · outbound

This paper cites Appendix E develops the technical details behind the statistical analysis in the main text.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Appendix E develops the technical details behind the statistical analysis in the main text

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T10:07:34.528922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:cc64b521ad3fd28f7b6da2e0287bc6be152bc460fb793772f05cf4f5a74f423d

Observation 547ffe04-b74c-4323-8039-c3461eec9345 · outbound

This paper cites an unresolved cited work.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-26T10:07:34.535327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:19e9c97dd20ea31922ce7a8973af5a68c0a76bc7f2749cb0ffa0fe2123f455e1

Observation e6476b21-681f-466c-8ac3-196244b87de2 · outbound

This paper cites [2026a] whenever applicable so that the comparison against prior baselines isolates the effect of the proposed PBSD objective as cleanly as possible.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization [2026a] whenever applicable so that the comparison against prior baselines isolates the effect of the proposed PBSD objective as cleanly as possible

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T10:07:34.532204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:331cc299144408ed63ecf8ffd1c2337949071eaba4056deea54ec3fdafd03466

Observation 517b923f-e1e3-4f87-8ba7-675bb2a4e767 · outbound

This paper cites Tool-use data.For the additional tool-use study, we follow the setup in Shenfeld et al.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Tool-use data.For the additional tool-use study, we follow the setup in Shenfeld et al

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T10:07:34.538558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:3447090136d7a037773f01c58a2b216e59ac15ad4ab114ee6b93bdd9aed8708c

Observation c11ee6f9-3076-42b1-bd70-682e600ffc2c · outbound

This paper cites Base (Student).

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Base (Student)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T10:07:34.541532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:62bc59e5a1cb6c369f5d99ade75e5b654331f4b688a3719bd9b488091807bf21

Pith citing papers

Observation dc822383-633f-4f61-a170-3acb28b804b1 · inbound

A Brief Overview: On-Policy Self-Distillation In Large Language Models cites this paper.

A Brief Overview: On-Policy Self-Distillation In Large Language Models Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:08:09.883187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T09:05:30.262601Z digest=sha256:cbeec70a663b7cc8a7714be9a953a557b1f0e0e873b8bd2bca94bd5979c57090

Observation 6159c1c3-988f-4682-a15d-d069dcea0db3 · inbound

A Brief Overview: On-Policy Self-Distillation In Large Language Models cites this paper.

A Brief Overview: On-Policy Self-Distillation In Large Language Models Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:54:46.934280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T09:51:52.886663Z digest=sha256:c12f2c067f2ae05e236f72f16ce3b546116833ebeb5c5e566628a0e892df97c2

Observation e63f3d9a-5097-4b76-87f0-b60c9afd40c1 · inbound

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation cites this paper.

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:07:48.375369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T10:28:18.490452Z digest=sha256:4a629276cb77baec921d4d28d132b231da5150384c35b86bad211744b23b8e74

Observation c1e27acb-7add-46b5-b717-b3c636aedb14 · inbound

DemoPSD: Disagreement-Modulated Policy Self-Distillation cites this paper.

DemoPSD: Disagreement-Modulated Policy Self-Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T16:28:38.436240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T16:21:02.422775Z digest=sha256:6f9900a02f74b0976a9bbae6f605a9c45d21c5d9d83c2f4ddc0671e3a61e8d64

Observation 685b8c5c-e0e9-4cd6-a1bb-a24693448c5e · inbound

DemoPSD: Disagreement-Modulated Policy Self-Distillation cites this paper.

DemoPSD: Disagreement-Modulated Policy Self-Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T08:00:10.945953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:00:10.945953Z digest=sha256:4b05bc361aeb282b080e128028d89bf970cbf685a97f94a5b0e9126b02478bae

Observation 9efda571-6255-4f55-ae5f-ead04887d3fa · inbound

DemoPSD: Disagreement-Modulated Policy Self-Distillation cites this paper.

DemoPSD: Disagreement-Modulated Policy Self-Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T16:40:00.821341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:40:00.821341Z digest=sha256:1066aa0247386da1ca7d1a39eed9243be50baee690b260020e0a78c442880383

Observation ed59648d-76be-4a61-9bd2-e75318123df3 · inbound

Contrastive On-Policy Distillation cites this paper.

Contrastive On-Policy Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.757803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.757803Z digest=sha256:ef2bd316b4368ec1a51aec995c5fc9b5ba72bf219c8aadb0435bfe19ad7aab46