Pith. sign in

Paper Citation Record · LEDGER

Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2603.24472.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.24472 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:18:15.268459Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T00:26:39.269322Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c1fcf864-8175-41a4-87a5-4721055c6e48 · inbound

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models cites this paper.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:25.667356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:45f43e49609d490a47e7c445995390f9c76d41393c07f15db2d6bf036dfbead0

Observation 63c34bef-947c-4ad0-a5d7-d3532cbff5bd · inbound

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe cites this paper.

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:11:06.654703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:05:41.683722Z digest=sha256:4b2e2d9686075db4761255d34b6d5af1d2bf4a1d1c411ddecfee6c74ac1a9f0c

Observation 9b3b0ec3-bb8d-4b92-913c-5786fc743a10 · inbound

How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data cites this paper.

How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:18:21.813100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:18:03.185565Z digest=sha256:24c7b11cb05c32466cc39215417e8ed248a57cd96a4d1be6dc2004dea78a1d14

Observation a50598b2-429e-4133-bab5-7ed85e944b32 · inbound

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency cites this paper.

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T08:22:37.723976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:08:50.330857Z digest=sha256:685c029195b2df12b80f08455e349c7c2b13a00b3b516198226cddbaafb60661

Observation d5c11abb-d573-4498-800d-80f459680e75 · inbound

Multilingual Safety Alignment via Self-Distillation cites this paper.

Multilingual Safety Alignment via Self-Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-09T05:45:22.455008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T19:35:56.059362Z digest=sha256:6d4f3ae111fedddb5dd4ca7144e19878bc3a98de627e88046a9e7df588e4b7bf

Observation f919ae42-9dd9-4994-8b35-2ac166d6f795 · inbound

Multilingual Safety Alignment via Self-Distillation cites this paper.

Multilingual Safety Alignment via Self-Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T04:41:00.870883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T01:08:32.264867Z digest=sha256:9fd16040946be797302a264e024b8c1c06b64fd4c5caf658abdd03cbbfcce99a

Observation 9d5a2af6-6dcb-4107-91c3-39c25dac3069 · inbound

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization cites this paper.

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:16:10.921197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T16:22:46.913172Z digest=sha256:cccebe2f8f23b13fd19b1d80593bf35d78cf8e7de0285f9cbb9a6d92d759a47a

Observation 866cb465-3091-4ebf-aacc-a8e3e1b27e57 · inbound

KL for a KL: On-Policy Distillation with Control Variate Baseline cites this paper.

KL for a KL: On-Policy Distillation with Control Variate Baseline Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.262998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:31:07.462474Z digest=sha256:90c9653ac94e0ddfd4d28e2e2e5100311cf393c05bebb4015428be5624a6656f

Observation e5cfe589-43af-45b6-91da-fe89ffc697c7 · inbound

On-Policy Distillation with Best-of-N Teacher Rollout Selection cites this paper.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:36:27.786285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:32:02.178344Z digest=sha256:5dac33b21833e38440318f295d018d30481bc94d4bbc2e036f5aae5f7c19ad38

Observation 2bd05964-0d53-4e10-8ae4-73fb5830aced · inbound

On-Policy Distillation with Best-of-N Teacher Rollout Selection cites this paper.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.057387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:88442f94700c2fa2218d44314a25adfd6da70739f175f6cfdc84b328b549c087

Observation 846a93f1-a330-40e3-be1f-8963f2aec8d6 · inbound

TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment cites this paper.

TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T02:56:18.634110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:56:09.363823Z digest=sha256:b115327ff45c22e73ed9da13b05dc8b8c060f63acd151c3a2fae2fde2ed41dd0

Observation 38c14afe-d670-4d9a-b719-93cf3d30400e · inbound

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR cites this paper.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:22.561357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:482c18947b1cf562b7ce953757d7de2e661d017ab7625f9dcdb43442a4f61098

Observation cb8accb8-4d0b-4ef8-98a9-10a793ffc49e · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:47:05.147366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:28:18.615371Z digest=sha256:6b3ff439d96ba7857b4eac1de2292574c7ddb7579968ba7fc9893dabe5842311

Observation 511120f4-4d18-4659-bbc0-44f098419d73 · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:22:59.323040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:20:24.066520Z digest=sha256:7ef8c5a18466c847dceaa9d3e31e667215e4d61dfdf9c05b7c546e75bab62f56

Observation 2c614824-c65c-41fe-b1fe-5985de517e59 · inbound

Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information cites this paper.

Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:47:04.756504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:30:03.483296Z digest=sha256:d25748ff8b85fc6b84b35f177167acf13869ec1c86522986d55d63623578f291

Observation 11a2c5f2-7eda-40f5-8100-c9dee6a87142 · inbound

Multi-Rollout On-Policy Distillation via Peer Successes and Failures cites this paper.

Multi-Rollout On-Policy Distillation via Peer Successes and Failures Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T21:32:59.876539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T21:29:36.803832Z digest=sha256:7d7656b91e0d455e64dab763fffa0664bdfd3e8cc4b82508375810268ff4c4a0

Observation 8568ff48-e04f-4c31-a6b2-f4fc27e1238d · inbound

Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-Distillation cites this paper.

Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:28:00.249118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:23:15.511083Z digest=sha256:61dfaef70e7be04bc01e222c0366a5392c37b5b643843d42aaae5b07fe090efe

Observation 05e5fcc2-7374-42cb-ae64-9390de3aa4b7 · inbound

Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning cites this paper.

Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:52:52.422843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:51:17.595707Z digest=sha256:703182f49456e4f310de9e45c40b8e0cbaacc6b1c9c0a426c0c67f28bd4b060f

Observation b4a81eaa-addc-484b-9a52-272695c65853 · inbound

Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards cites this paper.

Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:43:27.595678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:41:31.631505Z digest=sha256:1f928b826100b8de4433fef77e435bcff9e0e6c759b4c2f8b8f67174c4782412

Observation f0b142f2-ceb4-446d-8f56-b624fa9a0424 · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:39:00.362241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:7c2d8b2d3f61a8eca891dbdf69cd38fc8dd9f537afc9cd2d4504eb04b2be8748

Observation de5c9cf8-eee8-4fb5-85d3-896d0d9148eb · inbound

MixSD: Mixed Contextual Self-Distillation for Knowledge Injection cites this paper.

MixSD: Mixed Contextual Self-Distillation for Knowledge Injection Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T21:12:47.032286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T21:10:38.871570Z digest=sha256:130eb0557c7eb301d8ebaf5f063ee95680d667f8503fca7d449076d0d8822702

Observation 030e9b04-aa7e-48cd-9965-717c18376b76 · inbound

MixSD: Mixed Contextual Self-Distillation for Knowledge Injection cites this paper.

MixSD: Mixed Contextual Self-Distillation for Knowledge Injection Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:14:47.875482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-22T10:11:25.002891Z digest=sha256:43618c8394bc388c64c63f9cd16f7583192b53a7edc17c9642649794692bebde

Observation d16b4c0b-2ef3-4326-9b3f-594c76512169 · inbound

A Brief Overview: On-Policy Self-Distillation In Large Language Models cites this paper.

A Brief Overview: On-Policy Self-Distillation In Large Language Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:08:10.033457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T09:05:30.262601Z digest=sha256:62eaafeba57393404faecb53a8db4a5111977873215e7831f4824dd6c4df1210

Observation 1f65c38d-6664-4c02-9525-6095da4806d5 · inbound

A Brief Overview: On-Policy Self-Distillation In Large Language Models cites this paper.

A Brief Overview: On-Policy Self-Distillation In Large Language Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:54:47.006756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T09:51:52.886663Z digest=sha256:3aafdcf152ea3b5f341a37062ed214427435894986fd34533f0619dd2c19f9db

Observation 936d8f0b-b104-4917-8605-696a331409fa · inbound

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning cites this paper.

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:11:17.634261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:06:30.862911Z digest=sha256:3999d8995d597417eacd7497e75c69d7fdf31a955b517278bed18c8e45d94aa4

Observation 7ea9cf5b-6d39-4a07-85d6-c6c2a7bbd9d4 · inbound

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation cites this paper.

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:36:17.384806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T15:14:42.489647Z digest=sha256:741f9b041f85976b2295be20baf76c9435fdf872417c3918e757d9bcc6a3d775

Observation 05165148-a52e-4163-99fc-e3eeb7374f66 · inbound

Constitutional On-Policy Safe Distillation cites this paper.

Constitutional On-Policy Safe Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.519567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:47:14.135793Z digest=sha256:d45e7bef947cd2d9b41203df365a97bf3ae74b982a15459dc326742e3a4d6f0f

Observation 1334b41c-7138-4f18-95c9-016c8ae1dc90 · inbound

Physics-Guided Policy Optimization with Self-Distillation cites this paper.

Physics-Guided Policy Optimization with Self-Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:26.845621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:49:30.431905Z digest=sha256:521c7375766ff90b1b2948aa6ba7d6ff3011980ea09eeb39efd39a101cb9b982

Observation 4d36cec5-8546-4bcb-aeb3-ba631c61616e · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 129

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:26:26.754038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:56:13.058872Z digest=sha256:d5197d56f1304d4ee66bb64f7bc6a3540fb9f63e0c7559b2e5034bf56ea7542a

Observation 25815609-1393-4319-b04e-f9323e2e106f · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 129

Resolution
unresolved
no resolver link, observed 2026-07-13T07:44:25.325808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:44:25.325808Z digest=sha256:b6b1390ca709d7edb86ea81398026b5285e38de80f4462cb880bf89efe4aac8b

Observation 7bfbec51-a20d-4813-b33e-8b3f771fe877 · inbound

Reinforcement Learning from Rich Feedback with Distributional DAgger cites this paper.

Reinforcement Learning from Rich Feedback with Distributional DAgger Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:46:45.925082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T06:44:56.667364Z digest=sha256:b776f6a07d0c7a66aead6ee2c82f38c716e20e65788538d3a078588452dfc32d

Observation 25b44261-e5df-40ad-a3d3-904e1c611656 · inbound

OPRD: On-Policy Representation Distillation cites this paper.

OPRD: On-Policy Representation Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:56:55.253740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T02:46:52.991519Z digest=sha256:2cea2033aac59d011024e4e38e1c6ceca4a285e85823da1db992035752182508

Observation 70849500-5e7e-4f7d-aa6c-25caf480ac13 · inbound

OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing cites this paper.

OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T13:55:58.460765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:55:58.460765Z digest=sha256:4fe7782f04655da9aa30c6a601c59e0f63f3ea9b9a06a2a34ad8d6d04fd6ff8b

Observation 8e8bfeae-f655-4541-a86e-df00a45368be · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.953362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:942a3bab2bab2f418bd789b7cee32b5f6a3825d3a98fff26dd293ea1f5e20ca7

Observation 731c3730-8a9b-4913-a323-582ea2951abf · inbound

Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards cites this paper.

Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T23:39:04.438909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T21:52:29.017826Z digest=sha256:f9ffdd5702e325a4f8baab77863578338a0616f5b71e5196e46092af0a3a7b3d

Observation e6bbf87d-3bf1-408b-8cb8-8556994f040e · inbound

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models cites this paper.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.102957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:177fcd7e55b544dd7afc985a3ddb4079ea5cc4d26638298918165081419d3ded

Observation 21ecdc99-aac9-4a65-95ed-eb78c6da9aae · inbound

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think cites this paper.

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:58:21.048797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T13:56:13.827493Z digest=sha256:f7df7d0311ed82f81f15705a5fdbd1720ea8f2f45be30a07affb8db904bbdadf

Observation 9d40ba8a-d939-4b3a-b217-a7ecb28a0b09 · inbound

DemoPSD: Disagreement-Modulated Policy Self-Distillation cites this paper.

DemoPSD: Disagreement-Modulated Policy Self-Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:28:38.420145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T16:21:02.422775Z digest=sha256:87e5e832e08aad6016a9b97635fd838af67989cb198ca469d76d916a16dff092

Observation 1420105e-e1a2-4d48-8fe3-90accb33f07c · inbound

DemoPSD: Disagreement-Modulated Policy Self-Distillation cites this paper.

DemoPSD: Disagreement-Modulated Policy Self-Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T08:00:10.945953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:00:10.945953Z digest=sha256:ac5eef4ac3d6585f5c4e02296b3a2fa84906a3a0ed580361f2dfd24682ddf0e3

Observation 09fa4929-3897-42bc-b056-bd927d85679f · inbound

DemoPSD: Disagreement-Modulated Policy Self-Distillation cites this paper.

DemoPSD: Disagreement-Modulated Policy Self-Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T16:40:00.821341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:40:00.821341Z digest=sha256:5e1b6252db306c42dd27baead1338d9f9c025d85ef46f0ed55b74600b27f370d

Observation c3e91af8-c3b6-41ab-9971-e775278d404b · inbound

Rethinking On-Policy Self-Distillation for Thinking Models cites this paper.

Rethinking On-Policy Self-Distillation for Thinking Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.528143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:aa4604e9bbe0ca3910de22fd3a69aa5d596cb54ce790fa139a1ea8a85e5eca90

Observation f38f8daa-68dd-4570-abed-bb87ed99bee1 · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:33:45.168462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-07T12:31:42.224094Z digest=sha256:ca066449f1a62e3b6724145ef03d87042e4fd181af570347a1c05a59d76bb8a9

Observation c5cc2cab-e0d2-4d53-a4cd-ec2741f3039b · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:40b4e728632ba178f8d5af079886a94122e12cbe10d038c3c830b80f50c83cf9

Observation 927d8ccf-df69-42f1-9435-a32544abaf53 · inbound

Geometric Self-Distillation for Reasoning Generalization cites this paper.

Geometric Self-Distillation for Reasoning Generalization Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.270869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-10T00:23:06.432913Z digest=sha256:affd96a19aacfd1797412c9f94c7a2b8f07268ec0359f5bcbace73ad07efd063

Observation 46c5f84a-6302-4da1-b772-3ffcdda5e07e · inbound

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation cites this paper.

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T09:08:42.969885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:08:42.969885Z digest=sha256:1b92fe244b1388158374f2e841fb48d6ed34ccc51ef22c2e0f9da9f647ec5bc7

Observation ca6f1c13-f71e-4410-addf-1862f0d31f09 · inbound

On-Policy Delta Distillation cites this paper.

On-Policy Delta Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T00:00:10.604016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:00:10.604016Z digest=sha256:22679651f8f17af7dd8b12fa2ed7dbe7067d5fa4e69b8e7dfe936864155dd2fc

Observation dd30c851-1b36-49c9-be0f-c6f257d31214 · inbound

Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents? cites this paper.

Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents? Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T17:42:16.061395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:42:16.061395Z digest=sha256:14216046e3bee6f49fea53f199402bfd5a631655b76326521a2a048b42c79125

Observation 27557ec8-3224-4e3a-878c-f0633e72cae4 · inbound

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks cites this paper.

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:59.676822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:59.676822Z digest=sha256:08d31eed58f4d8f2a60acfb0ff27abb5048ad145978cf3bc522e4677a1ce193c

Observation ac387a91-c932-4b01-92b2-fc46337d6360 · inbound

$\beta$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation cites this paper.

$\beta$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T03:08:53.649699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:08:53.649699Z digest=sha256:ba58c6dbb9e4b4dc2b972c4935d9a92d0fde0ba552757188a9757e501e114f13

Observation 2b5dd57a-e0bf-4553-b330-d49629def9bd · inbound

Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning cites this paper.

Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T00:37:18.865054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:37:18.865054Z digest=sha256:0826ac320f36874ef663aa4dcc053209a54c96fc31e1682b6411a560bdd2d8ff

Observation 9024278d-9b20-4eaf-ac2d-ac6e9a26f4bb · inbound

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning cites this paper.

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:18:15.268459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:18:15.268459Z digest=sha256:3e2c9f47cbe9d03b5ee518379958c2dfaf20a748c41f440bdf2581b35237f11a