Pith. sign in

Paper Citation Record · LEDGER

Privileged Information Distillation for Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2602.04942.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.04942 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:07:21.573005Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T00:26:39.270366Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ffe965c4-710a-4edc-ade9-001c6692847f · inbound

Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation cites this paper.

Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation Privileged Information Distillation for Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T05:11:07.600760Z digest=sha256:c3b63cf3dbeaf6107e3a87b8dea048e993f6fd2645584f2faa3b292e4c288e4e

Observation 0ed377c2-72a4-4294-8b83-1c2f0f9d703d · inbound

Embarrassingly Simple Self-Distillation Improves Code Generation cites this paper.

Embarrassingly Simple Self-Distillation Improves Code Generation Privileged Information Distillation for Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T14:33:35.834383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:33:35.834383Z digest=sha256:bd076a4a3f2c5911db18f3a23dcdfd0d63dbb7bf991226c92b4016dbac252fbf

Observation 5e5dbd66-ba12-4559-bfd5-93997e57e823 · inbound

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents cites this paper.

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents Privileged Information Distillation for Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T15:28:07.981488Z digest=sha256:26464a9b8a2beced42214289fb995fd0e832fc5d9df0e365b4cae03b8a464b97

Observation ec78b8c4-8b51-4007-a19b-0137648ac39d · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation Privileged Information Distillation for Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:12:06.985609Z digest=sha256:35829a9963ab57d9232410a1a659454c4cf068a3d23e552d11399ac932eed249

Observation 5ebe4292-7ac0-45c8-937f-d7f52b6f0daf · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation Privileged Information Distillation for Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:04:12.454268Z digest=sha256:aba7a50f754704fe37bbed062715ecdd2c16b98a8600745efe9418877351404e

Observation 612a8927-5dae-4f46-8951-4ce08df33aaa · inbound

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data cites this paper.

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data Privileged Information Distillation for Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:43:09.581054Z digest=sha256:fb8254849d8a180b2d3f85fc383deb9345f0117121ecdd72f88671308019f37f

Observation bb77c09b-b739-4431-a4c2-b02badc0ce0b · inbound

Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models cites this paper.

Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models Privileged Information Distillation for Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T10:54:57.732141Z digest=sha256:a3d6611aa37e2ca15b25b2ff73bb9086f488bea544338ca2d1e3d50436ce5abc

Observation 78479ed0-6f56-4576-a841-cf6bf83ae808 · inbound

TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents cites this paper.

TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents Privileged Information Distillation for Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T04:39:01.283802Z digest=sha256:3271c22c9bda07383da88f33672e0d2c7a44a8e18d8ebcffc27367d29f044d07

Observation baaebeef-1d64-40e4-b31f-bffd6e6052f3 · inbound

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate cites this paper.

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate Privileged Information Distillation for Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T15:18:39.844534Z digest=sha256:c8533ad1fc3bb485c200b753ef92d9640a830481056b3273b429a43de81b3762

Observation a159b673-3049-435f-b698-ce87c363e5d2 · inbound

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models cites this paper.

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models Privileged Information Distillation for Language Models

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T23:18:35.390642Z digest=sha256:266e431bbbcc090e5a2dd27d48d6117e0d7887e1c9fd0415c3cfdc805891175b

Observation d388fb8a-6378-4287-a5fe-d4519f9ea38c · inbound

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models cites this paper.

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models Privileged Information Distillation for Language Models

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T23:45:07.665274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T23:44:10.302520Z digest=sha256:a2a1c5cd27ad77eb928b982b93656caf7cf9da255110868669acfd861abf409a

Observation 2316f8ab-056a-4058-9cdc-b4901d134cb5 · inbound

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models cites this paper.

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models Privileged Information Distillation for Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:46:22.602765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T09:44:55.200796Z digest=sha256:a93615b75fc76a36e9ae611e3e9ac39ec19c5a315d5b6d3f120a2b38c8e41d3e

Observation d2b5a07e-82c1-4f98-9349-b4b9f16a1ed1 · inbound

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair cites this paper.

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair Privileged Information Distillation for Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:21:43.166304Z digest=sha256:319d54e4dfed5751a7825b4a4a5ef2dcf235fad12c24c195f94cbb60684f7cb0

Observation b1e518f7-5015-4e73-8957-0c918aa3f206 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Privileged Information Distillation for Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:dd1e11dea8c89f5c4dce0e988ac459672404728f31dd324dcc9d42aaa86d1333

Observation 4cd43221-c830-45d1-8484-a6635327a95d · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Privileged Information Distillation for Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.197177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.197177Z digest=sha256:94a6ce5234ca12856c51f60bd755531ad0a0d72b3cf4b447f0f0d002fac426f2

Observation 6f6de619-4937-4576-bfe3-4636ff447e27 · inbound

TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment cites this paper.

TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment Privileged Information Distillation for Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T02:56:09.363823Z digest=sha256:e96bbd48ddc4e9005e80a6c604b2b04dc4f33a0a3eee76434b77edb63f6e9bd0

Observation e8a69d5c-9c9b-4028-b70f-ae8d3e965e64 · inbound

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why cites this paper.

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why Privileged Information Distillation for Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T04:28:58.679078Z digest=sha256:e1d3e38029a70f59ebd2f53059fe2333114a0c8bd73a7b1a43e9112a9bb60109

Observation 94e94df9-0b87-46f0-b2f4-a234f4f35026 · inbound

From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation cites this paper.

From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation Privileged Information Distillation for Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T01:24:34.093541Z digest=sha256:09a7e4490654bbe8f5d711d4ad2739c06e83432e46f40b37f7eb6b8b78037c9a

Observation 2e05cd63-f5bd-4c92-b777-47e963b74fb2 · inbound

Multi-Rollout On-Policy Distillation via Peer Successes and Failures cites this paper.

Multi-Rollout On-Policy Distillation via Peer Successes and Failures Privileged Information Distillation for Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T21:29:36.803832Z digest=sha256:9641c4f8d663693a20bf2f103a30e4ed3124986cb1853758db7281b585a01e79

Observation d708a2ee-5b78-40ba-8e42-5ebe73006382 · inbound

Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-Distillation cites this paper.

Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-Distillation Privileged Information Distillation for Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T21:23:15.511083Z digest=sha256:ccd82d8fd6b2d25524a588451d7702dfc9dec9f778ee8cf889f1c07e193a7f99

Observation 30691877-32e6-43e2-9e93-e04e7e774518 · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation Privileged Information Distillation for Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:23c034bec51d8e763cab1b753387e278219e2a35005b5a866c0dce8637e2ca5f

Observation c0ff8b8a-d77a-4176-9b12-ad6241303b2b · inbound

A Brief Overview: On-Policy Self-Distillation In Large Language Models cites this paper.

A Brief Overview: On-Policy Self-Distillation In Large Language Models Privileged Information Distillation for Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T09:05:30.262601Z digest=sha256:c6766f20791fe72ddfa5cfa3b44270a8a29643120cde091bf68808eda7f612ab

Observation 88a5c801-7be1-43af-8153-da58b2e61155 · inbound

A Brief Overview: On-Policy Self-Distillation In Large Language Models cites this paper.

A Brief Overview: On-Policy Self-Distillation In Large Language Models Privileged Information Distillation for Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:54:46.865979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T09:51:52.886663Z digest=sha256:1d4f023470192d264ddc720cdb72643519b97c0ef4d16da422f83c912e568fb7

Observation 46a7e490-75b2-42af-83ab-17634ea31430 · inbound

Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction cites this paper.

Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction Privileged Information Distillation for Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T02:50:44.932051Z digest=sha256:04cf4a9c31a41428bf9747ea312506c41872d4bd55b7c445a81bcc9dda81e437

Observation 2d2c7633-b550-4205-baa4-9fab09c1892c · inbound

Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction cites this paper.

Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction Privileged Information Distillation for Language Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:14:47.711714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T10:12:17.474628Z digest=sha256:9eb3b6cb54d5402735b584db2095eac37b60ba68037e576800073a774ec52918

Observation e0793516-e4ec-483a-b45b-b9f87305e1cf · inbound

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization cites this paper.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Privileged Information Distillation for Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:72f444a1efb19c478e232b476ddd43d888690410ed2bb67e1a0e1cf48ddd0dde

Observation 971f7dc3-1261-4dd5-b6c8-24ac4fb9a243 · inbound

Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL cites this paper.

Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL Privileged Information Distillation for Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-22T07:24:31.355799Z digest=sha256:d13dda6454b1335ae018f77ec6417d24ed331b7492c116e4f0b0ac8a50ed73bc

Observation 87490850-665c-46d0-beb9-236f9311db60 · inbound

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning cites this paper.

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning Privileged Information Distillation for Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T08:06:30.862911Z digest=sha256:e2e4045ac366495683712d72394f8c8edce8d17926c86f339375636c833409dd

Observation c175ca74-6d8a-460c-925f-0d49e92dc74b · inbound

Self-Policy Distillation via Capability-Selective Subspace Projection cites this paper.

Self-Policy Distillation via Capability-Selective Subspace Projection Privileged Information Distillation for Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T05:31:42.803465Z digest=sha256:f651e7fa4912ef61a15303c4cac381b2f554e4034eb8da239d4a341821957c3b

Observation 6874a8a3-0c4a-4256-9ad5-2d54906575b0 · inbound

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning cites this paper.

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning Privileged Information Distillation for Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:45.227482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T17:14:47.440914Z digest=sha256:8a02292f676cd6a71ec7d4d592374edd729b9c235237d2f95b4ad21411d0bcd4

Observation cefce76c-8e83-41a4-97f6-cd994bc81952 · inbound

Self-Distilled Policy Gradient cites this paper.

Self-Distilled Policy Gradient Privileged Information Distillation for Language Models

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:56:27.869924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T11:24:11.878439Z digest=sha256:f184220c8ae69d7fdb817454677722dce132990511951f7a3b43eb02219f6d3a

Observation b6790bb3-d48b-418a-adc8-bfc77713c643 · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Privileged Information Distillation for Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:07:27.961777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T17:30:57.001021Z digest=sha256:8a124e010b1805c0cf4e7f1021fa743054dddf5dbcd0a000a8eda0efc3f68c64

Observation 65f75f0f-f3c6-47b0-8184-50193743d983 · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Privileged Information Distillation for Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-15T10:53:37.186361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:53:37.186361Z digest=sha256:8545705741d5627cb98372ba4bb4386c99d066cf4107eb5ecced2f55625ec8ae

Observation ff700a35-f8c9-424c-8087-7e019b34f4dc · inbound

Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts cites this paper.

Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts Privileged Information Distillation for Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:47:37.751059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T13:42:21.520627Z digest=sha256:368c293181e7b55dc7e7a17641f327f5c71831876de6e2c88a5baf9cb400693d

Observation 6d35a4cc-40f9-419a-8642-65d44bb1c2e1 · inbound

Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation cites this paper.

Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation Privileged Information Distillation for Language Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:07:36.568006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T14:15:30.242783Z digest=sha256:4be546daab9fd89c6c56f67e55c8d4a7cafaddaecc4cd537482861cdd6682f5c

Observation c3d31d0d-f718-4252-ade1-f7950cf93721 · inbound

HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation cites this paper.

HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation Privileged Information Distillation for Language Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:27:56.073037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T10:03:48.678510Z digest=sha256:05a0b3f6e27a44d3785996ea8af6a351704880927c754cad48d3d54b2ab01300

Observation ec01ee55-5ce1-4f1f-a21d-dfd4df889de9 · inbound

ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models cites this paper.

ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models Privileged Information Distillation for Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:20:07.067987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T20:22:16.508280Z digest=sha256:3a38de9e7eaf56d69566e0b90a48dbdc08f59559815f1a9e8fe6d5b38af3a1ae

Observation 2cbdfe62-ef0a-4716-a818-828af8904b68 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Privileged Information Distillation for Language Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T20:50:12.726056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:cb2859ad06459640147cdcaed764553eaee077cb9711829ca99de8510a3920df

Observation a22a6ef2-263b-4dd7-943d-b8651e98514e · inbound

Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts cites this paper.

Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts Privileged Information Distillation for Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:24:26.996081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T06:05:47.948367Z digest=sha256:77d4dd4ecbd7306e76075526ddff4031c182f6b126f55d277890635a4f777cd2

Observation 0e8a175d-d3b7-4f5b-98c5-8336b591c47d · inbound

DOPD: Dual On-policy Distillation cites this paper.

DOPD: Dual On-policy Distillation Privileged Information Distillation for Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T05:54:18.576890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T05:51:18.037199Z digest=sha256:854f68b646f2e6ea1eedfb56b87bad1ddc68684691c442f1e682113d3bccbfc5

Observation f0e7e388-ff58-4425-a764-6195cda5f8ce · inbound

TREK: Distill to Explore, Reinforce to Refine cites this paper.

TREK: Distill to Explore, Reinforce to Refine Privileged Information Distillation for Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.067687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:8f5e8dba042478cfebea4380917f2d0b10db6b93147927a660b3be49c774a995

Observation 243a904d-3dc2-4e64-a5cf-e60a24035d5a · inbound

Geometric Self-Distillation for Reasoning Generalization cites this paper.

Geometric Self-Distillation for Reasoning Generalization Privileged Information Distillation for Language Models

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.271678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-10T00:23:06.432913Z digest=sha256:2b9a111d81a1fa870c0c7edd24864a7460a7a3b3358905445671e49e9e8712da

Observation 5b3b8c12-6253-4733-8524-2de3eadc56a3 · inbound

Contrastive On-Policy Distillation cites this paper.

Contrastive On-Policy Distillation Privileged Information Distillation for Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.685025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.685025Z digest=sha256:c0dc410e380e8c031715dab277b91e8d506b7d90af205970e8782a8d508a0fdd

Observation 0a9e510c-3637-4742-b0c4-f8c16dcd3e72 · inbound

Sample-Efficient Learning from Agent Experience cites this paper.

Sample-Efficient Learning from Agent Experience Privileged Information Distillation for Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:43:21.669819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:43:21.669819Z digest=sha256:5b7d1c5267561298418fe96c77670e9e83116817e827e91e7fc1e8e44a0afdad

Observation 2925f4a0-efba-41b0-9bf5-0385742deb9b · inbound

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation cites this paper.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Privileged Information Distillation for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.167530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.167530Z digest=sha256:d16bd9a8f628a535539381c7165465fab5cd615fec713747e281c1fd6b05c917

Observation d81f8ef2-0add-49a1-8a84-d1e750300254 · inbound

Pass the Baton: Trajectory-Relayed On-Policy Distillation cites this paper.

Pass the Baton: Trajectory-Relayed On-Policy Distillation Privileged Information Distillation for Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:49:26.988561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:49:26.988561Z digest=sha256:75da0a4596318d47f74e40a44a8010c1809e42dc37c44986ad1a0687aff061cf

Observation b1cba096-2ea0-40f9-8879-457da848da60 · inbound

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning cites this paper.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Privileged Information Distillation for Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.080946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.080946Z digest=sha256:46a3872118df700ab57bbb45b4984fb188b7803b4f9939e1ddb94ac9c97327ce

Observation b30ab58f-eed9-4e45-8784-87573d40aa35 · inbound

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models cites this paper.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Privileged Information Distillation for Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.469715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.469715Z digest=sha256:c7d0231349c1eec4b06612a0b05fef8a2d6c05679e2cf227ea8fa16aa9c6fba8

Observation 9f60ec44-9bb8-488f-b329-ee89d846c1d3 · inbound

Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning cites this paper.

Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning Privileged Information Distillation for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T00:37:18.978472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:37:18.978472Z digest=sha256:4c84fdc71d7fe28162d5bcf3273e7fc95e19604174c5e0fe9031ea9375edc367

Observation 0affb3db-dfba-4460-8545-87dfecc7d1ae · inbound

DAPD: Dual-Anchored Policy Distillation cites this paper.

DAPD: Dual-Anchored Policy Distillation Privileged Information Distillation for Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T22:02:31.427078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:02:31.427078Z digest=sha256:f1c0fd3e0664bad7ea93eb8167e33e3050da1ec0cb89c8deb92a0f0f2e8ef2c8

Observation 3b9a869b-3f76-4b13-8b96-023b04888c2d · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Privileged Information Distillation for Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T15:25:49.251304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:25:49.251304Z digest=sha256:eec3cd0cedb694519ae9a7f8558192500066ed23a2ca163fd3df972bcbf26a92

Observation 487ce935-df39-4d34-b567-48054be6a356 · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Privileged Information Distillation for Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:15.847567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:15.847567Z digest=sha256:71f10979e598b32c82a42a6ce60d981e42ec892bab034acb976fc18eff09aa1a

Observation 536e9971-7b0a-4e0d-87eb-1d46b5525292 · inbound

Self-Improving Large Language Models via Progressive Experience Evolution cites this paper.

Self-Improving Large Language Models via Progressive Experience Evolution Privileged Information Distillation for Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:12.116064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:12.116064Z digest=sha256:6056dc62507d2aa15f9fa235cd76bd5367130916898a5a2c05bdbf1fb02862ea

Observation 17bae05d-6cd2-4496-a84d-ee2afd62df82 · inbound

Self-Improving Large Language Models via Progressive Experience Evolution cites this paper.

Self-Improving Large Language Models via Progressive Experience Evolution Privileged Information Distillation for Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.667860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.667860Z digest=sha256:c3d135c9a4d18b776edb5db09fa5e19470150cfc553601af02a75e34aa040ce9

Observation 5d0201ab-a43f-49dd-9762-5644e0729f83 · inbound

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast cites this paper.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Privileged Information Distillation for Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.934410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.934410Z digest=sha256:758e80fe2924ef6048dfcfc42e83ad380c4d72ad7c0412c320986fa3102b89ab

Observation 13844ce5-9c4b-45f4-a5a6-165b9150862c · inbound

BOUND: Brief-Guided Corrective Preference Distillation at Search-Control Boundaries cites this paper.

BOUND: Brief-Guided Corrective Preference Distillation at Search-Control Boundaries Privileged Information Distillation for Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:30.562391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:30.562391Z digest=sha256:0418dbb9307ac19d6f4e2a49fa87a0aea957cd7ded44777e5117dd6d8308db67

Observation c974849f-93cb-4426-bf4e-9ba3419b14f9 · inbound

Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation cites this paper.

Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation Privileged Information Distillation for Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:10.334200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:10.334200Z digest=sha256:827e16864b246f487b5fb49ec8137a299fa7162e6b9aa3c2ad1753f5682b9fa1

Observation cf81325e-f6c6-49b4-98f6-d0a18497a1a4 · inbound

SR-OPSD: Self-Referenced On-Policy Self-Distillation cites this paper.

SR-OPSD: Self-Referenced On-Policy Self-Distillation Privileged Information Distillation for Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:45.743478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:37:45.743478Z digest=sha256:5f65e124c96fffc7a9fb746630cc7a441d320a88a2e279da40195d61bd21334b

Observation 9d7f7437-2607-476f-be08-137d76aed4b9 · inbound

ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection cites this paper.

ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection Privileged Information Distillation for Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:50.695524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:50:50.695524Z digest=sha256:017b717edf4fd73ce742aaaa92c214425ceb54c890cf2a55eefff505938697df

Observation 681049ca-33c3-44f4-b910-6864eea1494a · inbound

Latent On-Policy Self-Distillation cites this paper.

Latent On-Policy Self-Distillation Privileged Information Distillation for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:07:21.573005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:07:21.573005Z digest=sha256:9a1d243cef12c64fa676836fdfac4f597c8dd4a9144f103d4f9c4f118ccabf65

Observation c0d6e633-0499-4ac2-b14a-8fe661d1ea9b · inbound

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents cites this paper.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Privileged Information Distillation for Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.136223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.136223Z digest=sha256:6b396048144af85cd462386dcbb01f58d5f8e85cfdade2991dceeb3009cdfa82