Pith. sign in

Paper Citation Record · LEDGER

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation

As of 13 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2608.09826.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09826 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:57:51.075050Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ba62059-8747-4625-a908-af4364bfbe6c · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:50.994951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:50.994951Z digest=sha256:2990e1109053916541a5a6067c54b2a0849108fa4e30ecd736c7b2fb700f8df5

Observation 4f8ac343-3629-46c1-a33f-6d7448a3298e · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Entropy-Aware On-Policy Distillation of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.002061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.002061Z digest=sha256:efa746907fc6974eb16f3ea913c95afdf7e168a1ed4dc586c9373bedf38903f3

Observation 04e77260-9a20-408f-8ee9-8fbe1360e585 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.017525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.017525Z digest=sha256:ec6908e0f7cb6b6e3190c53cbc7ef7a8d993257ceb0453e38142301f0777e775

Observation fe98a91a-b63a-46ce-9cba-d6518db871bd · outbound

This paper cites InInternational Con- ference on Learning Representations, volume 2024, 39578– 39601.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation InInternational Con- ference on Learning Representations, volume 2024, 39578– 39601

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:57:51.384107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:57:51.021814Z digest=sha256:512678daa8652526e480c659cba0546565453de6bb5f08ce3b85e766068daf38

Observation 83618215-19e0-45f4-a1b5-4f4b6aa7dfde · outbound

This paper cites Nam, T.; Sun, S.-H.; Pertsch, K.; Hwang, S.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Nam, T.; Sun, S.-H.; Pertsch, K.; Hwang, S

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.030096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.030096Z digest=sha256:90a06fd36ed42b95970ed9a13b874d4d2ed6100d18e6f3ccba01a1d275f69f13

Observation b7a4e63d-38d9-4e0c-9d03-483c4181436f · outbound

This paper cites Skill-based Meta-Reinforcement Learning.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Skill-based Meta-Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.033925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.033925Z digest=sha256:8af03f059c4e14da0e1f6edce2c8cd38891945316741df2d37276925e4aa8d17

Observation 74c6dd90-6804-4140-8cff-c30037ed0391 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.038075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.038075Z digest=sha256:407bdd4d1542c4a68593705d6a2b9406f48228af8e1fbd6ca85c4c7869861aa8

Observation 0c257387-727d-4507-a38e-e0064d82d088 · outbound

This paper cites Skill-based Model-based Reinforcement Learning.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Skill-based Model-based Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.042725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.042725Z digest=sha256:09c273f1e5f33e36dd803cf19a42bfb0c399b7d4f6ac4f5bd973a7821bff335e

Observation d355fa74-b0e1-4ff6-a014-c3ee91c4fbb9 · outbound

This paper cites Learning by Distilling Context.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Learning by Distilling Context

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.046808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.046808Z digest=sha256:a785d2a883170207a7f0dbb1db2789be1355571041136e7e4a79e6290a74a017

Observation c01aace5-e8f3-4764-a66e-af538adb1447 · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.055042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.055042Z digest=sha256:241a4306b37149f2fb98fedb2018ed07fc9f3b319573ffd53748ba6fd36e2ef2

Observation 1ea6ecc2-2dfb-41ee-8fd8-45873a489136 · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.059264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.059264Z digest=sha256:76bf3f0d3d2a963a65775f743db133846b3ca0f3a8ef055bdf4b0397130ed2c3

Observation 3d608ef9-024e-4b91-9396-5790c6e51917 · outbound

This paper cites A Survey on Knowledge Distillation of Large Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation A Survey on Knowledge Distillation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.063267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.063267Z digest=sha256:be839853ba8c76d4a6ccba6c058cadd09f472940b29fb5b5dd1d25054e8ca559

Observation 91918b64-a48b-4a63-b639-a8d3dccac5e1 · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.067414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.067414Z digest=sha256:2ba2728215def1471a8d2edb2bfbe4e397b00dc9185c954643795f06bfdf64b2

Observation d568653b-c3d7-4471-be70-3169ce2aee47 · outbound

This paper cites OPSDL: On-Policy Self-Distillation for Long-Context Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation OPSDL: On-Policy Self-Distillation for Long-Context Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.071097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.071097Z digest=sha256:35b9bf557070efa4e1d73dab6175311b62d3539031a5c78a4ce0f75844044140

Observation 46216a6d-52ce-417d-8b25-0848781b78bc · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.075050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.075050Z digest=sha256:5de319644098d52a35c323cdfbb6bcb080602fae5f5c77a055fdda02c64fede3

Observation 49f09919-1614-428a-9f8e-00ce3e13d919 · outbound

This paper cites Rényidivergenceand Kullback-Leibler divergence.IEEE Transactions on Infor- mation Theory, 60(7): 3797–3820.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Rényidivergenceand Kullback-Leibler divergence.IEEE Transactions on Infor- mation Theory, 60(7): 3797–3820

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:57:51.372408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:57:51.051171Z digest=sha256:980dd73158e2ccdc6726081ced0fe1a67cfa170766464776bc46581534219eba

Observation abb5e02e-59b5-4365-9795-b310804e5112 · outbound

This paper cites In-context Reinforcement Learning with Algorithm Distillation.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation In-context Reinforcement Learning with Algorithm Distillation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.008080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.008080Z digest=sha256:b25179a916291b707258c0b7c852a742ae775187f8c5f1d07eb0bf624a70ca50

Observation 39e02caa-754d-4b65-9b75-46df2de1ce23 · outbound

This paper cites Gou,J.;Yu,B.;Maybank,S.J.;andTao,D.2021.Knowledge distillation: A survey.International journal of computer vision, 129(6): 1789–1819.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Gou,J.;Yu,B.;Maybank,S.J.;andTao,D.2021.Knowledge distillation: A survey.International journal of computer vision, 129(6): 1789–1819

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:57:51.395143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:57:50.990619Z digest=sha256:ec3db44a90e74a8fb15a8363eebb8d206ea6f828d53d1b1c0646e3904b012e6f

Observation 49f2b6e4-87a6-4f3f-ad81-d79f9f09dca1 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.012718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.012718Z digest=sha256:3708c26dca41ece7b3db9c69b9a6e0100576d1a739f51a803b4f7b426e9240b6

Observation 8c4ed339-f714-4e9d-a641-632708282d8a · outbound

This paper cites an unresolved cited work.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:57:51.406340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:57:50.980488Z digest=sha256:7ab2f0e202bb5e4594172d6c19d672c48d1f8b3e254bb432d54dd54de9d86d2d

Observation 74884d22-95b9-49bc-936c-5d299c6fcb1e · outbound

This paper cites Unifying distillation and privileged information.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Unifying distillation and privileged information

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.025787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.025787Z digest=sha256:8d61c999a908ef3596f48d372dffd55b3e371020cb75f779051b7da30c7f6fbb

Observation f082d908-3485-4b6e-bd7f-6227454726d8 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:50.985760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:50.985760Z digest=sha256:53e1f738df1f3470a60bb0b27047bc08d63b7d026ec1b3a2ee6836e7c0bdfd96

Pith citing papers

No inbound Pith citation observations are available.