Pith. sign in

Paper Citation Record · LEDGER

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation

As of 13 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2608.09826.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09826 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:57:51.075050Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ba62059-8747-4625-a908-af4364bfbe6c · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:50.994951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:50.994951Z digest=sha256:ce01ee0e4f197e70450c6cea50226a4f7e9fe4dcf277d8b9d1e5b9c2821d8fa0

Observation 4f8ac343-3629-46c1-a33f-6d7448a3298e · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Entropy-Aware On-Policy Distillation of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.002061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.002061Z digest=sha256:5aad167138a4f7ad66e18a7199876923257ffa9dbf3142944ddb38af28303b0c

Observation 04e77260-9a20-408f-8ee9-8fbe1360e585 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.017525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.017525Z digest=sha256:d169638291a99738cc97ea6305e52e4dd9faee16dcf1f607a0cfeccca74272dc

Observation fe98a91a-b63a-46ce-9cba-d6518db871bd · outbound

This paper cites InInternational Con- ference on Learning Representations, volume 2024, 39578– 39601.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation InInternational Con- ference on Learning Representations, volume 2024, 39578– 39601

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:57:51.384107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:57:51.021814Z digest=sha256:0b8db7162e2f4c82ff77559f88fbf1bd20d59bc71d2ead556725ead38cf50904

Observation 83618215-19e0-45f4-a1b5-4f4b6aa7dfde · outbound

This paper cites Nam, T.; Sun, S.-H.; Pertsch, K.; Hwang, S.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Nam, T.; Sun, S.-H.; Pertsch, K.; Hwang, S

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.030096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.030096Z digest=sha256:854cc9271f05daf86acc25aaa23f191b0400b4cc9de0a89d315b6cea46137ead

Observation b7a4e63d-38d9-4e0c-9d03-483c4181436f · outbound

This paper cites Skill-based Meta-Reinforcement Learning.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Skill-based Meta-Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.033925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.033925Z digest=sha256:4e4715e32d30b9afe794cdd6d54dc4fac175bc584902559fc444f7940fba07ce

Observation 74c6dd90-6804-4140-8cff-c30037ed0391 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.038075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.038075Z digest=sha256:0f689e28c3bcc7a562f297ba12bd04635c7333a6469e0d5cb236f5d62b12a042

Observation 0c257387-727d-4507-a38e-e0064d82d088 · outbound

This paper cites Skill-based Model-based Reinforcement Learning.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Skill-based Model-based Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.042725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.042725Z digest=sha256:88794a85ddb35d222f0698eb832de4b53c0842fdab51e058a202cd9f7a89c320

Observation d355fa74-b0e1-4ff6-a014-c3ee91c4fbb9 · outbound

This paper cites Learning by Distilling Context.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Learning by Distilling Context

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.046808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.046808Z digest=sha256:e7a42a45f75a45da53e773e12171aae2a5e1ccddf1d47a5e74b308cf558b1b00

Observation c01aace5-e8f3-4764-a66e-af538adb1447 · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.055042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.055042Z digest=sha256:f80346f5fdc7136d68f3f5c8c133199336e3af632534b787d262d48a7359a71d

Observation 1ea6ecc2-2dfb-41ee-8fd8-45873a489136 · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.059264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.059264Z digest=sha256:0d3cb8822dfe92a399eed7b47e4b105c5354c18be7cb419f4eb5c12908442674

Observation 3d608ef9-024e-4b91-9396-5790c6e51917 · outbound

This paper cites A Survey on Knowledge Distillation of Large Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation A Survey on Knowledge Distillation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.063267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.063267Z digest=sha256:e6e7c9aba168f54d94c44d5158ad82d7324ec40830842dd82112abac3b63911a

Observation 91918b64-a48b-4a63-b639-a8d3dccac5e1 · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.067414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.067414Z digest=sha256:b75f0739dd9238dbc14f10b034bad623df7f54f08f0ca5984fa787b0998ee131

Observation d568653b-c3d7-4471-be70-3169ce2aee47 · outbound

This paper cites OPSDL: On-Policy Self-Distillation for Long-Context Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation OPSDL: On-Policy Self-Distillation for Long-Context Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.071097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.071097Z digest=sha256:6ddd5443002bea7580b71e1ebcdaf0bd45161da71d7a0d1849b5103b5368388e

Observation 46216a6d-52ce-417d-8b25-0848781b78bc · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.075050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.075050Z digest=sha256:badf9811609c3fcabfc9ff387f7db644d6c75a5a8c4d32a3e5fe95ff1d907bf9

Observation 49f09919-1614-428a-9f8e-00ce3e13d919 · outbound

This paper cites Rényidivergenceand Kullback-Leibler divergence.IEEE Transactions on Infor- mation Theory, 60(7): 3797–3820.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Rényidivergenceand Kullback-Leibler divergence.IEEE Transactions on Infor- mation Theory, 60(7): 3797–3820

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:57:51.372408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:57:51.051171Z digest=sha256:378f34e14168f44af4d03b427b535b010917fbb0db7bbd99c1b7bae90703376d

Observation abb5e02e-59b5-4365-9795-b310804e5112 · outbound

This paper cites In-context Reinforcement Learning with Algorithm Distillation.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation In-context Reinforcement Learning with Algorithm Distillation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.008080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.008080Z digest=sha256:8620bc2a10b2b734256c1c1be76890b411a2f039c5ba94f8b6c597c1a695370b

Observation 39e02caa-754d-4b65-9b75-46df2de1ce23 · outbound

This paper cites Gou,J.;Yu,B.;Maybank,S.J.;andTao,D.2021.Knowledge distillation: A survey.International journal of computer vision, 129(6): 1789–1819.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Gou,J.;Yu,B.;Maybank,S.J.;andTao,D.2021.Knowledge distillation: A survey.International journal of computer vision, 129(6): 1789–1819

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:57:51.395143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:57:50.990619Z digest=sha256:edd458c80c5cd2002b4e41a91451607ff8c86e2db7bfc6f49aa1f7b5387acebe

Observation 49f2b6e4-87a6-4f3f-ad81-d79f9f09dca1 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.012718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.012718Z digest=sha256:a822865d586c2b27ec31b9d51b3e35b48f6277faefef8e5c2274fd146c7e87ed

Observation 8c4ed339-f714-4e9d-a641-632708282d8a · outbound

This paper cites an unresolved cited work.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:57:51.406340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:57:50.980488Z digest=sha256:3173b75cba53ddc87146eeba8397ee003223803b7ecaea90dda8cfed2f15a654

Observation 74884d22-95b9-49bc-936c-5d299c6fcb1e · outbound

This paper cites Unifying distillation and privileged information.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Unifying distillation and privileged information

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.025787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.025787Z digest=sha256:83153cb2c894064186389879481aedcec47228ab52894abce1c2232fd09c333b

Observation f082d908-3485-4b6e-bd7f-6227454726d8 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:50.985760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:50.985760Z digest=sha256:b8bcf515bf536d80bbf5e73587f8533275e3e35c91f5d9570f5c9c555f0b6ec2

Pith citing papers

No inbound Pith citation observations are available.