Pith. sign in

Paper Citation Record · LEDGER

On-Policy Self-Distillation without Any Supervision

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.06296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06296 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:43.839213Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9d279ec-09f0-4ddb-9790-a4860180521e · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

On-Policy Self-Distillation without Any Supervision Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.326448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.326448Z digest=sha256:11e55bb0078f05cee460ebbb68657e61c6f87bb1951f2692df8355f458550f47

Observation 1b7637fa-51c4-4807-a2f1-43ef153e9df6 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

On-Policy Self-Distillation without Any Supervision MiniLLM: On-Policy Distillation of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.411447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.411447Z digest=sha256:808b122b21516fe52e74a992cfe13a4ccfaf67fd0eef3dcc1f92fa23cd5a5c91

Observation f9cb2373-0e40-4d27-8f14-1197802d626e · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

On-Policy Self-Distillation without Any Supervision OpenThoughts: Data Recipes for Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.444364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.444364Z digest=sha256:da101521f5b5b2fedbd2ea71c65de08e9d8e789ebb82ed5c14c7c614db7d34da

Observation df86c3a3-12bb-4391-b060-2bf05ce0c6dc · outbound

This paper cites Large Language Models Can Self-Improve.

On-Policy Self-Distillation without Any Supervision Large Language Models Can Self-Improve

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.572940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.572940Z digest=sha256:e7556f94ee0a6959f2ce06acfa0d55dd9d0f1ce4ef14487feba730f9fb36071c

Observation a710c963-c998-47cc-9513-7d4135445f1a · outbound

This paper cites UniSD: Towards a Unified Self-Distillation Framework for Large Language Models.

On-Policy Self-Distillation without Any Supervision UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.728357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.728357Z digest=sha256:dad3b5f1e05ad3d90c7739c424b5067a5aa7e813db76de3f92d0d8e9d05fa4b6

Observation ae2517a6-bafe-45a7-9921-e2a36b8ff6e1 · outbound

This paper cites Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning.

On-Policy Self-Distillation without Any Supervision Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.784586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.784586Z digest=sha256:e16f267afc34fca135a76302ec1e1ddd5534993493c9f01d214d28e4d75b1c1b

Observation f2c04ddc-ed22-4a6b-8056-4cdb81fed65e · outbound

This paper cites Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models.

On-Policy Self-Distillation without Any Supervision Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.827630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.827630Z digest=sha256:4c19ca6e5fe419ae644cdc13bcd92ec150721b2c20ab1e448ddf27229759ae23

Observation d1caa6be-f65c-45c2-a1a6-d0a3c4949e81 · outbound

This paper cites Self-evolving visual questioner.arXiv preprint arXiv:2606.13929,.

On-Policy Self-Distillation without Any Supervision Self-evolving visual questioner.arXiv preprint arXiv:2606.13929,

Reference 13

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:19:44.930747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:41.853544Z digest=sha256:6933aa5755377a3567618f2e34e643d4e25837b13f1108a5f17c12bbeb463224

Observation cce98735-534e-4d7b-bec6-723f40a9ff5b · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,.

On-Policy Self-Distillation without Any Supervision Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.914010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.914010Z digest=sha256:b1611cde05be66e66301284847936cb06e1869c15ef1d35adba6611805006f3a

Observation dbbaa0a1-f975-4e45-b0a9-8fad67c2f673 · outbound

This paper cites HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation.

On-Policy Self-Distillation without Any Supervision HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:19:44.623767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:41.961817Z digest=sha256:7a5825a6d86059cb8ef52b96b905c80e0adf009ac04abcf29fa235f3aa731023

Observation 5409db2d-a287-478d-9367-4c9e5523aa91 · outbound

This paper cites URL https://thinkingmachines.ai/ blog/on-policy-distillation.

On-Policy Self-Distillation without Any Supervision URL https://thinkingmachines.ai/ blog/on-policy-distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.025433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.025433Z digest=sha256:b26ef77d0ab7801a36ca3d84d8d3e3925c77ccab86ce8c98aebcb8addb5e92f3

Observation 5ddeec39-3bf7-488d-a025-b484ce551768 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

On-Policy Self-Distillation without Any Supervision MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.100814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.100814Z digest=sha256:af375f4feed7b479dde204cc7a37840c603bcb250e1bee5ed0afe119dafe3adc

Observation fdd5552c-0f46-4046-b580-ea0c1db55611 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

On-Policy Self-Distillation without Any Supervision Maximizing Confidence Alone Improves Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.170600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.170600Z digest=sha256:5d56abaa5a4633ce1c1c5563e73c38dd7609a8c3399fbe30e5b0cc27cfdc2824

Observation 4caa0439-287f-4296-bd16-6aa4e452b899 · outbound

This paper cites Self-Consistency Preference Optimization.

On-Policy Self-Distillation without Any Supervision Self-Consistency Preference Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.234006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.234006Z digest=sha256:e6be5b88fbaf43bc7e5fa7d8dcc8fca1ef897a5c115d68850eac5ad3e9c463c3

Observation 7b608f62-0f24-4183-bf92-ecf1227a2398 · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

On-Policy Self-Distillation without Any Supervision CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.294503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.294503Z digest=sha256:49e5f26334f1f02e2c18db42d8a23f0ff1fae18545ca8fe004447a4aef8824a0

Observation 21f116ed-e599-4564-8d34-7d621f60020a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On-Policy Self-Distillation without Any Supervision DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.384243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.384243Z digest=sha256:d5b06b493790c1f13ba1ca849b6191d05d9e0c1323c207df813a715492b9e11a

Observation a9f0db98-ed67-46b7-8b0d-98bb297156c5 · outbound

This paper cites Self-Distillation Enables Continual Learning.

On-Policy Self-Distillation without Any Supervision Self-Distillation Enables Continual Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.449663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.449663Z digest=sha256:a022ddbaf2ce78681cc539575d8ab3580f218c8b3d2a3f622a7fb5afa2fbd43d

Observation d2ee794e-11d1-41ac-a219-3fda760934a1 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

On-Policy Self-Distillation without Any Supervision GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.513853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.513853Z digest=sha256:00a83234aaa41e5d850cf246bd07d762be3937d53602dbc2090e81c3085cd94d

Observation 2154647d-37e0-44a8-9cd8-93e4ab1e3cd2 · outbound

This paper cites Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations.

On-Policy Self-Distillation without Any Supervision Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:19:44.047486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:42.587810Z digest=sha256:804217557f565cbf09ad0ea7d60027042edc65a658eb1ac132ee7fecad666847

Observation 84d0f8aa-9ec0-4549-9c80-977a7ede8044 · outbound

This paper cites SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces.

On-Policy Self-Distillation without Any Supervision SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.700966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.700966Z digest=sha256:f1120d9756c07decc4d234bb9549eb6f0a3d5ae4315d623b58ec9da25b8a4db8

Observation 97b0ba77-58c6-42f0-9eea-bfb789c67920 · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

On-Policy Self-Distillation without Any Supervision Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.801108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.801108Z digest=sha256:619b4e0a3c23771d2a52aff2c95fb7bda0f73f7598213694c9725a953bd0dc64

Observation e5f4537e-e154-4c97-a652-cdf41624b582 · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

On-Policy Self-Distillation without Any Supervision Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.927212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.927212Z digest=sha256:e6411bcd071177d8ffba3be859c99206f6943293b8c874f9af8fa413cc5b02aa

Observation 2c84e049-02db-4bd7-b671-bc8815b22c4f · outbound

This paper cites Qwen3 Technical Report.

On-Policy Self-Distillation without Any Supervision Qwen3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.007032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.007032Z digest=sha256:bded68325161f7c06e73363ffb05b899a324c8faa91d943c21db1e782b7596b5

Observation 87977f5a-2953-48d6-aa83-22770c97d75b · outbound

This paper cites Snapshot Distillation: Teacher-Student Optimization in One Generation.

On-Policy Self-Distillation without Any Supervision Snapshot Distillation: Teacher-Student Optimization in One Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:19:44.305382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:43.092889Z digest=sha256:6df7ab2d02a7bdec34cd6a27e3c61aa9d15f418268607bb2b67e21ce5c845252

Observation cbaaed0d-6c3f-47f9-810d-868cc4bead47 · outbound

This paper cites On-Policy Context Distillation for Language Models.

On-Policy Self-Distillation without Any Supervision On-Policy Context Distillation for Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.250443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.250443Z digest=sha256:26df454c8aab47c6e7b01112d46621abc509bfa9e75ef5e2cf7ca4f706f10877

Observation bd14b0e8-7f9d-484b-841f-0b74a17ba5de · outbound

This paper cites Self-Rewarding Language Models.

On-Policy Self-Distillation without Any Supervision Self-Rewarding Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.347186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.347186Z digest=sha256:6d1036bd988c6881cf11989ad1c42aa6ac68ba3d09e1944e26e9fccb0c0bc9a9

Observation e37b18b6-0199-485d-863f-ee6df588524a · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

On-Policy Self-Distillation without Any Supervision STaR: Bootstrapping Reasoning With Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.516621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.516621Z digest=sha256:07ef223140108cf5ac5f17cd71d7c1fad4259a73b8298512c59369d04030e511

Observation 5f69390f-b506-483a-9d7d-9f611a7600fd · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

On-Policy Self-Distillation without Any Supervision Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.698887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.698887Z digest=sha256:50906433bb1b35617e418806c1f430cf906e189e76d49d1f14d496bb0019acdd

Observation 08631336-b861-4564-ab57-28cf45aba4d4 · outbound

This paper cites Learning to Reason without External Rewards.

On-Policy Self-Distillation without Any Supervision Learning to Reason without External Rewards

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.764127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.764127Z digest=sha256:955c3dd6bd48079eb41ca19d0c3507f0247bebad229e5d309acec1bfbf615781

Observation 3acd0707-5d9d-425b-82d2-f8b27eb2a3fd · outbound

This paper cites pub.” denotes the numbers published in the official OPSD repository; “ours.

On-Policy Self-Distillation without Any Supervision pub.” denotes the numbers published in the official OPSD repository; “ours

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:19:45.236912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:43.839213Z digest=sha256:33af2a7072f32ac123fc7c4a42f9baf537e9ec567e22603e922446db8ee021a2

Observation ceacb47f-efeb-4c0c-8e6a-932ff0750eef · outbound

This paper cites Self-Distilled RLVR.

On-Policy Self-Distillation without Any Supervision Self-Distilled RLVR

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.157443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.157443Z digest=sha256:3994ba983561b8e2a11880a87b942e53a4c95c7c6edaa0aad9a7520e7fe70db3

Observation 9cfbb828-d21e-44c6-8a5b-d32ece739d6f · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

On-Policy Self-Distillation without Any Supervision Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.621441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.621441Z digest=sha256:8591171f37607e83c5513fdc2e031adb2c5ddca542bd9e9e72fc8f52a61e6fe2

Observation 125cc2a1-4970-4931-8111-eb1d3c01d637 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

On-Policy Self-Distillation without Any Supervision R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.490605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.490605Z digest=sha256:b6b64a23245f5d06fbab2c83328dc867bcb598e203b370da7896448ddf160470

Observation 8b6d3f6b-7b5d-4a83-afb4-79b6cf5f0262 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

On-Policy Self-Distillation without Any Supervision Reinforcement Learning via Self-Distillation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.667277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.667277Z digest=sha256:61813bf503b4bb800d955ae0f9d07cb41c2e89a76c6ec181fa073097fe14ca6e

Observation 2a38042b-66e7-4535-a4a0-aa0bd7eafc88 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

On-Policy Self-Distillation without Any Supervision Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.431632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.431632Z digest=sha256:d55b6af66c0bf55950d2287214ca2568c0917f7abe71b136ab651c3d03092100

Observation 49934ccb-49e9-4533-8497-46dfbd014348 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On-Policy Self-Distillation without Any Supervision DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.249543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.249543Z digest=sha256:1d1e363851e034b71c55193ab3764932dd5afac35525d7e0460622a57d6bfa75

Observation 2a5413ea-f2e3-4e93-a015-1e7d2c3215b6 · outbound

This paper cites Serl: Self-play reinforcement learning for large language models with limited data.arXiv preprint arXiv:2505.20347,.

On-Policy Self-Distillation without Any Supervision Serl: Self-play reinforcement learning for large language models with limited data.arXiv preprint arXiv:2505.20347,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.277301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.277301Z digest=sha256:cf5d86bcc2fbef975e2c04c23b4f197b8e79d8402b1ff35b0450df70d4dffba7

Observation 2faf3dfb-1cba-4f8b-8eeb-6600154c7a29 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

On-Policy Self-Distillation without Any Supervision Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.381899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.381899Z digest=sha256:48eec39983967615a13211a211455c5e81e4f3fa5c983b637d75e77504b840b2

Pith citing papers

No inbound Pith citation observations are available.