Pith. sign in

Paper Citation Record · LEDGER

On-Policy Self-Distillation without Any Supervision

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.06296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06296 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:43.839213Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9d279ec-09f0-4ddb-9790-a4860180521e · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

On-Policy Self-Distillation without Any Supervision Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.326448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.326448Z digest=sha256:56e55dad88e709fb730377dbd06b8bed33ae168482ee2e1efc3002ac3ef9b87e

Observation 1b7637fa-51c4-4807-a2f1-43ef153e9df6 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

On-Policy Self-Distillation without Any Supervision MiniLLM: On-Policy Distillation of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.411447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.411447Z digest=sha256:0d7f42baacc2b9da185f7304d42bed2455bb9b2063f15164fa84f1dad8bf8164

Observation f9cb2373-0e40-4d27-8f14-1197802d626e · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

On-Policy Self-Distillation without Any Supervision OpenThoughts: Data Recipes for Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.444364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.444364Z digest=sha256:b40f8aa32485ed0de29fee11b5d591bb0b52afdb3da1fbfd3fcc4cde8ad92d30

Observation df86c3a3-12bb-4391-b060-2bf05ce0c6dc · outbound

This paper cites Large Language Models Can Self-Improve.

On-Policy Self-Distillation without Any Supervision Large Language Models Can Self-Improve

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.572940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.572940Z digest=sha256:329168332802ba0798a6f259dc44da99e0c55d9d44ffc0b0974217c7aa4adb8d

Observation a710c963-c998-47cc-9513-7d4135445f1a · outbound

This paper cites UniSD: Towards a Unified Self-Distillation Framework for Large Language Models.

On-Policy Self-Distillation without Any Supervision UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.728357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.728357Z digest=sha256:53dec79d455dcf47ea783815f82810478379af5e85d05a530b077c7d28996992

Observation ae2517a6-bafe-45a7-9921-e2a36b8ff6e1 · outbound

This paper cites Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning.

On-Policy Self-Distillation without Any Supervision Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.784586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.784586Z digest=sha256:3c07fd305bd817e8ef9a52c2bee9b50b9cc7537913accebb76121d3e12b51f65

Observation f2c04ddc-ed22-4a6b-8056-4cdb81fed65e · outbound

This paper cites Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models.

On-Policy Self-Distillation without Any Supervision Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.827630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.827630Z digest=sha256:5e046b2cd0e8017a007ed1ef883b4dfef4a0a3576ae35f40d19515fc7187f7db

Observation d1caa6be-f65c-45c2-a1a6-d0a3c4949e81 · outbound

This paper cites Self-evolving visual questioner.arXiv preprint arXiv:2606.13929,.

On-Policy Self-Distillation without Any Supervision Self-evolving visual questioner.arXiv preprint arXiv:2606.13929,

Reference 13

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:19:44.930747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:41.853544Z digest=sha256:61b20d113386398e08bf05135699ecdbc555a617c274d06d764d1a71697a8430

Observation cce98735-534e-4d7b-bec6-723f40a9ff5b · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,.

On-Policy Self-Distillation without Any Supervision Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.914010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.914010Z digest=sha256:ada6df19991fe850484e87d0369487c0c45573bb95a6476f0f511e02bd45fac1

Observation dbbaa0a1-f975-4e45-b0a9-8fad67c2f673 · outbound

This paper cites HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation.

On-Policy Self-Distillation without Any Supervision HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:19:44.623767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:41.961817Z digest=sha256:7aa7c5388afe6190fe88ceb0efeb2bf866b386aec143de0c699365886c562183

Observation 5409db2d-a287-478d-9367-4c9e5523aa91 · outbound

This paper cites URL https://thinkingmachines.ai/ blog/on-policy-distillation.

On-Policy Self-Distillation without Any Supervision URL https://thinkingmachines.ai/ blog/on-policy-distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.025433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.025433Z digest=sha256:640ef55a916838e53150cb17c1f344914bbd09ef8dd40a839817c291aa2cce2f

Observation 5ddeec39-3bf7-488d-a025-b484ce551768 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

On-Policy Self-Distillation without Any Supervision MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.100814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.100814Z digest=sha256:34f815b6876a5df4ee4dddec4683785ba0892400b67f6c0e5cb3a7bfb5f26edb

Observation fdd5552c-0f46-4046-b580-ea0c1db55611 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

On-Policy Self-Distillation without Any Supervision Maximizing Confidence Alone Improves Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.170600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.170600Z digest=sha256:c92a3b2fbe54c2d1f22859bea6718bfa8d53d5fbaec0e3a41f33443e20a4788a

Observation 4caa0439-287f-4296-bd16-6aa4e452b899 · outbound

This paper cites Self-Consistency Preference Optimization.

On-Policy Self-Distillation without Any Supervision Self-Consistency Preference Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.234006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.234006Z digest=sha256:6edb126a285e9ec9634583fbc6a56fd5acc4294adf3f4454084dea2d0e93ee6c

Observation 7b608f62-0f24-4183-bf92-ecf1227a2398 · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

On-Policy Self-Distillation without Any Supervision CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.294503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.294503Z digest=sha256:03a09755895760cc2ff737f47ab59315b8c4f4f7238bf9ed31fab36232ae33d3

Observation 21f116ed-e599-4564-8d34-7d621f60020a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On-Policy Self-Distillation without Any Supervision DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.384243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.384243Z digest=sha256:edf9403c51f71d4e285dda90e0983087af98c36e5acf54dfc73c3caade678902

Observation a9f0db98-ed67-46b7-8b0d-98bb297156c5 · outbound

This paper cites Self-Distillation Enables Continual Learning.

On-Policy Self-Distillation without Any Supervision Self-Distillation Enables Continual Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.449663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.449663Z digest=sha256:3a090a643864a14461c7cd5668ec487d56924c661bc84fb90b453a88798a5b35

Observation d2ee794e-11d1-41ac-a219-3fda760934a1 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

On-Policy Self-Distillation without Any Supervision GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.513853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.513853Z digest=sha256:473ca63a6b9e4c237f147e7255b31e977c1b6a703b0e606fe915d7a3bae9ea61

Observation 2154647d-37e0-44a8-9cd8-93e4ab1e3cd2 · outbound

This paper cites Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations.

On-Policy Self-Distillation without Any Supervision Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:19:44.047486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:42.587810Z digest=sha256:238e6561bdfdcb66e1f83a0dc2be226129c64025804a39cd1b066e765f17763a

Observation 84d0f8aa-9ec0-4549-9c80-977a7ede8044 · outbound

This paper cites SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces.

On-Policy Self-Distillation without Any Supervision SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.700966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.700966Z digest=sha256:67ea1ed2a4fabf0e06b5a010e86f7ed6178604c879f0900ec51f09a23e04404e

Observation 97b0ba77-58c6-42f0-9eea-bfb789c67920 · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

On-Policy Self-Distillation without Any Supervision Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.801108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.801108Z digest=sha256:618f71500877f7ef1b6e8c6e2ba66d424886fceafadc549ccfce50228aa8724b

Observation e5f4537e-e154-4c97-a652-cdf41624b582 · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

On-Policy Self-Distillation without Any Supervision Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.927212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.927212Z digest=sha256:485207c2b9afb93e1b90f031cbca3c2ed6d016d4bd97f1c0fa5543a50fc533f2

Observation 2c84e049-02db-4bd7-b671-bc8815b22c4f · outbound

This paper cites Qwen3 Technical Report.

On-Policy Self-Distillation without Any Supervision Qwen3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.007032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.007032Z digest=sha256:95e4cd0274244dd5c37709a0031ed546d4ab908eecfc67916d03348deed60e8f

Observation 87977f5a-2953-48d6-aa83-22770c97d75b · outbound

This paper cites Snapshot Distillation: Teacher-Student Optimization in One Generation.

On-Policy Self-Distillation without Any Supervision Snapshot Distillation: Teacher-Student Optimization in One Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:19:44.305382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:43.092889Z digest=sha256:5cc983bd82e1284057c4a058198308b32bcec5a35c75e5d8ad0ecde7877aac8d

Observation cbaaed0d-6c3f-47f9-810d-868cc4bead47 · outbound

This paper cites On-Policy Context Distillation for Language Models.

On-Policy Self-Distillation without Any Supervision On-Policy Context Distillation for Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.250443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.250443Z digest=sha256:8ff1fe8b5931edf976523ae36125b42e959094e3e18ce3f87d9ff85e2e22f81d

Observation bd14b0e8-7f9d-484b-841f-0b74a17ba5de · outbound

This paper cites Self-Rewarding Language Models.

On-Policy Self-Distillation without Any Supervision Self-Rewarding Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.347186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.347186Z digest=sha256:41e65062107bdda7a5e5953eaf94070279484167c80edb2baf556b2a5a97db7a

Observation e37b18b6-0199-485d-863f-ee6df588524a · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

On-Policy Self-Distillation without Any Supervision STaR: Bootstrapping Reasoning With Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.516621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.516621Z digest=sha256:886374c3d2387cbd5625fea960409637f1d19a48b725b7b3c67f18791e9a06e0

Observation 5f69390f-b506-483a-9d7d-9f611a7600fd · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

On-Policy Self-Distillation without Any Supervision Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.698887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.698887Z digest=sha256:1ca5a73e0b4f810e03b392a1127350df072728243bfdc214403cf417a6ee4fe7

Observation 08631336-b861-4564-ab57-28cf45aba4d4 · outbound

This paper cites Learning to Reason without External Rewards.

On-Policy Self-Distillation without Any Supervision Learning to Reason without External Rewards

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.764127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.764127Z digest=sha256:b0a4b706e4031d677a5ce03bf074d06eeb8643d2359e34b67feedbea5b604c0f

Observation 3acd0707-5d9d-425b-82d2-f8b27eb2a3fd · outbound

This paper cites pub.” denotes the numbers published in the official OPSD repository; “ours.

On-Policy Self-Distillation without Any Supervision pub.” denotes the numbers published in the official OPSD repository; “ours

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:19:45.236912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:43.839213Z digest=sha256:7f9518977333d9879e044ec830f0285db0864582f9e0deedffa76056ccabd88b

Observation ceacb47f-efeb-4c0c-8e6a-932ff0750eef · outbound

This paper cites Self-Distilled RLVR.

On-Policy Self-Distillation without Any Supervision Self-Distilled RLVR

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.157443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.157443Z digest=sha256:9c4d4f7a43658a8ef63723d6ea5ae00f890b3f8876f9a0b4812cf8dafddbd91f

Observation 9cfbb828-d21e-44c6-8a5b-d32ece739d6f · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

On-Policy Self-Distillation without Any Supervision Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.621441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.621441Z digest=sha256:2474f03b5843f8ca640be949cdb0de4fdac720f9e0e1b361423ba4d9cc5025b8

Observation 125cc2a1-4970-4931-8111-eb1d3c01d637 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

On-Policy Self-Distillation without Any Supervision R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.490605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.490605Z digest=sha256:a3b1897de57626633128b280d0775b89642645da3e5061e4de5d0da579fe144a

Observation 8b6d3f6b-7b5d-4a83-afb4-79b6cf5f0262 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

On-Policy Self-Distillation without Any Supervision Reinforcement Learning via Self-Distillation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.667277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.667277Z digest=sha256:1ac52acaa7f304f80142ca5c9dc8cd1d6a309631296a9386661e8885cffbad80

Observation 2a38042b-66e7-4535-a4a0-aa0bd7eafc88 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

On-Policy Self-Distillation without Any Supervision Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.431632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.431632Z digest=sha256:4c21dc3d03d126ed04b46a7ba08e302d37c862fb46f1216f8cf8a479200b6bdc

Observation 49934ccb-49e9-4533-8497-46dfbd014348 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On-Policy Self-Distillation without Any Supervision DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.249543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.249543Z digest=sha256:7f8e56cacb4ec16e82a64e8949316898bc0c551227f9d045462d6e7281e20ff2

Observation 2a5413ea-f2e3-4e93-a015-1e7d2c3215b6 · outbound

This paper cites Serl: Self-play reinforcement learning for large language models with limited data.arXiv preprint arXiv:2505.20347,.

On-Policy Self-Distillation without Any Supervision Serl: Self-play reinforcement learning for large language models with limited data.arXiv preprint arXiv:2505.20347,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.277301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.277301Z digest=sha256:d9f718e831e2a014dffe1a40a43c6a1bc689982803c4cf3fc9fa816bf331e3c9

Observation 2faf3dfb-1cba-4f8b-8eeb-6600154c7a29 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

On-Policy Self-Distillation without Any Supervision Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.381899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.381899Z digest=sha256:b8d3b9347f8e2c2b3ad13058f57cf4018a9d24ae32bcf73cfaf8383128a426dc

Pith citing papers

No inbound Pith citation observations are available.