Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:43.839213Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.06296.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:43.839213Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f9d279ec-09f0-4ddb-9790-a4860180521e · outbound
On-Policy Self-Distillation without Any Supervision Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b7637fa-51c4-4807-a2f1-43ef153e9df6 · outbound
On-Policy Self-Distillation without Any Supervision MiniLLM: On-Policy Distillation of Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9cb2373-0e40-4d27-8f14-1197802d626e · outbound
On-Policy Self-Distillation without Any Supervision OpenThoughts: Data Recipes for Reasoning Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df86c3a3-12bb-4391-b060-2bf05ce0c6dc · outbound
On-Policy Self-Distillation without Any Supervision Large Language Models Can Self-Improve
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a710c963-c998-47cc-9513-7d4135445f1a · outbound
On-Policy Self-Distillation without Any Supervision UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae2517a6-bafe-45a7-9921-e2a36b8ff6e1 · outbound
On-Policy Self-Distillation without Any Supervision Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c04ddc-ed22-4a6b-8056-4cdb81fed65e · outbound
On-Policy Self-Distillation without Any Supervision Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1caa6be-f65c-45c2-a1a6-d0a3c4949e81 · outbound
On-Policy Self-Distillation without Any Supervision Self-evolving visual questioner.arXiv preprint arXiv:2606.13929,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cce98735-534e-4d7b-bec6-723f40a9ff5b · outbound
On-Policy Self-Distillation without Any Supervision Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbbaa0a1-f975-4e45-b0a9-8fad67c2f673 · outbound
On-Policy Self-Distillation without Any Supervision HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5409db2d-a287-478d-9367-4c9e5523aa91 · outbound
On-Policy Self-Distillation without Any Supervision URL https://thinkingmachines.ai/ blog/on-policy-distillation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ddeec39-3bf7-488d-a025-b484ce551768 · outbound
On-Policy Self-Distillation without Any Supervision MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd5552c-0f46-4046-b580-ea0c1db55611 · outbound
On-Policy Self-Distillation without Any Supervision Maximizing Confidence Alone Improves Reasoning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4caa0439-287f-4296-bd16-6aa4e452b899 · outbound
On-Policy Self-Distillation without Any Supervision Self-Consistency Preference Optimization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b608f62-0f24-4183-bf92-ecf1227a2398 · outbound
On-Policy Self-Distillation without Any Supervision CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f116ed-e599-4564-8d34-7d621f60020a · outbound
On-Policy Self-Distillation without Any Supervision DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f0db98-ed67-46b7-8b0d-98bb297156c5 · outbound
On-Policy Self-Distillation without Any Supervision Self-Distillation Enables Continual Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ee794e-11d1-41ac-a219-3fda760934a1 · outbound
On-Policy Self-Distillation without Any Supervision GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2154647d-37e0-44a8-9cd8-93e4ab1e3cd2 · outbound
On-Policy Self-Distillation without Any Supervision Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84d0f8aa-9ec0-4549-9c80-977a7ede8044 · outbound
On-Policy Self-Distillation without Any Supervision SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97b0ba77-58c6-42f0-9eea-bfb789c67920 · outbound
On-Policy Self-Distillation without Any Supervision Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5f4537e-e154-4c97-a652-cdf41624b582 · outbound
On-Policy Self-Distillation without Any Supervision Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c84e049-02db-4bd7-b671-bc8815b22c4f · outbound
On-Policy Self-Distillation without Any Supervision Qwen3 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87977f5a-2953-48d6-aa83-22770c97d75b · outbound
On-Policy Self-Distillation without Any Supervision Snapshot Distillation: Teacher-Student Optimization in One Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cbaaed0d-6c3f-47f9-810d-868cc4bead47 · outbound
On-Policy Self-Distillation without Any Supervision On-Policy Context Distillation for Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd14b0e8-7f9d-484b-841f-0b74a17ba5de · outbound
On-Policy Self-Distillation without Any Supervision Self-Rewarding Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e37b18b6-0199-485d-863f-ee6df588524a · outbound
On-Policy Self-Distillation without Any Supervision STaR: Bootstrapping Reasoning With Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f69390f-b506-483a-9d7d-9f611a7600fd · outbound
On-Policy Self-Distillation without Any Supervision Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08631336-b861-4564-ab57-28cf45aba4d4 · outbound
On-Policy Self-Distillation without Any Supervision Learning to Reason without External Rewards
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3acd0707-5d9d-425b-82d2-f8b27eb2a3fd · outbound
On-Policy Self-Distillation without Any Supervision pub.” denotes the numbers published in the official OPSD repository; “ours
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ceacb47f-efeb-4c0c-8e6a-932ff0750eef · outbound
On-Policy Self-Distillation without Any Supervision Self-Distilled RLVR
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfbb828-d21e-44c6-8a5b-d32ece739d6f · outbound
On-Policy Self-Distillation without Any Supervision Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125cc2a1-4970-4931-8111-eb1d3c01d637 · outbound
On-Policy Self-Distillation without Any Supervision R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b6d3f6b-7b5d-4a83-afb4-79b6cf5f0262 · outbound
On-Policy Self-Distillation without Any Supervision Reinforcement Learning via Self-Distillation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a38042b-66e7-4535-a4a0-aa0bd7eafc88 · outbound
On-Policy Self-Distillation without Any Supervision Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49934ccb-49e9-4533-8497-46dfbd014348 · outbound
On-Policy Self-Distillation without Any Supervision DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5413ea-f2e3-4e93-a015-1e7d2c3215b6 · outbound
On-Policy Self-Distillation without Any Supervision Serl: Self-play reinforcement learning for large language models with limited data.arXiv preprint arXiv:2505.20347,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2faf3dfb-1cba-4f8b-8eeb-6600154c7a29 · outbound
On-Policy Self-Distillation without Any Supervision Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.