Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-14T21:11:56.772993Z
Paper Citation Record · LEDGER
As of 2 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2605.09725.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-14T21:11:56.772993Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T13:40:42.016315Z
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 33b5251b-b9ed-413b-b16a-015f9e4f471f · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection On-policy distillation of language models: Learning from self-generated mistakes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d4b2564c-a910-452c-8048-864f85a82aa8 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5d60505c-8aa4-432e-b377-3647b4a590bf · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Scheduled sampling for sequence prediction with recurrent neural networks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation fdf8c78b-9477-4dab-b4bb-f07ffaab32fb · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a0f870eb-4da2-4ac0-9b55-a932063805b6 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation be5d799d-5d2b-4562-bf22-54c2b5517c1f · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation aaa1408b-e8ff-4a8c-b606-b54a56cde74a · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Hdpo: Hybrid distillation policy optimization via privileged self-distillation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 08bafd69-3abd-41d8-b64e-2556d68ba3cc · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection RAFT: Reward ranked finetuning for generative foundation model alignment
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 31ed2a02-1463-4d64-9de4-ea2aeb46c4e1 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Specializing Smaller Language Models towards Multi-Step Reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a1000f0d-0a14-455b-ac56-1b39e35bca3e · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e0d94931-438d-450c-a7e5-515c4fbb9486 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection GLM-5: from Vibe Coding to Agentic Engineering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation eb968ee7-51d5-40eb-9385-12f68860475f · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Maybank, and Dacheng Tao
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3e47c3de-9a3a-4ba3-ad9f-aba08a3b021c · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection MiniLLM: Knowledge distillation of large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6faa0da4-2ae6-42ae-b052-dd789a0f3894 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection OpenThoughts: Data Recipes for Reasoning Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7321ec7c-a00b-4c42-8dd5-a7c8b2c3e62b · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning.Nature
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3c4291d1-aa3b-4f78-bf7b-5906eb689fee · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Justrl: Scaling a 1.5 b llm with a simple rl recipe
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d9a85f44-3721-4a43-809f-27cb519d4743 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Distilling the Knowledge in a Neural Network
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation fbf8e0a0-ebaf-436e-ad9c-cfe9ef495edb · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Reinforcement Learning via Self-Distillation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8f6ff3d9-e8cd-4b22-b58c-907c90dcb7b0 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Stable On-Policy Distillation through Adaptive Target Reformulation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a4e6f081-93e2-47b8-adb2-c1e669a2d710 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection TinyBERT: Distilling BERT for natural language understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1ffdee76-9c32-48b7-95b0-0b7b6ac5ffe3 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Entropy-Aware On-Policy Distillation of Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2bd05964-0d53-4e10-8ae4-73fb5830aced · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b3da5ea7-7c45-4f1e-a604-5728500af02c · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 51353a3b-b36b-4e3e-acaa-0d792f225525 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d20c781f-ed39-442f-b7bf-9457c77d025f · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f75542ef-3cfd-4640-a639-6fc78f2763c8 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1e2d0c86-7104-4a32-b3b8-b56511d20b44 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Small models struggle to learn from strong reasoners
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0f5071df-7e60-4ac8-a032-a9b875b8d573 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Let's Verify Step by Step
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cf5de36d-b2a4-4ac3-8e58-72852845a495 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection On-policy distillation.Thinking Machines Lab: Con- nectionism
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6ecf00bc-45b2-4ce9-9923-355168d3c607 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection An empirical study of catastrophic forgetting in large language models during continual fine-tuning.IEEE Transactions on Audio, Speech and Language Processing
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 46233be4-8fb6-4d81-8051-7a67a0a6c608 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection WebGPT: Browser-assisted question-answering with human feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7607ee17-05e6-4070-a640-f9181d7464e0 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Privileged information distillation for language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation af8bb8c4-2e4d-44d4-8849-77ec9b29d380 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection A reduction of imitation learning and structured prediction to no-regret online learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5c0db343-f418-436b-9f4a-15b2d61bc7a7 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cf909211-e443-42b4-8c4b-8088f68f8a32 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 07b27394-78b4-476b-949b-e0364a2cffed · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection RL's Razor: Why Online Reinforcement Learning Forgets Less
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 109b6a53-f35b-4f9d-b93b-e5ef755de091 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Self-Distillation Enables Continual Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 23bee9d7-64e5-46b4-b0f1-a43e959297b1 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 35fd46a9-9caa-4fad-9e68-aea5ccb063f4 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Learning by distilling context
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bd4969d8-fc64-4f4c-90e2-922ab5535483 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Christiano
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7afd1895-a95c-4d93-9f0d-f914aceb00af · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.Advances in Neural Information Processing Systems, 33:5776–5788
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ec52dde5-9128-48f3-aea1-79bf27499118 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 567179b8-47d0-40c6-982d-d7a19a377aa0 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection MiMo-V2-Flash Technical Report
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8b1e0ad3-3003-499d-a51a-b549ae43c5da · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Error bounds of imitating policies and environments.Advances in Neural Information Processing Systems, 33
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9d52d55c-7ef5-4a13-8cf3-af8594b7e0d7 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 522fe48f-c223-4eea-86f3-db5da151c74c · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Qwen3 Technical Report
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e7f76d52-484f-42dc-9cb6-152434a7aa1c · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Self-Distilled RLVR
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 46470e9e-6002-4ca9-9e4f-cb5869d8ab62 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 228aa3df-a52e-41bb-abb0-c269f54eb743 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Black-box on-policy distillation of large language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3fad5a0f-d677-486c-93ef-2e516987934e · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection On-Policy Context Distillation for Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e72c8497-b433-41bd-9236-9e8949d68498 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 48e88f44-d4e1-4f9b-84a7-b152916b69a1 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection STaR: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1673e108-308c-4548-9848-f1474f88a042 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Towards the law of capacity gap in distilling language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 20edf877-2e00-4854-9165-80666f074963 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2b944b22-8b4b-47a3-8d5a-25b1309cd2b8 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f50da723-c575-4651-b6ef-065fbe0cf00a · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8c1777fe-d217-43b6-bae1-4ce9d2543397 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2c76b486-ae0a-4391-8799-059784ea2247 · outbound
On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 38755796-b319-4f5e-9c80-a0dbf5a873ee · inbound
Contrastive On-Policy Distillation On-Policy Distillation with Best-of-N Teacher Rollout Selection
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.