Pith. sign in

Paper Citation Record · LEDGER

On-Policy Distillation with Best-of-N Teacher Rollout Selection

As of 2 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2605.09725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09725 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T21:11:56.772993Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:40:42.016315Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact28
  • verified fuzzy22
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33b5251b-b9ed-413b-b16a-015f9e4f471f · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

On-Policy Distillation with Best-of-N Teacher Rollout Selection On-policy distillation of language models: Learning from self-generated mistakes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.453644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:bb1e13421f57580f8102c45e4ecdbb07d94d3ef871359218d9266d7dd93d11e9

Observation d4b2564c-a910-452c-8048-864f85a82aa8 · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

On-Policy Distillation with Best-of-N Teacher Rollout Selection MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:10:14.938194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:742ac9c2404c15462572c9934eab4fa7db06f034042401ed2751cd169010ce0c

Observation 5d60505c-8aa4-432e-b377-3647b4a590bf · outbound

This paper cites Scheduled sampling for sequence prediction with recurrent neural networks.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Scheduled sampling for sequence prediction with recurrent neural networks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.468997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:2b19413476da8aa446be95ea9d5cbfdffd8f7478b31a2361821642abc0c5ee9f

Observation fdf8c78b-9477-4dab-b4bb-f07ffaab32fb · outbound

This paper cites Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:22.491699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:25e2efdda361bbc55877a0df6aabe3c26750455b03b7d1fb6fed7a9578b0f153

Observation a0f870eb-4da2-4ac0-9b55-a932063805b6 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

On-Policy Distillation with Best-of-N Teacher Rollout Selection SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T21:12:58.996247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:018ec6ec25cdfcb06d1bcb8a917fed38a16cc945a9f1e492602c9c8c8ea2a719

Observation be5d799d-5d2b-4562-bf22-54c2b5517c1f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.029572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:f4bdd7d025cc4cf1cdb99b5d32a39f22ef5df7ec2e101ce0fe244b7a4beb10f9

Observation aaa1408b-e8ff-4a8c-b606-b54a56cde74a · outbound

This paper cites Hdpo: Hybrid distillation policy optimization via privileged self-distillation.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Hdpo: Hybrid distillation policy optimization via privileged self-distillation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:59.015767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:6a8a1beff4cc604b752155cb6e5fcddb36c4a9989c392fc572403d3d5f987841

Observation 08bafd69-3abd-41d8-b64e-2556d68ba3cc · outbound

This paper cites RAFT: Reward ranked finetuning for generative foundation model alignment.

On-Policy Distillation with Best-of-N Teacher Rollout Selection RAFT: Reward ranked finetuning for generative foundation model alignment

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.491279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:15a456765c498bcd7f9722707b4da3bccb9b959d161eb68d10084e8f4f7a76fd

Observation 31ed2a02-1463-4d64-9de4-ea2aeb46c4e1 · outbound

This paper cites Specializing Smaller Language Models towards Multi-Step Reasoning.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Specializing Smaller Language Models towards Multi-Step Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.989926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:4a175b4719f95ac2f10d93cb451be591f8ba9631035e77d7f5c7c39560fa22fd

Observation a1000f0d-0a14-455b-ac56-1b39e35bca3e · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.069513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:3d8842338c3abe457de079ee9c0894eb6987786f003a3bd23abe8ab2cce9722d

Observation e0d94931-438d-450c-a7e5-515c4fbb9486 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

On-Policy Distillation with Best-of-N Teacher Rollout Selection GLM-5: from Vibe Coding to Agentic Engineering

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.054182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:c7b092a39d1923f9a856a08197cf33c03d0406a8861ba1663afe4b90fd7251c4

Observation eb968ee7-51d5-40eb-9385-12f68860475f · outbound

This paper cites Maybank, and Dacheng Tao.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Maybank, and Dacheng Tao

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.488607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:66c7fc7f490df3189b2be805425dbde4084bef3fc6dd35468719f5caa1693cdb

Observation 3e47c3de-9a3a-4ba3-ad9f-aba08a3b021c · outbound

This paper cites MiniLLM: Knowledge distillation of large language models.

On-Policy Distillation with Best-of-N Teacher Rollout Selection MiniLLM: Knowledge distillation of large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.498994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:6d5db50fcd50b73d2d06643b69d9febe646194b5e86cb9eb0addbc5c31b52c24

Observation 6faa0da4-2ae6-42ae-b052-dd789a0f3894 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

On-Policy Distillation with Best-of-N Teacher Rollout Selection OpenThoughts: Data Recipes for Reasoning Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.063295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:afc55f7b14509f6b6e48fe06b4cd356dc122ecb431fb54fcf6e8836885c3fb05

Observation 7321ec7c-a00b-4c42-8dd5-a7c8b2c3e62b · outbound

This paper cites DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning.Nature.

On-Policy Distillation with Best-of-N Teacher Rollout Selection DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning.Nature

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.518215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:3bfc73e8be36f1873f3c36a6ddb85b3268b6ffb4ab6c135746c307ce0a3d3e8a

Observation 3c4291d1-aa3b-4f78-bf7b-5906eb689fee · outbound

This paper cites Justrl: Scaling a 1.5 b llm with a simple rl recipe.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Justrl: Scaling a 1.5 b llm with a simple rl recipe

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:59.051417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:902d1140c5661cad2c27b4a083e8e26772842dd1a70194943f7c5440611e4fcb

Observation d9a85f44-3721-4a43-809f-27cb519d4743 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Distilling the Knowledge in a Neural Network

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.060394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:6c4c6a3031ff45cca59090d63f92ee68505839f049084bd01e652d1cc6fd62e5

Observation fbf8e0a0-ebaf-436e-ad9c-cfe9ef495edb · outbound

This paper cites Reinforcement Learning via Self-Distillation.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Reinforcement Learning via Self-Distillation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.036053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:7052a154a49e972ea8adb692db34adf6d6d7e1986bac76829b7a1035d3f07bda

Observation 8f6ff3d9-e8cd-4b22-b58c-907c90dcb7b0 · outbound

This paper cites Stable On-Policy Distillation through Adaptive Target Reformulation.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Stable On-Policy Distillation through Adaptive Target Reformulation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.026355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:74443e39ecbb6b824c8ef38ef7498a2756a38ebd1c1e35a0ba2f2eed6c643565

Observation a4e6f081-93e2-47b8-adb2-c1e669a2d710 · outbound

This paper cites TinyBERT: Distilling BERT for natural language understanding.

On-Policy Distillation with Best-of-N Teacher Rollout Selection TinyBERT: Distilling BERT for natural language understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.486531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:9684f2bb1c0265488b325a0e6f8730a2239421f29752bb808964396719266369

Observation 1ffdee76-9c32-48b7-95b0-0b7b6ac5ffe3 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Entropy-Aware On-Policy Distillation of Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:02:00.585154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:7c3595c4d71762a24c9e93a31bbe8c6eb636649328dc1dc0c0d3ce4a9815030e

Observation 2bd05964-0d53-4e10-8ae4-73fb5830aced · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.057387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:19e07bfb6f87c5eb017a7ad57813a339561319686c09ce9bbc61a24bdcb66b4b

Observation b3da5ea7-7c45-4f1e-a604-5728500af02c · outbound

This paper cites an unresolved cited work.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:10:34.481842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:abe6e87f061eec4c5fd711f133cd9a1e305f0a11d1ce3f8f4210507e0bec256e

Observation 51353a3b-b36b-4e3e-acaa-0d792f225525 · outbound

This paper cites Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:12:59.048163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:7add9df95e75c5eeec240d559d1e71f42df55cca3de087001062b33cbd3ae4e5

Observation d20c781f-ed39-442f-b7bf-9457c77d025f · outbound

This paper cites Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.484071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:7b3483198c9ec8af881ae173e42c70c938e53fd580e05a6b0f07cf77c7da58ec

Observation f75542ef-3cfd-4640-a639-6fc78f2763c8 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.044937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:992691e701b969ee2148f1fbed77d401c590caa9b623188b7a4b2d46c1698336

Observation 1e2d0c86-7104-4a32-b3b8-b56511d20b44 · outbound

This paper cites Small models struggle to learn from strong reasoners.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Small models struggle to learn from strong reasoners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.493756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:e9eb88dfe5c30ca8d8b4179affa29ff0d656d668c428921c8b2e467bb16ce8f0

Observation 0f5071df-7e60-4ac8-a032-a9b875b8d573 · outbound

This paper cites Let's Verify Step by Step.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Let's Verify Step by Step

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.032791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:44333a621e3b486e3b0198229fdaf92fff7a648308ee717c619159b353a237d7

Observation cf5de36d-b2a4-4ac3-8e58-72852845a495 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Con- nectionism.

On-Policy Distillation with Best-of-N Teacher Rollout Selection On-policy distillation.Thinking Machines Lab: Con- nectionism

Reference 29

Resolution
verified exact
doi, observed 2026-05-14T21:12:57.794166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:a665a6c41da2e61346cdde9caafee6aab21a066fab86a4e800f6a630af1dad14

Observation 6ecf00bc-45b2-4ce9-9923-355168d3c607 · outbound

This paper cites An empirical study of catastrophic forgetting in large language models during continual fine-tuning.IEEE Transactions on Audio, Speech and Language Processing.

On-Policy Distillation with Best-of-N Teacher Rollout Selection An empirical study of catastrophic forgetting in large language models during continual fine-tuning.IEEE Transactions on Audio, Speech and Language Processing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.496352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:2fbc7de6e362c10ac0b2998ff77018c6cb6cd5a13093d4ad127495317dc3dcfd

Observation 46233be4-8fb6-4d81-8051-7a67a0a6c608 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

On-Policy Distillation with Best-of-N Teacher Rollout Selection WebGPT: Browser-assisted question-answering with human feedback

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:58.999913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:4dd44e5e4edf3514d9af493ed7d36ef741fdb5c920eb1f8d563a0fa4cfa5965b

Observation 7607ee17-05e6-4070-a640-f9181d7464e0 · outbound

This paper cites Privileged information distillation for language models.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Privileged information distillation for language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.476754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:53ce2a6af68e4e3fbb91d30e8ae27d377b5e9c0b925d13185f665f19bc91a221

Observation af8bb8c4-2e4d-44d4-8849-77ec9b29d380 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

On-Policy Distillation with Best-of-N Teacher Rollout Selection A reduction of imitation learning and structured prediction to no-regret online learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.461044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:a743b57af94c5e081bbe5a4d06ce91ff664e1ccdfdb8f19446fac7b095d3a501

Observation 5c0db343-f418-436b-9f4a-15b2d61bc7a7 · outbound

This paper cites DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter.

On-Policy Distillation with Best-of-N Teacher Rollout Selection DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.512289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:984864a6629e1c68947ba9dc7cfb89e53694ae963e62dbfe99c1a427e58f13b9

Observation cf909211-e443-42b4-8c4b-8088f68f8a32 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On-Policy Distillation with Best-of-N Teacher Rollout Selection DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:58.981301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:c2cfbf7b4ae041c4b7dc240407e44d804f634cba98b9c44f5a5bba9c93826411

Observation 07b27394-78b4-476b-949b-e0364a2cffed · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

On-Policy Distillation with Best-of-N Teacher Rollout Selection RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:03:54.455790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:321ed45b9587156216f7ea8033e77ec8b6f383b4dec1d760092a365557fb9a7f

Observation 109b6a53-f35b-4f9d-b93b-e5ef755de091 · outbound

This paper cites Self-Distillation Enables Continual Learning.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Self-Distillation Enables Continual Learning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:58.992950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:f671ff1b264d418902bdc7b2f2d143bdc0b16ee3276e5b2a9ff1b5e54f58cfe7

Observation 23bee9d7-64e5-46b4-b0f1-a43e959297b1 · outbound

This paper cites Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.503855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:0d7cdf907da0d9c6b2a56a8a23f892ba282a640b247b47199b3e34382f09f038

Observation 35fd46a9-9caa-4fad-9e68-aea5ccb063f4 · outbound

This paper cites Learning by distilling context.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Learning by distilling context

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.471634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:838624a952da45a7c94629d07f047feef0e30ac3b5bb480c7437ace6aaf3f824

Observation bd4969d8-fc64-4f4c-90e2-922ab5535483 · outbound

This paper cites Christiano.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Christiano

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.515144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:e2ddeeb5225aa2192a91781e026e0c3d0ece0d797645bc911937b75825f31131

Observation 7afd1895-a95c-4d93-9f0d-f914aceb00af · outbound

This paper cites MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.Advances in Neural Information Processing Systems, 33:5776–5788.

On-Policy Distillation with Best-of-N Teacher Rollout Selection MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.Advances in Neural Information Processing Systems, 33:5776–5788

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.466058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:c6ccd820b1b08616f56cb68888c87eecc2dabfc1149ef420b20c0a34550b7927

Observation ec52dde5-9128-48f3-aea1-79bf27499118 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.509242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:335476f5664d018744fb1447a497d3b5c33d32decd714c32f000f9de06bc57a2

Observation 567179b8-47d0-40c6-982d-d7a19a377aa0 · outbound

This paper cites MiMo-V2-Flash Technical Report.

On-Policy Distillation with Best-of-N Teacher Rollout Selection MiMo-V2-Flash Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.042170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:23426710d6814c5b41fe9cd6e214621d4805dbb2fc54204cb6add425137c48c5

Observation 8b1e0ad3-3003-499d-a51a-b549ae43c5da · outbound

This paper cites Error bounds of imitating policies and environments.Advances in Neural Information Processing Systems, 33.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Error bounds of imitating policies and environments.Advances in Neural Information Processing Systems, 33

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.506494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:cd7843126a40e0612bc6e6a37785cf4a06abac450fc0da8a03607f9e24746483

Observation 9d52d55c-7ef5-4a13-8cf3-af8594b7e0d7 · outbound

This paper cites Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:59.012162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:b30869a9ff9a8df8325b6848df0ca318bba07a6f51815f0077833e495c5a57d0

Observation 522fe48f-c223-4eea-86f3-db5da151c74c · outbound

This paper cites Qwen3 Technical Report.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Qwen3 Technical Report

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.003858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:f82a774b4aa81d4d3afcbc9c93ad399c730410a6bfe505e62e57fae35e15bcde

Observation e7f76d52-484f-42dc-9cb6-152434a7aa1c · outbound

This paper cites Self-Distilled RLVR.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Self-Distilled RLVR

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.019206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:c26c4a1011e0ea1f47dab1991d5461d61b90edf3540401c200ee1735556f5a24

Observation 46470e9e-6002-4ca9-9e4f-cb5869d8ab62 · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:11:07.754859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:a74ceaada817cf0f6c39f815741fa934664d8e4a5b24aebeaeac04e72623d203

Observation 228aa3df-a52e-41bb-abb0-c269f54eb743 · outbound

This paper cites Black-box on-policy distillation of large language models.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Black-box on-policy distillation of large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.458443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:2927600496dcd06f129b7bc6dd4cb9853648624566d20abbda8a06865b0bf803

Observation 3fad5a0f-d677-486c-93ef-2e516987934e · outbound

This paper cites On-Policy Context Distillation for Language Models.

On-Policy Distillation with Best-of-N Teacher Rollout Selection On-Policy Context Distillation for Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:58.986825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:2c235fa7746a81e0e75efe0563a19a1374afb763b990fe6168c3ec2a8770f9aa

Observation e72c8497-b433-41bd-9236-9e8949d68498 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

On-Policy Distillation with Best-of-N Teacher Rollout Selection DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:58.984096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:502513e5ece4c8a4f1e2b46278bb7a9a762c2dafa117b60c98e8f38d59e1bee5

Observation 48e88f44-d4e1-4f9b-84a7-b152916b69a1 · outbound

This paper cites STaR: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems.

On-Policy Distillation with Best-of-N Teacher Rollout Selection STaR: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.463431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:29dfd6e991c55bd36ee97fe7a2163f182be139ad7c9be55eb08a6d80adcde35d

Observation 1673e108-308c-4548-9848-f1474f88a042 · outbound

This paper cites Towards the law of capacity gap in distilling language models.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Towards the law of capacity gap in distilling language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:10:34.479575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:a2600b0311c111b13be2cffb193b4739c1e86d8c1c4fb91cffa3006aedc80a0b

Observation 20edf877-2e00-4854-9165-80666f074963 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 55

Resolution
malformed identifier
local_arxiv, observed 2026-05-14T21:12:59.072341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:65ac0c167e01498301e03664349b0437874bc005bd23b855674ee3d07d4e02a1

Observation 2b944b22-8b4b-47a3-8d5a-25b1309cd2b8 · outbound

This paper cites an unresolved cited work.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:10:34.521046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:5681d09d6d8bcf94f02071c5beae96d01e13c41eafdc61a3f8ac18a93bc42dcb

Observation f50da723-c575-4651-b6ef-065fbe0cf00a · outbound

This paper cites an unresolved cited work.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:10:34.501406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:db9a0193e4151361ff7e2e37b917a675442ac52c36ec816c242c3501d36f90b5

Observation 8c1777fe-d217-43b6-bae1-4ce9d2543397 · outbound

This paper cites an unresolved cited work.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:10:34.474296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:6c5e6613665b73a3ed0a6f0ed38215cee88e4347663fa323e2d795780819b1bb

Observation 2c76b486-ae0a-4391-8799-059784ea2247 · outbound

This paper cites an unresolved cited work.

On-Policy Distillation with Best-of-N Teacher Rollout Selection Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:10:34.455957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T21:11:56.772993Z digest=sha256:16085b93777a7a3767b161de956b5ae77cdac7fdf2dd89bebfdf0fc7e75b570e

Pith citing papers

Observation 38755796-b319-4f5e-9c80-a0dbf5a873ee · inbound

Contrastive On-Policy Distillation cites this paper.

Contrastive On-Policy Distillation On-Policy Distillation with Best-of-N Teacher Rollout Selection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:42.016315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:42.016315Z digest=sha256:b8daf5aa34becf0332cb515c7bc39497314af7e1eb25d31deb829719b0f2e874