Pith. sign in

Paper Citation Record · LEDGER

Contrastive On-Policy Distillation

As of 19 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2607.19046.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19046 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:40:42.303146Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T02:56:58.020990Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30875ada-ecfa-400f-9ba8-660e9f1730b1 · outbound

This paper cites On-policy distillation of language models: Learning from self- generated mistakes.

Contrastive On-Policy Distillation On-policy distillation of language models: Learning from self- generated mistakes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:38.307846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:38.307846Z digest=sha256:d3465dd8d1a837b244b63a983077ac2d02da59f436702fa0ebe00428463ffa96

Observation 8c49b3ca-de7a-4177-aff0-e6af4e1790fd · outbound

This paper cites Token- budget-aware llm reasoning.

Contrastive On-Policy Distillation Token- budget-aware llm reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:38.846295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:38.846295Z digest=sha256:2048bcb089f51c7a674aca099302a20e6fa9bd7a2127efba867aadf7fc29fc31

Observation 903893c7-1265-44bf-8876-72235cb4514b · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Contrastive On-Policy Distillation Distilling the Knowledge in a Neural Network

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:38.957682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:38.957682Z digest=sha256:944aa819ec96f3b2877a803af292457827872befcfaa9cb75f19f6daefc7f8f6

Observation f1406273-a49b-48e4-8b0c-b97a6a496546 · outbound

This paper cites Towards reasoning in large language models: A survey.

Contrastive On-Policy Distillation Towards reasoning in large language models: A survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:39.315192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:39.315192Z digest=sha256:3797e9a48cbbfeae4d84f562292f65a8663e523277322460575f565b3b12a05f

Observation 7d1845b8-5392-4db0-9e63-e9c60616a0c9 · outbound

This paper cites Tinybert: Distilling bert for natural language understanding.

Contrastive On-Policy Distillation Tinybert: Distilling bert for natural language understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:39.479014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:39.479014Z digest=sha256:ed64167145f708bc024eb34dd58dd1f13f9ee7549fa63266f5e08f5b87aad475

Observation d58fab11-0a67-4879-9243-c240d0895db4 · outbound

This paper cites ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization.

Contrastive On-Policy Distillation ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:39.777071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:39.777071Z digest=sha256:a88e99cda296a09253e6982d96bc89be4388504fa33443714b07aad3a425ed0b

Observation 5c5f1829-1eee-49db-86f4-8c17142f2de1 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Contrastive On-Policy Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:39.884234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:39.884234Z digest=sha256:365d9730bf3fede384b093945e85e443e31b68b7865454a68e36cdecca7ac2bb

Observation 57c1594c-7118-48e8-aeb2-cc80aa204c6a · outbound

This paper cites Visual-Advantage On-Policy Distillation for Vision-Language Models.

Contrastive On-Policy Distillation Visual-Advantage On-Policy Distillation for Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:39.971688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:39.971688Z digest=sha256:518a00fc7f251fc11b8db11279216b089fd759273fc3854ac8b87ca4e1ea59d1

Observation e5a7d63c-ee68-4eef-bff6-80eb742f0be1 · outbound

This paper cites Dler: Doing length penalty right-incentivizing more intelligence per token via reinforcement learning.arXiv preprint arXiv:2510.15110,.

Contrastive On-Policy Distillation Dler: Doing length penalty right-incentivizing more intelligence per token via reinforcement learning.arXiv preprint arXiv:2510.15110,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.065268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.065268Z digest=sha256:8ea3ba1da650dc4b5d68a4148c77b4b0c72e23a3387d8ab5d0d7466e4bd12c1b

Observation aad17fb9-1e44-40e3-a6a7-ba6c09acaf89 · outbound

This paper cites Decoupled Weight Decay Regularization.

Contrastive On-Policy Distillation Decoupled Weight Decay Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.174032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.174032Z digest=sha256:45e8848169bdad5cedc31ccb1e124558f9f601a8d22b93992c017b879ef9296f

Observation 1f9feb8b-a01f-4f7a-96e8-4a0cec92c587 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Contrastive On-Policy Distillation Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.283470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.283470Z digest=sha256:529c9bb0544642f9a22dd5db86ee9a1286bbbf337e84635ca3512599eed88ab7

Observation 847d59a2-057f-4214-a84a-3f63f6449b3e · outbound

This paper cites Self- training elicits concise reasoning in large language models.

Contrastive On-Policy Distillation Self- training elicits concise reasoning in large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.393311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.393311Z digest=sha256:923bb41f6a7f2a2047147b47d77f5800823bfc68555c3cc4d310e3a162705e9b

Observation 9d6985be-4572-40fe-b205-e5ad8f2cbde7 · outbound

This paper cites Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost.

Contrastive On-Policy Distillation Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.501299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.501299Z digest=sha256:6fdcc1afcc4a6c7ba8e21824b65b210d3c3f770b75c93ebf5f5071b71a5eb7af

Observation 522d74e7-f8a8-414f-83c6-b1d0751d96be · outbound

This paper cites KL for a KL: On-Policy Distillation with Control Variate Baseline.

Contrastive On-Policy Distillation KL for a KL: On-Policy Distillation with Control Variate Baseline

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.576451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.576451Z digest=sha256:a1abe215b80b7b4824c307bdfb30069cab99fc87f6a74a2b6f3efb3965816b77

Observation 5b3b8c12-6253-4733-8524-2de3eadc56a3 · outbound

This paper cites Privileged Information Distillation for Language Models.

Contrastive On-Policy Distillation Privileged Information Distillation for Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.685025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.685025Z digest=sha256:c0dc410e380e8c031715dab277b91e8d506b7d90af205970e8782a8d508a0fdd

Observation df6d35f7-d248-414b-81a1-cefa399d0993 · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

Contrastive On-Policy Distillation CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.794672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.794672Z digest=sha256:b72a204e2e74ebccaad04c67fd88827f4cc758b49aa0c2da222e793c0fd32abc

Observation 5c89e43b-43ea-4b8f-846a-00eb38933f4d · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Contrastive On-Policy Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:40.904501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:40.904501Z digest=sha256:0b301566ef175fc3895497fd3a8c5e48373e17ff3acb86b243aea914e46f0b7e

Observation 7b4a38b5-d77d-4204-930e-d3b1fc7f4217 · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

Contrastive On-Policy Distillation A Survey of On-Policy Distillation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.241326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.241326Z digest=sha256:b687b2cd001314f23f602f5e314776c00fff8d37017beb711e60a277a2dacad2

Observation d0c01b0a-6ec4-4b85-9cab-755df1189ce5 · outbound

This paper cites Patient knowledge distillation for bert model compression.

Contrastive On-Policy Distillation Patient knowledge distillation for bert model compression

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.320027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.320027Z digest=sha256:a761ee9c36c852e036190f6794c8775eea08d86537e0161196a567aa25164290

Observation 802b0075-64d9-4aa9-b4cb-c6f5934b3d21 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Contrastive On-Policy Distillation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.385416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.385416Z digest=sha256:e4298663d5e1667c7bb974c1ad5cb3755de2b66c224b33e012ed14d7e1ca1cde

Observation 23c56f0b-43e9-44c7-b206-2c7f92f559b7 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Contrastive On-Policy Distillation LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.570637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.570637Z digest=sha256:26a137d359ed3a73a12caaf22e6e8d8930514692dd10fc766481f9ca2c2a01b1

Observation 3b02a138-aef3-482a-af09-b333fc7aaa85 · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Contrastive On-Policy Distillation Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.663241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.663241Z digest=sha256:04c2fb3d48338cf285b3f2475eb6672041e302427b332455b6924b48f9683e4c

Observation ed59648d-76be-4a61-9bd2-e75318123df3 · outbound

This paper cites Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization.

Contrastive On-Policy Distillation Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.757803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.757803Z digest=sha256:7f154fe09b1c8867d296950a4ff0b0c67322179f1f65addfc40cb3aebeeb0fe8

Observation 1322fb1a-f495-4520-8c46-7e9ccac9fd1c · outbound

This paper cites MMGist: A Comprehensive Multimodal Benchmark for 2027.

Contrastive On-Policy Distillation MMGist: A Comprehensive Multimodal Benchmark for 2027

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.853449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.853449Z digest=sha256:9174f3f4b032a7b049b5385c5ccde04061fb041d12133ce613f8203e3f82439f

Observation 64a3584e-895e-485a-9bd1-1fc5c6ebf4d6 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Contrastive On-Policy Distillation Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.931136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.931136Z digest=sha256:0f003ed7168e7df0bf4298ce9266d35d354e7133a260f3f452cb176030cf9a03

Observation 38755796-b319-4f5e-9c80-a0dbf5a873ee · outbound

This paper cites On-Policy Distillation with Best-of-N Teacher Rollout Selection.

Contrastive On-Policy Distillation On-Policy Distillation with Best-of-N Teacher Rollout Selection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:42.016315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:42.016315Z digest=sha256:b49f3bb2036784376e7d1aa50677f1afeefa4d4b0e9f0093f7c29d047dbc72b7

Observation c05e428a-a217-4cab-ab87-f693ae0da0f2 · outbound

This paper cites Are Full Rollouts Necessary for On-Policy Distillation?.

Contrastive On-Policy Distillation Are Full Rollouts Necessary for On-Policy Distillation?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:42.104085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:42.104085Z digest=sha256:1a0beeae0a8f771a423f0a1f03bd74cdb8763dcb45a9ba6691c9eb56c30f3f47

Observation e02bf4e0-af28-4418-8c1a-e0867d0bd2b7 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Contrastive On-Policy Distillation PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:42.188116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:42.188116Z digest=sha256:eeb41b52320390e85051b37d9c27e8ee9f577d00badf2ae98c774ca40b65144d

Observation 8e81c10a-9b37-4880-aeb6-baf2e04a98ba · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Contrastive On-Policy Distillation Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:42.303146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:42.303146Z digest=sha256:a8007b4b78fe17a87d7fa9d5d68230d08b050ba216115f5a5a72a5d8d723a2b7

Observation 792371da-71a4-4db6-be0c-81cb3a40188d · outbound

This paper cites Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe.

Contrastive On-Policy Distillation Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:39.133263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:39.133263Z digest=sha256:cdd7dab7101b9cc604acd06570b8f530fcf2784a127bbccb672a16cfe0d30217

Observation 1831506d-096c-4787-b5a1-039f46404933 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Contrastive On-Policy Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.134065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.134065Z digest=sha256:b47336ca9366605f5fe422b955ed74e55fdf39d55d00dda3e32eca324f802eb8

Observation 5819c389-0162-4f52-8f33-418f66d09db2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Contrastive On-Policy Distillation Proximal Policy Optimization Algorithms

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.009758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.009758Z digest=sha256:8529cce0cdf23670c9918fe3b60c5a903471706302a9ff4d8e38bdb9e37a27e3

Observation 7b2d1f50-d97f-44f2-b96b-b1c5f967de59 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Contrastive On-Policy Distillation Entropy-Aware On-Policy Distillation of Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:39.620771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:39.620771Z digest=sha256:0f26b31f1dc1a24ae31206058bd2a7d2d156d84b44a4280ea0a77b019698fac5

Observation d1b273f8-0ebb-45bd-b07e-7726f5597dcb · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Contrastive On-Policy Distillation Minillm: Knowledge distillation of large language models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:38.714350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:38.714350Z digest=sha256:826a81ade254c48c825c6ee53d604156a20abcd379222f2abace8495b4e6d879

Observation e90466b8-eb43-471a-95c0-01e3bcb98525 · outbound

This paper cites When More is Less: Understanding Chain-of-Thought Length in LLMs.

Contrastive On-Policy Distillation When More is Less: Understanding Chain-of-Thought Length in LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:41.475804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:41.475804Z digest=sha256:0853e0c9d6c08023cd1459ba9045c3c6e387c2636ad3fe7b057f5f492018b7b0

Observation 801402f3-4ca1-4cdd-8c5f-c4991e062efd · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Contrastive On-Policy Distillation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:38.423347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:38.423347Z digest=sha256:6c68b2915280c302e75639d14a4858b61004140d8e503c8804a1d1f455b0d6e1

Observation 8c70757a-9ef0-403b-b493-63747d8b6beb · outbound

This paper cites ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models.

Contrastive On-Policy Distillation ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:38.524310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:38.524310Z digest=sha256:005d3b4591354f16fd733f8ca1aac71e20bcaee295a1f481875717ce8421f486

Observation 5354084a-f257-4407-a231-cfd33b4776ed · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Contrastive On-Policy Distillation Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:38.612685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:38.612685Z digest=sha256:28fda844736eab1355d3717950025253412ddadc29e9e2f631bf3dd48ff758f0

Observation e6990f83-5cff-4127-8bcf-65fc1bc0012a · outbound

This paper cites Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes.

Contrastive On-Policy Distillation Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:39.197447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:39.197447Z digest=sha256:f4cb3286a19e9c6c133ba9dff885aa348c2917906ec22153d9a20d729e81986d

Pith citing papers

Observation b0a6f67f-5f37-47ee-b93b-757d35002dfa · inbound

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation cites this paper.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation Contrastive On-Policy Distillation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-31T02:56:58.020990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T02:56:58.020990Z digest=sha256:397a608ca3e9b4bb445de72b09df08fffba7553a72c6fc4f6c49f6b9a3fb80cf