Pith. sign in

Paper Citation Record · LEDGER

Reward-Gated On-Policy Distillation

As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2607.04037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04037 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T22:07:54.738215Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:19.124618Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T14:39:46.237545Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dabe41dd-4661-4ec6-b43a-5b36757cd657 · outbound

This paper cites Agarwal, N.

Reward-Gated On-Policy Distillation Agarwal, N

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:71d97163d579020f64ca9950948f843bb03d3e075c31dcbdb45edb40f49c9755

Observation ccd41238-23c4-4d4a-8b2a-57730a87a139 · outbound

This paper cites Austin, A.

Reward-Gated On-Policy Distillation Austin, A

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:faba1567a0329f3a14c661bdf511f18da40e139231bfa3a7bd77a03c4fc1f282

Observation 9d17f507-6cfe-46c1-b01f-cdcce3cea4da · outbound

This paper cites Cobbe, V.

Reward-Gated On-Policy Distillation Cobbe, V

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:900f02177ec50f66a9602af9eed0e4f9aa2f6c43d4656501b363dd1b76a6c529

Observation fd46cde3-a885-4b60-adef-e443b6873a85 · outbound

This paper cites Murphy: Reflectivemulti-turnreinforcementlearningforself-correctingcodegenerationinlargelanguage.

Reward-Gated On-Policy Distillation Murphy: Reflectivemulti-turnreinforcementlearningforself-correctingcodegenerationinlargelanguage

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:dd8fb14ba364368e7437441be50c39cf63b93777a014e6f52120773d5e2288bf

Observation 2532560c-5c4b-4f22-b791-3d031592b46c · outbound

This paper cites an unresolved cited work.

Reward-Gated On-Policy Distillation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:48ee16444d467bf987aace5312fd46a50d077a52704bd88cff977d817b3bb37d

Observation 5b91f8d8-cb5d-43ec-b5b0-d9d64d07ff2c · outbound

This paper cites an unresolved cited work.

Reward-Gated On-Policy Distillation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:14ff8732def0c4aa2fd8134a73df6d8e9491d42284e951218c9569f63908914a

Observation a8a9edb0-293c-48c4-9957-203196e074ea · outbound

This paper cites Hendrycks, C.

Reward-Gated On-Policy Distillation Hendrycks, C

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:be75640bd3e259f111fa147db43a1df12404ec42b71285ad88a72cef30033139

Observation 7a78c896-030a-4d8e-ae3e-9ce0cc87bdc5 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Reward-Gated On-Policy Distillation Distilling the Knowledge in a Neural Network

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:721c2abe5c9c1c64b295d899dc640e2b3acf5975c43a8305921bdee1fec98315

Observation 8f55a362-2f6b-4879-8bf4-01cd1d48bceb · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Reward-Gated On-Policy Distillation Reinforcement Learning via Self-Distillation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:1913ebd3fc8761dac7475d3aad82b2a6fd5152037f9553765404de9e7b939625

Observation 71b012c7-1bea-4ebd-a1e2-56cef830bd10 · outbound

This paper cites Explaininyourownwords: Improvingreasoningviatoken-selectivedual knowledgedistillation.

Reward-Gated On-Policy Distillation Explaininyourownwords: Improvingreasoningviatoken-selectivedual knowledgedistillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:a4f5a96fba2d77f3235ce3c107ab55fc0bc17fac7ffba641063a993415ecf87e

Observation 51dc7e38-14db-4036-bbcd-78e87c80f633 · outbound

This paper cites an unresolved cited work.

Reward-Gated On-Policy Distillation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:0ef960752b5be11405a7e8deda682d7fab705a94cff365ffa445d151ff8d8f87

Observation 5a73ff77-3912-4e04-9725-5f85ff5853ea · outbound

This paper cites Understandingr1-zero-like training: A critical perspective.

Reward-Gated On-Policy Distillation Understandingr1-zero-like training: A critical perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:88999b011491ec38f9a6edb0acc8fad930b2498dbd32b25c40f3e57bb60dbb1d

Observation 68d10316-3bc6-46ac-8449-1bf722e2f5da · outbound

This paper cites Schulman, F.

Reward-Gated On-Policy Distillation Schulman, F

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:aeac5fdadad0d9ac739ba16ab0e9bf37f105ff1db91041fe38831c56ffc72c8e

Observation f1695f4a-1a48-4b20-a5e5-bc7eeb7ae3cb · outbound

This paper cites an unresolved cited work.

Reward-Gated On-Policy Distillation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:e2f8aba8188e380479ae850d83b0384baae614e457738401df12e9124387a7ab

Observation cc686297-25b3-4863-bd86-477d70a428a1 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Reward-Gated On-Policy Distillation Self-Distillation Enables Continual Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:2c775e9b7de354f192f214c3a01653a64af8f073091b39f9e6ddca5e3b8f61d8

Observation 76c9cb66-bd19-43f3-970c-fd8aa827ba81 · outbound

This paper cites Sprague, X.

Reward-Gated On-Policy Distillation Sprague, X

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:136379d2023bf08244fdce0c6aee1df3437a8a9fa80a6a03572fe2dfbaa5368f

Observation 6b05d31b-7c0f-40b3-b5f5-e765a3a67cbd · outbound

This paper cites Suzgun, N.

Reward-Gated On-Policy Distillation Suzgun, N

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:60480ef762221f4a1063d6d47c456310fcb614be6d9f2873a9345f7b197fd7fe

Observation 269d5db7-efbe-458a-91f4-974b470407ea · outbound

This paper cites an unresolved cited work.

Reward-Gated On-Policy Distillation Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:a6e2fe2a996761e56a086eaf3be0274861ada15b3a0f80b0f5277e0f2bbe51e3

Observation 032cf95a-5b64-4edb-8622-4cb92742d0ac · outbound

This paper cites Welbl, N.

Reward-Gated On-Policy Distillation Welbl, N

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:57031c676f53744af7429c008eb945c806798b2676a2f69381b555fba53d068b

Observation 8d86fe6c-b348-4d9d-b850-0e963a146349 · outbound

This paper cites Kdrl: Post-training reasoning llms via unified knowledge distillation and reinforcement learning, 2025.

Reward-Gated On-Policy Distillation Kdrl: Post-training reasoning llms via unified knowledge distillation and reinforcement learning, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:133df9260791bc62241571e0afe19dfcab9d133f177c32fcdd6903f587a5b86a

Observation 61640c2d-4ec9-41eb-bb24-99695b356187 · outbound

This paper cites an unresolved cited work.

Reward-Gated On-Policy Distillation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:6fa73d913b2010daddc7834ae58a41d4426712b2cff9629eac4953886f776dee

Observation 511175c2-e0a7-4a45-807f-9aa3af34dad0 · outbound

This paper cites an unresolved cited work.

Reward-Gated On-Policy Distillation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:2b6d3ad298416278588e439546172105dc7f4a531ad944be125becf8fcd22c74

Observation b84562b4-57b0-4155-aa3f-68f84ee50733 · outbound

This paper cites Zhang, S.

Reward-Gated On-Policy Distillation Zhang, S

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:c435120dab442b6e05300f928246701a6e8ce65c34fa23bef26e00bc3799ee54

Observation 3bd2d94a-1d84-4e5d-80a6-4cd8b43010d1 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Reward-Gated On-Policy Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:3bd4e08c0ce570fb8e03640ca56a311c8a80533c5e4f68656ce9a203a65112cb

Observation 060f3a0b-77af-4efa-823e-54fcbe0954a7 · outbound

This paper cites Zheng, S.

Reward-Gated On-Policy Distillation Zheng, S

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:9d0f734eb966cb47bcf3a970d53a8c8e75f0525b293360f46a43711138138115

Observation 0e5fd19e-b112-4cea-9eb0-fe0e2621b12d · outbound

This paper cites an unresolved cited work.

Reward-Gated On-Policy Distillation Unresolved cited work

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-07-11T22:07:54.738215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:07:54.738215Z digest=sha256:936710c1ae1800a45fbb255e82fa386da2fccfc5f3bbc0fda0d13d988eb01e45

Pith citing papers

Observation e9198a1b-9d49-494d-b97d-d09a9df0c6d0 · inbound

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation cites this paper.

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation Reward-Gated On-Policy Distillation

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T14:39:46.245029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:39:46.145211Z digest=sha256:6484590503231ada903b87c100e256068aa8e13ba5f2bb4fbdd4f2982416b311

Observation 343ca932-becd-49a3-bfbd-a32217317fcd · inbound

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation cites this paper.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Reward-Gated On-Policy Distillation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.124618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.124618Z digest=sha256:eeb67ce9afcc7461a50c9f09b2eebc06f8335356b792fb872c1b514995f9960d