Pith. sign in

Paper Citation Record · LEDGER

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation

As of 16 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.11698.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11698 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:19.167249Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved17
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d2f49b6-af66-4f9a-bea5-9792c9e053c9 · outbound

This paper cites On-policy distillation of language mod- els: Learning from self-generated mistakes.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation On-policy distillation of language mod- els: Learning from self-generated mistakes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:19.520409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:35:19.054951Z digest=sha256:26c246eacad727f10213603b32d1eb6e89b603c4f5bf798dc236cf8479a823d1

Observation 79e2b58e-3191-428c-a018-304fcebf1f0b · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.060109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.060109Z digest=sha256:e15d935a01942eedf79382aefc79ea7a46d0dca2efc3aed4ffe8554b2fcc898d

Observation ed9ec3eb-7a5c-4d04-bc22-7cb208b97a1b · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.075167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.075167Z digest=sha256:bf19d8b688797c3575b7a6b72456625332175186be0abc970fc9234f426e0699

Observation 7f5e6aa7-4a84-4743-a369-f2480d7b71bb · outbound

This paper cites Distilling the Knowledge in a Neural Network.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Distilling the Knowledge in a Neural Network

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.080208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.080208Z digest=sha256:fb599e82589ac38d4b97775c5ec8783ffe0d85b4b8aaab74937b43fcc9851425

Observation b316f74c-d005-4fb4-8f7c-22dadafd967b · outbound

This paper cites an unresolved cited work.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.085157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.085157Z digest=sha256:8baa5b749829aed49255350da5ed6288d837406917a809c849f5053e033de7e4

Observation 4949b4f3-3c69-414b-98c1-b74f08159cfe · outbound

This paper cites MiniLLM: Knowledge distillation of large language mod- els.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation MiniLLM: Knowledge distillation of large language mod- els

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:19.504125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:35:19.089748Z digest=sha256:b7f97d200f2d5e2395b3a6e2d61ed28bb3ce3dc9354d4e0b3562ecd6c7b1a941

Observation a8c7552d-97dc-4320-baa3-2eacadd96557 · outbound

This paper cites Model extrapolation expedites alignment.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Model extrapolation expedites alignment

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-16T00:35:19.094622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.094622Z digest=sha256:a3ff0a91d85ad9be0b5123a9f75b3fa2ba9f19d62732e9525bfed24b27978bc7

Observation 46898c64-f0ea-44ef-b5db-7c92ab01bc04 · outbound

This paper cites LLM-oriented token-adaptive knowledge distillation.Pro- ceedings of the AAAI Conference on Artificial Intelli- gence, 40(40):34070–34078, 2026.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation LLM-oriented token-adaptive knowledge distillation.Pro- ceedings of the AAAI Conference on Artificial Intelli- gence, 40(40):34070–34078, 2026

Reference 9

Resolution
malformed identifier
no resolver link, observed 2026-08-16T00:35:19.101059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.101059Z digest=sha256:4863f168c23d12f2d40aa0148807cebe3eaf33cfb2bdd11852e6630d17ba08f5

Observation c23c527f-e31f-4cda-a179-1ae2fdd62788 · outbound

This paper cites ASKD: Reinforcement learning-style knowledge distillation with quality-adaptive skewness.Proceedings of the AAAI Conference on Ar- tificial Intelligence, 40(41):34781–34789, 2026.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation ASKD: Reinforcement learning-style knowledge distillation with quality-adaptive skewness.Proceedings of the AAAI Conference on Ar- tificial Intelligence, 40(41):34781–34789, 2026

Reference 10

Resolution
verified exact
doi, observed 2026-08-16T00:35:19.205464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:35:19.105626Z digest=sha256:a93c1569a84dd43aacbbdd4d7554d6b3abd5e0f0c05da8c9952b59ec910971b2

Observation 54419194-7d93-41c3-a444-1a01734cd859 · outbound

This paper cites Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.110461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.110461Z digest=sha256:6d9991362c30662d523ee68818af9046fd969d9136b8de6a3bf814d38da08b20

Observation d5b56c85-1b64-442b-bb03-1b6f5daffe81 · outbound

This paper cites SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.115157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.115157Z digest=sha256:3c626bd3b4af9bd3e65be9aa6db965a09b07531b9c374b6ef11311d5be78d76b

Observation 7ab18514-f7ae-4c58-b3cf-b0ca4203e93b · outbound

This paper cites SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.120078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.120078Z digest=sha256:7c0ef74060ab295b38b7639cffe9995bad754a7efe3ed6c8c155872cba021c40

Observation 343ca932-becd-49a3-bfbd-a32217317fcd · outbound

This paper cites Reward-Gated On-Policy Distillation.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Reward-Gated On-Policy Distillation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.124618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.124618Z digest=sha256:fd8abf6467235d34fcc7d8b902610c9943a74423b3b7739e44ecad9b6d76fd86

Observation 60ab21a5-297c-438c-b8c9-98d94c2130db · outbound

This paper cites Proximal Policy Optimization Algorithms.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.129403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.129403Z digest=sha256:49ee767a97babe80fc2556739a8b392e928cd9588f5b3e03d207ddbaaaf5e7f7

Observation 2cc351a7-de7f-42b0-8719-30b5195aac97 · outbound

This paper cites Qwen3 Technical Report.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.134187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.134187Z digest=sha256:47bc1ecbbe573287bf13ab23d1ad4ed04c00a015408193105c019524f7ebbf63

Observation d455d430-d7af-4006-9d85-f977df22ef75 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.138789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.138789Z digest=sha256:17edbce8260ebc9a8b9331a1c84b40e188da068388c690371bb761ba56c5a9ab

Observation bd8c6ada-e011-45c3-ac60-530c11e2b87b · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Advancing LLM Reasoning Generalists with Preference Trees

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.143448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.143448Z digest=sha256:6ce216aa6e1146ba62d8b4594acefb0bff537554055babdba971e3917bc51dd4

Observation 4df29590-f01d-4cb7-982a-e9b5594f2c56 · outbound

This paper cites Decoupled weight de- cay regularization.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Decoupled weight de- cay regularization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:19.488791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:35:19.148303Z digest=sha256:a3e1d229a06aea1defc7b711fae1aa47b81cbdbda4c29f3d1978bbb4c854e0fb

Observation 7b3b7e15-b9d8-4d91-a529-f2e623f53190 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Evaluating Large Language Models Trained on Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.152756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.152756Z digest=sha256:e95c5f69dd2dc96f2703aaec0bc504189a01d36848eb7f2483cf59c1e00ed2b1

Observation 1ec700ea-8cb8-495b-8eb0-b14d17c86318 · outbound

This paper cites Program Synthesis with Large Language Models.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Program Synthesis with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.157784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.157784Z digest=sha256:65cab2f2bd77d26401cc9aaaf2672364499dee9a43840437d5a702316b9ec9d2

Observation 0e08ed39-0c87-412e-9dba-2d8a1f67fe99 · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.162472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.162472Z digest=sha256:9093e356b04f7e5f613f39f9a0346bb0615bd8adccb19000a26060ee11616f99

Observation a93e8f58-8d3a-493b-a754-e4b59ae2cf4b · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.167249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.167249Z digest=sha256:2afa8e91d8310b58f738b3372576c1ece8a9a78ba33df5e6fca08b41c6b1b6a1

Observation ce6c5471-a501-4538-a6d5-e41177e99e63 · outbound

This paper cites TIP: Token Importance in On-Policy Distillation.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation TIP: Token Importance in On-Policy Distillation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.070631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.070631Z digest=sha256:78ca2a4b8f8d24b1bfb1d298a5bebdb07d28e4f95619d4cdd548963d76d80a3e

Pith citing papers

No inbound Pith citation observations are available.