Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:19.167249Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.11698.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:19.167249Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8d2f49b6-af66-4f9a-bea5-9792c9e053c9 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation On-policy distillation of language mod- els: Learning from self-generated mistakes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 79e2b58e-3191-428c-a018-304fcebf1f0b · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed9ec3eb-7a5c-4d04-bc22-7cb208b97a1b · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f5e6aa7-4a84-4743-a369-f2480d7b71bb · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Distilling the Knowledge in a Neural Network
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b316f74c-d005-4fb4-8f7c-22dadafd967b · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4949b4f3-3c69-414b-98c1-b74f08159cfe · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation MiniLLM: Knowledge distillation of large language mod- els
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a8c7552d-97dc-4320-baa3-2eacadd96557 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Model extrapolation expedites alignment
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46898c64-f0ea-44ef-b5db-7c92ab01bc04 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation LLM-oriented token-adaptive knowledge distillation.Pro- ceedings of the AAAI Conference on Artificial Intelli- gence, 40(40):34070–34078, 2026
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23c527f-e31f-4cda-a179-1ae2fdd62788 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation ASKD: Reinforcement learning-style knowledge distillation with quality-adaptive skewness.Proceedings of the AAAI Conference on Ar- tificial Intelligence, 40(41):34781–34789, 2026
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 54419194-7d93-41c3-a444-1a01734cd859 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5b56c85-1b64-442b-bb03-1b6f5daffe81 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab18514-f7ae-4c58-b3cf-b0ca4203e93b · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 343ca932-becd-49a3-bfbd-a32217317fcd · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Reward-Gated On-Policy Distillation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60ab21a5-297c-438c-b8c9-98d94c2130db · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc351a7-de7f-42b0-8719-30b5195aac97 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Qwen3 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d455d430-d7af-4006-9d85-f977df22ef75 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd8c6ada-e011-45c3-ac60-530c11e2b87b · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Advancing LLM Reasoning Generalists with Preference Trees
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df29590-f01d-4cb7-982a-e9b5594f2c56 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Decoupled weight de- cay regularization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7b3b7e15-b9d8-4d91-a529-f2e623f53190 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Evaluating Large Language Models Trained on Code
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ec700ea-8cb8-495b-8eb0-b14d17c86318 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Program Synthesis with Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e08ed39-0c87-412e-9dba-2d8a1f67fe99 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a93e8f58-8d3a-493b-a754-e4b59ae2cf4b · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce6c5471-a501-4538-a6d5-e41177e99e63 · outbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation TIP: Token Importance in On-Policy Distillation
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.