Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-08T01:04:59.662046Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2607.05184.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-08T01:04:59.662046Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:05.711475Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-04T22:02:35.499158Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cbba8739-0646-4fca-bb69-bfbedf14edfb · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Forking Paths in Neural Text Generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 396f6a89-249f-40bd-88a8-769fa5891127 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eb8f50b2-3593-4b2d-b1de-da1768b4882b · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Chakravarthy, Anikait Singh, Nathan Lile, and Noah D
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9146e4db-699d-489e-b982-0736d154a3fb · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f542e8fc-6641-4546-b5c8-3f025de3e865 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Reinforcement Learning via Self-Distillation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c3e91af8-c3b6-41ab-9971-e775278d404b · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bdadb457-eb31-4c85-8b4e-19aedac412ad · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e802dbe8-8581-4ab8-9b8b-9a0d04549f24 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Pope: Learning to reason on hard problems via privileged on-policy exploration
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36b68aad-dc12-4f62-80a3-5482845b9328 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be81aa2d-827e-4823-aac3-1f8519c3e5f3 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Self-Distillation Enables Continual Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a4d5b85-22a0-4110-8d15-efd0c35c8128 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Olmo 3
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afbf9a04-0dcf-4fd0-a8b4-6354c8bd52cf · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Understanding reasoning in thinking language models via steering vectors
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e148a67-9f7d-4410-9d2a-018a29fd6c7e · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Qwen3 Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c8272c1a-c400-4d83-9492-cf6d3c41167c · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Embarrassingly Simple Self-Distillation Improves Code Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a309d32f-8968-435e-b44c-7157562fe251 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 37258f2f-723a-488e-ad56-29a8a099477e · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52c3f718-42b9-4900-b8bb-3706cb76ae18 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd3ad507-5f89-409f-b287-639a407cfec3 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models The Average column averages the three benchmarks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7c7a81d-4e28-47fb-90ac-2186c707c26b · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Epistemic-token OPD applies the same loss only to tokens in the epistemic-marker set
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cbf32b64-d01a-4caf-b5e1-191965f8354f · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dcfd014f-39a6-41f3-bcd6-0904ee2b91f3 · outbound
Rethinking On-Policy Self-Distillation for Thinking Models Blue curves are base thinking models, orange curves are OPSD with full gold-demonstration context, and green curves are OPSD with final-answer-only privileged context
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 713f8264-9954-41ac-9f95-00c226482406 · inbound
DAPD: Dual-Anchored Policy Distillation Rethinking On-Policy Self-Distillation for Thinking Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de8a8bae-340f-406b-b2b5-d2ebd0a6a75e · inbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Rethinking On-Policy Self-Distillation for Thinking Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.