Pith. sign in

Paper Citation Record · LEDGER

Rethinking On-Policy Self-Distillation for Thinking Models

As of 22 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2607.05184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05184 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T01:04:59.662046Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:59:51.920749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T22:02:35.499158Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact7
  • verified fuzzy5
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbba8739-0646-4fca-bb69-bfbedf14edfb · outbound

This paper cites Forking Paths in Neural Text Generation.

Rethinking On-Policy Self-Distillation for Thinking Models Forking Paths in Neural Text Generation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.525524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:7a14a98b6506b0fb932bc142d340eb758c0feffa5a5a92cb50202129fa280a52

Observation 396f6a89-249f-40bd-88a8-769fa5891127 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Rethinking On-Policy Self-Distillation for Thinking Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.516491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:4b61b45d6ec6d2f2f3813b30c3b6dbe42809d752056ef7efd02cb6a5c2705f66

Observation eb8f50b2-3593-4b2d-b1de-da1768b4882b · outbound

This paper cites Chakravarthy, Anikait Singh, Nathan Lile, and Noah D.

Rethinking On-Policy Self-Distillation for Thinking Models Chakravarthy, Anikait Singh, Nathan Lile, and Noah D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.721941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:261a8724fa6ad36514a868bb6b120987594f8a4abd9f9c0e76afb7cb89803fc9

Observation 9146e4db-699d-489e-b982-0736d154a3fb · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

Rethinking On-Policy Self-Distillation for Thinking Models Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.515390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:d7aedd094ddaf370b5391e52c4ed33b4aa0c3b22f2f8c734c693450daad43f5e

Observation f542e8fc-6641-4546-b5c8-3f025de3e865 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Rethinking On-Policy Self-Distillation for Thinking Models Reinforcement Learning via Self-Distillation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.535875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:84fb84d4b82b888b5a37636312f83cdf61b1de84e63c15550f6bcc42c8fb9433

Observation c3e91af8-c3b6-41ab-9971-e775278d404b · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

Rethinking On-Policy Self-Distillation for Thinking Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.528143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:43b683e5ad1990235745f04ab219f4ed1e4bab5745e69f04806e7fff0e209ba4

Observation bdadb457-eb31-4c85-8b4e-19aedac412ad · outbound

This paper cites Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability.

Rethinking On-Policy Self-Distillation for Thinking Models Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.539267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:9f16180a3819d75f35929236edb8d8566793674217118e4c0097a9793fae56a0

Observation e802dbe8-8581-4ab8-9b8b-9a0d04549f24 · outbound

This paper cites Pope: Learning to reason on hard problems via privileged on-policy exploration.

Rethinking On-Policy Self-Distillation for Thinking Models Pope: Learning to reason on hard problems via privileged on-policy exploration

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:14:27.537447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:3dbb8d00b1d5560ff7b6c2479023c69ae226abd84367f8b44af179d89be5bd9d

Observation 36b68aad-dc12-4f62-80a3-5482845b9328 · outbound

This paper cites Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes.

Rethinking On-Policy Self-Distillation for Thinking Models Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:14:27.529342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:e1e5b94c21b6fd8a789d9cc9e9ead3fa11a7496bbba8b7452b3f946a3938f3cb

Observation be81aa2d-827e-4823-aac3-1f8519c3e5f3 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Rethinking On-Policy Self-Distillation for Thinking Models Self-Distillation Enables Continual Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.531170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:2b059b2f4f2084d07231f04fce1317fa35379c4c19d04fbc2dddc6ec1adcd272

Observation 1a4d5b85-22a0-4110-8d15-efd0c35c8128 · outbound

This paper cites Olmo 3.

Rethinking On-Policy Self-Distillation for Thinking Models Olmo 3

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.545307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:903a59707085fbe9a7833afccced9a4e9395d9d25061c6f7f36caaec228aec50

Observation afbf9a04-0dcf-4fd0-a8b4-6354c8bd52cf · outbound

This paper cites Understanding reasoning in thinking language models via steering vectors.

Rethinking On-Policy Self-Distillation for Thinking Models Understanding reasoning in thinking language models via steering vectors

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.732244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:e35e6f10b6b6e8a1929f4cc0ab2c79e5b5b2d4624907a1f6e20e15afb17a165f

Observation 9e148a67-9f7d-4410-9d2a-018a29fd6c7e · outbound

This paper cites Qwen3 Technical Report.

Rethinking On-Policy Self-Distillation for Thinking Models Qwen3 Technical Report

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.542787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:30440ab457ce5dc8a6f407ce1f390ed28ab9999190240600b0bcecf074ef1e8b

Observation c8272c1a-c400-4d83-9492-cf6d3c41167c · outbound

This paper cites Embarrassingly Simple Self-Distillation Improves Code Generation.

Rethinking On-Policy Self-Distillation for Thinking Models Embarrassingly Simple Self-Distillation Improves Code Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.525726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:7c66e97cfafd74bccb0c4a3bf9e547d457eae07f79c835decdcaa67de8dd7cdf

Observation a309d32f-8968-435e-b44c-7157562fe251 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Rethinking On-Policy Self-Distillation for Thinking Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.501845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:9a7e4dad8a792811b3767adb29f03337a74a85f4329ee0273fafcb4de8bbf771

Observation 37258f2f-723a-488e-ad56-29a8a099477e · outbound

This paper cites an unresolved cited work.

Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:14:27.730229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:d0d13a9570203165d26c8891e6479ee3d75bd6a3e1eab3925b92751683dd366e

Observation 52c3f718-42b9-4900-b8bb-3706cb76ae18 · outbound

This paper cites an unresolved cited work.

Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:14:27.720002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:6fbd2f873fa1557d8782b254337dfecf189f335ac98856dbf30541a46c08de3e

Observation dd3ad507-5f89-409f-b287-639a407cfec3 · outbound

This paper cites The Average column averages the three benchmarks.

Rethinking On-Policy Self-Distillation for Thinking Models The Average column averages the three benchmarks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.726381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:505eeec8952626fd3a74f97a40b1071d573d969328122e94c1762bc80fb83b86

Observation d7c7a81d-4e28-47fb-90ac-2186c707c26b · outbound

This paper cites Epistemic-token OPD applies the same loss only to tokens in the epistemic-marker set.

Rethinking On-Policy Self-Distillation for Thinking Models Epistemic-token OPD applies the same loss only to tokens in the epistemic-marker set

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.723944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:60dfc2d3d28fc7c0218affd4fe7da20058d601fed877bb4350bf0bb6f2fd5306

Observation cbf32b64-d01a-4caf-b5e1-191965f8354f · outbound

This paper cites an unresolved cited work.

Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:14:27.734544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:cf414d5dcf2a31a127ac028423507fd680f710e5ecce98ce7cce56a0e5c214ba

Observation dcfd014f-39a6-41f3-bcd6-0904ee2b91f3 · outbound

This paper cites Blue curves are base thinking models, orange curves are OPSD with full gold-demonstration context, and green curves are OPSD with final-answer-only privileged context.

Rethinking On-Policy Self-Distillation for Thinking Models Blue curves are base thinking models, orange curves are OPSD with full gold-demonstration context, and green curves are OPSD with final-answer-only privileged context

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.728454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:03848259b59173f24bdf0225c15a72d97248435dda68147feeb1dc9d85908d81

Pith citing papers

Observation 713f8264-9954-41ac-9f95-00c226482406 · inbound

DAPD: Dual-Anchored Policy Distillation cites this paper.

DAPD: Dual-Anchored Policy Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 76

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T22:02:35.540632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:02:33.791583Z digest=sha256:ac04db0b0ea0c0e41dc10a6d7b009d90b3224544ad98aaee823cea8fc30de365

Observation de8a8bae-340f-406b-b2b5-d2ebd0a6a75e · inbound

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation cites this paper.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.711475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.711475Z digest=sha256:486fcc0af75080f018b4fb68ff30bdc8fe89d87750ed98c2a9be7a24fe723817

Observation 75ab8712-8de3-4b5a-b064-92fb51e46615 · inbound

Simple-OPD: Demystifying Warm-up for On-policy Distillation cites this paper.

Simple-OPD: Demystifying Warm-up for On-policy Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:49.690977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:49.690977Z digest=sha256:a9341c009e35cbb35853e1801192e4449392772870fb5bd57d533d737c26db42

Observation 770dfc69-243a-40f7-97cf-8d7b3ebea9e2 · inbound

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ cites this paper.

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ Rethinking On-Policy Self-Distillation for Thinking Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T10:59:51.920749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:59:51.920749Z digest=sha256:eaf5171a5bc57a1ff8023ef9738b1bd13f247b72eac51ab5b0760d5efedd20f8