Pith. sign in

Paper Citation Record · LEDGER

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains

As of 7 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 3 inbound Pith citation observations for arXiv:2605.28014.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28014 v1

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T13:04:28.418120Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:02:30.513908Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T07:24:21.975072Z

Reference resolution

9 of 9 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 694fa727-c843-4f24-9b90-031be593f824 · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:13:27.608752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:04:28.418120Z digest=sha256:a83152d102224e25cf15d45ab0475a95ecb754446ab2b02b474b49c5a01ab3cb

Observation 33fc5d03-48c0-49bf-ac87-454fd7af2792 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Understanding R1-Zero-Like Training: A Critical Perspective

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:13:27.605975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:04:28.418120Z digest=sha256:41e72dbd8e23a6772e7f3d43e28f819defe3959d153116c97618c52cc76dd2fc

Observation 757a8a5d-70cb-46ba-a8cb-3b268cb123bb · outbound

This paper cites an unresolved cited work.

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T13:04:28.418120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:04:28.418120Z digest=sha256:318fb51fb2284b2bfabfb1a0327b21b6c376dc1752949466d11d7f7a1fc7ccc0

Observation e583b04d-c8f6-4aa6-b446-f8750df598ba · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:13:27.598454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:04:28.418120Z digest=sha256:e4ec074f3a19e8cb21e64ba698bd04282a6ad7ae81a0d0ee42bcd3b3609cf2ad

Observation f4896633-0afe-40a4-85b8-5d790a753626 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:13:27.600926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:04:28.418120Z digest=sha256:b6f3a1226f10619942af7cc5ffcc883e7013d354667b0528762d31e3be752150

Observation 07462d50-c2ce-4b3a-9c40-c63d1c7c4cd1 · outbound

This paper cites Qwen3 Technical Report.

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Qwen3 Technical Report

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:13:27.603192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:04:28.418120Z digest=sha256:db3f4ba08a52860767c4427f318f1a5446d2e95a4adc51d343364b4bb1b4ffd5

Observation df8451af-bf00-43f1-a93c-0c814ed66c77 · outbound

This paper cites We train for 10 epochs on the science question-answering tasks and 5 epochs on the ToolUse task, using a training batch size of.

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains We train for 10 epochs on the science question-answering tasks and 5 epochs on the ToolUse task, using a training batch size of

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T13:04:28.418120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:04:28.418120Z digest=sha256:0f658c2b2da7735cb6d20a15b6d45c57fb556c40c7fce5ad5fbd4438acf1f2ac

Observation b6d584bb-b163-4f66-980e-b89179123d6f · outbound

This paper cites All algorithms are eval- uated every 10 training steps.

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains All algorithms are eval- uated every 10 training steps

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T13:04:28.418120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:04:28.418120Z digest=sha256:10c36f40adfbddcf5d41e48938dc42c8a56eed8dba2c9c1a117c42ffb9d04d39

Observation a56641b6-f4fb-46fd-b6e2-b7a47ffb13ba · outbound

This paper cites Given a question and four options, please select the right answer.

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Given a question and four options, please select the right answer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T13:04:28.418120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:04:28.418120Z digest=sha256:522b6478b6013e029fe1f2aa9ddc6a4ed2ef7e9ede59e22d746e01e44aa271de

Pith citing papers

Observation c0065342-62a5-46b4-9a64-867016beeb23 · inbound

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation cites this paper.

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:21.976321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T07:15:44.387347Z digest=sha256:e298bd0749a86450cbdf523cc6b176a0b4e24e0c9bbe314814ae4f23a45a01ce

Observation aae79c50-4483-4cf4-ba8b-7467a79da262 · inbound

Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning cites this paper.

Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T22:30:30.736777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:30:30.736777Z digest=sha256:ec1175c81e7333344850d8d9df28cd9ce20f4f0342c99a0b2def4820162d21df

Observation 83bc18b7-8f62-4c40-8af9-570efdef245b · inbound

DAPD: Dual-Anchored Policy Distillation cites this paper.

DAPD: Dual-Anchored Policy Distillation ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T22:02:30.513908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:02:30.513908Z digest=sha256:1911ac7735c4e2bb1e59d892dd248cd62086491ee3aea09bb9ceb5554940d093