Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T13:04:28.418120Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 3 inbound Pith citation observations for arXiv:2605.28014.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T13:04:28.418120Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:02:30.513908Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T07:24:21.975072Z
9 of 9 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 694fa727-c843-4f24-9b90-031be593f824 · outbound
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 33fc5d03-48c0-49bf-ac87-454fd7af2792 · outbound
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Understanding R1-Zero-Like Training: A Critical Perspective
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 757a8a5d-70cb-46ba-a8cb-3b268cb123bb · outbound
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e583b04d-c8f6-4aa6-b446-f8750df598ba · outbound
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f4896633-0afe-40a4-85b8-5d790a753626 · outbound
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 07462d50-c2ce-4b3a-9c40-c63d1c7c4cd1 · outbound
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Qwen3 Technical Report
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df8451af-bf00-43f1-a93c-0c814ed66c77 · outbound
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains We train for 10 epochs on the science question-answering tasks and 5 epochs on the ToolUse task, using a training batch size of
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6d584bb-b163-4f66-980e-b89179123d6f · outbound
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains All algorithms are eval- uated every 10 training steps
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56641b6-f4fb-46fd-b6e2-b7a47ffb13ba · outbound
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains Given a question and four options, please select the right answer
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0065342-62a5-46b4-9a64-867016beeb23 · inbound
UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aae79c50-4483-4cf4-ba8b-7467a79da262 · inbound
Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83bc18b7-8f62-4c40-8af9-570efdef245b · inbound
DAPD: Dual-Anchored Policy Distillation ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.