Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T17:27:37.161657Z
Paper Citation Record · LEDGER
As of 29 July 2026, this Paper Citation Record lists 31 of 31 outbound references and 16 inbound Pith citation observations for arXiv:2604.08527.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T17:27:37.161657Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-28T06:31:03.373048+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T11:44:54.717393Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T16:09:56.995110Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b8abbb58-5eb4-4361-9904-10ebfe911eac · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Process Reinforcement through Implicit Rewards
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation c04e973e-a9d1-491a-ac75-c8ed7a09058b · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models MiniLLM: On-Policy Distillation of Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 2276962d-227c-4e08-842e-ec06895aab23 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 81e6947b-18df-4a56-a94d-6cd63946a289 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 40ac80b8-9e99-4fe2-986d-fe1ded2cef6d · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 59c794f9-6257-4678-b11f-b01ad56c2cf5 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Distilling the Knowledge in a Neural Network
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation bd51b5a3-88d4-451f-849e-2be9eec2343c · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation e9156fda-705c-4af2-8f32-7511be2ed81c · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Reinforcement Learning via Self-Distillation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation c1fcf864-8175-41a4-87a5-4721055c6e48 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation fc0f537e-63d8-4d7b-a27e-e65b487925ae · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models and Rush, A
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 881ca87b-091f-45a8-ab42-cd3a3ff62932 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Dual Policy Distillation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 8181b803-58df-495b-bc2e-104fe803fa31 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Solving Quantitative Reasoning Problems with Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 881f9d32-a398-4993-b33f-c78f028855dd · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Let's Verify Step by Step
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 2970a45c-60c2-4bf4-b696-e486ed827aa7 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Understanding R1-Zero-Like Training: A Critical Perspective
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 1c730f72-9819-4b20-b8a5-faa1acdf2f9e · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models 20251026
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 9cc87e02-f16b-4500-b8e8-8fe9f02b7009 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Policy Distillation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation eb0047b6-7d34-41bb-8d84-d375a8a67b2c · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation ca188012-58ae-4272-8959-49e766c006ec · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Proximal Policy Optimization Algorithms
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation b5f0a30a-7d66-4bde-9199-c3c70396ac03 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 6dc57683-d61d-453f-a22a-1cc743f085d3 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models MiMo-V2-Flash Technical Report
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 2be2d9c2-6261-4c1e-bdfe-9d85220e63c8 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Learning to Reason under Off-Policy Guidance
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 76a7a0f5-5d7e-4777-b998-045d16356060 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation c134a4ce-be53-4262-8b9f-e6cdc509d312 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Qwen3 Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 702bf885-7308-45e7-b79b-bde35918684c · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Black-box on-policy distillation of large language models.arXiv preprint
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation deb75af3-a730-4bbb-910a-c7712e56e9db · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 930f1c66-095d-4c67-a699-e775291e48e7 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 94c4fcd5-d4e2-4e3e-b450-c0537958cc2c · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models A Survey of Reinforcement Learning for Large Reasoning Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation a6f8d663-b49c-4d2f-8e24-bc00a140300f · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 936e03b1-0802-4aba-8450-526f25379efd · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Appendix B
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 36e8de24-1349-4f57-8d75-c720b8018423 · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation f8b954ce-30bb-46e4-bf29-10df51c99bcf · outbound
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation bffb3a39-fc4f-4bc8-aa3d-0a6d1dc7315e · inbound
UniSD: Towards a Unified Self-Distillation Framework for Large Language Models Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 5789fdb2-a335-4cac-a115-a4d4bcef9fb3 · inbound
UniSD: Towards a Unified Self-Distillation Framework for Large Language Models Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 9d78977c-c3e9-4d74-a39f-68247d4e9259 · inbound
The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation f1800a1b-5e5e-485b-a988-6c6fbb5b4e6c · inbound
The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 8e3ab833-e9bd-4aca-809f-72e9b236ea08 · inbound
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 93ef7496-5ceb-4dbe-b906-15219d6078a4 · inbound
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation d74c0cfe-4275-404a-b72d-b72efe45aa59 · inbound
Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 0047499b-cde7-48d1-8bcc-2310a655b21d · inbound
ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 9d842b35-1215-49f0-af28-22c9e37cf14e · inbound
SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 9a7e96b4-03d0-4a8d-bf88-d2c65942ea5b · inbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 4eef3ad0-b655-400e-9bbd-3ac125455005 · inbound
A Formula-Driven Survey and Research Agenda for On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 7e5ca99a-2617-4633-93af-f0db5377a306 · inbound
Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation be204bb8-3062-423f-84b3-3ce60c6f1b02 · inbound
Blockwise Policy-Drift Gating for On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 8501e135-1f97-47cf-b856-bfe3bf3fd30f · inbound
DanceOPD: On-Policy Generative Field Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 0eb77a78-3ce9-4d3a-a886-f0c4d2bc17b8 · inbound
DanceOPD: On-Policy Generative Field Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c825d042-3a3b-469e-a29b-82aeaefddc21 · inbound
Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.