Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T07:02:36.752153Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2607.28449.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T07:02:36.752153Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ddd7fc9e-a4d0-4546-a811-6f358c73364b · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models AIME 2024.https://huggingface.co/datasets/AI-MO/aimo-validation-aime,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd79266-7d44-432b-ad48-0a6dd58fc69f · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc5aab60-667a-4dd0-b780-aea286a13b4c · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models OpenThoughts: Data Recipes for Reasoning Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc1a036-9424-4bbf-b503-033469dd48cb · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5909f497-0756-4b3b-ad21-d3661e934ae0 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d5da68-1923-4559-a563-0905637bbfdc · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 975bbc5e-7d34-4bef-8d83-7775c76450b1 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Reinforcement Learning via Self-Distillation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2084da83-caa3-4c49-941a-7d1e56ac3087 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38dd32c-7209-41ae-af75-6a61e6b03265 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Entropy-Aware On-Policy Distillation of Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee7c5b74-6577-4cb7-87cc-c9ea7e5aa1ab · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42a823f1-4c4c-42f4-a412-dc775f6283bb · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b670b4-ea02-4b7d-afb4-9d72cc8f7a99 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Sequence-level knowledge distillation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7862e934-f55d-4afc-820c-8cc32614c136 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 807dbbc4-4eec-4054-bdb4-cf5469ab7aff · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ae4232-676a-4c37-96ab-9431c39276a7 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb182fe-34e0-412a-a434-eec9270291e4 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models DeepSeek-V3 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff7c33b-6005-4a1b-b8ee-c1adfba11976 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models UFT: Unifying supervised and reinforcement fine-tuning.arXiv preprint arXiv:2505.16984, 2025a
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b203fea0-f5bf-4942-a61e-1442a19c46d1 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models https://thinkingmachines.ai/blog/on-policy-distillation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a087ab-2523-4016-a877-a14e7cf84415 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae4651e4-33f5-43b2-9cd1-451b62bcc1d0 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models NVIDIA Nemotron 3: Efficient and Open Intelligence
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f19ca90c-1a11-44d1-8095-362418d35de6 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models AIME 2025.https://github.com/open-compass/opencompass,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89a053cc-c1f2-4931-8fd0-2c31e75c0ac3 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Revealing the power of post-training for small language models via knowledge distillation.arXiv preprint arXiv:2509.26497,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfac691-02da-416c-82de-29742d00f830 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Proximal Policy Optimization Algorithms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ee137f-c17f-450c-9951-754fbe620e4d · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models RL's Razor: Why Online Reinforcement Learning Forgets Less
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cccce64-d684-4371-acfd-817a38490354 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Self-Distillation Enables Continual Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deef98bc-7504-47bb-a9d0-3f9df3ab35df · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models OpenAI GPT-5 System Card
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d7d686-12d0-4d27-a4c7-8346910dee3b · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models A Survey of On-Policy Distillation for Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc31d68-6bdc-4893-bc14-e94e6b49a585 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Self-Supervised On-Policy Distillation for Reasoning Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e67e821-e832-4136-961d-3518d6577db3 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b99a255f-b801-4986-9efa-96857d746678 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Kimi K2.5: Visual Agentic Intelligence
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7ae3d3-3481-46ce-831b-d24f80758b7a · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c910826-67b5-4b69-8fe3-5f187761628a · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a39875db-6500-42a1-a564-561822931b6d · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models MiMo-V2-Flash Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ab9de9-2145-472b-89ed-1c45fef0917f · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models On the Position Bias of On-Policy Distillation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77ac798f-70e1-4acb-a73f-48793c51da64 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa89616-2045-4e00-a665-2d49fae5c1eb · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Learning to Reason under Off-Policy Guidance
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 588ac625-7646-4f1a-b16a-6753604ab89f · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c6b745-2daf-498b-9e44-a798cab22edd · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Qwen3 Technical Report
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 185ee781-bb66-4b78-bd75-72e87e492e14 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dbe2b1a-9005-4219-86c6-c9868c18dbe1 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eef9deb5-9289-4a84-8567-56f20e255c95 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01e5a407-f9da-4655-ae0e-699014a21359 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc9009ee-bca9-460e-8e6a-5dcf8a00d7c2 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models On-policy RL meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting.arXiv preprint arXiv:2508.11408,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb05be9-1614-43c3-8c22-408743b5e0a0 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Reinforcement-aware Knowledge Distillation for LLM Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c92d4674-260a-4c38-99d5-528589902195 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911d5221-e152-47a5-a0ad-af83f2ee42b4 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Group Sequence Policy Optimization
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05df5b25-d605-4c24-803e-01617e1caa9e · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee3de70-d338-4fa4-8329-73a1ae80d57c · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc3823e-a6b5-41ca-ae76-9a21050b90c0 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Process Reinforcement through Implicit Rewards
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b30ab58f-eed9-4e45-8784-87573d40aa35 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Privileged Information Distillation for Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41570081-83f2-4aaf-b8ca-184ecb71579d · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c070f38-ab33-438f-a32f-01cbf65e2b24 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4395ab5a-1095-460b-ab01-5d9f6faca5e6 · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9884baaf-7536-45d6-9bb8-5a78833016be · outbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.